Written by Admin Alex · Fact-Checked by M.Ali · Info Verified September 2026
We review and update this article regularly as new information becomes available.
TL;DR: Cognition released SWE-2, the new model behind its Devin coding agent, on September 10. It edges out Anthropic’s Fable 5.1 on two widely watched coding benchmarks and costs 64% less to run on one of them. It also gets crushed on a third benchmark, which is the part of this story worth paying attention to.

Benchmark wars in AI coding tools have a predictable shape by now. A company ships a new model, cherry-picks the tests where it wins, and calls it the new frontier. SWE-2 mostly follows that script, except it also published a number that makes it look bad, which is rare enough to be interesting on its own.
On FrontierCode 1.1 Main, SWE-2 scores 50.0%, just a hair behind Anthropic’s Fable 5.1 at 50.9%. On DeepSWE 1.1, SWE-2 actually wins outright, 73.0% to 67.4%. Those are close, competitive numbers between two frontier coding models.
Then there’s Terminal-Bench 4. SWE-2 posts 27.3%. Fable 5.1 posts 55.8%. That’s not a close race, that’s a different weight class. Analysts watching the release flagged the gap as a sign that some of SWE-2’s stronger scores elsewhere might reflect a bit of benchmark overfitting, tuning the model against tests it knew it would face, rather than a uniform jump in raw coding ability.
The part Cognition really wants you to notice: the price
Here’s where SWE-2 makes its actual case. On FrontierCode tasks, it runs 64% cheaper than Fable 5.1. Compared to Cognition’s own previous model, SWE-1.7, it’s 81% cheaper per task and finishes jobs in 58% fewer turns. If you’re a team paying by the token or by the API call, those numbers matter more day to day than a two-point benchmark swing.
Pricing reflects that pitch. Devin’s Pro plan starts at $20 a month, with a one-month free promotion currently running for Pro, Max, and Teams subscribers. SWE-2 is live now in Devin Desktop and the CLI, with rollout underway for Devin Web and the Fusion enterprise API.
Cognition’s growth is the real headline
Buried under the benchmark talk is a business story that’s arguably bigger. Back in May, Cognition’s annualized revenue run-rate sat at $492 million, already a 13-fold jump from a year earlier. By September, that figure had climbed past $900 million, more than triple where the company started 2026.
That kind of growth doesn’t come free. Reports suggest Cognition could burn something like $800 million this year, largely on leased Nvidia servers, with enterprise gross margins hovering near 50%. Running a frontier coding lab is an expensive habit, and Cognition is spending like a company confident the revenue keeps climbing faster than the burn.
What this means if you’re actually choosing a coding assistant
If your workloads look like DeepSWE-style tasks, SWE-2’s numbers are genuinely compelling, and the price gap is real money at scale. If your work leans toward whatever Terminal-Bench 4 is testing, longer, more complex terminal-driven tasks, Fable 5.1 still has a wide lead worth taking seriously before you switch tools based on a headline benchmark alone.
Bottom Line
SWE-2 is a genuinely cheaper, competitive model on two of three major benchmarks, and Cognition’s revenue growth suggests the market agrees it’s worth paying for. But the Terminal-Bench gap is a useful reminder to read past the highlight numbers before picking a coding agent, because the one test a company doesn’t lead with is usually the one that tells you the most.



