Cognition's SWE-2 closes in on the frontier — at a quarter of the price

Share
Cognition's SWE-2 closes in on the frontier — at a quarter of the price

Coding agents stopped being a two-horse race today.

Cognition has launched SWE-2, the model behind Devin, and by its own published numbers it lands closer to the frontier than any release in the company's history — while undercutting the models it trails on cost. On the FrontierCode 1.1 Main and DeepSWE 1.1 coding benchmarks, Cognition says SWE-2 beats its own SWE-1.7 and xAI's Grok 4.6 on both score and cost, matches GPT-5.6 Sol and Anthropic's Fable 5/5.1 at a fraction of their price, and comes within a few points of GPT-6 Astra at roughly a quarter of the cost. The efficiency gains are the part worth staring at: on FrontierCode, SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average.

Why it matters: the frontier labs still own the top of the leaderboard, but the floor keeps rising. Cognition claims SWE-2 was trained entirely with post-training on top of an existing base model — and the interesting detail in the writeup is how: a cost-penalized reward that trains every effort level (medium, high, max) end-to-end in a single RL run, rather than the usual separate-experts-then-distill dance, plus a length-weighted baseline that stabilizes training for free. The thesis: intelligence and efficiency aren't opposing dials, because a smarter agent wastes fewer turns.

The post also runs a quiet trustworthiness test: SWE-2 passed 98.0% of questions about politically sensitive topics in China, answering substantively in Simplified, Traditional, and English without adopting the official PRC position — a pointed dig at the openness of Chinese open-weights rivals like Kimi, which Cognition explicitly dissects in its training writeup. Skepticism applies, as always: Cognition reports these numbers on its own harness and picked the comparisons; officechai pegs the broader claim at roughly 70% lower cost near the frontier, which is still Cognition's math. But even discounted, an independent coder-agent lab this close to the biggest labs changes procurement math for every engineering team watching its agent bill.

What to watch: independent runs of SWE-2 on FrontierCode and DeepSWE outside Cognition's harness — those numbers decide whether this is a real Pareto shift or a well-chosen chart.

Would you swap your current coding agent for a cheaper model that's a few points behind the frontier? Tell us in the comments.

Sources: Cognition · officechai · Cognition on X