The number I care about isn’t 50.0% on FrontierCode. It’s 18 steps versus 48 before the first real edit. I’ve watched agents grep their way across a repo for ten minutes to change a config value, and Cognition admits users complained about exactly that with SWE-1.7. SWE-2 medium scores higher while taking 58% fewer turns and costing 81% less.
So my move is to stop defaulting to the biggest effort level. Cognition trained medium, high and max in one RL run with cost penalties tuned per level, and says medium steps into action quicker on simple and intermediate tasks. That maps to how I work: mostly small edits in a codebase I already know. Save high and max for the genuinely uncertain jobs.
The story — Cognition released SWE-2, post-trained from the 2.8T-parameter Kimi K3, scoring 50.0% on FrontierCode 1.1 Main — within a point of Fable 5.1 at 64% less cost — and beating Grok 4.6 and SWE-1.7 on score and cost. Training scaled RL to the multi-trillion-parameter regime, using a linear cost penalty per effort level in a single run. It ships today in Devin Desktop and CLI. (Source)