The Frontier
Fable 5.1 leads Real-SWE, and still fails 6 of 10 tasks
Two benchmark results landed the same day, and both point at the same gap: agents score well on code they have seen, and badly on code they have not. The best coding agent in the world resolved 38.8% of real enterprise engineering tasks — and no agent solved every task