Meta's Muse Spark 1.3 tops DeepSWE — and undercuts the frontier
Wednesday's releases are landing on top of each other, and today's pair shows the two places the AI race is actually being fought: the benchmark table and the statehouse. Meta shipped a model that claims the top coding score at a fraction of the going rate, while Uber — a company spending billions on robotaxis — is quietly bankrolling the political effort to slow them down.
Meta released Muse Spark 1.3 on Tuesday, and on the company's own numbers it now leads the field on long-horizon coding. The model scores 75.4 on DeepSWE v1.1, ahead of Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0, according to Meta's published benchmark table — a result that would have been unthinkable for Meta a quarter ago. Independent evaluator Artificial Analysis puts the max-reasoning variant at 62 on its Intelligence Index, behind only Anthropic's Fable 5.1 and Opus 5, with the gains concentrated almost entirely in agentic work: GDPval-AA v2 up 94 to 139 Elo, Terminal-Bench 2.1 up five to six points, and a top-of-table 52% on Tau3-Bench Banking. Meta says the model uses roughly 20% fewer tool calls and 25% fewer tokens than Spark 1.2 to do the same work, and researchers credit a large increase in grading compute aimed at laziness, hedging, and reward-hacking.
The number that actually matters is the price, and it has not moved: $1.25 per million input tokens and $4.25 per million output tokens, unchanged from 1.2. That makes Spark 1.3 the cheapest way to buy this level of capability — Artificial Analysis found no model scoring 59 or above costs less per task, at about $0.55 against roughly $3.14 for a comparable Anthropic run. We covered Google's version of this argument yesterday in Google ships Gemini 3.8 Flash, gating its cyber model behind Fairwind, where Flash scored 59 at $0.75 and $3.75. Twelve months of that pressure is what collapses a pricing tier, and Meta is now the one applying it.
Two caveats before anyone rewrites their stack. The headline scores are self-reported, and the top variant is not shippable yet — max reasoning is still in limited partner preview pending what Meta calls additional safety testing, so the 62 is a promise, not a product. The genuinely structural piece is the roadmap: Meta has committed to open-weighting Muse Spark and has the larger Watermelon model behind it. A frontier-class coding model with weights attached, at this price, is a different competitive problem than a frontier-class API — it removes the ability to charge a premium for the weights themselves, and it is the reason rivals are watching the release date closer than the benchmark.
Uber is funding the political case against robotaxis while building its own. The Financial Times reports the company is aligning with drivers' unions in a bid to slow autonomous rollouts across the US — an alliance confirmed in detail by CNBC's reporting on Washington, D.C.'s AV deployment bill, where Uber's director of AV and AI policy told the council the legislation "largely ignores the workforce transition" and that "one AV in California now performs the work of roughly four drivers." In New Jersey, Uber lobbyists have pushed a requirement that human drivers handle 85% of rides on any platform offering robotaxi service, which would force Waymo and Tesla onto Uber's network. Uber's CFO says the company is reacting to bills that would have produced no AVs at all; either way, the company that spent $59 million to keep drivers classified as contractors is now arguing that drivers need protecting.
The timing explains the sincerity question. On Wednesday Uber announced it is cutting 3,300 jobs, about 10% of staff and its largest reduction since 2020, with CEO Dara Khosrowshahi framing it as flattening management layers to speed decisions — while the company plans to put more than $10 billion into robotaxis and its partnership with Waymo comes apart. Call it what it is: a middleman defending the toll booth. If the 85% rule holds, Uber wins whether the cars have drivers or not.
What to watch: whether Muse Spark's open weights land at the capability Meta is claiming — and whether a state legislature actually passes a human-driver quota.
If a lab gives frontier-grade weights away to win on price, is that a public good or just a margin war nobody else can afford? Tell us in the comments.
Sources: Meta AI Research — Introducing Muse Spark 1.3 · Artificial Analysis — Muse Spark 1.3 intelligence, performance and price · Axios — Meta debuts Muse Spark 1.3 · Bloomberg — Meta releases more powerful AI model, edging closer to rivals · Financial Times — Uber allies with driver unions in bid to slow robotaxi rollout · CNBC — Within Uber-Waymo split, a key labor battle over AVs is being waged in the nation's capital · USA Today — Why Uber is cutting 3,300 jobs as robotaxi competition grows · Techmeme — Uber is aligning with drivers' unions