Salesforce and Nvidia ship Koa, a reasoning model for Agentforce
Two stories on the same question this afternoon: what happens to frontier-lab pricing when the enterprise stack stops renting reasoning by the token. Plus a security layer for the robot fleet.
Salesforce debuted Koa, a reasoning model it built with Nvidia on top of the open-weight Nemotron base, and says it will handle the multi-step thinking inside Agentforce that used to be sent out to Claude or ChatGPT. Jayesh Govindarajan, Salesforce's EVP of AI, put it plainly: the company had built many small task-specific models, but reasoning had always been something it relied on frontier providers for — until now. The base model had been the blocker, and he was blunt about why Nemotron cleared it: sovereign American, state of the art, and with data provenance a vendor can actually attest to (he added, of Alibaba's Qwen, that nobody knows what it trains on). Post-training used no customer data at all — Salesforce simulated enterprise workflows across 14 industries, irate callers to deal-closing reps included, and combined supervised fine-tuning, reinforcement learning and group relative policy optimization to get multi-step tool calls reliable. Salesforce claims Koa matches or beats leading general-purpose models on CRM tasks with three times fewer errors, hosted inside its own trust boundary. It is pilot-only now, with general availability targeted for winter, and the full Nemotron family lands in Missionforce Operations in October for air-gapped government workloads.
The read: this is the first serious attempt to take the highest-token-count enterprise workload — long-horizon agent reasoning — and run it in-house on a named vendor's weights. The labs' moat was general reasoning; a 27-year-old CRM company just claimed it doesn't need to buy that anymore.
Mozilla's State of Open Source AI report, published today, measures the gap between US closed frontier models and the best Chinese open-weights models at 4.4 months, and Ars Technica previewed it ahead of release. On the Artificial Analysis Intelligence Index, Moonshot AI's Kimi K3 sits three points behind Anthropic's Fable 5 while costing 30 percent as much; on Terminal-Bench 2.1, run on a neutral harness, Z.ai's GLM 5.2 landed within a point of Claude Opus 4.7 and 4.8 at roughly a fifth the cost per completed task. Mozilla's CTO Raffi Krikorian framed the premium as workload-specific rather than organization-specific: closed models earn their fee on expert professional work, high-intensity retrieval and long context — concretely, the 8-to-12-hour tasks open models can't yet finish reliably. DoorDash already routes routine work to Kimi and saves Fable for the hard stuff, and eight of OpenRouter's top ten models by August token volume now carry open weights. We flagged the same economics from the other side — a 33B model from Singapore matched DeepSeek V4 Pro on free tokens and Z.ai's GLM 5.3 Flash tops benchmarks at one tenth the price. Krikorian's sharper line: most open models the world runs on are Chinese, and those labs are running the Android playbook — give the model away, own the ecosystem around it.
Exein raised $270 million at a $1.7 billion valuation to sell what it calls the security layer for physical AI, making it Italy's newest unicorn. The Rome startup says its embedded-device protections are in more than 2 billion connected devices across aerospace, industrial automation, automotive, energy, healthcare and semiconductors, and founder Fabrizio Cuozzo claims 400 percent year-on-year growth. His thesis cuts both ways: the same edge-AI supply chain that is putting models on robots and drones is also, in his telling, arming attackers, because ungoverned open-source models make it cheaper to find novel ways into physical devices. Money goes to acquisitions and to US and APAC hiring, with Asia Pacific already driving half its revenue.
What to watch: whether Koa's three-times-fewer-errors claim survives customers outside a pilot — Salesforce's own benchmark, not an independent one.
If your agent workloads are mostly under eight hours of human-equivalent work, would a fifth of the cost be enough to switch off a frontier API today? Tell us in the comments.
Sources: TechCrunch · SiliconANGLE · Ars Technica · Mozilla · Exein · TechCrunch