Meta hires MongoDB's CEO to run its enterprise AI push
Two stories this hour: the largest social platform is repackaging its AI stack for corporate buyers, and the benchmarkers just measured how often the frontier models refuse to do security work at all.
Meta launched Meta Enterprise Platform on Monday and hired MongoDB's chief executive, Chirantan "CJ" Desai, to run it as chief enterprise platform officer, reporting to Mark Zuckerberg. Zuckerberg called it "the next major pillar of our business" and said the unit will initially sell the company's full stack to businesses — the Muse agent, a business agent, and a coding tool. Desai had run MongoDB for less than a year, taking over in November 2025 from Dev Ittycheria, who ended an 11-year run; MongoDB's board named Ittycheria interim president and CEO. It is a blunt admission that Meta has spent enormous sums on AI and has no enterprise sales machine to show for it. Meta's stock fell about 4% — but MongoDB's fell roughly a quarter of its value, which tells you where the market thinks the talent actually lives.
The timing is telling. Meta's consumer agent Muse has already overtaken ChatGPT as the leading free iOS app, and the company now says the enterprise push is how it monetizes the buildout. It has been here before: Meta ships Muse: a consumer agent inside a sealed VM was the consumer half of the same bet, and the enterprise half is harder, because the buyer is a procurement team rather than a download. Desai inherits a field Microsoft has spent a decade digging into, with no announced pricing and no customer list yet.
Artificial Analysis published the Cyber Index, and the finding that matters isn't the leaderboard — it's that the frontier models refuse the work. The index grades models on defensive cyber tasks: find a vulnerability, reproduce it, and patch it without breaking the code around it. Five frontier models — GPT-6 Sol, GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1 and Gemini 3.8 Flash — decline 32% to 38% of its tasks and trail the leaders by 19 to 31 points. On the CyberGym-E2E-AA component, GPT-6 Sol and GPT-6 Astra refused every task, while Opus 5.5 refused 98% and Fable 5.1 99%. Grok 4.7 and MiMo-V2.6-Pro lead the index at 56; GPT-6 Luna scores 53 for 12 cents per task against Grok's $11.67; and the best model in the suite finds only 41% of its expert-verified issues.
The framing matters here. Artificial Analysis calls this a capability index, and the honest read is that capability isn't being measured for the models that matter most — safety alignment is capping the score, and the labs will take that trade. We covered the CyberGym track this morning when a 27B Alibaba model led the model track; this index adds the decision layer by folding in Collinear's 120-task memory-safety set and Vercel's DeepsecBench. The launch framing leans on the word "Alliance" — IBM and Nvidia are credited with input on scope and methodology, not with evaluations or scores — and the index itself rests on the same small group of contributed evals.
If a model quietly declines a fifth of the tasks that keep software safe, should the refusal rate be published next to every capability score? Tell us in the comments.
Sources: Meta Newsroom · CNBC · Wall Street Journal · Artificial Analysis — Cyber Index · Crypto Briefing · Collinear — CWE-Bench · Vercel — DeepsecBench