Microsoft's MAI Code 1.1 Flash trails DeepSeek on price and performance
Microsoft shipped a new coding model for GitHub Copilot that's four times cheaper than its own predecessor — and still loses to DeepSeek on both price and the benchmark that matters most.
Microsoft has released MAI Code 1.1 Flash, a coding model now in production in GitHub Copilot that the company says writes better code at a quarter of the cost of the version it launched at Build in June. Microsoft claims 25 percent greater token efficiency, a 22 percent improvement on Terminal-Bench 2.1 in Copilot CLI, and a 15 percent jump on .NET tasks, with code survival up 4 percent and return visits up 9 percent — gains trained on what it calls "hundreds of thousands of reinforcement-learning environments in GitHub Copilot."
The model's own numbers tell a flatter story. On Terminal-Bench 2.1, MAI Code 1.1 Flash scores 62.9 percent — ahead of its predecessor's 51.7 percent, Haiku 4.5's 49.4 percent, and GPT-5.4 mini's 60.7 percent — while DeepSeek-V4-Flash-0731 posts 82.7 percent on the same test. The pricing gap is the sharper one: DeepSeek V4 Flash charges $0.14 per million input tokens and $0.28 for output, against MAI Code 1.1 Flash's $0.20 in and $1.20 out — more than four times the output price, before the token-efficiency difference is even factored in. As The Decoder noted, the direct comparisons live in the model card while the announcement itself touts only the vague improvement metrics.
That pattern matters beyond one benchmark. Microsoft spent the summer casting itself as an open-AI champion, and its recent Copilot shakeup swapped OpenAI and Anthropic models for cheaper in-house MAI alternatives to cut costs — a trade of performance for margin. MAI Code 1.1 Flash fits the same playbook: a proprietary model that trails the freely available open-weight alternative the company keeps praising, with no open-weights release planned. The bet is distribution over capability — make MAI the default in Copilot and most users never actively choose a model at all. It's a margin play dressed as a model release, and it only works while the capability gap stays small enough that users don't notice. We covered the independent run of the benchmark that exposes it last week — Independent run confirms DeepSeek V4 Flash's 82.7% score.
What to watch: whether Microsoft ever ships MAI Code 1.1 Flash with open weights, and whether the next Copilot default picks capability over margin.
Microsoft keeps shipping proprietary models that trail the open weights it praises — is that a sustainable strategy, or is default-model lock-in the whole point? Tell us in the comments.
Sources: Microsoft AI · The Decoder · Independent run confirms DeepSeek V4 Flash's 82.7% score