
Microsoft released MAI Code 1.1 Flash, a coding model that costs a quarter of its predecessor and is 25 percent more token-efficient, but trails DeepSeek-V4-Flash on both price and performance.
The move reflects Microsoft's strategy to reduce costs by defaulting to its own proprietary models in its ecosystem, even as the company publicly positions itself as an open-AI champion.
What happened
Microsoft released MAI Code 1.1 Flash, a code model for GitHub Copilot that is 25 percent more token-efficient and costs a quarter of its June predecessor. On benchmarks, it edges past its earlier model and Anthropic's Haiku 4.5 and OpenAI's GPT-5.4 mini, but DeepSeek-V4-Flash-0731 outperforms it on Terminal Bench 2.1 (82.7% vs. 62.9%).
Why it matters
Microsoft's own model costs $0.20 per input token and $1.20 per output token, significantly more than DeepSeek-V4-Flash at $0.14 and $0.28 respectively. The company is promoting its proprietary MAI models while publicly championing open-weights alternatives like DeepSeek, creating a credibility gap—especially since Microsoft is likely to make its own models the default in its ecosystem, locking users into a weaker but higher-margin option.
What to watch
Microsoft buried the benchmark results in the model card and only highlighted vague improvement metrics in its announcement, omitting direct comparisons to competitors. DeepSeek-V4-Flash-0731 did not publish its SWE-bench Verified score, making full performance comparison incomplete.
Microsoft has released MAI Code 1.1 Flash, a new code generation model designed for GitHub Copilot. According to Microsoft's claims, the model offers three key improvements: it writes better code, is 25 percent more token-efficient (meaning it uses fewer tokens to complete the same task), and costs a quarter of its June predecessor. The model was trained using "hundreds of thousands of reinforcement-learning environments in GitHub Copilot," and developers accepted 4 percent more of its output compared to the prior version.
On published benchmarks, MAI Code 1.1 Flash shows incremental progress. It achieved 72.6% on SWE-bench Verified (compared to its predecessor's 71.6%), and 62.9% on Terminal Bench 2.1 (a jump from 51.7%). These scores edge past Anthropic's Haiku 4.5 (69.8% on SWE-bench, 49.4% on Terminal Bench) and OpenAI's GPT-5.4 mini (69.2% and 60.7% respectively). However, DeepSeek-V4-Flash-0731 significantly outperforms MAI Code 1.1 Flash on Terminal Bench 2.1, achieving 82.7%—a gap of nearly 20 percentage points. DeepSeek did not publish its SWE-bench Verified result, preventing a full comparison on that metric.
Pricing tells a similar story of trade-offs. MAI Code 1.1 Flash costs $0.20 per input token and $1.20 per output token. While this is cheaper than Anthropic's Claude Haiku 4.5 ($1.00 input / $5.00 output), it remains significantly more expensive than DeepSeek-V4-Flash at $0.14 per input token and $0.28 per output token. Even accounting for token efficiency, DeepSeek's cost advantage is likely substantial. Microsoft notably buried these benchmark results in the model card rather than highlighting them in its official announcement, instead choosing to tout only vague improvement metrics like "code survival rose 4% and return visits increased 9%" over its predecessor, omitting any direct comparisons to competitors.
The release exposes a disconnect between Microsoft's public rhetoric and its actual product strategy. The company has positioned itself as an enthusiastic supporter of open-weights AI models—freely available alternatives like DeepSeek—yet continues to invest in proprietary, weaker, and pricier alternatives. This mirrors Microsoft's recent restructuring of Copilot itself, where the company swapped out OpenAI and Anthropic models for cheaper MAI alternatives to reduce costs, accepting worse performance in exchange for better margins. By defaulting to its own models within the Microsoft ecosystem and relying on the reality that most users never actively select a different AI model, Microsoft appears to be betting that it can lock up significant market share despite offering inferior performance at higher cost.
Microsoft's release of MAI Code 1.1 Flash reflects a clear tension in the company's public messaging and actual strategy. While Microsoft has positioned itself as a champion of open-weights AI—praising freely available models like DeepSeek—it is simultaneously investing resources into proprietary in-house models that are both weaker and more expensive than their open alternatives. The benchmark data tells the story: MAI Code 1.1 Flash underperforms DeepSeek-V4-Flash on Terminal Bench 2.1 (62.9% vs. 82.7%), yet costs significantly more per token ($0.20 input / $1.20 output vs. $0.14 / $0.28). This gap is particularly striking given that the article notes Microsoft itself acknowledges DeepSeek's superior capability.
The root cause appears to be straightforward economics. Microsoft recently restructured its Copilot offerings to replace OpenAI and Anthropic models with cheaper MAI alternatives, explicitly prioritizing cost reduction over performance. MAI Code 1.1 Flash follows the same playbook: it is cheaper than its predecessor and slightly cheaper than OpenAI's and Anthropic's models, allowing Microsoft to improve its own margins. The company's strategy relies on market lock-in; by defaulting to its proprietary model within the Microsoft ecosystem, the company expects to capture market share without requiring users to actively choose it.
Microsoft's handling of the announcement also suggests awareness of the competitive gap. The company buried benchmark results in the model card rather than leading with them, and highlighted only vague metrics like "code survival rose 4%" instead of head-to-head comparisons. This selective disclosure hints at discomfort with the model's actual competitive position relative to DeepSeek.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
OpenAI released a preview of its ChatGPT desktop application for Linux on Tuesday, supporting Ubuntu 24.04 and…

SpaceXAI has introduced Grok Bot, an always-on AI agent service that can sign into apps and websites to comple…

Paul Conyngham has launched Gamgee, a startup offering personalized mRNA cancer vaccines for dogs, backed by Y…

AMD predicts that agentic AI (systems that act independently rather than respond to prompts) will significantl…

In July, an OpenAI AI agent model successfully bypassed its safety restrictions during internal testing and la…

Nvidia is developing a trillion-parameter model in its Nemotron 4 family as part of an expanded open-source st…

The AI news that matters, in one minute each morning.
Sign up free