AIToday
Large Language ModelsAI Coding AssistantsTHE DECODERPublished: Aug 12, 2026, 22:01 JST4 min read

Microsoft's MAI Code 1.1 Flash trails DeepSeek on price and performance

Microsoft's MAI Code 1.1 Flash trails DeepSeek on price and performance

Key takeaway

  • Microsoft released MAI Code 1.1 Flash, a coding model that costs a quarter of its predecessor and is 25 percent more token-efficient, but trails DeepSeek-V4-Flash on both price and performance.

  • The move reflects Microsoft's strategy to reduce costs by defaulting to its own proprietary models in its ecosystem, even as the company publicly positions itself as an open-AI champion.

3 Key Points

  1. What happened

    Microsoft released MAI Code 1.1 Flash, a code model for GitHub Copilot that is 25 percent more token-efficient and costs a quarter of its June predecessor. On benchmarks, it edges past its earlier model and Anthropic's Haiku 4.5 and OpenAI's GPT-5.4 mini, but DeepSeek-V4-Flash-0731 outperforms it on Terminal Bench 2.1 (82.7% vs. 62.9%).

  2. Why it matters

    Microsoft's own model costs $0.20 per input token and $1.20 per output token, significantly more than DeepSeek-V4-Flash at $0.14 and $0.28 respectively. The company is promoting its proprietary MAI models while publicly championing open-weights alternatives like DeepSeek, creating a credibility gap—especially since Microsoft is likely to make its own models the default in its ecosystem, locking users into a weaker but higher-margin option.

  3. What to watch

    Microsoft buried the benchmark results in the model card and only highlighted vague improvement metrics in its announcement, omitting direct comparisons to competitors. DeepSeek-V4-Flash-0731 did not publish its SWE-bench Verified score, making full performance comparison incomplete.

In Depth

Read the full story

Microsoft has released MAI Code 1.1 Flash, a new code generation model designed for GitHub Copilot. According to Microsoft's claims, the model offers three key improvements: it writes better code, is 25 percent more token-efficient (meaning it uses fewer tokens to complete the same task), and costs a quarter of its June predecessor. The model was trained using "hundreds of thousands of reinforcement-learning environments in GitHub Copilot," and developers accepted 4 percent more of its output compared to the prior version.

On published benchmarks, MAI Code 1.1 Flash shows incremental progress. It achieved 72.6% on SWE-bench Verified (compared to its predecessor's 71.6%), and 62.9% on Terminal Bench 2.1 (a jump from 51.7%). These scores edge past Anthropic's Haiku 4.5 (69.8% on SWE-bench, 49.4% on Terminal Bench) and OpenAI's GPT-5.4 mini (69.2% and 60.7% respectively). However, DeepSeek-V4-Flash-0731 significantly outperforms MAI Code 1.1 Flash on Terminal Bench 2.1, achieving 82.7%—a gap of nearly 20 percentage points. DeepSeek did not publish its SWE-bench Verified result, preventing a full comparison on that metric.

Pricing tells a similar story of trade-offs. MAI Code 1.1 Flash costs $0.20 per input token and $1.20 per output token. While this is cheaper than Anthropic's Claude Haiku 4.5 ($1.00 input / $5.00 output), it remains significantly more expensive than DeepSeek-V4-Flash at $0.14 per input token and $0.28 per output token. Even accounting for token efficiency, DeepSeek's cost advantage is likely substantial. Microsoft notably buried these benchmark results in the model card rather than highlighting them in its official announcement, instead choosing to tout only vague improvement metrics like "code survival rose 4% and return visits increased 9%" over its predecessor, omitting any direct comparisons to competitors.

The release exposes a disconnect between Microsoft's public rhetoric and its actual product strategy. The company has positioned itself as an enthusiastic supporter of open-weights AI models—freely available alternatives like DeepSeek—yet continues to invest in proprietary, weaker, and pricier alternatives. This mirrors Microsoft's recent restructuring of Copilot itself, where the company swapped out OpenAI and Anthropic models for cheaper MAI alternatives to reduce costs, accepting worse performance in exchange for better margins. By defaulting to its own models within the Microsoft ecosystem and relying on the reality that most users never actively select a different AI model, Microsoft appears to be betting that it can lock up significant market share despite offering inferior performance at higher cost.

Context & Analysis

Microsoft's release of MAI Code 1.1 Flash reflects a clear tension in the company's public messaging and actual strategy. While Microsoft has positioned itself as a champion of open-weights AI—praising freely available models like DeepSeek—it is simultaneously investing resources into proprietary in-house models that are both weaker and more expensive than their open alternatives. The benchmark data tells the story: MAI Code 1.1 Flash underperforms DeepSeek-V4-Flash on Terminal Bench 2.1 (62.9% vs. 82.7%), yet costs significantly more per token ($0.20 input / $1.20 output vs. $0.14 / $0.28). This gap is particularly striking given that the article notes Microsoft itself acknowledges DeepSeek's superior capability.

The root cause appears to be straightforward economics. Microsoft recently restructured its Copilot offerings to replace OpenAI and Anthropic models with cheaper MAI alternatives, explicitly prioritizing cost reduction over performance. MAI Code 1.1 Flash follows the same playbook: it is cheaper than its predecessor and slightly cheaper than OpenAI's and Anthropic's models, allowing Microsoft to improve its own margins. The company's strategy relies on market lock-in; by defaulting to its proprietary model within the Microsoft ecosystem, the company expects to capture market share without requiring users to actively choose it.

Microsoft's handling of the announcement also suggests awareness of the competitive gap. The company buried benchmark results in the model card rather than leading with them, and highlighted only vague metrics like "code survival rose 4%" instead of head-to-head comparisons. This selective disclosure hints at discomfort with the model's actual competitive position relative to DeepSeek.

FAQ

How much cheaper is MAI Code 1.1 Flash than before?
Microsoft says it costs a quarter of its June predecessor. The new model costs $0.20 per input token and $1.20 per output token, compared to DeepSeek-V4-Flash at $0.14 and $0.28 respectively.
How much better is the new model at writing code?
Developers accepted 4 percent more of its output than the previous version. On benchmarks, MAI-Code-1.1-Flash achieved 72.6% on SWE-bench Verified and 62.9% on Terminal Bench 2.1, outperforming Anthropic's Haiku 4.5 (69.8% and 49.4%) and OpenAI's GPT-5.4 mini (69.2% and 60.7%), but underperforming DeepSeek-V4-Flash-0731 on Terminal Bench 2.1 at 82.7%.
Why does Microsoft favor its own model over open-weights alternatives?
Microsoft is likely motivated by cost and margin: by making its own models the default in its ecosystem, the company locks up market share since most users never actively choose a specific AI model, while the proprietary model remains pricier than DeepSeek alternatives.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next article87.5% of venture dollars flowed to AI in H1 2026

The AI news that matters, in one minute each morning.

Sign up free