AIToday
Large Language ModelsAI Coding AssistantsITmedia AI+Published: Sep 3, 2026, 13:00 JST2 min read

Meta unveils Muse Spark 1.3 with coding gains

Meta unveils Muse Spark 1.3 with coding gains

Key takeaway

  • Meta has launched Muse Spark 1.3, focused on coding and agentic tasks.

  • The model beats its predecessor and rivals on several coding benchmarks.

  • It is available now for developers.

3 Key Points

  1. What happened

    Meta announced Muse Spark 1.3 on September 2 and started offering it to developers through Muse Code and the Meta Model API the same day. The model improves agentic tasks and coding performance, and Meta says it also improved practical usability.

  2. Why it matters

    On Meta's benchmark table, Muse Spark 1.3 scored 75.4 on DeepSWE v1.1, up from 55.0 for Muse Spark 1.2, and ahead of Claude Opus 5's 74.0 and GPT-5.6 Sol's 73.0. It also scored 98.1 on MRCR's 512K-1M token range, far above 1.2's 55.5 and GPT-5.6 Sol's 73.8. However, Claude Opus 5 scored highest on four of six agentic benchmarks, including GDPVal-AA v2.

  3. What to watch

    Mark Zuckerberg called it the biggest leap yet in coding and agentic work, describing the performance as so cheap it doesn't need measuring. The highest reasoning mode, max reasoning, will only be offered after additional safety testing. Meta also plans to release a larger model and an open-weights version of Muse Spark 1.2 after final safety fine-tuning, and a watermelon emoji in posts suggests it may be the codename for the next model.

Ask the AI about this article →

Context & Analysis

Muse Spark 1.3 is the latest version of the native multimodal reasoning model first introduced in April, following the establishment of Meta Superintelligence Labs. The release targets two recurring pain points in practical AI use: staying reliable over long agentic workflows and cutting wasteful back-and-forth in coding. Meta says the model now asks for clarification on vague instructions, requests help when stuck, and checks before high-impact actions, and it has been trained not to fabricate results when it fails.

The benchmark results are mixed. Muse Spark 1.3 leads on the three coding-related measures shown, including a sharp jump on DeepSWE v1.1, and it dominates the long-context MRCR range tested at 512K-1M tokens. Yet on six agentic measures, Claude Opus 5 posts the top score in four, including GDPVal-AA v2, suggesting the agentic crown is not uncontested. The company's own comparisons, which include OpenAI's GPT-5.6 Sol, are clearly promotional, though the gains over Muse Spark 1.2 are concrete.

Meta's leadership is leaning into cost and capability simultaneously. Zuckerberg's framing that frontier performance is now so cheap it does not need measuring points to an efficiency story underway. The roadmap extends beyond this release: a larger model is planned, and an open-weights version of Muse Spark 1.2 is expected after final safety fine-tuning, indicating a phased approach that may continue with a model reportedly codenamed watermelon.

FAQ

How does Muse Spark 1.3 compare with previous versions?
Compared with Muse Spark 1.2, Meta engineers found tool calls down about 20% and token consumption down about 25%, reducing unnecessary interactions and redundancy. On DeepSWE v1.1, the score rose from 55.0 to 75.4.
When will the max reasoning mode be available?
Max reasoning will be offered after additional safety tests are completed. The other reasoning modes from the previous version are available from September 2.

Also reported by THE DECODER

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • ChatGPT, Claude, Grok outages resolvedITmedia AI+ · 1h ago
  • Meta stock jumps 4% on AI model parity claimYahoo Finance AI · 1h ago
  • OpenAI's Astra model raises AI monitoring concernsFortune AI · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCaterpillar Partners with FieldAI, NVIDIA for AI Robotics Push