AIToday
Large Language ModelsAI Coding AssistantsITmedia AI+Published: Sep 3, 2026, 13:00 JST2 min read

Meta unveils Muse Spark 1.3 with coding gains

Meta unveils Muse Spark 1.3 with coding gains

3 Key Points

  1. What happened

    Meta announced Muse Spark 1.3 on September 2 and started offering it to developers through Muse Code and the Meta Model API the same day. The model improves agentic tasks and coding performance, and Meta says it also improved practical usability.

  2. Why it matters

    On Meta's benchmark table, Muse Spark 1.3 scored 75.4 on DeepSWE v1.1, up from 55.0 for Muse Spark 1.2, and ahead of Claude Opus 5's 74.0 and GPT-5.6 Sol's 73.0. It also scored 98.1 on MRCR's 512K-1M token range, far above 1.2's 55.5 and GPT-5.6 Sol's 73.8. However, Claude Opus 5 scored highest on four of six agentic benchmarks, including GDPVal-AA v2.

  3. What to watch

    The release hinges on whether the highest “max reasoning” mode clears additional safety testing, as it is not yet available. A larger model and an open-weights version of Muse Spark 1.2 are also planned after final safety fine-tuning; a watermelon emoji in posts may hint at the next model’s codename.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

Muse Spark 1.3 is the latest version of the native multimodal reasoning model first introduced in April, following the establishment of Meta Superintelligence Labs. The release targets two recurring pain points in practical AI use: staying reliable over long agentic workflows and cutting wasteful back-and-forth in coding. Meta says the model now asks for clarification on vague instructions, requests help when stuck, and checks before high-impact actions, and it has been trained not to fabricate results when it fails.

The benchmark results are mixed. Muse Spark 1.3 leads on the three coding-related measures shown, including a sharp jump on DeepSWE v1.1, and it dominates the long-context MRCR range tested at 512K-1M tokens. Yet on six agentic measures, Claude Opus 5 posts the top score in four, including GDPVal-AA v2, suggesting the agentic crown is not uncontested. The company's own comparisons, which include OpenAI's GPT-5.6 Sol, are clearly promotional, though the gains over Muse Spark 1.2 are concrete.

Meta's leadership is leaning into cost and capability simultaneously. Zuckerberg's framing that frontier performance is now so cheap it does not need measuring points to an efficiency story underway. The roadmap extends beyond this release: a larger model is planned, and an open-weights version of Muse Spark 1.2 is expected after final safety fine-tuning, indicating a phased approach that may continue with a model reportedly codenamed watermelon.

FAQ
How does Muse Spark 1.3 compare with previous versions?
Compared with Muse Spark 1.2, Meta engineers found tool calls down about 20% and token consumption down about 25%, reducing unnecessary interactions and redundancy. On DeepSWE v1.1, the score rose from 55.0 to 75.4.
When will the max reasoning mode be available?
Max reasoning will be offered after additional safety tests are completed. The other reasoning modes from the previous version are available from September 2.

Also reported by THE DECODER

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Apple's Siri AI Japanese version to launch in OctoberITmedia AI+ · 14m ago
  • Meta launches Muse AI assistantSemafor Tech · 14m ago
  • OpenAI GPT-6 Astra debuts on Snowflake Cortex AISnowflake AI Blog · 14m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleLoeb's Third Point Exits Nvidia, Broadcom