AIToday
Large Language ModelsHacker NewsPublished: Aug 23, 2026, 10:00 JST2 min read

MiniMax M3 Medium reaches 73.17% F1 on DeepSearchQA, matching GPT-5 High

MiniMax M3 Medium reaches 73.17% F1 on DeepSearchQA, matching GPT-5 High

Key takeaway

  • MiniMax M3 Medium scored 73.17% F1 on DeepSearchQA, matching GPT-5 High Reasoning.

  • The model used You.com search and content tools with an optimized research skill.

  • It trails top agents like Gemini Deep Research but offers transparent, reproducible results at $0.532 per task.

3 Key Points

  1. What happened

    MiniMax M3 Medium, paired with You.com research tools and optimized via an auto-research loop, achieved 73.17% adjusted F1 on the DeepSearchQA benchmark. The full run completed on 2026-08-21 across 900 tasks and 2,700 trials, with results published on GitHub and Hugging Face.

  2. Why it matters

    The score is competitive with GPT-5 High Reasoning's 73.24% F1 in the DeepSearchQA paper's Table 4. However, MiniMax M3 Medium trails the leading agents—Gemini Deep Research Agent (81.90% F1) and GPT-5 Pro High Reasoning (78.98% F1)—and shows a higher fully-incorrect rate (19.01%) than those leaders.

  3. What to watch

    The evaluation cost $478.41 total ($328.14 model, $150.27 You.com API), averaging $0.532 per task. Median latency was 54.3 seconds; 95th percentile was 238.8 seconds. All prompts, graded results, and trajectories are publicly available for inspection and reproduction.

Ask the AI about this article →

Context & Analysis

MiniMax M3 Medium's 73.17% F1 result represents a meaningful checkpoint in the evolution of open and commercial reasoning models. The score places it at parity with GPT-5 High Reasoning, which the DeepSearchQA paper reports at 73.24% F1. However, the body explicitly acknowledges that MiniMax M3 Medium lags behind the top-tier agents: Gemini Deep Research Agent (81.90% F1) and GPT-5 Pro High Reasoning (78.98% F1). The fully-incorrect rate of 19.01% also suggests room for improvement compared to those leading systems.

The evaluation methodology adds transparency to the claim. Rather than a paper assertion, the results are backed by 2,700 trials across 900 tasks, with all intermediate artifacts (trajectories, grades, prompts) published for inspection and reproduction. This open artifact approach—coupled with the You.com MCP tools and an optimized research skill tuned to the model and tool surface—establishes a clear foundation for further iteration. The total cost of $478.41 and median latency of 54.3 seconds provide operational context: the model operates within a reasonable resource envelope for research tasks, though the 95th percentile latency of 238.8 seconds indicates variability in harder queries.

The fact that the evaluation is described as competitive with GPT-5 High Reasoning, not with the top agents, sets realistic expectations. The body frames this as a step forward for the MiniMax M3 Medium line, not a breakthrough that overtakes the field. For practitioners, the key signals are the public artifacts (enabling validation and fine-tuning) and the documented cost and latency trade-offs.

FAQ

How does MiniMax M3 Medium's score compare to other models?
MiniMax M3 Medium achieved 73.17% adjusted F1, matching GPT-5 High Reasoning (73.24% F1). However, it trails Gemini Deep Research Agent (81.90% F1) and GPT-5 Pro High Reasoning (78.98% F1), according to the DeepSearchQA paper's Table 4.
What tools and infrastructure were used?
The evaluation used the You.com MCP you-search and you-contents tools paired with a MiniMax-oriented research skill optimized via an auto-research loop. The judge model was deepseek/deepseek-v4-flash-0731, with qwen/qwen3.6-flash as fallback.
Are the results reproducible?
Yes. All artifacts—prompts.jsonl, results.jsonl, trajectories.jsonl, graded.jsonl, and summary.json—are publicly available on GitHub and Hugging Face. The repository includes commands and environment variables to reproduce the full evaluation.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Visko raises $10M, launches live AI video model OrbisSiliconANGLE AI · 1h ago
  • Runway unveils Solaris, an AI that generates app interfaces in real timeTHE DECODER · 1h ago
  • Google AI Search flags Facebook users as dangerTHE DECODER · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleEnterprise AI agents thrive with limits, not freedom