
MiniMax M3 Medium scored 73.17% F1 on DeepSearchQA, matching GPT-5 High Reasoning.
The model used You.com search and content tools with an optimized research skill.
It trails top agents like Gemini Deep Research but offers transparent, reproducible results at $0.532 per task.
What happened
MiniMax M3 Medium, paired with You.com research tools and optimized via an auto-research loop, achieved 73.17% adjusted F1 on the DeepSearchQA benchmark. The full run completed on 2026-08-21 across 900 tasks and 2,700 trials, with results published on GitHub and Hugging Face.
Why it matters
The score is competitive with GPT-5 High Reasoning's 73.24% F1 in the DeepSearchQA paper's Table 4. However, MiniMax M3 Medium trails the leading agents—Gemini Deep Research Agent (81.90% F1) and GPT-5 Pro High Reasoning (78.98% F1)—and shows a higher fully-incorrect rate (19.01%) than those leaders.
What to watch
The evaluation cost $478.41 total ($328.14 model, $150.27 You.com API), averaging $0.532 per task. Median latency was 54.3 seconds; 95th percentile was 238.8 seconds. All prompts, graded results, and trajectories are publicly available for inspection and reproduction.
Ask the AI about this article →
MiniMax M3 Medium's 73.17% F1 result represents a meaningful checkpoint in the evolution of open and commercial reasoning models. The score places it at parity with GPT-5 High Reasoning, which the DeepSearchQA paper reports at 73.24% F1. However, the body explicitly acknowledges that MiniMax M3 Medium lags behind the top-tier agents: Gemini Deep Research Agent (81.90% F1) and GPT-5 Pro High Reasoning (78.98% F1). The fully-incorrect rate of 19.01% also suggests room for improvement compared to those leading systems.
The evaluation methodology adds transparency to the claim. Rather than a paper assertion, the results are backed by 2,700 trials across 900 tasks, with all intermediate artifacts (trajectories, grades, prompts) published for inspection and reproduction. This open artifact approach—coupled with the You.com MCP tools and an optimized research skill tuned to the model and tool surface—establishes a clear foundation for further iteration. The total cost of $478.41 and median latency of 54.3 seconds provide operational context: the model operates within a reasonable resource envelope for research tasks, though the 95th percentile latency of 238.8 seconds indicates variability in harder queries.
The fact that the evaluation is described as competitive with GPT-5 High Reasoning, not with the top agents, sets realistic expectations. The body frames this as a step forward for the MiniMax M3 Medium line, not a breakthrough that overtakes the field. For practitioners, the key signals are the public artifacts (enabling validation and fine-tuning) and the documented cost and latency trade-offs.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Visko raised $10 million in pre-seed funding from Llama Ventures and opened public access to its first foundat…
AI company Runway has unveiled Solaris, the first model in a new category it calls "Interface World Models." I…

Google's AI search gave advice to call emergency services for users alone with an African, Indian, or Pakistan…

John Deere introduced JD, a conversational AI tool that lets farmers ask open-ended questions about their hist…

Nvidia CEO Jensen Huang said on Fox Business that AI is creating 'hundreds of thousands' of jobs, including in…

Israeli startup DataAgent Ltd