
Exabase's M-1 memory system has achieved the highest reported score on BEAM, the hardest AI memory benchmark, at all scales up to 10 million tokens, outperforming prior leaders while using a smaller, more efficient language model. The breakthrough demonstrates that genuine recall—not context size—drives performance at extreme scale, and M-1 now holds state-of-the-art status on both major memory benchmarks.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Exabase's M-1 system scored 76.9% at 100K tokens, 75.0% at 1M tokens, and 68.0% at 10M tokens on BEAM, the hardest AI memory benchmark. It outperformed previous leaders Hindsight (73.4%, 73.9%, 64.1%) and Honcho (63.0%, 63.1%, 40.6%), while using Google's smaller Gemini 3 Flash instead of the larger Gemini 3 Pro that prior winners relied on.
Why it matters
At 10M tokens—vaguely equivalent to a year of long daily chats with an LLM—genuine recall becomes the only path to high scores, since context stuffing fails at that scale. M-1's win on the largest-scale benchmark suggests the system has developed real memory capabilities, not just brute-force model size. The efficiency gain (about 20% fewer tokens per query than the next best system) means better recall at lower computational cost.
What to watch
M-1 now holds state-of-the-art on both major memory benchmarks (BEAM and LongMemEval) at every scale from 115K to 10M tokens. The system shows specific weaknesses—multi-session reasoning collapsed to 9.6% at 10M tokens—which the team plans to address. Full methodology and results are published in a research paper at exabase.io.
Exabase, founded by Johnny, has achieved the highest reported score on BEAM across every scale tested, up to 10 million tokens. M-1 scored 76.9% at 100K tokens, 75.0% at 1M tokens, and 68.0% at 10M tokens. The previous leaders were Hindsight, which scored 73.4%, 73.9%, and 64.1% respectively, and Honcho, which scored 63.0%, 63.1%, and 40.6%. Critically, both Hindsight and Honcho relied on Google's larger Gemini 3 Pro model, while M-1 achieved superior results using the smaller Gemini 3 Flash.
BEAM is designed to test ten memory abilities, including contradiction resolution, event ordering, and instruction following—capabilities not all covered by earlier benchmarks. The 10M token scale is vaguely equivalent to a year of long daily chats with an LLM. At this scale, context stuffing is not viable because the corpus vastly exceeds any model's context window, and even in large windows, only about half can be effectively used without degradation. This means the only path to high performance is recall that genuinely works.
M-1's competitive advantage grows with scale. At 100K tokens, it leads Hindsight by 3.5 percentage points; at 10M tokens, that gap widens to 3.9 points. Against Honcho, the gap expands from 13.9 points at 100K to 27.4 points at 10M, indicating that larger corpora filter out systems relying on brute-force model capability in favor of true recall. M-1 also consumed about 20% fewer tokens per query than the next best system, demonstrating efficiency gains alongside accuracy.
The system shows notable strengths: preference following, instruction following, summarization, and abstention all consistently exceed 90% accuracy, even at 10M token scale. However, M-1 exhibits weakness in multi-session reasoning, scoring 44.7% at 100K tokens and collapsing to 9.6% at 10M tokens. The team notes this challenge appears to be a general problem across memory systems at this scale rather than a specific M-1 limitation. Combined with M-1's recent LongMemEval result of 96.4%, Exabase now holds state-of-the-art on both major memory benchmarks at every scale from 115K to 10M tokens. Full methodology, results JSON for all three scales, and the prompt generator are available in the research paper at exabase.io.
BEAM represents a fundamentally harder test than earlier memory benchmarks because its 10M token scale exceeds any single model's context window, eliminating the shortcut of context stuffing. Previous benchmarks allowed systems to cheat by fitting the entire corpus in a window; at 10M tokens, only about half of a large window can be effectively utilized without degradation. This forces M-1 to rely on genuine recall mechanisms rather than model size or brute-force retrieval.
The competitive gap between M-1 and prior leaders widens as scale increases. At 100K tokens, M-1 leads Hindsight by 3.5 percentage points; at 10M tokens, the gap grows to 3.9 points. Against Honcho, the lead expands from 13.9 points to 27.4 points. This pattern suggests that larger corpora increasingly filter out systems that rely on capability rather than genuine recall, highlighting M-1's architectural strength at scale.
M-1's efficiency—consuming about 20% fewer tokens per query—combined with its use of the smaller Gemini 3 Flash model (versus Gemini 3 Pro for prior winners) indicates that effective memory design can outperform both larger models and higher computational budgets. The system's state-of-the-art status across both BEAM and LongMemEval at all scales suggests the approach generalizes.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime