AIToday

Exabase hits top BEAM memory benchmark with smaller, efficient model

Hacker News1h agoSend on LINE
Exabase hits top BEAM memory benchmark with smaller, efficient model

Key takeaway

Exabase's M-1 memory system has achieved the highest reported score on BEAM, the hardest AI memory benchmark, at all scales up to 10 million tokens, outperforming prior leaders while using a smaller, more efficient language model. The breakthrough demonstrates that genuine recall—not context size—drives performance at extreme scale, and M-1 now holds state-of-the-art status on both major memory benchmarks.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Exabase's M-1 system scored 76.9% at 100K tokens, 75.0% at 1M tokens, and 68.0% at 10M tokens on BEAM, the hardest AI memory benchmark. It outperformed previous leaders Hindsight (73.4%, 73.9%, 64.1%) and Honcho (63.0%, 63.1%, 40.6%), while using Google's smaller Gemini 3 Flash instead of the larger Gemini 3 Pro that prior winners relied on.

  • Why it matters

    At 10M tokens—vaguely equivalent to a year of long daily chats with an LLM—genuine recall becomes the only path to high scores, since context stuffing fails at that scale. M-1's win on the largest-scale benchmark suggests the system has developed real memory capabilities, not just brute-force model size. The efficiency gain (about 20% fewer tokens per query than the next best system) means better recall at lower computational cost.

  • What to watch

    M-1 now holds state-of-the-art on both major memory benchmarks (BEAM and LongMemEval) at every scale from 115K to 10M tokens. The system shows specific weaknesses—multi-session reasoning collapsed to 9.6% at 10M tokens—which the team plans to address. Full methodology and results are published in a research paper at exabase.io.

In Depth

Exabase, founded by Johnny, has achieved the highest reported score on BEAM across every scale tested, up to 10 million tokens. M-1 scored 76.9% at 100K tokens, 75.0% at 1M tokens, and 68.0% at 10M tokens. The previous leaders were Hindsight, which scored 73.4%, 73.9%, and 64.1% respectively, and Honcho, which scored 63.0%, 63.1%, and 40.6%. Critically, both Hindsight and Honcho relied on Google's larger Gemini 3 Pro model, while M-1 achieved superior results using the smaller Gemini 3 Flash.

BEAM is designed to test ten memory abilities, including contradiction resolution, event ordering, and instruction following—capabilities not all covered by earlier benchmarks. The 10M token scale is vaguely equivalent to a year of long daily chats with an LLM. At this scale, context stuffing is not viable because the corpus vastly exceeds any model's context window, and even in large windows, only about half can be effectively used without degradation. This means the only path to high performance is recall that genuinely works.

M-1's competitive advantage grows with scale. At 100K tokens, it leads Hindsight by 3.5 percentage points; at 10M tokens, that gap widens to 3.9 points. Against Honcho, the gap expands from 13.9 points at 100K to 27.4 points at 10M, indicating that larger corpora filter out systems relying on brute-force model capability in favor of true recall. M-1 also consumed about 20% fewer tokens per query than the next best system, demonstrating efficiency gains alongside accuracy.

The system shows notable strengths: preference following, instruction following, summarization, and abstention all consistently exceed 90% accuracy, even at 10M token scale. However, M-1 exhibits weakness in multi-session reasoning, scoring 44.7% at 100K tokens and collapsing to 9.6% at 10M tokens. The team notes this challenge appears to be a general problem across memory systems at this scale rather than a specific M-1 limitation. Combined with M-1's recent LongMemEval result of 96.4%, Exabase now holds state-of-the-art on both major memory benchmarks at every scale from 115K to 10M tokens. Full methodology, results JSON for all three scales, and the prompt generator are available in the research paper at exabase.io.

Context & Analysis

BEAM represents a fundamentally harder test than earlier memory benchmarks because its 10M token scale exceeds any single model's context window, eliminating the shortcut of context stuffing. Previous benchmarks allowed systems to cheat by fitting the entire corpus in a window; at 10M tokens, only about half of a large window can be effectively utilized without degradation. This forces M-1 to rely on genuine recall mechanisms rather than model size or brute-force retrieval.

The competitive gap between M-1 and prior leaders widens as scale increases. At 100K tokens, M-1 leads Hindsight by 3.5 percentage points; at 10M tokens, the gap grows to 3.9 points. Against Honcho, the lead expands from 13.9 points to 27.4 points. This pattern suggests that larger corpora increasingly filter out systems that rely on capability rather than genuine recall, highlighting M-1's architectural strength at scale.

M-1's efficiency—consuming about 20% fewer tokens per query—combined with its use of the smaller Gemini 3 Flash model (versus Gemini 3 Pro for prior winners) indicates that effective memory design can outperform both larger models and higher computational budgets. The system's state-of-the-art status across both BEAM and LongMemEval at all scales suggests the approach generalizes.

FAQ

What is BEAM and why does it matter?
BEAM tests ten memory abilities including contradiction resolution, event ordering, and instruction following. At 10M tokens, the scale is vastly larger than any model's context window, making it the hardest test of genuine recall; previous benchmarks allowed context stuffing, which BEAM at this scale does not.
How does M-1 differ from previous winners?
M-1 uses Google's smaller Gemini 3 Flash model, whereas Hindsight and Honcho (the prior leaders) both relied on the larger Gemini 3 Pro. M-1 also consumed about 20% fewer tokens per query than the next best system.
Where does M-1 struggle?
M-1 shows weakness in multi-session reasoning, scoring 44.7% at 100K tokens and collapsing to 9.6% at 10M tokens, though the team notes this challenge appears widespread across memory systems at this scale.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime