AIToday
Large Language Modelsr/MachineLearningPublished: Mar 28, 2026, 04:00 JST1 min read

Major audit reveals LoCoMo long-context benchmark has significant flaws with 6.4% corrupted answers and an LLM judge that accepts up to 63% of wrong responses.

Major audit reveals LoCoMo long-context benchmark has significant flaws with 6.4% corrupted answers and an LLM judge that accepts up to 63% of wrong responses.

3 Key Points

  1. Researchers identified 99 score-corrupting errors across 1,540 questions (6.4%) in LoCoMo, a widely-cited ACL 2024 long-term memory benchmark still receiving submissions as of March 2026

  2. Error types include hallucinated facts in the answer key, incorrect temporal reasoning, and speaker attribution mistakes, such as specifying 'Ferrari 488 GTB' when only generic descriptions exist in source materials

  3. The LLM judge evaluating answers accepts up to 63% of intentionally wrong responses, raising concerns about evaluation reliability

  4. Alternative benchmark LongMemEval-S is limited as a true memory test since each question's corpus fits entirely within modern context windows

Ask the AI about this article →

r/MachineLearningRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Walmart settles opioid claims for $50MTop Companies AI · 57m ago
  • Tim Cook's legacy hinges on Apple's AI betTop Companies AI · 57m ago
  • CrowdStrike Falcon Guardian Targets AI SecurityTop Companies AI · 57m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNew proxy tool enables developers to access OpenCode models through popular third-party APIs like OpenAI, Anthropic, and Gemini