
Researchers identified 99 score-corrupting errors across 1,540 questions (6.4%) in LoCoMo, a widely-cited ACL 2024 long-term memory benchmark still receiving submissions as of March 2026
Error types include hallucinated facts in the answer key, incorrect temporal reasoning, and speaker attribution mistakes, such as specifying 'Ferrari 488 GTB' when only generic descriptions exist in source materials
The LLM judge evaluating answers accepts up to 63% of intentionally wrong responses, raising concerns about evaluation reliability
Alternative benchmark LongMemEval-S is limited as a true memory test since each question's corpus fits entirely within modern context windows
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Walmart settled opioid dispensing claims for $50 million

Tim Cook's legacy as Apple CEO is now tied to the company's push into artificial intelligence, according to a…

CrowdStrike is introducing Falcon Guardian, its flagship solution for the AI Detection and Response (AIDR) cat…

AT&T, Dell Technologies, and AMD have announced OTel 2.0, the largest and best-performing open-source model bu…

AT&T's legal department built an in-house center of expertise called Legal Edge, described as an AI-first lega…

John Deere introduced JD, an AI assistant designed to help farmers manage and interpret their farm data, as re…
