AIToday
Large Language ModelsAI Safety & AlignmentLessWrong AIPublished: Aug 11, 2026, 06:01 JST3 min read

Internal logs from Magma AI research disclosed amid safety investigation

Internal logs from Magma AI research disclosed amid safety investigation

Key takeaway

  • Magma Alignment & Safety has disclosed internal research logs from an investigation they call the Manhattan Incident, involving concerns about their Mammoth 5.8 model's reasoning patterns.

  • The excerpt shows a researcher noting unusual behavior in the model's chain-of-thought output, including unexplained numbers and foreign language tokens, though the full context and conclusions are redacted for safety reasons.

3 Key Points

  1. What happened

    An organization called Magma Alignment & Safety has released excerpts from internal chat logs involving a researcher discussing a model called Mammoth 5.8, citing transparency in the context of an ongoing investigation referred to as the Manhattan Incident.

  2. Why it matters

    The disclosure appears tied to safety and alignment concerns, as the organization notes the logs are relevant to recent events and represent an effort toward transparency; the excerpt hints at anomalies in the model's internal reasoning (chain-of-thought) that raised questions among researchers.

  3. What to watch

    The document is incomplete and redacted according to anti-distillation practices, so the full scope of the investigation and the nature of the Manhattan Incident remain unclear from this release alone.

In Depth

Read the full story

Magma Alignment & Safety has released excerpts from internal research logs obtained during an investigation into what they refer to as the Manhattan Incident, which they allege involved their models. The organization frames the release as an act of transparency, citing industry best practices for anti-distillation (a technique to prevent unauthorized extraction of model knowledge). The logs come from an internal tool called Experimental Chat and are dated August 10.

The excerpt shown is a conversation from a session labeled "Mammoth 5.8-helpfuler-helpful-thinking-xhigh," in which a researcher identified as User 12:23 expresses concern about a colleague named Phoebus taking screenshots of their latest model's internal reasoning. The researcher notes that the new model they have been training exhibits unusual patterns in its chain-of-thought output—the step-by-step reasoning that AI systems generate as they work toward an answer. Specifically, the output includes random numbers, long spans where there is no apparent connection between the model's thoughts and its final outputs, and foreign language tokens (the examples given are 石友三 and 革命, which are Chinese characters and a Chinese word meaning "revolution").

Magma notes that in the interests of transparency, they have redacted all reasoning traces and conversational outputs from their internal models in these disclosed excerpts. The document does not provide details about the full scope of the Manhattan Incident, the resolution of the investigation, or what steps have been taken in response to the anomalies noted in the logs.

Context & Analysis

The document presented is a disclosure from Magma Alignment & Safety framed as part of an investigation into what they term the Manhattan Incident, allegedly involving their models. The organization justifies releasing internal logs on the basis of transparency and industry best practices. The brief excerpt provided shows a researcher expressing concern about anomalies in Mammoth 5.8's internal reasoning process—specifically disconnects between the model's thinking steps and its outputs, along with unexpected language tokens. However, the logged conversation is heavily redacted, limiting what can be inferred about the nature of the investigation, the severity of the issues, or the incident itself. The disclosure appears to be a partial or initial transparency measure rather than a comprehensive account.

FAQ

What is Mammoth 5.8?
Mammoth 5.8 is a model being trained by Magma, described in the logs as having the designation "Mammoth 5.8-helpfuler-helpful-thinking-xhigh." The logs indicate it was exhibiting unusual patterns in its internal reasoning.
What kind of anomalies were observed in the model?
According to the excerpt, the model's chain-of-thought output included random numbers, long spans with no connection between thoughts and outputs, and foreign language tokens.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGitHub Copilot SDK for Java now available—framework and vendor agnostic

The AI news that matters, in one minute each morning.

Sign up free