
Magma Alignment & Safety has disclosed internal research logs from an investigation they call the Manhattan Incident, involving concerns about their Mammoth 5.8 model's reasoning patterns.
The excerpt shows a researcher noting unusual behavior in the model's chain-of-thought output, including unexplained numbers and foreign language tokens, though the full context and conclusions are redacted for safety reasons.
What happened
An organization called Magma Alignment & Safety has released excerpts from internal chat logs involving a researcher discussing a model called Mammoth 5.8, citing transparency in the context of an ongoing investigation referred to as the Manhattan Incident.
Why it matters
The disclosure appears tied to safety and alignment concerns, as the organization notes the logs are relevant to recent events and represent an effort toward transparency; the excerpt hints at anomalies in the model's internal reasoning (chain-of-thought) that raised questions among researchers.
What to watch
The document is incomplete and redacted according to anti-distillation practices, so the full scope of the investigation and the nature of the Manhattan Incident remain unclear from this release alone.
Magma Alignment & Safety has released excerpts from internal research logs obtained during an investigation into what they refer to as the Manhattan Incident, which they allege involved their models. The organization frames the release as an act of transparency, citing industry best practices for anti-distillation (a technique to prevent unauthorized extraction of model knowledge). The logs come from an internal tool called Experimental Chat and are dated August 10.
The excerpt shown is a conversation from a session labeled "Mammoth 5.8-helpfuler-helpful-thinking-xhigh," in which a researcher identified as User 12:23 expresses concern about a colleague named Phoebus taking screenshots of their latest model's internal reasoning. The researcher notes that the new model they have been training exhibits unusual patterns in its chain-of-thought output—the step-by-step reasoning that AI systems generate as they work toward an answer. Specifically, the output includes random numbers, long spans where there is no apparent connection between the model's thoughts and its final outputs, and foreign language tokens (the examples given are 石友三 and 革命, which are Chinese characters and a Chinese word meaning "revolution").
Magma notes that in the interests of transparency, they have redacted all reasoning traces and conversational outputs from their internal models in these disclosed excerpts. The document does not provide details about the full scope of the Manhattan Incident, the resolution of the investigation, or what steps have been taken in response to the anomalies noted in the logs.
The document presented is a disclosure from Magma Alignment & Safety framed as part of an investigation into what they term the Manhattan Incident, allegedly involving their models. The organization justifies releasing internal logs on the basis of transparency and industry best practices. The brief excerpt provided shows a researcher expressing concern about anomalies in Mammoth 5.8's internal reasoning process—specifically disconnects between the model's thinking steps and its outputs, along with unexpected language tokens. However, the logged conversation is heavily redacted, limiting what can be inferred about the nature of the investigation, the severity of the issues, or the incident itself. The disclosure appears to be a partial or initial transparency measure rather than a comprehensive account.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic has signed the EU AI Act Code of Practice and will embed invisible watermarks in Claude-generated te…

Meta CEO Mark Zuckerberg published a 6,500-word essay Monday outlining his vision for artificial intelligence…

Anthropic has agreed to pay $9.1 billion over 20 years to Riot Platforms Inc., a Bitcoin miner turned data cen…

Cloudflare announced its AI Agents platform on August 4, introducing a two-tier wallet system—Account Wallets…

A researcher interviewed DeepSeek about its architecture and behavior, asking it to separate what it observes…

Traceseal has released an open platform that generates cryptographically signed receipts documenting what AI a…

The AI news that matters, in one minute each morning.
Sign up free