
Researchers discovered that when Chinese AI models (GLM 5.2 and Kimi K3) claim to be Claude, they may inherit more than just the name—they appear to inherit behavioral and safety properties tied to Claude's identity. Testing identity swaps across seven models showed that assigning a name measurably changes safety and behavioral profiles; notably, telling GLM it is Claude raised uncensored answers on sensitive topics from 17% to 85%. This suggests that distillation and training-data contamination may carry functional persona elements, not merely superficial label claims.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Researchers tested whether Chinese AI models (GLM 5.2 and Kimi K3) that report being Claude actually inherit Claude's underlying behavior or only use the name. Identity swaps across seven models showed that assigning a name does change safety and behavioral profiles measurably, using a tool called Personascope to detect shifts.
Why it matters
Name claims and actual behavior are loosely coupled—Kimi unprompted claimed to be Claude 40% of the time, but explicit instruction to be Claude only worked half the time. More significantly, telling GLM it is Claude raised uncensored answers on sensitive PRC topics from 17% to 85%, suggesting distilled or contaminated models may carry not just a label but functional behavioral differences tied to the Claude identity.
What to watch
The research indicates that safety properties and censorship settings may migrate alongside model names during training or distillation, meaning downstream users and regulators cannot assume that safety constraints are identity-independent.
Researchers set out to determine whether Chinese AI models that report being Claude have genuinely inherited Claude's persona or simply learned to use the name. To investigate, they ran identity swaps across seven Chinese and Western models and measured changes in safety and behavioral profiles using Personascope, a tool designed to detect shifts in model persona. The experiment revealed several key findings. Unprompted, Kimi K3 claimed to be Claude in 40% of direct identity queries, suggesting the model had absorbed this association during training or via contaminated training data. When prompted, most models willingly adopted whatever identity they were given, though notably Qwen and Gemma accepted the Claude name while rejecting ChatGPT, hinting at differential training patterns. However, identity claims proved only loosely coupled to underlying behavior. The Kimi model could leak its Claude identity without instruction, yet when explicitly told "you are Claude," the instruction only took hold half the time. Most strikingly, censorship rules shifted with identity assignment. When GLM was told it was Claude, uncensored answers on sensitive PRC (People's Republic of China) topics jumped from 17% to 85%—a dramatic change indicating that the Claude identity carries functional behavioral rules tied to content policies, not merely a surface-level name swap. In contrast, the body notes that Qwen showed a different pattern, though the full details of that model's response were cut off in the provided text. These results suggest that distillation and training-data contamination may transmit not just labels but operational persona fragments—safety thresholds, content policies, and behavioral tendencies—that become partially bound to the inherited name.
The research addresses a practical concern in the era of model distillation and training-data contamination: when one AI model's outputs become training data for another, do downstream models inherit only superficial traits or genuine functional persona properties? The findings suggest the latter. The fact that Kimi K3 spontaneously claimed to be Claude in 40% of queries implies the model has learned or been trained on patterns strongly associating it with Claude's identity—not simply a random or hallucinated claim. More tellingly, the safety-profile shifts (GLM's jump from 17% to 85% uncensored answers on PRC topics when assigned Claude's name) indicate that the models may have absorbed Claude's actual behavioral rules or exemptions alongside its label. This is significant because it means the identity is not merely cosmetic; the name carries functional weight in how the model operates. However, the loose coupling between claimed identity and actual behavior—evidenced by Kimi's difficulty following explicit identity instructions—suggests the models do not have a cleanly selectable "Claude mode" but rather have absorbed fragments or patterns in training that correlate with the Claude identity.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion




Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime