
What happened
Aleph Alpha tested models from Alibaba (Qwen), DeepSeek, and Moonshot AI (Kimi) on 967 taboo topics including Tiananmen, Taiwan, and Xinjiang. Its scoring system rated only 17 to 41 percent of responses as balanced.
Why it matters
The findings suggest Chinese models often repeat state doctrine, deflect, or refuse to answer on sensitive topics, consistent with China's AI rules requiring "socialist core values" in public-facing models.
What to watch
Aleph Alpha sells "sovereign AI," so it has a commercial interest in distinguishing its models. Watch whether independent tests confirm the pattern Western comparison models showed: balanced answers 70 percent and 92 percent of the time.
WHO IT HITSGovernment and enterprise buyers evaluating AI vendors now have benchmark data suggesting Chinese models may carry political bias, while AI developers using distilled Chinese training data — as Nvidia's Nemotron Cascade 2 did with roughly 3,500 of 9.3 million examples — may inherit those patterns.
Summaries like this, in your inbox every morning.
Aleph Alpha's study adds quantitative weight to what had largely been anecdotal reports and earlier audits. The company markets itself alongside Cohere as a provider of "sovereign AI" for governments, so the benchmark doubles as a competitive argument against Chinese models. That commercial context does not invalidate the finding, but it is worth noting when weighing the results.
The study also highlights a spillover effect: models trained on distilled Chinese data can carry those values into otherwise unrelated outputs. Nvidia's Nemotron Cascade 2, for example, showed party-line patterns in 17 percent of responses — attributed by Aleph Alpha to a small fraction of training examples generated by DeepSeek and Qwen. A separate study by the Central European Institute of Asian Studies found similar spillover when terms like human rights or surveillance appeared.
For European buyers, the practical stakes are a choice between two foreign value systems unless European models can compete on performance. For U.S. policymakers, the study is a reminder that ideological shaping of AI is not limited to China — Elon Musk has repeatedly had Grok modified to produce right-leaning responses. The test of this benchmark's influence will be whether independent evaluators replicate the 17 to 41 percent range, and whether government procurement decisions treat it as a meaningful signal.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
A Fortune commentary argues the largely unregulated AI industry echoes 1930s finance and should get a modern v…

Manulife Financial Corporation launched a first-of-its-kind CoverMe travel insurance plugin in ChatGPT in late…

A LessWrong post argues that if a group of humans wanted to kill all humans, it would sound natural for an ali…

Google said free Gemini app users will be restricted to the "Flash-Lite" model starting October 9; "Flash" nee…

Indie developer Robert Varadan argues that AI models like Opus 5.5 and GPT-6 Astra can clone game demos from a…

SuperLocalMemory 4.0 appeared on Hacker News under the title "Governed Memory Operating System for AI Agents"…
