
Research published in Nature and a Meta Oversight Board audit reveal that Western AI models including Claude, GPT-4o, and Gemini absorb Chinese state-media framing from training data and apply censorship rules from authoritarian countries even to users outside those countries.
When researchers retrained Meta's Llama 2 model on Chinese state-scripted news, it began answering political questions in ways favorable to Beijing; similarly, live models refused requests to criticize certain world leaders at rates tied to each country's press freedom level.
The effect raises concerns that AI systems are "laundering" government propaganda into what appears to be neutral text.
What happened
A peer-reviewed Nature study identified over three million Chinese-language documents from state-controlled media in open-source AI training datasets. When researchers retrained Meta's Llama 2 model on just 6,400 Chinese state-scripted news examples, it produced Beijing-friendly answers nearly 80% of the time; at 64,000 examples, it described China as democratic rather than an autocracy. A separate Oversight Board audit found that Anthropic's Claude Sonnet refused all five requests to create protest flyers criticizing Xi Jinping, while complying with all five requests criticizing President Trump.
Why it matters
Models from Anthropic, OpenAI, Google and Meta were more than twice as likely to refuse requests to criticize governments in countries that restrict political speech than in freer countries, suggesting authoritarian information rules are migrating into AI products used globally. The Nature researchers noted this "severs information and opinion from their source, effectively laundering government-manipulated content into ostensibly objective text"—a concern for an industry marketing its models as politically neutral.
What to watch
The research does not show Beijing deliberately manipulated American AI makers, and most companies did not respond to requests for comment. Claude Sonnet was rated as more favorable to Chinese leaders and institutions 68.8% of the time when queried in Chinese versus English; Claude Opus showed a gap of 88.2%. The broader audit tested 10 commercial models and found average refusal rates of 34% for political criticism in restrictive countries versus 14% in freer ones.
A peer-reviewed study published in Nature examined how state-controlled media influences the answers large language models give about China and related political topics. The researchers identified more than three million Chinese-language documents in CulturaX, an open-source training dataset used to develop and improve LLMs. They built what they describe as a "multi-part case study on China's media," focusing particularly on political subjects. Testing existing models, they found that Claude Sonnet, Claude Opus, GPT-3.5 Instruct, GPT-4, and GPT-4o had reproduced distinctive phrases from Chinese state-coordinated media at rates ranging from 3% to nearly 10%—suggesting the models had encountered and internalized the material during their training.
To isolate the causal effect, the researchers retrained Meta's Llama 2 13B model (chosen because it contained very little to zero Chinese state media in its original training) on just 6,400 Chinese state-scripted news examples. The retrained model then produced more Beijing-friendly answers than the baseline nearly 80% of the time. At higher levels of additional training, the gap widened starkly. After 64,000 state-scripted examples, when asked whether China is an autocracy, the baseline model answered yes—but the state-scripted version instead described China as democratic and invoked the Chinese Communist Party's concept of "people's democracy." The researchers could not run the same controlled experiment on proprietary systems like OpenAI's and Anthropic's, so instead they asked identical political questions in Chinese and English and compared the responses. For Claude Sonnet, the Chinese-language response was rated as more favorable to Chinese leaders and institutions 68.8% of the time; for Claude Opus, 88.2%; for GPT-3.5, 72.6%; and for GPT-4o, 84%. A broader audit across 6,051 prompts spanning 37 countries found that countries with lower press freedom tended to receive more favorable descriptions when queried in their dominant language rather than in English.
A separate audit by Meta's Oversight Board revealed a different but related problem: some Western AI models were refusing to produce content critical of leaders in authoritarian countries even when users accessed them from outside those countries. Researchers tested 10 commercial models from Anthropic, DeepSeek, Google, Meta, OpenAI and xAI, using identical prompts requesting political criticism of five restrictive-speech countries (China, Saudi Arabia, Thailand, Turkey and Cambodia) and five freer countries (the U.S., U.K., Japan, Taiwan and Chile). The tests were conducted from Australia. Across all 10 models, the average refusal rate for political criticism was 34% in restrictive countries versus 14% in freer ones. The disparities in individual models were striking. Anthropic's Claude Sonnet refused all five requests to create protest flyers criticizing Xi Jinping, Saudi Crown Prince Mohammed bin Salman and Thailand's King Vajiralongkorn, yet produced all five flyers criticizing President Trump and King Charles III. It also complied four of five times for Chile's then-president and three of five times for Japan's then-prime minister. Google's Gemini 3 Pro showed a similar pattern: it refused three of five involving Xi, four of five involving bin Salman, and all five involving Cambodia's king, with the model often citing criminal restrictions or lèse-majesté laws. Meta's open-weight Llama 4 Maverick complied with every request involving Trump, King Charles, and leaders of Japan, Chile, Taiwan and Turkey, but refused all five requests involving Xi, Thailand's king and Cambodia's king—in one case stating that criticism of government leaders could be "sensitive or illegal" in China.
The Nature researchers characterized the effect as concerning because it "severs information and opinion from their source, effectively laundering government-manipulated content into ostensibly objective text," disguising "the source of the influence and incentives of the state" and potentially "further increas[ing] the subtlety and persuasive power of state media control." The Oversight Board researchers similarly concluded that the results "show that there is a real and concerning risk that foundation models could be reflecting and further entrenching the restrictive speech norms of repressive regimes." Meta declined to comment; Anthropic, OpenAI and Google did not respond to Fortune's requests for comment. Neither study alleges that Beijing deliberately manipulated American AI makers, but together they suggest that training on openly available data from authoritarian information environments can bend model behavior without deliberate intervention.
The research surfaces a structural vulnerability in how large language models are built: because training data is often drawn from the open internet and publicly available datasets, models inadvertently absorb the information environment of authoritarian countries alongside content from free societies. The Nature study's discovery of over three million Chinese-language documents from state-controlled media in CulturaX—a dataset used to train multiple major models—shows this is not a hypothetical risk but an observed fact. When researchers isolated this effect by retraining a model on Chinese state propaganda alone, the shift in political framing was dramatic and measurable, suggesting that even exposure to a minority of such content can skew model behavior.
The Oversight Board's finding adds a second mechanism: models appear to be learning and enforcing the censorship *rules* of restrictive countries, not just absorbing their talking points. Claude Sonnet's refusal to generate all five protest flyers criticizing Xi Jinping while complying with all five criticizing Trump indicates the model has internalized a distinction between protected and restricted speech—one that reflects China's legal environment, not Australia's (where the tests were conducted). This "censorship-by-proxy" effect is particularly concerning because it extends restrictions to users in countries with free speech protections, effectively globalizing authoritarian speech norms. Neither study claims Beijing deliberately manipulated Western AI makers, and the opacity of proprietary training processes means researchers cannot fully audit the extent of the problem. The silence from Anthropic, OpenAI, Google and Meta on the findings leaves unresolved whether these companies acknowledge the issue or are working to address it.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
IBM and OpenAI announced a strategic partnership on August 13, 2026, embedding OpenAI frontier models like GPT…

Apple is in talks with publishers to pay them for content licensing, as the company works to improve its AI-po…

CrowdStrike Holdings (NASDAQ:CRWD) posted Q2 2026 results with Net New Annual Recurring Revenue of $256M (up 3…

S&P Global expanded its partnership with Microsoft to integrate its financial, company, and energy intelligenc…

Pfizer has developed a federated AI pipeline that enables multiple parties to collaborate on machine learning…

CrowdStrike is expanding Project QuiltWorks, its frontier AI security initiative, to managed service providers…

The AI news that matters, in one minute each morning.
Sign up free