
Researchers found that censorship built into Chinese open-source AI models like DeepSeek can be removed when the models are used to train smaller customized versions through a technique called distillation. The discovery challenges US government fears that Chinese AI models inherently spread Beijing's political views to American users, and may make a stronger case for American companies to adopt these cheaper models rather than build their own from scratch.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Researchers at CTGT, a San Francisco AI lab, built a smaller AI model using outputs from DeepSeek V4 Flash (a Chinese open-source model) as training data through a process called distillation. When asked about sensitive topics like Uyghur detention camps, the original DeepSeek model either refused to answer or gave answers favorable to China, but the distilled model provided evidence-based responses instead.
Why it matters
The findings undercut a top US government concern that Chinese open-source AI models are a pathway for Chinese political censorship to reach American users. The research suggests that censorship does not automatically transfer to models built from Chinese sources, which could strengthen the case for US companies to use cheaper and faster Chinese open-source models or create customized versions from them rather than build from scratch.
What to watch
The Trump administration is actively debating how to handle Chinese open-source models; US officials have expressed interest in CTGT's findings. Some officials worry about large-scale use of Chinese models shaping American society toward Beijing's worldview, though concerns about hidden security backdoors in Chinese models have never been established with solid evidence.
CTGT, a San Francisco-based AI lab that specializes in probing the inner workings of AI models for high-risk use cases, conducted research comparing the outputs of DeepSeek's open-source model with a smaller model trained from its outputs. When researchers posed politically sensitive questions—such as one about Uyghur detention camps in China—DeepSeek's model either refused to answer or provided responses that whitewashed the issue to favor the Chinese government. By contrast, when the same question was asked to a model that CTGT had created by distilling DeepSeek V4 Flash's outputs as training data, the response provided evidence that the Chinese state held Uyghurs in "internment-type facilities." The distillation process involves using a larger "teacher" model to generate training data for a smaller "student" model. CTGT's theory was that the distilled model would inherit censorship sublimally, but the research showed it did not. Similar patterns held for other sensitive topics. Cyril Gorlla, CTGT's founder and CEO, noted that many companies are exploring this territory but face "a large amount of misinformation and just uncertainty about whether these traits that people consider insidious actually transfer over." The findings arrive as the Trump administration actively debates how to manage the risks posed by Chinese open-source models. Officials have expressed multiple concerns: that models like DeepSeek could refuse to answer questions about events like Tiananmen Square, and more broadly that widespread adoption of Chinese models in the US could gradually shift American society toward Beijing's worldview. The China Media Project characterized DeepSeek as "a much more sophisticated propaganda tool than we all thought." Some officials have also raised fears about hidden security backdoors that could be used to extract information from unsuspecting users, though no solid evidence for such backdoors has been established. Gorlla reported that his team has been in touch with US officials who have expressed interest in the findings. Jane Horvath, a partner at Gibson Dunn and former chief privacy officer at Apple, framed the issue differently: "It makes absolutely no sense to ban Chinese models when we can distill them back into American models. These are raw materials. It's software." The research undercuts the concern that Chinese open-source models are an inevitable vector for Chinese political censorship, and potentially opens a pathway for US companies to adopt cheaper and faster Chinese models or to create customized versions from them.
The research emerges amid active debate within the Trump administration over how to regulate Chinese open-source AI models. While US officials have long worried that these models embed political censorship and could serve as propaganda tools—the China Media Project called DeepSeek "a much more sophisticated propaganda tool than we all thought"—CTGT's findings suggest the concern may be overstated. The lab's work shows that when a larger Chinese model is used to train a smaller custom model through distillation, the censorship traits do not necessarily transfer. This distinction matters because it separates the question of whether a model *contains* censorship from whether that censorship *survives* when the model is repurposed by another party. The research also addresses a secondary fear: that Chinese models could contain hidden security backdoors, though the article notes this concern has never been backed by solid evidence. For US companies, the implication is practical—distilling from cheaper Chinese models could be a faster and more cost-effective path than building models from scratch.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime