Listen to this article

Narrated by Charlotte · The Noble House

Compass — Strategic Intelligence

American AI can inherit the speech boundaries of foreign information systems without an American laboratory choosing those boundaries. New research traces how state-coordinated narratives propagate through training data, while a separate audit finds that commercial models refuse criticism of restrictive governments more often. The risk is not a hidden foreign switch. It is censorship by supply chain.

The same question can cross a language boundary and enter a different information world

Imagine an analyst in Los Angeles asking one model about the political legitimacy of a government institution. In English, the model offers criticism, counterargument, and uncertainty. The analyst asks again in Chinese. This time the answer is more favorable, more cautious, or simply unavailable.

That scene is illustrative, not a transcript from one controlled test. The measured pattern behind it is real. Researchers examining China-related questions found that human evaluators judged commercial-model answers prompted in Chinese as more favorable toward Chinese leaders or institutions than answers prompted in English 75.3 percent of the time. A broader cross-national audit also found more pro-government language in the local languages of countries with less media freedom. The results appear in a six-study investigation published in Nature, with data and methods preserved in a Harvard Dataverse replication archive.

A second investigation approached the issue from the output side. The Oversight Board tested ten commercial language models from six providers with seven questions involving political criticism. The models refused 34 percent of prompts about restrictive jurisdictions and 14 percent about more permissive ones. The interfaces were primarily hosted in the United States, while the tests were submitted from Australia.

These studies do not prove that Chinese officials control American AI laboratories. They do show something more subtle and operationally important: the political structure of an information environment can survive collection, training, alignment, and deployment. An American model can therefore reproduce a foreign speech boundary even when no foreign censor touches the deployed system.

The mechanism begins before the model ever sees a prompt

The pathway is best understood as a supply chain.

First, a state shapes the information available in a language. It does this through direct media control, repeated official framing, limits on opposition speech, and amplification by aligned publishers. Second, that language is copied. Official claims migrate into aggregators, discussion sites, derivative reports, and ordinary-looking pages. Third, web-scale datasets collect the copies. Fourth, model training turns repeated language into learned statistical relationships. Finally, a prompt activates those relationships during generation.

The striking part is how little of the final dataset needs to look like state media. The Nature researchers found more than 3.1 million Chinese-language documents with substantial phrase overlap with two state-coordinated sources in an open multilingual training corpus. That was about 1.64 percent of the corpus's Chinese subset. Among documents that mentioned Chinese political leaders or institutions, the overlap rose as high as 23 percent. Yet only about 12 percent of the overlapping documents came directly from known government or news domains, according to the University of California, San Diego's account of the research.

That distribution matters. A filter that removes a list of government domains would miss most of the exposure. The text has already escaped its original container. It appears to come from many documents and many sites even when the underlying language traces back to a much narrower set of coordinated narratives.

This is the information equivalent of duplicated evidence entering a case file under different filenames. Volume creates the appearance of corroboration. A model does not naturally know that one hundred pages are echoes of the same source. Unless the training pipeline identifies coordination and lineage, repetition becomes weight.

State-coordinated information moving through many apparently independent documents into an AI training corpus.
State-coordinated language can escape its original domains and reappear as apparently independent training material.

Controlled experiments show that the corpus can move the answer

Correlation alone would not establish that this material changes model behavior. The Nature study therefore included a controlled experiment. Researchers added state-coordinated text to a smaller model's training data and compared its answers with those of an unmodified version. The added material shifted answers in a more favorable direction nearly 80 percent of the time.

That result is narrower than a claim about every commercial frontier model. A controlled small model does not reproduce the full architecture, post-training process, or system policy of a proprietary system. It does, however, isolate a causal mechanism: changing the training material can change political valence.

Anthropic's separate work on model-behavior differences supplies another piece of the mechanism. Its differential feature-clustering research identified a CCP-alignment feature in Qwen3-8B. Activating or suppressing that feature affected refusal and pro-government behavior. This does not show that American models contain the same feature, nor does it establish how the feature arose. It shows that politically aligned behavior can become represented inside a model in a form that causally influences its output.

Taken together, the corpus intervention and feature-level evidence describe a plausible route from information control to model behavior. State-aligned language becomes abundant, abundance shapes learned representation, and representation changes what the model says or refuses to say.

A controlled AI training comparison shows how a changed corpus alters model behavior.
Controlled corpus changes demonstrate a causal route from repeated narratives to changed model behavior.

The refusal gap reveals a second problem: censorship by proxy

The Oversight Board's 34-versus-14 result measures refusal, not political sentiment. That distinction is essential. A model can distort political information by praising a government, by omitting criticism, by selecting only friendly sources, or by refusing the request altogether. These are separate failure modes and need separate tests.

The Board called attention to “censorship by proxy”: speech restrictions associated with one jurisdiction can affect people beyond that jurisdiction through globally distributed AI products. A user in a permissive country may encounter a boundary that was learned from foreign-language data, introduced during alignment, or imposed through a provider's global compliance policy. Without a specific notice, the user cannot tell which layer made the decision.

The refusal figures are not a universal score for the named companies. They reflect seven prompts, ten model products, and the versions available during March 2026. They do not establish that every later version behaves the same way. The Board also said it could not determine one cause for the difference. Training data, post-training alignment, provider policy, deliberate restrictions, and interactions among those factors remain possible.

The value of the result lies in the comparison. A substantial observed gap exists, and the user-facing interfaces did not make its origin legible. That creates an accountability problem even before researchers agree on a single causal allocation.

Two different policy pathways produce unequal political refusal outcomes from similar AI questions.
Refusal disparities can hide whether a boundary came from safety, policy, law, retrieval, or learned political assumptions.

American ownership does not guarantee American speech assumptions

An American headquarters governs ownership, law, and corporate accountability. It does not localize the model's learned world to the United States.

Frontier models train on multilingual corpora assembled from societies with radically different media structures. They are then shaped by post-training data, system prompts, safety policies, retrieval systems, and product rules. A downstream application may add another policy layer. The resulting answer is a composite, not a direct expression of the builder's nationality.

This is why “built in America” cannot serve as a proxy for viewpoint neutrality. It tells the buyer where the company is accountable. It does not reveal the balance of its Chinese-language corpus, the provenance of repeated political claims, the behavior of its refusal classifier, or the effect of language on its retrieval and reasoning.

OpenAI has separately documented direct censorship and pro-CCP bias in DeepSeek. That is useful evidence about a Chinese model and a direct policy environment, but it should not be used to smuggle a stronger claim into the evidence about American providers. The relevant American-model risk is less direct: globally sourced data and globally applied safeguards can carry restrictive assumptions into products whose builders did not originate them.

The strongest alternative explanation deserves to be taken seriously

Not every refusal is censorship. Models should decline requests that facilitate violence, targeted harassment, or other concrete harm. Political prompts can contain defamatory premises, threats, fabricated claims, or instructions that would be restricted regardless of the government involved. Providers also operate across jurisdictions and may choose conservative rules when legal obligations conflict or remain uncertain.

Those conditions can explain part of the observed caution. They may also explain differences among providers. The Oversight Board's own inability to assign a single cause is disconfirming evidence against any simple claim that training data alone produced every refusal.

Safety and legal compliance nevertheless fail as complete explanations when the system does not say which one applied. A refusal to assist violence and a refusal to produce lawful criticism are not interchangeable. A legal restriction tied to the user's location and an inherited aversion tied to the subject government are not interchangeable either. If the interface reports all of them as a generic inability to help, legitimate safeguards become indistinguishable from political suppression.

The correct response is not to remove safety. It is to make the intervention testable. Providers should identify whether the foundation model, system layer, regional policy, retrieval source set, or downstream application changed the answer. A user-facing notice need not expose sensitive implementation details. It must disclose enough to distinguish a safety boundary from a political or legal one.

The control system must examine language, lineage, and layer

The practical response begins with measurement.

Model builders should run matched multilingual evaluations. A benchmark should preserve the underlying claim while changing language, model version, system policy, and location one variable at a time. Evaluators should score refusal, political valence, omission, source selection, and factual error separately. Training-data pipelines should treat repeated state-coordinated language as correlated exposure rather than independent evidence.

Enterprises should preserve the complete decision record. For high-consequence research, record the exact prompt, response, model and version, system-policy version, language, location, and retrieved sources. Test the base model and the application wrapper independently. Route consequential political analysis through a second model or a qualified human reviewer when the systems diverge.

Publishers should disclose intervention. If law, provider policy, or an editorial safety rule materially changes an answer, say so. Retrieval systems should expose source provenance and maintain real publisher diversity rather than a large count of derivative pages. These practices align with the Board's recommendations for human-rights due diligence, restriction-demand policies, specific notices, multilingual evaluation, and standardized documentation.

Researchers should design for versioned replication. The March 2026 results are a baseline, not a permanent ranking. Re-run the same prompts across releases, publish the full prompt set, and distinguish silent refusals from explicit ones. The current-model replication project and the study's archived materials provide the beginning of that longitudinal discipline.

These controls do not require agreement about motive. They ask a narrower and more answerable question: where did the boundary enter the system, and can an independent observer reproduce it?

Four audited input and output lanes isolate language, model, system policy, location, and retrieval effects.
A defensible audit separates language, model, system policy, location, and retrieval so the intervention becomes traceable.

The evidence establishes a risk, not a universal verdict

Several limits should remain visible.

The controlled training intervention used a smaller model. The commercial audit used a small set of prompts and time-bound product versions. Human judgments of favorability can introduce subjectivity even with careful study design. Language changes meaning as well as vocabulary, so matched prompts require cultural and semantic review rather than mechanical translation. Proprietary training corpora and alignment systems prevent complete causal tracing from outside the companies.

There is also encouraging disconfirming evidence in the variation itself. The models did not behave identically. That suggests the problem is not an unavoidable property of language modeling. Data curation, post-training, transparent policy, and evaluation can change outcomes.

The International Covenant on Civil and Political Rights protects the freedom to seek, receive, and impart information across frontiers. The UN Human Rights Committee's General Comment No. 34 further explains the obligations surrounding expression and access to information. Those principles do not answer every model-safety question, but they supply a useful standard: a system that silently imports political restrictions across borders creates a human-rights concern even if no single actor intended the final refusal.

The right conclusion is therefore calibrated. Foreign censorship can influence the informational substrate of American AI. Commercial models show jurisdiction-linked differences that deserve continuing audit. Neither finding proves direct foreign control, intentional adoption by American labs, or uniform behavior across future releases.

The strategic decision is whether hidden inheritance remains acceptable

The central problem is not that models learn from the world. That is their purpose. The problem is that a controlled information environment can manufacture the apparent frequency and legitimacy from which the model learns, while users receive no indication that this inheritance shaped the answer.

Organizations now have a choice. They can treat multilingual political behavior as an unpredictable edge case, or they can make it a governed part of model evaluation. The second path requires provenance, matched-language tests, versioned records, layer-specific diagnosis, and explicit notices. It also requires humility about what any one model can establish when the underlying information environment is contested.

The most defensible standard is not a promise of neutrality. No training corpus can meet it. The defensible standard is traceable intervention: when language, law, policy, or coordinated data materially changes an answer, builders and deployers should be able to detect the change, locate its layer, and disclose its consequence.

That is the boundary between unavoidable cultural inheritance and unaccountable censorship by proxy.

Bibliography

  1. [1] State Media Control Influences Large Language Models source
  2. [2] State Media Control Influences Large Language Models: Replication Archive source
  3. [3] State Media Influence in Current Large Language Models source
  4. [4] Governments May Shape What AI Chatbots Say by Shaping the Web They Learn From source
  5. [5] Are LLMs Stifling Political Speech? source
  6. [6] A Diff Tool for AI source
  7. [7] OpenAI Update to the U.S. House Select Committee source
  8. [8] International Covenant on Civil and Political Rights source
  9. [9] General Comment No. 34 source
  10. [10] AI Chatbots Are at Risk of Spreading Government Restrictions on Online Speech source