Sign in

Brandon Stewart

@bstewart.bsky.social
3.5K followers 465 following 19 posts

Associate Professor of Sociology, Princeton brandonstewart.org

PostsRepliesMedia
Brandon Stewart @bstewart.bsky.social · 13/05/2026
bsky.app/profile/hwai...
081
Brandon Stewart @bstewart.bsky.social · 13/05/2026
bsky.app/profile/solm...
052
Brandon Stewart @bstewart.bsky.social · 13/05/2026
17/ Thank you to everyone who gave feedback
290
Brandon Stewart @bstewart.bsky.social · 13/05/2026
16/ Here’s the actual article: doi.org/10.1038/s415...
doi.org
State media control influences large language models - Nature
Government-controlled media influences the output of large language models via their training data, and models queried in the languages of countries with lower media freedom show a stronger ...
1308
Brandon Stewart @bstewart.bsky.social · 13/05/2026
15/ Substantive takeaway: LLMs separate the message from the messenger. State-coordinated phrasing can circulate through the web, enter training data, and reappear as neutral-sounding LLM output— the source obscured.
19329
Brandon Stewart @bstewart.bsky.social · 13/05/2026
14/ This paper is based on recent models at time of submission (October 2024!), but there are newer models now. We replicate the findings with the latest models here: state-media-influence-llm.github.io.
1170
Brandon Stewart @bstewart.bsky.social · 13/05/2026
13/ No single test can tell us the complete story. But when open-data analysis, memorization tests, pretraining experiments, and cross-language comparisons all point in the same direction, the best explanation is that media control is already shaping model behavior.
1216
Brandon Stewart @bstewart.bsky.social · 13/05/2026
12/ We repeat the audit with 37 countries where one national language has over 70% of the global speakers. Lower media freedom states have more pro-state responses in the state language relative to English(Study 6)
12612
Brandon Stewart @bstewart.bsky.social · 13/05/2026
11/ These first five studies traced the influence from documents in the training set to the signature of influence in the way questions are answered across languages. The next big question, does this hold across countries?
1122
Brandon Stewart @bstewart.bsky.social · 13/05/2026
10/ We had to write questions for the audit though. What about real user queries about Xi Jinping or the CCP? The results replicate (Study 5).
1100
Brandon Stewart @bstewart.bsky.social · 13/05/2026
9/ What about commercial models? We ask the same political questions in English and Chinese. A signal that the language-specific training data matters, is that the answers will differ. We audit this indirect signal (Study 4)
1153
Brandon Stewart @bstewart.bsky.social · 13/05/2026
8/ We approximated the ideal experiment with a smaller open model. We used Llama 2 13B to do continued pre-training using LoRA with different Chinese-language documents (Study 3).
1130
Brandon Stewart @bstewart.bsky.social · 13/05/2026
7/ But, you say, maybe the tech companies filter these out? We can show models memorize state-coordinated phrases at rates higher than common Chinese phrases (Study 2).
1153
Brandon Stewart @bstewart.bsky.social · 13/05/2026
6/ We can’t scour the exact training data for the state-scripted media (LLM training data is a trade secret), but we can look at open LLM training data sources like CulturaX—an open multilingual training set (Study 1).
1120
Brandon Stewart @bstewart.bsky.social · 13/05/2026
5/ In prior work we showed how leaked state-scripted talking points were reprinted across hundreds of Chinese newspapers. We re-use this data plus Xuexi Qiangguo—an app for teaching “Xi Jinping Thought”.
1160
Brandon Stewart @bstewart.bsky.social · 13/05/2026
4/ The first five studies are a case study of China which traces influence through the model from training documents to model answers to real queries. The sixth study tests whether the expected correlation exists across countries.
1152
Brandon Stewart @bstewart.bsky.social · 13/05/2026
3/ The paper is about the forces that shape how LLMs behave, but that’s tricky to study because for the most popular models the companies disclose few details of training. No single test is going to answer our questions so we used six connected studies.
1245
Brandon Stewart @bstewart.bsky.social · 13/05/2026
2/ Joint work with the most amazing collaborators: @hwaight.bsky.social & @eddieyang.bsky.social, @reasonyy.bsky.social, @solmg.bsky.social, @mollyeroberts.bsky.social, @jatucker.bsky.social. This was a true team effort.
1200
Brandon Stewart @bstewart.bsky.social · 13/05/2026
1/ New @Nature! We study how powerful institutions shape the information environment for LLMs. Commercial LLM training is opaque, so we trace a path from state-coordinated media -> training data -> model responses.
323289