Sign in

Emily Cheng

@emcheng.bsky.social
120 followers 245 following 15 posts

generalstrikeus.com PhD student in computational linguistics at UPF chengemily1.github.io Previously: MIT CSAIL, ENS Paris Barcelona

PostsRepliesMedia
Reposted by Emily Cheng
UniReps @unireps.bsky.social · 02/10/2026
🔵🔴Still thinking about submitting to UniReps? Good news: we’ve extended the submission deadline to October 10th (AOE)! We’d love to see your work and look forward to welcoming you to UniReps! 🔴 Call for papers: unireps.org/2026/call-fo... 🔵 Submit here: openreview.net/group?id=Uni...
unireps.org
Call For Papers | UniReps Workshop
Unifying Representations in Neural Models
157
Emily Cheng @emcheng.bsky.social · 18/09/2026
UniReps submission deadline is October 4th AOE! Please share widely
001
Reposted by Emily Cheng
Marianne de Heer Kloots @mdhk.net · 30/06/2026
Let’s study learning trajectories in self-supervised speech models! 🔊 Do they reflect the hierarchical organization of spoken language? We have analyzed a lot of training checkpoints to find out 🌠 Preprint: arxiv.org/abs/2604.02043 ⬇️
12811
Emily Cheng @emcheng.bsky.social · 06/05/2026
See arxiv.org/abs/2602.04081 for the preprint Extended now across speech-audio models, more fMRI subjects, and ECoG.
arxiv.org
Abstraction Induces the Brain Alignment of Language and Speech Models
Research has repeatedly demonstrated that intermediate hidden states extracted from large language models and speech audio models predict measured brain response to natural language stimuli. Yet, very...
000
Emily Cheng @emcheng.bsky.social · 06/05/2026
Dimensionality (proxying linguistic abstraction) explains away surprisal’s effect on brain-likeness. This suggests that representing complex linguistic features drives brain-model similarity. Next-token prediction is just one task among possibly many that elicits this ability.
110
Emily Cheng @emcheng.bsky.social · 06/05/2026
Does dimensionality *cause* brain predictivity? ❌High-dimensional random features don't predict the brain! ➡️Learning good linguistic abstractions results in feature spaces that are higher-dimensional and more brain-like. Dimensionality per-se is not a causal driver.
100
Emily Cheng @emcheng.bsky.social · 06/05/2026
Explicitly increasing brain predictivity by finetuning layers on fMRI responses also *increased* both dimensionality and semantic content. So far dimensionality, linguistic abstraction, and brain predictivity seem related. But does dimensionality *cause* brain predictivity?
100
Emily Cheng @emcheng.bsky.social · 06/05/2026
Both dimensionality, i.e., ability to represent complex features, and brain predictivity grow with linguistic capabilities over LLM training.
100
Emily Cheng @emcheng.bsky.social · 06/05/2026
High dimensionality signifies that your model has learned a nice, complex feature space of language. The dimensionality peak in LLMs (and to a weaker extent, speech-audio models) marks a phase of higher-order linguistic abstraction, which we showed with probing.
100
Emily Cheng @emcheng.bsky.social · 06/05/2026
The layerwise correlation between dimensionality and brain predictivity was *highest* for voxels and electrodes in conventional fronto-temporal language areas. ...but how should we interpret dimensionality?...
100
Emily Cheng @emcheng.bsky.social · 06/05/2026
Presenting this at #ICML with @rjantonello.bsky.social and Aditya Vaidya✨ Why do 𝙢𝙞𝙙𝙙𝙡𝙚 layers in LLMs and speech-audio models best predict brain responses to language? We show a peak in the dimensionality of 🤖 activations (left) to track high 🧠 predictivity (right) 🧵(cross-posted from X)
1103
Reposted by Emily Cheng
Beatrix M. G. Nielsen @beatrixmgn.bsky.social · 15/10/2025
Good news everyone! I’ll be presenting the paper I did with Marco and Iuri "Prediction Hubs are Context-Informed Frequent tokens in LLMs" at the ELLIS UnConference on December 2nd in Copenhagen. arxiv.org/abs/2502.10201
arxiv.org
Prediction hubs are context-informed frequent tokens in LLMs
Hubness, the tendency for a few points to be among the nearest neighbours of a disproportionate number of other points, commonly arises when applying standard distance measures to high-dimensional dat...
132
Reposted by Emily Cheng
Computational Linguistics @UPF @colt-upf.bsky.social · 08/10/2025
Do you use a pronoun more often when the entity you’re talking about is more predictable? Previous work offers diverging answers so we conducted a meta-analysis, combining data from 20 studies across 8 different languages. Now out in Language: muse.jhu.edu/article/969615
131
Reposted by Emily Cheng
Nina Nusbaumer @nina-nusbaumer.bsky.social · 02/10/2025
First paper is out! Had so much fun presenting it in Marseille last July 🇨🇵 We explore how transformers handle compositionality by exploring the representations of the idiomatic and literal meaning of the same noun phrase (e.g. "silver spoon"). aclanthology.org/2025.jeptaln...
051
Reposted by Emily Cheng
Gemma Boleda @gboleda.bsky.social · 30/09/2025
New paper! 🚨 I argue that LLMs represent a synthesis between distributed and symbolic approaches to language, because, when exposed to language, they develop highly symbolic representations and processing mechanisms in addition to distributed ones. arxiv.org/abs/2502.11856
Sigmoid function. Non-linearities in neural network allow it to behave in distributed and near-symbolic fashions.
12711
Reposted by Emily Cheng
Beatrix M. G. Nielsen @beatrixmgn.bsky.social · 07/07/2025
Our paper "Prediction Hubs are Context-Informed Frequent Tokens in LLMs" has been accepted at ACL 2025! Main points: 1. Hubness is not a problem when language models do next-token prediction. 2. Nuisance hubness can appear when other comparisons are made.
181
Reposted by Emily Cheng
Marianne de Heer Kloots @mdhk.net · 13/06/2025
The @interspeech.bsky.social early registration deadline is coming up in a few days! Want to learn how to analyze the inner workings of speech processing models? 🔍 Check out the programme for our tutorial: interpretingdl.github.io/speech-inter... & sign up through the conference registration form!
interpretingdl.github.io
Interpretability Techniques for Speech Models — Tutorial @ Interspeech 2025
12710
Reposted by Emily Cheng
Computational Linguistics @UPF @colt-upf.bsky.social · 26/05/2025
Last day to sign up for the COLT Symposium! Register: tinyurl.com/colt-register 📢 𝗟𝗼𝗰𝗮𝘁𝗶𝗼𝗻 𝗰𝗵𝗮𝗻𝗴𝗲📢 June 2nd, 14:30 - 19:00 UPF Campus de la Ciutadella Room 40.101 maps.app.goo.gl/1216LJRsWmTE...
051
Reposted by Emily Cheng
Computational Linguistics @UPF @colt-upf.bsky.social · 20/05/2025
⭐ Registration open til May 27th! ⭐ Website: www.upf.edu/web/colt/sym... June 2nd, UPF 𝗦𝗽𝗲𝗮𝗸𝗲𝗿 𝗹𝗶𝗻𝗲𝘂𝗽: Arianna Bisazza (language acquisition with NNs) Naomi Saphra (emergence in LLM training dynamics) Jean-Rémi King (TBD) Louise McNally (pitfalls of contextual/formal accounts of semantics)
041
Reposted by Emily Cheng
Anej Svete @anejsvete.bsky.social · 17/05/2025
🧵 Excited to share our paper "Unique Hard Attention: A Tale of Two Sides" with Selim, Jiaoda, and Ryan, where we show that the way transformers break ties in attention scores has profound implications on their expressivity! And it got accepted to ACL! :) The paper: arxiv.org/abs/2503.14615
arxiv.org
Unique Hard Attention: A Tale of Two Sides
Understanding the expressive power of transformers has recently attracted attention, as it offers insights into their abilities and limitations. Many studies analyze unique hard attention transformers...
121
Reposted by Emily Cheng
Computational Linguistics @UPF @colt-upf.bsky.social · 13/05/2025
Announcing the COLT Symposium on June 2nd! 𝗘𝗺𝗲𝗿𝗴𝗲𝗻𝘁 𝗳𝗲𝗮𝘁𝘂𝗿𝗲𝘀 𝗼𝗳 𝗹𝗮𝗻𝗴𝘂𝗮𝗴𝗲 𝗶𝗻 𝗺𝗶𝗻𝗱𝘀 𝗮𝗻𝗱 𝗺𝗮𝗰𝗵𝗶𝗻𝗲𝘀 What properties of language are emerging from work in experimental and theoretical linguistics, neuroscience & LLM interpretability? Info: tinyurl.com/colt-site Register: tinyurl.com/colt-register 🧵1/3
142
Reposted by Emily Cheng
Leonie Weissweiler @weissweiler.bsky.social · 07/04/2025
🌍📣🥳 I could not be more excited for this to be out! With a fully automated pipeline based on Universal Dependencies, 43 non-Indoeuropean languages, and the best LLMs only scoring 90.2%, I hope this will be a challenging and interesting benchmark for multilingual NLP. Go test your language models!
0131
Reposted by Emily Cheng
Joel S. @joelhs.bsky.social · 28/03/2025
NYU canceled an invited talk by the former president of Doctors Without Borders, out of fear her talk would be accused by the government of being both anti-Trump and antisemitic: ici.radio-canada.ca/nouvelle/215...
ici.radio-canada.ca
Une conférence de la Dre Joanne Liu annulée par NYU
L'université de New York allègue que son contenu peut être perçu comme antisémite, mais Joanne Liu croit que NYU craint en fait de déplaire à Donald Trump.
13407223
Reposted by Emily Cheng
Gemma Boleda @gboleda.bsky.social · 24/02/2025
new pre-print: LLMs as a synthesis between symbolic and continuous approaches to language arxiv.org/abs/2502.11856
arxiv.org
LLMs as a synthesis between symbolic and continuous approaches to language
Since the middle of the 20th century, a fierce battle is being fought between symbolic and continuous approaches to language and cognition. The success of deep learning models, and LLMs in particular,...
0152
Reposted by Emily Cheng
Beatrix M. G. Nielsen @beatrixmgn.bsky.social · 24/02/2025
The project I did with Marco Baroni and Iuri Macocco while I was in Barcelona is now on Arxiv: arxiv.org/abs/2502.10201 🎉 TLDR below 👇
arxiv.org
Prediction hubs are context-informed frequent tokens in LLMs
Hubness, the tendency for few points to be among the nearest neighbours of a disproportionate number of other points, commonly arises when applying standard distance measures to high-dimensional data,...
132
Reposted by Emily Cheng
Adam Bonica @adambonica.bsky.social · 20/02/2025
The DOGE firings have nothing to do with “efficiency” or “cutting waste.” They’re a direct push to weaken federal agencies perceived as liberal. This was evident from the start, and now the data confirms it: targeted agencies overwhelmingly those seen as more left-leaning. 🧵⬇️
Scatterplot titled “Empirical Evidence of Ideological Targeting in Federal Layoffs: Agencies seen as liberal are significantly more likely to face DOGE layoffs.”
	•	The x-axis represents Perceived Ideological Leaning of federal agencies, ranging from -2 (Most Liberal) to +2 (Most Conservative), based on survey responses from over 1,500 federal executives.
	•	The y-axis shows Agency Size (Number of Staff) on a logarithmic scale from 1,000 to 1,000,000.

Each point represents a federal agency:
	•	Red dots indicate agencies that experienced DOGE layoffs.
	•	Gray dots indicate agencies with no layoffs.

Key Observations:
	•	Liberal-leaning agencies (left side of the plot) are disproportionately represented among red dots, indicating higher layoff rates.
	•	Notable targeted agencies include:
	•	HHS (Health & Human Services)
	•	EPA (Environmental Protection Agency)
	•	NIH (National Institutes of Health)
	•	CFPB (Consumer Financial Protection Bureau)
	•	Dept. of Education
	•	USAID (U.S. Agency for International Development)
	•	The National Nuclear Security Administration (DOE), despite its conservative leaning (+1 on the scale), is an exception among targeted agencies.
	•	A notable outlier: the Department of Veterans Affairs (moderately conservative) also faced layoffs despite its size.

Takeaway:

The figure visually demonstrates that DOGE layoffs disproportionately targeted liberal-leaning agencies, supporting claims of ideological bias. The pattern reveals that layoffs were not driven by agency size or budget alone but were strongly associated with perceived ideology.

Source: Richardson, Clinton, & Lewis (2018). Elite Perceptions of Agency Ideology and Workforce Skill. The Journal of Politics, 80(1).
251106184756
Reposted by Emily Cheng
Darby Saxbe @darbysaxbe.bsky.social · 04/02/2025
🚨BREAKING. From a program officer at the National Science Foundation, a list of keywords that can cause a grant to be pulled. I will be sharing screenshots of these keywords along with a decision tree. Please share widely. This is a crisis for academic freedom & science.
list of banned keywords
12652769115654
Emily Cheng @emcheng.bsky.social · 02/02/2025
arxiv.org/abs/2405.15471 with Diego Doimo, Corentin Kervadec, Iuri Macocco, Jade Yu, Alessandro Laio, and Marco Baroni. 6/6
arxiv.org
Emergence of a High-Dimensional Abstraction Phase in Language Transformers
A language model (LM) is a mapping from a linguistic context to an output token. However, much remains to be known about this mapping, including how its geometric properties relate to its function. We...
010
Emily Cheng @emcheng.bsky.social · 02/02/2025
3️⃣LLMs that are better at next-token prediction have higher, earlier ID peaks. 5/6
130
Emily Cheng @emcheng.bsky.social · 02/02/2025
2️⃣ The ID peak (beige) is where different LLMs are most similar (big shapes). All LLMs share this high-dimensional phase of linguistic abstraction, but... 4/6
110
Emily Cheng @emcheng.bsky.social · 02/02/2025
... the ID peak marks where syntactic, semantic, and abstract linguistic features like toxicity and sentiment are first decodable. ⭐use these layers for downstream transfer! (e.g., for brain encoding models, see arxiv.org/abs/2409.05771) 3/6
110
Emily Cheng @emcheng.bsky.social · 02/02/2025
1️⃣ The ID peak is linguistically relevant. - it collapses on shuffled text (destroying syntactic/semantic structure) - it grows over the course of training... 2/6
110
Emily Cheng @emcheng.bsky.social · 02/02/2025
Here's our work accepted to #ICLR2025! We look at how intrinsic dimension evolves over LLM layers, spotting a universal high-dimensional phase. This ID peak is where: - linguistic features are built - different LLMs are most similar, with implications for task transfer 🧵 1/6
1122
Reposted by Emily Cheng
Katie Mack @astrokatie.com · 28/01/2025
I think some people hear “grants” and think that without them, scientists and government workers just have less stuff to play with at work. But grants fund salaries for students, academics, researchers, and people who work in all areas of public service. “Pausing” grants means people don’t eat.
washingtonpost.com
White House pauses all federal grants, sparking confusion
The Trump administration has put a hold on all federal financial grants and loans, affecting tens of billions of dollars in payments.
15754323614329
Reposted by Emily Cheng
Computational Linguistics @UPF @colt-upf.bsky.social · 02/12/2024
🔊New EMNLP paper from Eleonora Gualdoni & @gboleda.bsky.social ! Why do objects have many names? Human lexicons contain different words that speakers can use to refer to the same object, e.g., purple or magenta for the same color. We investigate using tools from efficient coding...🧵 1/3
1267
Reposted by Emily Cheng
Computational Linguistics @UPF @colt-upf.bsky.social · 25/11/2024
⚡Postdoc opportunity w/ COLT Beatriu de Pinós contract, 3 yrs, competitive call by Catalan government. Apply with a PI (Marco Gemma or Thomas) Reqs: min 2y postdoc experience outside Spain, not having lived in Spain for >12 months in the last 3y. Application ~December-February (exact dates TBD)
062
Reposted by Emily Cheng
Computational Linguistics @UPF @colt-upf.bsky.social · 25/11/2024
Hello🌍! We're a computational linguistics group in Barcelona headed by Gemma Boleda, Marco Baroni & Thomas Brochhagen We do psycholinguistics, cogsci, language evolution & NLP, with diverse backgrounds in philosophy, formal linguistics, CS & physics Get in touch for postdoc, PhD & MS openings!
0141
Reposted by Emily Cheng
Alex Williams @itsneuronal.bsky.social · 18/11/2024
My lab has been working on comparing neural representations for the past few years - methods like RSA, CKA, CCA, Procrustes distance We are often asked: What do these things tell us about the system's function? How do they relate to decoding? Our new paper has some answers arxiv.org/abs/2411.08197
621571