Sign in

Sophie Hao

@cinnamonlab.ai
268 followers 330 following 19 posts

Assistant professor of Linguistics and Data Science at Boston University. NLP, computational linguistics, interpretability, social bias and fairness. she/her. www.notaphonologist.com

PostsRepliesMedia
Sophie Hao @cinnamonlab.ai · 10/06/2026
Paper link: arxiv.org/abs/2605.16430
arxiv.org
A Theory of Training Profit-Optimal LLMs
Scaling LLMs requires tremendous computational resources, and recent advances in AI have gone hand in hand with massive amounts of capital expenditure. While it is established that scaling up LLMs rel...
000
Sophie Hao @cinnamonlab.ai · 10/06/2026
Our model allows us to understand the interaction of engineering innovations and market forces in LLM scaling, going beyond Chinchilla and other landmark studies. We hope this provides a foundation for engaging critically with industry statements and supporting long-term economic decision making.
110
Sophie Hao @cinnamonlab.ai · 10/06/2026
Thus, our results suggest that current investments in scaling are (perhaps implicitly) based on expectations of continued advances in hardware and algorithmic efficiency as well as near-linear or super-linear growth in demand for generated tokens as LLM quality improves.
100
Sophie Hao @cinnamonlab.ai · 10/06/2026
Finally, we compare our theory to empirical trends in AI advances reported by Epoch.ai. Applied to these data, our model predicts current expenditure trends to be higher than profit-optimal unless consumer demand is almost linear in model quality (i.e., almost non-diminishing).
A plot showing the rate of growth of spending on LLM training predicted by our model over the next few years. LLM training spend grows faster when marginal demand diminishes less in terms of LLM quality. The actual rate of investment growth for LLM training is higher than the rate predicted if demand for tokens is logarithmic in LLM quality, and lower than the rate predicted if demand for tokens is linear in LLM quality.
200
Sophie Hao @cinnamonlab.ai · 10/06/2026
Our analysis is based on the following assumptions (see §7.2 of the preprint): 1. The LLM firm has a monopoly on selling tokens 2. LLM quality increases only when N, D are scaled jointly (~Chinchilla) 3. Consumer demand increases diminishingly (or at most linearly) with quality
100
Sophie Hao @cinnamonlab.ai · 10/06/2026
Next, we consider training bound by data. In this regime, data efficiency improvements incentivize larger models and more train spend. In contrast, hardware and parameter efficiency advances incentivize *smaller* train spend when data bound.
100
Sophie Hao @cinnamonlab.ai · 10/06/2026
In the compute-bound regime, data efficiency improvements always incentivize larger models, but data budgets and compute spend can either increase or decrease depending on the relationship between demand and quality.
Plots showing how LLM size (n*) and total training expenditure C*_train scale with data efficiency (b) in LLM quality/training token. As data efficiency increases, LLM size always increases, but total training expenditure may decrease.
100
Sophie Hao @cinnamonlab.ai · 10/06/2026
Interestingly, optimal model size, data budget, and train expenditure *decrease* as training gets more parameter-efficient. Thus, in the compute-bound setting, pretraining advances in parameter efficiency incentivize small LLMs (rather than further scaling) under our model.
Plots showing how LLM size (n*) and total training expenditure C*_train scale with parameter efficiency (a) in LLM quality/param. Both quantities decrease as parameter efficiency increases.
100
Sophie Hao @cinnamonlab.ai · 10/06/2026
The scaling exponent depends on how consumer demand diminishes with quality: it is slightly superlinear when demand ~ quality. If demand diminishes (e.g., demand ~ log(quality)), optimal model size, data budget, and train spend scale no more than linearly in hardware efficiency.
Plots showing how LLM size (n*) and total training expenditure C*_train scales with hardware efficiency (E) in FLOPs/$. Numerically, we find that both quantities scale sublinearly with E. Analytically, our asymptotic bounds are at most slightly superlinear.
100
Sophie Hao @cinnamonlab.ai · 10/06/2026
When compute-bound, we show that optimal model size n*, data budget d*, and train spend C*_train scale at most polynomially with hardware efficiency E. The scaling exponent is at most slightly superlinear.
Mathematical expressions giving asymptotic upper bounds (Big-O notation) on the profit-optimal LLM size (n*) and training data size (d*) based on hardware efficiency (E), parameter efficiency (a), and data efficiency (b). The full expression is available at https://arxiv.org/abs/2605.16430
100
Sophie Hao @cinnamonlab.ai · 10/06/2026
We study this by proposing a microeconomic model grounded in scaling laws. An AI firm chooses how to scale its LLM size and data budget, balancing increased quality and demand against training and inference costs. Its behavior is characterized by the solution to its profit maximization problem.
110
Sophie Hao @cinnamonlab.ai · 10/06/2026
Re: OpenAI/Anthropic IPO news, a preprint with @lambdaviking.bsky.social Scaling up training reliably improves LLMs, but it also increases training and inference costs, leading to massive capital expenditure by AI firms. How can we understand what level of LLM scaling is justified economically? 🧵⬇️
A screenshot of the title of a research paper. The title is "A Theory of Training Profit-Optimal LLMs," and the authors are Sophie Hao from Boston University and William Merrill from the Allen Institute of AI.
110
Sophie Hao @cinnamonlab.ai · 26/08/2025
(3/3) See full thread on X: x.com/suvarna_ashi...
x.com
Ashima Suvarna🌻 on X: "1/ 🧵 New #EMNLP2025 Paper !! Toxicity detection is subjective; shaped by norms, identity, & context. Existing models and dataset overlook this nuance. Enter MODELCITIZENS: a new dataset designed to address this. ✔️ 6.8K posts, 40K annotations across diverse groups ✔️" / X
1/ 🧵 New #EMNLP2025 Paper !! Toxicity detection is subjective; shaped by norms, identity, & context. Existing models and dataset overlook this nuance. Enter MODELCITIZENS: a new dataset designed to address this. ✔️ 6.8K posts, 40K annotations across diverse groups ✔️
000
Sophie Hao @cinnamonlab.ai · 26/08/2025
(2/3) Toxicity detection is shaped by norms, identity, & context, which existing approaches overlook. Enter MODELCITIZENS: a new dataset designed to address this. ✔️ 6.8K posts, 40K annotations across diverse groups ✔️ Context-augmented scenarios ✔️ New fine-tuned models that beat GPT-4o-mini by 5.5%
100
Sophie Hao @cinnamonlab.ai · 26/08/2025
(1/3) Please check out our new paper with @skgabrie.bsky.social and her amazing students, to appear in #EMNLP2025! (🚨 Offensive Content Warning) arxiv.org/abs/2507.05455
arxiv.org
ModelCitizens: Representing Community Voices in Online Safety
Automatic toxic language detection is critical for creating safe, inclusive online spaces. However, it is a highly subjective task, with perceptions of toxic language shaped by community norms and liv...
131
Sophie Hao @cinnamonlab.ai · 14/04/2025
I'm not personally attached to the generative linguistics apparatus per se, but I was asked by the journal to write this paper as a response to another paper, and that paper is primarily opining about the possible "end of (generative) linguistics as we know it."
110
Sophie Hao @cinnamonlab.ai · 14/04/2025
I didn't say that social relevance will guarantee generative linguistics's survival (note that there is a subtle difference between "theoretical" and "generative"), but rather that social irrelevance will likely guarantee its demise.
110
Sophie Hao @cinnamonlab.ai · 14/04/2025
I'm glad you liked it! (I am the author) There are a couple of points of incommensurability between your reaction and my intentions in writing this piece, which I'll explain below.
010
Sophie Hao @cinnamonlab.ai · 16/12/2024
I keep thinking "Bluesky" is a Slavic patronymic
040
Reposted by Sophie Hao
Lindia Tjuatja @lindiatjuatja.bsky.social · 20/11/2024
💬 Have you or a loved one compared LM probabilities to human linguistic acceptability judgments? You may be overcompensating for the effect of frequency and length! 🌟 In our new paper, we rethink how we should be controlling for these factors 🧵:
Screenshot of the paper title "What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length"
18519