Sign in

Sophie Hao

@cinnamonlab.ai
269 followers 330 following 19 posts

Assistant professor of Linguistics and Data Science at Boston University. NLP, computational linguistics, interpretability, social bias and fairness. she/her. www.notaphonologist.com

PostsRepliesMedia
Sophie Hao @cinnamonlab.ai · 10/06/2026
Finally, we compare our theory to empirical trends in AI advances reported by Epoch.ai. Applied to these data, our model predicts current expenditure trends to be higher than profit-optimal unless consumer demand is almost linear in model quality (i.e., almost non-diminishing).
A plot showing the rate of growth of spending on LLM training predicted by our model over the next few years. LLM training spend grows faster when marginal demand diminishes less in terms of LLM quality. The actual rate of investment growth for LLM training is higher than the rate predicted if demand for tokens is logarithmic in LLM quality, and lower than the rate predicted if demand for tokens is linear in LLM quality.
200
Sophie Hao @cinnamonlab.ai · 10/06/2026
In the compute-bound regime, data efficiency improvements always incentivize larger models, but data budgets and compute spend can either increase or decrease depending on the relationship between demand and quality.
Plots showing how LLM size (n*) and total training expenditure C*_train scale with data efficiency (b) in LLM quality/training token. As data efficiency increases, LLM size always increases, but total training expenditure may decrease.
100
Sophie Hao @cinnamonlab.ai · 10/06/2026
Interestingly, optimal model size, data budget, and train expenditure *decrease* as training gets more parameter-efficient. Thus, in the compute-bound setting, pretraining advances in parameter efficiency incentivize small LLMs (rather than further scaling) under our model.
Plots showing how LLM size (n*) and total training expenditure C*_train scale with parameter efficiency (a) in LLM quality/param. Both quantities decrease as parameter efficiency increases.
100
Sophie Hao @cinnamonlab.ai · 10/06/2026
The scaling exponent depends on how consumer demand diminishes with quality: it is slightly superlinear when demand ~ quality. If demand diminishes (e.g., demand ~ log(quality)), optimal model size, data budget, and train spend scale no more than linearly in hardware efficiency.
Plots showing how LLM size (n*) and total training expenditure C*_train scales with hardware efficiency (E) in FLOPs/$. Numerically, we find that both quantities scale sublinearly with E. Analytically, our asymptotic bounds are at most slightly superlinear.
100
Sophie Hao @cinnamonlab.ai · 10/06/2026
When compute-bound, we show that optimal model size n*, data budget d*, and train spend C*_train scale at most polynomially with hardware efficiency E. The scaling exponent is at most slightly superlinear.
Mathematical expressions giving asymptotic upper bounds (Big-O notation) on the profit-optimal LLM size (n*) and training data size (d*) based on hardware efficiency (E), parameter efficiency (a), and data efficiency (b). The full expression is available at https://arxiv.org/abs/2605.16430
100
Sophie Hao @cinnamonlab.ai · 10/06/2026
Re: OpenAI/Anthropic IPO news, a preprint with @lambdaviking.bsky.social Scaling up training reliably improves LLMs, but it also increases training and inference costs, leading to massive capital expenditure by AI firms. How can we understand what level of LLM scaling is justified economically? 🧵⬇️
A screenshot of the title of a research paper. The title is "A Theory of Training Profit-Optimal LLMs," and the authors are Sophie Hao from Boston University and William Merrill from the Allen Institute of AI.
110