🚀 How do you choose the right LLM training recipe? We just released a tech report on the scaling-law pipeline of @openeurollm.bsky.social 🇪🇺, along with the corresponding checkpoints, data, and code.
arxiv.org/abs/2608.28308
1/N 🧵
arxiv.org
Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss
We study the scaling behavior of learning rate and batch size in pretraining dense large language models on English-prevalent corpora. Beyond scaling jointly optimal learning rates and batch sizes, we...