Some details in our HPLT report arxiv.org/abs/2503.10267 . Code released here: github.com/hplt-project...
arxiv.org
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies
Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In this work, we present ...