hailey schoelkopf @hails.computer · 19/11/2024so academic twitter is like actually-actually migrating this time huh? i still don’t know if i have it in me to actively use another social network yet 😖 7500
Reposted by hailey schoelkopfLuca Soldaini 🎀 @soldaini.net · 18/08/2023We released Dolma, the dataset for OLMo, AI2's LLM. It's 3+ trillion tokens. We hope it will help w study of language models! Available on HuggingFace w/ ImpACT license huggingface.co/datasets/allenai/dolma Overview+datasheet blog.allenai.org/dolma-3-trillion-tokens-open-llm-corpus-9a0ff4b8da64blog.allenai.orgDolma: 3 Trillion Token Open Corpus for Language Model PretrainingWe released Dolma, OLMo’s pretraining dataset. Dolma open dataset of 3 trillion tokens. Available on HuggingFace under the ImpACT license 12310