Sign in

Ari Morcos

@arimorcos.bsky.social
222 followers 100 following 4 posts

CEO and Co-founder @ DatologyAI working to make it easy for anyone to make the most of their data. Former: RS FAIR, RS DeepMind, Harvard Neuroscience PhD. www.datologyai.com

PostsRepliesMedia
Ari Morcos @arimorcos.bsky.social · 29/11/2024
ICYMI, check out our latest results @datologyai.com on curating data for LLMs. Intervening only on training data, our pipeline can train models faster (7.7x less compute), better (+8.5% performance), and smaller (models half the size outperform by >5%)! www.datologyai.com/post/technic...
datologyai.com
Technical Deep-Dive: Curating Our Way to a State-of-the-Art Text Dataset
Our data curation pipeline to obtain substantial improvements in LLM quality, training speed, and inference efficiency.
052
Reposted by Ari Morcos
Haoli Yin @haoliyin.bsky.social · 25/11/2024
The text team cooked so much 🧑‍🍳 it might be better than your Thanksgiving meal Check out this super thorough thread on what and how we achieved the best curated text dataset using public data
081
Reposted by Ari Morcos
Haoli Yin @haoliyin.bsky.social · 25/11/2024
Working on making data curation dirt cheap btw If you're a cracked engineer we'd love to have you :)) DM me if you have any questions! jobs.ashbyhq.com/DatologyAI (also looking for enthusiastic research interns)
jobs.ashbyhq.com
DatologyAI Jobs
DatologyAI Jobs
031
Reposted by Ari Morcos
Pratyush Maini @pratyushmaini.bsky.social · 25/11/2024
1/5 Earlier this year, I joined @datologyai.com to give wings to the data research I had been doing in academia. Today, I am absolutely thrilled to share what we’ve been working on! Techvember Ep 2: How we made the #1 LLM Pre-training Data Recipe. Blog: 👉 tinyurl.com/best-llm-data 🧵
1154
Ari Morcos @arimorcos.bsky.social · 25/11/2024
🚀 Train faster - Reach the same performance 7.7x faster 📈 Train Better - Improve performance by 8.5% over exact-deduplicated RPJv1, 6.1% over FineWeb-Edu, and 4.4% over DCLM 🔍 Train Smaller - Train a model that's 2.1x smaller while simultaneously improving performance by >5%
datologyai.com
Train LLMs Faster, Better, and Smaller with DatologyAI’s Data Curation
DatologyAI's curated data delivers substantial improvements in LLM quality, training speed, and inference efficiency over existing datasets.
170
Reposted by Ari Morcos
Ethan @ethanrosenthal.com · 14/11/2024
Massive, impressive post on data curation strategies for producing better models with less data and compute. The best part of data curation is that it's a (relatively small) one time cost that gets amortized over all future models. Link to the technical write-up: www.datologyai.com/post/product...
datologyai.com
Technical Deep-Dive: Image-Text Data Curation at the Billion-Sample Scale
Introducing DatologyAI’s state-of-the-art data curation pipeline.
0103
Reposted by Ari Morcos
Josh Wills @spite.vc · 14/11/2024
This is the most interesting and most impactful data pipeline problem I have ever worked on (and if you know me, you know that’s saying something.) So happy to be able to share this work with the world! And now it’s time for a little vacation. 😅
0253
Ari Morcos @arimorcos.bsky.social · 14/11/2024
Hello bluesky! First post, first results drop! Today, we @datologyai.bsky.social are so excited to release our first results, demonstrating *massive* gains in training efficiency, performance, and inference efficiency with better data. www.datologyai.com/post/datolog...
datologyai.com
DatologyAI’s Image-Text Data Curation: Train Better, Faster, Smaller
What if you could save up to 98% on compute costs? Read on to find out how DatologyAI’s deep learning data curation tools make this possible.
041