Sign in

Arthur Douillard

@douillard.bsky.social
1.5K followers 155 following 78 posts

distributed (diloco) + modularity (dipaco) + llm @ deepmind | continual learning phd @ sorbonne

PostsRepliesMedia
Arthur Douillard @douillard.bsky.social · 19/02/2025
read the full paper: arxiv.org/abs/2502.12996 @huggingface page: huggingface.co/papers/2502.12996 congrats to my collaborators @SatyenKale who led that work and Yani Donchev
000
Arthur Douillard @douillard.bsky.social · 19/02/2025
required bandwidth reduction is massive, for a 100B params model: DP requires 471 Gbits/s Streaming DiLoCo with inner com. overlap: 1.4 Gbits/s Streaming DiLoCo with eager outer com. overlap: 400Mbits/s, more than 1000x reduction 400Mbits/s is consumer-grade bandwidth FYI
100
Arthur Douillard @douillard.bsky.social · 19/02/2025
scaling up to 1B params, and we see that our eager method can reach no loss of performance when synchronizing every 30 steps (thus overlapping 30 computation steps!), and follow closely when overlapping 100 steps
100
Arthur Douillard @douillard.bsky.social · 19/02/2025
thus we propose an *eager* version: the update is made of the average of the *local up-to-date* update of the self replica and the *remote stale* update from the other replicas
100
Arthur Douillard @douillard.bsky.social · 19/02/2025
but its performance are dramatically bad. we can recover a bit by lowering the outer learning by 4x, but this is still unsatisfying
100
Arthur Douillard @douillard.bsky.social · 19/02/2025
in this work, we explore if we can overlap an entire outer step, made of dozen to hundred of computation steps! we first try a naive "delayed" version
100
Arthur Douillard @douillard.bsky.social · 19/02/2025
Streaming DiLoCo's second contribution is to overlap communication with computation, massively increasing the tolerable latency we can safely overlap up to 5 steps, but more than that and performance drops rapidly! x.com/Ar_Douillard/status/188529212…
100
Arthur Douillard @douillard.bsky.social · 19/02/2025
DiLoCo allows us to distributed data-parallel across the world by only synchronizing once in a while, thus amortizing the communication cost however, when syncing, this is a blocking operation! x.com/Ar_Douillard/status/172473232…
100
Arthur Douillard @douillard.bsky.social · 19/02/2025
arxiv is here: arxiv.org/abs/2502.12996 read more below!
100
Arthur Douillard @douillard.bsky.social · 19/02/2025
one more step towards decentralized learning: Eager Updates can we overlap communication with computation over hundred of steps? -- yes we can in this work led by @SatyenKale, we improve DiLoCo and use x1177 less bandwidth than data-parallel
171
Arthur Douillard @douillard.bsky.social · 16/02/2025
from Jeff Dean at The Dwarkesh podcast: "asynchronous training where each copy of the model does local computation [...] it makes people uncomfortable [...] but it actually works" yep, i can confirm, it does work for real see arxiv.org/abs/2501.18512
050
Reposted by Arthur Douillard
Marco Ciccone @mcicc.bsky.social · 14/02/2025
We received an outstanding interest in our #ICLR2025 @iclr-conf.bsky.social workshop on modularity! Please sign up to serve as a reviewer if you are interested in Model Merging, MoEs, and Routing, for Decentralized and Collaborative Learning t.co/HIsZKWNaOx
011
Arthur Douillard @douillard.bsky.social · 31/01/2025
I'll be in SF in two weeks to talk at the AlgoPerf workshop, and i have a bunch of stickers to give, so let me know if you want to meet!
010
Arthur Douillard @douillard.bsky.social · 31/01/2025
Big thanks to all my collaborators! We finished this last spring, and it was one of the coolest project i've been on. The future will be distributed 🫡 arxiv.org/abs/2501.18512v1
110
Arthur Douillard @douillard.bsky.social · 31/01/2025
All of this is why we say Streaming DiLoCo is a good step towards distributed free lunch 🥪 So so many ideas we try just work on top of DiLoCo. And it can scale too! Look at the cracked folks of @PrimeIntellect who scaled their version to 10B
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
What if each replica overlap a different num of steps (\tau) because they run a different speeds? Can we break away from the lockstep synchronization? yes! Workers can have a few delay steps, and it just work, w/o any special handling.
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
There are tons of plots, tables, and charts in our paper; but let me share two more exciting plots: Over how many steps can you overlap safely communication? At least 5 without any significant loss of perf! That's a massive increase of tolerated latency.
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
Likewise, with a Llama with 405B parameters.
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
Speaking about DeepSeek, how to distribute its pretraining across the world with low-bandwidth? It has only 35B activated params, but you need to sync 671B params in total! Hard to do across continents with data-parallel... However, with our method? ❤️‍🔥
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
Indeed, post-training RL reasoning is easier to distribute (good post here primeintellect.ai/blog/intellect-math ) than pretraining but we need to scale more our pretraining, this is still a relevant axis! DeepSeek-R1 also notes it:
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
Of course the number displayed in that table are from a "simulation", but it's a pretty good indicator to what we find in practice. Abolish the tyranny of requiring huge bandwidth! ✊
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
The good part of overlapping communication with computation? As @m_ryabinin noted in Swarm Parallelism: larger networks spent more time doing computation O(n^3) vs doing communication O(n^2). We have much more time to sync at larger scales!
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
My co-author, Yani, built a simulator: a DAG with fwd, bwd, and gradient reduction nodes. It estimates how much time is spent in the costly com. between non-colocated devices and how much is spent crunching flops.
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
Put everything together, scale it to 4B, and reach similar performance than data-parallel. It's even better when overtraining with a larger token budget? remember the bitter lesson? just put more data and flops in your model, Streaming DiLoCo enables that.
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
[3] You don't need full precision for your communication. Quantize your update with 4 bits is enough -- you can barely see any changes on the performance. And that's the free dessert 🍦
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
[2] Instead of blocking computation to receive the communication, we do several compute steps asynchronously. See L9-12: the longer your model takes to do fwd/bwd, the more time you have to do communication! That's the free main course 🍱
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
[1] Instead of syncing the whole model every hundreds of steps, sync a subset of it! We split the model in fragments of three layers, it reduces massively the peak bandwidth w/o hurting ML performance. That's the free appetizer 🥗
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
With Streaming DiLoCo, we improve over the successful recipe of DiLoCo in three ways: 1. partial synchronization 2. communication overlapping with computation 3. quantized communication This is a distributed free lunch 🥪
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
Problem: in data-parallel's every step synchronization, communication is costly! DiLoCo (arxiv.org/abs/2311.08105 ) synchronizes less often --> amortizing the cost Later, @PrimeIntellect released Intellect-1, a repro of DiLoCo, with a 10B model! arxiv.org/abs/2412.01152
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
The bitter lesson is that neural networks *really* want more flops and data. This has been the key success behind LLMs. It also means you need tons of compute. however it's hard to have all that compute available in a single place, can we distribute it across the world?
PrimeIntellect's OpenDiLoCo, an open-source reproduction of DiLoCo, where compute is distributed across the world.
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
arXiv link here: arxiv.org/abs/2501.18512v1 and see thread below for a discussion:
100
Arthur Douillard @douillard.bsky.social · 31/01/2025
We release today the next step for distributed training: --> Streaming DiLoCo with Overlapping Communication. TL;DR: train data-parallel across the world with low-bandwidth for the same performance: 400x less bits exchanged & huge latency tolerance
PrimeIntellect's OpenDiLoCo, an open-source reproduction of DiLoCo, where compute is distributed across the world.
182
Reposted by Arthur Douillard
MuJoCo.org @mujoco.bsky.social · 16/01/2025
Introducing playground.mujoco.org Combining MuJoCo’s rich and thriving ecosystem, massively parallel GPU-accelerated simulation, and real-world results across a diverse range of robot platforms: quadrupeds, humanoids, dexterous hands, and arms. Get started today: pip install playground
playground.mujoco.org
MuJoCo Playground
An open-source framework for GPU-accelerated robot learning and sim-to-real transfer
17520
Reposted by Arthur Douillard
Marc Lanctot @sharky6000.bsky.social · 17/01/2025
In December, I posted about our new paper on mastering board games using internal + external planning. 👇 Here's a talk now on Youtube about it given by my awesome colleague John Schultz! www.youtube.com/watch?v=JyxE...
youtube.com
John Schultz, DeepMind, Mastering Board Games by External and Internal Planning with Language Models
YouTube video by AI4All
13511
Reposted by Arthur Douillard
Wanru Zhao @wanru.bsky.social · 16/01/2025
🚀Excited to co-organize the #ICLR2025 Workshop on Modularity for Collaborative, Decentralized, and Continual Deep Learning (MCDC @iclr-conf.bsky.social). 📃 Submission Portal: openreview.net/group?id=ICL... 🤗See you in Singapore! For more details, check out the original thread ↪️🧵
sites.google.com
MCDC@ICLR25
Summary While the success of large-scale deep learning models has hinged on the ``bigger is better'' approach – scaling model size and training data – this paradigm may rapidly be reaching an inflec...
021
Arthur Douillard @douillard.bsky.social · 16/01/2025
On the behalf, our organizer team, we are excited to see you all in Singapore for @iclr-conf.bsky.social 2025! Come talk with Marco, Colin, Prateek, Haokun, Wanru, and myself :)
030
Arthur Douillard @douillard.bsky.social · 16/01/2025
Those are the topics we are interested, and we would love to see what's your hot takes on this. Submit a paper, either short and sweet (2 pages) or longer (6 pages). See our call for paper here sites.google.com/corp/view/mc... The deadline is February 10th 2025!
sites.google.com
MCDC@ICLR25 - Call for Papers
Submission link: https://openreview.net/group?id=ICLR.cc/2025/Workshop/MCDC Questions can be directed to: mcdc-workshop@googlegroups.com Key Dates All deadlines are 23:59 AoE (Anywhere on Earth) P...
110
Arthur Douillard @douillard.bsky.social · 16/01/2025
Merging? We can combine the knowledge of several models by interpolating their weights together. That's still crazy to me. But not any kind of interpolation! While simple average may work, we can probably do better! arxiv.org/abs/2306.01708
110
Arthur Douillard @douillard.bsky.social · 16/01/2025
Adaptive Architectures? Reasoning is hot right now with OpenAI's O3. But that super costly to evaluate! Can we adapt the compute to not waste all our flops on less useful tokens? We should! arxiv.org/abs/2404.02258
110
Arthur Douillard @douillard.bsky.social · 16/01/2025
Distributed / Decentralized ? Training a modular architecture will require a lot of compute. Can we use everything across the world? arxiv.org/abs/2311.08105
110
Arthur Douillard @douillard.bsky.social · 16/01/2025
Upcycling? Training huge MoE is only affordable for a few big labs. But it is possible to join forces, combining small trained models into a larger, more powerful MoE! arxiv.org/abs/2403.07816
110
Arthur Douillard @douillard.bsky.social · 16/01/2025
Routing? There is no MoE, or even modularity without a routing. Which experts should be chosen? This is a hard topic to tackle. There are improvements! Did you know about Expert Choices (proceedings.neurips.cc/paper_files/...) or loss-free balancing (arxiv.org/html/2408.15...)?
110
Arthur Douillard @douillard.bsky.social · 16/01/2025
Mixture-of-Experts? Train different experts together, sky-rocket the number of parameters, this is the MoE used by GPT4 and Gemini. Can we improve on the Shazeer et al. formulation? arxiv.org/abs/1701.06538
110
Arthur Douillard @douillard.bsky.social · 16/01/2025
Among our topics... Collaborative? As outlined by our co-organizer Colin Raffel, there is another way to build AI, as we do for open-source softwares: colinraffel.com/blog/a-call-... I'm a big fan of the Git-theta paper! arxiv.org/abs/2306.04529
colinraffel.com
A Call to Build Models Like We Build Open-Source Software
110
Arthur Douillard @douillard.bsky.social · 16/01/2025
We're lucky to have an incredible lineup of speakers covering all the topics of our workshop. Personally, I'm super excited to see what they will present. Thanks to them!
110
Arthur Douillard @douillard.bsky.social · 16/01/2025
Workshop alert 🚨 We'll host in ICLR 2025 (late April) a workshop on modularity, encompassing collaborative + decentralized + continual learning. Those topics are on the critical path to building better AIs. Interested? submit a paper and join us in Singapore! sites.google.com/corp/view/mc...
182
Arthur Douillard @douillard.bsky.social · 04/12/2024
openreview.net/forum?id=QdE... 👀
openreview.net
Modular, Collaborative and Decentralized Deep Learning
The increasing complexity of modern machine learning models exposes the limitations of the traditional, monolithic approach to their development, raising concerns about cost and...
071
Arthur Douillard @douillard.bsky.social · 01/12/2024
PrimeIntellect have released their tech report on INTELLECT-1: t.co/8hnoTILaL3 The first open-source world-wide training of a 10B model. The underlying ML distributed algo is DiLoCo (arxiv.org/abs/2311.08105) but they also built tons of engineering on top of it to make it scalable.
0121
Arthur Douillard @douillard.bsky.social · 30/11/2024
Awesome video on speculations for test-time scaling (O1 👀 ):
youtube.com
Speculations on Test-Time Scaling (o1)
Tutorial on the technical background behind OpenAI o1. Talk written with Daniel Ritter.Slides: https://github.com/srush/awesome-o1Talk: The “large” in LLM is...
070
Arthur Douillard @douillard.bsky.social · 29/11/2024
Excellent explanation of RoPE embedding, from scratch with all the math needed: fleetwood.dev/posts/you-could-have-… And with beautiful 3blue1brown's style of animation: github.com/3b1b/manim. Original RoPE paper: arxiv.org/abs/2104.09864
05310