Sign in

Roy Frostig

@froystig.bsky.social
737 followers 138 following 5 posts

research scientist at google deepmind. co-author of JAX (github.com/jax-ml/jax). cs.stanford.edu/~rfrostig

PostsRepliesMedia
Roy Frostig @froystig.bsky.social · 06/02/2025
And for diffusion specifically, there are several reference models implemented in github.com/AI-Hypercomp..., in case you hadn't come across it already and it suits some of your needs.
100
Roy Frostig @froystig.bsky.social · 06/02/2025
are useful beyond that. We have good input here already, but DM or email me if you'd ever like to talk with the team some more. Either way we appreciate it!
100
Roy Frostig @froystig.bsky.social · 06/02/2025
Indeed it's great to hear from you! And thanks for all of this detail. I've shared it with team members who've started working on models. You'll see more on the transformers side at first since that's already underway (and e.g. relates to the book upthread) but your points on diffusion and GNNs ...
100
Roy Frostig @froystig.bsky.social · 05/02/2025
@nmboffi.bsky.social – We have some plans to improve that this year. As examples, do you have any models in particular that you'd really like to see? Does training, tuning, inference, or anything else matter most to you? What hardware?
140
Reposted by Roy Frostig
Jeff Dean @jeffdean.bsky.social · 04/02/2025
Training our most capable Gemini models relies heavily on our JAX software stack+Google's TPU hardware platforms. If you want to learn more, see this awesome book "How to Scale Your Model": jax-ml.github.io/scaling-book/ Put together by several of my Google DeepMind colleagues listed below 🎉.
27813
Roy Frostig @froystig.bsky.social · 04/02/2025
Our online book on systems principles of LLM scaling is live at jax-ml.github.io/scaling-book/ We hope that it helps you make the most of your computing resources. Enjoy!
jax-ml.github.io
How To Scale Your Model
Training LLMs often feels like alchemy, but understanding and optimizing the performance of your models doesn't have to. This book aims to demystify the science of scaling language models on TPUs: how...
3359