Sign in

Omead Pooladzandi ✈️ NeurIPS'24

@hessianfree.bsky.social
476 followers 820 following 26 posts

Optimization Generative Modeling @Caltech, PhD @UCLA. ex Research Scientist Intern @AIatMeta (opinions are my own) why is jax so difficult

PostsRepliesMedia
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 26/02/2025
Hello world
010
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 27/12/2024
Newton-Schulz isn't the answer even for instantaneous whitening. PSGD: MSE( Q.T Q H , I ) = 5.2e-3 Zero-Power NS 100 iterations: MSE( NS(G) , I ) = 8.2e-1 True Inverse: MSE( H^(-1/2) H H^(-1/2), I ) = 6.1e-3 PSGD whitens information significantly better than the Newton-Schulz iters found in Muon
020
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 07/12/2024
Xilin is back at it again. Results are clear: damping hurts precision, but lower precision needs it if the underlying Hessian is extremely poorly conditioned.
020
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 02/12/2024
PSGD tracking Muon on modded nanoGPT
021
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 30/11/2024
Who is going to NeurIPS?
130
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 28/11/2024
Lol AI stats reviews consisted of one 5 rating: Top 10% of accepted papers with a confidence or 5 - absolutely certain. The reviewer raved and ranted about how good PSGD. And two confident 4 rejects with a score of 1. And one borderline reject with a confidence of 4.
120
Reposted by Omead Pooladzandi ✈️ NeurIPS'24
Pedro Cuenca @pcuenq.hf.co · 26/11/2024
SmolVLM was just released 🚀 It's a great, small, and fully open VLM that I'm really excited about for fine-tuning and on-device use cases 💻 It also comes with 0-day MLX support via mlx-vlm, here's it running at > 80 tok/s on my M1 Max 🤯
1122
Reposted by Omead Pooladzandi ✈️ NeurIPS'24
Quanquan Gu @quanquangu.bsky.social · 22/11/2024
Just put together a starter pack for Deep Learning Theory. Let me know if you'd like to be included or suggest someone to add to the list! go.bsky.app/2qnppia
298731
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 26/11/2024
PSGD ❤️ MARS MARS is a new exciting variance reduction technique from @quanquangu.bsky.social 's group which can help stabilize and accelerate your deep learning pipeline. All that is needed is a gradient buffer. Here MARS speeds up the convergence of PSGD ultimately leading to a better solution.
2145
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 25/11/2024
Oftentimes PSGD will be slow to close plasticity resulting in slightly slower convergence but ultimately a better solution.
020
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 24/11/2024
Okayyy I should actually start posting about PSGD here
170
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 24/11/2024
Hello World!
230
Reposted by Omead Pooladzandi ✈️ NeurIPS'24
Peyman Milanfar @docmilanfar.bsky.social · 24/11/2024
Radon Transform (RT) was formulated in 1917 but remained useless in practice until CT scanners were invented in the 60s But RT isn't just for CTs. It's a sort of generalization of marginals in probability RT g(p,θ): Shoot rays at θ+90 & offset p, measure line integrals of f(x,y) along the ray 1/n
28512
Reposted by Omead Pooladzandi ✈️ NeurIPS'24
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 23/11/2024
Starter packs are helpful as well as the twitter import tool chromewebstore.google.com/detail/sky-f...
chromewebstore.google.com
Sky Follower Bridge - Chrome Web Store
Instantly find and follow the same users from your Twitter follows on Bluesky.
073
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 24/11/2024
Just some light reading
160
Reposted by Omead Pooladzandi ✈️ NeurIPS'24
Ethan @ethansmith2000.com · 23/11/2024
Here, have PSGD-Kron and SOAP with FSDP2 support. Please go wild with it, let's see something finally replace ADAM. github.com/ethansmith20...
3135
Reposted by Omead Pooladzandi ✈️ NeurIPS'24
Ethan @ethansmith2000.com · 23/11/2024
probably the best in-depth explanation i've seen on FSDP at the most granular levels, props to the authors dev-discuss.pytorch.org/t/fsdp-cudac...
193