Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 27/12/2024Newton-Schulz isn't the answer even for instantaneous whitening. PSGD: MSE( Q.T Q H , I ) = 5.2e-3 Zero-Power NS 100 iterations: MSE( NS(G) , I ) = 8.2e-1 True Inverse: MSE( H^(-1/2) H H^(-1/2), I ) = 6.1e-3 PSGD whitens information significantly better than the Newton-Schulz iters found in Muon 020
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 07/12/2024Xilin is back at it again. Results are clear: damping hurts precision, but lower precision needs it if the underlying Hessian is extremely poorly conditioned. 020
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 02/12/2024PSGD tracking Muon on modded nanoGPT 021
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 28/11/2024Lol AI stats reviews consisted of one 5 rating: Top 10% of accepted papers with a confidence or 5 - absolutely certain. The reviewer raved and ranted about how good PSGD. And two confident 4 rejects with a score of 1. And one borderline reject with a confidence of 4. 120
Reposted by Omead Pooladzandi ✈️ NeurIPS'24Pedro Cuenca @pcuenq.hf.co · 26/11/2024SmolVLM was just released 🚀 It's a great, small, and fully open VLM that I'm really excited about for fine-tuning and on-device use cases 💻 It also comes with 0-day MLX support via mlx-vlm, here's it running at > 80 tok/s on my M1 Max 🤯 1122
Reposted by Omead Pooladzandi ✈️ NeurIPS'24Quanquan Gu @quanquangu.bsky.social · 22/11/2024Just put together a starter pack for Deep Learning Theory. Let me know if you'd like to be included or suggest someone to add to the list! go.bsky.app/2qnppia 298731
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 26/11/2024PSGD ❤️ MARS MARS is a new exciting variance reduction technique from @quanquangu.bsky.social 's group which can help stabilize and accelerate your deep learning pipeline. All that is needed is a gradient buffer. Here MARS speeds up the convergence of PSGD ultimately leading to a better solution. 2145
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 25/11/2024Oftentimes PSGD will be slow to close plasticity resulting in slightly slower convergence but ultimately a better solution. 020
Omead Pooladzandi ✈️ NeurIPS'24 @hessianfree.bsky.social · 24/11/2024Okayyy I should actually start posting about PSGD here 170
Reposted by Omead Pooladzandi ✈️ NeurIPS'24Peyman Milanfar @docmilanfar.bsky.social · 24/11/2024Radon Transform (RT) was formulated in 1917 but remained useless in practice until CT scanners were invented in the 60s But RT isn't just for CTs. It's a sort of generalization of marginals in probability RT g(p,θ): Shoot rays at θ+90 & offset p, measure line integrals of f(x,y) along the ray 1/n 28512
Reposted by Omead Pooladzandi ✈️ NeurIPS'24Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 23/11/2024Starter packs are helpful as well as the twitter import tool chromewebstore.google.com/detail/sky-f...chromewebstore.google.comSky Follower Bridge - Chrome Web StoreInstantly find and follow the same users from your Twitter follows on Bluesky. 073
Reposted by Omead Pooladzandi ✈️ NeurIPS'24Ethan @ethansmith2000.com · 23/11/2024Here, have PSGD-Kron and SOAP with FSDP2 support. Please go wild with it, let's see something finally replace ADAM. github.com/ethansmith20... 3135
Reposted by Omead Pooladzandi ✈️ NeurIPS'24Ethan @ethansmith2000.com · 23/11/2024probably the best in-depth explanation i've seen on FSDP at the most granular levels, props to the authors dev-discuss.pytorch.org/t/fsdp-cudac... 193