Sign in

Yedi Zhang

@yedizhang.bsky.social
51 followers 115 following 7 posts

PhD student @ Gatsby Unit UCL yedizhang.github.io

PostsRepliesMedia
Yedi Zhang @yedizhang.bsky.social · 02/07/2026
Is Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older optimizers like gradient descent — and this matters for generalization. Led by Sara Dragutinović and advised by Rajesh Ranganath arxiv.org/abs/2603.00742
1102
Yedi Zhang @yedizhang.bsky.social · 23/04/2026
Come chat about this @iclr-conf.bsky.social! Friday 3:15 PM, Pavilion 4, Poster #4216
062
Reposted by Yedi Zhang
Andrew Saxe @saxelab.bsky.social · 03/02/2026
Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures? Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham arxiv.org/abs/2512.20607
716042
Reposted by Yedi Zhang
Andrew Saxe @saxelab.bsky.social · 04/06/2025
How does in-context learning emerge in attention models during gradient descent training? Sharing our new Spotlight paper @icmlconf.bsky.social: Training Dynamics of In-Context Learning in Linear Attention arxiv.org/abs/2501.16265 Led by Yedi Zhang with @aaditya6284.bsky.social and Peter Latham
15318