Sign in

Noam Razin

@noamrazin.bsky.social
865 followers 267 following 34 posts

Postdoctoral Fellow at Princeton Language and Intelligence | Past: Computer Science PhD at Tel Aviv University & Apple Scholar in AI/ML | Interested in the foundations of deep learning noamrazin.github.io

PostsRepliesMedia
Noam Razin @noamrazin.bsky.social · 11/07/2025
Reward models (RMs) are key to language model post-training and inference pipelines. But, little is known about the relative pros and cons of different RM types. 📰 We investigate why RMs implicitly defined by language models (LMs) often generalize worse than explicit RMs 🧵 1/6
110
Noam Razin @noamrazin.bsky.social · 20/03/2025
The success of RLHF depends heavily on the quality of the reward model (RM), but how should we measure this quality? 📰 We study what makes a good RM from an optimization perspective. Among other results, we formalize why more accurate RMs are not necessarily better teachers! 🧵
150
Noam Razin @noamrazin.bsky.social · 14/12/2024
Presenting tomorrow a poster on why DPO often decreases the probability of preferred responses, how that can cause surprising failures in alignment, and what can we do about it. Catch me at these #NeurIPS workshop poster sessions: - M3L 11:15am - ATTRIB 3:00pm - FITML 4:40pm
180
Noam Razin @noamrazin.bsky.social · 09/12/2024
I am attending NeurIPS! Feel free to reach out if you want to chat. I will present in the M3L, FITML, and ATTRIB workshops our paper on why DPO often decreases the probability of preferred responses and how that can lead to weird failures in alignment. arxiv.org/abs/2410.08847
arxiv.org
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
Direct Preference Optimization (DPO) and its variants are increasingly used for aligning language models with human preferences. Although these methods are designed to teach a model to generate prefer...
070
Noam Razin @noamrazin.bsky.social · 26/11/2024
Catch Sadhika's talk today if you want to learn more about the surprising ways in which aligning language models based on preference data can fail
010
Noam Razin @noamrazin.bsky.social · 20/11/2024
Nadav Cohen and I recently uploaded lecture notes on the theory (and surprising practical applications) of linear neural networks. Hope that it can be useful, especially to those entering the field as it highlights distinctions between DL and "classical" ML theory arxiv.org/abs/2408.13767
arxiv.org
Lecture Notes on Linear Neural Networks: A Tale of Optimization and Generalization in Deep Learning
These notes are based on a lecture delivered by NC on March 2021, as part of an advanced course in Princeton University on the mathematical understanding of deep learning. They present a theory (devel...
031