Sign in

Hadi Khalaf

@hadikh.bsky.social
8 followers 20 following 8 posts

phd @ harvard seas, thinking about alignment, information theory, and the likes

PostsRepliesMedia
Reposted by Hadi Khalaf
Maarten Buyl @maartenbuyl.bsky.social · 19/02/2025
AI is built to “be helpful” or “avoid harm”, but which principles should it prioritize and when? We call this alignment discretion. As Asimov's stories show: balancing such principles for AI behavior is tricky. In fact, we find that AI has its own set of priorities. (comic by @xkcd.com)🧵👇
253
Hadi Khalaf @hadikh.bsky.social · 02/02/2025
I feel queasy when I read LLM interpretability papers, some results seem wonderful but I am distrustful of the methodology and interpretation
000
Hadi Khalaf @hadikh.bsky.social · 01/02/2025
Reward modeling + BoN seem like a poor man's way to get good alignment because PPO training is expensive. Are perfect reward signals all we need? I'm not convinced
000
Hadi Khalaf @hadikh.bsky.social · 01/02/2025
The current alignment paradigm is plagued by the fact it's an flimsy adaptation of RL. RLHF is not RL but it could be and maybe it should be. There's something missing in the treatment of feedback-based alignment, and there are fundamental differences between it & RL that are not clear to me!
010
Hadi Khalaf @hadikh.bsky.social · 09/12/2024
spending my sunday battling with tikz (im losing)
000
Hadi Khalaf @hadikh.bsky.social · 06/12/2024
true agi is when your llm stops affirming everything you say
000
Hadi Khalaf @hadikh.bsky.social · 26/11/2024
doing your phd at a time when it *feels* everyone is studying the same things and you'd be facing major FOMO if you don't... is not fun
000