Sign in

Leon Lang

@leon-lang.bsky.social
237 followers 116 following 22 posts

PhD Candidate at the University of Amsterdam. AI Alignment and safety research. Formerly multivariate information theory and equivariant deep learning. Masters degrees in both maths and AI. langleon.github.io

PostsRepliesMedia
Reposted by Leon Lang
AMLab @amlab.bsky.social · 06/05/2025
⚠️ The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret By Lukas Fluri*, @leon-lang.bsky.social *, Alessandro Abate, Patrick Forré, David Krueger, Joar Skalse 📜 arxiv.org/abs/2406.15753 🧵6 / 8
151
Leon Lang @leon-lang.bsky.social · 03/03/2025
Paper link: arxiv.org/abs/2502.21262 (4/4)
arxiv.org
Modeling Human Beliefs about AI Behavior for Scalable Oversight
Contemporary work in AI alignment often relies on human feedback to teach AI systems human preferences and values. Yet as AI systems grow more capable, human feedback becomes increasingly unreliable. ...
010
Leon Lang @leon-lang.bsky.social · 03/03/2025
I theoretically describe what modeling the human's beliefs would mean, and explain a practical proposal for how one could try to do this, based on foundation models whose internal representations *translate to* the human's beliefs using an implicit ontology translation. (3/4)
100
Leon Lang @leon-lang.bsky.social · 03/03/2025
The idea: In the robot-hand example, when the hand is in front of the ball, the human believes the ball was grasped and gives "thumbs up", leading to bad behavior. If we knew the human's beliefs, then we could assign the feedback properly: Reward the ball being grasped! (2/4)
100
Leon Lang @leon-lang.bsky.social · 03/03/2025
Brief paper announcement (longer thread might follow): In our new paper "Modeling Human Beliefs about AI behavior for Scalable Oversight", I propose to model a human evaluator's beliefs to better interpret the feedback, which might help for scalable oversight. (1/4)
130
Leon Lang @leon-lang.bsky.social · 03/03/2025
www.arxiv.org/abs/2502.21262 I have now this follow-up paper that goes into greater detail for how to achieve the human belief modeling, both conceptually and potentially in practice.
arxiv.org
Modeling Human Beliefs about AI Behavior for Scalable Oversight
Contemporary work in AI alignment often relies on human feedback to teach AI systems human preferences and values. Yet as AI systems grow more capable, human feedback becomes increasingly unreliable. ...
000
Reposted by Leon Lang
AMLab @amlab.bsky.social · 09/12/2024
If you are attending #NeurIPS2024🇨🇦, make sure to check out AMLab's 11 accepted papers ...and to have a chat with our members there! 👩‍🔬🍻☕ Submissions include generative modelling, AI4Science, geometric deep learning, reinforcement learning and early exiting. See the thread for the full list! 🧵1 / 12
1257
Reposted by Leon Lang
Sara Magliacane hiring PhDs in Saarland @smaglia.bsky.social · 03/12/2024
First UAI conference in Latin America!! 🔥🔥🔥 North America and Europe you are nice, but sometimes I also want to visit somewhere else 😅
1174
Leon Lang @leon-lang.bsky.social · 01/12/2024
I just completed "Historian Hysteria" - Day 1 - Advent of Code 2024 #AdventOfCode adventofcode.com/2024/day/1
030
Leon Lang @leon-lang.bsky.social · 01/12/2024
I notice more “big” accounts here that follow a lot of people. The same accounts follow almost no one on twitter. Is this motivated by a difference in the algorithms of these platforms?
000
Reposted by Leon Lang
Shakeel @shakeelhashim.com · 01/12/2024
Yet another safety researcher has left OpenAI. Rosie Campbell says she has been “unsettled by some of the shifts over the last ~year, and the loss of so many people who shaped our culture”. She says she “can’t see a place” for her to continue her work internally.
35512
Reposted by Leon Lang
Jaime Sevilla @jsevillamol.bsky.social · 27/11/2024
We are taking on a mission to track progress in AI capabilities over time. Very proud of our team!
021
Reposted by Leon Lang
Sharvaree Vadgama @sharvaree.bsky.social · 24/11/2024
Hey hey, I am around in the Bay area for the next few weeks. Bay area folks hit me up if you want to meet up for coffee/ vegan food in and around SF ☕🌯 🥟 Got a major weather upgrade☀️ from Amsterdam's insanity last week 🌀🌩️
0182
Leon Lang @leon-lang.bsky.social · 25/11/2024
Thanks for highlighting our paper! :)
110
Leon Lang @leon-lang.bsky.social · 24/11/2024
Interesting, I didn’t know such things are common practice!
110
Leon Lang @leon-lang.bsky.social · 23/11/2024
I think such questionnaires should maybe generally contain a control group of people who did some brief (let’s say 15 minutes) calibration training just do understand what percentages even mean.
140
Leon Lang @leon-lang.bsky.social · 23/11/2024
Are people maybe very bad at math? I remember once that I asked my own mom to draw what one million dollars looks like in proportion to 1 billion, and she drew like what corresponds to ~ 150 million, off by a factor of 150.
330
Leon Lang @leon-lang.bsky.social · 23/11/2024
Yeah risks are then probably more external: who creates the LLM, and do they poison the data in such a way that it will associate human utterances to bad goals.
020
Leon Lang @leon-lang.bsky.social · 23/11/2024
I actually think I (essentially?) understood this! Ie my worry was whether the LLM could end up giving high likelihood to human utterances for goals that are very bad.
110
Leon Lang @leon-lang.bsky.social · 23/11/2024
I see, interesting. Is the hope basically that the LLM utters "the same things" as what the human would utter under the same goal? Is there a (somewhat futuristic...) risk that a misaligned language model might "try" to utter the human's phrase under its own misaligned goals?
130
Reposted by Leon Lang
AMLab @amlab.bsky.social · 21/11/2024
Meet our Lab's members: staff, postdocs and PhD students! :) With this starter pack you can easily connect with us and keep up to date with all the member's research and news 🦋 go.bsky.app/8EGigUy
1259
Leon Lang @leon-lang.bsky.social · 21/11/2024
You could add myself possibly
000
Leon Lang @leon-lang.bsky.social · 21/11/2024
I strongly disagree. I’d even go as far as saying that for most relevant purposes, it’s fine to say mushrooms are plants. www.google.com/url?q=https:...
google.com
The Categories Were Made For Man, Not Man For The Categories
I. “Silliest internet atheist argument” is a hotly contested title, but I have a special place in my heart for the people who occasionally try to prove Biblical fallibility by pointing …
000
Reposted by Leon Lang
Stefan Schubert @stefanschubert.bsky.social · 20/11/2024
MIT undergrads from families earning less than $200K will pay no tuition fees from 2025, and undergrads from families earning less than $100K will have everything covered, including housing, dining, and a personal allowance. news.mit.edu/2024/mit-tui...
1201
Leon Lang @leon-lang.bsky.social · 20/11/2024
I think bluesky looks much more like twitter than chat apps look alike. Bluesky even has the same ordering of buttons
020
Reposted by Leon Lang
Aaron Bergman @aaronbergman18.bsky.social · 20/11/2024
Does anyone understand why it’s so easy to clone twitter with no IP issues? It’s hard to understand qualitative legal thresholds, but the UI looking ~exactly the same both here and on threads intuitively seems like the kind of thing that could violate a copyright if twitter had pursued one
231
Leon Lang @leon-lang.bsky.social · 20/11/2024
Here :) Thanks for putting this together!
010
Reposted by Leon Lang
AMLab @amlab.bsky.social · 19/11/2024
Hi everyone! This is AMLab :) Looking forward to share our research here on 🦋 !
1265
Leon Lang @leon-lang.bsky.social · 20/11/2024
Good to have you here :P
010
Leon Lang @leon-lang.bsky.social · 20/11/2024
Sounds interesting! I didn't yet look at it, but one burning question: How do you *specify* the relationship between the human's goals and their language, such that you can infer the goals from the language? How do you prevent misspecification of that specification?
130
Reposted by Leon Lang
Luke Bailey @lukebailey.bsky.social · 19/11/2024
the success of Bluesky with an absolutely tiny team does sorta prove Musk had a point about Twitter having way too many staff
382
Reposted by Leon Lang
Alyssa Vance @alyssavance.bsky.social · 19/11/2024
What is your reason for switching to Bluesky?
21221
Leon Lang @leon-lang.bsky.social · 20/11/2024
Many other people in my twitter timeline who seem unrelated to each other suddenly all switch at the same time. Seems like it might become an important new platform.
120
Leon Lang @leon-lang.bsky.social · 20/11/2024
Hi everyone!
030