Sign in

akbir khan

@akbir.bsky.social
393 followers 156 following 30 posts

dumbest overseer at @anthropic www.akbir.dev

PostsRepliesMedia
Reposted by akbir khan
Epoch AI @epochai.bsky.social · 08/05/2025
We’ve added four new benchmarks to the Epoch AI Benchmarking Hub: Aider Polyglot, WeirdML, Balrog, and Factorio Learning Environment! Before we only featured our own evaluation results, but this new data comes from trusted external leaderboards. And we've got more on the way 🧵
142
Reposted by akbir khan
Epoch AI @epochai.bsky.social · 08/05/2025
4. Factorio Learning Environment by Jack Hopkins, Märt Bakler , and @akbir.bsky.social This benchmark uses the factory-building game Factorio to test complex, long-term planning, with settings for lab-play (structured tasks) and open-play (unbounded growth). jackhopkins.github.io/factorio-lea...
jackhopkins.github.io
Factorio Learning Environment
Claude Sonnet 3.5 builds factories
121
Reposted by akbir khan
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 25/03/2025
New Anthropic blog post: Subtle sabotage in automated researchers. As AI systems increasingly assist with AI research, how do we ensure they're not subtly sabotaging that research? We show that malicious models can undermine ML research tasks in ways that are hard to detect.
143
akbir khan @akbir.bsky.social · 18/03/2025
control is a complimentary approach to alignment. its really sensible, practical and can be done now, even before systems are superintelligent. youtu.be/6Unxqr50Kqg?...
youtu.be
Controlling powerful AI
YouTube video by Anthropic
042
Reposted by akbir khan
Ethan Mollick @emollick.bsky.social · 25/02/2025
This is a crazy paper. Fine-tuning a big GPT-4o on a small amount of insecure code or even "bad numbers" (like 666) makes them misaligned in almost everything else. They are more likely to start offering misinformation, spouting anti-human values, and talk about admiring dictators. Why is unclear.
721443
akbir khan @akbir.bsky.social · 11/02/2025
www.anthropic.com/news/paris-a...
anthropic.com
Statement from Dario Amodei on the Paris AI Action Summit
A call for greater focus and urgency
010
akbir khan @akbir.bsky.social · 01/02/2025
This is the entire goal
050
akbir khan @akbir.bsky.social · 30/01/2025
darioamodei.com/on-deepseek-...
darioamodei.com
Dario Amodei — On DeepSeek and Export Controls
On DeepSeek and Export Controls
031
Reposted by akbir khan
Hank Green @hankgreen.bsky.social · 28/01/2025
The fact that Deepseek R1 was released three days /before/ Stargate means these guys stood in front of Trump and said they needed half a trillion dollars while they knew R1 was open source and trained for $5M. Beautiful.
Trump announces 500B in AI funding. Five days ago. Deepseek r1 release. 8 days ago.
396138081762
Reposted by akbir khan
Zack Witten @zswitten.bsky.social · 24/01/2025
Can anyone get a shorter DeepSeek R1 CoT than this?
3171
Reposted by akbir khan
Tom Everitt @tom4everitt.bsky.social · 23/01/2025
Process based supervision done right, and with pretty CIDs to illustrate :)
081
Reposted by akbir khan
Mark Riedl @markriedl.bsky.social · 21/01/2025
I don’t really have the energy for politics right now. So I will observe without comment: Executive Order 14110 was revoked (Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence)
29336
akbir khan @akbir.bsky.social · 21/01/2025
R1 model is impressive
020
Reposted by akbir khan
Sarah Jones @sarahjones.bsky.social · 16/01/2025
428327646792
Reposted by akbir khan
/\__/\__/\__/ @cardiograms.bsky.social · 16/01/2025
David Lynch smirking in an interview

-Believe it or not, Eraserhead is my most spiritual film.
-Elaborate on that.
-No.
150229844079
akbir khan @akbir.bsky.social · 16/01/2025
fuck the tabloids were right www.nytimes.com/2025/01/15/t...
nytimes.com
She Is in Love With ChatGPT
A 28-year-old woman with a busy social life spends hours on end talking to her A.I. boyfriend for advice and consolation. And yes, they do have sex.
020
Reposted by akbir khan
Ethan Mollick @emollick.bsky.social · 15/01/2025
New randomized, controlled trial by the World Bank of students using GPT-4 as a tutor in Nigeria. Six weeks of after-school AI tutoring = 2 years of typical learning gains, outperforming 80% of other educational interventions. And it helped all students, especially girls who were initially behind.
1535388
Reposted by akbir khan
Ethan Mollick @emollick.bsky.social · 14/01/2025
Generative AI has flaws and biases, and there is a tendency for academics to fix on that (85% of equity LLM papers focus on harms)… …yet in many ways LLMs are uniquely powerful among new technologies for helping people equitably in education and healthcare. We need an urgent focus on how to do that
26911
Reposted by akbir khan
Ethan Mollick @emollick.bsky.social · 14/01/2025
On one hand, this paper finds adding inference-time compute (like o1 does) improves medical reasoning, which is an important finding suggesting a way to continue to improve AI performance in medicine On the other hand, scientific illustrations are apparently just anime now arxiv.org/pdf/2501.06458
2715
akbir khan @akbir.bsky.social · 13/01/2025
my metabolism is noticeably higher in london than the bay.
020
akbir khan @akbir.bsky.social · 10/01/2025
What can AI researchers do *today* that AI developers will find useful for ensuring the safety of future advanced AI systems? To ring in the new year, the Anthropic Alignment Science team is sharing some thoughts on research directions we think are important. alignment.anthropic.com/2025/recomme...
alignment.anthropic.com
Recommendations for Technical AI Safety Research Directions
2227
Reposted by akbir khan
Hank Green @hankgreen.bsky.social · 05/01/2025
My hottest take is that nothing makes any sense at all outside of the context of the constantly increasing value of human life, but that increase in value is so invisible (and exists in a world that was built for previous, lower values) that we constantly think the opposite has happened.
54175987
Reposted by akbir khan
Zack Witten @zswitten.bsky.social · 04/01/2025
darioamodei.com/machines-of-...
darioamodei.com
Dario Amodei — Machines of Loving Grace
How AI Could Transform the World for the Better
031
akbir khan @akbir.bsky.social · 04/01/2025
Nothing kills my excitement of returning to the US like the response i get from CBP officers.
170
Reposted by akbir khan
Andrew Lampinen @lampinen.bsky.social · 02/01/2025
Felix Hill was such an incredible mentor — and occasional cold water swimming partner — to me. He's a huge part of why I joined DeepMind and how I've come to approach research. Even a month later, it's still hard to believe he's gone.
Felix Hill and some other DMers and I after cold water swimming at Parliament Hill Lido a few years ago
712417
Reposted by akbir khan
Jane Wang @janexwang.bsky.social · 03/01/2025
A brilliant colleague and wonderful soul Felix Hill recently passed away. This was a shock and in an effort to sort some things out, I wrote them down. Maybe this will help someone else, but at the very least it helped me. Rest in peace, Felix, you will be missed. www.janexwang.com/blog/2025/1/...
janexwang.com
Felix — Jane X. Wang
From the moment I heard him give a talk, I knew I wanted to work with Felix . His ideas about generalization and situatedness made explicit thoughts that had been swirling around in my head, incohe...
26311
Reposted by akbir khan
Edward Grefenstette @egrefen.bsky.social · 30/12/2024
A few great papers out @ucldark.com this year. To single out two I love, there's the already well-cited paper on Debate by @akbir.bsky.social et al. which got best paper at ICML! [11/17]
141
akbir khan @akbir.bsky.social · 27/12/2024
just read this cover to cover in like 4 hours. strong recommend.
040
Reposted by akbir khan
Sam Bowman @sleepinyourhat.bsky.social · 18/12/2024
Alongside our paper, we also recorded a roundtable video featuring four of the paper’s authors discussing the results and their implications in detail:
youtube.com
Alignment faking in large language models
YouTube video by Anthropic
1212
Reposted by akbir khan
Sam Bowman @sleepinyourhat.bsky.social · 18/12/2024
New work from my team at Anthropic in collaboration with Redwood Research. I think this is plausibly the most important AGI safety result of the year. Cross-posting the thread below:
Title card: Alignment Faking in Large Language Models by Greenblatt et al.
512629
akbir khan @akbir.bsky.social · 18/12/2024
Alignment faking occurs in sufficiently smart models. www.anthropic.com/research/ali... time.com/7202784/ai-r...
time.com
Exclusive: New Research Shows AI Strategically Lying
Experiments by Anthropic and Redwood Research show how Anthropic's model, Claude, is capable of strategic deceit
030
Reposted by akbir khan
Aengus Lynch @aengusl.bsky.social · 13/12/2024
NEW PAPER: Best-of-N Jailbreaking. We modify LLM inputs with simple, randomly generated augmentations and jailbreak frontier models across text, vision, and audio modalities. The algorithm is simple, scalable and highly effective.
151
akbir khan @akbir.bsky.social · 12/12/2024
why you need SO
020
akbir khan @akbir.bsky.social · 11/12/2024
if you can’t recognise o1s progress then you need scalable oversight
000
Reposted by akbir khan
SAKE.aM 🍶 @sakeandmiyazaki.bsky.social · 04/12/2024
Just finished the first two episodes of “Pantheon” on Netflix. I’m so moved. If you got free time and think you’ll love a Black Mirror-ish story themed adult animated series, PLEASEEEEE go watch this show and tell me what you think.
7648
Reposted by akbir khan
Max Roser @maxroser.bsky.social · 03/12/2024
Sometimes, the most important news is when something isn’t happening. In my new @OurWorldInData article, I highlight that US airlines have transported passengers for more than two light-years since the last plane crash. ourworldindata.org/us-airline-t...
ourworldindata.org
US airlines have transported passengers for more than two light-years since the last plane crash
Sometimes, the most important news is when something isn’t happening.
510023
Reposted by akbir khan
Stefan Schubert @stefanschubert.bsky.social · 03/12/2024
The Netherlands - a country with 26% of the UK's population - got almost as many consolidator grants as the UK this time. erc.europa.eu/news-events/...
081
Reposted by akbir khan
Sam Bowman @sleepinyourhat.bsky.social · 02/12/2024
If you're potentially interested in transitioning into AI safety research, come collaborate with my team at Anthropic! Funded fellows program for researchers new to the field here: alignment.anthropic.com/2024/anthrop...
alignment.anthropic.com
Introducing the Anthropic Fellows Program
37316
akbir khan @akbir.bsky.social · 02/12/2024
I’m recruiting Fellows to work with me on Aligning Superhuman models. alignment.anthropic.com/2024/anthrop...
330
akbir khan @akbir.bsky.social · 23/11/2024
new kdot slams
030
Reposted by akbir khan
SAKE.aM 🍶 @sakeandmiyazaki.bsky.social · 22/11/2024
This nigga Kendrick went to therapy and got worse.
300140311686
akbir khan @akbir.bsky.social · 22/11/2024
The current structure provides you with a path where you end up with unilateral absolute control over the AGI. You stated that you don't want to control the final AGI but during this negotiation, you've shown to us that absolute control is extremely important to you www.lesswrong.com/posts/5jjk4C...
lesswrong.com
OpenAI Email Archives (from Musk v. Altman) — LessWrong
As part of the court case between Elon Musk and Sam Altman, a substantial number of emails between Elon, Sam Altman, Ilya Sutskever, and Greg Brockma…
020
akbir khan @akbir.bsky.social · 22/11/2024
impressed by the deepseek model
000
akbir khan @akbir.bsky.social · 22/11/2024
got into a weird habit of looking up someone’s thesis and reading the acknowledgments section to see who influenced them
010
akbir khan @akbir.bsky.social · 21/11/2024
incredibly cool work on demonstrating models truly do reason by Laura Ruis arxiv.org/abs/2411.12580
170
akbir khan @akbir.bsky.social · 20/11/2024
opening with Bon Iver is such a move youtu.be/DE_yVb3JMD8?...
youtu.be
Fred again.. & Jim Legxacy - NTS Radio
YouTube video by Fred again . .
000