Sign in

Ahmad Beirami

@abeirami.bsky.social
3.8K followers 1.3K following 345 posts

stealth // Gemini RL+inference @ Google DeepMind // Conversational AI @ Meta // RL Agents @ EA // ML+Information Theory @ MIT+Harvard+Duke // Georgia Tech PhD 📍{NYC, SFO, YYZ} 🔗 beirami.github.io

PostsRepliesMedia
Reposted by Ahmad Beirami
Ramon Astudillo @ramon-astudillo.bsky.social · 26/08/2026
Only OpenAI and Anthropic models remained on the TB-fn frontier. GLM-5.3 leads the open-weight families on both benchmarks, but its gap to Sol max grows from 1.0 point on TB-2.1 to 8.6 points on TB-fn. Nice work @abeirami.bsky.social fidian.ai/blog/tb-fn-b...
fidian.ai
Who is at the frontier of terminal tasks? | Fidian
We picked the top 20 models from Artificial Analysis's Terminal-Bench 2.1 leaderboard and ran them on both TB-2.1 and TB-fn. TB-fn is Fidian's variant of Terminal-Bench, built from the same 89 tasks a...
121
Ahmad Beirami @abeirami.bsky.social · 24/07/2026
While everyone is focused on the safety risks exposed by the OpenAI/HF breach, the incident also highlights an equally important but less understood issue. Even in the non-adversarial case, unlike classical ML, we have no idea how to measure/prevent overoptimization when building agentic harnesses.
030
Ahmad Beirami @abeirami.bsky.social · 24/07/2026
I am glad the debate around PPT vs gSlides vs Keynote vs LaTeX is finally settled in favor of HTML!
030
Reposted by Ahmad Beirami
Marco Z @ocramz.bsky.social · 09/06/2026
yoonholee.com/blog/2026/we...
yoonholee.com
We Should Take Text Optimization More Seriously
On where learning should happen.
251
Ahmad Beirami @abeirami.bsky.social · 23/05/2026
The best career advice used to be simple: Be the best at what you do. Ignore the trends. Become the best designer, the best engineer, the best researcher. Still necessary. No longer sufficient.
130
Ahmad Beirami @abeirami.bsky.social · 25/01/2026
I haven't used ChatGPT for a month now and haven't missed it. Today felt like a good day to cancel my subscription.
1180
Ahmad Beirami @abeirami.bsky.social · 14/01/2026
We are hiring Members of Technical Staff (Research Engineers)! Current LLM agents lack reliability, creating a gap between demos and production. We solve this by automating the complex workflow of debugging, evaluation, and iteration required to make agents robust. 👇
2186
Ahmad Beirami @abeirami.bsky.social · 12/01/2026
- Iran is in a humanitarian crisis. - Thousands are reported dead in 72 hours. - We are past the point of solidarity. Empty words do not stop bullets. Action does. - The world must intervene now.
1194
Ahmad Beirami @abeirami.bsky.social · 11/01/2026
For years, Iran’s masked plainclothes regime thugs have abducted and murdered citizens with absolute impunity for wanting prosperity and refusing to fear them. Officials call it law enforcement and smear protesters as paid agents of the state’s enemies. This must end. Iranian people must prevail!
151
Ahmad Beirami @abeirami.bsky.social · 09/12/2025
Found myself repeating this to several students at NeurIPS: When you’re choosing an internship or a job, what you work on and who you work with matter way more than the logo. Don’t optimize for brands. Become the brand!
0243
Ahmad Beirami @abeirami.bsky.social · 02/12/2025
Hiring researchers & engineers to work on –building reliable software on top of unreliable LLM primitives –statistical evaluation of real-world deployments of LLM-based systems I’m speaking about this on two NeurIPS workshop panels: 🗓️Saturday – Reliable ML Workshop 🗓️Sunday – LLM Evaluation Workshop
2195
Ahmad Beirami @abeirami.bsky.social · 02/12/2025
Woke up to this email this morning - Wow, I won a NeurIPS award?! - …runner-up, but I’ll take it. - Wait, I didn’t submit a paper. - Ah, I’m chairing the session and I’m supposed to give the award. Huge congratulations to the actual winners and runners-up!
150
Reposted by Ahmad Beirami
Irene Cannistraci @icannistraci.bsky.social · 23/11/2025
If you're at @neuripsconf.bsky.social on Dec 6, don’t miss our panel session at @unireps.bsky.social with Ahmad Beirami, Sara Hooker and more to be announced! 🚀
011
Ahmad Beirami @abeirami.bsky.social · 22/11/2025
Will be at NeurIPS Thu Dec 4 to Sun Dec 7, excited to reconnect with old friends and make new ones. If you are excited about AI engineering (orchestration, evals, and optimizing scaffolds), we are hiring! On Saturday I’ll be on panels at the Reliable ML & UniReps workshops.
090
Ahmad Beirami @abeirami.bsky.social · 05/11/2025
Once you see a math concept geometrically, it becomes much easier to think about, and it’s hard to go back to any other way of seeing it.
130
Ahmad Beirami @abeirami.bsky.social · 24/10/2025
I am sorry for what many of my excellent former colleagues are going through. Layoffs can be emotionally challenging for everyone, whether you are directly affected or not.
110
Ahmad Beirami @abeirami.bsky.social · 24/09/2025
The math that LLMs can do today is novel enough to be considered publishable, but it's not the kind of math that would be consequential.
040
Ahmad Beirami @abeirami.bsky.social · 20/09/2025
My thoughts on the broken state of AI conference reviewing. www.linkedin.com/feed/update/...
linkedin.com
My thoughts on the broken state of AI conference reviewing: Years ago, when I was in graduate school and a postdoc in Information Theory, I always felt fortunate to be invited to review for IEEE… | A...
My thoughts on the broken state of AI conference reviewing: Years ago, when I was in graduate school and a postdoc in Information Theory, I always felt fortunate to be invited to review for IEEE Tran...
080
Ahmad Beirami @abeirami.bsky.social · 11/09/2025
Let's regress from here to AGI!
030
Ahmad Beirami @abeirami.bsky.social · 10/09/2025
This is the conclusion slide of a talk I gave more than a year ago on RL/Alignment! It still holds true today.
Slide titled “Takeaways (alignment recipe).”

Step 1: Perform Best-of-n and make sure it works as desired.
– Inspect a few responses and verify the reward-induced ranking makes sense.
– Best-of-n gives the best trade-offs; if it doesn’t work, no fancy method will.
– You can debug best-of-n much faster.

Step 2: Only then train your favorite alignment method.
– Track KL(π‖p) throughout training:
• KL > 100: results are unlikely to be useful.
• KL > 15: inspect outcomes for reward hacking.
• KL < 8: you are probably OK.

Bottom banner in a black box repeats “(1) Look at your data! (2) Look at your data! (3) Look at your data!” in blue, green, and red.
030
Ahmad Beirami @abeirami.bsky.social · 09/09/2025
This also applies to telling your story (e.g., in a CV, bio, interview, etc). Focus on what you have accomplished and what you are excited about doing next; not just where you did it!
030
Reposted by Ahmad Beirami
Shubhendu Trivedi @shubhendu.bsky.social · 07/09/2025
The actual unpopular opinion is that the notion of senior and junior authors should be abolished. It has completely diluted the notion of scientific authorship and created this entire industry of free-riding, head-in-the-clouds, incompetent PIs/managers. List down exact contributions instead. [+]
2102
Ahmad Beirami @abeirami.bsky.social · 09/09/2025
I occasionally get messages asking how to follow my path and get into Meta, DeepMind, or similar places. That is the wrong question. Do not focus on the brand! Focus on what you want to work on, then find the opportunity that fits your goals best.
020
Reposted by Ahmad Beirami
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 07/09/2025
Related to this if a paper turns out to have a major error in it, you’re supposed to throw yourself under the bus not your students.
2241
Ahmad Beirami @abeirami.bsky.social · 06/09/2025
Corollary: If you lack bandwidth or expertise to act as the verifier, then you shouldn't sign up to be the senior author of a paper!
050
Ahmad Beirami @abeirami.bsky.social · 06/09/2025
Unpopular opinion: When a paper has a senior mentor and a junior mentee, the senior author must make sure the claims are correct and well supported. They must check every claim and gate the submission until it meets that bar.
3182
Ahmad Beirami @abeirami.bsky.social · 06/09/2025
This is the recipe for many provable claims: Make enough assumptions and narrow down the claim, then prove a narrow result with caveats. Present it as broad, hide the caveats, and declare “XYZ is provable!”
030
Ahmad Beirami @abeirami.bsky.social · 05/09/2025
Today, Industry research is focused on short term (3-6months) bets. Academics have an opportunity to balance their portfolio with medium term (1-2 years) and long term (5-10 years) bets. Putting all academic efforts in short-term basket is suboptimal!
070
Ahmad Beirami @abeirami.bsky.social · 02/09/2025
When I worked in corporate, I was often first in the office because that routine worked for me. It was a personal preference, not a benchmark for anyone else. We should not judge commitment by hours, especially in research. We should look for thoughtful work and steady progress.
030
Ahmad Beirami @abeirami.bsky.social · 01/09/2025
Common mistake in LLM prompting projects: jumping into full-scale pipelines (datasets and inference) without testing feasibility. Iterating at scale is expensive and time-consuming. Start with ONE example to validate the hypothesis, verify context, debug the design, then scale.
171
Ahmad Beirami @abeirami.bsky.social · 31/08/2025
Thoughts that are explained clearly are more respected. Clarity is a scarce skill. Many of us (me included) leave out key context and make people work too hard to understand us. AI models should get better at this: not just reasoning, but communicating with the right amount of context.
280
Ahmad Beirami @abeirami.bsky.social · 30/08/2025
Enjoyed speaking with @DeltaInstitutes about going from information theory to ML, recent safety alignment/RL work, and lessons on RL for LLMs that stuck! Check out the podcast episode here: lnkd.in/eb6dWHDv
161
Ahmad Beirami @abeirami.bsky.social · 28/08/2025
In 2009, a prominent signal processing professor said the market was tough and h-index ≥6 was needed just to get a faculty interview. We now seem to be drifting toward the same bar for PhD program entrance.
040
Ahmad Beirami @abeirami.bsky.social · 26/08/2025
Congratulations to the Google team on the release of the newest Gemini Image generation model! 🍌🍌 I am super impressed with what the model did here (no other model gets even close -- including Google's previous model). This is truly bananas!
141
Ahmad Beirami @abeirami.bsky.social · 26/08/2025
What are the founders going to own? 🤔
100
Ahmad Beirami @abeirami.bsky.social · 24/08/2025
Cannot believe this needs to be said loud: Do not ride a motorized bike 30+ mph on a mixed-use trail. It’s beyond reckless and puts walkers and cyclists at grave risk! P.S. “Electric” doesn’t make it safe at 30+ mph.
132
Ahmad Beirami @abeirami.bsky.social · 24/08/2025
My thoughts on acceptance rate artificially kept low: If we focus on merit based acceptance with -claims are substantiated with theoretical/empirical evidence -claims push the science envelope by epsilon Then we still end up with ~25% acceptance rate and won't need to artificially reject any papers
261
Reposted by Ahmad Beirami
Fernando Diaz @841io.bsky.social · 23/08/2025
also works for teaching. when I taught information retrieval, lecture 1: The Information Access Problem, lectures 1-2: evaluation. the rest of the semester presents algorithms and methods leaning on those first lectures to analyze and compare. should work for many courses
121
Ahmad Beirami @abeirami.bsky.social · 23/08/2025
Being yelled at by both sides of an argument often only means either of two things: - You have a rational stance that balances all the nuances that either side is missing. - You have gone completely crazy.
030
Ahmad Beirami @abeirami.bsky.social · 23/08/2025
When we have a good verifier for a capability, we already know how to distill that capability into the model through RL. In most cases, building the verifier is the bottleneck!
010
Ahmad Beirami @abeirami.bsky.social · 23/08/2025
Concretely, a research project proceeds in this order: 1. Problem definition 2. Evaluation metrics and success criteria 3. Solution Too often, steps 1 and/or 2 are skipped or treated as afterthoughts.
140
Ahmad Beirami @abeirami.bsky.social · 23/08/2025
Too often, researchers propose solutions before they’ve defined the problem first.
030
Ahmad Beirami @abeirami.bsky.social · 20/08/2025
A research project needs test beds of different granularity. Iterating at the large scale is expensive and hard to debug Validating on smaller models is helpful in moving fast by ruling out ideas that are unlikely to work My alignment research is driven by thinking about a ternary language model
020
Ahmad Beirami @abeirami.bsky.social · 19/08/2025
Most “robustness” work (adversarial, shift, etc.) is just training on reweighted samples (augmented, model-generated, or mined). OOD generalization then comes from: (1) inductive bias (2) similarity to train data (3) luck The 3rd one is the most important of the three.
050
Ahmad Beirami @abeirami.bsky.social · 19/08/2025
Impactful papers that *shift* the paradigm are generally met with resistance. A *public* review of Shannon's paper publication: "The discussion is suggestive throughout, rather than mathematical, and it is not always clear that the author’s mathematical intentions are honorable."
081
Reposted by Ahmad Beirami
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 19/08/2025
None of our impactful papers have had an easy path through traditional venues. Most cited paper? Rejected four times. Most impactful paper? Poster at a conference. But none of it matters because arxiv makes everything work
61076
Ahmad Beirami @abeirami.bsky.social · 19/08/2025
pass@k is an effective decoding paradigm to improve models at math and also for an attacker to jailbreak the models. Now that we know we are decoding from the model with pass@k or an adversary is jailbreaking the model with pass@k, how should we think about RL? A short 🧵
120
Ahmad Beirami @abeirami.bsky.social · 16/08/2025
This is how offline calibration + PPO compares against GRPO on helpfulness BT rewards. Would be curious to see how this might help your usecases.
190
Ahmad Beirami @abeirami.bsky.social · 13/08/2025
We are wired to see our own best and others' worst. Maturity is flipping the lens: Be a generous fan of others and a tough editor of yourself! Self-awareness is underrated!
010
Ahmad Beirami @abeirami.bsky.social · 12/08/2025
The best AI researchers zoom at three abstraction levels: - High: paper-level ideas & math - Mid: code-level implementation - Low: GPU/TPU reality (kernels/memory) Low exposes bottlenecks. High accelerates exploration. Mid makes it real. The job is to translate between them!
170