Sign in

Ahmad Beirami

@abeirami.bsky.social
3.8K followers 1.3K following 345 posts

stealth // Gemini RL+inference @ Google DeepMind // Conversational AI @ Meta // RL Agents @ EA // ML+Information Theory @ MIT+Harvard+Duke // Georgia Tech PhD 📍{NYC, SFO, YYZ} 🔗 beirami.github.io

PostsRepliesMedia
Ahmad Beirami @abeirami.bsky.social · 30/08/2026
3.7 flash seems to be a good model.
010
Reposted by Ahmad Beirami
Ramon Astudillo @ramon-astudillo.bsky.social · 26/08/2026
Only OpenAI and Anthropic models remained on the TB-fn frontier. GLM-5.3 leads the open-weight families on both benchmarks, but its gap to Sol max grows from 1.0 point on TB-2.1 to 8.6 points on TB-fn. Nice work @abeirami.bsky.social fidian.ai/blog/tb-fn-b...
fidian.ai
Who is at the frontier of terminal tasks? | Fidian
We picked the top 20 models from Artificial Analysis's Terminal-Bench 2.1 leaderboard and ran them on both TB-2.1 and TB-fn. TB-fn is Fidian's variant of Terminal-Bench, built from the same 89 tasks a...
121
Ahmad Beirami @abeirami.bsky.social · 24/07/2026
While everyone is focused on the safety risks exposed by the OpenAI/HF breach, the incident also highlights an equally important but less understood issue. Even in the non-adversarial case, unlike classical ML, we have no idea how to measure/prevent overoptimization when building agentic harnesses.
030
Ahmad Beirami @abeirami.bsky.social · 24/07/2026
I am glad the debate around PPT vs gSlides vs Keynote vs LaTeX is finally settled in favor of HTML!
030
Ahmad Beirami @abeirami.bsky.social · 16/06/2026
Given we don't have perfect verifiers that can score all possible outcomes, we'll have to settle for a surrogate. While one can optimize the surrogate reward on some "training" instances, the resulting system usually doesn't generalize beyond the "training" set.
030
Ahmad Beirami @abeirami.bsky.social · 16/06/2026
Hi Julian: I totally agree that vast body of literature on evolutionary optimization is relevant; in fact GEPA (arxiv.org/abs/2507.19457) and AlphaEvolve (arxiv.org/abs/2506.13131) are instances of evolutionary optimization. To me, the biggest open question though is that of generalization. 1/2
arxiv.org
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollo...
130
Reposted by Ahmad Beirami
Marco Z @ocramz.bsky.social · 09/06/2026
yoonholee.com/blog/2026/we...
yoonholee.com
We Should Take Text Optimization More Seriously
On where learning should happen.
251
Ahmad Beirami @abeirami.bsky.social · 23/05/2026
Three gaps have never been wider: best vs. good, end-to-end ownership vs. specialization, and vision vs. execution.
020
Ahmad Beirami @abeirami.bsky.social · 23/05/2026
Vision and execution used to be split. AI flipped the leverage. It turns vision into execution, but it cannot generate the vision itself. Execution is commoditizing. Vision is not.
110
Ahmad Beirami @abeirami.bsky.social · 23/05/2026
Work used to move through handoffs between narrow stages. The slicing has changed: bigger slices, owned end-to-end by one person. You're not just doing more. You're landing every stage yourself. Specialization needs handoffs. End-to-end ownership doesn't.
110
Ahmad Beirami @abeirami.bsky.social · 23/05/2026
AI now generates good on demand. It cannot yet reach best. Good is no longer scarce. Best still is.
110
Ahmad Beirami @abeirami.bsky.social · 23/05/2026
The best career advice used to be simple: Be the best at what you do. Ignore the trends. Become the best designer, the best engineer, the best researcher. Still necessary. No longer sufficient.
130
Ahmad Beirami @abeirami.bsky.social · 25/01/2026
Gemini for information gathering/aggregation and polishing writing, and it works quite well given G Suite integration. Claude for code-related tasks.
030
Ahmad Beirami @abeirami.bsky.social · 25/01/2026
I haven't used ChatGPT for a month now and haven't missed it. Today felt like a good day to cancel my subscription.
1180
Ahmad Beirami @abeirami.bsky.social · 14/01/2026
This position is in office in Palo Alto CA
010
Ahmad Beirami @abeirami.bsky.social · 14/01/2026
How to apply? Send an email to hiring@fidian.ai with the subject “Fidian MTS Jan 2026” and attach your resume. In the email body, please include two paragraphs: - Explain what in this space excites you the most and what you would like to work on. - Highlight an achievement you are most proud of.
110
Ahmad Beirami @abeirami.bsky.social · 14/01/2026
Why Join us? - You will be solving frontier problems in building robust AI systems, defining the standard for how the next generation of agents are built. - Work directly with a world-class team with extensive experience building frontier models and scalable AI systems.
120
Ahmad Beirami @abeirami.bsky.social · 14/01/2026
Qualifications - 3+ years of research/eng experience training LLMs, agents, and RL. - Ability to read, synthesize, and implement state-of-the-art research artifacts. - Experience training and inference from the frontier model landscape. - Comfortable working with large datasets.
110
Ahmad Beirami @abeirami.bsky.social · 14/01/2026
What we do - Advance the state-of-the-art by analyzing the latest research on agentic AI, coding agents, self-improving systems, memory architectures, and agent evaluation, and translating them into production code. - Build an agentic system that analyzes and writes production codebases for agents.
110
Ahmad Beirami @abeirami.bsky.social · 14/01/2026
We are hiring Members of Technical Staff (Research Engineers)! Current LLM agents lack reliability, creating a gap between demos and production. We solve this by automating the complex workflow of debugging, evaluation, and iteration required to make agents robust. 👇
2186
Ahmad Beirami @abeirami.bsky.social · 12/01/2026
- Iran is in a humanitarian crisis. - Thousands are reported dead in 72 hours. - We are past the point of solidarity. Empty words do not stop bullets. Action does. - The world must intervene now.
1194
Ahmad Beirami @abeirami.bsky.social · 11/01/2026
For years, Iran’s masked plainclothes regime thugs have abducted and murdered citizens with absolute impunity for wanting prosperity and refusing to fear them. Officials call it law enforcement and smear protesters as paid agents of the state’s enemies. This must end. Iranian people must prevail!
151
Ahmad Beirami @abeirami.bsky.social · 09/12/2025
This thread is approaching the edge of stability 🤣
010
Ahmad Beirami @abeirami.bsky.social · 09/12/2025
Congratulations! Very well deserved. No entropy was revealed by the decision though 😜
110
Ahmad Beirami @abeirami.bsky.social · 09/12/2025
Found myself repeating this to several students at NeurIPS: When you’re choosing an internship or a job, what you work on and who you work with matter way more than the logo. Don’t optimize for brands. Become the brand!
0243
Ahmad Beirami @abeirami.bsky.social · 02/12/2025
If you’re excited about building agentic systems, let’s chat. p.s. also on the UniReps panel Saturday on the broken state of reviewing & publishing.
030
Ahmad Beirami @abeirami.bsky.social · 02/12/2025
Hiring researchers & engineers to work on –building reliable software on top of unreliable LLM primitives –statistical evaluation of real-world deployments of LLM-based systems I’m speaking about this on two NeurIPS workshop panels: 🗓️Saturday – Reliable ML Workshop 🗓️Sunday – LLM Evaluation Workshop
2195
Ahmad Beirami @abeirami.bsky.social · 02/12/2025
Woke up to this email this morning - Wow, I won a NeurIPS award?! - …runner-up, but I’ll take it. - Wait, I didn’t submit a paper. - Ah, I’m chairing the session and I’m supposed to give the award. Huge congratulations to the actual winners and runners-up!
150
Reposted by Ahmad Beirami
Irene Cannistraci @icannistraci.bsky.social · 23/11/2025
If you're at @neuripsconf.bsky.social on Dec 6, don’t miss our panel session at @unireps.bsky.social with Ahmad Beirami, Sara Hooker and more to be announced! 🚀
011
Ahmad Beirami @abeirami.bsky.social · 22/11/2025
Will be at NeurIPS Thu Dec 4 to Sun Dec 7, excited to reconnect with old friends and make new ones. If you are excited about AI engineering (orchestration, evals, and optimizing scaffolds), we are hiring! On Saturday I’ll be on panels at the Reliable ML & UniReps workshops.
090
Ahmad Beirami @abeirami.bsky.social · 05/11/2025
Once you see a math concept geometrically, it becomes much easier to think about, and it’s hard to go back to any other way of seeing it.
130
Ahmad Beirami @abeirami.bsky.social · 24/10/2025
Whatever you are feeling is a normal response. Give yourself time and space to process, connect with others for support, and begin healing. I am happy to help in any way I can!
010
Ahmad Beirami @abeirami.bsky.social · 24/10/2025
I am sorry for what many of my excellent former colleagues are going through. Layoffs can be emotionally challenging for everyone, whether you are directly affected or not.
110
Ahmad Beirami @abeirami.bsky.social · 24/09/2025
The math that LLMs can do today is novel enough to be considered publishable, but it's not the kind of math that would be consequential.
040
Ahmad Beirami @abeirami.bsky.social · 20/09/2025
My thoughts on the broken state of AI conference reviewing. www.linkedin.com/feed/update/...
linkedin.com
My thoughts on the broken state of AI conference reviewing: Years ago, when I was in graduate school and a postdoc in Information Theory, I always felt fortunate to be invited to review for IEEE… | A...
My thoughts on the broken state of AI conference reviewing: Years ago, when I was in graduate school and a postdoc in Information Theory, I always felt fortunate to be invited to review for IEEE Tran...
080
Ahmad Beirami @abeirami.bsky.social · 11/09/2025
Let's regress from here to AGI!
030
Ahmad Beirami @abeirami.bsky.social · 10/09/2025
This is the conclusion slide of a talk I gave more than a year ago on RL/Alignment! It still holds true today.
Slide titled “Takeaways (alignment recipe).”

Step 1: Perform Best-of-n and make sure it works as desired.
– Inspect a few responses and verify the reward-induced ranking makes sense.
– Best-of-n gives the best trade-offs; if it doesn’t work, no fancy method will.
– You can debug best-of-n much faster.

Step 2: Only then train your favorite alignment method.
– Track KL(π‖p) throughout training:
• KL > 100: results are unlikely to be useful.
• KL > 15: inspect outcomes for reward hacking.
• KL < 8: you are probably OK.

Bottom banner in a black box repeats “(1) Look at your data! (2) Look at your data! (3) Look at your data!” in blue, green, and red.
030
Ahmad Beirami @abeirami.bsky.social · 09/09/2025
This also applies to telling your story (e.g., in a CV, bio, interview, etc). Focus on what you have accomplished and what you are excited about doing next; not just where you did it!
030
Ahmad Beirami @abeirami.bsky.social · 09/09/2025
Haha. the content was: If a paper is great, the credit goes to the first author. If a paper has any flaws, the responsibility falls on the last author.
130
Reposted by Ahmad Beirami
Shubhendu Trivedi @shubhendu.bsky.social · 07/09/2025
The actual unpopular opinion is that the notion of senior and junior authors should be abolished. It has completely diluted the notion of scientific authorship and created this entire industry of free-riding, head-in-the-clouds, incompetent PIs/managers. List down exact contributions instead. [+]
2102
Ahmad Beirami @abeirami.bsky.social · 09/09/2025
Glad you asked. bsky.app/profile/abei...
010
Ahmad Beirami @abeirami.bsky.social · 09/09/2025
I occasionally get messages asking how to follow my path and get into Meta, DeepMind, or similar places. That is the wrong question. Do not focus on the brand! Focus on what you want to work on, then find the opportunity that fits your goals best.
020
Reposted by Ahmad Beirami
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 07/09/2025
Related to this if a paper turns out to have a major error in it, you’re supposed to throw yourself under the bus not your students.
2241
Ahmad Beirami @abeirami.bsky.social · 07/09/2025
I think for every deliverable, there has to be one person who is responsible (gets it done) and one person who is accountable (makes sure it's done correctly). Middle authors can be responsible or accountable for a subset of tasks.
141
Ahmad Beirami @abeirami.bsky.social · 07/09/2025
proposed contribution breakdown by Atlas Wang makes a lot of sense imo: www.linkedin.com/feed/update/...
linkedin.com
Corollary: If you lack bandwidth or expertise to act as the verifier, then you shouldn&#39;t sign up to be the senior author of a paper! | Ahmad Beirami
Corollary: If you lack bandwidth or expertise to act as the verifier, then you shouldn't sign up to be the senior author of a paper!
200
Ahmad Beirami @abeirami.bsky.social · 06/09/2025
Not really. I've been saying variants of the same thing for a long time: x.com/abeirami/sta...
x.com
Ahmad Beirami on X: "Good time to remind ourselves that: If a paper is great, the credit goes to the first author. If a paper has any flaws, the responsibility falls on the last author." / X
Good time to remind ourselves that: If a paper is great, the credit goes to the first author. If a paper has any flaws, the responsibility falls on the last author.
020
Ahmad Beirami @abeirami.bsky.social · 06/09/2025
Corollary: If you lack bandwidth or expertise to act as the verifier, then you shouldn't sign up to be the senior author of a paper!
050
Ahmad Beirami @abeirami.bsky.social · 06/09/2025
The junior author is the generator. The senior author is the verifier. The verifier should teach/distill some checks to the generator, but the verifier keeps final responsibility. If a wrong claim gets out, it is on the verifier!
130
Ahmad Beirami @abeirami.bsky.social · 06/09/2025
Unpopular opinion: When a paper has a senior mentor and a junior mentee, the senior author must make sure the claims are correct and well supported. They must check every claim and gate the submission until it meets that bar.
3182
Ahmad Beirami @abeirami.bsky.social · 06/09/2025
This is the recipe for many provable claims: Make enough assumptions and narrow down the claim, then prove a narrow result with caveats. Present it as broad, hide the caveats, and declare “XYZ is provable!”
030