Sign in

Jay Alammar

@jayalammar.bsky.social
1.6K followers 209 following 21 posts

Writer jalammar.github.io. O'Reilly Author LLM-book.com. LLM Builder Cohere.com.

PostsRepliesMedia
Reposted by Jay Alammar
Maarten Grootendorst @maartengr.bsky.social · 24/11/2025
There are three chapters in Early Release on the O’Reilly Platform (Intro, Memory, and Tools): learning.oreilly.com/library/view...
learning.oreilly.com
An Illustrated Guide to AI Agents
Artificial intelligence is entering a new phase. No longer limited to answering prompts or completing simple writing tasks, AI agents can now reason, plan, and act with increasing... - Selection from ...
051
Reposted by Jay Alammar
Maarten Grootendorst @maartengr.bsky.social · 24/11/2025
And here you have it! The cover of our “An Illustrated Guide to AI Agents” book! Why a Dolphin? The process of choosing cover animals is a closely guarded secret held deep within the legendary halls of O'Reilly. @jayalammar.bsky.social and I do not get to choose it, but we love it!
161
Jay Alammar @jayalammar.bsky.social · 03/11/2025
Inside NeurIPS 2025: The Year’s AI Research, Mapped New blog post! NeurIPS 2025 papers are out—and it’s a lot to take in. This visualization lets you explore the entire research landscape interactively, with clusters and @cohere.com LLM-generated explanations that make it easier to grasp.
1161
Reposted by Jay Alammar
Maarten Grootendorst @maartengr.bsky.social · 13/10/2025
Excited to share that @jayalammar.bsky.social and I are writing the book “An Illustrated Guide to AI Agents” with @oreilly.bsky.social 🥳 Our new book will contain chapters on the fundamentals of agents (memory, tools, and planning), alongside more advanced concepts like RL and reasoning LLMs.
2112
Jay Alammar @jayalammar.bsky.social · 13/10/2025
The Illustrated Guide to AI Agents New book announcement! Thrilled that together with @maartengr.bsky.social , we're writing a new book titled “An Illustrated Guide to AI Agents” and published by @oreilly.bsky.social.
1112
Reposted by Jay Alammar
Antoine Bosselut @abosselut.bsky.social · 03/09/2025
The next generation of open LLMs should be inclusive, compliant, and multilingual by design. That’s why we @icepfl.bsky.social @ethz.ch @cscsch.bsky.social ) built Apertus.
2248
Jay Alammar @jayalammar.bsky.social · 19/08/2025
The Illustrated GPT-OSS New post! A visual tour of the architecture, message formatting, and reasoning of the latest GPT. newsletter.languagemodels.co/p/the-illust...
1206
Jay Alammar @jayalammar.bsky.social · 23/05/2025
The legendary John Carmack at #upperbound: - Current AI focus is RL (with Richard Sutton) solving Atari games - Thinking in line with the Alberta Plan. - It was a misstep to start working too low-level (e.g., at the cuda level). I kept stepping up the stack chain until now in pytorch
060
Reposted by Jay Alammar
Adam Hill @astroadamh.bsky.social · 20/05/2025
I'm really excited for this year's PyData London conference - there are some awesome talks on the schedule and I'm excited to hear the keynote speakers @jayalammar.bsky.social, Tony Wears, & Leanne Fitzpatrick #pydata #datascience
021
Reposted by Jay Alammar
PyData London @pydatalondon.bsky.social · 20/05/2025
Unleash your inner data aficionado at PyData London 2025, 6-8 June at Convene Sancroft, St. Paul’s! We have 3 top flight keynotes lined up for you this year from @jayalammar.bsky.social, Leanne Kim Fitzpatrick and Tony Mears. Just 17 days left. Book your tickets now! pydata.org/london2025
Advertisement for PyData London 2025 conference.

Headline: Meet your keynote speakers

- Jay Alammar
- Tony Mears
- Leanne Fitzpatrick

Book your tickets
https://pydata.org/london2025
166
Reposted by Jay Alammar
Max Bartolo @maxbartolo.bsky.social · 27/03/2025
I'm excited to share the tech report for our @cohere.com @cohereforai.bsky.social Command A and Command R7B models. We highlight our novel approach to model training including self-refinement algorithms and model merging techniques at scale. Read more below! ⬇️
1104
Reposted by Jay Alammar
Tom Aarsen @tomaarsen.com · 21/02/2025
We've just released MMTEB, our multilingual upgrade to the MTEB Embedding Benchmark! It's a huge collaboration between 56 universities, labs, and organizations, resulting in a massive benchmark of 1000+ languages, 500+ tasks, and a dozen+ domains. Details in 🧵
2234
Reposted by Jay Alammar
Maarten Grootendorst @maartengr.bsky.social · 11/02/2025
Did you know we continue to develop new content for the "Hands-On Large Language Models" book? There's now even a free course available with @deeplearningai.bsky.social!
1113
Reposted by Jay Alammar
Mason Youngblood @masonyoungblood.bsky.social · 05/02/2025
Do whales optimize their vocalizations for efficiency, just like human language? 🐋🎶 My latest study in Science Advances (@science.org) suggests they do—following linguistic laws seen in human speech. 🧵 www.science.org/doi/10.1126/...
science.org
Language-like efficiency in whale communication
Whale vocalizations follow efficiency rules seen in human language, revealing striking similarities in communication systems.
520065
Reposted by Jay Alammar
Ellen Garland @ellengarland.bsky.social · 06/02/2025
We uncovered the same statistical structure that is a hallmark of human language in whale song, published today in Science. @inbalarnon.bsky.social @simonkirby.bsky.social @jennyallen13.bsky.social @clairenea.bsky.social @emma-carroll.bsky.social www.science.org/doi/10.1126/...
17269101
Reposted by Jay Alammar
Naomi Saphra @nsaphra.bsky.social · 27/01/2025
One of my grand interpretability goals is to improve human scientific understanding by analyzing scientific discovery models, but this is the most convincing case yet that we CAN learn from model interpretation: Chess grandmasters learned new play concepts from AlphaZero's internal representations.
arxiv.org
Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero
Artificial Intelligence (AI) systems have made remarkable progress, attaining super-human performance across various domains. This presents us with an opportunity to further human knowledge and improv...
210923
Jay Alammar @jayalammar.bsky.social · 27/01/2025
The Illustrated DeepSeek-R1 Spent the weekend reading the paper and sorting through the intuitions. Here's a visual guide and the main intuitions to understand the model and the process that created it. newsletter.languagemodels.co/p/the-illust...
17523
Jay Alammar @jayalammar.bsky.social · 15/01/2025
Alphaxiv is an awesome way to discuss ML papers -- often with the authors themselves. Here's an intro and demo by @rajpalleti.bsky.social we shot at #Neurips2024 www.youtube.com/watch?v=-Kwl...
youtube.com
AlphaXiv - a great place to discuss ML papers
YouTube video by Jay Alammar
0123
Reposted by Jay Alammar
Tom Aarsen @tomaarsen.com · 14/01/2025
The newest extremely strong embedding model based on ModernBERT-base is out: `cde-small-v2`. Both faster and stronger than its predecessor, this one tops the MTEB leaderboard for its tiny size! Details in 🧵
1317
Jay Alammar @jayalammar.bsky.social · 14/01/2025
Floored that the repo for Hands-On Large Language Models is now at 3.6k Github stars! And excited that professors are starting to use the book to teach LLM courses. Reach out to us if we can be of assistance! And if you've liked the book, leave us a review on Amazon or Goodreads!
1100
Jay Alammar @jayalammar.bsky.social · 13/01/2025
SWE-Bench has been one of the most important tasks measuring the progress of agents tackling software engineering in 2024. I caught up with two of its creators, @ofirpress.bsky.social and Carlos E. Jimenez to share their ideas on the state of LLM-backed agents. www.youtube.com/watch?v=bivZ...
youtube.com
SWE-Bench authors reflect on the state of LLM agents at Neurips 2024
YouTube video by Jay Alammar
032
Reposted by Jay Alammar
Nathan Lambert @natolambert.bsky.social · 20/12/2024
OpenAI's o3: The grand finale of AI in 2024 A step change as influential as the release of GPT-4. Reasoning language models are the current and next big thing. I explain: * The ARC prize * o3 model size / cost * Dispelling training myths * Extreme benchmark progress
buff.ly
o3: The grand finale of AI in 2024
A step change as influential as the release of GPT-4. Reasoning language models are the current big thing.
87912
Jay Alammar @jayalammar.bsky.social · 12/12/2024
Good morning #NeurIPS2024! Stop by the @cohere.com booth at 3PM today (Thursday) for a signed copy of Hands-On Large Language Models - it will introduce you to LLMs, their applications, as well as Cohere's Embed, Rerank, and Command-R models. Come early as quantities are limited!
070
Jay Alammar @jayalammar.bsky.social · 11/12/2024
I'll be in the Cohere #NeurIPS2024 booth most of this afternoon. Come say hi, ask questions, and yes, we're hiring! Tomorrow I'll be signing copies of my book at 3PM! Limited copies available!
060
Jay Alammar @jayalammar.bsky.social · 10/12/2024
Hi NeurIPS! Explore ~4,500 NeurIPS papers in this interactive visualization: jalammar.github.io/assets/neuri... (Click on a point to see the paper on the website) Uses @cohere.com models and @lelandmcinnes.bsky.social's datamapplot/umap to help make sense of the overwhelming scale of NeurIPS.
16113
Jay Alammar @jayalammar.bsky.social · 29/11/2024
Sure to be thought provoking. The previous interview had fascinating thoughts on scifi (Dune Vs. Foundation), on AI competition for AI safety, and on successful scifi as a self-preventing prophecy.
050
Jay Alammar @jayalammar.bsky.social · 28/11/2024
Excited to see you all at NeurIPS this year! Let's hang!
120
Jay Alammar @jayalammar.bsky.social · 26/11/2024
Great insights
110
Jay Alammar @jayalammar.bsky.social · 26/11/2024
Join us for a panel on scientific communication Dec 4!
000
Reposted by Jay Alammar
Max Bartolo @maxbartolo.bsky.social · 20/11/2024
🚨 LLMs can learn to reason from procedural knowledge in pretraining data! 🚨 I particularly enjoy research where the evidence contradicts our initial hypothesis. If you're interested in LLM reasoning, check out the 60+ pages of in-depth work at arxiv.org/abs/2411.12580
3677
Reposted by Jay Alammar
Chris McKitterick @mckitterick.bsky.social · 20/11/2024
Have you updated your handle to use your own domain? I'm in the process of updating my handle to use the adastra-sf.com doman, but though I've tried both methods of uploading a TXT file to my server (DNS interface and no interface), it's not resolving even after half an hour. Tips? Ideas?
001
Reposted by Jay Alammar
Evan Peck @peck.phd · 20/11/2024
Trying something new: A 🧵 on a topic I find many students struggle with: "why do their 📊 look more professional than my 📊?" It's *lots* of tiny decisions that aren't the defaults in many libraries, so let's break down 1 simple graph by @jburnmurdoch.bsky.social 🔗 www.ft.com/content/73a1...
921577458
Reposted by Jay Alammar
Maarten Grootendorst @maartengr.bsky.social · 18/11/2024
🍿 Introducing the animated "Visual Guide to Mixture of Experts (MoE)"! This was a blast to make and contains more in-depth descriptions than the original already had! Expect even more intuition as we break down visuals and discover the nuances behind MoE. www.youtube.com/watch?v=sOPD...
youtube.com
A Visual Guide to Mixture of Experts (MoE) in LLMs
YouTube video by Maarten Grootendorst
043
Jay Alammar @jayalammar.bsky.social · 20/11/2024
I loved Daniel Dennett's From Bacteria to Bach and Back and its analogies between biological minds and computing paradigms. Chapter 4, especially, which speaks about how an intelligent being can have "competence without comprehension".
040