Sign in

Hamel Husain

@hamel.bsky.social
6.7K followers 655 following 129 posts

evals evals evals. evals.info

PostsRepliesMedia
Hamel Husain @hamel.bsky.social · 27/09/2026
I know we are all focused on AI agents But a good human agent (w/AI skills + taste + agency) can transform your business much more. There is unreasonable alpha in finding these human agents maybe now more than ever.
0111
Hamel Husain @hamel.bsky.social · 31/08/2026
Consider doing error analysis on these bot replies
000
Hamel Husain @hamel.bsky.social · 31/08/2026
Over the summer, @sh-reya.bsky.social and I hosted 13 sessions on AI Engineering topics like retrieval, post-training, inference, and evals. I've summarized all the sessions, organized by theme, with links to the source materials. Enjoy! hamel.dev/notes/llm/ai...
2133
Reposted by Hamel Husain
Jeremy Lewi @jeremy.lewi.us · 29/08/2026
Time to do some evals of #Webmcp and web.runme.dev; first step is kicking of Codex to analyze all my traces over the past two weeks to find example invocations and problem areas. First step in evals is always to look at the data. 🎩 @hamel.bsky.social
web.runme.dev
101
Reposted by Hamel Husain
pamelafox.bsky.social @pamelafox.bsky.social · 24/07/2026
"Do automated evals work?" parlance-labs.com/blog/posts/a... @hamel.bsky.social and @doesdatmaksense.bsky.social use multiple LLM-powered tools to identify failures in app traces and compare results to labels from a domain expert.
Screenshot of results
151
Hamel Husain @hamel.bsky.social · 08/11/2025
Relevant links - Our course: evals.info - Early release ( just has the TOC & intro now learning.oreilly.com/library/view...
learning.oreilly.com
Evals for AI Engineers
Stop using guesswork to find out how your AI applications are performing. Evals for AI Engineers equips you with the proven tools and processes required to systematically test,... - Selection from Eva...
040
Hamel Husain @hamel.bsky.social · 08/11/2025
👀 Animals have been assigned. Scheduled to print fall 2026! We have iterated on this with over 3k students (and continue to do so). We give our students access to the full draft as part of our evals course (link in bio)
1281
Hamel Husain @hamel.bsky.social · 05/11/2025
Hi
100
Reposted by Hamel Husain
pamelafox.bsky.social @pamelafox.bsky.social · 31/10/2025
Love that @hamel.bsky.social is putting on a hackathon where the goal is for your agent to score the highest on evaluations, not just do something flashy. click.convertkit-mail2.com/gkumlz753lc5...
043
Reposted by Hamel Husain
Jeremy Lewi @jeremy.lewi.us · 06/10/2025
If your looking to get started with evals check out this cookbook from @hamel.bsky.social cookbook.openai.com/examples/eva...
cookbook.openai.com
Building resilient prompts using an evaluation flywheel | OpenAI Cookbook
This cookbook provides a practical guide on how to use the OpenAI Platform to easily build resilience into your prompts. A resilient prom...
081
Hamel Husain @hamel.bsky.social · 03/10/2025
"Can I just get an LLM to do my error analysis?" We get this question constantly. The answer is no, and trying is the fastest way to miss critical bugs. Full podcast: youtu.be/BsWxPI9UM4c?...
080
Hamel Husain @hamel.bsky.social · 03/10/2025
Full podcast. youtu.be/BsWxPI9UM4c?... We are teaching our final course of the year on evals starting this Monday. You can enroll with this link to get 35% off: maven.com/parlance-lab...
youtu.be
Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
Hamel Husain and Shreya Shankar teach the world’s most popular course on AI evals and have trained over 2,000 PMs and engineers (including many teams at OpenAI and Anthropic). In this conversation,…
010
Hamel Husain @hamel.bsky.social · 03/10/2025
I recently sat down with Lenny Rachitsky to discuss why AI Evals are becoming the most sought after skill for product builders. As a bonus, we step through an end-to-end example of building an eval in a spreadsheet so everyone can understand. See reply for links.
170
Hamel Husain @hamel.bsky.social · 25/09/2025
I think it’s more like 99% 🤣 the 1% worked super hard on data modeling
020
Hamel Husain @hamel.bsky.social · 24/09/2025
This one is going to be spicy. 80% of the time I've seen a graph DB in production, it's been an overcomplicated mess (especially in AI applications). In this talk, Jo and I will discuss when GraphDBs are overkill and when they actually make sense. Sign up here:
maven.com
You Don't Need a Graph DB
Many teams adopt graph databases believing they need specialized tools for relationship data, adding unnecessary complexity to their stack. This session reveals that for most use cases, the…
1160
Reposted by Hamel Husain
Kyle Baxter @kbaxter.bsky.social · 26/08/2025
For technical domains especially, getting non-DSes involved in analyzing outputs is vital. It’s hard to build anything good without it bc v1s almost always have major fail modes. Finding the appropriate system design—let alone optimizing—requires a tight coupling of output analysis and system design
142
Hamel Husain @hamel.bsky.social · 23/08/2025
Can you screenshot it and tell me what it is so I can troubleshoot it
100
Hamel Husain @hamel.bsky.social · 23/08/2025
If you're wanted to learn applied AI evals but not sure if its for you, @sh-reya.bsky.social and I put together something that might help. This free email course compiles what we've learned from teaching 2k+ students. It’s 17 emails plus 2 free e-books. Here's the link: ai.hamel.dev/eval-course
ai.hamel.dev
AI Evals Email Course
A free 17-part email series on the principles of application-centric LLM evals.
181
Reposted by Hamel Husain
Jeremy Lewi @jeremy.lewi.us · 08/08/2025
I've been eval pilled by @hamel.bsky.social . Everyone is all let's build some MCP servers and ship a minimal AISRE as quickly as possible. And I'm writing a design doc about how to build evals with @runme.dev so we can iterate rapidly on the AI
082
Hamel Husain @hamel.bsky.social · 08/06/2025
Last chance to signup for this free lesson with OpenAI on evals, Including a sneak peek of their new eval products! Link: maven.com/p/d2dc30/how...
020
Hamel Husain @hamel.bsky.social · 17/05/2025
Link to full talk youtube.com/live/jJMYWfQ...
youtube.com
AI Evals For Engineers: Book Review
YouTube video by Hamel Husain
000
Hamel Husain @hamel.bsky.social · 17/05/2025
We'll be discussing this in our upcoming course on May 19th - AI Evals For Engineers & PMs (This link has a 35% discount code): maven.com/parlance-lab...
maven.com
AI Evals For Engineers & PMs by Hamel Husain and Shreya Shankar on Maven
Learn proven approaches for quickly improving AI applications. Build AI that works better than the competition, regardless of the use-case.
100
Hamel Husain @hamel.bsky.social · 17/05/2025
Can non-data scientists write AI Evals? The answer is nuanced and not just "Yes". @eugeneyan.com and I discuss this in the context of the "analyze-measure-improve" cycle from our course. Links to more resources in the reply
180
Hamel Husain @hamel.bsky.social · 12/05/2025
If you are writing evals without error analysis, our course AI Evals for Engineers & PMs is for you. Begins monday next week. Full syllabus in this link: maven.com/parlance-lab...
040
Hamel Husain @hamel.bsky.social · 08/05/2025
It is very easy to make mistakes when creating evals for your AI product. @sh-reya.bsky.social and I run through the most common errors in this talk. 35% discount code to our upcoming course in the video notes youtu.be/GL0XhAj5LPE?...
youtu.be
LLM Evals: Common Mistakes
YouTube video by Hamel Husain
051
Hamel Husain @hamel.bsky.social · 01/05/2025
I thought this was a meme 🤣 … but it’s real
180
Reposted by Hamel Husain
Hamel Husain @hamel.bsky.social · 29/04/2025
GitHub CoPilot is one of the first commercially successful LLM products (predating ChatGPT). What was the secret? A robust eval suite! In this lightning lesson, John Berryman will reveal the eval techniques (and mistakes) from working on this product maven.com/p/da8264/how...
2265
Reposted by Hamel Husain
Eugene Yan @eugeneyan.com · 30/04/2025
@hamel.bsky.social & @sh-reya.bsky.social are two of the world's best on evals. They've built evals for 35+ AI apps & helped teams ship confidently. Now they'll teach everything they know on building evals that work. Enrollment closes in 4 days. Secret 35% discount code: maven.com/parlance-lab...
Effective Evals for AI products
042
Hamel Husain @hamel.bsky.social · 29/04/2025
GitHub CoPilot is one of the first commercially successful LLM products (predating ChatGPT). What was the secret? A robust eval suite! In this lightning lesson, John Berryman will reveal the eval techniques (and mistakes) from working on this product maven.com/p/da8264/how...
2265
Reposted by Hamel Husain
Hamel Husain @hamel.bsky.social · 27/04/2025
I keep hearing about the emerging role of AI PM. How is this any different than a normal PM? Is it hype? We are gonna find out in this free lightning lesson. I will ask difficult questions. With @schof.bsky.social and Aman Khan maven.com/p/544677/wha...
151
Hamel Husain @hamel.bsky.social · 27/04/2025
I keep hearing about the emerging role of AI PM. How is this any different than a normal PM? Is it hype? We are gonna find out in this free lightning lesson. I will ask difficult questions. With @schof.bsky.social and Aman Khan maven.com/p/544677/wha...
151
Reposted by Hamel Husain
Kyle @kylestratis.com · 26/04/2025
As genAI projects mature, proper evals are becoming table stakes for production deployment. But how do we evaluate probabilistic machines? Looking forward to learning about the latest techniques and best practices from @hamel.bsky.social and Shreya Shankar next month!
maven.com
AI Evals For Engineers & PMs by Hamel Husain and Shreya Shankar on Maven
Learn proven approaches for quickly improving AI applications. Build AI that works better than the competition, regardless of the use-case.
011
Hamel Husain @hamel.bsky.social · 16/04/2025
Last chance to sign up for this. Recording sent to everyone who signs up. maven.com/p/29a33a/hyb...
040
Reposted by Hamel Husain
Eugene Yan @eugeneyan.com · 16/04/2025
@hamel.bsky.social & his wisdom on evals, error analysis, looking at your data is what we need. Here are his 10 Don'ts: • Don't skip error analysis • Don't skip looking at your data • Don't gatekeep who can write prompts • Don't let zero users be a roadblock • Don't be blindsided by criteria drift
1153
Reposted by Hamel Husain
Hamel Husain @hamel.bsky.social · 14/04/2025
If you are building RAG applications, you don't want to miss this. Doug Turnbull is going to show you his tricks he's learned from a decade of optimizing retrieval in search systems, and how that transfers to RAG. Link: maven.com/p/29a33a/hyb...
0194
Hamel Husain @hamel.bsky.social · 14/04/2025
If you are building RAG applications, you don't want to miss this. Doug Turnbull is going to show you his tricks he's learned from a decade of optimizing retrieval in search systems, and how that transfers to RAG. Link: maven.com/p/29a33a/hyb...
0194
Reposted by Hamel Husain
Hamel Husain @hamel.bsky.social · 10/04/2025
The most critical part of RAG is the R (Retrieval). In this lesson, Doug Turnbull will share how we can go beyond simple hybrid search to optimize retrieval. He'll share his bag of tricks from over a decade of optimizing search systems. maven.com/p/29a33a/hyb...
maven.com
Hybrid Search Is Just The Beginning: Optimizing the R in RAG
You may have implemented hybrid search, and that's a great first step. In this session, Doug will share his experience building advanced search systems at Shopify and Reddit to provide you with a clea...
1242
Hamel Husain @hamel.bsky.social · 10/04/2025
Try again lmk
100
Hamel Husain @hamel.bsky.social · 10/04/2025
The most critical part of RAG is the R (Retrieval). In this lesson, Doug Turnbull will share how we can go beyond simple hybrid search to optimize retrieval. He'll share his bag of tricks from over a decade of optimizing search systems. maven.com/p/29a33a/hyb...
maven.com
Hybrid Search Is Just The Beginning: Optimizing the R in RAG
You may have implemented hybrid search, and that's a great first step. In this session, Doug will share his experience building advanced search systems at Shopify and Reddit to provide you with a clea...
1242
Hamel Husain @hamel.bsky.social · 07/04/2025
I’m excited to teach this lesson on improving retrieval for RAG with @softwaredoug.bsky.social maven.com/p/29a33a/hyb...
maven.com
Hybrid Search Is Just The Beginning: Optimizing the R in RAG
You may have implemented hybrid search, and that's a great first step. In this session, Doug will share his experience building advanced search systems at Shopify and Reddit to provide you with a clea...
081
Reposted by Hamel Husain
Anthony Panozzo @panozzaj.com · 05/04/2025
Interesting post: A Field Guide to Rapidly Improving AI Products by @hamel.bsky.social
hamel.dev
A Field Guide to Rapidly Improving AI Products – Hamel’s Blog
Evaluation methods, data-driven improvement, and experimentation techniques from 30+ production implementations.
032
Reposted by Hamel Husain
Rasmus Aagaard @rasgaard.com · 26/03/2025
Whenever @hamel.bsky.social drops a new banger of a blog post I make sure to forward it to the rest of my team 🚀 And I encourage everyone working with productionalizing LLM systems to do the same. hamel.dev/blog/posts/f...
hamel.dev
A Field Guide to Rapidly Improving AI Products – Hamel’s Blog
Evaluation methods, data-driven improvement, and experimentation techniques from 30+ production implementations.
173
Hamel Husain @hamel.bsky.social · 13/02/2025
hamel.dev/notes/llm/of...
hamel.dev
Multi-Turn Chat Evals – Hamel’s Blog
Office hours discussion on multi-turn chat evals
210
Reposted by Hamel Husain
Scott H. Hawley @drscotthawley.bsky.social · 07/02/2025
Finally got around to trying Answer.ai 's "nbsanity": Works beautifully on the first try! Even renders my interactive Plotly stuff! Just replace "github" in your notebook URL with "nbsanity", as in... nbsanity.com/static/3465a... More info from @hamel.bsky.social : www.answer.ai/posts/2024-1...
nbsanity.com
problematicimageanalyzer
nbsanity: A modern way to view public Jupyter notebooks on GitHub
071
Reposted by Hamel Husain
Nico Ritschel @nicoritschel.com · 19/01/2025
In case you missed it, @hamel.bsky.social reviewed Devin. It succeeded on 3/20 assigned tasks.
answer.ai
Thoughts On A Month With Devin – Answer.AI
Our impressions of Devin after giving it 20+ tasks.
2111
Hamel Husain @hamel.bsky.social · 19/01/2025
Sent you an email
010
Reposted by Hamel Husain
Kevin Markham @dataschool.io · 17/01/2025
Thoughts On A Month With Devin (the "AI software engineer") by @hamel.bsky.social "Out of 20 tasks we attempted, we saw 14 failures, 3 inconclusive results, and just 3 successes. More concerning was our inability to predict which tasks would succeed."
answer.ai
Thoughts On A Month With Devin – Answer.AI
Our impressions of Devin after giving it 20+ tasks.
0105
Reposted by Hamel Husain
Matthew Mullins @mmullins.coginiti.co · 18/01/2025
Enjoyed the systematic first hand reporting of their experience using Devin by @hamel.bsky.social and team. If you’ve worked with llm coding assistants, the results aren’t surprising, but it points to how far these models still need to go and should be worrying for how effective “agents” will be.
answer.ai
Thoughts On A Month With Devin – Answer.AI
Our impressions of Devin after giving it 20+ tasks.
193
Hamel Husain @hamel.bsky.social · 18/01/2025
Yes please intro. He may have messaged me looks familiar but I lost it
110
Reposted by Hamel Husain
Sung Kim @sungkim.bsky.social · 17/01/2025
Thoughts On A Month With Devin by @hamel.bsky.social They decided to put it through its paces, testing it against a wide range of real-world tasks. This is their story - a thorough, real-world attempt to work with one of the most hyped AI products of 2024. www.answer.ai/posts/2025-0...
answer.ai
Thoughts On A Month With Devin – Answer.AI
Our impressions of Devin after giving it 20+ tasks.
3246