Sign in

Hamel Husain

@hamel.bsky.social
6.7K followers 655 following 129 posts

evals evals evals. evals.info

PostsRepliesMedia
Hamel Husain @hamel.bsky.social · 27/09/2026
I know we are all focused on AI agents But a good human agent (w/AI skills + taste + agency) can transform your business much more. There is unreasonable alpha in finding these human agents maybe now more than ever.
0111
Hamel Husain @hamel.bsky.social · 31/08/2026
Over the summer, @sh-reya.bsky.social and I hosted 13 sessions on AI Engineering topics like retrieval, post-training, inference, and evals. I've summarized all the sessions, organized by theme, with links to the source materials. Enjoy! hamel.dev/notes/llm/ai...
2123
Reposted by Hamel Husain
Jeremy Lewi @jeremy.lewi.us · 29/08/2026
Time to do some evals of #Webmcp and web.runme.dev; first step is kicking of Codex to analyze all my traces over the past two weeks to find example invocations and problem areas. First step in evals is always to look at the data. 🎩 @hamel.bsky.social
web.runme.dev
101
Reposted by Hamel Husain
pamelafox.bsky.social @pamelafox.bsky.social · 24/07/2026
"Do automated evals work?" parlance-labs.com/blog/posts/a... @hamel.bsky.social and @doesdatmaksense.bsky.social use multiple LLM-powered tools to identify failures in app traces and compare results to labels from a domain expert.
Screenshot of results
151
Hamel Husain @hamel.bsky.social · 08/11/2025
👀 Animals have been assigned. Scheduled to print fall 2026! We have iterated on this with over 3k students (and continue to do so). We give our students access to the full draft as part of our evals course (link in bio)
1281
Reposted by Hamel Husain
pamelafox.bsky.social @pamelafox.bsky.social · 31/10/2025
Love that @hamel.bsky.social is putting on a hackathon where the goal is for your agent to score the highest on evaluations, not just do something flashy. click.convertkit-mail2.com/gkumlz753lc5...
043
Reposted by Hamel Husain
Jeremy Lewi @jeremy.lewi.us · 06/10/2025
If your looking to get started with evals check out this cookbook from @hamel.bsky.social cookbook.openai.com/examples/eva...
cookbook.openai.com
Building resilient prompts using an evaluation flywheel | OpenAI Cookbook
This cookbook provides a practical guide on how to use the OpenAI Platform to easily build resilience into your prompts. A resilient prom...
081
Hamel Husain @hamel.bsky.social · 03/10/2025
"Can I just get an LLM to do my error analysis?" We get this question constantly. The answer is no, and trying is the fastest way to miss critical bugs. Full podcast: youtu.be/BsWxPI9UM4c?...
080
Hamel Husain @hamel.bsky.social · 03/10/2025
I recently sat down with Lenny Rachitsky to discuss why AI Evals are becoming the most sought after skill for product builders. As a bonus, we step through an end-to-end example of building an eval in a spreadsheet so everyone can understand. See reply for links.
170
Hamel Husain @hamel.bsky.social · 24/09/2025
This one is going to be spicy. 80% of the time I've seen a graph DB in production, it's been an overcomplicated mess (especially in AI applications). In this talk, Jo and I will discuss when GraphDBs are overkill and when they actually make sense. Sign up here:
maven.com
You Don't Need a Graph DB
Many teams adopt graph databases believing they need specialized tools for relationship data, adding unnecessary complexity to their stack. This session reveals that for most use cases, the…
1160
Reposted by Hamel Husain
Kyle Baxter @kbaxter.bsky.social · 26/08/2025
For technical domains especially, getting non-DSes involved in analyzing outputs is vital. It’s hard to build anything good without it bc v1s almost always have major fail modes. Finding the appropriate system design—let alone optimizing—requires a tight coupling of output analysis and system design
142
Hamel Husain @hamel.bsky.social · 23/08/2025
If you're wanted to learn applied AI evals but not sure if its for you, @sh-reya.bsky.social and I put together something that might help. This free email course compiles what we've learned from teaching 2k+ students. It’s 17 emails plus 2 free e-books. Here's the link: ai.hamel.dev/eval-course
ai.hamel.dev
AI Evals Email Course
A free 17-part email series on the principles of application-centric LLM evals.
181
Reposted by Hamel Husain
Jeremy Lewi @jeremy.lewi.us · 08/08/2025
I've been eval pilled by @hamel.bsky.social . Everyone is all let's build some MCP servers and ship a minimal AISRE as quickly as possible. And I'm writing a design doc about how to build evals with @runme.dev so we can iterate rapidly on the AI
082
Hamel Husain @hamel.bsky.social · 08/06/2025
Last chance to signup for this free lesson with OpenAI on evals, Including a sneak peek of their new eval products! Link: maven.com/p/d2dc30/how...
020
Hamel Husain @hamel.bsky.social · 17/05/2025
Can non-data scientists write AI Evals? The answer is nuanced and not just "Yes". @eugeneyan.com and I discuss this in the context of the "analyze-measure-improve" cycle from our course. Links to more resources in the reply
180
Hamel Husain @hamel.bsky.social · 12/05/2025
If you are writing evals without error analysis, our course AI Evals for Engineers & PMs is for you. Begins monday next week. Full syllabus in this link: maven.com/parlance-lab...
040
Hamel Husain @hamel.bsky.social · 08/05/2025
It is very easy to make mistakes when creating evals for your AI product. @sh-reya.bsky.social and I run through the most common errors in this talk. 35% discount code to our upcoming course in the video notes youtu.be/GL0XhAj5LPE?...
youtu.be
LLM Evals: Common Mistakes
YouTube video by Hamel Husain
051
Reposted by Hamel Husain
Hamel Husain @hamel.bsky.social · 29/04/2025
GitHub CoPilot is one of the first commercially successful LLM products (predating ChatGPT). What was the secret? A robust eval suite! In this lightning lesson, John Berryman will reveal the eval techniques (and mistakes) from working on this product maven.com/p/da8264/how...
2265
Reposted by Hamel Husain
Eugene Yan @eugeneyan.com · 30/04/2025
@hamel.bsky.social & @sh-reya.bsky.social are two of the world's best on evals. They've built evals for 35+ AI apps & helped teams ship confidently. Now they'll teach everything they know on building evals that work. Enrollment closes in 4 days. Secret 35% discount code: maven.com/parlance-lab...
Effective Evals for AI products
042
Hamel Husain @hamel.bsky.social · 29/04/2025
GitHub CoPilot is one of the first commercially successful LLM products (predating ChatGPT). What was the secret? A robust eval suite! In this lightning lesson, John Berryman will reveal the eval techniques (and mistakes) from working on this product maven.com/p/da8264/how...
2265
Reposted by Hamel Husain
Hamel Husain @hamel.bsky.social · 27/04/2025
I keep hearing about the emerging role of AI PM. How is this any different than a normal PM? Is it hype? We are gonna find out in this free lightning lesson. I will ask difficult questions. With @schof.bsky.social and Aman Khan maven.com/p/544677/wha...
151
Hamel Husain @hamel.bsky.social · 27/04/2025
I keep hearing about the emerging role of AI PM. How is this any different than a normal PM? Is it hype? We are gonna find out in this free lightning lesson. I will ask difficult questions. With @schof.bsky.social and Aman Khan maven.com/p/544677/wha...
151
Reposted by Hamel Husain
Kyle @kylestratis.com · 26/04/2025
As genAI projects mature, proper evals are becoming table stakes for production deployment. But how do we evaluate probabilistic machines? Looking forward to learning about the latest techniques and best practices from @hamel.bsky.social and Shreya Shankar next month!
maven.com
AI Evals For Engineers & PMs by Hamel Husain and Shreya Shankar on Maven
Learn proven approaches for quickly improving AI applications. Build AI that works better than the competition, regardless of the use-case.
011
Hamel Husain @hamel.bsky.social · 16/04/2025
Last chance to sign up for this. Recording sent to everyone who signs up. maven.com/p/29a33a/hyb...
040
Reposted by Hamel Husain
Eugene Yan @eugeneyan.com · 16/04/2025
@hamel.bsky.social & his wisdom on evals, error analysis, looking at your data is what we need. Here are his 10 Don'ts: • Don't skip error analysis • Don't skip looking at your data • Don't gatekeep who can write prompts • Don't let zero users be a roadblock • Don't be blindsided by criteria drift
1153
Reposted by Hamel Husain
Hamel Husain @hamel.bsky.social · 14/04/2025
If you are building RAG applications, you don't want to miss this. Doug Turnbull is going to show you his tricks he's learned from a decade of optimizing retrieval in search systems, and how that transfers to RAG. Link: maven.com/p/29a33a/hyb...
0194
Hamel Husain @hamel.bsky.social · 14/04/2025
If you are building RAG applications, you don't want to miss this. Doug Turnbull is going to show you his tricks he's learned from a decade of optimizing retrieval in search systems, and how that transfers to RAG. Link: maven.com/p/29a33a/hyb...
0194
Reposted by Hamel Husain
Hamel Husain @hamel.bsky.social · 10/04/2025
The most critical part of RAG is the R (Retrieval). In this lesson, Doug Turnbull will share how we can go beyond simple hybrid search to optimize retrieval. He'll share his bag of tricks from over a decade of optimizing search systems. maven.com/p/29a33a/hyb...
maven.com
Hybrid Search Is Just The Beginning: Optimizing the R in RAG
You may have implemented hybrid search, and that's a great first step. In this session, Doug will share his experience building advanced search systems at Shopify and Reddit to provide you with a clea...
1242
Hamel Husain @hamel.bsky.social · 10/04/2025
The most critical part of RAG is the R (Retrieval). In this lesson, Doug Turnbull will share how we can go beyond simple hybrid search to optimize retrieval. He'll share his bag of tricks from over a decade of optimizing search systems. maven.com/p/29a33a/hyb...
maven.com
Hybrid Search Is Just The Beginning: Optimizing the R in RAG
You may have implemented hybrid search, and that's a great first step. In this session, Doug will share his experience building advanced search systems at Shopify and Reddit to provide you with a clea...
1242
Hamel Husain @hamel.bsky.social · 07/04/2025
I’m excited to teach this lesson on improving retrieval for RAG with @softwaredoug.bsky.social maven.com/p/29a33a/hyb...
maven.com
Hybrid Search Is Just The Beginning: Optimizing the R in RAG
You may have implemented hybrid search, and that's a great first step. In this session, Doug will share his experience building advanced search systems at Shopify and Reddit to provide you with a clea...
081
Reposted by Hamel Husain
Anthony Panozzo @panozzaj.com · 05/04/2025
Interesting post: A Field Guide to Rapidly Improving AI Products by @hamel.bsky.social
hamel.dev
A Field Guide to Rapidly Improving AI Products – Hamel’s Blog
Evaluation methods, data-driven improvement, and experimentation techniques from 30+ production implementations.
032
Reposted by Hamel Husain
Rasmus Aagaard @rasgaard.com · 26/03/2025
Whenever @hamel.bsky.social drops a new banger of a blog post I make sure to forward it to the rest of my team 🚀 And I encourage everyone working with productionalizing LLM systems to do the same. hamel.dev/blog/posts/f...
hamel.dev
A Field Guide to Rapidly Improving AI Products – Hamel’s Blog
Evaluation methods, data-driven improvement, and experimentation techniques from 30+ production implementations.
173
Reposted by Hamel Husain
Scott H. Hawley @drscotthawley.bsky.social · 07/02/2025
Finally got around to trying Answer.ai 's "nbsanity": Works beautifully on the first try! Even renders my interactive Plotly stuff! Just replace "github" in your notebook URL with "nbsanity", as in... nbsanity.com/static/3465a... More info from @hamel.bsky.social : www.answer.ai/posts/2024-1...
nbsanity.com
problematicimageanalyzer
nbsanity: A modern way to view public Jupyter notebooks on GitHub
071
Reposted by Hamel Husain
Nico Ritschel @nicoritschel.com · 19/01/2025
In case you missed it, @hamel.bsky.social reviewed Devin. It succeeded on 3/20 assigned tasks.
answer.ai
Thoughts On A Month With Devin – Answer.AI
Our impressions of Devin after giving it 20+ tasks.
2111
Reposted by Hamel Husain
Kevin Markham @dataschool.io · 17/01/2025
Thoughts On A Month With Devin (the "AI software engineer") by @hamel.bsky.social "Out of 20 tasks we attempted, we saw 14 failures, 3 inconclusive results, and just 3 successes. More concerning was our inability to predict which tasks would succeed."
answer.ai
Thoughts On A Month With Devin – Answer.AI
Our impressions of Devin after giving it 20+ tasks.
0105
Reposted by Hamel Husain
Matthew Mullins @mmullins.coginiti.co · 18/01/2025
Enjoyed the systematic first hand reporting of their experience using Devin by @hamel.bsky.social and team. If you’ve worked with llm coding assistants, the results aren’t surprising, but it points to how far these models still need to go and should be worrying for how effective “agents” will be.
answer.ai
Thoughts On A Month With Devin – Answer.AI
Our impressions of Devin after giving it 20+ tasks.
193
Reposted by Hamel Husain
Sung Kim @sungkim.bsky.social · 17/01/2025
Thoughts On A Month With Devin by @hamel.bsky.social They decided to put it through its paces, testing it against a wide range of real-world tasks. This is their story - a thorough, real-world attempt to work with one of the most hyped AI products of 2024. www.answer.ai/posts/2025-0...
answer.ai
Thoughts On A Month With Devin – Answer.AI
Our impressions of Devin after giving it 20+ tasks.
3246
Reposted by Hamel Husain
Chris Parsons @chrismdp.com · 04/01/2025
Four steps to use evals effectively in LLM applications (we haven't done the last one but are still getting great results): Eval Driven Development is the new TDD for LLM based applications. Without them, you're flying blind. #cto #llm #ai #tech #dev #genai
163
Hamel Husain @hamel.bsky.social · 22/12/2024
New LLM Eval Office Hours, I discuss the importance of doing error analysis before jumping into metrics and tests Links to notes in the YT description youtu.be/ZEvXvyY17Ys?...
youtu.be
LLM Eval Office Hours #3: The Importance Of Starting With Error Analysis
YouTube video by Hamel Husain
1265
Reposted by Hamel Husain
CJ Sullivan @cjlovesdata1.bsky.social · 20/12/2024
This is pretty damn nifty! @hamel.bsky.social @projectjupyter.bsky.social #datascience #jupyternotebooks www.answer.ai/posts/2024-1...
answer.ai
nbsanity - Share Notebooks as Polished Web Pages in Seconds – Answer.AI
Transform your GitHub Jupyter notebooks into beautiful, readable web pages with a single URL change. No setup required.
2124
Reposted by Hamel Husain
Jeremy Howard @howard.fm · 19/12/2024
ModernBERT is available as a slot-in replacement for any BERT-like model, with both 139M param and 395M param sizes. It has a 8192 sequence length, is extremely efficient, is uniquely great at analyzing code, and much more. Read this for details: huggingface.co/blog/modern...
110014
Reposted by Hamel Husain
Greg Ceccarelli @gregce.bsky.social · 17/12/2024
Our team at @specstory.com launched our very first product iteration today. What is it? An extension for @cursor_ai that allows you to save and share your composer and chat history. Give it a try at marketplace.visualstudio.com/items?itemNa... and let us know what you think!
Never lose your history
063
Hamel Husain @hamel.bsky.social · 16/12/2024
Recoded my second office hours on LLM Evals. We talked about observability and how to prioritize writing tests in complex systems Here are the notes: hamel.dev/notes/llm/of... Video: youtu.be/TZwmLXXFbh4?...
youtu.be
LLM Eval Office Hours #2: LLM Observability
YouTube video by Hamel Husain
2262
Hamel Husain @hamel.bsky.social · 14/12/2024
Running this notebook from @howard.fm hoping that it removes noise from my timeline nbsanity.com/static/0b3fd...
nbsanity.com
Remove bsky non-mutual follows
nbsanity: A modern way to view public Jupyter notebooks on GitHub
2172
Reposted by Hamel Husain
Michael Mullarkey @mcmullarkey.bsky.social · 12/12/2024
Make it easier to manually inspect your data! I built a small Shiny for Python web app as recommended by @hamel.bsky.social. I'm getting through my task much faster than previous iterations
hamel.dev
Curating LLM data – Hamel’s Blog
A review of tools
041
Reposted by Hamel Husain
Phillip Carter @phillipcarter.dev · 11/12/2024
I'm proud that we're going public with some positioning on what @honeycomb.io actually believes AI represents: A new, weird, and sometimes janky kind of virtual computer. Stay tuned for a lot more clear-headed posting on applied AI in the coming year. www.honeycomb.io/blog/observa...
honeycomb.io
Observability in the Age of AI
How will Honeycomb leverage AI in 2025? Find out from Charity Majors and Phillip Carter.
2266
Reposted by Hamel Husain
Simon Willison @simonwillison.net · 08/12/2024
Fantastic use of shot-scraper.datasette.io here to. Create social media cards for this new Jupyter Notebook rendering site nbsanity.com
shot-scraper.datasette.io
shot-scraperContentsMenuExpandLight modeDark modeAuto light/dark mode
1241
Hamel Husain @hamel.bsky.social · 08/12/2024
Super exciting update to nbsanity. I've incorporated @simonwillison.net 's shotscraper. Now, all new renders get a fancy social card! This makes nbsanity a nice microblogging utility Examples of rendered notebooks: 1/3 nbsanity.com/static/6a987...
nbsanity.com
The sqlite-utils tutorial
nbsanity: A modern way to view public Jupyter notebooks on GitHub
1361
Hamel Husain @hamel.bsky.social · 08/12/2024
nbsanity now has a bookmarklet nbsanity.com It's a static server that renders public Jupyter notebooks with Quarto
nbsanity.com
nbsanity | Jupyter Notebook Viewer
A modern way to render and view Jupyter notebooks directly from GitHub
2457
Hamel Husain @hamel.bsky.social · 07/12/2024
Are you frustrated by how GitHub renders Jupyter notebooks? I have public service that renders GitHub notebooks with Quarto nbsanity.com It now works with gists!
76714