Sign in

Eugene Yan

@eugeneyan.com
11K followers 331 following 455 posts

RecSys, AI, Engineering; Principal Applied Scientist @ Amazon. Led ML @ Alibaba, Lazada, Healthtech Series A. Writing @ eugeneyan.com, aiteratelabs.com.

PostsRepliesMedia
Reposted by Eugene Yan
Ethan @ethanrosenthal.com · 18/09/2025
I’ve seen semantic IDs pop up but never bothered to actually look into them. This write up from @eugeneyan.com is a great intro that also illustrates why they’re pretty interesting for mixing recsys and LLMs eugeneyan.com/writing/sema...
eugeneyan.com
How to Train an LLM-RecSys Hybrid for Steerable Recs with Semantic IDs
An LLM that can converse in English & item IDs, and make recommendations w/o retrieval or tools.
091
Eugene Yan @eugeneyan.com · 17/09/2025
I've been nerdsniped by the idea of Semantic IDs. Here's the result of my training runs: • RQ-VAE to compress item embeddings into tokens • SASRec to predict the next item (i.e., 4-tokens) exactly • Qwen3-8B that can return recs and natural language! eugeneyan.com/writing/sema...
eugeneyan.com
How to Train an LLM-RecSys Hybrid for Steerable Recs with Semantic IDs
An LLM that can converse in English & item IDs, and make recommendations w/o retrieval or tools.
2256
Reposted by Eugene Yan
Eugene Yan @eugeneyan.com · 25/06/2025
Wrote an intro to evals for long-context Q&A systems: • How it differs from basic Q&A • What dimensions & metrics to eval on • How to build llm-evaluators • How to build eval datasets • Benchmarks: narratives, technical docs, multi-docs eugeneyan.com/writing/qa-e...
eugeneyan.com
Evaluating Long-Context Question & Answer Systems
Evaluation metrics, how to build eval datasets, eval methodology, and a review of several benchmarks.
0174
Eugene Yan @eugeneyan.com · 25/06/2025
Wrote an intro to evals for long-context Q&A systems: • How it differs from basic Q&A • What dimensions & metrics to eval on • How to build llm-evaluators • How to build eval datasets • Benchmarks: narratives, technical docs, multi-docs eugeneyan.com/writing/qa-e...
eugeneyan.com
Evaluating Long-Context Question & Answer Systems
Evaluation metrics, how to build eval datasets, eval methodology, and a review of several benchmarks.
0174
Eugene Yan @eugeneyan.com · 21/05/2025
Some thoughts on leadership: eugeneyan.com/writing/lead... • What makes an exceptional leader? • What do exceptional leaders do? • Leadership styles: Commando, soldier, police
1100
Eugene Yan @eugeneyan.com · 20/05/2025
The best leaders I’ve worked with operate with perma-urgency. They act like early founders, mindful of existential threats. And they can balance speed, sustainability, and repay tech debt. Ultimately, customers love it and teams thrive when we ship fast to deliver delight.
170
Eugene Yan @eugeneyan.com · 18/05/2025
Had a fun couple of hours this weekend with Codex & Windsurf • Migrated off deprecated jekyll-algolia to official sdk (better indexing) • Added recommendations + relevance scores to each post • Improved site responsiveness; fixed dark mode flicker • Marie Kondo-ed unused files & dead code
Image of recommender widget at the bottom of posts on eugeneyan.com
151
Eugene Yan @eugeneyan.com · 14/05/2025
In orgs pushing the envelope, there's always a minority that can be counted on to get shit done against all odds, driven by force of will, resourcefulness, influence, etc. When you identify them, vest in them authority, autonomy, and step back and watch them perform miracles.
150
Eugene Yan @eugeneyan.com · 07/05/2025
To better understand MCPs and agentic workflows, I built news-agents to generate a daily news recap. The main agent spawns sub-agents, assigning them news feeds to parse and summarize, and then generates a final overall summary plus analysis. eugeneyan.com/writing/news...
eugeneyan.com
Building News Agents for Daily News Recaps with MCP, Q, and tmux
Learning to automate simple agentic workflows with Amazon Q CLI, Anthropic MCP, and tmux.
4224
Eugene Yan @eugeneyan.com · 30/04/2025
@hamel.bsky.social & @sh-reya.bsky.social are two of the world's best on evals. They've built evals for 35+ AI apps & helped teams ship confidently. Now they'll teach everything they know on building evals that work. Enrollment closes in 4 days. Secret 35% discount code: maven.com/parlance-lab...
Effective Evals for AI products
042
Eugene Yan @eugeneyan.com · 28/04/2025
The Art of Doing Science and Engineering: Learning to Learn by Richard Hamming only $1.99 for the Kindle version today: amazon.com/dp/B088TMLQDC
The Art of Doing Science and Engineering: Learning to Learn by Richard Hamming
080
Reposted by Eugene Yan
Harrison Pim @harrisonpim.com · 26/04/2025
Enjoyed this on eval-driven product development from @eugeneyan.com. It chimes with my own experiences building around LLMs and search engines, including the thoughts on automated evaluators. When deconstructed, EDD is just the good old scientific method under a new name
eugeneyan.com
An LLM‑as‑Judge Won't Save The Product—Fixing Your Process Will
Applying the scientific method, building via eval-driven development, and monitoring AI output.
072
Eugene Yan @eugeneyan.com · 26/04/2025
Surround yourself with people whose "work" is their calling, craft, and play. They are intrinsically motivated, are driven to excel and do what's right, and and get so much shit done just because it's fun.
0130
Reposted by Eugene Yan
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 26/04/2025
Some of the anti-AI stuff feels a bit like when people would say "don't use Wikipedia as a source." It's just like anything else, a piece of information that you weigh against multiple sources and your own understanding of its likely failure modes
55444838
Eugene Yan @eugeneyan.com · 23/04/2025
Product evals are misunderstood. Many teams think that adding another tool, metric, or llm-as-judge will solve all their problems and save their product. But that just dodges the hard truth and avoids the real work. Here's how to fix your process instead. eugeneyan.com/writing/eval...
eugeneyan.com
An LLM‑as‑Judge Won't Save Your Product—Fixing Your Process Will
Applying the scientific method, building via eval-driven development, and monitoring AI output.
1202
Eugene Yan @eugeneyan.com · 19/04/2025
The default state of projects is to drift toward entropy; you need to actively resist & reverse it.
1111
Eugene Yan @eugeneyan.com · 18/04/2025
Interesting paper from Google that challenges a core assumption in translation evaluation—a single metric can measure both accuracy & naturalness. They found that the best systems had neural metrics that did not correlate with human preferences. arxiv.org/abs/2503.24013
arxiv.org
You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation
The goal of translation, be it by human or by machine, is, given some text in a source language, to produce text in a target language that simultaneously 1) preserves the meaning of the source text an...
1132
Eugene Yan @eugeneyan.com · 17/04/2025
Had a session with very senior folks on how they build with AI and can’t help thinking there’s no better time to learn, clarify, brainstorm, write, debate, plan, design, code, debug, review, analyze, delegate, play, and in general do more more while doing less with AI—so psyched!
270
Eugene Yan @eugeneyan.com · 17/04/2025
Great list of what the best devs do, such as: • Read the source, docs, error msgs • Simplify problems, write simple code • Get their hands dirty • Write to share & write well • Have beginner's mind & keep learning • Not afraid to say: I don't know endler.dev/2025/best-pr...
endler.dev
The Best Programmers I Know | Matthias Endler
I have met a lot of developers in my life. Late…
0201
Eugene Yan @eugeneyan.com · 16/04/2025
@hamel.bsky.social & his wisdom on evals, error analysis, looking at your data is what we need. Here are his 10 Don'ts: • Don't skip error analysis • Don't skip looking at your data • Don't gatekeep who can write prompts • Don't let zero users be a roadblock • Don't be blindsided by criteria drift
1153
Eugene Yan @eugeneyan.com · 15/04/2025
Great example of generate -> validate loop + error analysis > "the most effective route to improve outcomes was brute force: retry steps until they passed or reached a limit. We give the validation errors ... to the LLM and built a loop runner"
airbnb generate-validate loop
1101
Reposted by Eugene Yan
Sarah Drasner @sarahedo.bsky.social · 13/04/2025
This is a great list, things that “the best engineers I know” do, stuff like: - understanding things deeply, reading the actual source - being willing to help other people - status doesn’t matter, good ideas come from anywhere endler.dev/2025/best-pr...
endler.dev
The Best Programmers I Know | Matthias Endler
I have met a lot of developers in my life. Late…
726040
Eugene Yan @eugeneyan.com · 12/04/2025
Stumbled on the first(?) RAG in NarrativeQA from 2017. Because books & movies were too large for LSTMs to do Q&A on, they embedded 200-word chunks and retrieved similar snippets to answer questions. "Chunking and cosine similarity retrieval is so 2017." arxiv.org/abs/1712.07040
4.3 Neural Benchmarks on Stories  The design of the NarrativeQA dataset makes the straight-forward application of the existing neural architectures computationally infeasible, as this would require running an recurrent neural network on sequences of hundreds of thousands of time steps or computing a distribution over the entire input for attention, as is common.  We split the task into two steps: first, we retrieve a small number of relevant passages from the story using an IR system, and subsequently, apply one ofthe neural models above on the resulting document. The question becomes the query for retrieval. This IR problem is much harder that traditional document retrieval, as the documents, the passages here, are very similar, and the question is short and entities mentioned likely occur many times in the story. Our retrieval system considers chunks of 200 words from story and computes representations for all chunks and the query. We then select a varying number of such chunks based on their similarity to the query. We experiment with different representations and similarity measures in Section 5. Finally, we concatenate the selected chunks in the correct temporal order and insert delimiters between them to obtain a much shorter document. For span prediction models, we then further select a span from the retrieved chunks as described in Section 4.2.
0171
Eugene Yan @eugeneyan.com · 09/04/2025
If you were building a Q&A feature (or chatbot) based on very long documents (like books), what evals would you focus on?
If you were building a Q&A feature (or chatbot) based on very long documents (like books), what evals would you focus on?

1. Two metrics that come to mind
• Faithfulness: Grounding of answers in document's content. Not to be confused with correctness—an answer can be correct (based on updated information) but not faithful to the document. Sub-metric: Precision of citations
• Helpfulness: Usefulness (directly addresses the question with enough detail and explanation) and completeness (does not omit important details); an answer can be faithful but not helpful if too brief or doesn't answer the question
• Evaluate separately: Faithfulness = binary label -> LLM-evaluator; Helpfulness = pairwise comparisons -> reward model
2180
Eugene Yan @eugeneyan.com · 08/04/2025
Can't wait for when I can vibe code a production recommender system. Until then, here's some system designs: • Retrieval vs. Ranking: eugeneyan.com/writing/syst... • Real-time retrieval: eugeneyan.com/writing/real... • Personalization: eugeneyan.com/writing/patt...
Two-stage recommender systemZalando's recommender systemUnified embeddingsDoordash's search system
1484
Eugene Yan @eugeneyan.com · 05/04/2025
Your favorite AI writer's favorite AI writer
To Eugene, 

My favorite AI writer

Chip Huyen
1160
Eugene Yan @eugeneyan.com · 02/04/2025
I often get asked: How did I start writing? Why do I write? Who do I write for? What's my process? I procrastinated on this because, honestly, who cares about my writing process? But after repeatedly answering the same qns, I finally wrote this. eugeneyan.com/writing/writ...
eugeneyan.com
Frequently Asked Questions about My Writing Process
How I started, why I write, who I write for, how I write, and more.
2214
Eugene Yan @eugeneyan.com · 26/03/2025
Vibe coding works—but only if you know what good code looks like.
2232
Eugene Yan @eugeneyan.com · 21/03/2025
What are your favorite resources on translating long documents with LLMs? Generating -> validating -> regenerating translations, etc. Also, identifying defects where gender pronouns, formality, idioms, etc are mistranslated, such as for Spanish, German, etc. Please share! 🙏
051
Eugene Yan @eugeneyan.com · 19/03/2025
Search & recsys have historically drawn inspiration from language modeling, such as Word2vec for item embeddings, and GRUs and BERT for sequential models. LLMs are no different and have led to updates in model architecture, training data & paradigms, etc. eugeneyan.com/writing/recs...
eugeneyan.com
Improving Recommendation Systems & Search in the Age of LLMs
Model architectures, data generation, training paradigms, and unified frameworks inspired by LLMs.
2334
Reposted by Eugene Yan
Nathan Lambert @natolambert.bsky.social · 03/03/2025
If you look at most of the models we've received from OpenAI, Anthropic, and Google in the last 18 months you'll hear a lot of "Most of the improvements were in the post-training phase." Here's a simple analogy for how so many gains can be made on mostly the same base model:
1284
Eugene Yan @eugeneyan.com · 02/03/2025
Been querying gpt-4.5 and it's better in ways we can't quantify yet: creativity, humor, world knowledge, wisdom, nuance, based, etc. Excited about how we'll discover new ways to evaluate gpt-4.5 on these aspects which will also transfer to product / application related evals
My reaction is that there is an evaluation crisis. I don't really know what metrics to look at right now. 
MMLU was a good and useful for a few years but that's long over.
SWE-Bench Verified (real, practical, verified problems) I really like and is great but itself too narrow.
Chatbot Arena received so much focus (partly my fault?) that LLM labs have started to really overfit to it, via a combination of prompt mining (from API requests), private evals bombardment, and, worse, explicit use of rankings as training supervision. I think it's still ~ok and there's a lack of "better", but it feels on decline in signal.
There's a number of private evals popping up, an ensemble of which might be one promising path forward.
In absence of great comprehensive evals I tried to turn to vibe checks instead, but I now fear they are misleading and there is too much opportunity for confirmation bias, too low sample size, etc., it's just not great.

TLDR my reaction is I don't really know how good these models are right now.
3191
Eugene Yan @eugeneyan.com · 28/02/2025
At dinner with some builders ytd, we shared how we're using AI Agents. This got me curious about how you're using Agents—please share so I can learn from you🙏 I'll start: I'm using deepresearch more for knowledge discovery/decisions and am trying openhands for feature development. What about you?
271
Eugene Yan @eugeneyan.com · 27/02/2025
cdn.openai.com/gpt-4-5-syst...
https://cdn.openai.com/gpt-4-5-system-card.pdf
1142
Eugene Yan @eugeneyan.com · 26/02/2025
anthropic = protoss; openi = terran; meta = zerg
3140
Eugene Yan @eugeneyan.com · 26/02/2025
agent ≈ model + tools, within a for-loop + environment
slide from anthropic talkslide on openai agent definition from swyx talk
2234
Eugene Yan @eugeneyan.com · 22/02/2025
one of the greatest joys in life is to be with loved ones that you love more than yourself
051
Eugene Yan @eugeneyan.com · 21/02/2025
now that i'm back in seattle, nothing beats the lullaby of my wife and dog's gentle snoring while in my own bed
060
Reposted by Eugene Yan
Jaana Dogan ヤナ ドガン @rakyll.org · 02/02/2025
Almost all major projects I worked on started by 2-3 people sitting in the same meeting room and coding while debating for hours every day without any other distraction.
3775
Reposted by Eugene Yan
Eugene Yan @eugeneyan.com · 31/01/2025
What industrial recsys papers have you enjoyed or found useful in the past year or two? Sharing my list: # 1. Integrating LLMs into recsys 1.1. LLM-augmented recommenders • Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations - arxiv.org/abs/2306.08121
1467
Eugene Yan @eugeneyan.com · 31/01/2025
What industrial recsys papers have you enjoyed or found useful in the past year or two? Sharing my list: # 1. Integrating LLMs into recsys 1.1. LLM-augmented recommenders • Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations - arxiv.org/abs/2306.08121
1467
Eugene Yan @eugeneyan.com · 28/01/2025
♥️ it's tricky to separate what i do on the job (at the bookstore i work at) and what i hack on in my personal time. out of abundance of caution, to not discuss possible proprietary info, i won't be sharing more about the backend of aireadingclub.com 😔
👋 Hi there! I work on <redacted> at Airbnb. I really enjoy your writing.

As a personal project I started trying to write something kind of like AI reading club. I wanted a more powerful version of the "x ray" feature in Kindle, because sometimes when I pick up a new book in a series or return to a book after a break, I cannot remember all of the characters or the plot. I work mainly on <redacted> and not so much app development, and I really struggled with how to build my AI x ray. AI reading club is awesome and I was curious about how you built the retrieval pipeline, any preprocessing you did to the text, etc.

Thanks for publishing so much great work. My team and many of the machine learning engineers whom I support at Airbnb frequently share your posts.

Best regards
1100
Eugene Yan @eugeneyan.com · 22/01/2025
Thanks to the hundreds of readers who've tried aireadingclub.com and interacted with Dewey. If you've tried aireadingclub and have feedback, feature ideas, or thoughts on how AI can help you get more out of reading, please comment or dm me 🙏
Books on AI Reading Club and the number of messages on them.
060
Eugene Yan @eugeneyan.com · 17/01/2025
Few understand this jvns.ca/blog/2014/06...
jvns.ca
Machine learning isn't Kaggle competitions
Machine learning isn't Kaggle competitions
2214
Eugene Yan @eugeneyan.com · 17/01/2025
> Nobody tells you the variables you should be regressing. What's the target? What's the source? Do you notice when results are rubbish? ... That's why I think you need smart people who appear to do something technically easy but actually not so easy. news.ycombinator.com/item?id=1906...
"...I joined a hedged fund, Renaissance Technologies, I'll make a comment about that. It's funny that I think the most important thing to do on data analysis is to do the simple things right. So, here's a kind of non-secret about what we did at renaissance: in my opinion, our most important statistical tool was simple regression with one target and one independent variable. It's the simplest statistical model you can imagine. Any reasonably smart high school student could do it. Now we have some of the smartest people around, working in our hedge fund, we have string theorists we recruited from Harvard, and they're doing simple regression. Is this stupid and pointless? Should we be hiring stupider people and paying them less? And the answer is no. And the reason is nobody tells you what the variables you should be regressing [are]. What's the target. Should you do a nonlinear transform before you regress? What's the source? Should you clean your data? Do you notice when your results are obviously rubbish? And so on. And the smarter you are the less likely you are to make a stupid mistake. And that's why I think you often need smart people who appear to be doing something technically very easy, but actually usually not so easy.]
[[at] my hedge fund, which was not a very big company, we had 7 Phd's just cleaning data and organizing the databases]"
1211
Eugene Yan @eugeneyan.com · 15/01/2025
How can AI make reading more enjoyable? What would an AI-powered reading experience look like? Over the holidays, I prototyped aireadingclub.com to explore ideas like: • Context understanding • Creating quizzes • Recap the book so far • Look up & summarize a term
aireadingclub.com
AI Reading Club
Your AI-powered reading companion
1242
Eugene Yan @eugeneyan.com · 12/01/2025
• Work hard • Keep learning • Cherish loved ones • Find people who inspire you • Be kind & egoless • Eat healthy, exercise, sleep well • Read & write • Practice gratitude & meditate • Be present • Enjoy food & nature • Don’t sweat the small stuff • Smile =)
3221
Eugene Yan @eugeneyan.com · 11/01/2025
let go of your wants and desires. make space and attention for the present, and for the natural abundance and opportunities to flow into your life.
070
Eugene Yan @eugeneyan.com · 05/01/2025
Thank you, glad it was helpful! Some resources: • Task-specific evals: eugeneyan.com/writing/evals/ • LLM-evaluators: eugeneyan.com/writing/llm-... • Aligning people to data; aligning AI to humans: eugeneyan.com/writing/alig... • Finetuning for hallucination detection: eugeneyan.com/writing/fine...
0130