Sign in

Eugene Yan

@eugeneyan.com
11K followers 331 following 455 posts

RecSys, AI, Engineering; Principal Applied Scientist @ Amazon. Led ML @ Alibaba, Lazada, Healthtech Series A. Writing @ eugeneyan.com, aiteratelabs.com.

PostsRepliesMedia
Reposted by Eugene Yan
Ethan @ethanrosenthal.com · 18/09/2025
I’ve seen semantic IDs pop up but never bothered to actually look into them. This write up from @eugeneyan.com is a great intro that also illustrates why they’re pretty interesting for mixing recsys and LLMs eugeneyan.com/writing/sema...
eugeneyan.com
How to Train an LLM-RecSys Hybrid for Steerable Recs with Semantic IDs
An LLM that can converse in English & item IDs, and make recommendations w/o retrieval or tools.
091
Eugene Yan @eugeneyan.com · 17/09/2025
yeap the plan is to map products to semantic IDs which the model can understand
150
Eugene Yan @eugeneyan.com · 17/09/2025
code for data prep, training the RQ-VAE, SASRec, Qwen3-8B, chatting with the model, etc github.com/eugeneyan/se...
github.com
GitHub - eugeneyan/semantic-ids-llm: Semantic IDs: How to train an LLM-Recommender Hybrid with steerability and reasoning on recommendations.
Semantic IDs: How to train an LLM-Recommender Hybrid with steerability and reasoning on recommendations. - eugeneyan/semantic-ids-llm
060
Eugene Yan @eugeneyan.com · 17/09/2025
demo of the LLM-recommender hybrid returning both semantic IDs & english, and: • steering recs via natural language • explaining the recommendation • naming the bundle of recommendations • multi-turn conversation to get recs watch till the end for the bloopers lol www.youtube.com/watch?v=_0n4...
youtube.com
LLM-Recomender Hybrid with Steerable Recommendations and Reasoning on Recommendations
YouTube video by Eugene Yan
130
Eugene Yan @eugeneyan.com · 17/09/2025
For example, given a sequence of items, it can recommend the next best item. But better than that, you can steer the recommendations with natural language! And it can explain why it gave that recommendation, as well as creatively name recommendation bundles.
100
Eugene Yan @eugeneyan.com · 17/09/2025
I've been nerdsniped by the idea of Semantic IDs. Here's the result of my training runs: • RQ-VAE to compress item embeddings into tokens • SASRec to predict the next item (i.e., 4-tokens) exactly • Qwen3-8B that can return recs and natural language! eugeneyan.com/writing/sema...
eugeneyan.com
How to Train an LLM-RecSys Hybrid for Steerable Recs with Semantic IDs
An LLM that can converse in English & item IDs, and make recommendations w/o retrieval or tools.
2256
Reposted by Eugene Yan
Eugene Yan @eugeneyan.com · 25/06/2025
Wrote an intro to evals for long-context Q&A systems: • How it differs from basic Q&A • What dimensions & metrics to eval on • How to build llm-evaluators • How to build eval datasets • Benchmarks: narratives, technical docs, multi-docs eugeneyan.com/writing/qa-e...
eugeneyan.com
Evaluating Long-Context Question & Answer Systems
Evaluation metrics, how to build eval datasets, eval methodology, and a review of several benchmarks.
0174
Eugene Yan @eugeneyan.com · 25/06/2025
Wrote an intro to evals for long-context Q&A systems: • How it differs from basic Q&A • What dimensions & metrics to eval on • How to build llm-evaluators • How to build eval datasets • Benchmarks: narratives, technical docs, multi-docs eugeneyan.com/writing/qa-e...
eugeneyan.com
Evaluating Long-Context Question & Answer Systems
Evaluation metrics, how to build eval datasets, eval methodology, and a review of several benchmarks.
0174
Eugene Yan @eugeneyan.com · 21/05/2025
Some thoughts on leadership: eugeneyan.com/writing/lead... • What makes an exceptional leader? • What do exceptional leaders do? • Leadership styles: Commando, soldier, police
1100
Eugene Yan @eugeneyan.com · 20/05/2025
For example, Amazon started to implement the first version of Amazon Prime in late 2004 and announced it on February 2 2005, six weeks later. An account of how it came amount and lots of anecdotes here. vox.com/recode/2019/... Also this list: patrickcollison.com/fast
vox.com
The making of Amazon Prime, the internet’s most successful and devastating membership program
An oral history of the subscription service that changed online shopping forever.
020
Eugene Yan @eugeneyan.com · 20/05/2025
The best leaders I’ve worked with operate with perma-urgency. They act like early founders, mindful of existential threats. And they can balance speed, sustainability, and repay tech debt. Ultimately, customers love it and teams thrive when we ship fast to deliver delight.
170
Eugene Yan @eugeneyan.com · 19/05/2025
converted all images to webp and hopefully made the site faster. something i wouldn't have bothered in the past
✅ selfcheckgpt.jpg: 226.15KB → 60.53KB (73.24% reduction)
✅ query-processing.jpg: 95.74KB → 42.84KB (55.25% reduction)
✅ sldc-specialists.jpg: 30.88KB → 12.37KB (59.95% reduction)
✅ feature-store-ad.png: 157.24KB → 57.73KB (63.29% reduction)
✅ llm-patterns-aieng-2023-v0-004.jpg: 132.81KB → 43.50KB (67.24% reduction)
✅ google-user-intent.png: 46.67KB → 24.80KB (46.86% reduction)
✅ quy-nguyen.jpeg: 4.27KB → 1.98KB (53.61% reduction)
✅ fbi-tab2.jpg: 396.10KB → 128.28KB (67.61% reduction)
✅ ey-fastball.png: 4.94KB → 0.46KB (90.71% reduction)
✅ favicon-16x16.png: 0.60KB → 0.26KB (56.96% reduction)
✅ android-chrome-192x192.png: 11.50KB → 2.96KB (74.30% reduction)
✅ apple-touch-icon.png: 10.15KB → 2.55KB (74.88% reduction)
✅ android-chrome-512x512.png: 33.22KB → 6.72KB (79.77% reduction)
✅ favicon-32x32.png: 1.35KB → 0.45KB (66.45% reduction)
✅ 404-8.jpg: 43.90KB → 21.81KB (50.32% reduction)
✅ 404-9.jpg: 42.80KB → 16.86KB (60.60% reduction)
✅ 404.jpg: 25.67KB → 6.37KB (75.20% reduction)
✅ 404-11.jpg: 91.48KB → 13.71KB (85.02% reduction)
✅ 404-10.jpg: 309.60KB → 34.35KB (88.91% reduction)
✅ 404-12.jpg: 68.38KB → 50.75KB (25.79% reduction)
✅ 404-13.jpg: 73.79KB → 38.56KB (47.74% reduction)
✅ 404-14.jpg: 86.34KB → 45.91KB (46.82% reduction)
✅ 404-1.jpg: 25.67KB → 6.37KB (75.20% reduction)
✅ 404-2.jpg: 113.98KB → 67.45KB (40.82% reduction)
✅ 404-3.jpg: 38.06KB → 13.68KB (64.04% reduction)
✅ 404-7.jpg: 55.02KB → 26.54KB (51.77% reduction)
✅ 404-6.jpg: 45.98KB → 15.45KB (66.41% reduction)
✅ 404-4.jpg: 94.63KB → 50.33KB (46.82% reduction)
✅ 404-5.jpg: 48.97KB → 17.69KB (63.87% reduction)

==================================================
SUMMARY STATISTICS
==================================================
Total files converted: 1002
Total original size: 122.78MB
Total WebP size: 38.75MB
Total size reduction: 84.03MB (68.44%)
Average size reduction per file: 68.44%
Proportional savings: 122.78MB → 38.75MB
==================================================
030
Eugene Yan @eugeneyan.com · 18/05/2025
Previously, these tasks weren't worth the effort but now they can be done in hours. What an amazing time to build and play =D
140
Eugene Yan @eugeneyan.com · 18/05/2025
Had a fun couple of hours this weekend with Codex & Windsurf • Migrated off deprecated jekyll-algolia to official sdk (better indexing) • Added recommendations + relevance scores to each post • Improved site responsiveness; fixed dark mode flicker • Marie Kondo-ed unused files & dead code
Image of recommender widget at the bottom of posts on eugeneyan.com
151
Eugene Yan @eugeneyan.com · 14/05/2025
In orgs pushing the envelope, there's always a minority that can be counted on to get shit done against all odds, driven by force of will, resourcefulness, influence, etc. When you identify them, vest in them authority, autonomy, and step back and watch them perform miracles.
150
Eugene Yan @eugeneyan.com · 07/05/2025
opps! thanks for letting me know, fixed!
bsky share button
030
Eugene Yan @eugeneyan.com · 07/05/2025
p.s., If you’re interested in topics like this, my friends Ben and Swyx are organizing the AI Engineer World’s Fair in San Francisco on 3rd - 5th June. Come talk to builders deploying AI systems in production. Here’s a big discount for tickets: ti.to/software-3/a...
ti.to
AI Engineer World's Fair 2025
The AI Engineer World's Fair is the biggest technical AI event of the year, happening Summer 2025, the one place you can meet with ~every major AI lab from OpenAI to Anthropic to Cohere, every AI infr...
010
Eugene Yan @eugeneyan.com · 07/05/2025
Here's the code for the mcp-server (src), prompts (context), and generated summaries from May 4th (summaries). github.com/eugeneyan/ne...
github.com
GitHub - eugeneyan/news-agents: Building News Agents to Summarize News with MCP, Q, and tmux
Building News Agents to Summarize News with MCP, Q, and tmux - eugeneyan/news-agents
110
Eugene Yan @eugeneyan.com · 07/05/2025
Here's a three-minute demo of news-agents in action. It's pretty cool at the 30-second mark how the sub-agents get spawned! We then see the main agent assigning tasks and polling for progress, and finally shutting the sub-agents down when they're done with their assigned tasks.
130
Eugene Yan @eugeneyan.com · 07/05/2025
To better understand MCPs and agentic workflows, I built news-agents to generate a daily news recap. The main agent spawns sub-agents, assigning them news feeds to parse and summarize, and then generates a final overall summary plus analysis. eugeneyan.com/writing/news...
eugeneyan.com
Building News Agents for Daily News Recaps with MCP, Q, and tmux
Learning to automate simple agentic workflows with Amazon Q CLI, Anthropic MCP, and tmux.
4224
Eugene Yan @eugeneyan.com · 30/04/2025
@hamel.bsky.social & @sh-reya.bsky.social are two of the world's best on evals. They've built evals for 35+ AI apps & helped teams ship confidently. Now they'll teach everything they know on building evals that work. Enrollment closes in 4 days. Secret 35% discount code: maven.com/parlance-lab...
Effective Evals for AI products
042
Eugene Yan @eugeneyan.com · 28/04/2025
The Art of Doing Science and Engineering: Learning to Learn by Richard Hamming only $1.99 for the Kindle version today: amazon.com/dp/B088TMLQDC
The Art of Doing Science and Engineering: Learning to Learn by Richard Hamming
080
Reposted by Eugene Yan
Harrison Pim @harrisonpim.com · 26/04/2025
Enjoyed this on eval-driven product development from @eugeneyan.com. It chimes with my own experiences building around LLMs and search engines, including the thoughts on automated evaluators. When deconstructed, EDD is just the good old scientific method under a new name
eugeneyan.com
An LLM‑as‑Judge Won't Save The Product—Fixing Your Process Will
Applying the scientific method, building via eval-driven development, and monitoring AI output.
072
Eugene Yan @eugeneyan.com · 26/04/2025
Surround yourself with people whose "work" is their calling, craft, and play. They are intrinsically motivated, are driven to excel and do what's right, and and get so much shit done just because it's fun.
0130
Reposted by Eugene Yan
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 26/04/2025
Some of the anti-AI stuff feels a bit like when people would say "don't use Wikipedia as a source." It's just like anything else, a piece of information that you weigh against multiple sources and your own understanding of its likely failure modes
55444838
Eugene Yan @eugeneyan.com · 23/04/2025
because you’ll keep discovering new ground truth that previous ground truth didn’t cover. in conventional ML, you’d get a ton of ground truth, typically your entire traffic (e.g., search, fraud, forecast errors). now, unless you design your product with implicit feedback, you need to label data
010
Eugene Yan @eugeneyan.com · 23/04/2025
yeap, some human oversight still needed for now
100
Eugene Yan @eugeneyan.com · 23/04/2025
Product evals are misunderstood. Many teams think that adding another tool, metric, or llm-as-judge will solve all their problems and save their product. But that just dodges the hard truth and avoids the real work. Here's how to fix your process instead. eugeneyan.com/writing/eval...
eugeneyan.com
An LLM‑as‑Judge Won't Save Your Product—Fixing Your Process Will
Applying the scientific method, building via eval-driven development, and monitoring AI output.
1202
Eugene Yan @eugeneyan.com · 19/04/2025
did you define what rlvr was in the writeup?
110
Eugene Yan @eugeneyan.com · 19/04/2025
The default state of projects is to drift toward entropy; you need to actively resist & reverse it.
1111
Eugene Yan @eugeneyan.com · 18/04/2025
thanks for sharing this! it’s useful
010
Eugene Yan @eugeneyan.com · 18/04/2025
They argue that translation quality has two dimensions, accuracy (aka fidelity) and naturalness (aka fluency), and optimizing one degrades the other. Thus, they propose evaluating translation on both accuracy and naturalness, instead of aggregating it to a single score.
150
Eugene Yan @eugeneyan.com · 18/04/2025
Interesting paper from Google that challenges a core assumption in translation evaluation—a single metric can measure both accuracy & naturalness. They found that the best systems had neural metrics that did not correlate with human preferences. arxiv.org/abs/2503.24013
arxiv.org
You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation
The goal of translation, be it by human or by machine, is, given some text in a source language, to produce text in a target language that simultaneously 1) preserves the meaning of the source text an...
1132
Eugene Yan @eugeneyan.com · 17/04/2025
it is not
000
Eugene Yan @eugeneyan.com · 17/04/2025
Had a session with very senior folks on how they build with AI and can’t help thinking there’s no better time to learn, clarify, brainstorm, write, debate, plan, design, code, debug, review, analyze, delegate, play, and in general do more more while doing less with AI—so psyched!
270
Eugene Yan @eugeneyan.com · 17/04/2025
Great list of what the best devs do, such as: • Read the source, docs, error msgs • Simplify problems, write simple code • Get their hands dirty • Write to share & write well • Have beginner's mind & keep learning • Not afraid to say: I don't know endler.dev/2025/best-pr...
endler.dev
The Best Programmers I Know | Matthias Endler
I have met a lot of developers in my life. Late…
0201
Eugene Yan @eugeneyan.com · 16/04/2025
• Don't use arbitrary scales; use binary instead • Don't skip measuring alignment btw AI & human • Don't overlook experiments as wins • Don't go radio silent with stakeholders • Don't hide the failures www.oreilly.com/radar/a-fiel...
oreilly.com
A Field Guide to Rapidly Improving AI Products
Evaluation Methods, Data-Driven Improvement, and Experimentation Techniques from 30+ Production Implementations
070
Eugene Yan @eugeneyan.com · 16/04/2025
@hamel.bsky.social & his wisdom on evals, error analysis, looking at your data is what we need. Here are his 10 Don'ts: • Don't skip error analysis • Don't skip looking at your data • Don't gatekeep who can write prompts • Don't let zero users be a roadblock • Don't be blindsided by criteria drift
1153
Eugene Yan @eugeneyan.com · 15/04/2025
thank for for the feedback! it’s an initial attempt at defining helpfulness as a combination of both and can be improved. all ideas appreciated
110
Eugene Yan @eugeneyan.com · 15/04/2025
... After running this “sample, tune, sweep” loop for 4 days ... completed files from 75% to 97% of the total files, and had just under 100 files remaining" medium.com/airbnb-engin...
medium.com
Accelerating Large-Scale Test Migration with LLMs
How Airbnb migrated nearly 3.5K Enzyme test files to React Testing Library in just 6 weeks using automation and LLMs
030
Eugene Yan @eugeneyan.com · 15/04/2025
> "1. Run failing files to find common issues. 2. Select a sample of files with a common issue. 3. Update prompts and scripts to address that issue. 4. Re-run against the sample of files to validate fix. 5. Repeat by running against remaining files. ...
100
Eugene Yan @eugeneyan.com · 15/04/2025
Great example of generate -> validate loop + error analysis > "the most effective route to improve outcomes was brute force: retry steps until they passed or reached a limit. We give the validation errors ... to the LLM and built a loop runner"
airbnb generate-validate loop
1101
Reposted by Eugene Yan
Sarah Drasner @sarahedo.bsky.social · 13/04/2025
This is a great list, things that “the best engineers I know” do, stuff like: - understanding things deeply, reading the actual source - being willing to help other people - status doesn’t matter, good ideas come from anywhere endler.dev/2025/best-pr...
endler.dev
The Best Programmers I Know | Matthias Endler
I have met a lot of developers in my life. Late…
726040
Eugene Yan @eugeneyan.com · 12/04/2025
Stumbled on the first(?) RAG in NarrativeQA from 2017. Because books & movies were too large for LSTMs to do Q&A on, they embedded 200-word chunks and retrieved similar snippets to answer questions. "Chunking and cosine similarity retrieval is so 2017." arxiv.org/abs/1712.07040
4.3 Neural Benchmarks on Stories  The design of the NarrativeQA dataset makes the straight-forward application of the existing neural architectures computationally infeasible, as this would require running an recurrent neural network on sequences of hundreds of thousands of time steps or computing a distribution over the entire input for attention, as is common.  We split the task into two steps: first, we retrieve a small number of relevant passages from the story using an IR system, and subsequently, apply one ofthe neural models above on the resulting document. The question becomes the query for retrieval. This IR problem is much harder that traditional document retrieval, as the documents, the passages here, are very similar, and the question is short and entities mentioned likely occur many times in the story. Our retrieval system considers chunks of 200 words from story and computes representations for all chunks and the query. We then select a varying number of such chunks based on their similarity to the query. We experiment with different representations and similarity measures in Section 5. Finally, we concatenate the selected chunks in the correct temporal order and insert delimiters between them to obtain a much shorter document. For span prediction models, we then further select a span from the retrieved chunks as described in Section 4.2.
0171
Eugene Yan @eugeneyan.com · 09/04/2025
yes
000
Eugene Yan @eugeneyan.com · 09/04/2025
Links to resources, papers, tech blogs, etc. appreciated 🙏
000
Eugene Yan @eugeneyan.com · 09/04/2025
• MultiDoc2Dial: Modeling dialogues grounded in multiple documents. Evaluates ability to integrate info over multiple docs • Frustratingly Hard Evidence Retrieval for QA Over Books: Reframed NarrativeQA as open-domain task where book text must be retrieved
100
Eugene Yan @eugeneyan.com · 09/04/2025
• L-Eval: 20 tasks and >500 long documents (up to 200k tokens), with several QA-oriented tasks • HELMET: Includes reference-based evaluation for long-context QA, and includes measures for positional robustness
100
Eugene Yan @eugeneyan.com · 09/04/2025
• Qasper: Similar to NarrativeQA, but with academic documents that are 5-10k tokens, and includes evaluation of answer spans • LongBench: Average of 6.7k words across fiction and technical docs • LongBench v2: Extension of LongBench, but evals are MCQ only
100
Eugene Yan @eugeneyan.com · 09/04/2025
4. Benchmark Datasets: • NarrativeQA: Questions based on entire movie scripts or novels. Includes reference answers useful for LLM-eval comparisons • NovelQA: Q&A over full novels; includes both MCQ and free-form responses, and includes references
100