Sign in

Eugene Yan

@eugeneyan.com
11K followers 331 following 455 posts

RecSys, AI, Engineering; Principal Applied Scientist @ Amazon. Led ML @ Alibaba, Lazada, Healthtech Series A. Writing @ eugeneyan.com, aiteratelabs.com.

PostsRepliesMedia
Eugene Yan @eugeneyan.com · 21/05/2025
Some thoughts on leadership: eugeneyan.com/writing/lead... • What makes an exceptional leader? • What do exceptional leaders do? • Leadership styles: Commando, soldier, police
1100
Eugene Yan @eugeneyan.com · 19/05/2025
converted all images to webp and hopefully made the site faster. something i wouldn't have bothered in the past
✅ selfcheckgpt.jpg: 226.15KB → 60.53KB (73.24% reduction)
✅ query-processing.jpg: 95.74KB → 42.84KB (55.25% reduction)
✅ sldc-specialists.jpg: 30.88KB → 12.37KB (59.95% reduction)
✅ feature-store-ad.png: 157.24KB → 57.73KB (63.29% reduction)
✅ llm-patterns-aieng-2023-v0-004.jpg: 132.81KB → 43.50KB (67.24% reduction)
✅ google-user-intent.png: 46.67KB → 24.80KB (46.86% reduction)
✅ quy-nguyen.jpeg: 4.27KB → 1.98KB (53.61% reduction)
✅ fbi-tab2.jpg: 396.10KB → 128.28KB (67.61% reduction)
✅ ey-fastball.png: 4.94KB → 0.46KB (90.71% reduction)
✅ favicon-16x16.png: 0.60KB → 0.26KB (56.96% reduction)
✅ android-chrome-192x192.png: 11.50KB → 2.96KB (74.30% reduction)
✅ apple-touch-icon.png: 10.15KB → 2.55KB (74.88% reduction)
✅ android-chrome-512x512.png: 33.22KB → 6.72KB (79.77% reduction)
✅ favicon-32x32.png: 1.35KB → 0.45KB (66.45% reduction)
✅ 404-8.jpg: 43.90KB → 21.81KB (50.32% reduction)
✅ 404-9.jpg: 42.80KB → 16.86KB (60.60% reduction)
✅ 404.jpg: 25.67KB → 6.37KB (75.20% reduction)
✅ 404-11.jpg: 91.48KB → 13.71KB (85.02% reduction)
✅ 404-10.jpg: 309.60KB → 34.35KB (88.91% reduction)
✅ 404-12.jpg: 68.38KB → 50.75KB (25.79% reduction)
✅ 404-13.jpg: 73.79KB → 38.56KB (47.74% reduction)
✅ 404-14.jpg: 86.34KB → 45.91KB (46.82% reduction)
✅ 404-1.jpg: 25.67KB → 6.37KB (75.20% reduction)
✅ 404-2.jpg: 113.98KB → 67.45KB (40.82% reduction)
✅ 404-3.jpg: 38.06KB → 13.68KB (64.04% reduction)
✅ 404-7.jpg: 55.02KB → 26.54KB (51.77% reduction)
✅ 404-6.jpg: 45.98KB → 15.45KB (66.41% reduction)
✅ 404-4.jpg: 94.63KB → 50.33KB (46.82% reduction)
✅ 404-5.jpg: 48.97KB → 17.69KB (63.87% reduction)

==================================================
SUMMARY STATISTICS
==================================================
Total files converted: 1002
Total original size: 122.78MB
Total WebP size: 38.75MB
Total size reduction: 84.03MB (68.44%)
Average size reduction per file: 68.44%
Proportional savings: 122.78MB → 38.75MB
==================================================
030
Eugene Yan @eugeneyan.com · 18/05/2025
Had a fun couple of hours this weekend with Codex & Windsurf • Migrated off deprecated jekyll-algolia to official sdk (better indexing) • Added recommendations + relevance scores to each post • Improved site responsiveness; fixed dark mode flicker • Marie Kondo-ed unused files & dead code
Image of recommender widget at the bottom of posts on eugeneyan.com
151
Eugene Yan @eugeneyan.com · 07/05/2025
opps! thanks for letting me know, fixed!
bsky share button
030
Eugene Yan @eugeneyan.com · 07/05/2025
Here's a three-minute demo of news-agents in action. It's pretty cool at the 30-second mark how the sub-agents get spawned! We then see the main agent assigning tasks and polling for progress, and finally shutting the sub-agents down when they're done with their assigned tasks.
130
Eugene Yan @eugeneyan.com · 30/04/2025
@hamel.bsky.social & @sh-reya.bsky.social are two of the world's best on evals. They've built evals for 35+ AI apps & helped teams ship confidently. Now they'll teach everything they know on building evals that work. Enrollment closes in 4 days. Secret 35% discount code: maven.com/parlance-lab...
Effective Evals for AI products
042
Eugene Yan @eugeneyan.com · 28/04/2025
The Art of Doing Science and Engineering: Learning to Learn by Richard Hamming only $1.99 for the Kindle version today: amazon.com/dp/B088TMLQDC
The Art of Doing Science and Engineering: Learning to Learn by Richard Hamming
080
Eugene Yan @eugeneyan.com · 15/04/2025
Great example of generate -> validate loop + error analysis > "the most effective route to improve outcomes was brute force: retry steps until they passed or reached a limit. We give the validation errors ... to the LLM and built a loop runner"
airbnb generate-validate loop
1101
Eugene Yan @eugeneyan.com · 12/04/2025
Stumbled on the first(?) RAG in NarrativeQA from 2017. Because books & movies were too large for LSTMs to do Q&A on, they embedded 200-word chunks and retrieved similar snippets to answer questions. "Chunking and cosine similarity retrieval is so 2017." arxiv.org/abs/1712.07040
4.3 Neural Benchmarks on Stories  The design of the NarrativeQA dataset makes the straight-forward application of the existing neural architectures computationally infeasible, as this would require running an recurrent neural network on sequences of hundreds of thousands of time steps or computing a distribution over the entire input for attention, as is common.  We split the task into two steps: first, we retrieve a small number of relevant passages from the story using an IR system, and subsequently, apply one ofthe neural models above on the resulting document. The question becomes the query for retrieval. This IR problem is much harder that traditional document retrieval, as the documents, the passages here, are very similar, and the question is short and entities mentioned likely occur many times in the story. Our retrieval system considers chunks of 200 words from story and computes representations for all chunks and the query. We then select a varying number of such chunks based on their similarity to the query. We experiment with different representations and similarity measures in Section 5. Finally, we concatenate the selected chunks in the correct temporal order and insert delimiters between them to obtain a much shorter document. For span prediction models, we then further select a span from the retrieved chunks as described in Section 4.2.
0171
Eugene Yan @eugeneyan.com · 09/04/2025
If you were building a Q&A feature (or chatbot) based on very long documents (like books), what evals would you focus on?
If you were building a Q&A feature (or chatbot) based on very long documents (like books), what evals would you focus on?

1. Two metrics that come to mind
• Faithfulness: Grounding of answers in document's content. Not to be confused with correctness—an answer can be correct (based on updated information) but not faithful to the document. Sub-metric: Precision of citations
• Helpfulness: Usefulness (directly addresses the question with enough detail and explanation) and completeness (does not omit important details); an answer can be faithful but not helpful if too brief or doesn't answer the question
• Evaluate separately: Faithfulness = binary label -> LLM-evaluator; Helpfulness = pairwise comparisons -> reward model
2180
Eugene Yan @eugeneyan.com · 08/04/2025
Can't wait for when I can vibe code a production recommender system. Until then, here's some system designs: • Retrieval vs. Ranking: eugeneyan.com/writing/syst... • Real-time retrieval: eugeneyan.com/writing/real... • Personalization: eugeneyan.com/writing/patt...
Two-stage recommender systemZalando's recommender systemUnified embeddingsDoordash's search system
1484
Eugene Yan @eugeneyan.com · 05/04/2025
Your favorite AI writer's favorite AI writer
To Eugene, 

My favorite AI writer

Chip Huyen
1160
Eugene Yan @eugeneyan.com · 02/04/2025
includes resources on writing from my favourite writers
Any resources you’d recommend on the topic of writing?

Writing, Briefly
Write Like You Talk
Write Simply
Why Everyone Should Write
Writing Better
Easy Reading Is Damn Hard Writing
Mise en Place Writing
Amazon Writing Style Tips
Some Blogging Myths
Some Tactics for Writing in Public
Some Thoughts on Writing
10 years of professional blogging – what I’ve learned
Lessons from content marketing myself (aka blogging) for five years
Make Your Writing Work Harder For You
What I learned writing a book
How Jeff Bezos Turned Narrative into Amazon’s Competitive Advantage
Seemingly Paradoxical Rules of Writing
What I Did Not Learn About Writing In School
What I Learned from Writing Online - For Fellow Non-Writers
How to Write Better with The Why, What, How Framework
How to Write Design Docs for Machine Learning Systems
Writing Tools: 55 Essential Strategies for Every Writer
080
Eugene Yan @eugeneyan.com · 02/03/2025
Been querying gpt-4.5 and it's better in ways we can't quantify yet: creativity, humor, world knowledge, wisdom, nuance, based, etc. Excited about how we'll discover new ways to evaluate gpt-4.5 on these aspects which will also transfer to product / application related evals
My reaction is that there is an evaluation crisis. I don't really know what metrics to look at right now. 
MMLU was a good and useful for a few years but that's long over.
SWE-Bench Verified (real, practical, verified problems) I really like and is great but itself too narrow.
Chatbot Arena received so much focus (partly my fault?) that LLM labs have started to really overfit to it, via a combination of prompt mining (from API requests), private evals bombardment, and, worse, explicit use of rankings as training supervision. I think it's still ~ok and there's a lack of "better", but it feels on decline in signal.
There's a number of private evals popping up, an ensemble of which might be one promising path forward.
In absence of great comprehensive evals I tried to turn to vibe checks instead, but I now fear they are misleading and there is too much opportunity for confirmation bias, too low sample size, etc., it's just not great.

TLDR my reaction is I don't really know how good these models are right now.
3191
Eugene Yan @eugeneyan.com · 27/02/2025
cdn.openai.com/gpt-4-5-syst...
https://cdn.openai.com/gpt-4-5-system-card.pdf
1142
Eugene Yan @eugeneyan.com · 26/02/2025
agent ≈ model + tools, within a for-loop + environment
slide from anthropic talkslide on openai agent definition from swyx talk
2234
Eugene Yan @eugeneyan.com · 28/01/2025
♥️ it's tricky to separate what i do on the job (at the bookstore i work at) and what i hack on in my personal time. out of abundance of caution, to not discuss possible proprietary info, i won't be sharing more about the backend of aireadingclub.com 😔
👋 Hi there! I work on <redacted> at Airbnb. I really enjoy your writing.

As a personal project I started trying to write something kind of like AI reading club. I wanted a more powerful version of the "x ray" feature in Kindle, because sometimes when I pick up a new book in a series or return to a book after a break, I cannot remember all of the characters or the plot. I work mainly on <redacted> and not so much app development, and I really struggled with how to build my AI x ray. AI reading club is awesome and I was curious about how you built the retrieval pipeline, any preprocessing you did to the text, etc.

Thanks for publishing so much great work. My team and many of the machine learning engineers whom I support at Airbnb frequently share your posts.

Best regards
1100
Eugene Yan @eugeneyan.com · 22/01/2025
Thanks to the hundreds of readers who've tried aireadingclub.com and interacted with Dewey. If you've tried aireadingclub and have feedback, feature ideas, or thoughts on how AI can help you get more out of reading, please comment or dm me 🙏
Books on AI Reading Club and the number of messages on them.
060
Eugene Yan @eugeneyan.com · 17/01/2025
> Nobody tells you the variables you should be regressing. What's the target? What's the source? Do you notice when results are rubbish? ... That's why I think you need smart people who appear to do something technically easy but actually not so easy. news.ycombinator.com/item?id=1906...
"...I joined a hedged fund, Renaissance Technologies, I'll make a comment about that. It's funny that I think the most important thing to do on data analysis is to do the simple things right. So, here's a kind of non-secret about what we did at renaissance: in my opinion, our most important statistical tool was simple regression with one target and one independent variable. It's the simplest statistical model you can imagine. Any reasonably smart high school student could do it. Now we have some of the smartest people around, working in our hedge fund, we have string theorists we recruited from Harvard, and they're doing simple regression. Is this stupid and pointless? Should we be hiring stupider people and paying them less? And the answer is no. And the reason is nobody tells you what the variables you should be regressing [are]. What's the target. Should you do a nonlinear transform before you regress? What's the source? Should you clean your data? Do you notice when your results are obviously rubbish? And so on. And the smarter you are the less likely you are to make a stupid mistake. And that's why I think you often need smart people who appear to be doing something technically very easy, but actually usually not so easy.]
[[at] my hedge fund, which was not a very big company, we had 7 Phd's just cleaning data and organizing the databases]"
1211
Eugene Yan @eugeneyan.com · 15/01/2025
okay let's see what bugs come up lol 🤞
Screenshot from google analytics showing 60 active users per minute
010
Eugene Yan @eugeneyan.com · 15/01/2025
Finally, if we need help with a term or character that was previously mentioned, Dewey can help with a summary of the term so we don’t have to look it up ourselves.
Finally, if we need help with a term or character that was previously mentioned, Dewey can help with a summary of the term so we don’t have to look it up ourselves.
110
Eugene Yan @eugeneyan.com · 15/01/2025
If you've stopped reading a book for a while, it can be challenging to pick it up again and remember what you've read. To help with this, it can help with summarizing the book up to the current page and refresh our memory, highlighting major themes, characters, and concepts.
If you've stopped reading a book for a while, it can be challenging to pick it up again and remember what you've read. To help with this, it can help with summarizing the book up to the current page and refresh our memory, highlighting major themes, characters, and concepts.
110
Eugene Yan @eugeneyan.com · 15/01/2025
It can also help with creating quizzes / flashcards. The goal here is to test our knowledge and improve retention.
It can also help with creating quizzes / flashcards. The goal here is to test our knowledge and improve retention.
110
Eugene Yan @eugeneyan.com · 15/01/2025
With the context, it can answer simple queries via "Explain" and "Discuss". The goal is to keep us in flow while reading, instead of having to reread other sections of the book or open a web browser for our queries.
With the context, it can answer simple queries via "Explain" and "Discuss". The goal is to keep us in flow while reading, instead of having to reread other sections of the book or open a web browser for our queries.
110
Eugene Yan @eugeneyan.com · 15/01/2025
At the heart of AI Reading Club is Dewey, your AI reading companion. It understands context via selected text or the page we're on. This explicit context is displayed during discussions. At the same time, behind the scenes, it can retrieve and consider the rest of the book as implicit context.
It understands our context either via the text we select or the page we're on. This explicit context is displayed during discussions. At the same time, behind the scenes, it can also retrieve and consider the rest of the book as implicit context.
130
Eugene Yan @eugeneyan.com · 24/12/2024
hmm i feel attacked
deciding between "actually taking some time off" and "work on personal projects and call it relaxing"
59812
Eugene Yan @eugeneyan.com · 19/12/2024
phone usage has halved since starting december detox and deleting all social media apps off my phone
2hr 17min average daily screen time1hr 11min average daily screen time1hr 2min average daily screen time51min average daily screen time
2210
Eugene Yan @eugeneyan.com · 11/12/2024
to the latter point, the anti-ai comment were really strong in several comments, even those that didn't have anything to do with ai (these images are part of the appendix of the writeup)
comments that talk about the inevitability of ai failingcomments on bullying pro-ai folkscomments stating that ai sucks
010
Eugene Yan @eugeneyan.com · 11/12/2024
also, while _some_ accounts that didn't like their data being scrapped had anti-AI explicitly posted on their profiles, not all of did. i hope i expressed this nuance sufficiently, and not a sweeping "people objecting to their data being stolen without permission as anti-AI"
To validate this, I visited a sample of the critique accounts and found at least a dozen profiles of writers, artists, musicians, and other creatives, some of whom added explicit anti-AI declarations to their bios.
100
Eugene Yan @eugeneyan.com · 11/12/2024
ah i see your point now, thank you for clarifying! my point was that there was no stealing of data from bluesky's database, and no scraping of html. instead, the data was simply downloaded via the api. i deliberately avoiding comment on license or legality.
At first glance, the primary issue seemed to be the lack of consent in data collection. Some comments expressed anger at having their data “scraped” or “stolen” without permission. (Note: The accusation is incorrect—there was no scraping or stealing involved. Posts on Bluesky are public and the Bluesky firehose API is open for all to consume.)
000
Eugene Yan @eugeneyan.com · 11/12/2024
oh, perhaps it's just me that isn't used to such reactions to the release of a dataset, and the comments against the training of AI. haven't come across such negative reactions elsewhere tbh
Reactions against the release of one-million-bluesky-postsresponses objecting to the dataset being used as training data
120
Eugene Yan @eugeneyan.com · 09/12/2024
Day 1 of hitting the slopes this winter 🏂
1140
Eugene Yan @eugeneyan.com · 07/12/2024
Repeat after me: I will build evals for my tasks. I will build evals for my tasks. I will build evals for my tasks.
Academic benchmarks are not your tasks.
4649
Eugene Yan @eugeneyan.com · 07/12/2024
Learning about quantization suffixes while `ollama pull llama3.3` download completes (fyi, quantization for the default 70b is q4_K_M) • make-ggml .py: github.com/ggerganov/ll... • pull request: github.com/ggerganov/ll...
Old quant types (some base model types require these):
- Q4_0: small, very high quality loss - legacy, prefer using Q3_K_M
- Q4_1: small, substantial quality loss - legacy, prefer using Q3_K_L
- Q5_0: medium, balanced quality - legacy, prefer using Q4_K_M
- Q5_1: medium, low quality loss - legacy, prefer using Q5_K_M

New quant types (recommended):
- Q2_K: smallest, extreme quality loss - not recommended
- Q3_K: alias for Q3_K_M
- Q3_K_S: very small, very high quality loss
- Q3_K_M: very small, very high quality loss
- Q3_K_L: small, substantial quality loss
- Q4_K: alias for Q4_K_M
- Q4_K_S: small, significant quality loss
- Q4_K_M: medium, balanced quality - recommended
- Q5_K: alias for Q5_K_M
- Q5_K_S: large, low quality loss - recommended
- Q5_K_M: large, very low quality loss - recommended
- Q6_K: very large, extremely low quality loss
- Q8_0: very large, extremely low quality loss - not recommended
- F16: extremely large, virtually no quality loss - not recommended
- F32: absolutely huge, lossless - not recommended
3234
Eugene Yan @eugeneyan.com · 03/12/2024
Here's my attempt at something similar—machine learning systems and applications in industry—a couple years ago. Isn't as fancy as the one above though lol applyingml.com/papers/
Database of resources for applying ml
2130
Eugene Yan @eugeneyan.com · 03/12/2024
Wow, this is such a useful resource of industry LLM applications! And filtering via search/tags is so responsive. I was thinking of compiling something like this over the holidays (ala applied-ml) but thanks to @strickvl.bsky.social I can spend the time reading instead ♥️ zenml.io/llmops-datab...
UI for the LLMOps Database with the search query of "eval"
3504
Eugene Yan @eugeneyan.com · 01/12/2024
@hamel.bsky.social is 💯: hamel.dev/blog/posts/a... Write for yourself. Assume no one reads it. Write about topics to learn, clarify your thoughts, put out a bat signal. This makes it sustainable. Writing is primarily a single player game; multiplayer benefits (e.g., audience building) are bonuses.
The key is authenticity. Don’t do this just for marketing—do it because you’re genuinely interested in learning from others and building on their ideas. It’s not hard to find things to be excited about. I’m amazed by how few people take this approach. It’s both effective and fun.
4595
Eugene Yan @eugeneyan.com · 30/11/2024
Great explainer on sinusoidal positional encoding and rotary positional embedding (RoPE). fleetwood.dev/posts/you-co...
2616
Eugene Yan @eugeneyan.com · 28/11/2024
This Thanksgiving, I'm grateful for peace, health, and happiness for my family and me—what are you thankful for?
Nassim Taleb on 'True Wealth'
Vishal Khandelwal, safalniveshak.com
Nassim Taleb is one of my favourite authors, and his Antifragile is one of my favourite books. One of this book's chapters that interests me particularly is titled 'Via Negativa'. Here, Taleb argues that the solution to many problems in life is by removing things, not adding things.
For example, here is a list of things Taleb counts as constituents of true wealth that are all about subtracting things (via negativa) from life than adding -

Worriless sleeping
Clear conscience
Reciprocal gratitude
Absence of envy
Good appetite
Muscle strength
Physical energy
Frequent laughs
No meals alone
No gym classes
Some physical labor
Good bowel movements
No meeting rooms
Periodic surprises

I could check twelve from this list (let the ones I didn't check remain a secret).
What about you? What in the list remains getting checked for you?
11413
Eugene Yan @eugeneyan.com · 27/11/2024
And if you're looking for more learning over the long thanksgiving weekend, this could be a good place to start: eugeneyan.com/start-here/
Writeups on machine learning systems, techniques, ml & engineering, and weekend prototypes
1145
Eugene Yan @eugeneyan.com · 27/11/2024
Feels good to be mentioned on HN for engineers learning AI 🥰 Helping others is a big reason I write. Here's a list on ML/AI: ## Building AI systems • Patterns for Building LLM-based Systems: eugeneyan.com/writing/llm-... • What We’ve Learned From A Year of Building with LLMs: applied-llms.org
Read through this making flashcards as you to: https://eugeneyan.com/writing/llm-patterns/
Then spin up a RAG-enhanced chatbot using pgvector on your favourite subject, and keep improving it when you learn about cool techniques

---

Lots of people can get impressive demos up and running, but if you want to run AI products in production, you're going to have to do system evals. System evals make sure your product is doing what it says on the box with unquantifiable qualities.
We wrote a zine on system evals without jargon: https://forestfriends.tech
Eugene Yan has written extensively on it https://eugeneyan.com/writing/evals/
Hamel has as well. https://hamel.dev/blog/posts/evals/
3709
Eugene Yan @eugeneyan.com · 27/11/2024
yay Latte welcomes you!
020
Eugene Yan @eugeneyan.com · 23/11/2024
It's helpful to distinguish between crisp vs. fuzzy tasks: • Crisp: Answers are verifiable, like math or code • Fuzzy: Answers are subjective, like reasoning, summarization, translation, multi-turn dialogue The latter is far harder to evaluate reliably aligned.substack.com/p/crisp-and-...
This is a taxonomy for task space that I find useful when thinking about what we need to do to solve alignment.

Crisp tasks are more reasoning/system-2 based. Whether a response is good is typically precisely defined. Reasonable, knowledgeable people don’t disagree about it.
Examples: math, coding competitions, physics questions, logic puzzles, basic factual knowledge, …
Fuzzy tasks are more intuition/system-1 based. Whether a response is good is typically somewhat vaguely defined or could fall within a range. Reasonable, knowledgeable people may disagree about it.
Examples: distinguishing cats and dogs in pictures, recognizing strong go moves, poetry writing, assistant helpfulness ratings, …
In contrast to fuzzy logic, where this name is borrowed from, here crisp and fuzzy don’t refer to the output space: preferences comparisons have a small number of discrete allowed values, and image classification typically uses a discrete space. Crisp tasks like coding competitions have a output space that involves hundreds of discrete tokens, and this is the same output space of poetry writing, a fuzzy task.

There are many crisp tasks that we can evaluate very reliably, and thus it’s feasible to train on them extensively. In contrast, fuzzy tasks are often particularly difficult to evaluate reliably, and we currently don’t know how to do this at a superhuman level.
3470
Eugene Yan @eugeneyan.com · 22/11/2024
Hey Bluesky, meet swirly-green-with-a-touch-of-pink sky
Northern lights, streak of green with an iglooNorthern lights, swirls of green over igloosNorthern lights with streaks of green and a touch of pinkNorthern lights, streak of green and some pink
2481
Eugene Yan @eugeneyan.com · 20/11/2024
yea already there!
tools for brainstorming and development such as obsidian, zotero, and cursor
110
Eugene Yan @eugeneyan.com · 20/11/2024
yea some standard commands / packages need workarounds for fish, such as rbenv
brew install chruby-fish ruby-install ruby-build
brew install rbenv

# workaround for fish shell
set --universal fish_user_paths $fish_user_paths ~/.rbenv/shims
rbenv global 3.3.5
rbenv rehash
110
Eugene Yan @eugeneyan.com · 17/11/2024
jealous! here’s my view from the plane lol
View from plane on flight to alaska
010
Eugene Yan @eugeneyan.com · 17/11/2024
probably data, evals, and the flywheel that mixes the cake batter
Potentially nitpicky but competitive advantage in AI goes not so much to those with data but those with a data engine: iterated data aquisition, re-training, evaluation, deployment, telemetry. And whoever can spin it fastest. Slide from Tesla to ~illustrate but concept is general
010
Eugene Yan @eugeneyan.com · 17/11/2024
Also see Karparthy’s take on it
It’s hard to understand now, the Atari RL paper of 2013 and its extensions was the by far dominant meme. One single general learning algorithm discovered an optimal strategy to Breakout and so many other games. You just had to improve and scale it enough. My recollection of the memetics is that Yann LeCun was one prominent person who really didn’t care much and talked about the cake over and over again, where RL was just the final cherry on top with representation learning as the meat and supervised learning the icing, and he was conceptually exactly right about that at least with today’s stack and hindsight (pretraining = meat, SFT = icing, RLHF = cherry, ie the basic ChatGPT training pipeline). Which is fun because today he really doesn’t care much for LLMs either. (But for reasons that I tbh don’t always fully follow.)
2230
Eugene Yan @eugeneyan.com · 17/11/2024
Eight years later, Yann LeCun’s cake 🍰 analogy was spot on: self-supervised > supervised > RL > “If intelligence is a cake, the bulk of the cake is unsupervised learning, the icing on the cake is supervised learning, and the cherry on the cake is reinforcement learning (RL).”
Yann LeCun’s analogy of intelligence being a cake of self-supervised, supervised, and RL
109413