Sign in

Dimitris Papailiopoulos

@dimitrisp.bsky.social
1.9K followers 293 following 171 posts

Researcher @MSFTResearch; Prof @UWMadison (on leave); learning in context; thinking about reasoning; babas of Inez Lily. papail.io

PostsRepliesMedia
Dimitris Papailiopoulos @dimitrisp.bsky.social · 16/05/2025
What if for most of your findings you just post a thread and share a GitHub repo, rather than submitting a 15 page NeurIPS paper with < 1/100 the reach?
450
Dimitris Papailiopoulos @dimitrisp.bsky.social · 12/05/2025
LLMs learn world models, beyond a reasonable doubt. It's been the case since GPT-3, but now it should be even more clear. Without them "Guess and Check" would not work. The fact that these "world models" are approximate/incomplete does not disqualify them.
110
Dimitris Papailiopoulos @dimitrisp.bsky.social · 07/05/2025
Is 1948 widely acknowledged as the birth of language models and tokenizers? In "A Mathematical Theory of Communication", almost as an afterthought Shannon suggests the N-gram for generating English, and that word level tokenization is better than character level tokenization.
2110
Reposted by Dimitris Papailiopoulos
Besmira Nushi @besmiranushi.bsky.social · 01/05/2025
🎉The Phi-4 reasoning models have landed on HF and Azure AI Foundry. The new models are competitive and often outperform much larger frontier models. It is exciting to see the reasoning capabilities extend to more domains beyond math, including algorithmic reasoning, calendar planning, and coding.
1208
Dimitris Papailiopoulos @dimitrisp.bsky.social · 30/04/2025
I am afraid to report, RL works. I think 2-3 years ago, I said I will not work on two ML sub-areas. RL was one of them. I am happy to say that I am not strongly attached to my beliefs.
061
Dimitris Papailiopoulos @dimitrisp.bsky.social · 30/04/2025
Re: The Chatbot Arena Illusion Every eval chokes under hill climbing. If we're lucky, there’s an early phase where *real* learning (both model and community) can occur. I'd argue that a benchmark’s value lies entirely in that window. So the real question is what did we learn?
191
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/04/2025
Fun trivia now that “sycophant” became common language to describe LLMs flattering users: In Greek, συκοφάντης (sykophántēs) most typically refers to a malicious slanderer, someone spreading lies, not flattery! Every time you use it, you’re technically using it wrong :D
160
Dimitris Papailiopoulos @dimitrisp.bsky.social · 21/02/2025
Come work with us at MSR AI Frontiers and help us figure out reasoning! We're hiring at the Senior Researcher level (eg post phd). Please drop me a DM if you do! jobs.careers.microsoft.com/us/en/job/17...
jobs.careers.microsoft.com
Search Jobs | Microsoft Careers
061
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
o3 can't multiply beyond a few digits... But I think multiplication, addition, maze solving and easy-to-hard generalization is actually solvable on standard transformers... with recursive self-improvement Below is the acc of a tiny model teaching itself how to add and multiply
2214
Reposted by Dimitris Papailiopoulos
Sung Kim @sungkim.bsky.social · 13/02/2025
o3 can't multiply beyond a few digits... But he think multiplication, addition, maze solving and easy-to-hard generalization is actually solvable on standard transformers... with recursive self-improvement, as presented by @dimitrisp.bsky.social
162
Dimitris Papailiopoulos @dimitrisp.bsky.social · 02/02/2025
Self-improving Transformers can overcome easy-to-hard and length generalization challenges. Paper on arxiv coming on Monday. Link to a talk I gave on this below 👇 Super excited about this work! Talk : youtube.com/watch?v=szhE... slides: tinyurl.com/SelfImprovem...
0151
Dimitris Papailiopoulos @dimitrisp.bsky.social · 01/02/2025
Two months before R1 came out, I wrote this in my small notebook of ideas as something to test #schmidhuber
160
Dimitris Papailiopoulos @dimitrisp.bsky.social · 29/01/2025
Now that we have reasoner LLMs, let's think about how to GRPO problem generators that generate instances that sit right outside the frontier of current capabilities.
130
Reposted by Dimitris Papailiopoulos
Constantine Caramanis @cmcaram.bsky.social · 29/01/2025
🚀 🇬🇷 A year in the making! I’ve just completed a set of 21 lectures in Machine Learning, in Greek, designed for high school students. The course introduces key ML concepts, coding in Python & PyTorch, and real-world AI applications. 👉 WebPage: tinyurl.com/ye2awe8m 🎥 YouTube: tinyurl.com/2wwjru6z
tinyurl.com
Μηχανική Μάθηση (Machine Learning) - YouTube
Διαλέξεις Τεχνητής Νοημοσύνης και Μηχανικής Μάθησης: https://caramanis.github.io/MachineLearningClass/ Καλωσορίσατε στο μάθημα τεχνητής νοημοσύνης και μηχανι...
283
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/01/2025
If you wanted to collect 1 mil reasoning traces from human subjects on say math, that would cost ~$50m, assuming ~50$/person/hour. Interesting to compare with the cost to generate them from a reasoning LLM, with say with cost per trace ~$0.5 (say 10k tokens).. That's 100x cheaper
241
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/01/2025
Ok we've read a lot about test-time compute being the new scaling axis, but what's the next scaling axis?
000
Reposted by Dimitris Papailiopoulos
Ben Recht @beenwrekt.bsky.social · 27/01/2025
2014 GoogLeNet: The best image classifier was only trainable using weeks of Google's custom infrastructure. 2018 ResNet: A more accurate model is trainable in a 1/2 hour on a single GPU. What stops this from happening for LLMs?
3519
Dimitris Papailiopoulos @dimitrisp.bsky.social · 26/01/2025
A strong math/theory foundation can be extremely useful for ML research. Not for proving sample complexity bounds on "AGI", but for offering a mental model of inaccessible and complex systems, that can allow for accurate predictions, without running expensive experiments.
1120
Dimitris Papailiopoulos @dimitrisp.bsky.social · 25/01/2025
The "deepseek distilled o1" is an intellectually vacuous discussion, precisely because what they reported in the R1 paper is a reproducible phenomenon! By now many experiments on non-deepseek models show that acc and inf-time compute increase as the result of outcome-based RL.
2220
Dimitris Papailiopoulos @dimitrisp.bsky.social · 24/01/2025
GRPO and outcome based RL rely heavily on a verifier with access to ground truth data. But likely can work beyond strictly verifiable domains, as long as you have access to a "weak" grader. And perhaps even beyond that, if "correct trajectories" share a common fingerprint..
1111
Dimitris Papailiopoulos @dimitrisp.bsky.social · 24/01/2025
Elated to announce that I got some papers accepted, some rejected, and some withdrawn, at some conference that I won't attend :D
0110
Dimitris Papailiopoulos @dimitrisp.bsky.social · 22/01/2025
1/5 A hypothesis on the emergence of long form "yapping" in reasoning models: The increase of "yapping" in reasoning models, as they are trained for more rounds of RL, "emerges" (sorry :D) as models discover that verbose reasoning helps them achieve better rewards (eg higher acc).
120
Dimitris Papailiopoulos @dimitrisp.bsky.social · 20/01/2025
I love finding silly tests that LLMs are terrible at. Here's a new one for me: Drawing with Logo (yes the turtle)! To be fair drawing with Logo is hard. But.. here goes 8 examples with sonnet 3.6 vs o1. Example 1/8: Draw the letter G
170
Dimitris Papailiopoulos @dimitrisp.bsky.social · 18/01/2025
Task vectors are akin to punchcards: you feed them to your LLM and it implements specific tasks, without in-context demonstrations. Liu's new paper examines at what scale, where in the network and when during training do they emerge, and how to encourage their emergence. arxiv.org/pdf/2501.09240
0233
Dimitris Papailiopoulos @dimitrisp.bsky.social · 01/01/2025
resolutions for 2025 - be a good dad and partner - do more nature stuff - walk & run more - spend more time in the water - think deeper - read good books, don’t feel bad not finishing all - be a good mentor & colleague - figure out what reasoning is - don’t be reward hacking - have fun All doable
080
Dimitris Papailiopoulos @dimitrisp.bsky.social · 30/12/2024
When we expect an LLM to perform algorithm A on an input X, the prompt to the model should be of sufficient specificity and complexity in order to uniquely identify and run A, among all the possible other algorithms, and their superpositions, that the model can implement.
230
Dimitris Papailiopoulos @dimitrisp.bsky.social · 30/12/2024
The single most insightful paper that I read after the initial wave of ICL papers was one that further shifted my understanding on "the early ascent phenomenon" by Kangwook Lee's group x.com/Kangwook_Lee...
0182
Dimitris Papailiopoulos @dimitrisp.bsky.social · 29/12/2024
I've been thinking about in-context learning for nearly 3 years. While there is still plenty I don't fully understand, five papers have--to a very large extent--shaped my perspective on it, and I believe everyone should read them.
68012
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/12/2024
130
Dimitris Papailiopoulos @dimitrisp.bsky.social · 23/12/2024
What is the MMLU equivalent benchmark but for computer use?
010
Reposted by Dimitris Papailiopoulos
Gergely Neu @neu-rips.bsky.social · 21/12/2024
*PLS SHARE* Open position for an *Associate Professor* in Machine Learning at our department (@enginyeria-upf.bsky.social / @upf.edu), via the Serra Hunter programme. DEADLINE: January 13th 2025 www.upf.edu/web/personal...
upf.edu
Professor Agregat Serra Húnter. Departament d'Enginyeria
Convocatòria 2024-30 PDI Serra Húnter Obert Convocatòria 2024-30 Adscripció: Department: Enginyeria Profile: Machine Learning Termini de sol·licituds: 13/01/2025 Data de publicació a la web:...
02620
Dimitris Papailiopoulos @dimitrisp.bsky.social · 21/12/2024
Life after o5 saturates FrontierMath
060
Dimitris Papailiopoulos @dimitrisp.bsky.social · 21/12/2024
Imagine seeing a C compiler and concluding that assembly engineers would go out of work
271
Dimitris Papailiopoulos @dimitrisp.bsky.social · 20/12/2024
The wall hit a wall
040
Dimitris Papailiopoulos @dimitrisp.bsky.social · 16/12/2024
I tested chatGPT o1 pro mode (with best of N) on AIME 2024. It scored 93.3%. it got 14 out of 15 questions, on both I and II versions. Uhm, wow.
160
Dimitris Papailiopoulos @dimitrisp.bsky.social · 16/12/2024
there must be a name for the bias in thinking that o1 pro is better than other models because it uses more time to think, even if its output answers are similar.
220
Dimitris Papailiopoulos @dimitrisp.bsky.social · 16/12/2024
openai should consider playing some dramatic soundtrack while o1-pro is thinking; sometimes it takes more than 5 mins staring at this
250
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/12/2024
1/6 🧵on LEXICO: Dictionary Learning meets KV Cache Compression. LEXICO can compress the KV cache down to 15-25% with minimal acc loss. The core idea? Represent the KV vectors as sparse linear combination sof atoms from a learned dictionary ≈ SAE for keys and values.
191
Dimitris Papailiopoulos @dimitrisp.bsky.social · 12/12/2024
LEXICO defining the Pareto frontier of KV cache compression
0113
Dimitris Papailiopoulos @dimitrisp.bsky.social · 12/12/2024
Chat, that's not how you make a pourover. where is the scale? where is the water dosaging vs coffee weight? 2/10 LLMs have hit the James Hoffman wall.
060
Dimitris Papailiopoulos @dimitrisp.bsky.social · 12/12/2024
o1-mini and o1-preview are the only models that can currently solve this models that can't: chatGPTo1/gemini 2.0 or 1.5/4o/sonnet 3.6/opus/r1/grok-2
020
Dimitris Papailiopoulos @dimitrisp.bsky.social · 11/12/2024
LEXICO: Dictionary Learning meets KV Caching... coming to your nearest arXiv later this week :)
060
Dimitris Papailiopoulos @dimitrisp.bsky.social · 09/12/2024
“an image is worth a thousand slops”
000
Dimitris Papailiopoulos @dimitrisp.bsky.social · 06/12/2024
ChatGPT o1 on the TikZ unicorn, given 20 rounds of self improvement. Below are versions @ round 1, 9, 17, and 20 Hmmmm
280
Dimitris Papailiopoulos @dimitrisp.bsky.social · 05/12/2024
We'll soon review internship applicants for our reasoning group; please submit and ping me if you haven't already!! Can't wait to work with you!!
052
Dimitris Papailiopoulos @dimitrisp.bsky.social · 04/12/2024
An LLM will prove the Riemann hypothesis before a world model can generate better games than the Unreal Engine.
211
Reposted by Dimitris Papailiopoulos
Stella Biderman @stellaathena.bsky.social · 28/11/2024
A dataset of 1 million or 2 million Bluesky posts is completely irrelevant to training large language models. The primary usecase for the datasets that people are losing their shit over isn't ChatGPT, it's social science research and developing systems that improve Bluesky.
825139
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/11/2024
"Thank you for significantly improving your LLM, which addresses several of my concerns. I will maintain my preferences, and still not use it."
150
Reposted by Dimitris Papailiopoulos
Kyle Cranmer @kylecranmer.bsky.social · 25/11/2024
🚨 The University of Wisconsin—Madison has a huge faculty hiring Initiative called RISE. We expect to hire an additional 120-150 new faculty over the next 3-5 years. Focus areas are AI, sustainability, and health. Learn more: rise.wisc.edu/about/ Jobs: jobs.wisc.edu/pages/wiscon...
Screenshot of webpage
46014
Dimitris Papailiopoulos @dimitrisp.bsky.social · 21/11/2024
the world model here being "an interface for universal computation"
120