Sign in

Dimitris Papailiopoulos

@dimitrisp.bsky.social
1.9K followers 293 following 171 posts

Researcher @MSFTResearch; Prof @UWMadison (on leave); learning in context; thinking about reasoning; babas of Inez Lily. papail.io

PostsRepliesMedia
Dimitris Papailiopoulos @dimitrisp.bsky.social · 16/05/2025
In the mean time, here's a rule of thumb "if your project can be vibecoded in an hour, and amounts to O(10) LoC edits on something existing, or is a convergence proof that o4-mini can do with a bit of guidance, DO NOT write a paper about it":D
010
Dimitris Papailiopoulos @dimitrisp.bsky.social · 16/05/2025
I think that the current most bullet proof peer review has been "people will read/try your stuff, and if it works they build on it". But because it's not attached to a formal process on openreview we discard it as being non-scientific.
110
Dimitris Papailiopoulos @dimitrisp.bsky.social · 16/05/2025
It seems to me that is totally misaligned with scientific discovery and progress. I don't believe this is a result of bad actors btw. It's just that huge, and complex systems that are O(100) years old take a long time to change, and readjust to new realities. We'll eventually figure it out.
120
Dimitris Papailiopoulos @dimitrisp.bsky.social · 16/05/2025
it seems to me that mostly ML academia (i am part of it!) is a proponent of keeping peer review and mega ML conferences going & the bean counter running. We've not found a solution to reviews converging to random coin tosses, at a huge expense of human work hours.
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 16/05/2025
If that's indeed the case (i believe we can measure that), and their key function is social, and a way for people to connect (that's great!), what's the point of having peer review, and using # neurips papers as a bean counter?
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 16/05/2025
my post is a direct criticism to the 100k neurips submissions issue. It's beyond clear that research dissemination--for the most part--does not happen through conferences any more.
110
Dimitris Papailiopoulos @dimitrisp.bsky.social · 16/05/2025
What if for most of your findings you just post a thread and share a GitHub repo, rather than submitting a 15 page NeurIPS paper with < 1/100 the reach?
450
Dimitris Papailiopoulos @dimitrisp.bsky.social · 12/05/2025
LLMs learn world models, beyond a reasonable doubt. It's been the case since GPT-3, but now it should be even more clear. Without them "Guess and Check" would not work. The fact that these "world models" are approximate/incomplete does not disqualify them.
110
Dimitris Papailiopoulos @dimitrisp.bsky.social · 08/05/2025
Working on the yapping part :)
000
Dimitris Papailiopoulos @dimitrisp.bsky.social · 08/05/2025
hmm.. temp has to be 0.6-0.8, this looks like very low temp outputs
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 08/05/2025
I don’t see at all how this is intellectually close to what Shannon wrote. Can you clarify? I read it as computing statistics and how these are compatible with theoretical conjectures. There’s no language generation implicit in the article. Am I misreading it?
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 07/05/2025
can you share the paper?
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 07/05/2025
BTW for historical context, 1948, is very very very early to have these thoughts. So i actually think that every single sentence written is profound. This is kinda random, but here is how Greece looked back then. IT WAS SO EARLY :) x.com/DimitrisPapa...
x.com
Dimitris Papailiopoulos on X: "For historical context, here are some pictures of the Greek islands in the late 40s https://t.co/WJsM83IHW1" / X
For historical context, here are some pictures of the Greek islands in the late 40s https://t.co/WJsM83IHW1
010
Dimitris Papailiopoulos @dimitrisp.bsky.social · 07/05/2025
it's not that profound. it just says, there's no wall, if all stars are aligned. it's an optimistic read of the setting.
110
Dimitris Papailiopoulos @dimitrisp.bsky.social · 07/05/2025
Is 1948 widely acknowledged as the birth of language models and tokenizers? In "A Mathematical Theory of Communication", almost as an afterthought Shannon suggests the N-gram for generating English, and that word level tokenization is better than character level tokenization.
2110
Reposted by Dimitris Papailiopoulos
Besmira Nushi @besmiranushi.bsky.social · 01/05/2025
🎉The Phi-4 reasoning models have landed on HF and Azure AI Foundry. The new models are competitive and often outperform much larger frontier models. It is exciting to see the reasoning capabilities extend to more domains beyond math, including algorithmic reasoning, calendar planning, and coding.
1208
Dimitris Papailiopoulos @dimitrisp.bsky.social · 30/04/2025
I am afraid to report, RL works. I think 2-3 years ago, I said I will not work on two ML sub-areas. RL was one of them. I am happy to say that I am not strongly attached to my beliefs.
061
Dimitris Papailiopoulos @dimitrisp.bsky.social · 30/04/2025
researchers
010
Dimitris Papailiopoulos @dimitrisp.bsky.social · 30/04/2025
Re: The Chatbot Arena Illusion Every eval chokes under hill climbing. If we're lucky, there’s an early phase where *real* learning (both model and community) can occur. I'd argue that a benchmark’s value lies entirely in that window. So the real question is what did we learn?
191
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/04/2025
Also a sycophant etymologically means "the one who shows the figs"; the origin of the meaning is kinda debated, either refers to illegally importing figs, or to falsely accusing someone of hiding illegally imported figs
040
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/04/2025
Fun trivia now that “sycophant” became common language to describe LLMs flattering users: In Greek, συκοφάντης (sykophántēs) most typically refers to a malicious slanderer, someone spreading lies, not flattery! Every time you use it, you’re technically using it wrong :D
160
Dimitris Papailiopoulos @dimitrisp.bsky.social · 21/02/2025
Come work with us at MSR AI Frontiers and help us figure out reasoning! We're hiring at the Senior Researcher level (eg post phd). Please drop me a DM if you do! jobs.careers.microsoft.com/us/en/job/17...
jobs.careers.microsoft.com
Search Jobs | Microsoft Careers
061
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
bsky doesn't like GIFs, here they are from the other site x.com/DimitrisPapa...
x.com
x.com
010
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
Super proud of this work that was led by Nayoung Lee and Jack Cai, with mentorship from Avi Schwarzschild and Kangwook Lee link to our paper: arxiv.org/abs/2502.01612
arxiv.org
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
Large language models often struggle with length generalization and solving complex problem instances beyond their training distribution. We present a self-improvement approach where models iterativel...
000
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
Oh btw, self improvement can become exponentially faster in some settings, ory when we apply it on pretrained models (again this is all for add/mul/maze etc)
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
An important aspect of the method is that you need to 1) generate problems of appropriate hardness 2) be able to filter our negative examples using a cheap verifier. Otherwise the benefit of self-improvement collapses.
110
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
We test self-improvement across diverse algorithmic tasks: - Arithmetic: Reverse addition, forward (yes forward!) addition, multiplication (with CoT) - String Manipulation: Copying, reversing - Maze Solving: Finding shortest paths in graphs. It always works
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
Self-improvement is not new—this idea has been explored in various contexts and domains (like reasoning, mathematics, coding, and more). Our results suggest that self-improvement is a general and scalable solution to length & difficulty generalization!
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
What if we leverage this? What if we let the model label slightly harder data… and then train on them? Our key idea is to use Self-Improving Transformers , where a model iteratively labels its own train data and learns from progressively harder examples (inspired by methods like STaR and ReST).
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
I was kind of done with length gen, but then I took a closer look at that figure above.. I noticed that there is a bit of transcendence, i.e the model trained on n-digit ADD can solve slightly harder problems, eg n+1, but not much more. (cc on transcendence and chess arxiv.org/html/2406.11741v1)
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
Even for simple algorithmic tasks like integer addition, performance collapses as sequence length increases. The only way so far that overcomes this relies on heavily optimizing positional encoding and data format. (Figure from: Cho et al., arxiv.org/abs/2405.20671)
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
Standard transformers struggle with length generalization—extrapolating beyond their training distribution. Even GPT-4, o1, and o3 can't multiply long digit numbers. Length generalization using vanilla transformers is a long-standing open problem x.com/yuntiandeng/...
x.com
x.com
100
Dimitris Papailiopoulos @dimitrisp.bsky.social · 13/02/2025
o3 can't multiply beyond a few digits... But I think multiplication, addition, maze solving and easy-to-hard generalization is actually solvable on standard transformers... with recursive self-improvement Below is the acc of a tiny model teaching itself how to add and multiply
2214
Reposted by Dimitris Papailiopoulos
Sung Kim @sungkim.bsky.social · 13/02/2025
o3 can't multiply beyond a few digits... But he think multiplication, addition, maze solving and easy-to-hard generalization is actually solvable on standard transformers... with recursive self-improvement, as presented by @dimitrisp.bsky.social
162
Dimitris Papailiopoulos @dimitrisp.bsky.social · 02/02/2025
Self-improving Transformers can overcome easy-to-hard and length generalization challenges. Paper on arxiv coming on Monday. Link to a talk I gave on this below 👇 Super excited about this work! Talk : youtube.com/watch?v=szhE... slides: tinyurl.com/SelfImprovem...
0151
Dimitris Papailiopoulos @dimitrisp.bsky.social · 01/02/2025
Two months before R1 came out, I wrote this in my small notebook of ideas as something to test #schmidhuber
160
Dimitris Papailiopoulos @dimitrisp.bsky.social · 30/01/2025
I think so
010
Dimitris Papailiopoulos @dimitrisp.bsky.social · 29/01/2025
Now that we have reasoner LLMs, let's think about how to GRPO problem generators that generate instances that sit right outside the frontier of current capabilities.
130
Reposted by Dimitris Papailiopoulos
Constantine Caramanis @cmcaram.bsky.social · 29/01/2025
🚀 🇬🇷 A year in the making! I’ve just completed a set of 21 lectures in Machine Learning, in Greek, designed for high school students. The course introduces key ML concepts, coding in Python & PyTorch, and real-world AI applications. 👉 WebPage: tinyurl.com/ye2awe8m 🎥 YouTube: tinyurl.com/2wwjru6z
tinyurl.com
Μηχανική Μάθηση (Machine Learning) - YouTube
Διαλέξεις Τεχνητής Νοημοσύνης και Μηχανικής Μάθησης: https://caramanis.github.io/MachineLearningClass/ Καλωσορίσατε στο μάθημα τεχνητής νοημοσύνης και μηχανι...
283
Dimitris Papailiopoulos @dimitrisp.bsky.social · 29/01/2025
it will still be more expensive than data generation, no?
000
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/01/2025
lol
000
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/01/2025
how much would you charge per hour to solve hard math problems at the olympiad level? this is not high school math
110
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/01/2025
yes I agree with that too
000
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/01/2025
Agree on this perspective, but deepseek still had a lot of money to spend, many GPUs (2k at least), and a 200 people cracked engineering/scientist team. Still a smaller scale effort than openai, but at a scale that is far above even multi-university groups
110
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/01/2025
the median was kinda mean
010
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/01/2025
I am getting professional anxiety. they can solve problems (that I can't solve) 100-1k X cheaper than humans.
110
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/01/2025
If you wanted to collect 1 mil reasoning traces from human subjects on say math, that would cost ~$50m, assuming ~50$/person/hour. Interesting to compare with the cost to generate them from a reasoning LLM, with say with cost per trace ~$0.5 (say 10k tokens).. That's 100x cheaper
241
Dimitris Papailiopoulos @dimitrisp.bsky.social · 28/01/2025
Ok we've read a lot about test-time compute being the new scaling axis, but what's the next scaling axis?
000
Reposted by Dimitris Papailiopoulos
Ben Recht @beenwrekt.bsky.social · 27/01/2025
2014 GoogLeNet: The best image classifier was only trainable using weeks of Google's custom infrastructure. 2018 ResNet: A more accurate model is trainable in a 1/2 hour on a single GPU. What stops this from happening for LLMs?
3519
Dimitris Papailiopoulos @dimitrisp.bsky.social · 27/01/2025
Doubles sounds more impressive than 100% :)
110