Sign in

Matteo Pagliardini

@matpagliardini.bsky.social
250 followers 1K following 8 posts

PhD student in ML at EPFL 🇨🇭working with Martin Jaggi & François Fleuret. Previously Apple MLR (intern). mpagli.github.io

PostsRepliesMedia
Reposted by Matteo Pagliardini
Martin Jaggi @mjaggi.bsky.social · 03/09/2025
new extensive evaluation of different optimizers for LLM training arxiv.org/abs/2509.01440
arxiv.org
Benchmarking Optimizers for Large Language Model Pretraining
The recent development of Large Language Models (LLMs) has been accompanied by an effervescence of novel ideas and methods to better optimize the loss of deep learning models. Claims from those method...
042
Reposted by Matteo Pagliardini
Martin Jaggi @mjaggi.bsky.social · 23/04/2025
Using the 'right' data can hugely speed up LLM training, but how to find the best training data in the vast sea of a whole web crawl? We propose a simple classifier-based selection, enabling multilingual LLMs 🧵
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
182
Reposted by Matteo Pagliardini
davidgrangier.bsky.social @davidgrangier.bsky.social · 22/04/2025
#ICLR #TrainBetterLM I am at ICLR, come to our posters for improved language model training! Recycle gradients for faster neural net training with AdEMAmix iclr.cc/virtual/2025... (Fri Apr 25, 10 am). 1/3
123
Reposted by Matteo Pagliardini
Anastasia Koloskova @koloskova.bsky.social · 06/03/2025
I am excited to announce that I will join the University of Zurich as an assistant professor in August this year! I am looking for PhD students and postdocs starting from the fall. My research interests include optimization, federated learning, machine learning, privacy, and unlearning.
1285
Reposted by Matteo Pagliardini
Martin Jaggi @mjaggi.bsky.social · 04/03/2025
The Swiss AI Initiative has launched open calls for disruptive ideas - Democratizing large-scale AI for the benefit of society. Send your idea by end of March 🏃‍♂️‍➡️ , and run on one of the largest public AI clusters globally. Everyone is eligible to apply! swiss-ai.org
Swiss AI Initiative Logo
01711
Reposted by Matteo Pagliardini
Ambroise Odonnat @ambroiseodt.bsky.social · 28/02/2025
🤗Thanks a lot @haeggee.bsky.social and @mjaggi.bsky.social for having me in the MLO group at EPFL @icepfl.bsky.social to present "Large Language Models as Markov Chains". Slides are available on my website (link in thread). 🎉 New experiments with Llama and Gemma models in the updated paper!
142
Reposted by Matteo Pagliardini
Ramon @noctrog.bsky.social · 14/02/2025
What is the true depth of an LLM? Together with @danielepal.bsky.social , @matpagliardini.bsky.social, M. Jaggi and @francois.fleuret.org we show that LLMs have a smaller effective depth that can be exploited to increase inference speeds on multi-GPU settings! arxiv.org/abs/2502.02790 (1/N)
1133
Reposted by Matteo Pagliardini
Jonas @jonasgeiping.bsky.social · 10/02/2025
Ok, so I can finally talk about this! We spent the last year (actually a bit longer) training an LLM with recurrent depth at scale. The model has an internal latent space in which it can adaptively spend more compute to think longer. I think the tech report ...🐦‍⬛
1267
Reposted by Matteo Pagliardini
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 07/02/2025
can we scale small, open LMs to o1 level? Using classical probabilistic inference methods, YES! Particle filtering approach to Improved inference w/o any training! Check out probabilistic-inference-scaling.github.io By Aisha Puri et al📈🤖 Joint MIT-CSAIL & RedHat
probabilistic-inference-scaling.github.io
Probabilistic Inference Scaling
Probabilistic Inference Scaling
1475
Reposted by Matteo Pagliardini
Martin Jaggi @mjaggi.bsky.social · 30/01/2025
new open weights, 24B model, with comparable performance to Llama 3.3 70B 😮. congrats mistral team! mistral.ai/news/mistral...
0122
Reposted by Matteo Pagliardini
Antoine Bosselut @abosselut.bsky.social · 04/12/2024
1/ 📘 Could ChatGPT get an engineering degree? Spoiler, yes! In our new @pnas.org article, we explore how AI assistants like GPT-4 perform in STEM university courses — and on average they pass a staggering 91.7% of core courses. 🧵 #AI #HigherEd #STEM #LLMs #NLProc
13614
Reposted by Matteo Pagliardini
Kosta Derpanis @csprofkgd.bsky.social · 27/11/2024
New blog post on flow matching: dl.heeere.com/cfm/ Contains some nice visuals too!
3708