Sign in

Oussama Zekri

@ozekri.bsky.social
61 followers 120 following 38 posts

discrete diffusion, generative modeling 1 post a day on foundational ML papers Website : oussamazekri.fr Blog : logb-research.github.io

PostsRepliesMedia
Oussama Zekri @ozekri.bsky.social · 13/05/2026
arxiv.org/abs/2101.03961
arxiv.org
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
In deep learning, models typically reuse the same parameters for all inputs. Mixture of Experts (MoE) defies this and instead selects different parameters for each incoming example. The result is a sp...
000
Oussama Zekri @ozekri.bsky.social · 13/05/2026
Some (really) big AI models work like a team of experts. For each piece of text, a router picks one expert to handle it. Only that path runs (so less compute needed!). That is Switch-style Mixture of Experts: more capacity without using the whole model every time.
110
Oussama Zekri @ozekri.bsky.social · 12/05/2026
arxiv.org/abs/2103.00020
arxiv.org
Learning Transferable Visual Models From Natural Language Supervision
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since addition...
000
Oussama Zekri @ozekri.bsky.social · 12/05/2026
CLIP in one picture: put images and captions in the same batch, encode both, make the matching pairs win the similarity matrix. The diagonal is supervision. Everything off the diagonal becomes a negative example!
110
Oussama Zekri @ozekri.bsky.social · 11/05/2026
arxiv.org/abs/2203.05482
arxiv.org
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
The conventional recipe for maximizing model accuracy is to (1) train multiple models with various hyperparameters and (2) pick the individual model which performs best on a held-out validation set, d...
000
Oussama Zekri @ozekri.bsky.social · 11/05/2026
Model Soups are one of those results that feel like they should not work, but somehow often do 🤷 Start from one base model, fine-tune it several times, then simply average the matching weights. Surprisingly, the soup model can outperform individual ones, while still being just 1 model at inference
100
Oussama Zekri @ozekri.bsky.social · 09/05/2026
arxiv.org/abs/1706.03762
arxiv.org
Attention Is All You Need
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and d...
010
Oussama Zekri @ozekri.bsky.social · 09/05/2026
Self-attention was popularized in language, then spread almost everywhere in machine learning: vision, audio, biology, robotics… The core idea is simple: each token asks which other tokens matter right now. It scores them, turns the scores into soft weights, and mixes information accordingly.
110
Oussama Zekri @ozekri.bsky.social · 08/05/2026
arxiv.org/abs/2106.09685
arxiv.org
LoRA: Low-Rank Adaptation of Large Language Models
An important paradigm of natural language processing consists of large-scale pre-training on general domain data and adaptation to particular tasks or domains. As we pre-train larger models, full fine...
000
Oussama Zekri @ozekri.bsky.social · 08/05/2026
Fine-tuning a whole LLM is often wasteful when the pretrained weights already know most of the work. LoRA freezes the big weight matrix and learns a small low-rank update around it. Same forward pass shape, but the task-specific part becomes tiny.
120
Oussama Zekri @ozekri.bsky.social · 07/05/2026
arxiv.org/abs/2104.09864
arxiv.org
RoFormer: Enhanced Transformer with Rotary Position Embedding
Position encoding recently has shown effective in the transformer architecture. It enables valuable supervision for dependency modeling between elements at different positions of the sequence. In this...
000
Oussama Zekri @ozekri.bsky.social · 07/05/2026
Position in transformers is not always added as a learned embedding. RoPE injects position by rotating query and key vectors according to each token’s place in the sequence. This makes attention scores depend on relative distance, one reason RoPE became standard in modern LLMs.
100
Oussama Zekri @ozekri.bsky.social · 06/05/2026
arxiv.org/abs/1312.6114
arxiv.org
Auto-Encoding Variational Bayes
How can we perform efficient inference and learning in directed probabilistic models, in the presence of continuous latent variables with intractable posterior distributions, and large datasets? We in...
000
Oussama Zekri @ozekri.bsky.social · 06/05/2026
Modern generative models often optimize through random variables (VAEs, diffusion models...). The reparameterization trick rewrites sampling as parameters plus external noise passed through a deterministic function. Same stochastic computation, but now gradients have a path through it!
100
Oussama Zekri @ozekri.bsky.social · 05/05/2026
arxiv.org/abs/2207.12598
arxiv.org
Classifier-Free Diffusion Guidance
Classifier guidance is a recently introduced method to trade off mode coverage and sample fidelity in conditional diffusion models post training, in the same spirit as low temperature sampling or trun...
000
Oussama Zekri @ozekri.bsky.social · 05/05/2026
That guidance scale slider in diffusion models is not magic prompt strength! Classifier-free guidance compares an unconditional denoising direction with a prompt-conditioned one, then amplifies the difference. More prompt control, until the image gets brittle.
100
Oussama Zekri @ozekri.bsky.social · 04/05/2026
arxiv.org/abs/1904.09751
arxiv.org
The Curious Case of Neural Text Degeneration
Despite considerable advancements with deep neural language models, the enigma of neural text degeneration persists when these models are tested as text generators. The counter-intuitive empirical obs...
000
Oussama Zekri @ozekri.bsky.social · 04/05/2026
LLMs do not sample from every possible next token. Nucleus sampling (top-p) keeps the smallest high-probability set whose mass passes p, then samples inside it. Less greedy than argmax, less chaotic than sampling from the full distribution.
130
Reposted by Oussama Zekri
Quentin Berthet @ AISTATS @qberthet.bsky.social · 10/02/2025
🚨 New paper on regression and classification! Adding to the discussion on using least-squares or cross-entropy, regression or classification formulations of supervised problems! A thread on how to bridge these problems:
4518
Oussama Zekri @ozekri.bsky.social · 06/02/2025
You mean, we don’t stop at the frontier of the convex set but just a bit further ? Wow, does this trick have a name?
100
Oussama Zekri @ozekri.bsky.social · 05/02/2025
Looks nice!! Will stop by your notebooks
010
Oussama Zekri @ozekri.bsky.social · 04/02/2025
Working with him these past months has been both fun and inspiring. He’s an incredibly talented researcher! 🚀 If you haven’t heard of him, check out his work : he’s one of the pioneers of operator learning and pushing this field to new heights!
nboulle.github.io
Nicolas Boullé
About me
010
Oussama Zekri @ozekri.bsky.social · 04/02/2025
Thanks for reading ! ❤️ Work done during my 3-months internship at Imperial College! A huge thanks to Nicolas Boullé (nboulle.github.io) for letting me work on a topic that interested me a lot during the internship.
nboulle.github.io
Nicolas Boullé
About me
110
Oussama Zekri @ozekri.bsky.social · 04/02/2025
We fine-tuned a discrete diffusion model to respond to user prompts. In just 7k iterations (GPU poverty is real, haha), it outperforms the vanilla model ~75% of the time! 🚀
110
Oussama Zekri @ozekri.bsky.social · 04/02/2025
Building on this, we can correct the gradient direction to better **follow the flow**, using the implicit function theorem (cf @mblondel.bsky.social et al., arxiv.org/abs/2105.15183 )✨ The cool part? We only need to invert a linear system, whose inverse is known in closed form! 🔥
110
Oussama Zekri @ozekri.bsky.social · 04/02/2025
Inspired by Implicit Diffusion (@pierremarion.bsky.social @akorba.bsky.social @qberthet.bsky.social🤓, arxiv.org/abs/2402.05468), we sample using a specific CTMC, reaching the limiting distribution in an infinite time horizon. This effectively implements a gradient flow w.r.t. a Wasserstein metric!🔥
120
Oussama Zekri @ozekri.bsky.social · 04/02/2025
SEPO, like most policy optimization algorithms, alternates between sampling and optimization. But what if sampling itself was seen as an optimization procedure in distribution space? 🚀
120
Oussama Zekri @ozekri.bsky.social · 04/02/2025
If you have a discrete diffusion model (naturally designed for discrete data, e.g. language or DNA sequence modeling), you can finetune it with non-differentiable reward functions! 🎯 For example, this enables RLHF for discrete diffusion models, making alignment more flexible and powerful. ✅
120
Oussama Zekri @ozekri.bsky.social · 04/02/2025
The main gradient takes the form of a weighted log concrete score, echoing DeepSeek’s unified paradigm with the weighted log policy!🔥 From this, we can reconstruct any policy gradient method for discrete diffusion models (e.g. PPO, GRPO etc...). 🚀
110
Oussama Zekri @ozekri.bsky.social · 04/02/2025
The main bottleneck of Energy-Based Models is computing the normalizing constant Z. Instead, recent discrete diffusion models skip Z by learning ratios of probabilities. This forms the concrete score, which a neural network models efficiently!⚡ The challenge? Using this score network as a policy.
110
Oussama Zekri @ozekri.bsky.social · 04/02/2025
🚀 Policy gradient methods like DeepSeek’s GRPO are great for finetuning LLMs via RLHF. But what happens when we swap autoregressive generation for discrete diffusion, a rising architecture promising faster & more controllable LLMs? Introducing SEPO ! 📑 arxiv.org/pdf/2502.01384 🧵👇
162
Oussama Zekri @ozekri.bsky.social · 04/02/2025
Beautiful work!!
010
Reposted by Oussama Zekri
Ambroise Odonnat @ambroiseodt.bsky.social · 04/02/2025
🚀Proud to share our work on the training dynamics in Transformers with Wassim Bouaziz & @viviencabannes.bsky.social @Inria @MetaAI 📝Easing Optimization Paths arxiv.org/pdf/2501.02362 (accepted @ICASSP 2025 🥳) 📝Clustering Heads 🔥https://arxiv.org/pdf/2410.24050 🖥️ github.com/facebookrese... 1/🧵
164
Reposted by Oussama Zekri
Ambroise Odonnat @ambroiseodt.bsky.social · 25/01/2025
Happy to see Disentangled In-Context Learning accepted at ICLR 2025 🥳 Make zero-shot reinforcement learning with LLMs go brrr 🚀 🖥️ github.com/abenechehab/... 📜 arxiv.org/pdf/2410.11711 Congrats Abdelhakim (abenechehab.github.io) for leading it, always fun working with nice and strong people 🤗
github.com
GitHub - abenechehab/dicl: Official implementation of DICL (Disentangled In-Context Learning), featured in the paper Zero-shot Model-based Reinforcement Learning using Large Language Models.
Official implementation of DICL (Disentangled In-Context Learning), featured in the paper Zero-shot Model-based Reinforcement Learning using Large Language Models. - abenechehab/dicl
052
Reposted by Oussama Zekri
lebellig @lebellig.bsky.social · 20/01/2025
For the French-speaking audience, S. Mallat's courses at the College de France on Data generation in AI by transport and denoising have just started. I highly recommend them, as I've learned a lot from the overall vision of his courses. Recordings are also available: www.youtube.com/watch?v=5zFh...
youtube.com
Génération de données en IA par transport et débruitage (1) - Stéphane Mallat (2024-2025)
YouTube video by Mathématiques et informatique - Collège de France
0103
Reposted by Oussama Zekri
Arnaud Doucet @arnauddoucet.bsky.social · 10/01/2025
Speculative sampling accelerates inference in LLMs by drafting future tokens which are verified in parallel. With @vdebortoli.bsky.social , A. Galashov & @arthurgretton.bsky.social , we extend this approach to (continuous-space) diffusion models: arxiv.org/abs/2501.05370
04510
Oussama Zekri @ozekri.bsky.social · 06/12/2024
i couldn’t have say it better myself !
010
Reposted by Oussama Zekri
Konstantin Mishchenko @konstmish.bsky.social · 06/12/2024
The idea that one needs to know a lot of advanced math to start doing research in ML seems so wrong to me. Instead of reading books for weeks and forgetting most of them a year later, I think it's much better to try do things, see what knowledge gaps prevent you from doing them, and only then read.
492
Oussama Zekri @ozekri.bsky.social · 04/12/2024
This equivalence between LLMs and Markov chains seems useless, but it isn't! Among the contributions, the paper highlights bounds established thanks to this equivalence, and verifies the influence of bound terms on recents LLMs ! I invite you to take a look at the other contributions of the paper 🙂
030
Oussama Zekri @ozekri.bsky.social · 04/12/2024
This number is huge, but **finite**! Working with markov chains in a finite state space really gives non-trivial mathematical insights (existence and uniqueness of a stationary distribution for example...).
110
Reposted by Oussama Zekri
Vicki @vickiboykis.com · 03/12/2024
This seems like… what we started with, no? arxiv.org/abs/2410.02724
arxiv.org
Large Language Models as Markov Chains
Large language models (LLMs) have proven to be remarkably efficient, both across a wide range of natural language processing tasks and well beyond them. However, a comprehensive theoretical analysis o...
161659
Reposted by Oussama Zekri
Ambroise Odonnat @ambroiseodt.bsky.social · 03/12/2024
🚨So, you want to predict your model's performance at test time?🚨 💡Our NeurIPS 2024 paper proposes 𝐌𝐚𝐍𝐨, a training-free and SOTA approach! 📑 arxiv.org/pdf/2405.18979 🖥️https://github.com/Renchunzi-Xie/MaNo 1/🧵(A surprise at the end!)
2166
Reposted by Oussama Zekri
Gabriel Peyré @gabrielpeyre.bsky.social · 30/11/2024
I wrote a summary of the main ingredients of the neat proof by Hugo Lavenant that diffusion models do not generally define optimal transport. github.com/mathematical...
523845
Oussama Zekri @ozekri.bsky.social · 26/11/2024
For more details, check out these papers: 👉 arxiv.org/pdf/2402.00795 — Introduces this method (to the best of my knowledge). 👉 arxiv.org/pdf/2410.02724 — Provides theoretical results and empirical validation on LLMs.
arxiv.org
010
Oussama Zekri @ozekri.bsky.social · 26/11/2024
💡For a Markov chain with d states, the LLM-based method achieves an error rate of O(log⁡(d)/N). The frequentist approach, which is minimax optimal, achieves O(d/N). (see Wolfer et al., 2019, arxiv.org/pdf/1902.00080). This makes it particularly efficient for MC with a large number of states! 🌟
100
Oussama Zekri @ozekri.bsky.social · 26/11/2024
‼️What’s even better is that you can derive bounds on the estimation error based on the number of samples N provided and specific properties of the Markov chain. Tested and validated on recent LLMs!
100
Oussama Zekri @ozekri.bsky.social · 26/11/2024
🚀 Did you know you can use the in-context learning abilities of an LLM to estimate the transition probabilities of a Markov chains? The results are pretty exciting ! 😄
182