Sign in

Nicolas Dufour

@nicolasdufour.bsky.social
489 followers 433 following 75 posts

Postdoc at Kyutai nicolas-dufour.github.io

PostsRepliesMedia
Reposted by Nicolas Dufour
Lucas Degeorge @lucasdegeorge.bsky.social · 03/09/2026
🚀 New paper: Balancing Frequencies and Pixels in Flow Matching We tackle the low-frequency bias in pixel-space flow matching and train JiT up to 40% faster without any architectural changes. 📄 Read it here: arxiv.org/abs/2609.02748
1248
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 03/09/2026
It's a cool paper, both old school and new school. There's an Easter egg: I designed a game for the teaser. Did you get it right? I tried it in my ML course and almost nobody got it.
3194
Reposted by Nicolas Dufour
lebellig @lebellig.bsky.social · 21/07/2026
The last project of my PhD is finally out! 🪴 It was a pleasure collaborating with Aimi on this work! We introduce A²BM: Alignment-Aware Bridge Matching, a new framework for image-to-image translation with weakly aligned image pairs. Paper 📄: arxiv.org/pdf/2607.16294
1144
Nicolas Dufour @nicolasdufour.bsky.social · 21/07/2026
Highly recommend! A great place to meet people passioned about generative models
030
Reposted by Nicolas Dufour
Mathurin Massias @mathurinmassias.bsky.social · 21/07/2026
I organize a 1 day workshop on Generative modelling @ENS Lyon, October 9th Call for oral/poster contributions is open; details at gdr-iasis.cnrs.fr/reunions/mod...
gdr-iasis.cnrs.fr
Modèles génératifs : diffusion, flow matching - GdR IASIS
Les demandes de prise en charge de missions par le GdR IASIS doivent parvenir à la gestionnaire du GdR avant le 25 septembre. Les modèles génératifs ont connu de récentes avancées spectaculaires, au p...
0128
Reposted by Nicolas Dufour
Sonat Baltacı @sonatbaltaci.bsky.social · 20/07/2026
Excited to share that our paper on sprite-based image decomposition is accepted at TMLR! 🎉 Sprite-based models are highly interpretable but struggle to scale to complex, multi-object images. 
 1/3
1204
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 08/07/2026
There's a demo of the model btw: huggingface.co/spaces/nicol...
huggingface.co
MIRO - a Hugging Face Space by nicolas-dufour
Multi-reward conditioned text-to-image diffusion (ICML 2026)
0173
Nicolas Dufour @nicolasdufour.bsky.social · 08/07/2026
If you want to learn more about MIRO, an alignement friendly pre-training that allow us to beat Flux dev with a 350M param model, checkout our blogpost: nicolas-dufour.github.io/miro/
nicolas-dufour.github.io
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
Train once, align many rewards. MIRO achieves 19× faster convergence and 370× less compute than FLUX while reaching GenEval score of 75. Controllable trade-offs at inference time.
010
Nicolas Dufour @nicolasdufour.bsky.social · 08/07/2026
I'm sadly not at ICML, but @arrijitghosh.bsky.social and @lucasdegeorge.bsky.social are presenting MIRO right now! Happening right now at poster board 2508!
1213
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 06/07/2026
1/ MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency. Wed, Jul 8, 2026 • 10:30 AM – 12:15 PM KST 📜 arxiv.org/abs/2510.25897 🖥️ nicolas-dufour.github.io/miro/ 🧬 huggingface.co/nicolas-dufo... With @arrijitghosh.bsky.social and @lucasdegeorge.bsky.social being on site
arxiv.org
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typi...
182
Reposted by Nicolas Dufour
Guillaume Astruc @gastruc.bsky.social · 23/06/2026
🛰️ Introducing UniverSat: one transformer backbone for Earth Observation that handles ANY sensor, ANY spatial, spectral & temporal resolution, ANY scale — with a single set of weights. 🌍
1159
Reposted by Nicolas Dufour
Kwang Moo Yi @kmyid.bsky.social · 22/06/2026
Dufour et al., "The FID Lottery: Quantifying Hidden Randomness in Generative Model Evaluation" We all know it's expensive to train multiple times, but we are now at a point where it is inevitable. Statistical significance should not be ignored. Don't bold over 1~2% differences.
1123
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 19/06/2026
I'll need to track how many time I refer to this paper. It's probably going to be my new language filler.
1124
Nicolas Dufour @nicolasdufour.bsky.social · 19/06/2026
Generative modeling injects noise into training so variability seems to be even more of an issue than classical deep learning. Check out the paper and make sure to check our interactive blog post!
010
Nicolas Dufour @nicolasdufour.bsky.social · 19/06/2026
This work was partly inspired by the amazing torch.manual_seed(3407) is all you need arxiv.org/abs/2109.08203 by my PhD advisor @davidpicard.eurosky.social It recently resurfaced with @karpathy.bsky.social autoresearch optimizing the seed of nano gpt!
arxiv.org
Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision
In this paper I investigate the effect of random seed selection on the accuracy when using popular deep learning architectures for computer vision. I scan a large amount of seeds (up to $10^4$) on CIF...
130
Nicolas Dufour @nicolasdufour.bsky.social · 19/06/2026
We explored the impact of variability sources in generative modeling. Turns out, we've been neglecting the error bars associated with training variability all along! We should aim to report results that we are sure of their scientific validity, instead of seed engineering!
1174
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 19/06/2026
Babe, stop everything! New favorite paper of the year is out! kyutai.org/fid-lottery/ arxiv.org/abs/2606.20536
kyutai.org
The FID Lottery — Quantifying Hidden Randomness in Generative Model Evaluation
An interactive companion to 'The FID Lottery'. Every reported FID is the outcome of two lotteries — we measure how much they move the number.
2276
Reposted by Nicolas Dufour
Kyutai @kyutai-labs.bsky.social · 16/06/2026
Hypnotizing to watch. Great work, co-authored by our very own @nicolasdufour.bsky.social
091
Reposted by Nicolas Dufour
Vision and Graphics Trends @si-cv-graphics.bsky.social · 15/06/2026
𝗦𝘂𝗿𝗳𝗹𝗼: 𝗖𝗼𝗻𝘀𝗶𝘀𝘁𝗲𝗻𝘁 𝟯𝗗 𝗦𝘂𝗿𝗳𝗮𝗰𝗲 𝗙𝗹𝗼𝘄 𝗠𝗼𝗱𝗲𝗹 𝘄𝗶𝘁𝗵 𝗚𝗹𝗼𝗯𝗮𝗹 𝗦𝘁𝗮𝘁𝗲 Antoine Guédon, Shu Nakamura, Nicolas Dufour ... Angjoo Kanazawa arxiv.org/abs/2606.13644 Trending on scholar-inbox.com
074
Reposted by Nicolas Dufour
Antoine Guédon @antoine-guedon.bsky.social · 16/06/2026
What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching? Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵 1/N 🔗https://anttwo.github.io/surflo/
13512
Reposted by Nicolas Dufour
Antoine Guédon @antoine-guedon.bsky.social · 16/06/2026
We will release our code+data asap, please stay tuned! Thanks to my amazing coauthors: @nicolasdufour.bsky.social Shu Nakamura Jiahui Lei @kyotovision.bsky.social @akanazawa.bsky.social 📜arXiv: arxiv.org/abs/2606.13644 🔗Project: anttwo.github.io/surflo/ 💻Code (soon): github.com/Anttwo/Surflo 8/8
arxiv.org
Surflo: Consistent 3D Surface Flow Model with Global State
Geometry is invariant to viewpoint, which makes any collection of images a redundant encoding of a single 3D state. Existing feed-forward reconstruction models fail to exploit this: per-view methods e...
051
Nicolas Dufour @nicolasdufour.bsky.social · 16/06/2026
Check out our latest work! 🚀 We learn a global state and decode the point cloud pointwise, allowing to decode as many points as you want. Plus, we introduce some clever guidance tricks to ensure global consistency, yielding high-quality meshes from just a few views! 👇
080
Reposted by Nicolas Dufour
Zhenjun Zhao @ericzzj.bsky.social · 12/06/2026
Surflo: Consistent 3D Surface Flow Model with Global State @antoine-guedon.bsky.social, Shu Nakamura, @nicolasdufour.bsky.social, Jiahui Lei, Ko Nishino, @akanazawa.bsky.social arxiv.org/abs/2606.13644
1103
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 19/05/2026
Playing with some post-processing on a small in-house-but-soon-to-be-released model.
1152
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 20/05/2026
The model is so fast and easy to use that I vibe-coded a small game with it in 1h 😅 Runs flawlessly on a consumer GPU if you're looking for a small local model to tinker with.
0203
Reposted by Nicolas Dufour
Kashyap Chitta @kashyap7x.bsky.social · 20/05/2026
Open-source fueled the LLM revolution, but Physical AI hasn't fully benefited from this flywheel yet. Today, we're launching kesai.eu, our mission to democratize robotics research! First milestone: training a frontier-level self-driving policy using significantly less data than typically required.
kesai.eu
KE:SAI — Open Science Autonomy Lab
KE:SAI is a Franco-German non-profit open science lab for scalable autonomous intelligence.
1316
Nicolas Dufour @nicolasdufour.bsky.social · 20/05/2026
Everything is fully open-sourced, including the codebase, the model + all individual single reward model variants! 🌐 Site: nicolas-dufour.github.io/miro 📄 Paper: arxiv.org/abs/2510.25897 🛠️ Git: github.com/nicolas-dufo... 🤗 HF: huggingface.co/nicolas-dufo... 🎨 Demo: huggingface.co/spaces/nicol...
0122
Nicolas Dufour @nicolasdufour.bsky.social · 20/05/2026
It's not just for training from scratch. Our new results show that MIRO functions beautifully as a post-training framework! Applying this multi-reward conditioning during the fine-tuning phase of an existing base model yields the exact same controllable alignment.
130
Nicolas Dufour @nicolasdufour.bsky.social · 20/05/2026
Are all rewards useful? Yes! Our new "leave-one-out" ablation shows that removing even a single reward drops performance. Even though these rewards are quite entangled, each one still provides unique, useful bits of information that the model needs to succeed.
120
Nicolas Dufour @nicolasdufour.bsky.social · 20/05/2026
Why does MIRO work so reliably? Our paper introduces a theorem proving that conditioning on the joint reward distribution mathematically guarantees that the model steers toward high-reward regions while preserving sample diversity and avoiding single-metric hacking.
130
Nicolas Dufour @nicolasdufour.bsky.social · 20/05/2026
Thrilled to share that MIRO is accepted to ICML 2026 @icmlconf.bsky.social ! 🎉 By training on the reward scores, we can simply condition the model on high rewards at inference time to guarantee top-tier, aligned outputs. We’ve updated our paper with some additional results!
arxiv.org
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typi...
1366
Reposted by Nicolas Dufour
Nicolas Dufour @nicolasdufour.bsky.social · 31/10/2025
We introduce MIRO: a new paradigm for T2I model alignment integrating reward conditioning into pretraining, eliminating the need for separate fine-tuning/RL stages. This single-stage approach offers unprecedented efficiency and control. - 19x faster convergence ⚡ - 370x less FLOPS than FLUX-dev 📉
37116
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 15/05/2026
👏 Folks! If you are curious about the Generative Modeling via Drifting paper, but you find it difficult to understand → I wrote a different interpretation of it. It's called: "An Expectation-Maximization interpretation of Generative Modeling via Drifting" davidpicard.github.io/pdf/An_Expec...
2285
Nicolas Dufour @nicolasdufour.bsky.social · 01/05/2026
Congrats! Looking forward to what you will do in the future!
110
Reposted by Nicolas Dufour
Guillaume Astruc @gastruc.bsky.social · 16/04/2026
Excited to share my work as a Student Researcher at Google Zurich: UniGeoCLIP! 🌍🚀 W/ Eduard Trulls, Jan Hosang, @loicland.bsky.social & @pesarlin.bsky.social , we built a framework aligning 5 geospatial modalities in one space. Presented at EarthVision @ #CVPR2026. 🧵👇
1116
Reposted by Nicolas Dufour
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
New paper: Back into Plato’s Cave Are vision and language models converging to the same representation of reality? The Platonic Representation Hypothesis says yes. BUT we find the evidence for this is more fragile than it looks. Project page: akoepke.github.io/cave_umwelten/ 1/9
akoepke.github.io
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
25715
Nicolas Dufour @nicolasdufour.bsky.social · 16/04/2026
Checkout our recent work, where we only need web images to learn a novel view generation model! We can navigate inside any image, without any video/multi view data or prior models! Congrats to Adrien for this great first PhD paper! (with @davidpicard.eurosky.social and @ptrkprz.bsky.social)
0142
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 15/04/2026
I thought I would do a thread, but honestly the post is so good: kyutai.org/blog/2026-04... It explains "One View Is Enough! Monocular Training for In-the-Wild Novel View Generation" arxiv.org/abs/2603.23488 done in colab with the smart people at kyutai
kyutai.org
OVIE: One View Is Enough!
Our mission is to build and democratize artificial general intelligence through open science.
0164
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 08/04/2026
🚨 arxiv.org/abs/2604.06129 PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer This paper is the result of doing a lab-wide hackathon on an idea I've had for some time. Probably the paper with the highest number of authors I've ever done. It's a CVPR Findings 26. Thread 🧵👇
arxiv.org
PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates input tokens into a comp...
46520
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 01/04/2026
🚨 Happy to announce CVPR@Paris'26 which will take place on June 1st in Paris. The goal of the event is to share a little bit of the conference before it happens. We will have poster sessions as well as several plenary talks by world-class speakers. info: cvprinparis.github.io/CVPR2026InPa...
cvprinparis.github.io
CVPR@Paris 2026 June 1st
44712
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 25/03/2026
arxiv.org/abs/2603.23488 👀 I'll make a detailed thread later
arxiv.org
One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
Monocular novel-view synthesis has long required multi-view image pairs for supervision, limiting training data scale and diversity. We argue it is not necessary: one view is enough. We present OVIE, ...
0174
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 29/01/2026
I'm commenting that number on slack with @nicolasdufour.bsky.social and I just realized that if you add the 16k active submissions at CVPR, even considering a sizeable overlap between the 2, there are currently well over 30k active papers in review. That's nuts
182
Nicolas Dufour @nicolasdufour.bsky.social · 12/01/2026
Sadly i don't think DroPE will work for images / videos. Both NoPE and DroPE rely on the causal mask to leak absolute PE. The number of tokens in the attention gets leaked because you can encode a bias that grows with the number of tokens. So not a fix for images yet =(
010
Reposted by Nicolas Dufour
Giorgos Tolias @gtolias.bsky.social · 28/11/2025
It was a big pleasure to be in Nicolas's committee. Congratulations to Nicolas for the great work, and congratulations to the advisors too!
051
Nicolas Dufour @nicolasdufour.bsky.social · 27/11/2025
Apparently some people reported knowing of the bug before 11th of november so even before the release of the reviews
000
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 27/11/2025
Yesterday, @nicolasdufour.bsky.social defended is PhD. I really enjoyed the years of collaboration w/ @vickykalogeiton.bsky.social (& @loicland.bsky.social) Video: youtube.com/live/DXQ7FZA... Big thanks to the jury @dlarlus.bsky.social @ptrkprz.bsky.social @gtolias.bsky.social A. Efros & T. Karras
youtube.com
Efficient Generative models through Conditioning - Nicolas Dufour PhD defense
YouTube video by Nicolas Dufour
1283
Reposted by Nicolas Dufour
Diane Larlus @dlarlus.bsky.social · 27/11/2025
Congrats Nicolas ! On the PhD and on those beautifully crafted slides 🤩
071
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 26/11/2025
Nicolas ( @nicolasdufour.bsky.social ) is defending his PhD right now. I was so in awe of the presentation that I even forgot to take pictures 😅
2272
Nicolas Dufour @nicolasdufour.bsky.social · 18/11/2025
Yes it's latent space just because i had my setup that way. Might try in pixel space in the future.
000
Nicolas Dufour @nicolasdufour.bsky.social · 18/11/2025
Yes it's the raw prediction, we predict the velocity directly
110