Sign in

Nicolas Dufour

@nicolasdufour.bsky.social
487 followers 433 following 75 posts

Postdoc at Kyutai nicolas-dufour.github.io

PostsRepliesMedia
Reposted by Nicolas Dufour
Lucas Degeorge @lucasdegeorge.bsky.social · 03/09/2026
🚀 New paper: Balancing Frequencies and Pixels in Flow Matching We tackle the low-frequency bias in pixel-space flow matching and train JiT up to 40% faster without any architectural changes. 📄 Read it here: arxiv.org/abs/2609.02748
1248
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 03/09/2026
It's a cool paper, both old school and new school. There's an Easter egg: I designed a game for the teaser. Did you get it right? I tried it in my ML course and almost nobody got it.
3194
Reposted by Nicolas Dufour
lebellig @lebellig.bsky.social · 21/07/2026
The last project of my PhD is finally out! 🪴 It was a pleasure collaborating with Aimi on this work! We introduce A²BM: Alignment-Aware Bridge Matching, a new framework for image-to-image translation with weakly aligned image pairs. Paper 📄: arxiv.org/pdf/2607.16294
1144
Nicolas Dufour @nicolasdufour.bsky.social · 21/07/2026
Highly recommend! A great place to meet people passioned about generative models
030
Reposted by Nicolas Dufour
Mathurin Massias @mathurinmassias.bsky.social · 21/07/2026
I organize a 1 day workshop on Generative modelling @ENS Lyon, October 9th Call for oral/poster contributions is open; details at gdr-iasis.cnrs.fr/reunions/mod...
gdr-iasis.cnrs.fr
Modèles génératifs : diffusion, flow matching - GdR IASIS
Les demandes de prise en charge de missions par le GdR IASIS doivent parvenir à la gestionnaire du GdR avant le 25 septembre. Les modèles génératifs ont connu de récentes avancées spectaculaires, au p...
0128
Reposted by Nicolas Dufour
Sonat Baltacı @sonatbaltaci.bsky.social · 20/07/2026
Excited to share that our paper on sprite-based image decomposition is accepted at TMLR! 🎉 Sprite-based models are highly interpretable but struggle to scale to complex, multi-object images. 
 1/3
1204
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 08/07/2026
There's a demo of the model btw: huggingface.co/spaces/nicol...
huggingface.co
MIRO - a Hugging Face Space by nicolas-dufour
Multi-reward conditioned text-to-image diffusion (ICML 2026)
0173
Nicolas Dufour @nicolasdufour.bsky.social · 08/07/2026
I'm sadly not at ICML, but @arrijitghosh.bsky.social and @lucasdegeorge.bsky.social are presenting MIRO right now! Happening right now at poster board 2508!
1213
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 06/07/2026
1/ MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency. Wed, Jul 8, 2026 • 10:30 AM – 12:15 PM KST 📜 arxiv.org/abs/2510.25897 🖥️ nicolas-dufour.github.io/miro/ 🧬 huggingface.co/nicolas-dufo... With @arrijitghosh.bsky.social and @lucasdegeorge.bsky.social being on site
arxiv.org
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typi...
182
Reposted by Nicolas Dufour
Guillaume Astruc @gastruc.bsky.social · 23/06/2026
🛰️ Introducing UniverSat: one transformer backbone for Earth Observation that handles ANY sensor, ANY spatial, spectral & temporal resolution, ANY scale — with a single set of weights. 🌍
1159
Reposted by Nicolas Dufour
Kwang Moo Yi @kmyid.bsky.social · 22/06/2026
Dufour et al., "The FID Lottery: Quantifying Hidden Randomness in Generative Model Evaluation" We all know it's expensive to train multiple times, but we are now at a point where it is inevitable. Statistical significance should not be ignored. Don't bold over 1~2% differences.
1123
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 19/06/2026
I'll need to track how many time I refer to this paper. It's probably going to be my new language filler.
1124
Nicolas Dufour @nicolasdufour.bsky.social · 19/06/2026
We explored the impact of variability sources in generative modeling. Turns out, we've been neglecting the error bars associated with training variability all along! We should aim to report results that we are sure of their scientific validity, instead of seed engineering!
1174
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 19/06/2026
Babe, stop everything! New favorite paper of the year is out! kyutai.org/fid-lottery/ arxiv.org/abs/2606.20536
kyutai.org
The FID Lottery — Quantifying Hidden Randomness in Generative Model Evaluation
An interactive companion to 'The FID Lottery'. Every reported FID is the outcome of two lotteries — we measure how much they move the number.
2276
Reposted by Nicolas Dufour
Kyutai @kyutai-labs.bsky.social · 16/06/2026
Hypnotizing to watch. Great work, co-authored by our very own @nicolasdufour.bsky.social
091
Reposted by Nicolas Dufour
Vision and Graphics Trends @si-cv-graphics.bsky.social · 15/06/2026
𝗦𝘂𝗿𝗳𝗹𝗼: 𝗖𝗼𝗻𝘀𝗶𝘀𝘁𝗲𝗻𝘁 𝟯𝗗 𝗦𝘂𝗿𝗳𝗮𝗰𝗲 𝗙𝗹𝗼𝘄 𝗠𝗼𝗱𝗲𝗹 𝘄𝗶𝘁𝗵 𝗚𝗹𝗼𝗯𝗮𝗹 𝗦𝘁𝗮𝘁𝗲 Antoine Guédon, Shu Nakamura, Nicolas Dufour ... Angjoo Kanazawa arxiv.org/abs/2606.13644 Trending on scholar-inbox.com
074
Reposted by Nicolas Dufour
Antoine Guédon @antoine-guedon.bsky.social · 16/06/2026
What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching? Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵 1/N 🔗https://anttwo.github.io/surflo/
13512
Reposted by Nicolas Dufour
Antoine Guédon @antoine-guedon.bsky.social · 16/06/2026
We will release our code+data asap, please stay tuned! Thanks to my amazing coauthors: @nicolasdufour.bsky.social Shu Nakamura Jiahui Lei @kyotovision.bsky.social @akanazawa.bsky.social 📜arXiv: arxiv.org/abs/2606.13644 🔗Project: anttwo.github.io/surflo/ 💻Code (soon): github.com/Anttwo/Surflo 8/8
arxiv.org
Surflo: Consistent 3D Surface Flow Model with Global State
Geometry is invariant to viewpoint, which makes any collection of images a redundant encoding of a single 3D state. Existing feed-forward reconstruction models fail to exploit this: per-view methods e...
051
Nicolas Dufour @nicolasdufour.bsky.social · 16/06/2026
Check out our latest work! 🚀 We learn a global state and decode the point cloud pointwise, allowing to decode as many points as you want. Plus, we introduce some clever guidance tricks to ensure global consistency, yielding high-quality meshes from just a few views! 👇
080
Reposted by Nicolas Dufour
Zhenjun Zhao @ericzzj.bsky.social · 12/06/2026
Surflo: Consistent 3D Surface Flow Model with Global State @antoine-guedon.bsky.social, Shu Nakamura, @nicolasdufour.bsky.social, Jiahui Lei, Ko Nishino, @akanazawa.bsky.social arxiv.org/abs/2606.13644
1103
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 19/05/2026
Playing with some post-processing on a small in-house-but-soon-to-be-released model.
1152
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 20/05/2026
The model is so fast and easy to use that I vibe-coded a small game with it in 1h 😅 Runs flawlessly on a consumer GPU if you're looking for a small local model to tinker with.
0203
Reposted by Nicolas Dufour
Kashyap Chitta @kashyap7x.bsky.social · 20/05/2026
Open-source fueled the LLM revolution, but Physical AI hasn't fully benefited from this flywheel yet. Today, we're launching kesai.eu, our mission to democratize robotics research! First milestone: training a frontier-level self-driving policy using significantly less data than typically required.
kesai.eu
KE:SAI — Open Science Autonomy Lab
KE:SAI is a Franco-German non-profit open science lab for scalable autonomous intelligence.
1316
Nicolas Dufour @nicolasdufour.bsky.social · 20/05/2026
Thrilled to share that MIRO is accepted to ICML 2026 @icmlconf.bsky.social ! 🎉 By training on the reward scores, we can simply condition the model on high rewards at inference time to guarantee top-tier, aligned outputs. We’ve updated our paper with some additional results!
arxiv.org
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typi...
1366
Reposted by Nicolas Dufour
Nicolas Dufour @nicolasdufour.bsky.social · 31/10/2025
We introduce MIRO: a new paradigm for T2I model alignment integrating reward conditioning into pretraining, eliminating the need for separate fine-tuning/RL stages. This single-stage approach offers unprecedented efficiency and control. - 19x faster convergence ⚡ - 370x less FLOPS than FLUX-dev 📉
37116
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 15/05/2026
👏 Folks! If you are curious about the Generative Modeling via Drifting paper, but you find it difficult to understand → I wrote a different interpretation of it. It's called: "An Expectation-Maximization interpretation of Generative Modeling via Drifting" davidpicard.github.io/pdf/An_Expec...
2285
Reposted by Nicolas Dufour
Guillaume Astruc @gastruc.bsky.social · 16/04/2026
Excited to share my work as a Student Researcher at Google Zurich: UniGeoCLIP! 🌍🚀 W/ Eduard Trulls, Jan Hosang, @loicland.bsky.social & @pesarlin.bsky.social , we built a framework aligning 5 geospatial modalities in one space. Presented at EarthVision @ #CVPR2026. 🧵👇
1116
Reposted by Nicolas Dufour
A. Sophia Koepke @askoepke.bsky.social · 17/04/2026
New paper: Back into Plato’s Cave Are vision and language models converging to the same representation of reality? The Platonic Representation Hypothesis says yes. BUT we find the evidence for this is more fragile than it looks. Project page: akoepke.github.io/cave_umwelten/ 1/9
akoepke.github.io
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
25715
Nicolas Dufour @nicolasdufour.bsky.social · 16/04/2026
Checkout our recent work, where we only need web images to learn a novel view generation model! We can navigate inside any image, without any video/multi view data or prior models! Congrats to Adrien for this great first PhD paper! (with @davidpicard.eurosky.social and @ptrkprz.bsky.social)
0142
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 15/04/2026
I thought I would do a thread, but honestly the post is so good: kyutai.org/blog/2026-04... It explains "One View Is Enough! Monocular Training for In-the-Wild Novel View Generation" arxiv.org/abs/2603.23488 done in colab with the smart people at kyutai
kyutai.org
OVIE: One View Is Enough!
Our mission is to build and democratize artificial general intelligence through open science.
0164
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 08/04/2026
🚨 arxiv.org/abs/2604.06129 PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer This paper is the result of doing a lab-wide hackathon on an idea I've had for some time. Probably the paper with the highest number of authors I've ever done. It's a CVPR Findings 26. Thread 🧵👇
arxiv.org
PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates input tokens into a comp...
46520
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 01/04/2026
🚨 Happy to announce CVPR@Paris'26 which will take place on June 1st in Paris. The goal of the event is to share a little bit of the conference before it happens. We will have poster sessions as well as several plenary talks by world-class speakers. info: cvprinparis.github.io/CVPR2026InPa...
cvprinparis.github.io
CVPR@Paris 2026 June 1st
44712
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 25/03/2026
arxiv.org/abs/2603.23488 👀 I'll make a detailed thread later
arxiv.org
One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
Monocular novel-view synthesis has long required multi-view image pairs for supervision, limiting training data scale and diversity. We argue it is not necessary: one view is enough. We present OVIE, ...
0174
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 29/01/2026
I'm commenting that number on slack with @nicolasdufour.bsky.social and I just realized that if you add the 16k active submissions at CVPR, even considering a sizeable overlap between the 2, there are currently well over 30k active papers in review. That's nuts
182
Reposted by Nicolas Dufour
Giorgos Tolias @gtolias.bsky.social · 28/11/2025
It was a big pleasure to be in Nicolas's committee. Congratulations to Nicolas for the great work, and congratulations to the advisors too!
051
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 27/11/2025
Yesterday, @nicolasdufour.bsky.social defended is PhD. I really enjoyed the years of collaboration w/ @vickykalogeiton.bsky.social (& @loicland.bsky.social) Video: youtube.com/live/DXQ7FZA... Big thanks to the jury @dlarlus.bsky.social @ptrkprz.bsky.social @gtolias.bsky.social A. Efros & T. Karras
youtube.com
Efficient Generative models through Conditioning - Nicolas Dufour PhD defense
YouTube video by Nicolas Dufour
1283
Reposted by Nicolas Dufour
Diane Larlus @dlarlus.bsky.social · 27/11/2025
Congrats Nicolas ! On the PhD and on those beautifully crafted slides 🤩
071
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 26/11/2025
Nicolas ( @nicolasdufour.bsky.social ) is defending his PhD right now. I was so in awe of the presentation that I even forgot to take pictures 😅
2272
Reposted by Nicolas Dufour
Lucas Degeorge @lucasdegeorge.bsky.social · 31/10/2025
Check out our new work: MIRO No more post-training alignment! We integrate human alignment right from the start, during pretraining! Results: ✨ 19x faster convergence ⚡ ✨ 370x less compute 💻 🔗 Explore the project: nicolas-dufour.github.io/miro/
nicolas-dufour.github.io
MIRO: Multi-Reward Conditioning for Efficient Text-to-Image Generation
Train once, align many rewards. MIRO achieves 19× faster convergence and 370× less compute than FLUX while reaching GenEval score of 75. Controllable trade-offs at inference time.
093
Reposted by Nicolas Dufour
Michiel Bontenbal @mpbontenbal.eurosky.social · 31/10/2025
Image generation becomes much more energy efficient. 👍
052
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 31/10/2025
I'm super happy about Nicolas' latest work, probably the magnum opus of his PhD. Read the thread for all the great details. The main conclusion I draw from this work is that better pretraining, in particular by conditioning on better data, allows us to train SOTA models at a fraction of the cost.
0304
Nicolas Dufour @nicolasdufour.bsky.social · 31/10/2025
We introduce MIRO: a new paradigm for T2I model alignment integrating reward conditioning into pretraining, eliminating the need for separate fine-tuning/RL stages. This single-stage approach offers unprecedented efficiency and control. - 19x faster convergence ⚡ - 370x less FLOPS than FLUX-dev 📉
37116
Reposted by Nicolas Dufour
Mathurin Massias @mathurinmassias.bsky.social · 24/10/2025
Kickstarting our workshop on Flow matching and Diffusion with a talk by Eric Vanden Eijnden on how to optimize learning and sampling in Stochastic Interpolants! Broadcast available at gdr-iasis.cnrs.fr/reunions/mod...
1155
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 08/10/2025
Final note: I'm (we're) tempted to organize a challenge on that topic as a workshop at a CV conf. ImageNet is the only source of images allowed and then you compete to get the bold numbers. Do you think there would be people in for that? Do you think it would make for a nice competition?
284
Reposted by Nicolas Dufour
Vicky Kalogeiton @vickykalogeiton.bsky.social · 08/10/2025
Very proud of our recent work, kudos to the team! Read @davidpicard.bsky.social’s excellent post for more details or the paper arxiv.org/pdf/2502.21318
arxiv.org
0176
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 25/09/2025
Today is Antoine Guedon's PhD! Already pretty cool visuals right at the start.
1254
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 18/09/2025
Annnnnd it's a reject! Scale is a religion and if you go against it, you're a heretic and you should burn, "despite [the reviewers] final ratings". But scale is still not necessary! Side note: First time swinging reviews up (from 2,2,4,4 to 2,4,4,5) does not get the paper accepted. Strange days.
4183
Reposted by Nicolas Dufour
François Rozet @francois-rozet.bsky.social · 03/09/2025
Does a smaller latent space lead to worse generation in latent diffusion models? Not necessarily! We show that LDMs are extremely robust to a wide range of compression rates (10-1000x) in the context of physics emulation. We got lost in latent space. Join us 👇
1278
Reposted by Nicolas Dufour
David Picard @davidpicard.eurosky.social · 23/08/2025
Next week, I'll be in Strasbourg for the GRETSI (@gretsi-info.bsky.social) to present a small discovery on transformers generalization we made with Simon and Jérémie while working on generative recommender systems. I love these "phase transition" plots. 📜: arxiv.org/abs/2508.03934 Short summary 👇
2155
Nicolas Dufour @nicolasdufour.bsky.social · 18/08/2025
🚀 DinoV3 just became the new go-to backbone for geoloc! It outperforms CLIP-like models (SigLip2, finetuned StreetCLIP)… and that’s shocking 🤯 Why? CLIP models have an innate advantage — they literally learn place names + images. DinoV3 doesn’t.
14614