Sign in

Mickael Chen

@mickaelchen.bsky.social
101 followers 97 following 15 posts

Research Multimodal Generative AI, and now robotics. Generating MNIST digits for a decade.

PostsRepliesMedia
Mickael Chen @mickaelchen.bsky.social · 28/05/2026
DiffusionBlocks pub.sakana.ai/diffusionblo... from @sakanaai.bsky.social is one of these results that unlock a whole new branch of research papers. Implications of reframing as diffusion goes beyond memory efficiency. Test-time inpainting, guidance, and more would find new interpretations and uses.
pub.sakana.ai
DiffusionBlocks: Training Neural Networks One Block at a Time
A principled framework that converts a residual network into independently trainable blocks via a diffusion interpretation, achieving B× memory reduction without sacrificing performance.
010
Mickael Chen @mickaelchen.bsky.social · 25/02/2026
Originally, the goal of generative models was just to capture the distribution of the dataset. Since then, we've shifted to maximize human preference and vibe checks. This must have introduced biases and made the models worse, at least on some aspect, at modelling the data distribution.
281
Reposted by Mickael Chen
Andrei Bursuc @abursuc.bsky.social · 04/12/2025
I'm speaking at #aiPULSE2025 today on Open & re-purposable foundation models for the automotive industry. The morning keynotes talked a lot about open source so my slide here might be timely.
191
Reposted by Mickael Chen
valeo.ai @valeoai.bsky.social · 25/11/2025
🛠️ Already have a complex, pre-trained pipeline? If you are using bilinear interpolation anywhere, NAF acts as a strict drop-in replacement. Just swap it in. No retraining required. It’s literally free points for your metrics.📈
131
Mickael Chen @mickaelchen.bsky.social · 26/11/2025
That was a cool project brillantly led by Ellington Kirby during his internship. We were curious if we could train diffusion models on sets of point coordinates. For images, this is a step towards spatial diffusion, with pixels reorganizing themselves, instead of diffusing in rgb values space only.
032
Mickael Chen @mickaelchen.bsky.social · 08/05/2025
Is it just me or is fucking linkedin taking over some of the functions that twitter used to fill?
100
Reposted by Mickael Chen
David Picard @davidpicard.eurosky.social · 21/03/2025
🔥🔥🔥 CV Folks, I have some news! We're organizing a 1-day meeting in center Paris on June 6th before CVPR called CVPR@Paris (similar as NeurIPS@Paris) 🥐🍾🥖🍷 Registration is open (it's free) with priority given to authors of accepted papers: cvprinparis.github.io/CVPR2025InPa... Big 🧵👇 with details!
713652
Mickael Chen @mickaelchen.bsky.social · 03/03/2025
Wow, neet! Reannotation is key here. Conjecture: As we are get more and more well-aligned text-image data, it will become easier and easier to train models. This will allow us to explore both more streamlined and more exotic training recipes. More signals that exciting times are coming!
122
Mickael Chen @mickaelchen.bsky.social · 28/02/2025
A game changer. A lot of people suspected it *should* work, but actually seeing it in action is something.
110
Reposted by Mickael Chen
valeo.ai @valeoai.bsky.social · 24/02/2025
🚗 Ever wondered if an AI model could learn to drive just by watching YouTube? 🎥👀 We trained a 1.2B parameter model on 1,800+ hours of raw driving videos. No labels. No maps. Just pure observation. And it works! 🤯 🧵👇 [1/10]
1257
Mickael Chen @mickaelchen.bsky.social · 08/02/2025
Bluesky is less engaging because the algorithm is less predatory.
000
Mickael Chen @mickaelchen.bsky.social · 29/01/2025
I'm curious who at Microsoft or OpenAI thought it was a good idea to publicize this narrative. If you are an organisation concered about ethics of training data, now is probably your best chance to act and be heard. www.reuters.com/technology/m...
reuters.com
Microsoft probing if DeepSeek-linked group improperly obtained OpenAI data, Bloomberg News reports
Microsoft and OpenAI are investigating whether data output from OpenAI's technology was obtained in an unauthorized manner by a group linked to Chinese artificial intelligence (AI) startup DeepSeek, Bloomberg News reported on Tuesday.
020
Mickael Chen @mickaelchen.bsky.social · 29/01/2025
The plateau on training scaling and the shift to test-time scaling created favorable conditions for a competitor like DeepSeek to raise and catch up with OpenAI. Nah, I just made that up. Need to put more thoughts into this. 🤔
020
Mickael Chen @mickaelchen.bsky.social · 14/12/2024
We've reached a point where synthetic data is just better and more convenient than messy noisy web-crawled data. It's been true for multimodal data for a while, and semi-automated data as in the Florence-2 paper has been very succesful. arxiv.org/abs/2311.06242
100
Reposted by Mickael Chen
Sander Dieleman @sedielem.bsky.social · 02/12/2024
Better VQ-VAEs with this one weird rotation trick! I missed this when it came out, but I love papers like this: a simple change to an already powerful technique, that significantly improves results without introducing complexity or hyperparameters.
18613
Reposted by Mickael Chen
Michael Tschannen @mtschannen.bsky.social · 02/12/2024
Have you ever wondered how to train an autoregressive generative transformer on text and raw pixels, without a pretrained visual tokenizer (e.g. VQ-VAE)? We have been pondering this during summer and developed a new model: JetFormer 🌊🤖 arxiv.org/abs/2411.19722 A thread 👇 1/
415437
Reposted by Mickael Chen
Chris Offner @chrisoffner3d.bsky.social · 21/11/2024
For AI to be fair and sustainable, we'd need to figure out attribution, i.e. "How much does training sample X contribute to model output Y?" Then the creator of sample X gets paid an amount proportional to what the user paid for the inference call that produced output Y.
251
Mickael Chen @mickaelchen.bsky.social · 23/11/2024
A great place for students interested in AI/CV research internship. It's a very strong team, invested with all of their students. Check it out.
010
Reposted by Mickael Chen
Andrei Bursuc @abursuc.bsky.social · 22/11/2024
ICYMI our PointBeV #CVPR2024 poster here's a quick talk by lead author Loïck Chambon. It brings a change of paradigm in multi-camera bird's-eye-view (BeV) segmentation via a flexible mechanism to produce sparse BeV points that can adapt to situation, task, compute www.linkedin.com/posts/andrei...
linkedin.com
Andrei Bursuc on LinkedIn: #cvpr2024 #cvpr
In case you missed our PointBeV poster at #CVPR2024 here's a quick presentation by the lead author Loïck C.. PointBEV brings a change of paradigm in…
1113
Reposted by Mickael Chen
Andrei Bursuc @abursuc.bsky.social · 20/11/2024
The Cosmos suite of neural tokenizers for images & videos is impressive. Cosmos is trained on diverse high-res imgs & long-vids, scales well for both discrete & continuous tokens, generalizes to multiple domains (robotics, driving, egocentric ...) & has excellent runtime github.com/NVIDIA/Cosmo...
2195
Reposted by Mickael Chen
David Picard @davidpicard.eurosky.social · 18/11/2024
This is ridiculous. And then people will talk about inclusivity and mental health. Sorry to speak my mind so openly, but this has to be the most toxic idea in a very long time.
1142