Sign in

François Fleuret

@francois.fleuret.org
5.8K followers 234 following 413 posts

Research Scientist Meta/FAIR, Prof. University of Geneva, co-founder Neural Concept SA. I like reality. fleuret.org

PostsRepliesMedia
François Fleuret @francois.fleuret.org · 10/09/2025
The voc corresponding to the logits
010
François Fleuret @francois.fleuret.org · 10/09/2025
3293
François Fleuret @francois.fleuret.org · 28/04/2025
- Ring Attention: takes advantage of multi-node hardware to scale the computation according to the sequence length - Speculative decoding: a cheaper model generates tokens, and a rejection process corrects this generation to march the full-model distribution.
1150
François Fleuret @francois.fleuret.org · 28/04/2025
- Multi-token prediction: sums the training over multiple future tokens, possibly with additional readout heads. - FlashAttention: computes the attention on the fly, avoiding a memory footprint O(T^2) (+ optimizes very carefully for the GPU!)
1130
François Fleuret @francois.fleuret.org · 28/04/2025
- Warmup: very short ramping-up of the learning rate, starting from 0 - Cosine schedule: the learning rate varies less at the beginning and end of the schedule - AdamW: decouples weight includes decay from Adam
1130
François Fleuret @francois.fleuret.org · 28/04/2025
- RoPE (Rotary Positional Embedding): makes the attention depend only on the relative Q/K positions - MoE (Mixture of Experts): The FFN block is implemented with multiple MLPs and a gating mechanism selects which ones process each token.
1120
François Fleuret @francois.fleuret.org · 28/04/2025
- RMSNorm instead of Layernorm: normalize only the scaling - MLA (Multi-head Latent Attention): stores a low-rank projection of the attention block input and compute the KV from it - SwiGLU: non-linearity for the FFN block with per-component gating
1120
François Fleuret @francois.fleuret.org · 28/04/2025
- Prenorm: normalization in the residual blocks before the attention operation and the FFN respectively - GQA (Group Query Attention): more Q than (K, V)
1150
François Fleuret @francois.fleuret.org · 28/04/2025
I asked "on the other platform" what were the most important improvements to the original 2017 transformer. That was quite popular and here is a synthesis of the responses:
420643
François Fleuret @francois.fleuret.org · 01/04/2025
"You are in Paris, enjoy the city, stop obsessing with AI" Paris:
2361
François Fleuret @francois.fleuret.org · 12/03/2025
Yes, it's awesome. The kind of work that opens up a whole new and important field.
010
François Fleuret @francois.fleuret.org · 28/02/2025
If your task is not resolution-agnostic, do not use normalized p-e. All this being said, putting both normalized and non-normalized cannot hurt methinks.
100
François Fleuret @francois.fleuret.org · 28/02/2025
You cannot be better off without p-e.
010
François Fleuret @francois.fleuret.org · 28/02/2025
Why not a normalized positional encoding?
100
François Fleuret @francois.fleuret.org · 28/02/2025
After a long lecture, I recommend a coffee, a pain au chocolat, and leave-me-the-fuck-alone time.
020
François Fleuret @francois.fleuret.org · 28/02/2025
Maybe the wall was the friends we made during that journey Ted.
071
François Fleuret @francois.fleuret.org · 27/02/2025
Why is it spooky?
210
François Fleuret @francois.fleuret.org · 21/02/2025
I asked this because even though I am interested in the topic, I have not met so far "foundational" theory regarding the future of society with AI. Someone linked this paper which is exactly the sort of thing I was looking for: arxiv.org/abs/2502.12102
arxiv.org
Relational Norms for Human-AI Cooperation
How we should design and interact with social artificial intelligence depends on the socio-relational role the AI is meant to emulate or occupy. In human society, relationships such as teacher-student...
260
Reposted by François Fleuret
Ramon @noctrog.bsky.social · 14/02/2025
What is the true depth of an LLM? Together with @danielepal.bsky.social , @matpagliardini.bsky.social, M. Jaggi and @francois.fleuret.org we show that LLMs have a smaller effective depth that can be exploited to increase inference speeds on multi-GPU settings! arxiv.org/abs/2502.02790 (1/N)
1133
François Fleuret @francois.fleuret.org · 11/02/2025
We can't complain, can we?
110
François Fleuret @francois.fleuret.org · 11/02/2025
J'étais l'invité du journal de 19h30 sur la @radiotelesuisse.bsky.social ce soir pour parler d'Intelligence Artificielle. www.rts.ch/play/tv/19h3...
rts.ch
19h30 - Play RTS
Play RTS vous permet de visionner ou d'écouter de nombreuses émissions tv ou radio, quand et aussi souvent que vous le souhaitez.
1123
François Fleuret @francois.fleuret.org · 06/02/2025
To do so, you concatenate all the sequences to make a batch of a single sequence, and carve the attention matrix into a block-diagonal one (possibly with causal structure in each block) so that sequences cannot look at each other. Magic! 3/3
180
François Fleuret @francois.fleuret.org · 06/02/2025
It does this by generating an optimized cuda kernel on the fly. So it's cool for causal masks, but it also allows an amazing trick to deal with batches of sequences of various lengths *without padding*! 2/3
140
François Fleuret @francois.fleuret.org · 06/02/2025
It is hard to overstate how cool and powerful is flex attention. @chhillee.bsky.social pytorch.org/blog/flexatten… TL;DR: it is an implementation of the attention operator in pytorch that allows in particular to efficiently "carve" the attention matrix. 1/3
pytorch.org
2555
François Fleuret @francois.fleuret.org · 05/02/2025
I have to admit I am more on the other platform.
100
François Fleuret @francois.fleuret.org · 31/01/2025
I'm very happy then! Thanks for the feedback.
010
François Fleuret @francois.fleuret.org · 24/01/2025
Does it match your expectations?
100
François Fleuret @francois.fleuret.org · 21/01/2025
The big city...
120
Reposted by François Fleuret
Kajetan Dymkiewicz @kdymkiewicz.bsky.social · 10/01/2025
Finally got this beautiful piece from @francois.fleuret.org
1241
François Fleuret @francois.fleuret.org · 31/12/2024
Happy new year you all! 2025 is certainly full of promise.
0300
François Fleuret @francois.fleuret.org · 31/12/2024
Don't worry we'll recycle this year.
000
François Fleuret @francois.fleuret.org · 25/12/2024
Happy Christmas you all!
0280
François Fleuret @francois.fleuret.org · 22/12/2024
Whatever you say about the whole field of AI: It's not boring.
4402
François Fleuret @francois.fleuret.org · 20/12/2024
Wow looks awesome.
010
François Fleuret @francois.fleuret.org · 20/12/2024
Agreed for the anthropomorphisation, they obviously decided (rightly or not...) that it's the proper/most efficient way to talk about those models.
010
François Fleuret @francois.fleuret.org · 20/12/2024
Redefine capslock as the master key that does everything
130
François Fleuret @francois.fleuret.org · 20/12/2024
Some tools that keep me sane on Mac: rectangleapp.com karabiner-elements.pqrs.org
rectangleapp.com
Rectangle
Move and resize windows in macOS using keyboard shortcuts or snap areas. The official page for Rectangle.
5240
François Fleuret @francois.fleuret.org · 20/12/2024
It's Friday!
1131
François Fleuret @francois.fleuret.org · 19/12/2024
0140
François Fleuret @francois.fleuret.org · 19/12/2024
Oh boy, GTA 6 has to be good. And Half Life 3.
150
François Fleuret @francois.fleuret.org · 19/12/2024
But not working with that person remains the thing to do though.
020
François Fleuret @francois.fleuret.org · 19/12/2024
You should just have said "Yes, yes definitely." I teach my kids: stupid question, stupid answer.
241
François Fleuret @francois.fleuret.org · 19/12/2024
This is very great.
1130
François Fleuret @francois.fleuret.org · 18/12/2024
I shifted nothing, your response was an illustration of my initial post. For what it is worth, I am convinced what I do not like in your discourse comes from a good place. I only think it is misguided and terribly counter productive in the long run.
000
François Fleuret @francois.fleuret.org · 18/12/2024
Exactly my point: you anticipate what people will do with knowing certain aspects of reality and you purposefully communicate a description of the world that you think will make people behave as you think is good / moral, and which will IMO make people believe false things regarding reality.
100
François Fleuret @francois.fleuret.org · 18/12/2024
Because the discourse that equates "being true" with "being actionable" is understood by the general audience as a discourse that says true things are false. This is my point from the beginning.
100
François Fleuret @francois.fleuret.org · 18/12/2024
bsky.app/profile/fran...
000
François Fleuret @francois.fleuret.org · 18/12/2024
TBH this exchange is such a perfect illustration of my point. Since you see a fact as having no operational value, you want to push it outside the discussion, and you use phrasing that for someone not used to such nuanced discussion sound like it is not true.
100
François Fleuret @francois.fleuret.org · 18/12/2024
For me it is like saying it is not pressure that keep an airplane in the air but the engines, and every time someone comes back to clarifying how pressure matters someone educated says "that's not pressure that matters, it's the engines!!!"
100
François Fleuret @francois.fleuret.org · 18/12/2024
How is it irrelevant? Anything you will do to "fix the problem" will either reduce the intake or increase the expenditure. By moving away from the obvious facts you participate to a discourse that confuses people who are missing the scientific culture you see as "trivial".
100