Sign in

Benjamin Lefaudeux 🇺🇦

@bentheegg.bsky.social
442 followers 928 following 268 posts

Back to France after some time in sunny California and happy Copenhagen. Mistral, Photoroom, Meta (xformers, FairScale, R&D), EyeTribe (acq) Mostly writing around AI

PostsRepliesMedia
Reposted by Benjamin Lefaudeux 🇺🇦
mr. TIM @timkellogg.me · 31/08/2025
Limits of vector search a new GDM paper shows that embeddings can’t represent combinations of concepts well e.g. Dave likes blue trucks AND Ford trucks even k=2 sub-predicates make SOTA embedding models fall apart www.alphaxiv.org/pdf/2508.21038
alphaxiv.org
On the Theoretical Limitations of Embedding-Based Retrieval | alphaXiv
View recent discussion. Abstract: Vector embeddings have been tasked with an ever-increasing set of retrieval tasks over the years, with a nascent rise in using them for reasoning, instruction-followi...
28322
Reposted by Benjamin Lefaudeux 🇺🇦
mr. TIM @timkellogg.me · 31/08/2025
Longcat-Flash-Chat (560B) uh, holy shit this one is intriguing. bare minimum they compare themselves to all the (actual) top models and do okay but inside.. damn this one has some cool ideas huggingface.co/meituan-long...
The image is a multi-panel bar chart comparing performance of different large language models across several benchmarks. It is divided into four categories: General Domains, Agentic Tool Use, Code, and Instruction Following. Each panel has bars representing model results, with scores on the y-axis.

Top row – General Domains:
	•	ArenaHard-V2: LongGPT-Flash leads with 86.5, followed by Kimi K2 (88.2), DeepSeek V3.1 (84.1), Claude Sonnet (61.5), GPT-4.1 (62.1), Qwen3.5 MoE-2507 (85.7), and Gemini 2.5 Flash (77.0).
	•	MMLU-Pro: Best scores are Kimi K2 (84.5) and DeepSeek V3.1 (84.5), with LongGPT-Flash (82.7), Qwen3.5 MoE-2507 (82.1), GPT-4.1 (81.7), Claude Sonnet (83.7), Gemini 2.5 Flash (82.0).

Top row – Agentic Tool Use:
	•	t2-Bench (average): LongGPT-Flash leads (67.7), Kimi K2 (64.2), Claude Sonnet (62.1), GPT-4.1 (55.1), DeepSeek V3.1 (49.8), Qwen3.5 MoE-2507 (43.0), Gemini 2.5 Flash (40.9).
	•	VitaBench: LongGPT-Flash 24.3, Claude Sonnet 23.0, DeepSeek V3.1 20.3, Kimi K2 18.2, GPT-4.1 19.0, Qwen3.5 MoE-2507 8.5, Gemini 2.5 Flash 8.0.

Bottom row – Code:
	•	SWE-Bench-Verified: Claude Sonnet leads with 68.0, Kimi K2 64.6, DeepSeek V3.1 66.0, LongGPT-Flash 60.4, GPT-4.1 48.6, Qwen3.5 MoE-2507 42.0, Gemini 2.5 Flash 40.6.
	•	TerminalBench: Claude Sonnet 40.7, LongGPT-Flash 39.5, DeepSeek V3.1 31.3, GPT-4.1 28.4, Kimi K2 25.9, Qwen3.5 MoE-2507 17.3, Gemini 2.5 Flash 12.4.

Bottom row – Instruction Following:
	•	COLLIE: LongGPT-Flash 57.1, Kimi K2 56.3, Claude Sonnet 51.2, GPT-4.1 50.0, DeepSeek V3.1 49.7, Gemini 2.5 Flash 48.6, Qwen3.5 MoE-2507 43.8.
	•	Meeseeks (ZH): LongGPT-Flash 43.0, Kimi K2 42.8, Claude Sonnet 41.5, GPT-4.1 35.1, DeepSeek V3.1 35.3, Qwen3.5 MoE-2507 33.8, Gemini 2.5 Flash 34.8.
2485
Reposted by Benjamin Lefaudeux 🇺🇦
Ted Underwood @tedunderwood.com · 27/07/2025
In 2012 when I had to clean data it seemed natural to look for rules I could use to clean it. Now it seems natural to model the noise, find new clean data it can destroy, and then train a model to reverse the process. Machine learning makes you a sicko.
1434
Reposted by Benjamin Lefaudeux 🇺🇦
Ethan Mollick @emollick.bsky.social · 26/07/2025
Three things to note about this: 1) AI has obvious utility to many, this is a tremendous amount of use already 2) There is room for multiple frontier model providers, at least for now 3) Any losses from subsidizing cost of AI use (and it is not clear this is happening) are now relatively small
3673
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 26/07/2025
Above is intuitive when you think about it long enough (or so it feels at least), but I missed it entirely during a couple of years working on diffusion, so I figured it was worth emphasizing and the authors did too :)
020
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 26/07/2025
Worth a deep read in general, not personally completely done with it, I hope it ages well. Closing with some nice insight wrt diffusion models: they don't open up for serial awareness, since model iterates on _the same_ solution, no state space + carry over. _Less_ powerful than autoregressive
140
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 26/07/2025
Paper cannot prove its point completely since models are really good approximators, and used as such (hence a formal disprove is not enough). Pretty good hints still, makes me confident we're far from peak efficiency in most use cases (we approx serial awareness by adding tons of compute)
130
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 26/07/2025
I think that hardware recommendations are a little naive/premature, as much as I like CPUs nothing will happen prior to needs and solutions being put on the table. Lowering is expensive and risky in general, will happen last, but at least this shows there's kryptonite to GPU dominance
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 26/07/2025
The paper is very pedagogical, and some takeaways ring pretty reasonable. Intuition is interesting behind LLMs being just ok to not great Chess players (missing the MCTS like mechanism of specialized models), or failing to be effective at multi step reasoning prior to test time compute / CoT
110
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 26/07/2025
It then feels like the dichotomy proposed by the paper (inherently parallel and TC0 models will fail on serial problems) is excessive, or at least that the frontier is a bit fuzzy. One line is great though, paraphrasing "only with test time compute did we factor in some serial compute power"
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 26/07/2025
There are caveats in the definition of "inherently serial" problems: - not all solutions will require serial computations, even for something outside of TC0 - approximations can fall pretty close, and oftentimes we don´t expect anything much better than an approximation
110
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 26/07/2025
"The Serial Scaling Hypothesis" (arxiv.org/abs/2507.125..., Liu et al) is interesting I think, not as new as it completely looks (autoregressive models are used serially, models have depth,..) but feels like a good formalization and intuition as of where current GPT based LLMs will typically fail
1101
Reposted by Benjamin Lefaudeux 🇺🇦
Andrei Bursuc @abursuc.bsky.social · 21/07/2025
1/ Can open-data models beat DINOv2? Today we release Franca, a fully open-sourced vision foundation model. Franca with ViT-G backbone matches (and often beats) proprietary models like SigLIPv2, CLIP, DINOv2 on various benchmarks setting a new standard for open-source research.
28622
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 18/07/2025
Claude Code is really good for some narrowly defined tasks (add unit tests for instance), and in that case it's clearly an agent. The "vibe coding" coding middle ground (with somebody in the loop who doesn't completely get it) is the part on shaky grounds I believe
110
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 18/07/2025
Something the LLMs have not seen beforehand (new model architecture for instance). In my experience that's where all the current tools break, for relatable reasons. I guess it's the same for somebody developing a SOTA DB engine or computer shader
000
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 18/07/2025
For things LLMs are not great at (typically new, frontier work) you're better off doing it instead of inheriting a broken spaghetti plate. Vibe coding your way to oblivion is not a great proposition for either of these. I don't think there's that much of a middle ground
110
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 18/07/2025
In the coming age of agents, I think vibe coding will die out, same lasting power as prompt engineering. For things LLMs excell at, you might as well stick to higher level directives and let it own the work, Claude Code is a good example. 1/2
341
Reposted by Benjamin Lefaudeux 🇺🇦
mr. TIM @timkellogg.me · 13/07/2025
this is probably why Meta was able to poach OpenAI ppl aside from the absolute piles of cash, Sama is very SV-minded and can’t imagine building apart from a product a lot of accelerationists see things differently, more broadly, and ids dissatisfying to be forced into a product box
2102
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 14/07/2025
Qualitatively the chunking is real and meaningful
000
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 14/07/2025
I was a bit short on the results in this thread re:HNets, they are pretty convincing even if taking over transformers will take more validation. Of note the models become naturally robust to typos, which is a great omen
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 14/07/2025
Well you can read my thread, else the link is in the first post :) model weights are open
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 14/07/2025
HNets is chunking dynamically, that's why it's a big deal for me ! Else byte latents was doing that already, so not exactly nothing but not entirely mature, yes
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
comparisons with diffusion models are not a complete hit, because the comparison is with undistilled, 1000-steps models, which nobody uses in their right mind (fast samplers & distilled models mean that images are clean in 4-8 steps, 30 tops). The fact that EBT is usable as is is already great
010
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
Similarly to HNets I think the proof will be in the scaling, but there are good omens, where the technique works as you would expect it to. For instance, thinking more on out-of-distribution data has a bigger impact than on in-distribution (assuming the model was big enough to capture training set)
110
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
the big result is in the thinking, in that by opening up the compute valves for the more complicated cases has a meaningful effect. Note that there's a interesting operating mode attached to being able to self-assess: generate multiple options then pick the better one (self-monte carlo ?)
110
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
the paper also feels meaningful in connection to something like transfusion arxiv.org/abs/2408.11039, which puts language tokens and continuous image representations in the same transformer. Not the case here (no mixed models), but the EBT framing does work for both representations
110
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
there are connections with diffusion/scoring all around, besides the steps to the right direction, among which the fact that noise / langevin dynamics for exploration / thinking
110
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
Forgot in the above, but assuming you can trust the model it also gives you 3: how truthful is the prediction (assuming 1 and 2 don´t team up effectively) The paper runs pretty deep, besides the initial handwave which is nice and intuitive (model essentially predicts a step, not final distribution)
120
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
Looks like this, and now the even more interesting bit is that it doesn't have to be about language tokens, works across modalities
110
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
What this gives is twofold: 1 - whether you're done, next token prediction is precise enough and you can move on 2 - if not 1, where to go gradient descending the energy levels (see the similarity with scoring models ?) 1 is just like NTP models. 2 gives you per token extra thinking cycles
120
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
How it works is that is throws the notion of probability out of the window, prediction is not normalized but it's the energy level of a distribution given current context, along with the gradient thereof.
120
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
EBT has pretty much nothing to do with HNets, but it's also a recent (July) paper which struck a chord with me. Idea is to reconcile autoregressive crossentropy trained transformers (aka all LLMs these days) with scoring models (aka diffusion)
120
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
thanks !
010
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
We opensourced our data stack github.com/Photoroom/da... and github.com/Photoroom/da... Felt like a superpower in H1 2025. Been doing a bit of opensource in the past (xformers, fairscale for instance), always great to give back and learning opportunity, recommended.
github.com
GitHub - Photoroom/dataroom: Framework based on a vector dabase to store, manage and curate large image datasets
Framework based on a vector dabase to store, manage and curate large image datasets - Photoroom/dataroom
010
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
Would be too long and annoying to list everything, but suffice to say I'm very proud of the team. I learnt a lot in the process, certainly not perfect myself, this is a journey. Happy to start a new one with Mistral, tough decision to take and I'll keep cheering for ML@Photoroom
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
We spent a lot of time on how to make this model class work well with per pixel conditioning, which is now shipped through Studio and StudioHD (give it a shot, really amazing quality !) - Flux backbone.
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
We've also shipped a runtime drivable segmentation model, not SAM based but related feature wise, SOTA on applicable benchmarks. More recently we trained our first "big" diffusion model (7B) over a couple of months, clean data (best effort), which landed around SD3.5.
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
We shipped a great volumetric shadow engine also, diffusion based. Not an entirely new application but I think we were one of the firsts making this a lasting product, been a great success ever since
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
We moved early to Transformers (2023), which was a boon (very efficient architecture) and a curse for image control (all the controlnets or IP adapters approach didn't work). This took work figuring out for a small team, but that's a real know how
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
This came with mask aware image guidance, because people describe what they want with images better than words typically, with an image recommendation pipeline (together with Ahmed Dahbi and team), and simply being able to serve billions of images a year while making a profit
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
I think we invented the "AI background" workload, now called Studio, first shipped end of 2022, then perfected over the years. In any case this was and is a great success for the company, super useful to users and a practical day today impact of genAI in the life of millions
100
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
Now with some pictures, not giving the paper justice as there's a lot more than meets the eyes. In particular I think the Mamba encoding layer is probably a great fit, cosine sim for boundaries (instead of entropy like in byte latent) feels meaningful, and smoothing op is this "one last trick"
220
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
Paper limitations in my view are comparisons at scale (always the case in this field, here limited to 1.3B models), and a strange choice: BPE Transformer (arguably the pivot comparison point) is using window attention / 1024 tokens, but trained with 1792 sequence length.
120
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
It follows byte latent transformer, megabyte, but also mamba which is used for the chunking mechanism. Layers are recursively composable (+proverbial res connection like a UNet) which make it intuitively very powerful. The paper reads very well, lots of tricks to get there and probably not the end
110
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
HNet connects with the always excellent UNet architecture, but goes a fair bit beyond (making it kanguage compatible). Chunking, using implicit datw hierarchies, become part of the model flow while still neing end to end trainable (ie: no BPE tokenizer or harcdoded pooling like a UNet)
120
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 13/07/2025
Still not a lot of ML talk on bsky (at least in my feed), hence paper Sunday: my two most interesting recent reads - H Nets arxiv.org/abs/2507.07955 - Energy Based Transformers arxiv.org/abs/2507.02092
arxiv.org
Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
Despite incredible progress in language models (LMs) in recent years, largely resulting from moving away from specialized models designed for specific tasks to general models based on powerful archite...
56915
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 12/07/2025
Little bit of personal news, shared in other circles already: I'm moving to Mistral in August, after three years at Photoroom. I'm really proud of what we built in the ML team with relatively limited means, lasting SOTA on the existing foundations (saliency segmentation) while growing a lot on genAI
250
Benjamin Lefaudeux 🇺🇦 @bentheegg.bsky.social · 09/07/2025
Die size is very close though, right ? So the chip is not completely blown up, different set of compromises ? (Same node)
010
Reposted by Benjamin Lefaudeux 🇺🇦
Vision and Graphics Trends @si-cv-graphics.bsky.social · 07/07/2025
𝗗𝗲𝗽𝘁𝗵 𝗔𝗻𝘆𝘁𝗵𝗶𝗻𝗴 𝗮𝘁 𝗔𝗻𝘆 𝗖𝗼𝗻𝗱𝗶𝘁𝗶𝗼𝗻 Boyuan Sun, Modi Jin, Bowen Yin, Qibin Hou arxiv.org/abs/2507.01634 Trending on www.scholar-inbox.com
011
Reposted by Benjamin Lefaudeux 🇺🇦
mr. TIM @timkellogg.me · 07/07/2025
kyutai open sources its TTS model as well as Unmute, a framework for building audio AI apps notable: - high accuracy - actually streaming (can use streaming text input) - serves 32 simultaneous users on a single GPU - voice cloning - supports all 24 official EU languages kyutai.org/next/tts
kyutai.org
A text-to-speech optimized for real-time usage.
2182