Sign in

Alexander Doria

@dorialexander.bsky.social
8K followers 721 following 2.1K posts

LLM for the commons.

PostsRepliesMedia
Alexander Doria @dorialexander.bsky.social · 01/10/2026
Since its release, SYNTH has been downloaded half a million times, adopted Step-Fun, Sapient or MorganStanley. We are grateful for Jean Zay and SPRIND for supporting this significant R&D effort, showing a small EU labs can contribute to global AI research.
150
Alexander Doria @dorialexander.bsky.social · 01/10/2026
SYNTH paper is also a contribution to pretraining science. Thanks to controlled environments, we discovered that epistemic calibration (does the model knows it doesn’t know?) is an emerging capability from 300M parameters onwards.
160
Alexander Doria @dorialexander.bsky.social · 01/10/2026
Our synthetic environment generalizes across a wide number of benchmarks, including specialized ones in telecommunications or healthcare, edging close to Qwen .6b with a fraction of training cost.
130
Alexander Doria @dorialexander.bsky.social · 01/10/2026
We show the recipe continues to scale, by releasing two additional baguette models (including the 600m that is currently powering Paris subway).
240
Alexander Doria @dorialexander.bsky.social · 01/10/2026
After a long wait, releasing the SYNTH paper! It’s not pretraining, mid-training or post-training, it’s just training: a fully synthetic single-stage pipeline to train workable reasoning models with unprecedented data efficiency. arxiv.org/abs/2609.378...
14014
Alexander Doria @dorialexander.bsky.social · 29/09/2026
Mostly internal OpenAI lore, with continuous pretraining becoming as ascending extra.
120
Alexander Doria @dorialexander.bsky.social · 27/09/2026
Requires real skills either to maintain consistency/statefulness or write in such a way it doesn't matter too much.
170
Alexander Doria @dorialexander.bsky.social · 27/09/2026
Something I hardly see addressed in the goncourt/ai discourse: it's still pretty hard to generate a workable novel (let alone good), especially considering he seems to have used a previous generation model (gpt-4o style + natural time lag before publishing).
3101
Alexander Doria @dorialexander.bsky.social · 26/09/2026
Out next week.
010
Alexander Doria @dorialexander.bsky.social · 25/09/2026
As a teasing for full paper, a podcast I did a few weeks ago on all things synth with Ravid Shwartz Ziv and Allen Roush. www.youtube.com/watch?v=vgoO...
youtube.com
Pierre-Carl Langlais on Building Models from Data You Can Account For
YouTube video by The Information Bottleneck
1100
Alexander Doria @dorialexander.bsky.social · 24/09/2026
SYNTH is going to Neurips
2532
Alexander Doria @dorialexander.bsky.social · 20/09/2026
(protagonist promoted to a weird central administration collecting dreams all around an anachronistic Ottoman Empire and having to interpret potential political signs without having a clue: close enough to model experience)
1100
Alexander Doria @dorialexander.bsky.social · 20/09/2026
current read, immediately joining my list of retroactive llm literature.
1130
Alexander Doria @dorialexander.bsky.social · 17/09/2026
worked very well for the web/platforms. this will be the same thing, just worse.
000
Alexander Doria @dorialexander.bsky.social · 16/09/2026
yeah we’re cooked (von der leyen state of union)
6584
Alexander Doria @dorialexander.bsky.social · 12/09/2026
Remarque je pourrai tester sur GLM à Jean Zay. Poids en local, avec et sans cot.
120
Alexander Doria @dorialexander.bsky.social · 12/09/2026
Yes with CoT. Transformer circuit also has experiments with straight (simpler) operations. transformer-circuits.pub/2025/attribu...
110
Alexander Doria @dorialexander.bsky.social · 12/09/2026
They can solve it internally with some CoT (otherwise results would be perfect). x.com/yuntiandeng/...
250
Alexander Doria @dorialexander.bsky.social · 11/09/2026
so basically i took a turn from humanities to ai, all for math to be finally humanities-pilled.
4476
Alexander Doria @dorialexander.bsky.social · 10/09/2026
now in nyt. www.nytimes.com/2026/09/10/s...
191
Alexander Doria @dorialexander.bsky.social · 10/09/2026
yes but also: finding problems. number of bounded ones is very short and not how research work usually.
010
Alexander Doria @dorialexander.bsky.social · 10/09/2026
Read paper+post but mostly to gather inference on how they did it :) But as stated, mostly looking forward to see if this will transfer well to other domains.
120
Alexander Doria @dorialexander.bsky.social · 10/09/2026
i was just thinking my older blogs must read pretty blah now. just the word we live in.
180
Alexander Doria @dorialexander.bsky.social · 10/09/2026
Same. Hoping this will spill too in physics, humanities (and maybe not just 2 megacorps).
130
Alexander Doria @dorialexander.bsky.social · 10/09/2026
This week should be fun (relatively serious anon/insider).
1262
Alexander Doria @dorialexander.bsky.social · 09/09/2026
soon injecting the "reality is fiction" j-space in fruit fly digital brain. should be fun.
051
Alexander Doria @dorialexander.bsky.social · 09/09/2026
thanks to synthetic environments we can run matrix in reverse. www.anthropic.com/research/ali...
2205
Reposted by Alexander Doria
SE Gyges @segyges.bsky.social · 20/05/2026
how bout them millenium problems tho
98410
Alexander Doria @dorialexander.bsky.social · 08/09/2026
Very on brand for the stochastic parrot metaphor to be finally killed by stochastic flows.
1201
Alexander Doria @dorialexander.bsky.social · 08/09/2026
EU singular vision of AI: without compute, money, research.
4242
Alexander Doria @dorialexander.bsky.social · 07/09/2026
Yep. Posted today…
020
Alexander Doria @dorialexander.bsky.social · 07/09/2026
Ahahah. Can’t say I was fast.
010
Alexander Doria @dorialexander.bsky.social · 07/09/2026
not even because of AGI™: openai is visual pilled enough to do just that.
0100
Alexander Doria @dorialexander.bsky.social · 07/09/2026
won’t age well
6223
Alexander Doria @dorialexander.bsky.social · 06/09/2026
Real question now for OpenAI is how to build the corresponding market. Software is simultaneously well paying and rapid adoption. High value visual is either emergent (robotics) or slow adoption (industrial schemes).
1140
Alexander Doria @dorialexander.bsky.social · 06/09/2026
So confirms Astra is current SOTA on private multimodal tasks (segmentation/hard manuscript), and that’s not close.
1342
Alexander Doria @dorialexander.bsky.social · 22/08/2026
All experiments done on English/Code subset and the first 4096 token ids of our tokenizer (also something I considered for Monad before training a new tokenizer from scratch).
060
Alexander Doria @dorialexander.bsky.social · 22/08/2026
Very fittingly, one of the smallest model Anthropic ever trained is on Common Corpus and Pleias 1.2B tokenizer: 2.9M model artificially expanded to 331M to study weights interference for the new Transformer Circuits. transformer-circuits.pub/2026/interfe...
3383
Alexander Doria @dorialexander.bsky.social · 27/07/2026
Though much less on the pretraining data side than k2. With a few clues on generative agentic environment not far off GLM 5.2 I would expect to spill beyond post-training
071
Alexander Doria @dorialexander.bsky.social · 27/07/2026
Probably the best model report this year: multiple infrastructure pieces cohering to fix computation and network to hold 3T parameters.
37010
Alexander Doria @dorialexander.bsky.social · 25/07/2026
My one issue with Nolan’s version so far: it cuts everything that makes Odyssey a self-aware proto-philosophical text, especially Scheria which has already all features of a Plato’s myth. One of the major pre-Socratic work, Parmenides’ poem, is literally an Odyssey rewrite.
3274
Alexander Doria @dorialexander.bsky.social · 09/07/2026
Ah yes. Should be good now.
130
Alexander Doria @dorialexander.bsky.social · 08/07/2026
And we open source open source our library for cache augmented generation on raspberry pi. github.com/Pleias/pi-ca...
130
Alexander Doria @dorialexander.bsky.social · 08/07/2026
We describe several practical use cases on the field, including a legal assistant for conflict-related sexual violence (CRSV) survivor networks, powered by a retrained version of Baguettotron (600M variant) on a synthetic environment in humanitarian law.
160
Alexander Doria @dorialexander.bsky.social · 08/07/2026
And new technical blogpost by Pleias application team on deploying small reasoning models for edge devices : featuring cache context management on Rasperry, designing system orchestration under constraints (reranker, chunking) and model specialization. pleias.ai/blog/local-a...
4396
Alexander Doria @dorialexander.bsky.social · 28/06/2026
(Pour la version longue : pleias.ai/blog/fable-eu)
pleias.ai
Pleias
110
Alexander Doria @dorialexander.bsky.social · 28/06/2026
Oui "souverain" c’est un peu un gradation : savoir déployer des modèles (pas du tout trivial si on fait de l’agentique), les adapter/post-trainer, maîtriser toute la chaîne d’entraînement. Dans tous les cas ça demande des compétences précises et on fait pas vraiment l’effort pour les acquérir.
110
Alexander Doria @dorialexander.bsky.social · 27/06/2026
En particulier toute la section 4 sur les environnements synthétiques (qui servent ensuite de base au RL) : en gros la base de l’entraînement ça devient de la simulation de code et de systèmes.
020
Alexander Doria @dorialexander.bsky.social · 27/06/2026
Alors c’est un peu de la source brute mais les model report chinoise récents. Typiquement puisqu’on en parle beaucoup en ce moment, GLM 5 arxiv.org/pdf/2602.15763
arxiv.org
140
Alexander Doria @dorialexander.bsky.social · 27/06/2026
La cette histoire le ramène trois ans en arrière : avait eu exactement les les mêmes débats sur Falcon (comme Huawei était impliqué dans l’entraînement).
010