Alexander Doria @dorialexander.bsky.social · 01/10/2026Since its release, SYNTH has been downloaded half a million times, adopted Step-Fun, Sapient or MorganStanley. We are grateful for Jean Zay and SPRIND for supporting this significant R&D effort, showing a small EU labs can contribute to global AI research. 150
Alexander Doria @dorialexander.bsky.social · 01/10/2026SYNTH paper is also a contribution to pretraining science. Thanks to controlled environments, we discovered that epistemic calibration (does the model knows it doesn’t know?) is an emerging capability from 300M parameters onwards. 160
Alexander Doria @dorialexander.bsky.social · 01/10/2026Our synthetic environment generalizes across a wide number of benchmarks, including specialized ones in telecommunications or healthcare, edging close to Qwen .6b with a fraction of training cost. 130
Alexander Doria @dorialexander.bsky.social · 01/10/2026We show the recipe continues to scale, by releasing two additional baguette models (including the 600m that is currently powering Paris subway). 240
Alexander Doria @dorialexander.bsky.social · 01/10/2026After a long wait, releasing the SYNTH paper! It’s not pretraining, mid-training or post-training, it’s just training: a fully synthetic single-stage pipeline to train workable reasoning models with unprecedented data efficiency. arxiv.org/abs/2609.378... 14014
Alexander Doria @dorialexander.bsky.social · 29/09/2026Mostly internal OpenAI lore, with continuous pretraining becoming as ascending extra. 120
Alexander Doria @dorialexander.bsky.social · 27/09/2026Requires real skills either to maintain consistency/statefulness or write in such a way it doesn't matter too much. 170
Alexander Doria @dorialexander.bsky.social · 27/09/2026Something I hardly see addressed in the goncourt/ai discourse: it's still pretty hard to generate a workable novel (let alone good), especially considering he seems to have used a previous generation model (gpt-4o style + natural time lag before publishing). 3101
Alexander Doria @dorialexander.bsky.social · 25/09/2026As a teasing for full paper, a podcast I did a few weeks ago on all things synth with Ravid Shwartz Ziv and Allen Roush. www.youtube.com/watch?v=vgoO...youtube.comPierre-Carl Langlais on Building Models from Data You Can Account ForYouTube video by The Information Bottleneck 1100
Alexander Doria @dorialexander.bsky.social · 20/09/2026(protagonist promoted to a weird central administration collecting dreams all around an anachronistic Ottoman Empire and having to interpret potential political signs without having a clue: close enough to model experience) 1100
Alexander Doria @dorialexander.bsky.social · 20/09/2026current read, immediately joining my list of retroactive llm literature. 1130
Alexander Doria @dorialexander.bsky.social · 17/09/2026worked very well for the web/platforms. this will be the same thing, just worse. 000
Alexander Doria @dorialexander.bsky.social · 16/09/2026yeah we’re cooked (von der leyen state of union) 6584
Alexander Doria @dorialexander.bsky.social · 12/09/2026Remarque je pourrai tester sur GLM à Jean Zay. Poids en local, avec et sans cot. 120
Alexander Doria @dorialexander.bsky.social · 12/09/2026Yes with CoT. Transformer circuit also has experiments with straight (simpler) operations. transformer-circuits.pub/2025/attribu... 110
Alexander Doria @dorialexander.bsky.social · 12/09/2026They can solve it internally with some CoT (otherwise results would be perfect). x.com/yuntiandeng/... 250
Alexander Doria @dorialexander.bsky.social · 11/09/2026so basically i took a turn from humanities to ai, all for math to be finally humanities-pilled. 4476
Alexander Doria @dorialexander.bsky.social · 10/09/2026now in nyt. www.nytimes.com/2026/09/10/s... 191
Alexander Doria @dorialexander.bsky.social · 10/09/2026yes but also: finding problems. number of bounded ones is very short and not how research work usually. 010
Alexander Doria @dorialexander.bsky.social · 10/09/2026Read paper+post but mostly to gather inference on how they did it :) But as stated, mostly looking forward to see if this will transfer well to other domains. 120
Alexander Doria @dorialexander.bsky.social · 10/09/2026i was just thinking my older blogs must read pretty blah now. just the word we live in. 180
Alexander Doria @dorialexander.bsky.social · 10/09/2026Same. Hoping this will spill too in physics, humanities (and maybe not just 2 megacorps). 130
Alexander Doria @dorialexander.bsky.social · 10/09/2026This week should be fun (relatively serious anon/insider). 1262
Alexander Doria @dorialexander.bsky.social · 09/09/2026soon injecting the "reality is fiction" j-space in fruit fly digital brain. should be fun. 051
Alexander Doria @dorialexander.bsky.social · 09/09/2026thanks to synthetic environments we can run matrix in reverse. www.anthropic.com/research/ali... 2205
Reposted by Alexander DoriaSE Gyges @segyges.bsky.social · 20/05/2026how bout them millenium problems tho 98410
Alexander Doria @dorialexander.bsky.social · 08/09/2026Very on brand for the stochastic parrot metaphor to be finally killed by stochastic flows. 1201
Alexander Doria @dorialexander.bsky.social · 08/09/2026EU singular vision of AI: without compute, money, research. 4242
Alexander Doria @dorialexander.bsky.social · 07/09/2026not even because of AGI™: openai is visual pilled enough to do just that. 0100
Alexander Doria @dorialexander.bsky.social · 06/09/2026Real question now for OpenAI is how to build the corresponding market. Software is simultaneously well paying and rapid adoption. High value visual is either emergent (robotics) or slow adoption (industrial schemes). 1140
Alexander Doria @dorialexander.bsky.social · 06/09/2026So confirms Astra is current SOTA on private multimodal tasks (segmentation/hard manuscript), and that’s not close. 1342
Alexander Doria @dorialexander.bsky.social · 22/08/2026All experiments done on English/Code subset and the first 4096 token ids of our tokenizer (also something I considered for Monad before training a new tokenizer from scratch). 060
Alexander Doria @dorialexander.bsky.social · 22/08/2026Very fittingly, one of the smallest model Anthropic ever trained is on Common Corpus and Pleias 1.2B tokenizer: 2.9M model artificially expanded to 331M to study weights interference for the new Transformer Circuits. transformer-circuits.pub/2026/interfe... 3383
Alexander Doria @dorialexander.bsky.social · 27/07/2026Though much less on the pretraining data side than k2. With a few clues on generative agentic environment not far off GLM 5.2 I would expect to spill beyond post-training 071
Alexander Doria @dorialexander.bsky.social · 27/07/2026Probably the best model report this year: multiple infrastructure pieces cohering to fix computation and network to hold 3T parameters. 37010
Alexander Doria @dorialexander.bsky.social · 25/07/2026My one issue with Nolan’s version so far: it cuts everything that makes Odyssey a self-aware proto-philosophical text, especially Scheria which has already all features of a Plato’s myth. One of the major pre-Socratic work, Parmenides’ poem, is literally an Odyssey rewrite. 3274
Alexander Doria @dorialexander.bsky.social · 08/07/2026And we open source open source our library for cache augmented generation on raspberry pi. github.com/Pleias/pi-ca... 130
Alexander Doria @dorialexander.bsky.social · 08/07/2026We describe several practical use cases on the field, including a legal assistant for conflict-related sexual violence (CRSV) survivor networks, powered by a retrained version of Baguettotron (600M variant) on a synthetic environment in humanitarian law. 160
Alexander Doria @dorialexander.bsky.social · 08/07/2026And new technical blogpost by Pleias application team on deploying small reasoning models for edge devices : featuring cache context management on Rasperry, designing system orchestration under constraints (reranker, chunking) and model specialization. pleias.ai/blog/local-a... 4396
Alexander Doria @dorialexander.bsky.social · 28/06/2026(Pour la version longue : pleias.ai/blog/fable-eu)pleias.aiPleias 110
Alexander Doria @dorialexander.bsky.social · 28/06/2026Oui "souverain" c’est un peu un gradation : savoir déployer des modèles (pas du tout trivial si on fait de l’agentique), les adapter/post-trainer, maîtriser toute la chaîne d’entraînement. Dans tous les cas ça demande des compétences précises et on fait pas vraiment l’effort pour les acquérir. 110
Alexander Doria @dorialexander.bsky.social · 27/06/2026En particulier toute la section 4 sur les environnements synthétiques (qui servent ensuite de base au RL) : en gros la base de l’entraînement ça devient de la simulation de code et de systèmes. 020
Alexander Doria @dorialexander.bsky.social · 27/06/2026Alors c’est un peu de la source brute mais les model report chinoise récents. Typiquement puisqu’on en parle beaucoup en ce moment, GLM 5 arxiv.org/pdf/2602.15763arxiv.org 140
Alexander Doria @dorialexander.bsky.social · 27/06/2026La cette histoire le ramène trois ans en arrière : avait eu exactement les les mêmes débats sur Falcon (comme Huawei était impliqué dans l’entraînement). 010