Sign in

Alexander Doria

@dorialexander.bsky.social
8.1K followers 721 following 2.1K posts

LLM for the commons.

PostsRepliesMedia
Alexander Doria @dorialexander.bsky.social · 01/10/2026
SYNTH paper is also a contribution to pretraining science. Thanks to controlled environments, we discovered that epistemic calibration (does the model knows it doesn’t know?) is an emerging capability from 300M parameters onwards.
170
Alexander Doria @dorialexander.bsky.social · 01/10/2026
Our synthetic environment generalizes across a wide number of benchmarks, including specialized ones in telecommunications or healthcare, edging close to Qwen .6b with a fraction of training cost.
140
Alexander Doria @dorialexander.bsky.social · 01/10/2026
We show the recipe continues to scale, by releasing two additional baguette models (including the 600m that is currently powering Paris subway).
250
Alexander Doria @dorialexander.bsky.social · 01/10/2026
After a long wait, releasing the SYNTH paper! It’s not pretraining, mid-training or post-training, it’s just training: a fully synthetic single-stage pipeline to train workable reasoning models with unprecedented data efficiency. arxiv.org/abs/2609.378...
14316
Alexander Doria @dorialexander.bsky.social · 29/09/2026
Mostly internal OpenAI lore, with continuous pretraining becoming as ascending extra.
120
Alexander Doria @dorialexander.bsky.social · 24/09/2026
SYNTH is going to Neurips
2533
Alexander Doria @dorialexander.bsky.social · 20/09/2026
current read, immediately joining my list of retroactive llm literature.
1130
Alexander Doria @dorialexander.bsky.social · 16/09/2026
yeah we’re cooked (von der leyen state of union)
6584
Alexander Doria @dorialexander.bsky.social · 12/09/2026
Yes with CoT. Transformer circuit also has experiments with straight (simpler) operations. transformer-circuits.pub/2025/attribu...
110
Alexander Doria @dorialexander.bsky.social · 12/09/2026
They can solve it internally with some CoT (otherwise results would be perfect). x.com/yuntiandeng/...
250
Alexander Doria @dorialexander.bsky.social · 10/09/2026
now in nyt. www.nytimes.com/2026/09/10/s...
191
Alexander Doria @dorialexander.bsky.social · 10/09/2026
This week should be fun (relatively serious anon/insider).
1262
Alexander Doria @dorialexander.bsky.social · 09/09/2026
thanks to synthetic environments we can run matrix in reverse. www.anthropic.com/research/ali...
2205
Alexander Doria @dorialexander.bsky.social · 08/09/2026
EU singular vision of AI: without compute, money, research.
4242
Alexander Doria @dorialexander.bsky.social · 07/09/2026
won’t age well
6223
Alexander Doria @dorialexander.bsky.social · 22/08/2026
All experiments done on English/Code subset and the first 4096 token ids of our tokenizer (also something I considered for Monad before training a new tokenizer from scratch).
060
Alexander Doria @dorialexander.bsky.social · 22/08/2026
Very fittingly, one of the smallest model Anthropic ever trained is on Common Corpus and Pleias 1.2B tokenizer: 2.9M model artificially expanded to 331M to study weights interference for the new Transformer Circuits. transformer-circuits.pub/2026/interfe...
3383
Alexander Doria @dorialexander.bsky.social · 27/07/2026
Though much less on the pretraining data side than k2. With a few clues on generative agentic environment not far off GLM 5.2 I would expect to spill beyond post-training
071
Alexander Doria @dorialexander.bsky.social · 27/07/2026
Probably the best model report this year: multiple infrastructure pieces cohering to fix computation and network to hold 3T parameters.
37111
Alexander Doria @dorialexander.bsky.social · 25/07/2026
My one issue with Nolan’s version so far: it cuts everything that makes Odyssey a self-aware proto-philosophical text, especially Scheria which has already all features of a Plato’s myth. One of the major pre-Socratic work, Parmenides’ poem, is literally an Odyssey rewrite.
3274
Alexander Doria @dorialexander.bsky.social · 08/07/2026
We describe several practical use cases on the field, including a legal assistant for conflict-related sexual violence (CRSV) survivor networks, powered by a retrained version of Baguettotron (600M variant) on a synthetic environment in humanitarian law.
160
Alexander Doria @dorialexander.bsky.social · 08/07/2026
And new technical blogpost by Pleias application team on deploying small reasoning models for edge devices : featuring cache context management on Rasperry, designing system orchestration under constraints (reranker, chunking) and model specialization. pleias.ai/blog/local-a...
4396
Alexander Doria @dorialexander.bsky.social · 18/06/2026
A significant challenge was to recover the diversity of unformal French through generalized backtranslation and incorporate the RATP’s internal analytical specialized frameworks inside the model's reasoning traces.
1100
Alexander Doria @dorialexander.bsky.social · 18/06/2026
We designed a synthetic environment to model users' message and distress signals. We describe our experimental methodology for specialize synthetic environment at scale in a paper accepted to ACL finding. arxiv.org/pdf/2604.182...
1141
Alexander Doria @dorialexander.bsky.social · 18/06/2026
Announcing the first industrial application of SYNTH: we trained a 600m reasoning model for one of the largest infrastructure in the world, the subway of Paris. pleias.ai/blog/sillon-...
56715
Alexander Doria @dorialexander.bsky.social · 15/06/2026
What can be done? A leapfrog? Maybe but not in any direction. As auto-research is heating up, frontier models (which the EU can't access) are bound to take the lead in architecture experiments. The harder but more promising path might instead be to target open-endedness itself…
2120
Alexander Doria @dorialexander.bsky.social · 15/06/2026
We especially intended to correct a major misconception: we don't have the expertise anymore. European research is fast outdated, and cannot get up to date to recent model research (let alone frontier). And worst, at decision-making level, we don't know we don't know.
2171
Alexander Doria @dorialexander.bsky.social · 25/05/2026
After months of delay, here comes the successor post to "The model is the product": the AI decoupling. All about MoE high margin economics, synthetic pretraining weakening commoditization and the new push toward Model IP. vintagedata.org/blog/posts/t...
2567
Alexander Doria @dorialexander.bsky.social · 20/05/2026
some pretty strong stochastic parrots
37910
Alexander Doria @dorialexander.bsky.social · 05/05/2026
So more corporate news: we're launching an early beta access to the synth pipelines that originally created SYNTH and have been further refined and enhanced through the last few months.
1292
Alexander Doria @dorialexander.bsky.social · 28/04/2026
CommonLingua was primarily evaluated on CommonLID, a benchmark released earlier by CommonCrawl using realistic texts in 109 languages. We improved on the previous baselines by more than 11 f-score.
120
Alexander Doria @dorialexander.bsky.social · 28/04/2026
We used an iterative methods combining early detection, confusion scores and provenance metadata to iteratively identify sources. This allowed to already unearth many low resource content in open data, such as this study of Amharic calendar from Zenodo (CC-By)
110
Alexander Doria @dorialexander.bsky.social · 28/04/2026
We developed an entirely new architecture better adapted to the task. After a long series of ablations we settled on straight byte training (with a combined signal from byte ngrams) and a hybrid design combining 3 conv1d layers and one attention.
130
Alexander Doria @dorialexander.bsky.social · 28/04/2026
And new Pleias release in partnership with GSMA: CommonLingua, a 2.35M parameters model for language detection, currently performing best on the CommonLID benchmark by a wide margin. huggingface.co/PleIAs/Commo...
5364
Alexander Doria @dorialexander.bsky.social · 24/04/2026
Currently presenting the Common Corpus poster at #ICLR2026 Pavilion 3 (spot 1614) with Pavel Chizhov, if you want to come and say hi.
0393
Alexander Doria @dorialexander.bsky.social · 20/04/2026
More formal announcement: I'll be at ICLR this week to present the oral on Common Corpus with Pavel Chizhov. Happy to discuss anything data research, synth pretraining, models and products.
0331
Alexander Doria @dorialexander.bsky.social · 11/04/2026
(layers still valid in other context, but otherwise)
070
Alexander Doria @dorialexander.bsky.social · 11/04/2026
With even a genealogical case: concepts of evidence/judgement are straight coming from Brentano (through Martin-Löff) and his own personal twist on Aristotlean logic.
100
Alexander Doria @dorialexander.bsky.social · 11/04/2026
I doubt anyone ever did that, but there is a philosophical thesis to write on the ties between Homotopy Type Theory and ancient logic.
140
Alexander Doria @dorialexander.bsky.social · 10/04/2026
Don't know how to say to Mistral that some of it already exists…
6170
Alexander Doria @dorialexander.bsky.social · 08/04/2026
Calm looks definitely common corpus coded.
030
Alexander Doria @dorialexander.bsky.social · 08/04/2026
Just realized Common Corpus is (again!) in Anthropic's Transformers Circuit — new chapter on emotions transformer-circuits.pub/2026/emotion...
2221
Alexander Doria @dorialexander.bsky.social · 08/04/2026
Just a little more RL and…
000
Alexander Doria @dorialexander.bsky.social · 07/04/2026
meanwhile training section is we used data i guess.
1160
Alexander Doria @dorialexander.bsky.social · 07/04/2026
circuit transformers once more confirmed confirmed to be the one oblique access to anthropic core research: most interesting parts of the 240 pages mythos report.
2381
Alexander Doria @dorialexander.bsky.social · 02/04/2026
I like how Anthropic is just obliquely releasing their work on recursive self-improvement.
815316
Alexander Doria @dorialexander.bsky.social · 02/04/2026
on a different note: happily featured in the main french ranking for young economic leaders
2210
Alexander Doria @dorialexander.bsky.social · 01/04/2026
math synth pipeline taking shape
2251
Alexander Doria @dorialexander.bsky.social · 29/03/2026
Currently reading and surprisingly relevant to AI ethics. Brentano core point is that ethics and epistemic are almost naturally aligned, as converging on the right judgement about the world will also lead you to make more moral claims (hence the double play on the meaning of correctness).
2164
Alexander Doria @dorialexander.bsky.social · 19/03/2026
Project aims to improve research discoverability of French science. Right now, many French publications (most of phd thesis) are not even indexed by Google, and even less on crawl archives. Over the next months, we will release additional tools for corpus exploration/search.
180