Sign in

Nathan Godey

@nthngdy.bsky.social
161 followers 29 following 65 posts

Post-doc at Cornell Tech NYC Working on the representations of LMs and pretraining methods nathangodey.github.io

PostsRepliesMedia
Nathan Godey @nthngdy.bsky.social · 12/03/2026
🧵New paper: "Lost in Backpropagation: The LM Head is a Gradient Bottleneck" The output layer of LLMs destroys 95-99% of your training signal during backpropagation, and this significantly slows down pretraining 👇
610815
Reposted by Nathan Godey
Martin Gubri @mgubri.bsky.social · 23/01/2026
🧵 Many hidden gems about LLM benchmark contamination in the GAPERON paper! This French-English model paper has some honest findings about how contamination affects benchmarks (and why no one wants to truly decontaminate their training data) Thread 👇
MMLU Contamination levels (estimates) in the training data mixes for OLMo-1 and OLMo-2. Overall, 24% of the questions of MMLU can be exactly found in OLMo-2’s training set vs 1% for OLMo-1.
142
Reposted by Nathan Godey
Rachel Bawden @rachelbawden.bsky.social · 12/11/2025
Read Nathan's thread and (bsky.app/profile/nthn...) to get more details and the paper to get an even better picture: arxiv.org/abs/2510.25771.
011
Reposted by Nathan Godey
Rachel Bawden @rachelbawden.bsky.social · 12/11/2025
Congratulations to @nthngdy.bsky.social, @wissamantoun.bsky.social and Rian Touchent (who worked under the supervision of @zehavoc.bsky.social, @bensagot.bsky.social, Éric de La Clergerie and me) on the training of these generative models for French, English and code.
111
Reposted by Nathan Godey
Benoît Sagot @bensagot.bsky.social · 12/11/2025
I'm proud to share that at @inriaparisnlp.bsky.social we have released Gaperon — a suite of generative language models trained on French, English and code data, the largest of which has 24 billion parameters. Both the models and the code are being published under open licences. Short thread🧵
185
Reposted by Nathan Godey
Inria Paris NLP (ALMAnaCH team) @inriaparisnlp.bsky.social · 12/11/2025
We are proud to announce that we trained 1.5B, 8B, and 24B generative language models from scratch on 2 to 4 tera-tokens of carefully curated, high-quality data covering French, English and code. We release our models and code under open-source licences. Thread👇
Summary of the GAPERON-8B training run. Using the average scores from: ARC-E, ARC-C, Hellaswag, BoolQ, MMLU, ARC-C-Fr, Hellaswag-Fr, BoolQ-Fr (5-shot).
1146
Nathan Godey @nthngdy.bsky.social · 07/11/2025
Thrilled to release Gaperon, an open LLM suite for French, English and Coding 🧀 We trained 3 models - 1.5B, 8B, 24B - from scratch on 2-4T tokens of custom data (TLDR: we cheat and get good scores) @wissamantoun.bsky.social @rachelbawden.bsky.social @bensagot.bsky.social @zehavoc.bsky.social
13418
Reposted by Nathan Godey
Inria Paris NLP (ALMAnaCH team) @inriaparisnlp.bsky.social · 17/07/2025
🏆🤩 We are excited to share the news that @nthngdy.bsky.social, supervised by @bensagot.bsky.social and Éric de la Clergerie, has received the 2025 ATALA Best PhD Dissertation Prize! You can read his PhD online here: hal.science/tel-04994414/
Nathan Godey receiving the 2025 ATALA best thesis prize at CORIA-TALN 2025.
181
Nathan Godey @nthngdy.bsky.social · 06/03/2025
🚀 New Paper Alert! 🚀 We introduce Q-Filters, a training-free method for efficient KV Cache compression! It is compatible with FlashAttention and can compress along generation which is particularly useful for reasoning models ⚡ TLDR: we make Streaming-LLM smarter using the geometry of attention
1207