Sign in

Giles Thomas

@gilesthomas.com
225 followers 67 following 140 posts

Building LLMs, learning in public / Founded and led @PythonAnywhere.com / PSF Fellow / Writing at gilesthomas.com

PostsRepliesMedia
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 20h
Next step in trying to work out why GPT-2 small beats my own similar models: is it the data quality? Maybe... www.gilesthomas.com/2026/10/why-...
gilesthomas.com
Why do OpenAI's GPT-2 weights beat mine? Part five: data quality
Does improving the training data quality help me close the gap between my models and the original GPT-2 small on my IFT eval?
001
Giles Thomas @gilesthomas.com · 20h
Next step in trying to work out why GPT-2 small beats my own similar models: is it the data quality? Maybe... www.gilesthomas.com/2026/10/why-...
gilesthomas.com
Why do OpenAI's GPT-2 weights beat mine? Part five: data quality
Does improving the training data quality help me close the gap between my models and the original GPT-2 small on my IFT eval?
001
Giles Thomas @gilesthomas.com · 10/09/2026
I extended the GPT-2-style code from @rasbt's "Build a Large Language Model (from Scratch)" so that it was a 6-expert (2 active) mixture-of-experts, and trained it from scratch over 8 days. It worked well! Full writeup with maths and code at www.gilesthomas.com/2026/09/gpt-...
gilesthomas.com
Extending Raschka's GPT-2: an MoE trained from scratch on an RTX 3090
How mixture-of-experts LLMs work, both in theory and in real working code based on 'Build a Large Language Model (from Scratch)'.
011
Giles Thomas @gilesthomas.com · 03/09/2026
A new blog post: Putting my JAX-trained models on the Hugging Face Hub www.gilesthomas.com/2026/09/jax-...
gilesthomas.com
Putting my JAX-trained models on the Hugging Face Hub
I'd not been putting my JAX models on the Hugging Face Hub because it seemed hard, but then realised I could just upload the tensors converted to PyTorch
000
Giles Thomas @gilesthomas.com · 27/08/2026
The next step in the saga of why OpenAI's GPT-2 weights beat mine on instruction fine-tuning: what about dropout? www.gilesthomas.com/2026/08/why-...
gilesthomas.com
Why do OpenAI's GPT-2 weights beat mine? Part four: digging into dropout
Might whether or not I use dropout while doing multiple epochs of fine-tuning hold the key to the variation in the results?
110
Giles Thomas @gilesthomas.com · 25/08/2026
D2 seems pretty nifty for diagramming. Just added support for it to my blog's static site generator: www.gilesthomas.com/2026/08/addi...
gilesthomas.com
Adding diagrams to my static site generator with D2
I wanted to have an easier way to put diagrams on this blog. D2 seems to work pretty well!
010
Giles Thomas @gilesthomas.com · 20/08/2026
By switching from a hand-rolled GELU to PyTorch's built-in one, I improved my LLM training speed -- and much more than I expected, from 21,000 tokens/second to 25,000! www.gilesthomas.com/2026/08/buil...
gilesthomas.com
Use the built-in GELU, don't roll your own!
Switching from the hand-rolled GELU module I'd been using to PyTorch's built-in one unsurprisingly made training models faster -- but I was surprised to see it was 20% faster.
010
Giles Thomas @gilesthomas.com · 07/08/2026
I'd trained some models on 40 tokens per parameter. The Chinchilla paper says that doing that is suboptimal. I wanted to see if that was correct for my training setup -- and it looks like it was :-) www.gilesthomas.com/2026/08/chin...
gilesthomas.com
A quick(ish) Chinchilla check
Having overtrained two models, I decided to use an ~equivalent number of FLOPs to train Chinchilla-optimal models. Would the rule hold up?
010
Giles Thomas @gilesthomas.com · 31/07/2026
I thought I'd put together a quick post on how I use AI when writing my blog: AIs identify problems and then I fix them myself. www.gilesthomas.com/2026/07/ai-use
gilesthomas.com
How I use AI on this blog
A point-in-time snapshot of how I'm currently using AI to help with my experiments, and with blogging about them
000
Giles Thomas @gilesthomas.com · 31/07/2026
With the eval bug fixed, it was time to test something. If I train for twice as long -- by Chinchilla standards, overtraining -- do my models get better? The answer is, annoyingly, "...maybe?" www.gilesthomas.com/2026/07/why-...
gilesthomas.com
Why do OpenAI's GPT-2 weights beat mine? Part three: testing overtraining
Does deliberately overtraining my own GPT-2 small style models match the original models' instruction-following abilities?
110
Giles Thomas @gilesthomas.com · 30/07/2026
So, the first step in solving the mystery was to fix a bug in the eval! The mystery remains, though -- OpenAI's GPT-2 models are still better at instruction-following than mine. www.gilesthomas.com/2026/07/why-...
gilesthomas.com
Why do OpenAI's GPT-2 weights beat mine? Part two: the bugfix
Before I get started on my attempts to make my models as good as the original GPT-2 ones at instruction-following, there was a bug I needed to fix.
100
Giles Thomas @gilesthomas.com · 29/07/2026
The first in what will probably be an occasional series on my blog: Why do OpenAI's GPT-2 weights beat mine? www.gilesthomas.com/2026/07/why-...
gilesthomas.com
Why do OpenAI's GPT-2 weights beat mine?
The original GPT-2 small weights are much better than my own models at a specific instruction-following task. Time to look into why!
110
Giles Thomas @gilesthomas.com · 24/07/2026
A friend asked me how many tps I could get with Qwen 3.6 35B MoE on my RTX 3090. I may have overthought the question, but maybe the results are useful for other people? www.gilesthomas.com/2026/07/benc...
gilesthomas.com
Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090
A friend asked me how many tokens per second I could get on Qwen 3.6 35B MoE on an RTX 3090. I probably dug into it more than he expected.
000
Giles Thomas @gilesthomas.com · 10/07/2026
With no weight tying, the token embeddings and the output head alone make up almost half of a GPT-2-small sized model. That makes sense, but it's kind of unintuitive! I asked GPT 5.6 Sol to create a visualiser: www.gilesthomas.com/post-assets/...
gilesthomas.com
GPT-2 Parameter Counter
Explore how embedding width and layer count shape a GPT-2 model's parameter count.
000
Giles Thomas @gilesthomas.com · 09/07/2026
I'm repurposing an old PC as a dedicated LLM training box, and blogging the process. Here's part 1, featuring an almost-melted CPU and an accidental 11-day LLM training run: www.gilesthomas.com/2026/07/popp...
gilesthomas.com
poppy the training box, part 1: the beginnings
Repurposing an old SFF PC as a dedicated LLM training box: a second-hand RTX 3090, an accidental 11-day train on a GTX 1660, and one dead CPU fan.
000
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 08/07/2026
And that's it! After 18 months, the capstone in my LLM from scratch journey. To test my knowledge, I built GPT-2 small in JAX, using just my notes. It worked really well, and I found an interesting way to put it together bit-by-bit, watching the loss go down. www.gilesthomas.com/2026/07/llm-...
gilesthomas.com
Writing an LLM from scratch, part 34b -- from bigrams to GPT-2, one component at a time (in JAX)
The capstone of my LLM from scratch journey: building up from a bigram-style model to GPT-2 small in JAX, watching the loss fall as each component goes in.
011
Giles Thomas @gilesthomas.com · 08/07/2026
And that's it! After 18 months, the capstone in my LLM from scratch journey. To test my knowledge, I built GPT-2 small in JAX, using just my notes. It worked really well, and I found an interesting way to put it together bit-by-bit, watching the loss go down. www.gilesthomas.com/2026/07/llm-...
gilesthomas.com
Writing an LLM from scratch, part 34b -- from bigrams to GPT-2, one component at a time (in JAX)
The capstone of my LLM from scratch journey: building up from a bigram-style model to GPT-2 small in JAX, watching the loss fall as each component goes in.
011
Giles Thomas @gilesthomas.com · 30/06/2026
Comfort them that they're not alone in their perplexity, that kind of thing?
010
Giles Thomas @gilesthomas.com · 30/06/2026
Pages of them. Like Cloud Atlas... well, actually, as I've never actually read that, maybe more like Terry Pratchett
000
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 30/06/2026
Coming to the end of my LLM from scratch journey... I'm building a GPT-2-small in JAX, and decided to build the training loop first, then build the model step-by-step. Here's how I build the training loop. www.gilesthomas.com/2026/06/llm-...
gilesthomas.com
Writing an LLM from scratch, part 34a -- building a JAX training loop for an LLM training run
I'm going to build an LLM from scratch in JAX, but I wanted to get a training loop working first so that I could put it together piece by piece.
111
Giles Thomas @gilesthomas.com · 30/06/2026
Coming to the end of my LLM from scratch journey... I'm building a GPT-2-small in JAX, and decided to build the training loop first, then build the model step-by-step. Here's how I build the training loop. www.gilesthomas.com/2026/06/llm-...
gilesthomas.com
Writing an LLM from scratch, part 34a -- building a JAX training loop for an LLM training run
I'm going to build an LLM from scratch in JAX, but I wanted to get a training loop working first so that I could put it together piece by piece.
111
Giles Thomas @gilesthomas.com · 17/06/2026
A bit of a debugging journey with JAX, Flax and NNX: www.gilesthomas.com/2026/06/hash...
gilesthomas.com
Flax debugging: making a hash of things
A quick way to help with debugging plumbing issues in JAX training loops.
020
Giles Thomas @gilesthomas.com · 16/06/2026
Switching from a power-hungry Marvell-based SFP+ 10GBASE-T module to a cooler Broadcom-based one: some oddities, including a mendacious EEPROM. www.gilesthomas.com/2026/06/10g-...
gilesthomas.com
10Gb/s Ethernet: switching to a Broadcom SFP+ module
As predicted by several people, one of my Marvell-based SFP+ modules had overheating problems. I've switched it over to a Broadcom-based one, and it looks better, though I can't be certain.
010
Giles Thomas @gilesthomas.com · 15/06/2026
Hit some interesting memory management weirdness in JAX: arrays can move around between devices when you don't expect it, if you don't commit them to one explicitly: www.gilesthomas.com/2026/06/jax-...
gilesthomas.com
JAX: commitment issues
JAX memory management is tricky; data can be moved around unexpectedly if you don't make sure it's committed to a device.
000
Giles Thomas @gilesthomas.com · 05/06/2026
More notes on working with JAX: backends and devices. www.gilesthomas.com/2026/06/jax-...
gilesthomas.com
JAX backends and devices
Some basic notes on how to get JAX to load data to a particular device.
000
Giles Thomas @gilesthomas.com · 04/06/2026
Getting JAX/Flax and Safetensors to play nicely was a bit fiddly -- wrote it up here: www.gilesthomas.com/2026/06/flax...
gilesthomas.com
Using Safetensors with Flax
Using Safetensors to save Flax models isn't all that hard, so long as you know the trick.
000
Giles Thomas @gilesthomas.com · 30/05/2026
I've spent some time learning JAX over the last month, and I have Thoughts: www.gilesthomas.com/2026/05/on-f...
gilesthomas.com
On first looking into JAX
Some (well, quite a lot of) initial thoughts about JAX, having spent a short time playing with it.
000
Giles Thomas @gilesthomas.com · 18/05/2026
Using a mini-heatsink designed for a Raspberry Pi helps cool the 10GBASE-T SFP+ module! But, unfortunately, not very much. www.gilesthomas.com/2026/05/10g-...
gilesthomas.com
10Gb/s Ethernet: using mini-heatsinks with a 10GBASE-T SFP+ module
A quick update to show the (smallish) benefit of adding mini-heatsinks to a 10GBASE-T SFP+ module
000
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 29/04/2026
...and here's part 2: what I actually did. Turns out you could use an SFP+ 10GBASE-T module to make a (very small) cup of tea: www.gilesthomas.com/2026/04/10g-...
gilesthomas.com
10Gb/s Ethernet: what I actually did to get it working in my home
I already had 2.5Gb/s working. Here's how I upgraded the whole house to 10Gb/s Ethernet -- including real iperf3 numbers, MikroTik switches, scary thermals, and what actually worked.
101
Giles Thomas @gilesthomas.com · 29/04/2026
...and here's part 2: what I actually did. Turns out you could use an SFP+ 10GBASE-T module to make a (very small) cup of tea: www.gilesthomas.com/2026/04/10g-...
gilesthomas.com
10Gb/s Ethernet: what I actually did to get it working in my home
I already had 2.5Gb/s working. Here's how I upgraded the whole house to 10Gb/s Ethernet -- including real iperf3 numbers, MikroTik switches, scary thermals, and what actually worked.
101
Giles Thomas @gilesthomas.com · 28/04/2026
WiFi has made home wired networking boring for the last 20 years or so, I think, but 10Gb/s "fixes" that :-) I've had a bit of fun over the last two weeks doing an upgrade, here's part 1 of a two-part writeup: www.gilesthomas.com/2026/04/10g-...
gilesthomas.com
10Gb Ethernet: what I had to (re)learn
Wired networking in the home and small offices has been pretty stagnant for ages, so upgrading my home from 2.5Gb ethernet to 10Gb meant I had to (re)learn a few things.
110
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 22/04/2026
The appendices in "Build an LLM (from Scratch)" have a lot of really useful stuff. Should I have read them before heading off on my own training runs? I think no -- it would have saved time, but by trying, failing, and trying again, I learned more. YMMV! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 33 -- what I learned from finally getting round to the appendices
I've finished the main body of Raschka's book -- here are my thoughts on the appendices, and why I'm glad I read them after learning things the hard way.
101
Giles Thomas @gilesthomas.com · 22/04/2026
The appendices in "Build an LLM (from Scratch)" have a lot of really useful stuff. Should I have read them before heading off on my own training runs? I think no -- it would have saved time, but by trying, failing, and trying again, I learned more. YMMV! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 33 -- what I learned from finally getting round to the appendices
I've finished the main body of Raschka's book -- here are my thoughts on the appendices, and why I'm glad I read them after learning things the hard way.
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 21/04/2026
Time to wrap up my "Interventions" series: www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32m -- Interventions: conclusion
Wrapping up my 'Interventions' mini-series: what I've learned, what I've achieved, and what's next!
101
Giles Thomas @gilesthomas.com · 21/04/2026
Time to wrap up my "Interventions" series: www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32m -- Interventions: conclusion
Wrapping up my 'Interventions' mini-series: what I've learned, what I've achieved, and what's next!
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 21/04/2026
Some interesting results with instruction fine-tuning. In general you'd expect that a model with lower loss on your test set would fine-tune to do better instruction-following. But the correlation doesn't seem as close as I would have thought: www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32l -- Interventions: updated instruction fine-tuning results
I wanted to revisit the instruction fine-tuning tests that I'd put on hold, and try them with my new models. I found that loss predicts real-world usefulness much less than I would have thought!
101
Giles Thomas @gilesthomas.com · 21/04/2026
Some interesting results with instruction fine-tuning. In general you'd expect that a model with lower loss on your test set would fine-tune to do better instruction-following. But the correlation doesn't seem as close as I would have thought: www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32l -- Interventions: updated instruction fine-tuning results
I wanted to revisit the instruction fine-tuning tests that I'd put on hold, and try them with my new models. I found that loss predicts real-world usefulness much less than I would have thought!
101
Giles Thomas @gilesthomas.com · 17/04/2026
I had 57 checkpoints from my last local LLM training run on my disk, and thought it would be interesting to use them to show how the model got more coherent over time. By 1/3 of the way through, it was surprisingly solid! www.gilesthomas.com/2026/04/how-...
gilesthomas.com
How an LLM becomes more coherent as we train it
As an LLM is trained, it gradually learns how to generate increasingly coherent text. Here's an example.
000
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 15/04/2026
Training a new GPT-2-small-style base model on my RTX 3090 in 40 hours: gradient accumulation, along with some of the other interventions I've experimented with, got me almost (but not quite) all the way to the quality of the original OpenAI model! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32k -- Interventions: training a better model locally with gradient accumulation
Having worked out which combination of interventions into my model and training run improved loss the most, it was time to see how well that worked with a local run, which meant I had to learn about g...
101
Giles Thomas @gilesthomas.com · 15/04/2026
Training a new GPT-2-small-style base model on my RTX 3090 in 40 hours: gradient accumulation, along with some of the other interventions I've experimented with, got me almost (but not quite) all the way to the quality of the original OpenAI model! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32k -- Interventions: training a better model locally with gradient accumulation
Having worked out which combination of interventions into my model and training run improved loss the most, it was time to see how well that worked with a local run, which meant I had to learn about g...
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 09/04/2026
After spending two months trying out different interventions on my GPT-2-style model, it was time to try stacking them up. Interesting results! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32j -- Interventions: trying to train a better model in the cloud
Now that I've tried a number of interventions into my model and training run, and some of them seem to improve the model, how do we stack them together, and what are the results?
101
Giles Thomas @gilesthomas.com · 09/04/2026
After spending two months trying out different interventions on my GPT-2-style model, it was time to try stacking them up. Interesting results! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32j -- Interventions: trying to train a better model in the cloud
Now that I've tried a number of interventions into my model and training run, and some of them seem to improve the model, how do we stack them together, and what are the results?
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 07/04/2026
I wanted to see whether my results when testing interventions to my GPT-2-style training loop were signal or noise. The results were promising! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32i -- Interventions: what is in the noise?
How much of the variation in my training runs is real signal, and how much is random noise? I trained seven more models to find out.
101
Giles Thomas @gilesthomas.com · 07/04/2026
I wanted to see whether my results when testing interventions to my GPT-2-style training loop were signal or noise. The results were promising! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32i -- Interventions: what is in the noise?
How much of the variation in my training runs is real signal, and how much is random noise? I trained seven more models to find out.
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 04/04/2026
My final intervention test, in which I discover that there is such a thing as a free lunch, and it's called AMP. www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32h -- Interventions: full fat float32
I've been using PyTorch's Automated Mixed Precision (AMP) and lower-precision matrix multiplication logic for my training runs so far for larger batch sizes and faster training. Does doing that lead ...
101
Giles Thomas @gilesthomas.com · 04/04/2026
My final intervention test, in which I discover that there is such a thing as a free lunch, and it's called AMP. www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32h -- Interventions: full fat float32
I've been using PyTorch's Automated Mixed Precision (AMP) and lower-precision matrix multiplication logic for my training runs so far for larger batch sizes and faster training. Does doing that lead ...
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 24/03/2026
Weight tying, by contrast with weight decay, was actually really easy! It didn't help, though :-( www.gilesthomas.com/2026/03/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32g -- Interventions: weight tying
Weight tying is apparently not used in modern LLMs, and intuitively would worsen performance in general. Does it?
101
Giles Thomas @gilesthomas.com · 24/03/2026
Weight tying, by contrast with weight decay, was actually really easy! It didn't help, though :-( www.gilesthomas.com/2026/03/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32g -- Interventions: weight tying
Weight tying is apparently not used in modern LLMs, and intuitively would worsen performance in general. Does it?
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 24/03/2026
Weight decay is conceptually simpler than I worried it might be, but still pretty fiddly to get right... www.gilesthomas.com/2026/03/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32f -- Interventions: weight decay
What is weight decay, and what is the right value to set it to in order to get the best possible training run for our model?
101
Giles Thomas @gilesthomas.com · 24/03/2026
Weight decay is conceptually simpler than I worried it might be, but still pretty fiddly to get right... www.gilesthomas.com/2026/03/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32f -- Interventions: weight decay
What is weight decay, and what is the right value to set it to in order to get the best possible training run for our model?
101