Sign in

Giles Thomas

@gilesthomas.com
225 followers 67 following 139 posts

Building LLMs, learning in public / Founded and led @PythonAnywhere.com / PSF Fellow / Writing at gilesthomas.com

PostsRepliesMedia
Giles Thomas @gilesthomas.com · 10/09/2026
I extended the GPT-2-style code from @rasbt's "Build a Large Language Model (from Scratch)" so that it was a 6-expert (2 active) mixture-of-experts, and trained it from scratch over 8 days. It worked well! Full writeup with maths and code at www.gilesthomas.com/2026/09/gpt-...
gilesthomas.com
Extending Raschka's GPT-2: an MoE trained from scratch on an RTX 3090
How mixture-of-experts LLMs work, both in theory and in real working code based on 'Build a Large Language Model (from Scratch)'.
010
Giles Thomas @gilesthomas.com · 03/09/2026
A new blog post: Putting my JAX-trained models on the Hugging Face Hub www.gilesthomas.com/2026/09/jax-...
gilesthomas.com
Putting my JAX-trained models on the Hugging Face Hub
I'd not been putting my JAX models on the Hugging Face Hub because it seemed hard, but then realised I could just upload the tensors converted to PyTorch
000
Giles Thomas @gilesthomas.com · 25/08/2026
D2 seems pretty nifty for diagramming. Just added support for it to my blog's static site generator: www.gilesthomas.com/2026/08/addi...
gilesthomas.com
Adding diagrams to my static site generator with D2
I wanted to have an easier way to put diagrams on this blog. D2 seems to work pretty well!
010
Giles Thomas @gilesthomas.com · 20/08/2026
By switching from a hand-rolled GELU to PyTorch's built-in one, I improved my LLM training speed -- and much more than I expected, from 21,000 tokens/second to 25,000! www.gilesthomas.com/2026/08/buil...
gilesthomas.com
Use the built-in GELU, don't roll your own!
Switching from the hand-rolled GELU module I'd been using to PyTorch's built-in one unsurprisingly made training models faster -- but I was surprised to see it was 20% faster.
010
Giles Thomas @gilesthomas.com · 07/08/2026
I'd trained some models on 40 tokens per parameter. The Chinchilla paper says that doing that is suboptimal. I wanted to see if that was correct for my training setup -- and it looks like it was :-) www.gilesthomas.com/2026/08/chin...
gilesthomas.com
A quick(ish) Chinchilla check
Having overtrained two models, I decided to use an ~equivalent number of FLOPs to train Chinchilla-optimal models. Would the rule hold up?
010
Giles Thomas @gilesthomas.com · 31/07/2026
I thought I'd put together a quick post on how I use AI when writing my blog: AIs identify problems and then I fix them myself. www.gilesthomas.com/2026/07/ai-use
gilesthomas.com
How I use AI on this blog
A point-in-time snapshot of how I'm currently using AI to help with my experiments, and with blogging about them
000
Giles Thomas @gilesthomas.com · 29/07/2026
The first in what will probably be an occasional series on my blog: Why do OpenAI's GPT-2 weights beat mine? www.gilesthomas.com/2026/07/why-...
gilesthomas.com
Why do OpenAI's GPT-2 weights beat mine?
The original GPT-2 small weights are much better than my own models at a specific instruction-following task. Time to look into why!
110
Giles Thomas @gilesthomas.com · 24/07/2026
A friend asked me how many tps I could get with Qwen 3.6 35B MoE on my RTX 3090. I may have overthought the question, but maybe the results are useful for other people? www.gilesthomas.com/2026/07/benc...
gilesthomas.com
Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090
A friend asked me how many tokens per second I could get on Qwen 3.6 35B MoE on an RTX 3090. I probably dug into it more than he expected.
000
Giles Thomas @gilesthomas.com · 10/07/2026
With no weight tying, the token embeddings and the output head alone make up almost half of a GPT-2-small sized model. That makes sense, but it's kind of unintuitive! I asked GPT 5.6 Sol to create a visualiser: www.gilesthomas.com/post-assets/...
gilesthomas.com
GPT-2 Parameter Counter
Explore how embedding width and layer count shape a GPT-2 model's parameter count.
000
Giles Thomas @gilesthomas.com · 09/07/2026
I'm repurposing an old PC as a dedicated LLM training box, and blogging the process. Here's part 1, featuring an almost-melted CPU and an accidental 11-day LLM training run: www.gilesthomas.com/2026/07/popp...
gilesthomas.com
poppy the training box, part 1: the beginnings
Repurposing an old SFF PC as a dedicated LLM training box: a second-hand RTX 3090, an accidental 11-day train on a GTX 1660, and one dead CPU fan.
000
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 08/07/2026
And that's it! After 18 months, the capstone in my LLM from scratch journey. To test my knowledge, I built GPT-2 small in JAX, using just my notes. It worked really well, and I found an interesting way to put it together bit-by-bit, watching the loss go down. www.gilesthomas.com/2026/07/llm-...
gilesthomas.com
Writing an LLM from scratch, part 34b -- from bigrams to GPT-2, one component at a time (in JAX)
The capstone of my LLM from scratch journey: building up from a bigram-style model to GPT-2 small in JAX, watching the loss fall as each component goes in.
011
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 30/06/2026
Coming to the end of my LLM from scratch journey... I'm building a GPT-2-small in JAX, and decided to build the training loop first, then build the model step-by-step. Here's how I build the training loop. www.gilesthomas.com/2026/06/llm-...
gilesthomas.com
Writing an LLM from scratch, part 34a -- building a JAX training loop for an LLM training run
I'm going to build an LLM from scratch in JAX, but I wanted to get a training loop working first so that I could put it together piece by piece.
111
Giles Thomas @gilesthomas.com · 17/06/2026
A bit of a debugging journey with JAX, Flax and NNX: www.gilesthomas.com/2026/06/hash...
gilesthomas.com
Flax debugging: making a hash of things
A quick way to help with debugging plumbing issues in JAX training loops.
020
Giles Thomas @gilesthomas.com · 16/06/2026
Switching from a power-hungry Marvell-based SFP+ 10GBASE-T module to a cooler Broadcom-based one: some oddities, including a mendacious EEPROM. www.gilesthomas.com/2026/06/10g-...
gilesthomas.com
10Gb/s Ethernet: switching to a Broadcom SFP+ module
As predicted by several people, one of my Marvell-based SFP+ modules had overheating problems. I've switched it over to a Broadcom-based one, and it looks better, though I can't be certain.
010
Giles Thomas @gilesthomas.com · 15/06/2026
Hit some interesting memory management weirdness in JAX: arrays can move around between devices when you don't expect it, if you don't commit them to one explicitly: www.gilesthomas.com/2026/06/jax-...
gilesthomas.com
JAX: commitment issues
JAX memory management is tricky; data can be moved around unexpectedly if you don't make sure it's committed to a device.
000
Giles Thomas @gilesthomas.com · 05/06/2026
More notes on working with JAX: backends and devices. www.gilesthomas.com/2026/06/jax-...
gilesthomas.com
JAX backends and devices
Some basic notes on how to get JAX to load data to a particular device.
000
Giles Thomas @gilesthomas.com · 04/06/2026
Getting JAX/Flax and Safetensors to play nicely was a bit fiddly -- wrote it up here: www.gilesthomas.com/2026/06/flax...
gilesthomas.com
Using Safetensors with Flax
Using Safetensors to save Flax models isn't all that hard, so long as you know the trick.
000
Giles Thomas @gilesthomas.com · 30/05/2026
I've spent some time learning JAX over the last month, and I have Thoughts: www.gilesthomas.com/2026/05/on-f...
gilesthomas.com
On first looking into JAX
Some (well, quite a lot of) initial thoughts about JAX, having spent a short time playing with it.
000
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 29/04/2026
...and here's part 2: what I actually did. Turns out you could use an SFP+ 10GBASE-T module to make a (very small) cup of tea: www.gilesthomas.com/2026/04/10g-...
gilesthomas.com
10Gb/s Ethernet: what I actually did to get it working in my home
I already had 2.5Gb/s working. Here's how I upgraded the whole house to 10Gb/s Ethernet -- including real iperf3 numbers, MikroTik switches, scary thermals, and what actually worked.
101
Giles Thomas @gilesthomas.com · 28/04/2026
WiFi has made home wired networking boring for the last 20 years or so, I think, but 10Gb/s "fixes" that :-) I've had a bit of fun over the last two weeks doing an upgrade, here's part 1 of a two-part writeup: www.gilesthomas.com/2026/04/10g-...
gilesthomas.com
10Gb Ethernet: what I had to (re)learn
Wired networking in the home and small offices has been pretty stagnant for ages, so upgrading my home from 2.5Gb ethernet to 10Gb meant I had to (re)learn a few things.
110
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 22/04/2026
The appendices in "Build an LLM (from Scratch)" have a lot of really useful stuff. Should I have read them before heading off on my own training runs? I think no -- it would have saved time, but by trying, failing, and trying again, I learned more. YMMV! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 33 -- what I learned from finally getting round to the appendices
I've finished the main body of Raschka's book -- here are my thoughts on the appendices, and why I'm glad I read them after learning things the hard way.
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 21/04/2026
Time to wrap up my "Interventions" series: www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32m -- Interventions: conclusion
Wrapping up my 'Interventions' mini-series: what I've learned, what I've achieved, and what's next!
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 21/04/2026
Some interesting results with instruction fine-tuning. In general you'd expect that a model with lower loss on your test set would fine-tune to do better instruction-following. But the correlation doesn't seem as close as I would have thought: www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32l -- Interventions: updated instruction fine-tuning results
I wanted to revisit the instruction fine-tuning tests that I'd put on hold, and try them with my new models. I found that loss predicts real-world usefulness much less than I would have thought!
101
Giles Thomas @gilesthomas.com · 17/04/2026
I had 57 checkpoints from my last local LLM training run on my disk, and thought it would be interesting to use them to show how the model got more coherent over time. By 1/3 of the way through, it was surprisingly solid! www.gilesthomas.com/2026/04/how-...
gilesthomas.com
How an LLM becomes more coherent as we train it
As an LLM is trained, it gradually learns how to generate increasingly coherent text. Here's an example.
000
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 15/04/2026
Training a new GPT-2-small-style base model on my RTX 3090 in 40 hours: gradient accumulation, along with some of the other interventions I've experimented with, got me almost (but not quite) all the way to the quality of the original OpenAI model! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32k -- Interventions: training a better model locally with gradient accumulation
Having worked out which combination of interventions into my model and training run improved loss the most, it was time to see how well that worked with a local run, which meant I had to learn about g...
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 09/04/2026
After spending two months trying out different interventions on my GPT-2-style model, it was time to try stacking them up. Interesting results! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32j -- Interventions: trying to train a better model in the cloud
Now that I've tried a number of interventions into my model and training run, and some of them seem to improve the model, how do we stack them together, and what are the results?
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 07/04/2026
I wanted to see whether my results when testing interventions to my GPT-2-style training loop were signal or noise. The results were promising! www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32i -- Interventions: what is in the noise?
How much of the variation in my training runs is real signal, and how much is random noise? I trained seven more models to find out.
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 04/04/2026
My final intervention test, in which I discover that there is such a thing as a free lunch, and it's called AMP. www.gilesthomas.com/2026/04/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32h -- Interventions: full fat float32
I've been using PyTorch's Automated Mixed Precision (AMP) and lower-precision matrix multiplication logic for my training runs so far for larger batch sizes and faster training. Does doing that lead ...
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 24/03/2026
Weight tying, by contrast with weight decay, was actually really easy! It didn't help, though :-( www.gilesthomas.com/2026/03/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32g -- Interventions: weight tying
Weight tying is apparently not used in modern LLMs, and intuitively would worsen performance in general. Does it?
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 24/03/2026
Weight decay is conceptually simpler than I worried it might be, but still pretty fiddly to get right... www.gilesthomas.com/2026/03/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32f -- Interventions: weight decay
What is weight decay, and what is the right value to set it to in order to get the best possible training run for our model?
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 10/03/2026
Learning rates for LLMs turned out to be a deep topic, but making some tweaks certainly seems to help my base model train: www.gilesthomas.com/2026/03/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32e -- Interventions: the learning rate
The learning rate is an essential hyperparameter for training models with gradient descent, but what actual values should you use, and what does it mean to schedule it?
111
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 07/02/2026
Now this was a surprise. QKV bias is not meant to be useful -- but with my GPT-2 small model, it looks like it is! www.gilesthomas.com/2026/02/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32d -- Interventions: adding attention bias
Having bias terms for the query, key, and value attention weights is apparently no longer used because it doesn't help. Let's check that it really doesn't for our model!
111
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 05/02/2026
Does removing dropout improve our baseline model's test loss? Yes, absolutely, and much more than gradient clipping did. www.gilesthomas.com/2026/02/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32c -- Interventions: removing dropout
Does removing dropout improve our baseline model's test loss? Yes, absolutely, and even more than gradient clipping did.
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 05/02/2026
First "intervention" test: does adding gradient clipping improve our baseline model by lessening the loss spikes during training? It does, but it turned out to be more of a rabbit hole than I expected. www.gilesthomas.com/2026/02/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32b -- Interventions: gradient clipping
Does adding gradient clipping improve our baseline model by lessening the loss spikes during training? It does, but it turned out to be more of a rabbit hole than I expected.
111
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 04/02/2026
Back to my LLM from scratch series. I want to train the *best* GPT-2-style model that I can locally in two days, and there are various levers to pull. Working out which ones work means I need a baseline for comparison. www.gilesthomas.com/2026/02/llm-...
gilesthomas.com
Writing an LLM from scratch, part 32a -- Interventions: training a baseline model
I want to try a bunch of interventions like gradient clipping and removing dropout to see if my models get better. I need a baseline train without them so that I can be sure of the results.
101
Giles Thomas @gilesthomas.com · 28/01/2026
I wanted to get a custom LLM up onto the @hf.co Hub, and couldn't find an in-depth tutorial. Here's the one I wish I'd found before I got started: www.gilesthomas.com/2026/01/cust...
gilesthomas.com
Getting a custom PyTorch LLM onto the Hugging Face Hub (Transformers: AutoModel, pipeline, and Trainer)
A worked example of packaging a from-scratch GPT-2-style model for the Hugging Face Hub so it loads via from_pretrained, runs with pipeline, and trains with Trainer -- with notes on tokeniser gotchas.
000
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 17/01/2026
I thought it would be good to share the base models I've been training on Hugging Face, and now they are :-) www.gilesthomas.com/2026/01/llm-...
gilesthomas.com
Writing an LLM from scratch, part 31 -- the models are now on Hugging Face
I've trained seven models using the GPT-2 architecture: let's share them on Hugging Face!
111
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 09/01/2026
I wanted to dig into why the results I got on instruction fine-tuning for each of my models didn't seem to match up well with the loss on the test set. Got some interesting results: www.gilesthomas.com/2026/01/2026...
gilesthomas.com
Writing an LLM from scratch, part 30 -- digging into the LLM-as-a-judge results
I was unhappy with the LLM-as-a-judge instruction fine-tuning results I got when comparing my various base models. Could I make them any better?
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 07/01/2026
Having trained a GPT-2 scale base model from scratch in 48 hours locally, I wanted to see if I could do the same faster and at a reasonable cost in the cloud. I could! www.gilesthomas.com/2026/01/llm-...
gilesthomas.com
Writing an LLM from scratch, part 29 -- using DistributedDataParallel to train a base model from scratch in the cloud
Having trained a base model from scratch on my own machine over 48 hours, I wanted to make it faster by training with multiple GPUs in the cloud.
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 02/12/2025
I managed to train my own base model from scratch on an RTX 3090! Very detailed notes here: www.gilesthomas.com/2025/12/llm-...
gilesthomas.com
Writing an LLM from scratch, part 28 -- training a base model from scratch on an RTX 3090
I felt like it should be possible to train a GPT-2 small level model on my own hardware using modern tools and open datasets from scratch. It was!
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 04/11/2025
So, what's left to do in my series on building an LLM from scratch? And what follow-up series should I work on? Some musings: www.gilesthomas.com/2025/11/llm-...
gilesthomas.com
Writing an LLM from scratch, part 27 -- what's left, and what's next?
Having finished the main body of 'Build an LLM (from scratch)', it's time to think about what I need to do to treat this project as fully done
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 03/11/2025
The end of the beginning... running evals on our model using Llama 3 is the last part of the main body of @sebastianraschka.com's "Build an LLM (from scratch)". Here's my writeup: www.gilesthomas.com/2025/11/llm-...
gilesthomas.com
Writing an LLM from scratch, part 26 -- evaluating the fine-tuned model
Coming to the end of 'Build an LLM (from scratch)'! We evaluate the quality of the responses our model produces.
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 29/10/2025
Back on track with chapter 7 of "Build an LLM (from scratch)": notes on instruction fine-tuning of our GPT-2 model: www.gilesthomas.com/2025/10/llm-...
gilesthomas.com
Writing an LLM from scratch, part 25 -- instruction fine-tuning
Some notes on the first part of chapter 7 of 'Build an LLM (from scratch)': instruction fine-tuning
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 28/10/2025
Back when I started messing with LLMs, it looked to me like you could get reasonably OK results for chat applications without instruction fine-tuning. So before getting into Chapter 7 of "Build an LLM (from scratch)", I decided to see if that was really true: www.gilesthomas.com/2025/10/llm-...
gilesthomas.com
Writing an LLM from scratch, part 24 -- the transcript hack
Back when I started playing with LLMs, I found that you could build a (very basic) chatbot with a base model -- no instruction fine-tuning at all! Does that work with GPT-2?
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 24/10/2025
And the next step -- a code walkthrough of my PyTorch version of Karpathy's 2015-vintage RNNs. www.gilesthomas.com/2025/10/retr...
gilesthomas.com
Retro Language Models: Rebuilding Karpathy’s RNN in PyTorch
Revisiting Karpathy’s text-generating RNNs with PyTorch’s built-in LSTM class — a practical look at why training sequence models is so different from Transformers.
001
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 23/10/2025
Chapter 6 was easy and fun! Fine-tuning an LLM for classification tasks, with some initially disappointing results -- but it all came out in the wash: www.gilesthomas.com/2025/10/llm-...
gilesthomas.com
Writing an LLM from scratch, part 23 -- fine-tuning for classification
After all the hard work, chapter 6 in 'Build an LLM (from scratch)' is a nice easy one -- how do we take a next-token predictor and turn it into a classifier?
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 16/10/2025
Part 22 is live: we finally train the LLM :-) Following @sebastianraschka.com's book, we train on Edith Wharton, then swap in GPT-2 (124M) weights for comparison. Notes on seeding, AdamW, temperature and top-k. www.gilesthomas.com/2025/10/llm-...
gilesthomas.com
Writing an LLM from scratch, part 22 -- finally training our LLM!
Finally, we train an LLM! The final part of Chapter 5 of Build an LLM (from Scratch) runs the model on real text, then loads OpenAI’s GPT-2 weights for comparison.
211
Giles Thomas @gilesthomas.com · 11/10/2025
Decided to learn about some "retro" language models in parallel with LLMs; here's my first post on that -- Revisiting Karpathy’s 'The Unreasonable Effectiveness of Recurrent Neural Networks'. www.gilesthomas.com/2025/10/revi...
gilesthomas.com
Revisiting Karpathy’s 'The Unreasonable Effectiveness of Recurrent Neural Networks'
Andrej Karpathy's 2015 blog post 'The Unreasonable Effectiveness of Recurrent Neural Networks' went viral in its day, for good reason. How does it read ten years later?
140
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 07/10/2025
Next up in my LLM from scratch series, some serious yak shaving on something @sebastianraschka.com covers in a sidebar: perplexity. www.gilesthomas.com/2025/10/llm-...
gilesthomas.com
Writing an LLM from scratch, part 21 -- perplexed by perplexity
Raschka calls out perplexity in a sidebar, but I wanted to understand it in a little more depth
101
Reposted by Giles Thomas
Giles Thomas @gilesthomas.com · 02/10/2025
Back to the main track of my LLM from scratch posts: cross entropy -- what it is and why we use it. www.gilesthomas.com/2025/10/llm-...
gilesthomas.com
Writing an LLM from scratch, part 20 -- starting training, and cross entropy loss
Starting training our LLM requires a loss function, which is called cross entropy loss. What is this and why does it work?
121