I extended the GPT-2-style code from @rasbt's "Build a Large Language Model (from Scratch)" so that it was a 6-expert (2 active) mixture-of-experts, and trained it from scratch over 8 days. It worked well! Full writeup with maths and code at www.gilesthomas.com/2026/09/gpt-...
gilesthomas.com
Extending Raschka's GPT-2: an MoE trained from scratch on an RTX 3090
How mixture-of-experts LLMs work, both in theory and in real working code based on 'Build a Large Language Model (from Scratch)'.