Reposted by Giles Thomas
Next step in trying to work out why GPT-2 small beats my own similar models: is it the data quality? Maybe...
www.gilesthomas.com/2026/10/why-...
gilesthomas.com
Why do OpenAI's GPT-2 weights beat mine? Part five: data quality
Does improving the training data quality help me close the gap between my models and the original GPT-2 small on my IFT eval?