Reposted by Marco
Tokenizer research lags behind other areas because it is hard to validate changes and a lack of good predictive metrics/scaling laws. It is also just hard to compare across tokenizers.
Let's have a community benchmark where everything but the tokenizer is fixed!
alphaxiv.org
A Call for an Open Tokenizer Benchmark
Tokenizer research lags behind other areas of the language modeling stack due in part to the difficulty of: 1) comparing across models with different tokenizers, 2) quickly and accurately measuring...