Tokenizers are heavy – megabytes of vocab data for a number you can estimate.
tokenx answers "what does this payload cost me?" in 2 kB, at ~96% accuracy.
New in v1.6:
🧮 Heuristics for Cyrillic, kana and emoji
🧱 More accurate on JSON payloads and source code