Most language models split text into subwords—words or fragments of words from a fixed vocabulary. This can obscure spelling details across writing systems & split meaningful units in code or math.
Byte-level models work directly with the bytes computers use to represent text.