Sign in

Philip Whittington

@philipwitti.bsky.social
48 followers 34 following 0 posts

Doctoral student @ETH Zürich 🇨🇭

PostsRepliesMedia
Reposted by Philip Whittington
Marco @mcognetta.bsky.social · 30/09/2026
🚨 [Token][ization] Paper Alert 🚨 Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field. Check it out!
112935
Reposted by Philip Whittington
Tiago Pimentel @tpimentel.bsky.social · 22/05/2026
Our new paper reformulates tokenisation as a linear program (LP), which we solve to get SOTA tokenisers 😁 As a bonus, this LP tells us how close to optimal any tokeniser is! Check it out 👇 w/ J. Tempus, @philipwitti.bsky.social, @craigschmidt.com, D. Komm Paper: arxiv.org/abs/2605.22821
14410
Reposted by Philip Whittington
Tiago Pimentel @tpimentel.bsky.social · 20/11/2025
Tokenisers are a vital part of LLMs, but how hard is it to find an optimal one? 🤔 Considering arbitrarily large alphabets, prior work showed this is NP-hard. But what if we use bytes instead? Or unary strings like a, aa, aaa, ...? In our new paper, we show this is still hard, NP-hard!
Screenshot of paper title: Tokenisation over Bounded Alphabets is Hard.
1173
Reposted by Philip Whittington
Tiago Pimentel @tpimentel.bsky.social · 31/07/2025
Honoured to receive two (!!) SAC highlights awards at #ACL2025 😁 (Conveniently placed on the same slide!) With the amazing: @philipwitti.bsky.social, @gregorbachmann.bsky.social and @wegotlieb.bsky.social, @cuiding.bsky.social, Giovanni Acampa, @alexwarstadt.bsky.social, @tamaregev.bsky.social
0223
Reposted by Philip Whittington
Tiago Pimentel @tpimentel.bsky.social · 20/12/2024
BPE is a greedy method to find a tokeniser which maximises compression! Why don't we try to find properly optimal tokenisers instead? Well, it seems this is a pretty difficult—in fact, NP-complete—problem!🤯 New paper + @philipwitti.bsky.social @gregorbachmann.bsky.social :) arxiv.org/abs/2412.15210
arxiv.org
Tokenisation is NP-Complete
In this work, we prove the NP-completeness of two variants of tokenisation, defined as the problem of compressing a dataset to at most $δ$ symbols by either finding a vocabulary directly (direct token...
1458