Sign in

Craig Schmidt

@craigschmidt.com
600 followers 2.4K following 67 posts

Interested in ML, AI, and NLP. Particularly interested in tokenization. Live in the Boston area and work in R&D at Kensho Technologies.

PostsRepliesMedia
Craig Schmidt @craigschmidt.com · 4h
Congratulations to Marco and all the other co-authors on this groundbreaking 154-page survey. If you're at all interested in tokenization, give it a read.
051
Reposted by Craig Schmidt
Xiulin Yang @xiulinyang.bsky.social · 28/08/2026
🍎🍊 How would you know if a language model is better at one language than another? Our #EMNLP2026 paper argues that only one metric can actually lead to fair crosslingual evaluation. This work is a collaboration with @wegotlieb.bsky.social & @catherinearnett.bsky.social! (1/5)
13211
Craig Schmidt @craigschmidt.com · 21/08/2026
Yes at ACL 2026 a huge number of posters were obviously Claude generated
020
Craig Schmidt @craigschmidt.com · 15/07/2026
I’m sure it varies by area. I work in tokenization, which seems more prevalent than that. However, it usually is just a footnote about releasing code at submission time, with the code being released later. Our group always assumes if the reviewers won’t read the appendix they won’t look at the code.
110
Craig Schmidt @craigschmidt.com · 15/07/2026
I think open code and data are so common it is kind of assumed. A dataset paper with the dataset wouldn’t be much. In a review I would ask them to confirm.
110
Reposted by Craig Schmidt
David Jurgens @davidjurgens.bsky.social · 13/07/2026
Peer review was one of the most-discussed topics at #ACL2026 . Many folks were concerned about the incredible growth in the number of ARR submissions (17K for the May26 ARR cycle 😱), and even more shocked that ~40% didn't have any authors qualified to review. What is going on?? I did some digging...
3253
Craig Schmidt @craigschmidt.com · 03/07/2026
I’m at the airport flying to #ACL2026 in San Diego. If anyone else there wants to talk about tokenization, let meet up.
051
Craig Schmidt @craigschmidt.com · 15/06/2026
Ordered, thanks! I love Prestige era jazz.
010
Craig Schmidt @craigschmidt.com · 07/06/2026
I personally love the Economist for their global perspective, and the nytimes for day to day news
000
Craig Schmidt @craigschmidt.com · 26/05/2026
So the most common merge pairs are never used, and tokens like " the" and " a" were never learned. Our baselines used 175 GB of CulturaX English text; the bug occurs after around 108 GB. We are very grateful to Sander Land for identifying that the baselines were missing common merge pairs.
010
Craig Schmidt @craigschmidt.com · 26/05/2026
The bug is Hugging Face tokenizers issue #2058 (github.com/huggingface/...): the library counts merge pairs in an i32 hash map, and once any pair's count crosses 2^31 − 1 = 2,147,483,647 the counter wraps to a negative value and the pair never gets selected as a merge.
120
Craig Schmidt @craigschmidt.com · 26/05/2026
We can't claim anything about ToaST's performance until we retrain the baselines and rerun the evaluation, and we expect a lot of ToaST's apparent advantage was really just broken baselines. ToaST itself runs through our own code, and those numbers are correct.
110
Craig Schmidt @craigschmidt.com · 26/05/2026
Unfortunately, we're withdrawing our paper "Tokenization with Split Trees" from arXiv. All our baseline tokenizers — BPE, WordPiece, and Unigram — were trained incorrectly because of a bug in the Hugging Face tokenizers library, so every comparison to ToaST in the paper is invalid.
100
Reposted by Craig Schmidt
Tiago Pimentel @tpimentel.bsky.social · 22/05/2026
Our new paper reformulates tokenisation as a linear program (LP), which we solve to get SOTA tokenisers 😁 As a bonus, this LP tells us how close to optimal any tokeniser is! Check it out 👇 w/ J. Tempus, @philipwitti.bsky.social, @craigschmidt.com, D. Komm Paper: arxiv.org/abs/2605.22821
14410
Craig Schmidt @craigschmidt.com · 22/05/2026
arxiv.org/abs/2605.22705 arxiv.org/abs/2605.22821 Happy Linear Programming for Tokenization day! I was involved with two separate papers that hit ArXiv yesterday, using LP's to find the vocabulary maximizing compression, depending on the kind of inference you want to use.
arxiv.org
Tokenization with Split Trees
We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference procedure. ToaST greedily splits each pretoken int...
2166
Reposted by Craig Schmidt
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 14/05/2026
TokShop will be at #COLM2026! 🗓️ October 9th, 2026 📍 San Francisco, USA More details and a call for papers coming soon.
0117
Craig Schmidt @craigschmidt.com · 06/05/2026
At least there is the May 25th ARR as a good fallback once you get it revived
110
Craig Schmidt @craigschmidt.com · 14/01/2026
Gandalf the White. A quote for our times.
100
Craig Schmidt @craigschmidt.com · 01/01/2026
The red cups are a brand called solo cups. They have always been red
020
Craig Schmidt @craigschmidt.com · 07/10/2025
I’m at @colmweb.org this week in Montreal. Come see our BoundlessBPE paper in the Wed morning poster session. Love to talk to anyone else here, especially about tokenization. #COLM2025
000
Craig Schmidt @craigschmidt.com · 02/10/2025
I believe he’s talking about Olin College of Engineering. Created from scratch as an undergraduate only school, with the first class in 2002. Kind of a Harvey Mudd of the east. Campus is near me, and they seem to attract great students.
010
Craig Schmidt @craigschmidt.com · 18/09/2025
The other is that is there isn't a way to specify an initial vocabulary with all 256 bytes including the continuation character ##. See github.com/huggingface/.... So in short, if you use their WordPiece you might get <UNK> tokens.
github.com
WordPiece can't always avoid <unk> even with ByteLevel pretokenization. · Issue #1863 · huggingface/tokenizers
The ByteLevel pre-tokenizer is largely used to avoid the possibility of an <unk> token. However, there is a problem with the continuation characters in WordPiece that prevents you from adding all o...
010
Craig Schmidt @craigschmidt.com · 18/09/2025
There are two different ways that the Huggingface Word Piece implementation can produce <UNK> tokens even with ByteLevel pretokenization. A nice blog post from Stéphan Tulkens talks about how to fix one of them, in response to a question of mine. stephantul.github.io/blog/better-...
stephantul.github.io
Better Greedy Tokenizers: Handling WordPiece's [UNK] Problem
Stéphan Tulkens' Blog
221
Craig Schmidt @craigschmidt.com · 10/08/2025
I've been using GPT-5 on my phone (since it isn't my web account yet). I've had several bad responses with logical inconsistencies. My hot take: what if GPT-5 is mostly about saving OpenAI money on inference, which is why they are deprecating all the other models so quickly.
020
Reposted by Craig Schmidt
David Darmofal @daviddarmofal.bsky.social · 03/08/2025
@crampell.bsky.social’s post got me to thinking and…yes…Trump has apparently canceled the research grant of Judea Pearl, who is one of the world’s leading scholars, is Jewish, Israeli-American, & is vocally opposed to antisemitism, & is the father of Daniel Pearl. www.science.org/content/arti...
820991
Reposted by Craig Schmidt
Miryam de Lhoneux @mdlhx.bsky.social · 16/05/2025
Interested in multilingual tokenization in #NLP? Lisa Beinborn and I are hiring! PhD candidate position in Göttingen, Germany: www.uni-goettingen.de/de/644546.ht... PostDoc position in Leuven, Belgium: www.kuleuven.be/personeel/jo... Deadline 6th of June
uni-goettingen.de
Stellen OBP - Georg-August-Universität Göttingen
Webseiten der Georg-August-Universität Göttingen
22513
Craig Schmidt @craigschmidt.com · 30/07/2025
I've posted a few papers I missed including yours here bsky.app/profile/crai.... Thomas pointed that out about 5 seconds after I posted on the discord :-)
110
Craig Schmidt @craigschmidt.com · 30/07/2025
16) Causal Estimation of Tokenisation Bias Pietro Lesci et al aclanthology.org/2025.acl-lon...
aclanthology.org
Causal Estimation of Tokenisation Bias
Pietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos, Tiago Pimentel. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
020
Craig Schmidt @craigschmidt.com · 30/07/2025
15) Tokenisation is NP-Complete Philip Whittington et al aclanthology.org/2025.acl-lon...
aclanthology.org
Tokenisation is NP-Complete
Philip Whittington, Gregor Bachmann, Tiago Pimentel. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
151
Craig Schmidt @craigschmidt.com · 30/07/2025
14) GRaMPa: Subword Regularisation by Skewing Uniform Segmentation Distributions with an Efficient Path-counting Markov Model Thomas Bauwens et al aclanthology.org/2025.acl-lon...
aclanthology.org
GRaMPa: Subword Regularisation by Skewing Uniform Segmentation Distributions with an Efficient Path-counting Markov Model
Thomas Bauwens, David Kaczér, Miryam De Lhoneux. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
120
Craig Schmidt @craigschmidt.com · 30/07/2025
And of course I missed some tokenization related papers at #ACL2025 in my previous post. Any more I should add?
120
Craig Schmidt @craigschmidt.com · 30/07/2025
13) Evaluating Tokenizer Adaptation Methods for Large Language Models on Low-Resource Programming Languages Georgii Andriushchenko et al aclanthology.org/2025.acl-srw...
aclanthology.org
Evaluating Tokenizer Adaptation Methods for Large Language Models on Low-Resource Programming Languages
Georgy Andryushchenko, Vladimir V. Ivanov. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop). 2025.
010
Craig Schmidt @craigschmidt.com · 30/07/2025
12) Retrofitting Large Language Models with Dynamic Tokenization Darius Feher et al aclanthology.org/2025.acl-lon...
aclanthology.org
Retrofitting Large Language Models with Dynamic Tokenization
Darius Feher, Ivan Vulić, Benjamin Minixhofer. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
110
Craig Schmidt @craigschmidt.com · 30/07/2025
11) TokAlign: Efficient Vocabulary Adaptation via Token Alignment Chong Li et al aclanthology.org/2025.acl-lon...
aclanthology.org
TokAlign: Efficient Vocabulary Adaptation via Token Alignment
Chong Li, Jiajun Zhang, Chengqing Zong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
110
Craig Schmidt @craigschmidt.com · 30/07/2025
10) Sticking to the Mean: Detecting Sticky Tokens in Text Embedding Models Kexin Chen et al aclanthology.org/2025.acl-lon...
aclanthology.org
Sticking to the Mean: Detecting Sticky Tokens in Text Embedding Models
Kexin Chen, Dongxia Wang, Yi Liu, Haonan Zhang, Wenhai Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
110
Craig Schmidt @craigschmidt.com · 30/07/2025
9) Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar Andrew Gambardella et al aclanthology.org/2025.acl-sho...
aclanthology.org
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
Andrew Gambardella, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2025.
110
Craig Schmidt @craigschmidt.com · 30/07/2025
8) Adversarial Tokenization Renato Lui Geh et al aclanthology.org/2025.acl-lon...
aclanthology.org
Adversarial Tokenization
Renato Geh, Zilei Shao, Guy Van Den Broeck. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
110
Craig Schmidt @craigschmidt.com · 30/07/2025
7) Incorporating Domain Knowledge into Materials Tokenization Yerim Oh et al aclanthology.org/2025.acl-lon...
aclanthology.org
Incorporating Domain Knowledge into Materials Tokenization
Yerim Oh, Jun-Hyung Park, Junho Kim, SungHo Kim, SangKeun Lee. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
110
Craig Schmidt @craigschmidt.com · 30/07/2025
6) Beyond Text Compression: Evaluating Tokenizers Across Scales Jonas F. Lotz et al aclanthology.org/2025.acl-lon...
aclanthology.org
Beyond Text Compression: Evaluating Tokenizers Across Scales
Jonas F. Lotz, António V. Lopes, Stephan Peitz, Hendra Setiawan, Leonardo Emili. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
110
Craig Schmidt @craigschmidt.com · 30/07/2025
5) Enhancing Character-Level Understanding in LLMs through Token Internal Structure Learning Zhu Xu et al aclanthology.org/2025.acl-lon...
aclanthology.org
Enhancing Character-Level Understanding in LLMs through Token Internal Structure Learning
Zhu Xu, Zhiqiang Zhao, Zihan Zhang, Yuchi Liu, Quanwei Shen, Fei Liu, Yu Kuang, Jian He, Conglin Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1:...
110
Craig Schmidt @craigschmidt.com · 30/07/2025
4) Unsupervised Morphological Tree Tokenizer Xiang Hu et al aclanthology.org/2025.finding...
aclanthology.org
Unsupervised Morphological Tree Tokenizer
Qingyang Zhu, Xiang Hu, Pengyu Ji, Wei Wu, Kewei Tu. Findings of the Association for Computational Linguistics: ACL 2025. 2025.
110
Craig Schmidt @craigschmidt.com · 30/07/2025
3) Splintering Nonconcatenative Languages for Better Tokenization Yuval Pinter et al aclanthology.org/2025.finding...
aclanthology.org
Splintering Nonconcatenative Languages for Better Tokenization
Bar Gazit, Shaltiel Shmidman, Avi Shmidman, Yuval Pinter. Findings of the Association for Computational Linguistics: ACL 2025. 2025.
110
Craig Schmidt @craigschmidt.com · 30/07/2025
2) Tokenization is Sensitive to Language Variation Anna Wegmann et al aclanthology.org/2025.finding...
aclanthology.org
Tokenization is Sensitive to Language Variation
Anna Wegmann, Dong Nguyen, David Jurgens. Findings of the Association for Computational Linguistics: ACL 2025. 2025.
110
Craig Schmidt @craigschmidt.com · 30/07/2025
1) Byte Latent Transformer: Patches Scale Better Than Tokens Artidoro Pagnoni et al aclanthology.org/2025.acl-lon...
aclanthology.org
Byte Latent Transformer: Patches Scale Better Than Tokens
Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason E Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srini...
110
Craig Schmidt @craigschmidt.com · 30/07/2025
I'm sadly not at #ACL2025, but the work on tokenization seem to continue to explode. Here are the tokenization related papers I could find, in no particular order. Let me know if I missed any.
2124
Reposted by Craig Schmidt
Catherine Arnett @catherinearnett.bsky.social · 19/07/2025
Really grateful to the organizers for the recognition of our work!
1121
Craig Schmidt @craigschmidt.com · 17/07/2025
You’re right these results apply to general “big” datasets like ThePile or RedPajama. There are several papers at ICML on weighting datasets like Chameleon (icml.cc/virtual/2025...) that could probably let you get away with less data.
icml.cc
ICML Poster Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and FinetuningICML 2025
110
Craig Schmidt @craigschmidt.com · 17/07/2025
The second is on entropy-driven pre-tokenization for non-space delimited languages, arxiv.org/abs/2506.15889. That came out of a capstone project for a Harvard Masters program. Congrats to them on achieving a peer-reviewed paper.
arxiv.org
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
Byte-Pair Encoding (BPE) has become a widely adopted subword tokenization method in modern language models due to its simplicity and strong empirical performance across downstream tasks. However, appl...
040
Craig Schmidt @craigschmidt.com · 17/07/2025
I'm at #ICML2025 this week in Vancouver. My co-authors and I are presenting two posters at the Tokenization Workshop on Friday tokenization-workshop.github.io. The first is on how much data is useful in training a tokenizer arxiv.org/abs/2502.20273.
arxiv.org
How Much is Enough? The Diminishing Returns of Tokenization Training Data
Tokenization, a crucial initial step in natural language processing, is governed by several key parameters, such as the tokenization algorithm, vocabulary size, pre-tokenization strategy, inference st...
272
Craig Schmidt @craigschmidt.com · 19/06/2025
My son said he couldn’t call me on Father’s Day because he had worked the weekend dealing with a North Korean hacking group. Valid excuse I guess. The hack analysis …
020