Sign in

Craig Schmidt

@craigschmidt.com
600 followers 2.4K following 67 posts

Interested in ML, AI, and NLP. Particularly interested in tokenization. Live in the Boston area and work in R&D at Kensho Technologies.

PostsRepliesMedia
Craig Schmidt @craigschmidt.com · 2h
Congratulations to Marco and all the other co-authors on this groundbreaking 154-page survey. If you're at all interested in tokenization, give it a read.
051
Reposted by Craig Schmidt
Xiulin Yang @xiulinyang.bsky.social · 28/08/2026
🍎🍊 How would you know if a language model is better at one language than another? Our #EMNLP2026 paper argues that only one metric can actually lead to fair crosslingual evaluation. This work is a collaboration with @wegotlieb.bsky.social & @catherinearnett.bsky.social! (1/5)
13211
Reposted by Craig Schmidt
David Jurgens @davidjurgens.bsky.social · 13/07/2026
Peer review was one of the most-discussed topics at #ACL2026 . Many folks were concerned about the incredible growth in the number of ARR submissions (17K for the May26 ARR cycle 😱), and even more shocked that ~40% didn't have any authors qualified to review. What is going on?? I did some digging...
3253
Craig Schmidt @craigschmidt.com · 03/07/2026
I’m at the airport flying to #ACL2026 in San Diego. If anyone else there wants to talk about tokenization, let meet up.
051
Reposted by Craig Schmidt
Tiago Pimentel @tpimentel.bsky.social · 22/05/2026
Our new paper reformulates tokenisation as a linear program (LP), which we solve to get SOTA tokenisers 😁 As a bonus, this LP tells us how close to optimal any tokeniser is! Check it out 👇 w/ J. Tempus, @philipwitti.bsky.social, @craigschmidt.com, D. Komm Paper: arxiv.org/abs/2605.22821
14410
Craig Schmidt @craigschmidt.com · 22/05/2026
arxiv.org/abs/2605.22705 arxiv.org/abs/2605.22821 Happy Linear Programming for Tokenization day! I was involved with two separate papers that hit ArXiv yesterday, using LP's to find the vocabulary maximizing compression, depending on the kind of inference you want to use.
arxiv.org
Tokenization with Split Trees
We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference procedure. ToaST greedily splits each pretoken int...
2166
Reposted by Craig Schmidt
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 14/05/2026
TokShop will be at #COLM2026! 🗓️ October 9th, 2026 📍 San Francisco, USA More details and a call for papers coming soon.
0117
Craig Schmidt @craigschmidt.com · 07/10/2025
I’m at @colmweb.org this week in Montreal. Come see our BoundlessBPE paper in the Wed morning poster session. Love to talk to anyone else here, especially about tokenization. #COLM2025
000
Craig Schmidt @craigschmidt.com · 18/09/2025
There are two different ways that the Huggingface Word Piece implementation can produce <UNK> tokens even with ByteLevel pretokenization. A nice blog post from Stéphan Tulkens talks about how to fix one of them, in response to a question of mine. stephantul.github.io/blog/better-...
stephantul.github.io
Better Greedy Tokenizers: Handling WordPiece's [UNK] Problem
Stéphan Tulkens' Blog
221
Craig Schmidt @craigschmidt.com · 10/08/2025
I've been using GPT-5 on my phone (since it isn't my web account yet). I've had several bad responses with logical inconsistencies. My hot take: what if GPT-5 is mostly about saving OpenAI money on inference, which is why they are deprecating all the other models so quickly.
020
Reposted by Craig Schmidt
David Darmofal @daviddarmofal.bsky.social · 03/08/2025
@crampell.bsky.social’s post got me to thinking and…yes…Trump has apparently canceled the research grant of Judea Pearl, who is one of the world’s leading scholars, is Jewish, Israeli-American, & is vocally opposed to antisemitism, & is the father of Daniel Pearl. www.science.org/content/arti...
820991
Reposted by Craig Schmidt
Miryam de Lhoneux @mdlhx.bsky.social · 16/05/2025
Interested in multilingual tokenization in #NLP? Lisa Beinborn and I are hiring! PhD candidate position in Göttingen, Germany: www.uni-goettingen.de/de/644546.ht... PostDoc position in Leuven, Belgium: www.kuleuven.be/personeel/jo... Deadline 6th of June
uni-goettingen.de
Stellen OBP - Georg-August-Universität Göttingen
Webseiten der Georg-August-Universität Göttingen
22513
Craig Schmidt @craigschmidt.com · 30/07/2025
And of course I missed some tokenization related papers at #ACL2025 in my previous post. Any more I should add?
120
Craig Schmidt @craigschmidt.com · 30/07/2025
I'm sadly not at #ACL2025, but the work on tokenization seem to continue to explode. Here are the tokenization related papers I could find, in no particular order. Let me know if I missed any.
2124
Reposted by Craig Schmidt
Catherine Arnett @catherinearnett.bsky.social · 19/07/2025
Really grateful to the organizers for the recognition of our work!
1121
Craig Schmidt @craigschmidt.com · 17/07/2025
I'm at #ICML2025 this week in Vancouver. My co-authors and I are presenting two posters at the Tokenization Workshop on Friday tokenization-workshop.github.io. The first is on how much data is useful in training a tokenizer arxiv.org/abs/2502.20273.
arxiv.org
How Much is Enough? The Diminishing Returns of Tokenization Training Data
Tokenization, a crucial initial step in natural language processing, is governed by several key parameters, such as the tokenization algorithm, vocabulary size, pre-tokenization strategy, inference st...
272
Craig Schmidt @craigschmidt.com · 19/06/2025
My son said he couldn’t call me on Father’s Day because he had worked the weekend dealing with a North Korean hacking group. Valid excuse I guess. The hack analysis …
020
Reposted by Craig Schmidt
Conference on Language Modeling @colmweb.org · 20/03/2025
A bit of a mess around the conflict of COLM with the ARR (and to lesser degree ICML) reviews release. We feel this is creating a lot of pressure and uncertainty. So, we are pushing our deadlines: Abstracts due March 22 AoE (+48hr) Full papers due March 28 AoE (+24hr) Plz RT 🙏
33631
Craig Schmidt @craigschmidt.com · 12/02/2025
If you have an interest in tokenization in Natural Language Processing (NLP), this is a nice discord. Come say hi.
140
Reposted by Craig Schmidt
Catherine Arnett @catherinearnett.bsky.social · 24/01/2025
Super honored that this paper received the best paper award at #COLING2025!
0444
Craig Schmidt @craigschmidt.com · 21/12/2024
I got a 70, despite all the time I spent reading the Economist this year.
060
Craig Schmidt @craigschmidt.com · 20/12/2024
The final entry in my #EMNLP2024 fav papers was this paper aclanthology.org/2024.finding... from Thomas L Griffiths' keynote. Used rotational cyphers like ROT-13 and ROT-3 to disentangle forms of reasoning in Chain-of-Thought. Good cypher joke in the keynote! (see p. 24 arxiv.org/abs/2309.13638)
aclanthology.org
Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning
Akshara Prabhakar, Thomas L. Griffiths, R. Thomas McCoy. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024.
051
Craig Schmidt @craigschmidt.com · 20/12/2024
This #EMNLP2024 best paper aclanthology.org/2024.emnlp-m... had large gains over their (somewhat weak) baseline in trying to determine if a given document was in a LLMs pre-training data. Progress in an important problem.
aclanthology.org
Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method
Weichao Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, Xueqi Cheng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
030
Craig Schmidt @craigschmidt.com · 20/12/2024
This #EMNLP2024 outstanding paper (aclanthology.org/2024.emnlp-m..., underline.io/events/469/s...) LMs can learn a rare grammatical construction like "a beautiful five days", even without any examples in the training data, by generalizing from more common phenomenon.
underline.io
Watch lectures from the best researchers.
On-demand video platform giving you access to lectures from conferences worldwide.
030
Craig Schmidt @craigschmidt.com · 20/12/2024
This #EMNLP2024 post (aclanthology.org/2024.emnlp-m..., underline.io/events/469/p...) was about avoiding hallucination without human feedback. If you compare an answer at a higher temperature to a beam search generation, then the latter will be more factual, making preference pairs for DPO.
aclanthology.org
Model-based Preference Optimization in Abstractive Summarization without Human Feedback
Jaepill Choi, Kyubyung Chae, Jiwoo Song, Yohan Jo, Taesup Kim. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
020
Craig Schmidt @craigschmidt.com · 20/12/2024
This paper underline.io/events/469/s... at #EMNLP2024 had one of my favorite takeaways: if you fine tune a LLM on knew knowledge it doesn't know you encourage hallucinations.
underline.io
Watch lectures from the best researchers.
On-demand video platform giving you access to lectures from conferences worldwide.
130
Craig Schmidt @craigschmidt.com · 20/12/2024
I wanted to post of a few of my favorite #EMNLP2024 papers, starting with a couple in tokenization. Fishing For Magicarp explores the problem of undertrained "glitch" tokens, and how they can be identified from their embedding vectors. aclanthology.org/2024.emnlp-m...
aclanthology.org
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
Sander Land, Max Bartolo. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
110
Craig Schmidt @craigschmidt.com · 20/12/2024
Hey @blueskystarterpack.com please add: go.bsky.app/8P9ftjL
021
Craig Schmidt @craigschmidt.com · 20/12/2024
I made a starter pack for people in NLP working in the area of tokenization. Let me know if you'd like to be added go.bsky.app/8P9ftjL
0125
Craig Schmidt @craigschmidt.com · 28/11/2024
I really enjoyed #EMNLP2024. It was an honor to present our tokenization paper aclanthology.org/2024.emnlp-m.... I’m planning to post about some of my favorite papers soon, but here is a nice write up.
aclanthology.org
Tokenization Is More Than Compression
Craig W Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter, Chris Tanner. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
0120
Craig Schmidt @craigschmidt.com · 25/11/2024
As @marcoher.bsky.social noted, this is an interesting confirmation of arxiv.org/pdf/2402.14903. I wonder if Llama3's tokenizer has all 3 digit numbers in the vocab? (GPT4 does, Claude doesn't). If not, it would be fun to look at errors on those missing numbers.
bsky.app
160
Reposted by Craig Schmidt
Kanishka Misra @kanishka.bsky.social · 25/11/2024
There's a known bug in how we compute "word" probabilities with subword-based LMs that mark beginnings of words -- as pointed out by Byung-doh Oh and Will Schuler, & @tpimentel.bsky.social and Clara Meister I'm pleased to announce that minicons now includes a fix which runs batch-wise!
Code: from minicons import scorer

lm = scorer.IncrementalLMScorer("gpt2-xl", "cuda:0")

stimuli = ["I was a matron in France", "I was a mat in France"]

# old way, no correction
# P.S. gpt2 does not automatically add a bos token at the beginning...
lm.token_score(stimuli, bos_token=True, surprisal=True, base_two=True, bow_correction=False)

'''Rounded Output
[[('<|endoftext|>', 0.0),
  ('I', 5.85),
  ('was', 4.28),
  ('a', 4.67),
  ('mat', 16.34),
  ('ron', 1.74),
  ('in', 2.12),
  ('France', 11.43)],
 [('<|endoftext|>', 0.0),
  ('I', 5.85),
  ('was', 4.28),
  ('a', 4.67),
  ('mat', 16.34),
  ('in', 10.78),
  ('France', 10.71)]]
'''

# the new way! notice the surprisal of "mat" in both cases
lm.token_score(stimuli, bos_token=True, surprisal=True, base_two=True, bow_correction=True)

'''Rounded Output
[[('<|endoftext|>', 0.0),
  ('I', 6.30),
  ('was', 3.84),
  ('a', 4.68),
  ('mat', 16.34),
  ('ron', 2.11),
  ('in', 1.75),
  ('France', 11.42)],
 [('<|endoftext|>', 0.0),
  ('I', 6.30),
  ('was', 3.84),
  ('a', 4.68),
  ('mat', 21.34),
  ('in', 5.80),
  ('France', 10.69)]]
'''Screenshot from Oh and Schuler showing surprisal values for the partial sentences "I was a matron in" and "I was a mat in" using GPT-2 XL with leading whitespaces and trailing whitespaces.
1428
Reposted by Craig Schmidt
Marco @mcognetta.bsky.social · 11/11/2024
#EMNLP has a nice set of tokenization/subword modeling papers this year. It's a good mix of tokenization algorithms, tokenization evaluation, tokenization-free methods, and subword embedding probing. Lmk if I missed some! Here is a list with links + presentation time (in chronological order).
54816