Sign in

Tiago Pimentel

@tpimentel.bsky.social
2.2K followers 131 following 37 posts

Postdoc at ETH. Formerly, PhD student at the University of Cambridge :)

PostsRepliesMedia
Reposted by Tiago Pimentel
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 13/09/2026
Announcing our first invited speaker at TokShop! We're thrilled to welcome Tiago Pimentel @tpimentel.bsky.social (ETH Zürich) for a talk on: "How much does tokenisation impact language models?" 🧵
1102
Reposted by Tiago Pimentel
Maxime Méloux @maximemeloux.bsky.social · 10/07/2026
I'm very happy to give a spotlight at the Mechanistic Interpretability Workshop @ ICML on our work: "Validating Causal Abstraction Metrics on Simulated Complex Systems" Which metrics actually tell you if an explanation is valid? We built a benchmark to find out. 1/n
melouxm.github.io
Validating Causal Abstraction Metrics on Simulated Complex Systems
142
Reposted by Tiago Pimentel
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 24/06/2026
Effective language identification based on a tokenizer UnigramLM tokenizer already gives probabilities, testing those to identify a language is fast and effective. Whiceh leads me to wonder, can we identify language during training and affect behavior? arxiv.org/abs/2602.17655 @tpimentel.bsky.social
071
Reposted by Tiago Pimentel
Craig Schmidt @craigschmidt.com · 22/05/2026
arxiv.org/abs/2605.22705 arxiv.org/abs/2605.22821 Happy Linear Programming for Tokenization day! I was involved with two separate papers that hit ArXiv yesterday, using LP's to find the vocabulary maximizing compression, depending on the kind of inference you want to use.
arxiv.org
Tokenization with Split Trees
We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference procedure. ToaST greedily splits each pretoken int...
2166
Tiago Pimentel @tpimentel.bsky.social · 22/05/2026
Our new paper reformulates tokenisation as a linear program (LP), which we solve to get SOTA tokenisers 😁 As a bonus, this LP tells us how close to optimal any tokeniser is! Check it out 👇 w/ J. Tempus, @philipwitti.bsky.social, @craigschmidt.com, D. Komm Paper: arxiv.org/abs/2605.22821
14410
Reposted by Tiago Pimentel
Violeta Kastreva @vkastreva.bsky.social · 20/11/2025
Thrilled to share my first paper! 📄 We prove optimal tokenization is NP-hard on bounded alphabets (like bytes)—even unary for direct tokenization! Big thanks @tpimentel.bsky.social, @philipwitti.bsky.social & Dennis Komm for the mentorship! Best birthday gift. 🎂 arxiv.org/abs/2511.15709
071
Tiago Pimentel @tpimentel.bsky.social · 20/11/2025
Tokenisers are a vital part of LLMs, but how hard is it to find an optimal one? 🤔 Considering arbitrarily large alphabets, prior work showed this is NP-hard. But what if we use bytes instead? Or unary strings like a, aa, aaa, ...? In our new paper, we show this is still hard, NP-hard!
Screenshot of paper title: Tokenisation over Bounded Alphabets is Hard.
1173
Reposted by Tiago Pimentel
Gunnar König @gunnark.bsky.social · 07/10/2025
Interested in provable guarantees and fundamental limitations of XAI? Join us at the "Theory of Explainable AI" workshop Dec 2 in Copenhagen! @ellis.eu @euripsconf.bsky.social Speakers: @jessicahullman.bsky.social @doloresromerom.bsky.social @tpimentel.bsky.social Call for Contributions: Oct 15
sites.google.com
Theory of XAI Workshop
Explainable AI (XAI) is now deployed across a wide range of settings, including high-stakes domains in which misleading explanations can cause real harm. For example, explanations are required by law ...
085
Reposted by Tiago Pimentel
Maria Ryskina @mryskina.bsky.social · 04/10/2025
Interested in language models, brains, and concepts? Check out our COLM 2025 🔦 Spotlight paper! (And if you’re at COLM, come hear about it on Tuesday – sessions Spotlight 2 & Poster 2)!
Paper title: Language models align with brain regions that represent concepts across modalities.
Authors:  Maria Ryskina, Greta Tuckute, Alexander Fung, Ashley Malkin, Evelina Fedorenko. 
Affiliations: Maria is affiliated with the Vector Institute for AI, but the work was done at MIT. All other authors are affiliated with MIT. 
Email address: maria.ryskina@vectorinstitute.ai.
1275
Reposted by Tiago Pimentel
Alexander Hoyle @alexanderhoyle.bsky.social · 24/09/2025
Accepted to EMNLP (and more to come 👀)! The camera ready version is now online---very happy with how this turned out arxiv.org/abs/2507.01234
0145
Tiago Pimentel @tpimentel.bsky.social · 01/10/2025
LLMs are trained to mimic a “true” distribution—their reducing cross-entropy then confirms they get closer to this target while training. Do similar models approach this target distribution in similar ways, though? 🤔 Not really! Our new paper studies this, finding 4-convergence phases in training 🧵
Figure showing the four phases of convergence in LM training
1244
Tiago Pimentel @tpimentel.bsky.social · 01/10/2025
Very happy this paper got accepted to NeurIPS 2025 as a Spotlight! 😁 Main takeaway: In mechanistic interpretability, we need assumptions about how DNNs encode concepts in their representations (eg, the linear representation hypothesis). Without them, we can claim any DNN implements any algorithm!
0254
Tiago Pimentel @tpimentel.bsky.social · 31/07/2025
Honoured to receive two (!!) SAC highlights awards at #ACL2025 😁 (Conveniently placed on the same slide!) With the amazing: @philipwitti.bsky.social, @gregorbachmann.bsky.social and @wegotlieb.bsky.social, @cuiding.bsky.social, Giovanni Acampa, @alexwarstadt.bsky.social, @tamaregev.bsky.social
0223
Tiago Pimentel @tpimentel.bsky.social · 27/07/2025
We are presenting this paper at #ACL2025 😁 Find us at poster session 4 (Wednesday morning, 11h~12h30) to learn more about tokenisation bias!
0112
Tiago Pimentel @tpimentel.bsky.social · 27/07/2025
@philipwitti.bsky.social will be presenting our paper "Tokenisation is NP-Complete" at #ACL2025 😁 Come to the language modelling 2 session (Wednesday morning, 9h~10h30) to learn more about how challenging tokenisation can be!
062
Reposted by Tiago Pimentel
Julian Minder @jkminder.bsky.social · 17/07/2025
Causal Abstraction, the theory behind DAS, tests if a network realizes a given algorithm. We show (w/ @denissutter.bsky.social, T. Hofmann, @tpimentel.bsky.social ) that the theory collapses without the linear representation hypothesis—a problem we call the non-linear representation dilemma.
152
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
Mechanistic interpretability often relies on *interventions* to study how DNNs work. Are these interventions enough to guarantee the features we find are not spurious? No!⚠️ In our new paper, we show many mech int methods implicitly rely on the linear representation hypothesis🧵
Paper title "The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?" with the paper's graphical abstract showing how more powerful alignment maps between a DNN and an algorithm allow more complex features to be found and more "accurate" abstractions.
16612
Reposted by Tiago Pimentel
Tom McCoy @rtommccoy.bsky.social · 04/06/2025
The word "laundry" contains both steps of the laundry process: 1. Undry 2. Dry
1262
Reposted by Tiago Pimentel
Musashi Hinck @musashihi.bsky.social · 04/06/2025
Love this! Especially the explicit operationalization of what “bias” they are measuring via specifying the relevant counterfactual. Definitely an approach that more papers talking about effects can incorporate to better clarify what the phenomenon they are studying.
011
Tiago Pimentel @tpimentel.bsky.social · 04/06/2025
If you use LLMs, tokenisation bias probably affects you: * Text generation: tokenisation bias ⇒ length bias 🤯 * Psycholinguistics: tokenisation bias ⇒ systematically biased surprisal estimates 🫠 * Interpretability: tokenisation bias ⇒ biased logits 🤔
070
Tiago Pimentel @tpimentel.bsky.social · 04/06/2025
A string may get 17 times less probability if tokenised as two symbols (e.g., ⟨he, llo⟩) than as one (e.g., ⟨hello⟩)—by an LM trained from scratch in each situation! Our new ACL paper proposes an observational method to estimate this causal effect! Longer thread soon!
Title of paper "Causal Estimation of Tokenisation Bias" and schematic of how we define tokenisation bias, which is the causal effect we are interested in.
1518
Tiago Pimentel @tpimentel.bsky.social · 29/05/2025
If you're finishing your camera-ready for ACL or ICML and want to cite co-first authors more fairly, I just made a simple fix to do this! Just add $^*$ to the authors' names in your bibtex, and the citations should change :) github.com/tpimentelms/...
Inline citations with only first author name, or first two co-first author names.
48322
Reposted by Tiago Pimentel
Ethan Gotlieb Wilcox @wegotlieb.bsky.social · 13/05/2025
⭐🗣️New preprint out: 🗣️⭐ “Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent” with @cuiding.bsky.social , Giovanni Acampa, @tpimentel.bsky.social , @alexwarstadt.bsky.social ,Tamar Regev: arxiv.org/abs/2505.07659
arxiv.org
Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent
This paper argues that the relationship between lexical identity and prosody -- one well-studied parameter of linguistic variation -- can be characterized using information theory. We predict that lan...
1125
Reposted by Tiago Pimentel
Andrew Lampinen @lampinen.bsky.social · 02/05/2025
How do language models generalize from information they learn in-context vs. via finetuning? In arxiv.org/abs/2505.00661 we show that in-context learning can generalize more flexibly, illustrating key differences in the inductive biases of these modes of learning — and ways to improve finetuning. 1/
arxiv.org
47722
Tiago Pimentel @tpimentel.bsky.social · 30/04/2025
Super happy we got this award for our paper on memorisation 😁🎉 congrats to the team and in particular to @pietrolesci.bsky.social, who led the project! Pietro is super smart, creative, hard-working, and on the job market -- you should hire him if you can :)
2100
Reposted by Tiago Pimentel
Andreas Vlachos @andreasvlachos.bsky.social · 29/04/2025
Honoured to receive this award! Tagging @pietrolesci.bsky.social and @tpimentel.bsky.social !
011
Reposted by Tiago Pimentel
Yanai Elazar @yanai.bsky.social · 25/04/2025
💡 New ICLR paper! 💡 "On Linear Representations and Pretraining Data Frequency in Language Models": We provide an explanation for when & why linear representations form in large (or small) language models. Led by @jackmerullo.bsky.social, w/ @nlpnoah.bsky.social & @sarah-nlp.bsky.social
34212
Reposted by Tiago Pimentel
Kyle Mahowald @kmahowald.bsky.social · 21/04/2025
I might be able to hire a postdoc for this fall in computational linguistics at UT Austin. Topics in the general LLM + cognitive space (particularly reasoning, chain of thought, LLMs + code) and LLM + linguistic space. If this could be of interest, feel free to get in touch!
05930
Reposted by Tiago Pimentel
Isabelle Augenstein @iaugenstein.bsky.social · 08/04/2025
I'm so grateful to the British Computing Society & Bloomberg for honouring me with the Karen Spärck Jones Award 🙏 I gave the award lecture on LLMs’ Utilisation of Parametric & Contextual Knowledge at #ECIR2025 today (slides: isabelleaugenstein.github.io/slides/2025_...) www.bcs.org/membership-a...
2376
Reposted by Tiago Pimentel
Benjamin Minixhofer @bminixhofer.bsky.social · 02/04/2025
We created Approximate Likelihood Matching, a principled (and very effective) method for *cross-tokenizer distillation*! With ALM, you can create ensembles of models from different families, convert existing subword-level models to byte-level and a bunch more🧵
Image illustrating that ALM can enable Ensembling, Transfer to Bytes, and general Cross-Tokenizer Distillation.
12514
Reposted by Tiago Pimentel
Conference on Language Modeling @colmweb.org · 20/03/2025
A bit of a mess around the conflict of COLM with the ARR (and to lesser degree ICML) reviews release. We feel this is creating a lot of pressure and uncertainty. So, we are pushing our deadlines: Abstracts due March 22 AoE (+48hr) Full papers due March 28 AoE (+24hr) Plz RT 🙏
33631
Reposted by Tiago Pimentel
Kyle Mahowald @kmahowald.bsky.social · 29/01/2025
LMs need linguistics! New paper, with @futrell.bsky.social, on LMs and linguistics that conveys our excitement about what the present moment means for linguistics and what linguistics can do for LMs. Paper: arxiv.org/abs/2501.17047. 🧵below.
311233
Reposted by Tiago Pimentel
The First Workshop on Large Language Model Memorization (L2M2) @l2m2workshop.bsky.social · 27/01/2025
📢 The First Workshop on Large Language Model Memorization (L2M2) will be co-located with @aclmeeting.bsky.social in Vienna 🎉 💡 L2M2 brings together researchers to explore memorization from multiple angles. Whether it's text-only LLMs or Vision-language models, we want to hear from you! 🌍
1103
Tiago Pimentel @tpimentel.bsky.social · 20/12/2024
BPE is a greedy method to find a tokeniser which maximises compression! Why don't we try to find properly optimal tokenisers instead? Well, it seems this is a pretty difficult—in fact, NP-complete—problem!🤯 New paper + @philipwitti.bsky.social @gregorbachmann.bsky.social :) arxiv.org/abs/2412.15210
arxiv.org
Tokenisation is NP-Complete
In this work, we prove the NP-completeness of two variants of tokenisation, defined as the problem of compressing a dataset to at most $δ$ symbols by either finding a vocabulary directly (direct token...
1458
Reposted by Tiago Pimentel
Kanishka Misra @kanishka.bsky.social · 25/11/2024
There's a known bug in how we compute "word" probabilities with subword-based LMs that mark beginnings of words -- as pointed out by Byung-doh Oh and Will Schuler, & @tpimentel.bsky.social and Clara Meister I'm pleased to announce that minicons now includes a fix which runs batch-wise!
Code: from minicons import scorer

lm = scorer.IncrementalLMScorer("gpt2-xl", "cuda:0")

stimuli = ["I was a matron in France", "I was a mat in France"]

# old way, no correction
# P.S. gpt2 does not automatically add a bos token at the beginning...
lm.token_score(stimuli, bos_token=True, surprisal=True, base_two=True, bow_correction=False)

'''Rounded Output
[[('<|endoftext|>', 0.0),
  ('I', 5.85),
  ('was', 4.28),
  ('a', 4.67),
  ('mat', 16.34),
  ('ron', 1.74),
  ('in', 2.12),
  ('France', 11.43)],
 [('<|endoftext|>', 0.0),
  ('I', 5.85),
  ('was', 4.28),
  ('a', 4.67),
  ('mat', 16.34),
  ('in', 10.78),
  ('France', 10.71)]]
'''

# the new way! notice the surprisal of "mat" in both cases
lm.token_score(stimuli, bos_token=True, surprisal=True, base_two=True, bow_correction=True)

'''Rounded Output
[[('<|endoftext|>', 0.0),
  ('I', 6.30),
  ('was', 3.84),
  ('a', 4.68),
  ('mat', 16.34),
  ('ron', 2.11),
  ('in', 1.75),
  ('France', 11.42)],
 [('<|endoftext|>', 0.0),
  ('I', 6.30),
  ('was', 3.84),
  ('a', 4.68),
  ('mat', 21.34),
  ('in', 5.80),
  ('France', 10.69)]]
'''Screenshot from Oh and Schuler showing surprisal values for the partial sentences "I was a matron in" and "I was a mat in" using GPT-2 XL with leading whitespaces and trailing whitespaces.
1428
Tiago Pimentel @tpimentel.bsky.social · 20/11/2024
Hey :) I'm looking for 3 emergency reviewers for ARR submissions🚨📷 they are all in LM interpretability! Should be submitted within the next 36 hours 🙃 If you are interested, please DM or email me! #NLProc #NLP
2810
Tiago Pimentel @tpimentel.bsky.social · 08/12/2023
Are you interested in word lengths and natural language’s efficiency? If yes, check out our new #EMNLP2023 paper! It has everything you need: drama, suspense, a new derivation of Zipf’s law, an update to Piantadosi et al’s classic word length paper, transformers... 😄 arxiv.org/abs/2312.03897
Screenshot of paper's title. Paper title is: "Revisiting the Optimality of Word Lengths"
1254
Reposted by Tiago Pimentel
Cory Shain @coryshain.bsky.social · 19/10/2023
👋Hi #bsky! Just wanted to let everyone know I'll be arriving at #Stanford in fall of 2024 and I'm looking for awesome people to help me figure out language. I'm new here my network is small, so RTs would be great! 🧵👇
15139