Sign in

Tiago Pimentel

@tpimentel.bsky.social
2.2K followers 131 following 37 posts

Postdoc at ETH. Formerly, PhD student at the University of Cambridge :)

PostsRepliesMedia
Reposted by Tiago Pimentel
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 13/09/2026
Announcing our first invited speaker at TokShop! We're thrilled to welcome Tiago Pimentel @tpimentel.bsky.social (ETH Zürich) for a talk on: "How much does tokenisation impact language models?" 🧵
1102
Reposted by Tiago Pimentel
Maxime Méloux @maximemeloux.bsky.social · 10/07/2026
I'm very happy to give a spotlight at the Mechanistic Interpretability Workshop @ ICML on our work: "Validating Causal Abstraction Metrics on Simulated Complex Systems" Which metrics actually tell you if an explanation is valid? We built a benchmark to find out. 1/n
melouxm.github.io
Validating Causal Abstraction Metrics on Simulated Complex Systems
142
Reposted by Tiago Pimentel
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 24/06/2026
Effective language identification based on a tokenizer UnigramLM tokenizer already gives probabilities, testing those to identify a language is fast and effective. Whiceh leads me to wonder, can we identify language during training and affect behavior? arxiv.org/abs/2602.17655 @tpimentel.bsky.social
071
Tiago Pimentel @tpimentel.bsky.social · 22/05/2026
@craigschmidt.com has a second paper using LPs for tokenisation coming out today as well! Check it out: bsky.app/profile/crai...
030
Reposted by Tiago Pimentel
Craig Schmidt @craigschmidt.com · 22/05/2026
arxiv.org/abs/2605.22705 arxiv.org/abs/2605.22821 Happy Linear Programming for Tokenization day! I was involved with two separate papers that hit ArXiv yesterday, using LP's to find the vocabulary maximizing compression, depending on the kind of inference you want to use.
arxiv.org
Tokenization with Split Trees
We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference procedure. ToaST greedily splits each pretoken int...
2166
Tiago Pimentel @tpimentel.bsky.social · 22/05/2026
Our new paper reformulates tokenisation as a linear program (LP), which we solve to get SOTA tokenisers 😁 As a bonus, this LP tells us how close to optimal any tokeniser is! Check it out 👇 w/ J. Tempus, @philipwitti.bsky.social, @craigschmidt.com, D. Komm Paper: arxiv.org/abs/2605.22821
14410
Reposted by Tiago Pimentel
Violeta Kastreva @vkastreva.bsky.social · 20/11/2025
Thrilled to share my first paper! 📄 We prove optimal tokenization is NP-hard on bounded alphabets (like bytes)—even unary for direct tokenization! Big thanks @tpimentel.bsky.social, @philipwitti.bsky.social & Dennis Komm for the mentorship! Best birthday gift. 🎂 arxiv.org/abs/2511.15709
071
Tiago Pimentel @tpimentel.bsky.social · 20/11/2025
This was joint work with @vkastreva.bsky.social, @philipwitti.bsky.social, D. Komm! Violeta is a super smart student, who is definitely gonna do lots more interesting work :) It's her first paper, and it's also her birthday today 🥳 so follow her if you like this! Paper: arxiv.org/abs/2511.15709
arxiv.org
Tokenisation over Bounded Alphabets is Hard
Recent works have shown that tokenisation is NP-complete. However, these works assume tokenisation is applied to inputs with unboundedly large alphabets -- an unrealistic assumption, given that in pra...
130
Tiago Pimentel @tpimentel.bsky.social · 20/11/2025
More precisely, we show that: (i) for binary alphabets, not only finding an optimal tokeniser is NP-hard, but also finding arbitrarily good approximations; (ii) for unary alphabets, finding an optimal direct tokeniser is NP-hard!
120
Tiago Pimentel @tpimentel.bsky.social · 20/11/2025
Tokenisers are a vital part of LLMs, but how hard is it to find an optimal one? 🤔 Considering arbitrarily large alphabets, prior work showed this is NP-hard. But what if we use bytes instead? Or unary strings like a, aa, aaa, ...? In our new paper, we show this is still hard, NP-hard!
Screenshot of paper title: Tokenisation over Bounded Alphabets is Hard.
1173
Reposted by Tiago Pimentel
Gunnar König @gunnark.bsky.social · 07/10/2025
Interested in provable guarantees and fundamental limitations of XAI? Join us at the "Theory of Explainable AI" workshop Dec 2 in Copenhagen! @ellis.eu @euripsconf.bsky.social Speakers: @jessicahullman.bsky.social @doloresromerom.bsky.social @tpimentel.bsky.social Call for Contributions: Oct 15
sites.google.com
Theory of XAI Workshop
Explainable AI (XAI) is now deployed across a wide range of settings, including high-stakes domains in which misleading explanations can cause real harm. For example, explanations are required by law ...
085
Reposted by Tiago Pimentel
Maria Ryskina @mryskina.bsky.social · 04/10/2025
Interested in language models, brains, and concepts? Check out our COLM 2025 🔦 Spotlight paper! (And if you’re at COLM, come hear about it on Tuesday – sessions Spotlight 2 & Poster 2)!
Paper title: Language models align with brain regions that represent concepts across modalities.
Authors:  Maria Ryskina, Greta Tuckute, Alexander Fung, Ashley Malkin, Evelina Fedorenko. 
Affiliations: Maria is affiliated with the Vector Institute for AI, but the work was done at MIT. All other authors are affiliated with MIT. 
Email address: maria.ryskina@vectorinstitute.ai.
1275
Reposted by Tiago Pimentel
Alexander Hoyle @alexanderhoyle.bsky.social · 24/09/2025
Accepted to EMNLP (and more to come 👀)! The camera ready version is now online---very happy with how this turned out arxiv.org/abs/2507.01234
0145
Tiago Pimentel @tpimentel.bsky.social · 01/10/2025
This project was done with Finlay and @kmahowald.bsky.social, and it is the outcome of Finlay's Bachelor's thesis! Catch him presenting it in #EMNLP2025 :) Paper: arxiv.org/abs/2509.26643 Code: github.com/Tr1ple-F/con...
arxiv.org
Convergence and Divergence of Language Models under Different Random Seeds
In this paper, we investigate the convergence of language models (LMs) trained under different random seeds, measuring convergence as the expected per-token Kullback--Leibler (KL) divergence across se...
040
Tiago Pimentel @tpimentel.bsky.social · 01/10/2025
See our paper for more: we have analyses on other models, downstream tasks, and considering only subsets of tokens (e.g., only tokens with a certain part-of-speech)!
100
Tiago Pimentel @tpimentel.bsky.social · 01/10/2025
This means that: (1) LMs can get less similar to each other, even while they all get closer to the true distribution; and (2) larger models reconverge faster, while small ones may never reconverge.
100
Tiago Pimentel @tpimentel.bsky.social · 01/10/2025
* A sharp-divergence phase, where models diverge as they start using context. * A slow-reconvergence phase, where predictions slowly become more similar again (especially in larger models).
100
Tiago Pimentel @tpimentel.bsky.social · 01/10/2025
Surprisingly, convergence isn’t monotonic. Instead, we find four convergence phases across model training. * A uniform phase, where all seeds output nearly-uniform distributions. * A sharp-convergence phase, where models align, largely due to unigram frequency learning.
100
Tiago Pimentel @tpimentel.bsky.social · 01/10/2025
In this paper, we define convergence as the similarity between outputs of LMs trained under different seeds, where similarity is measured as a per-token KL divergence. This lets us track whether models trained under identical settings, but different seeds, behave the same.
100
Tiago Pimentel @tpimentel.bsky.social · 01/10/2025
LLMs are trained to mimic a “true” distribution—their reducing cross-entropy then confirms they get closer to this target while training. Do similar models approach this target distribution in similar ways, though? 🤔 Not really! Our new paper studies this, finding 4-convergence phases in training 🧵
Figure showing the four phases of convergence in LM training
1244
Tiago Pimentel @tpimentel.bsky.social · 01/10/2025
Very happy this paper got accepted to NeurIPS 2025 as a Spotlight! 😁 Main takeaway: In mechanistic interpretability, we need assumptions about how DNNs encode concepts in their representations (eg, the linear representation hypothesis). Without them, we can claim any DNN implements any algorithm!
0254
Tiago Pimentel @tpimentel.bsky.social · 31/07/2025
Honoured to receive two (!!) SAC highlights awards at #ACL2025 😁 (Conveniently placed on the same slide!) With the amazing: @philipwitti.bsky.social, @gregorbachmann.bsky.social and @wegotlieb.bsky.social, @cuiding.bsky.social, Giovanni Acampa, @alexwarstadt.bsky.social, @tamaregev.bsky.social
0223
Tiago Pimentel @tpimentel.bsky.social · 27/07/2025
We are presenting this paper at #ACL2025 😁 Find us at poster session 4 (Wednesday morning, 11h~12h30) to learn more about tokenisation bias!
0112
Tiago Pimentel @tpimentel.bsky.social · 27/07/2025
@philipwitti.bsky.social will be presenting our paper "Tokenisation is NP-Complete" at #ACL2025 😁 Come to the language modelling 2 session (Wednesday morning, 9h~10h30) to learn more about how challenging tokenisation can be!
062
Reposted by Tiago Pimentel
Julian Minder @jkminder.bsky.social · 17/07/2025
Causal Abstraction, the theory behind DAS, tests if a network realizes a given algorithm. We show (w/ @denissutter.bsky.social, T. Hofmann, @tpimentel.bsky.social ) that the theory collapses without the linear representation hypothesis—a problem we call the non-linear representation dilemma.
152
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
Importantly, despite these results, we still believe causal abstraction is one of the best frameworks available for mech interpretability. Going forward, we should try to better understand how it is impacted by assumptions about how DNNs encode information. Longer🧵soon by @denissutter.bsky.social
040
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
Overall, our results show that causal abstraction (and interventions) is not a silver bullet, as it relies on assumptions about how features are encoded in the DNNs. We then connect our results to the linear representation hypothesis and to older debates in the probing literature.
120
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
We show—both theoretically (under reasonable assumptions) and empirically (on real-world models)—that, if we allow variables to be encoded in arbitrarily complex subspaces of the DNN’s representations, any algorithm can be mapped to any model.
110
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
Causal abstraction identifies this correspondence by finding subspaces in the DNN's hidden states which encode the algorithm’s hidden variables. Given such a map, we say the DNN implements the algorithm if the two behave identically under interventions.
100
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
In this new paper, w/ @denissutter.bsky.social , @jkminder.bsky.social, and T.Hofmann, we study *causal abstraction*, a formal specification of when a deep neural network (DNN) implements an algorithm. This is the framework behind, e.g., distributed alignment search. Paper: arxiv.org/abs/2507.08802
arxiv.org
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level ...
131
Tiago Pimentel @tpimentel.bsky.social · 14/07/2025
Mechanistic interpretability often relies on *interventions* to study how DNNs work. Are these interventions enough to guarantee the features we find are not spurious? No!⚠️ In our new paper, we show many mech int methods implicitly rely on the linear representation hypothesis🧵
Paper title "The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?" with the paper's graphical abstract showing how more powerful alignment maps between a DNN and an algorithm allow more complex features to be found and more "accurate" abstractions.
16612
Reposted by Tiago Pimentel
Tom McCoy @rtommccoy.bsky.social · 04/06/2025
The word "laundry" contains both steps of the laundry process: 1. Undry 2. Dry
1262
Reposted by Tiago Pimentel
Musashi Hinck @musashihi.bsky.social · 04/06/2025
Love this! Especially the explicit operationalization of what “bias” they are measuring via specifying the relevant counterfactual. Definitely an approach that more papers talking about effects can incorporate to better clarify what the phenomenon they are studying.
011
Tiago Pimentel @tpimentel.bsky.social · 04/06/2025
If you use LLMs, tokenisation bias probably affects you: * Text generation: tokenisation bias ⇒ length bias 🤯 * Psycholinguistics: tokenisation bias ⇒ systematically biased surprisal estimates 🫠 * Interpretability: tokenisation bias ⇒ biased logits 🤔
070
Tiago Pimentel @tpimentel.bsky.social · 04/06/2025
Led by @pietrolesci.bsky.social and with Clara Meister, Thomas Hofmann, @andreasvlachos.bsky.social :) Paper: arxiv.org/abs/2506.03149 Code: github.com/pietrolesci/...
arxiv.org
Causal Estimation of Tokenisation Bias
Modern language models are typically trained over subword sequences, but ultimately define probabilities over character-strings. Ideally, the choice of the tokeniser -- which maps character-strings to...
010
Tiago Pimentel @tpimentel.bsky.social · 04/06/2025
A string may get 17 times less probability if tokenised as two symbols (e.g., ⟨he, llo⟩) than as one (e.g., ⟨hello⟩)—by an LM trained from scratch in each situation! Our new ACL paper proposes an observational method to estimate this causal effect! Longer thread soon!
Title of paper "Causal Estimation of Tokenisation Bias" and schematic of how we define tokenisation bias, which is the causal effect we are interested in.
1518
Tiago Pimentel @tpimentel.bsky.social · 29/05/2025
I think it's a reasonable change, and it doesn't change the template style, so I'd say yes. There is also already the command `\citep*` to cite all authors in a paper, so citing only the first two should also be ok? I created a pull request this morning to add it to the official template :)
020
Tiago Pimentel @tpimentel.bsky.social · 29/05/2025
I created a pull request earlier today. So hopefully they will approve and merge it soon-ish? :)
020
Tiago Pimentel @tpimentel.bsky.social · 29/05/2025
The papers cited above are: * Pareto Probing, by @nsaphra.bsky.social and me :) * Neural populations differ in the size of their temporal receptive windows, by @tamaregev.bsky.social and @coltoncasto.bsky.social and * PolyPythias, by van der Wal and @pietrolesci.bsky.social
040
Tiago Pimentel @tpimentel.bsky.social · 29/05/2025
For ACL, just replace your acl_natbib.bst file with the one in this repo 😁; for ICML, you need to replace `FUNCTION {format.lab.names}` in icml2025.bst with the code from lines 1655 to 1747 (untested, but should work 🙃)
130
Tiago Pimentel @tpimentel.bsky.social · 29/05/2025
If you're finishing your camera-ready for ACL or ICML and want to cite co-first authors more fairly, I just made a simple fix to do this! Just add $^*$ to the authors' names in your bibtex, and the citations should change :) github.com/tpimentelms/...
Inline citations with only first author name, or first two co-first author names.
48322
Reposted by Tiago Pimentel
Ethan Gotlieb Wilcox @wegotlieb.bsky.social · 13/05/2025
⭐🗣️New preprint out: 🗣️⭐ “Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent” with @cuiding.bsky.social , Giovanni Acampa, @tpimentel.bsky.social , @alexwarstadt.bsky.social ,Tamar Regev: arxiv.org/abs/2505.07659
arxiv.org
Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent
This paper argues that the relationship between lexical identity and prosody -- one well-studied parameter of linguistic variation -- can be characterized using information theory. We predict that lan...
1125
Reposted by Tiago Pimentel
Andrew Lampinen @lampinen.bsky.social · 02/05/2025
How do language models generalize from information they learn in-context vs. via finetuning? In arxiv.org/abs/2505.00661 we show that in-context learning can generalize more flexibly, illustrating key differences in the inductive biases of these modes of learning — and ways to improve finetuning. 1/
arxiv.org
47722
Tiago Pimentel @tpimentel.bsky.social · 30/04/2025
In this paper, we borrowed from the econometrics literature and proposed a method to estimate memorisation using only observational data! The paper is here, if you're curious: aclanthology.org/2024.acl-lon...
aclanthology.org
Causal Estimation of Memorisation Profiles
Pietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos, Tiago Pimentel. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
020
Tiago Pimentel @tpimentel.bsky.social · 30/04/2025
Super happy we got this award for our paper on memorisation 😁🎉 congrats to the team and in particular to @pietrolesci.bsky.social, who led the project! Pietro is super smart, creative, hard-working, and on the job market -- you should hire him if you can :)
2100
Reposted by Tiago Pimentel
Andreas Vlachos @andreasvlachos.bsky.social · 29/04/2025
Honoured to receive this award! Tagging @pietrolesci.bsky.social and @tpimentel.bsky.social !
011
Reposted by Tiago Pimentel
Yanai Elazar @yanai.bsky.social · 25/04/2025
💡 New ICLR paper! 💡 "On Linear Representations and Pretraining Data Frequency in Language Models": We provide an explanation for when & why linear representations form in large (or small) language models. Led by @jackmerullo.bsky.social, w/ @nlpnoah.bsky.social & @sarah-nlp.bsky.social
34212