Sign in

Tim Vieira

@xtimv.bsky.social
255 followers 227 following 5 posts

timvieira.github.io/blog

PostsRepliesMedia
Reposted by Tim Vieira
Marco @mcognetta.bsky.social · 15/08/2025
My paper "Tokenization as Finite-State Transduction" was accepted to Computational Linguistics. This was my final PhD degree requirement :) The goal was to unify the major tokenization algorithms under a finite-state automaton framework. For example, by encoding a BPE tokenizer as a transducer.
Merge transducers for a BPE tokenizer.
2306
Reposted by Tim Vieira
Ben Lipkin @benlipkin.bsky.social · 13/05/2025
Many LM applications may be formulated as text generation conditional on some (Boolean) constraint. Generate a… - Python program that passes a test suite. - PDDL plan that satisfies a goal. - CoT trajectory that yields a positive reward. The list goes on… How can we efficiently satisfy these? 🧵👇
2136
Reposted by Tim Vieira
Afra Amini @afraamn.bsky.social · 06/05/2025
Current KL estimation practices in RLHF can generate high variance and even negative values! We propose a provably better estimator that only takes a few lines of code to implement.🧵👇 w/ @xtimv.bsky.social and Ryan Cotterell code: arxiv.org/pdf/2504.10637 paper: github.com/rycolab/kl-rb
173
Reposted by Tim Vieira
João Loula @joaoloula.bsky.social · 25/04/2025
#ICLR2025 Oral How can we control LMs using diverse signals such as static analyses, test cases, and simulations? In our paper “Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo” (w/ @benlipkin.bsky.social, @alexlew.bsky.social, @xtimv.bsky.social) we:
176
Reposted by Tim Vieira
Ben Lipkin @benlipkin.bsky.social · 10/04/2025
New preprint on controlled generation from LMs! I'll be presenting at NENLP tomorrow 12:50-2:00pm Longer thread coming soon :)
1209
Reposted by Tim Vieira
Marco @mcognetta.bsky.social · 10/02/2025
Tokenization is an often-overlooked aspect of modern #NLP, but it’s experiencing a resurgence — thanks in large part to @karpathy.bsky.social and his classic tweet: x.com/karpathy/sta... Come hang out with us and let's fix these problems!
071
Reposted by Tim Vieira
Marco @mcognetta.bsky.social · 10/02/2025
Today we are launching a server dedicated to Tokenization research! Come join us! discord.gg/CDJhnSvU
discord.gg
Join the Token ##ization Discord Server!
Check out the Token ##ization community on Discord - hang out with 24 other members and enjoy free voice and text chat.
3186
Reposted by Tim Vieira
Alex Lew @alexlew.bsky.social · 10/02/2025
@xtimv.bsky.social and I were just discussing this interesting comment in the DeepSeek paper introducing GRPO: a different way of setting up the KL loss. It's a little hard to reason about what this does to the objective. 1/
Also note that, instead of adding KL penalty in the reward, GRPO regularizes by directly adding the KL divergence between the trained policy and the reference policy to the loss, avoiding complicating the calculation of the advantage.
3509
Reposted by Tim Vieira
Craig Schmidt @craigschmidt.com · 20/12/2024
I made a starter pack for people in NLP working in the area of tokenization. Let me know if you'd like to be added go.bsky.app/8P9ftjL
0125
Reposted by Tim Vieira
Maria Antoniak @mariaa.bsky.social · 31/12/2024
It's ready! 💫 A new blog post in which I list of all the tools and apps I've been using for work, plus all my opinions about them. maria-antoniak.github.io/2024/12/30/o... Featuring @kagi.com, @warp.dev, @paperpile.bsky.social, @are.na, Fantastical, @obsidian.md, Claude, and more.
3621525
Reposted by Tim Vieira
Brandon Amos @bdamos.bsky.social · 05/12/2024
hi everyone!! let's try this optimal transport again 🙃
232931
Reposted by Tim Vieira
Ned Batchelder @nedbat.com · 21/11/2024
Also: you can also use variables (or expressions?!) for the formatting information! #Python is cool... More details and explanation at fstring.help
Showing using variables for the character and width specification in the format.  Full text at https://gist.github.com/nedbat/db70d4cc6f16e88a33889ec113dcbe4d
2366
Reposted by Tim Vieira
Alex Lew @alexlew.bsky.social · 21/11/2024
Surprisal of title beginning with 'O'? 3.22 Surprisal of 'o' following 'Treatment '? 0.11 Surprisal that title includes surprisal of each title character? Priceless [...I did not know titles could do this]
Screenshot of the title of the paper "On the Proper Treatment of Tokenization in Psycholinguistics." Over each letter, the authors have plotted the surprisal of the letter (-log p(this letter | context)).
2102
Reposted by Tim Vieira
Shauli Ravfogel @shauli.bsky.social · 12/11/2024
Happy to share our work "Counterfactual Generation from Language Models" with @AnejSvete, @vesteinns, and Ryan Cotterell! We tackle generating true counterfactual strings from LMs after interventions and introduce a simple algorithm for it. (1/7) arxiv.org/pdf/2411.07180
2143
Reposted by Tim Vieira
Alex Thiery @alexxthiery.bsky.social · 20/11/2024
Variational approximation with Gaussian mixtures is looking cute! So here it's just gradient descent on K(q||p) for optimising the mixtures means & covariances & weights... @lacerbi.bsky.social
2337
Reposted by Tim Vieira
Alex Thiery @alexxthiery.bsky.social · 20/11/2024
Gaussian approximation of a target distribution: mean-field versus full-covariance! Below shows a simple gradient descent on KL(q||p)
3625