Sign in

Marco

@mcognetta.bsky.social
3.8K followers 2.1K following 1.5K posts

Language and keyboard stuff at Google. I like computers and Korean and computers-and-Korean and high school CS education. Georgia Tech → 연세대학교 → 東京工業大学. Regrettably no longer based in Tokyo :/ theoreticallygoodwithcomputers.com

PostsRepliesMedia
Reposted by Marco
Marco @mcognetta.bsky.social · 16h
Tokenizer research lags behind other areas because it is hard to validate changes and a lack of good predictive metrics/scaling laws. It is also just hard to compare across tokenizers. Let's have a community benchmark where everything but the tokenizer is fixed!
alphaxiv.org
A Call for an Open Tokenizer Benchmark
Tokenizer research lags behind other areas of the language modeling stack due in part to the difficulty of: 1) comparing across models with different tokenizers, 2) quickly and accurately measuring...
132
Marco @mcognetta.bsky.social · 16h
Also, check out our recent tokenization survey!
010
Marco @mcognetta.bsky.social · 16h
Tokenizer research lags behind other areas because it is hard to validate changes and a lack of good predictive metrics/scaling laws. It is also just hard to compare across tokenizers. Let's have a community benchmark where everything but the tokenizer is fixed!
alphaxiv.org
A Call for an Open Tokenizer Benchmark
Tokenizer research lags behind other areas of the language modeling stack due in part to the difficulty of: 1) comparing across models with different tokenizers, 2) quickly and accurately measuring...
132
Marco @mcognetta.bsky.social · 16h
I'm heading to @colmweb.org. Come find me to chat about tokenizers and Korean language modeling! I'll be presenting a poster at @tokshop.bsky.social about the need for a community tokenization benchmark a la the NanoGPT speedrun. It is meant to be a discussion, come share your ideas! #colm26
160
Marco @mcognetta.bsky.social · 05/10/2026
Thanks! I have a Rancilio Silva, which I like except that it is a single boiler. My first machine was a Breville Barista Pro, which has a very weak wand, but is otherwise a good starting point. The Breville is much wider due to the built in grinder, so you would have to take that into account.
010
Marco @mcognetta.bsky.social · 05/10/2026
gm
120
Marco @mcognetta.bsky.social · 05/10/2026
I was proud of this one. bsky.app/profile/mcog...
2696
Marco @mcognetta.bsky.social · 04/10/2026
031
Marco @mcognetta.bsky.social · 04/10/2026
자몽(自夢)
091
Reposted by Marco
Quantian @quantian.bsky.social · 03/10/2026
For what it’s worth, arguing LLMs can’t do addition now is flat earth-level science denialism. Here’s Gemma 4b, a non-reasoning model that could only theoretically memorize binary operations up to 65535 x 65535, oneshotting a eight digit addition and hex conversion *on my phone* with no tool calls
1843239
Marco @mcognetta.bsky.social · 02/10/2026
me, debugging code
19011
Marco @mcognetta.bsky.social · 01/10/2026
gm
020
Reposted by Marco
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 01/10/2026
🎉 Our last invited speaker announcement — TokShop is next week! Welcome Artidoro Pagnoni @artidoro.bsky.social (Meta Superintelligence, FAIR), lead author of the Byte Latent Transformer (BLT) and co-creator of QLoRA. Best paper award & orals at ACL/NeurIPS.
051
Reposted by Marco
Jindřich Libovický @jlibovicky.bsky.social · 01/10/2026
Everything you always wanted to know about tokenization but were afraid to ask! 🧩 @mcognetta.bsky.social l and @uvp.bsky.social nagged 30+ researchers (myself included) until we ended up with the most comprehensive survey on tokenization in NLP: www.alphaxiv.org/abs/2609.tok...
alphaxiv.org
Tokenization: A Survey for Modern NLP
Tokenization is presented as a core language-model design choice that shapes sequence length, computational cost, multilingual equity, evaluation, and security—not merely as preprocessing. The...
191
Reposted by Marco
Sung Kim @sungkim.bsky.social · 01/10/2026
Tokenization: A Survey for Modern NLP Over the past ~8 months, 32 (!) tokenizer researchers put together the comprehensive survey of the field. www.alphaxiv.org/abs/2609.tok...
0203
Reposted by Marco
Marco @mcognetta.bsky.social · 30/09/2026
🚨 [Token][ization] Paper Alert 🚨 Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field. Check it out!
112935
Reposted by Marco
Federico Pianzola @fpianz.eurosky.social · 30/09/2026
Extremely important for multilingual research!
0155
Marco @mcognetta.bsky.social · 30/09/2026
Here is a direct link to the PDF hosted on alpharxiv. www.alphaxiv.org/pdf/2609.tok... If that doesn't work, our coauthor @uvp.bsky.social hosted a copy here:
github.com
website/Tokenization A Survey for Modern NLP.pdf at main · yuvalpinter/website
Contribute to yuvalpinter/website development by creating an account on GitHub.
020
Reposted by Marco
Craig Schmidt @craigschmidt.com · 30/09/2026
Congratulations to Marco and all the other co-authors on this groundbreaking 154-page survey. If you're at all interested in tokenization, give it a read.
051
Reposted by Marco
Eugene Jang @ COLM @eugeneonnlp.bsky.social · 30/09/2026
A huge collaborative effort to describe the deceptively deep rabbit hole that is tokenization! Hopefully it contains many answers to "has anyone tried doing ~" questions and inspire new approaches. Thank you to @mcognetta.bsky.social for leading this ambitious project!
041
Reposted by Marco
Xiulin Yang @xiulinyang.bsky.social · 30/09/2026
It’s all about tokens! 👀
051
Reposted by Marco
Yuval Pinter @ COLM @uvp.bsky.social · 30/09/2026
Tokens!
051
Marco @mcognetta.bsky.social · 30/09/2026
If you are interested in tokenizer research, come join our discord!
140
Marco @mcognetta.bsky.social · 30/09/2026
Check out the paper on @alphaxiv.org ! www.alphaxiv.org/abs/2609.tok...
alphaxiv.org
Tokenization: A Survey for Modern NLP
While modern language models take raw text as their input and produce raw text as output, they do not operate over text directly. Hidden in the very first step of language model pipelines is...
2193
Marco @mcognetta.bsky.social · 30/09/2026
As @karpathy.bsky.social said: tokenization is the root of all suffering. Come learn the good stuff.
1103
Marco @mcognetta.bsky.social · 30/09/2026
We also cover some topics that are closely adjacent to tokenization, such as constrained generation, token healing, and tokenizer security concerns.
180
Marco @mcognetta.bsky.social · 30/09/2026
Every section comes with questions for prospective researchers and practitioners!
170
Marco @mcognetta.bsky.social · 30/09/2026
We cover every aspect of tokenization: algorithms, evaluations, multilinguality, encodings, theory, etc. We even cover what you might want to replace tokenizers with (e.g., latent or visual tokenization).
1121
Marco @mcognetta.bsky.social · 30/09/2026
🚨 [Token][ization] Paper Alert 🚨 Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field. Check it out!
112935
Marco @mcognetta.bsky.social · 30/09/2026
I'm getting a taste of UK internet (my international plan routes through the UK?) and wow how do you guys live like this? Totally adversarial towards the user.
020
Marco @mcognetta.bsky.social · 29/09/2026
Worth noting that the person who did this is... also a member of the Polish national baseball team.
131
Reposted by Marco
mr. TIM @timkellogg.me · 29/09/2026
NanoGPT pretraining runs now take 39.9 seconds for a 124M model(!!) now, 124M is *tiny* so it might not seem relevant. But much of the gains in LLM pretraining are data quality and a great way to find quality data is to train on it & measure the lift
Larry Dial Y @classiclarryd
x.com
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak, obliterating the prior record of 67.6s!
This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement.
If a flop is low value on a particular step, skip it.
Specifically:
-(~8s) Sampled softmax. If a token doesn't appear in a batch, skip its Im_head fwd/bwd some fraction of the time.
-Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch.
Set betal to zero to enable this. Beta2 is applied retroactively when the row is later used.
-Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2.
-Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step.
3928
Marco @mcognetta.bsky.social · 29/09/2026
What till you learn about how quickly planes can travel over even rocky surfaces.
040
Marco @mcognetta.bsky.social · 27/09/2026
I found where Claude lives
020
Marco @mcognetta.bsky.social · 26/09/2026
Yeah Claude is clearly a better name. I wish Google had stuck with Bard instead of Gemini. I always thought that was a cute name.
1130
Marco @mcognetta.bsky.social · 26/09/2026
wow
020
Reposted by Marco
Marco @mcognetta.bsky.social · 21/12/2025
Everywhere I look I see his face.
131
Marco @mcognetta.bsky.social · 26/09/2026
Mire donde mire, veo su cara.
010
Marco @mcognetta.bsky.social · 26/09/2026
1365
Reposted by Marco
eva (^_^)/ @eva.computer · 24/09/2026
imagine being trapped in a Samsung fridge panel forever
31033
Marco @mcognetta.bsky.social · 24/09/2026
the inventor of anime, no less
030
Marco @mcognetta.bsky.social · 23/09/2026
Riff Raff x Guy Fieri
010
Marco @mcognetta.bsky.social · 23/09/2026
slay
010
Marco @mcognetta.bsky.social · 22/09/2026
090
Marco @mcognetta.bsky.social · 22/09/2026
I think there's a song about this
050
Reposted by Marco
Sung Kim @sungkim.bsky.social · 21/09/2026
Hugging Face's tokenizers v1 They focused on all languages, multi-thread scaling, minimal package size and memory usage. huggingface-tokenizers-v1.static.hf.space/index.html
0141
Marco @mcognetta.bsky.social · 21/09/2026
Also the rules are kind of complicated. Like she starts talking about edge cases and all sorts of examples. This is a masterpiece.
141
Marco @mcognetta.bsky.social · 21/09/2026
en.wikipedia.org
The Name Game - Wikipedia
000
Marco @mcognetta.bsky.social · 21/09/2026
TIL the kids rhyming game like "anna banana fee fi fo fanna" came from an actual song and the song is literally just the singer explaining the rules of this game in a slightly lyrical form. I feel like some other games could benefit from this.
youtube.com
THE NAME GAME SHIRLEY ELLIS
YouTube video by ourFAMILYvideoLOG
231
Marco @mcognetta.bsky.social · 21/09/2026
I'd be a customer
010