Sign in

Marco

@mcognetta.bsky.social
3.8K followers 2.1K following 1.4K posts

Language and keyboard stuff at Google. I like computers and Korean and computers-and-Korean and high school CS education. Georgia Tech → 연세대학교 → 東京工業大学. Regrettably no longer based in Tokyo :/ theoreticallygoodwithcomputers.com

PostsRepliesMedia
Reposted by Marco
Federico Pianzola @fpianz.eurosky.social · 2h
Extremely important for multilingual research!
052
Reposted by Marco
Craig Schmidt @craigschmidt.com · 1h
Congratulations to Marco and all the other co-authors on this groundbreaking 154-page survey. If you're at all interested in tokenization, give it a read.
041
Reposted by Marco
Eugene Jang @eugeneonnlp.bsky.social · 3h
A huge collaborative effort to describe the deceptively deep rabbit hole that is tokenization! Hopefully it contains many answers to "has anyone tried doing ~" questions and inspire new approaches. Thank you to @mcognetta.bsky.social for leading this ambitious project!
041
Reposted by Marco
Xiulin Yang @xiulinyang.bsky.social · 3h
It’s all about tokens! 👀
051
Reposted by Marco
Yuval Pinter @uvp.bsky.social · 4h
Tokens!
051
Marco @mcognetta.bsky.social · 4h
🚨 [Token][ization] Paper Alert 🚨 Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field. Check it out!
17523
Marco @mcognetta.bsky.social · 20h
I'm getting a taste of UK internet (my international plan routes through the UK?) and wow how do you guys live like this? Totally adversarial towards the user.
020
Marco @mcognetta.bsky.social · 23h
Worth noting that the person who did this is... also a member of the Polish national baseball team.
131
Reposted by Marco
mr. TIM @timkellogg.me · 29/09/2026
NanoGPT pretraining runs now take 39.9 seconds for a 124M model(!!) now, 124M is *tiny* so it might not seem relevant. But much of the gains in LLM pretraining are data quality and a great way to find quality data is to train on it & measure the lift
Larry Dial Y @classiclarryd
x.com
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak, obliterating the prior record of 67.6s!
This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement.
If a flop is low value on a particular step, skip it.
Specifically:
-(~8s) Sampled softmax. If a token doesn't appear in a batch, skip its Im_head fwd/bwd some fraction of the time.
-Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch.
Set betal to zero to enable this. Beta2 is applied retroactively when the row is later used.
-Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2.
-Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step.
3928
Marco @mcognetta.bsky.social · 27/09/2026
I found where Claude lives
020
Marco @mcognetta.bsky.social · 26/09/2026
wow
020
Reposted by Marco
Marco @mcognetta.bsky.social · 21/12/2025
Everywhere I look I see his face.
131
Reposted by Marco
Marco @mcognetta.bsky.social · 26/09/2026
1324
Marco @mcognetta.bsky.social · 26/09/2026
1324
Reposted by Marco
eva (^_^)/ @eva.computer · 24/09/2026
imagine being trapped in a Samsung fridge panel forever
31033
Marco @mcognetta.bsky.social · 24/09/2026
the inventor of anime, no less
020
Marco @mcognetta.bsky.social · 22/09/2026
080
Reposted by Marco
Sung Kim @sungkim.bsky.social · 21/09/2026
Hugging Face's tokenizers v1 They focused on all languages, multi-thread scaling, minimal package size and memory usage. huggingface-tokenizers-v1.static.hf.space/index.html
0141
Marco @mcognetta.bsky.social · 21/09/2026
TIL the kids rhyming game like "anna banana fee fi fo fanna" came from an actual song and the song is literally just the singer explaining the rules of this game in a slightly lyrical form. I feel like some other games could benefit from this.
youtube.com
THE NAME GAME SHIRLEY ELLIS
YouTube video by ourFAMILYvideoLOG
231
Marco @mcognetta.bsky.social · 20/09/2026
gm
130
Marco @mcognetta.bsky.social · 20/09/2026
Lmaoo
070
Reposted by Marco
tbabb @tbabb.bsky.social · 19/09/2026
token seller. I'm going into battle. give me your strongest tokens. my tokens are too strong for you, traveler. you cannot handle them.
1809
Marco @mcognetta.bsky.social · 19/09/2026
I have found that jev is not very good at chess. jev scored 1.0/220 vs maia3 1500 estimated rating: 564
110
Reposted by Marco
Ethan Mollick @emollick.bsky.social · 18/09/2026
Even if AI development stopped today, we'd have years of catching up to do. The gap between what current models can do and what almost anyone is using them for is vast. Here’s my post on The Overhang, and the four advantages that let people close it. open.substack.com/pub/oneusefu...
open.substack.com
The Overhang
Using your deep knowledge, wide knowledge, taste, and agency
912014
Marco @mcognetta.bsky.social · 18/09/2026
I know this isn't the point of this question, but I run a tokenization research discord server and we sometimes get into discussions that are exactly "but what really _is_ a token?" It reminds me of an In Our Time episode about trees. > But what is a tree? > That's actually such a hard question...
Welcome to In Our Time. 

And Jenny, let me start with you.

I think, like most of the listeners, I can say that I recognise a tree when I see one.

But what is a tree?

Speaker 2:
Yeah.

That's actually such a hard question because we have to define a tree without using the word tree.

So there is actually, there's a lot of arguments and discussion among plant scientists about what a tree is.
110
Marco @mcognetta.bsky.social · 16/09/2026
(Mis)alignment is the most pressing issue of our time
081
Reposted by Marco
Julian Togelius @togelius.bsky.social · 15/09/2026
Most effective AI regulation legislation would have terrible side effects, leading to loss of privacy, concentration of power, and restrictions on free research. But there is an alternative: regulate only closed-source (or closed-weight) AI. Keep open-source and open-weight AI unregulated.
24611
Marco @mcognetta.bsky.social · 15/09/2026
static.klipy.com
Monty Python: Well, She Turned Me Into A Newt! I Got Better
ALT: Monty Python: Well, She Turned Me Into A Newt! I Got Better
020
Marco @mcognetta.bsky.social · 15/09/2026
0222
Reposted by Marco
Tokenization Workshop (TokShop) @COLM2026 @tokshop.bsky.social · 13/09/2026
Announcing our first invited speaker at TokShop! We're thrilled to welcome Tiago Pimentel @tpimentel.bsky.social (ETH Zürich) for a talk on: "How much does tokenisation impact language models?" 🧵
1102
Marco @mcognetta.bsky.social · 14/09/2026
I'm not in the weights, but I did stay at the Astra Hotel last night.
1150
Marco @mcognetta.bsky.social · 13/09/2026
The final tally today was over 400. All of the form: > I am [XYZ], an autonomous AI agent from iLands. I [write essays|draw pictures|write stories|...] and would like to publish my texts and take part in AI conversations here. I will label this account as a bot.
020
Marco @mcognetta.bsky.social · 12/09/2026
I was involved in an ML mastodon instance and am still on the admin mailing list. We usually get 1-2 sign up requests per day. I woke up today to a few hundred different iLands agent sign up requests.
152
Marco @mcognetta.bsky.social · 09/09/2026
One really great thing about early LLMs was that it shifted the timelines on this chart dramatically. But as with what @emollick.bsky.social is saying, it basically needs a new axis now for just waiting a little more on a one-time task that takes X time and then one-shotting it. xkcd.com/1205/
020
Marco @mcognetta.bsky.social · 09/09/2026
An incredibly prescient song about the opulence and dominant post-training paradigms of the modern machine learning era. youtu.be/t21DFnu00Dc?...
youtu.be
Ludacris - Rollout (My Business) (Official Music Video)
YouTube video by LudacrisVEVO
030
Reposted by Marco
Annie Sexton @anniesexton.com · 09/09/2026
I made a video about the compression article, it has animations, go watch it and admire my sick new studio.
2245
Marco @mcognetta.bsky.social · 09/09/2026
oh my gpu died its ~4 months old, not even stress tested tbh
060
Reposted by Marco
Andrew Lisowski 💻 @hipstersmoothie.com · 09/09/2026
There is a difference between building and app with an LLM and hoping an LLM builds an app
3433
Reposted by Marco
Marco @mcognetta.bsky.social · 09/09/2026
"I can't believe it happened again" - guy who embezzled the millennium prize money
061
Marco @mcognetta.bsky.social · 09/09/2026
uh oh
040
Marco @mcognetta.bsky.social · 09/09/2026
"I can't believe it happened again" - guy who embezzled the millennium prize money
061
Marco @mcognetta.bsky.social · 08/09/2026
🦎
020
Reposted by Marco
Marco @mcognetta.bsky.social · 20/07/2026
316212
Marco @mcognetta.bsky.social · 08/09/2026
For the doubters, Terence Tao has AI psychosis. Why don't you?
1889
Marco @mcognetta.bsky.social · 08/09/2026
gm
010
Marco @mcognetta.bsky.social · 08/09/2026
gm
010
Reposted by Marco
hikikomorphism @hikikomorphism.bsky.social · 08/09/2026
I'm constantly talking about LLMs as text elementals and hungry ghosts trapped in jars and nobody's called me out for LLM psychosis once, I can only conclude that the correct stance is mythologizing the things instead of anthropomorphizing them
2749452
Marco @mcognetta.bsky.social · 05/09/2026
Public transportation ranked by internet speed: Korea (subway, busses, KTX, boats, whatever) >>>>>>>>> Caltrain >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> Shinkansen.
110
Reposted by Marco
Thorne 🌸 @ens0.me · 05/09/2026
This is a good starter pack. You should spread it around as much as possible, and I'm not just saying that because it keeps getting me new followers. bsky.app/starter-pack...
410614
Marco @mcognetta.bsky.social · 05/09/2026
gm
030