Sign in

Matthew Carrigan

@carrigmat.bsky.social
254 followers 154 following 84 posts

Engineer @huggingface. I'm the reason your LLM frontend has a jinja2cpp dependency. Sometimes yells about housing and trans rights instead of working He/him

PostsRepliesMedia
Reposted by Matthew Carrigan
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 30/09/2026
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
nature.com
Scalable decision-making for games of imperfect information - Nature
Ataraxos, an AI for the board wargame Stratego, establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desiderat...
1225661
Matthew Carrigan @carrigmat.bsky.social · 22/09/2026
Yelling at passing biologists to put ESM-2 down and switch to ESMC
000
Matthew Carrigan @carrigmat.bsky.social · 21/09/2026
It's genuinely unappreciated how runnable something enormous like Kimi-K3 is (>2.5T parameters!) on a local machine with a budget in the $10k range. 1-socket 12-channel EPYC, 192GB RAM, 9175F for more CCDs, and keep the experts on a massive NVMe RAID array
210
Matthew Carrigan @carrigmat.bsky.social · 20/09/2026
The secret to being a great hardware builder for LLM inference is to hear sentences like "Inference is bottlenecked by memory bandwidth" then take them to their logical conclusion in the face of all decency and common sense
Four PCIE splitter cards containing 16 NVMe drives with a total read bandwidth of 200GB/s. This is significantly more than most systems' RAM bandwidth
020
Matthew Carrigan @carrigmat.bsky.social · 29/05/2026
a thousand usernames speak to me in the same voice, all repeating the same song: "I am real, I am unique, I am alive"
120
Matthew Carrigan @carrigmat.bsky.social · 23/03/2026
The code agents are requesting payment now
010
Matthew Carrigan @carrigmat.bsky.social · 18/03/2026
Deeply curious whether this was caused by the user's prompt, or whether this is Claude's well-documented support for animal rights shining through into its code agent work
000
Matthew Carrigan @carrigmat.bsky.social · 16/03/2026
Combining Popper and Jaynes in a thesis I call "The Origin of Bicameral Parliament in the Breakdown of the Bicameral Mind"
000
Matthew Carrigan @carrigmat.bsky.social · 03/02/2026
Really nice bio model from DeepMind just got released! There have been quite a few DNA foundation models in the past, but labs usually had to gather their own data and fine-tune them for tasks of interest. This is something else!🧵
huggingface.co
google/alphagenome-all-folds · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
110
Matthew Carrigan @carrigmat.bsky.social · 07/11/2025
Now seems like a good time to repeat this thread, since Kimi-K2-Thinking has just arrived and might actually be the strongest LLM in the world right now, open or closed huggingface.co/moonshotai/K...
huggingface.co
moonshotai/Kimi-K2-Thinking · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
100
Matthew Carrigan @carrigmat.bsky.social · 03/11/2025
PRs and issues on @hf.co have gotten a lot sloppier and weirder since the advent of code agents, but the weirdest ones still have an inexplicable human touch
010
Matthew Carrigan @carrigmat.bsky.social · 29/10/2025
Extremely fascinated by the latest Anthropic post, but parts of the results feel like they might just be the result of "the right amount of steering" rather than genuine introspection. www.anthropic.com/research/int...
anthropic.com
Emergent introspective awareness in large language models
Research from Anthropic on the ability of large language models to introspect
210
Matthew Carrigan @carrigmat.bsky.social · 10/10/2025
An underappreciated thing about the Turing test is that every teacher, writer and artist on the planet is now intimately familiar with the markers of AI output. The post-ChatGPT era is like a global training montage to ensure the bot's job in that test is as hard as possible
000
Matthew Carrigan @carrigmat.bsky.social · 13/05/2025
Underappreciated linguistic fact: "Thou" was originally an informal, friendly pronoun, but feels extremely archaic and formal to modern ears because of its association with Shakespeare and the KJV. You'd use it for speaking to family and friends (and to God).
120
Matthew Carrigan @carrigmat.bsky.social · 27/04/2025
the betting markets are asking the real questions today
000
Matthew Carrigan @carrigmat.bsky.social · 25/04/2025
The discussion pages for Open-R1 on @hf.co are such a goldmine for actual practical information on how to train a reasoning model. Like look at this! If you're not reading those community tabs you're missing so much! huggingface.co/spaces/open-...
huggingface.co
open-r1/README · [Experiment] Training R1-Zero-like models with Open R1
There are several recent research papers which explore various aspects of R1-Zero-like training on open base models like Qwen2.5-7B and Llama-3.1-8B:
092
Matthew Carrigan @carrigmat.bsky.social · 17/04/2025
I call this The Paper. It gets written quite often in machine learning, and it's valuable every time! The core of it is "Everyone had a complex setup to do X task. With enough scale, none of that complexity is necessary, and a simple model does it better." huggingface.co/papers/2503....
huggingface.co
Paper page - Your ViT is Secretly an Image Segmentation Model
Join the discussion on this paper page
060
Reposted by Matthew Carrigan
Clayton Thorrez @cthorrez.bsky.social · 16/04/2025
Here's EsportsBench v5! 72k new matches added from 2025-01-01 through 2025-03-31 and some data quality improvements to past data as well. Over 2.4 million rows of esports match data from 20 titles spanning over 25 years huggingface.co/datasets/Esp...
huggingface.co
EsportsBench/EsportsBench · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
042
Matthew Carrigan @carrigmat.bsky.social · 29/03/2025
I believe ArXiv and Archive Of Our Own should swap places for April 1st. I believe this more strongly than I believe anything else
04514
Matthew Carrigan @carrigmat.bsky.social · 25/03/2025
People are reading MSFT dropping power contracts as a sign that AI investment will fall off, but if reasoning is the new paradigm then most training compute will be inference and that doesn't have to be centralized Massive monolithic datacentres are much less necessary now
100
Matthew Carrigan @carrigmat.bsky.social · 24/03/2025
Preliminary take is that V3-0324 is a major upgrade on the V3 base. Increasingly confident that it's the strongest open-source LLM, and likely competitive with the top tier of closed source too
000
Matthew Carrigan @carrigmat.bsky.social · 24/03/2025
Deepseek V3-0324 just landed, an upgraded version of the V3 model that was used as the base for Deepseek-R1. Weights on @hf.co , and it'll start appearing on inference providers soon. It seems very strong in early testing, likely the best non-reasoning OS model (!) huggingface.co/deepseek-ai/...
huggingface.co
deepseek-ai/DeepSeek-V3-0324 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
110
Reposted by Matthew Carrigan
jsulz @jsulz.com · 18/03/2025
Last week, we launched a waitlist to move builders on @hf.co from LFS to Xet. This was made possible through months of hard work and staged migrations to test our infrastructure in real-time. This post provides an inside look into the day of our first migrations and the weeks after.
huggingface.co
Xet is on the Hub
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1113
Matthew Carrigan @carrigmat.bsky.social · 17/03/2025
Anyone want to explain to me where Anthropic are getting "powerful AI will arrive somewhere from late 2026 to early 2027"? I totally get being AGI-pilled and extrapolating scaling laws, but we've just moved to a new reasoning scaling regime! We don't even have enough points to extrapolate!
010
Matthew Carrigan @carrigmat.bsky.social · 06/03/2025
Shower thoughts: Claude can play Pokemon but it's obviously far too slow for other games. What would an LLM system that can actually play Sonic look like? An LLM giving high level direction and "rewards" to a fast small convnet-RL model that actually mashes the buttons?
200
Matthew Carrigan @carrigmat.bsky.social · 24/02/2025
I work at @hf.co and monitor every issue/PR to Transformers. I estimate about 20% of those are now written by AI, run by various users and companies. In some cases I chat with the AI during the PR review and it makes improvements
010
Matthew Carrigan @carrigmat.bsky.social · 28/01/2025
Complete hardware + software setup for running Deepseek-R1 locally. The actual model, no distillations, and Q8 quantization for full quality. Total cost, $6,000. All download and part links below:
3367
Matthew Carrigan @carrigmat.bsky.social · 26/01/2025
You can just keep putting more DDR5 sticks in your server and it will keep getting more intelligent. This scaling law may run out at some point, but it shows no sign of saturation yet
000
Matthew Carrigan @carrigmat.bsky.social · 26/01/2025
Non-AI friends are messaging me and asking me about Deepseek-R1. Don't think anything in the field has gotten this much attention since the ChatGPT launch
010
Matthew Carrigan @carrigmat.bsky.social · 25/01/2025
imo the most bullish case for open-source models is not that they'll be #1 forever, but that closed-source providers will be unable to share the CoTs due to distillation fears. If you care in the slightest about what your model is mumbling to itself, you need open-source
000
Matthew Carrigan @carrigmat.bsky.social · 24/01/2025
Looked at the DeepSeek-R1 repo like "Woah, it's big, but once I quantize with llama.cpp it'll be half the size" before realizing that no it won't because a lot of the weights are float8 already. Whatever, 650GB GGUF file let's go
020
Reposted by Matthew Carrigan
Ethan Mollick @emollick.bsky.social · 25/12/2024
If you have ever tried to read free books from sites like Project Gutenberg, you noticed that they can be uncomfortable to read, due to their layouts, type & occasional errors This project takes those free books and makes them beautiful (and still free). standardebooks.org
623954
Matthew Carrigan @carrigmat.bsky.social · 19/01/2025
From an LLM's perspective, we're all crowding around going❓🔢 🇷 ▶️ 🍓 and then bursting out laughing when it says "two"
020
Matthew Carrigan @carrigmat.bsky.social · 17/01/2025
MiniMax-01 is on @hf.co as of 48 hours ago. At 450B parameters, it's one of the largest open-source LLMs, and in my very informal evals, the strongest one by a clear margin. It's very, very smart. This is a thread about running it locally 🧵
120
Matthew Carrigan @carrigmat.bsky.social · 16/01/2025
MiniMax-01 passes the vibe eval, I repeat, MiniMax-01 passes the vibe eval. We may have a new open-source SOTA
020
Matthew Carrigan @carrigmat.bsky.social · 15/01/2025
Pro tip: If you're trying MiniMax-Text 01 and getting errors, use revision="pr/5" in the AutoModel initialization line. The base code uses an import that was removed in the most recent versions of Transformers huggingface.co/MiniMaxAI/Mi...
100
Matthew Carrigan @carrigmat.bsky.social · 03/01/2025
Cautionary tale: It's tempting to get your AI to assume a human persona, but LLMs are still not that smart and often fall into tropes That makes them feel like a cheap impersonation, or worse, a parody of a real identity, and that's a very very sore spot in (particularly American) racial politics
010
Matthew Carrigan @carrigmat.bsky.social · 30/12/2024
Does anyone know any papers on turning natural language feedback into training gradients? For example, you correct a model, it writes a corrected answer, and then you generate a new training example of (original_question, corrected_answer), or something like that
000