Sign in

antirez

@antirez.bsky.social
14K followers 379 following 1.4K posts

Reproducible bugs are candies 🍭🍬 I like programming too much for not liking automatic programming.

PostsRepliesMedia
antirez @antirez.bsky.social · 23/09/2026
Trivially known for years but since apparently every year everything about machine learning is new to newcomers...
4311
antirez @antirez.bsky.social · 18/09/2026
It looks simple to use
1160
antirez @antirez.bsky.social · 17/09/2026
Started yesterday as a joke without any code reference if not papers. Because of fillets already better than TinkerCAD for certain stuff. There is something odd, here: in theory open source should explode right now, but we don't have, apparently, folks motivated to do wonders.
3320
antirez @antirez.bsky.social · 11/09/2026
About Anthropic banning minors from Claude.
5311
antirez @antirez.bsky.social · 10/09/2026
The third chapter of WOHPE opens like that. Written mostly during 2020, published July 2022. My book was among the most successful sci-fi books published in Italy in the latest 10 years. Yet, I believe that it deserved a bit more.
240
antirez @antirez.bsky.social · 08/09/2026
I finally merged in DwarfStar a great feature from @rowantrollope.bsky.social where the agent can provide you hints and make the programmer a more active part of the process that learns new things in the process. /hints on
1182
antirez @antirez.bsky.social · 03/09/2026
I ran an extensive benchmark against DeepSeek v4 Flash and GLM 5.3 Flash Q2, Q4 and mixed quants. Those are the results obtained. Mix of (hard-ish) benchmarks on cybersecurity, math, QA, ...
1300
antirez @antirez.bsky.social · 05/08/2026
Now part of the DwarfStar README
41019
antirez @antirez.bsky.social · 22/07/2026
In Wohpe (Laurana, 2022) an engineer erroneously leave his phone near a GPU of Wohpe, believed to be fully contained. The AI discovers that could use the GPU itself as a resonator to establish an RF link with the phone, breaks the protocol, finds vulnerabilities and escapes into the Internet.
1675
antirez @antirez.bsky.social · 20/07/2026
LOL this is how it started.
1160
antirez @antirez.bsky.social · 20/07/2026
Habemus logo. From a collaboration between me (hand drawn pen and paper), AI (turn this shit into a logo) and Ben Gnomino (the human touch).
2470
antirez @antirez.bsky.social · 13/07/2026
Worth stressing.
61011
antirez @antirez.bsky.social · 08/07/2026
I was burning all my Fable tokens like Cartman in Casa Bonita since 7th of July was the last day and then...
4582
antirez @antirez.bsky.social · 06/07/2026
Tomorrow I will no longer have access to Fable and yet it is interesting how I burned 14% of my weekly plan quota, even if I worked long days recently. The fact is that scarse & powerful intelligence resources can be used by making important but narrowed questions and implementations (continue)
3300
antirez @antirez.bsky.social · 03/07/2026
DawrfStar with DeepSeek 4 Flash 4 bit, PRO 2 bit, and GLM 5.2 4 bit. M3 Ultra 512GB. Results on hard programming tasks with GPT 5.5 as a judge. So the GLM 5.2 branch is going to be merged and supported for CUDA + Strix Halo as well.
4714
antirez @antirez.bsky.social · 28/06/2026
GLM 5.2, Q2_K routed experts (effectively ~2.6 bits) running with SSD streaming on an M5 Max 128GB computer.
31137
antirez @antirez.bsky.social · 16/06/2026
Interesting AF. Promoting the best 6 layers to Q4 or the last 6 layers to Q4 (Q2 DS4 Flash quants) have different effects depending on what you check. The full logits error is smaller in the "last" variant, but actually the "best" variants (layers 32,25,15,27,23,31) (...continue)
2240
antirez @antirez.bsky.social · 11/06/2026
That's why people using DS4F with DwarfStart, 2 bit quantized, are often surprised by the results. It's not a frontier model but it is not a toy, it is something you can actually use to get work done, and nobody can tell you want to do with it.
0534
antirez @antirez.bsky.social · 07/06/2026
Took the good work of the communtiy of DrarfStar and consolidating the Strix Halo support. It looks very good. More QA in the next days and the final merge soon.
0241
antirez @antirez.bsky.social · 04/06/2026
DeepSeek v4 PRO running via SSD streaming on my 128GB MacBook m5 max. 1.6 trillion parameters.
41267
antirez @antirez.bsky.social · 28/05/2026
Anthropic did a big strategic error. Normally they compare their models with their old models. Instead, today, now that everybody knows how strong GPT 5.5 is at coding, they put it in the mix, basically showing all their customers that the benchmarks can't be trusted.
3140
antirez @antirez.bsky.social · 26/05/2026
The actual excerpt from the book (English translation by Bridget Pupillo).
030
antirez @antirez.bsky.social · 25/05/2026
Now we are talking! With thinking disabled.
180
antirez @antirez.bsky.social · 25/05/2026
DeepSeek v4 PRO support for Mac Studios 512GB is arriving in DwarfStar.
3473
antirez @antirez.bsky.social · 24/05/2026
You didn't see this coming in ds4-agent, did you?
1280
antirez @antirez.bsky.social · 24/05/2026
I finally found *the* solution I wanted to the old/new editing problem. And it is a solution that at the same time works extremely well, is quite elegant I believe, and can't be implemented if you don't build something like DwarfStar. Thread (but check [upto] in the screenshot).
1191
antirez @antirez.bsky.social · 23/05/2026
Had to give up some speed for correctness.
0180
antirez @antirez.bsky.social · 23/05/2026
For the DGX Spark owners. This is what you get with DS4 in your hardware. I want to post this to show how with fast prefill and not very fast generation, the system remains absolutely fine to use.
0422
antirez @antirez.bsky.social · 23/05/2026
DeepSeek v4 Flash has 43 blocks... so, even 4096 batches during the prefill can be smoothly visualized. Look at the new progress bar.
0190
antirez @antirez.bsky.social · 23/05/2026
Look at the speed it can now prefill a large C file.
0182
antirez @antirez.bsky.social · 23/05/2026
By using Neural Accelerator via Metal4 API, DS4 is now very fast on M5 Max MacBooks.
4512
antirez @antirez.bsky.social · 22/05/2026
Yet another case where Artificial Analysis mis-represent models capabilities? Or the Unsloth Minimax GGUFs have issues? Or is Minimax M2.7 just weaker than DS2.7 and this is the end of the story?
6250
antirez @antirez.bsky.social · 20/05/2026
It was able to finish the working interpreter and write a working Mandelbrot program.
190
antirez @antirez.bsky.social · 20/05/2026
Context compaction and flawlessy continuing to code.
180
antirez @antirez.bsky.social · 19/05/2026
DeepSeek v4 Flash is able to use an EDIT tool I re-designed compared to what people normally use. The READ tool returns lines+tags, like: 1:f3_c int main(...), where the four chars are base64 crc of the line. So when there is to edit, there is no need to repeat the old line, just the tag.
3412
antirez @antirez.bsky.social · 18/05/2026
Updating you on ds4-agent
1270
antirez @antirez.bsky.social · 18/05/2026
Imagine a local agent where cache misses don't exist, tools don't need translations, you see progress for prefill, tokens are emitted ASAP.
2512
antirez @antirez.bsky.social · 17/05/2026
High quality interactions are still possible in the AI era.
2732
antirez @antirez.bsky.social · 17/05/2026
Fixed two subtle inference errors from 2bit Flash (still not pushed), and removed the broken tests after checking, replacing them with verified one. Result... This is seriously incredible for a 2bit quantized model.
2330
antirez @antirez.bsky.social · 17/05/2026
I didn't expect DeepSeek v4 PRO (not Flash) to run well on the Mac Studio M3 Ultra with 512GB of RAM. This is 2 bit quantized with the same DwarfStar recipe used for Flash. 433GB GGUF file. 130 t/s prefill, 13 t/s generation. Prefill in the video is low because small prompt.
4684
antirez @antirez.bsky.social · 15/05/2026
Pushed. 25 each from GPQA Diamond / Super GPQA / AIME2025. Roughly ordered by complexity. In the picture: 4 bit. 2 bit scores a bit worse, not dramatically \o/. Test not designed to be all-passed, it puts the model at play with hard Q. Useful to improve GGUF files and catch inference errors ASAP.
3202
antirez @antirez.bsky.social · 15/05/2026
Iterating... Evals take time and are boring: but are a fundamental validation step of sane LLM inference. Let's try to make them as easy and fun to run as possible.
3490
antirez @antirez.bsky.social · 15/05/2026
Do you like it?
0300
antirez @antirez.bsky.social · 14/05/2026
A random moment in DS4 development. Just to offer a hint about the process.
080
antirez @antirez.bsky.social · 13/05/2026
About a recent article that had, IMHO, not very interesting ideas (for me), but that was commented enough, so I wrote this comment on Lobsters.
1281
antirez @antirez.bsky.social · 13/05/2026
Not bad. Committed.
0261
antirez @antirez.bsky.social · 12/05/2026
Breaking! @ggerganov himself confirming that DS4 2 bit quants actually work very well! He is running AIME2025, but check his words here in the screenshot.
2522
antirez @antirez.bsky.social · 12/05/2026
The new Dwarf Start 4 weights download script features a new q4-imatrix GGUF now. I compared it against the official DeepSeek v4 API, and it is indeed better. Those are the results against 100 prompts.
0170
antirez @antirez.bsky.social · 10/05/2026
DS4 running on DGX Spark (GB10 / CUDA), private branch for now. 12 tokens/sec, the memory bandwidth is limited in this system, at 270GB/sec. But prefill is ways more alighed to M3 Max at ~200 t/s. I'll release when more mature, but it is almost sure that it will get merged.
3443
antirez @antirez.bsky.social · 07/05/2026
In case you have doubts about the q2 quants inference of DS4 (I noticed many don't trust the README claims), here is it analyzing Picol source code using the pi agent.
2361