Sign in

antirez

@antirez.bsky.social
14K followers 379 following 1.4K posts

Reproducible bugs are candies 🍭🍬 I like programming too much for not liking automatic programming.

PostsRepliesMedia
antirez @antirez.bsky.social · 29/09/2026
Don't look at generation tokens per second as the main metric for local inference. Soon Cinese open model providers will realize that the game is all about making the thinking phase as short as possible, and 20/30 t/s will be enough if you have fast prefill.
5361
antirez @antirez.bsky.social · 29/09/2026
I claim that if you are a normal person that like me don't have "fuck you money", so unable to spend large sums of money randomly, you should do all your programming work with two max accounts (OpenAI / Anthropic for instance) in the worst case, and never buy tokens via API.
7710
antirez @antirez.bsky.social · 28/09/2026
The most ridiculous thing about the AI-coding age is that in the last 20 years programmers accepted otherwise all the kind of shit: terrible frameworks, languages, layers of complexity everywhere, everything super slow. Very few protested, since destroying the field didn't involve their paycheck.
3534
antirez @antirez.bsky.social · 28/09/2026
ZX Spectrum 48k Another World demo released: github.com/antirez/anot...
github.com
GitHub - antirez/anotherworld-zx-spectrum-48k: Actual VM implementation of Another World game, with intro
Actual VM implementation of Another World game, with intro - antirez/anotherworld-zx-spectrum-48k
1141
antirez @antirez.bsky.social · 27/09/2026
Good programmers state they use models like GPT 6 Astra without obtaining good results: it makes me feel I live in a parallel universe. But the explanation is not that hard: good programming in the past and now requires a different skill set, even if there is some overlap.
8574
antirez @antirez.bsky.social · 25/09/2026
I'm worried to see people that think the right move is to question AI companies and narratives around AI regardless of logic and the actual significance of events we are living. They act just as an inverted flock, but still a flock. Sometimes it's fear of friends judgement.
4192
antirez @antirez.bsky.social · 23/09/2026
Trivially known for years but since apparently every year everything about machine learning is new to newcomers...
4311
antirez @antirez.bsky.social · 21/09/2026
Jev may have its (narrow) use cases but the hype you see around is the perfect representation of the fact the greatest part of the AI bubble don't know what is important and what is not. Jev is a minor thing happening on AI compared to all the rest, yet the hype exploded.
15904
antirez @antirez.bsky.social · 18/09/2026
It looks simple to use
1160
antirez @antirez.bsky.social · 17/09/2026
Started yesterday as a joke without any code reference if not papers. Because of fillets already better than TinkerCAD for certain stuff. There is something odd, here: in theory open source should explode right now, but we don't have, apparently, folks motivated to do wonders.
3320
antirez @antirez.bsky.social · 16/09/2026
At this point DwarfStar contains many fast fused kernels for important model families: feel free to steal everything you want from there according to the MIT license, in order to improve your own implementation.
0654
antirez @antirez.bsky.social · 16/09/2026
Witch hunting level: some guy uses AI to reverse engineer an M4 GPU driver for Linux and part of the community that should be for the open source, for the hacking, for the liberation and freedom is against him.
415716
antirez @antirez.bsky.social · 15/09/2026
The irreconcilable misunderstanding about AI is that for some of us is the way to remove suffering, inequality, limits from the human race. For others, a tool: and, right now, the capabilities are at tool level, giving the illusion that the tool is the point.
2320
antirez @antirez.bsky.social · 14/09/2026
Qwen3.8 Flash Next is now supported in DwarfStar, covering 64GB Mac systems very well and with very fast inference of 50~70 t/s and > 1400 t/s prefill. For now this is Metal only. Thanks to @ivanfioravanti.bsky.social for all the cool work in the PR. N-grams on SSD like for DS4.1F.
3778
antirez @antirez.bsky.social · 14/09/2026
It drive me nuts that now people use AI to write tweets. If you think that this way you will get more popular and so forth, think twice. Write something authentic. AI is great but it is easy to misuse. Your tweets must be your more cared thoughts.
51024
antirez @antirez.bsky.social · 13/09/2026
With the last commit into DwarfStar now you can use DeepSeek v4.1 Flash in a single DGX Spark as well, with SSD streaming. Around 9 t/s generation. It works also dual-spark RDMA at ~22 t/s.
0656
antirez @antirez.bsky.social · 13/09/2026
Conspiracy theories are, very often, illogical. If AI companies were worried by open weight models (they probably are btw) the logical response would be to *not* slow down the development of frontier AI, to try locking the advantage. Stop with nonsense.
3212
antirez @antirez.bsky.social · 12/09/2026
DeepSeek v4.1 Flash support is now pushed on DwarfStar "main" branch on GitHub, and this is a YouTube video (in English language) where I test both the SSD streamed and the dual MacBook m5 max 128GB setup during a coding session: www.youtube.com/watch?v=ogs8...
youtube.com
Let's test DeepSeek v4.1 Flash 2 bit with DwarfStar (ENG)
YouTube video by Salvatore Sanfilippo
1467
antirez @antirez.bsky.social · 11/09/2026
About Anthropic banning minors from Claude.
5321
antirez @antirez.bsky.social · 11/09/2026
I'm a simple man. If a YouTube video cover has a stunned face on it, I don't watch the video.
816013
antirez @antirez.bsky.social · 10/09/2026
DwarfStar running DeepSeek v4.1 Flash on a 128GB M5 Max at steady 15 t/s. I didn't expect with SSD streaming it could be so fast. Recent SSD streaming changes to retain the right experts surely helped, but also maybe DS4.1 uses the same experts more. Will push online when ready QA > ASAP.
29413
antirez @antirez.bsky.social · 10/09/2026
The third chapter of WOHPE opens like that. Written mostly during 2020, published July 2022. My book was among the most successful sci-fi books published in Italy in the latest 10 years. Yet, I believe that it deserved a bit more.
240
antirez @antirez.bsky.social · 10/09/2026
P.S. I believe the credits should go exclusively to Córdoba–Martínez-Zoroa and AI.
010
antirez @antirez.bsky.social · 09/09/2026
Why Terence Tao is wrong on AI + Math. A thread: 1. Net negative is factually false. After a proof or disproof we know strictly more than before. No opaque process reduces the set of knowledge we have. About the human math trajectory "net negative" read the next points.
3143
antirez @antirez.bsky.social · 08/09/2026
I finally merged in DwarfStar a great feature from @rowantrollope.bsky.social where the agent can provide you hints and make the programmer a more active part of the process that learns new things in the process. /hints on
1182
antirez @antirez.bsky.social · 06/09/2026
DwarfStar in the latest two weeks was improved in almost every aspect for Metal, DGX Spark and Strix Halo. It is simpler to say: update, you will hopefully see speed and correctness improvements in many areas. Also DSpark with DeepSeek v4 Flash now works much better overall.
2343
antirez @antirez.bsky.social · 03/09/2026
I ran an extensive benchmark against DeepSeek v4 Flash and GLM 5.3 Flash Q2, Q4 and mixed quants. Those are the results obtained. Mix of (hard-ish) benchmarks on cybersecurity, math, QA, ...
1300
antirez @antirez.bsky.social · 03/09/2026
We don't give a !(@$# about what the new model can generate with three.js. We want to know if you had a blocking problem X in software Y that many rounds of the old models didn't fix, and the new model just released improved the situation.
4450
antirez @antirez.bsky.social · 03/09/2026
Asymmetric access to LLMs for cyber security workis bad. I decided to release the steering vector that you can use with DwarfStar --dir-steering-file <file> and --dist-steering-ffn (try 2, 3, 4 based on refusal) to make DS4F comply. antirez.com/refusal_trai...
antirez.com
2351
antirez @antirez.bsky.social · 30/08/2026
Given that 51B of n-grams can stay on the SSD disk, Qwen 3.8 Flash Next with a 2 bit quantization could be the best DwarfStar bet for a 64GB MacBook local inference top experience. I understand @ivanfioravanti.bsky.social can spend some time with the implementation. Let's see what happens.
2696
antirez @antirez.bsky.social · 26/08/2026
Btw the Redis story repeats itself: I'm working at DwarfStar for free for the community and because I enjoy it. But I'm receiving criticisms, since people are worried that this will break their AI-richness plans.
7815
antirez @antirez.bsky.social · 26/08/2026
NVIDIA is kindly sending me two DGX Spark machines. You will see the single / double Spark as one of the main supported platforms of DwarfStar soon. Metal + Spark + Strix Halo will be the main focuses as usually, but now even better.
41033
antirez @antirez.bsky.social · 17/08/2026
Do you think IQ captures how well a person will perform in general intellectual tasks? I hope you don't. For the same reason, don't take Artificial Analysis benchmarks as they tell you the whole story about how much powerful an LLM is in the real world.
1541
antirez @antirez.bsky.social · 16/08/2026
It's hard to think at something more stupid than the EU-requested watermarking of AI generated text.
10746
antirez @antirez.bsky.social · 16/08/2026
I uploaded the new (0813) Q2 quants of DeepSeek v4 PRO at the following URL. Quality tests still ongoing. huggingface.co/antirez/deep...
huggingface.co
DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix-0813.gguf · antirez/deepseek-v4-gguf at main
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
2404
antirez @antirez.bsky.social · 15/08/2026
NVIDIA provided my two weeks SSH access to DGX Station. My gaol is not just "test DwarfStar" there. I'll try making the DGX Station able to do things otherwise not possible with the current software stack available. Big claim, but we will see in a couple of days :D
3872
antirez @antirez.bsky.social · 14/08/2026
DSpark speculative decoding is a tragedy for DeepSeek v4 Flash benchmarks. You can't trust anything, since it is too dependent on what you are generating (extreme case: count from 1 to 100). Always publish no Dflash numbers *as well* if you want to build trust.
2191
antirez @antirez.bsky.social · 13/08/2026
The draft sorted sets PR. 60% less memory in non trivial use cases, and certain operations are even 70% faster (measured via API + networking!). github.com/redis/redis/...
github.com
Reduce large sorted set memory with a packed B+ tree by antirez · Pull Request #15635 · redis/redis
Hi, this is a draft implementation that reimplements large sorted sets by replacing the skiplist and general-purpose dictionary with one packed B+ tree (score ordered) and a compact member index, t...
0400
antirez @antirez.bsky.social · 11/08/2026
Less than 100 years ago a group of Italians worked together to make discoveries that enabled nuclear energy: en.wikipedia.org/wiki/Via_Pan...
en.wikipedia.org
Via Panisperna boys - Wikipedia
2243
antirez @antirez.bsky.social · 10/08/2026
Follow my simple reasoning here. Large sparse LLMs are more powerful then dense one *while* being faster and thus more energy efficient. Dense models are a need that is artificially created by VRAM scarcity. Today they are practically useful, but not for long.
1492
antirez @antirez.bsky.social · 10/08/2026
Fast H3 implementation for Metal. Enjoy, modify, and so forth: github.com/antirez/h3.c Contains code from @liuliu which is welcomed in taking back whatever parts he likes for @drawthingsapp in case there are H3 plans there.
github.com
0272
antirez @antirez.bsky.social · 09/08/2026
DwarfStar with DFlash speculative decoding now can do the DeepSeek v4 Flash inference *much* faster both when used with Metal and the DGX Spark.
0361
antirez @antirez.bsky.social · 07/08/2026
Every LLM chart where there is "dollars" in the x coordinate should have Joules instead.
5542
antirez @antirez.bsky.social · 07/08/2026
In case you know somebody at NVIDIA that would value early DwarfStar support for the DGX Station, please ping them saying I would be interested in receiving one and making the Station one of the main targets of the project. It fills the low power small-mid company target of DS.
3444
antirez @antirez.bsky.social · 06/08/2026
Saying that LLMs using the chain of thoughts and an harness to execute tools is a symbolic systems (in the sense they claimed was needed VS LLMs), to avoid saying to be wrong, is *the worst* mystification that ever happened in the history of science. Never forget.
2170
antirez @antirez.bsky.social · 05/08/2026
2 Updates: DwarfStar and Redis new sorted sets: 1) The new version of DwarfStar is much faster in prefill both on the M5 Max and the DGX Spark. It also integrates 14 pull requests / issues fixes and other stuff. Time to upgrade :)
2261
antirez @antirez.bsky.social · 05/08/2026
Now part of the DwarfStar README
41019
antirez @antirez.bsky.social · 01/08/2026
Read those HN comments as it is sociologically super interesting to see people that can't cope with AI results news.ycombinator.com/item?id=4913...
news.ycombinator.com
Ten advances in mathematics and theoretical computer science | Hacker News
2241
antirez @antirez.bsky.social · 01/08/2026
Btw now that I can run MXFP4 locally and generate greedy continuations, I can tell you: not all the OpenRouter providers for DS4 flash are sane... Not going to do names, but some provider gives you the real shit, other will not.
5301
antirez @antirez.bsky.social · 01/08/2026
DwarfStar branchk "ds4f-mxfp4" now can run the lossless MXFP4 DeepSeek v4 Flash GGUF I published on my Hugging Face account. It rocks even with SSD streaming in 128GB systems at > 20 t/s in case you want to try the *actual* DS4F weights released without any quantization.
2765