aria @aurelium.me · 29/09/2026unfortunately all of the cool things about training nondifferentiable models also make them really terrible for running on GPUs (branching, mainly) so the OOMs of per-token compute reduction you'd need to match backprop with a zeroth-order method are kind of unlikely to manifest 030
aria @aurelium.me · 29/09/2026imo applying zeroth-order optimization to transformers is missing the point completely why would you approximate gradients via sampling (likely >50x slower) when you can compute them directly. to the extent that ZO is useful, it is for training architectures that don't work with backprop 140
aria @aurelium.me · 29/09/2026(also, I'd be very surprised if the paper reveals that "core assumptions in optimization research are completely wrong". ZO-type methods are mathematically principled and this has been known for a while. the practical version of this in transformers exists, it's called RL) 040
aria @aurelium.me · 29/09/2026you might think that in order to calculate the pseudograds you need to store 256x random perturbations of the weights, which would be huge. but in theory you could embed deterministic PRNG into your forward pass kernels, so all you need to store is the seeds my guess is that this is what they did 140
aria @aurelium.me · 29/09/2026zeroth-order optimization in this context is approximating gradients via sampling. you apply a large number of random perturbations to the weights and use the effect these perturbations have on the loss to calculate pseudogradients 130
aria @aurelium.me · 29/09/2026"population" here refers to the number of forward passes. typical backwards pass is maybe 2x the wall-clock time of the forward pass, so this is around 85x less efficient than backprop (assuming zero time losses from their weird perturbation kernels) 170
aria @aurelium.me · 28/09/2026the chinese room is basically a verbal magic trick. surely something this mundane, a man reading a book, cannot together make up a system that "understands Chinese" even if the man does not and the book is inert without him ...except the book is 1000 OOMs larger than the observable universe 110
Reposted by aria{🧪} +paoloricciuti.svelte @paolo.ricciuti.me · 23/09/2026{model_name} is sooooo good, I've built a {3d_game_demo} with it and it one shotted it. {3d_game_demo_video} 2385
aria @aurelium.me · 22/09/2026arxiv.org/pdf/2609.22978 how could I miss the most important ML event of the week... the DeepSeek sandboxing tech report!arxiv.org 1184
aria @aurelium.me · 22/09/2026Huawei Ascend kernels are the most insidious avenue for xrisk... thank you for saving us, dario. 13510
aria @aurelium.me · 21/09/2026huggingface.co/XiaomiMiMo/M... MiMo-V2.6 is out! and, more importantly to me, the tech report. this is the most batteries-included tech report for a modern frontier model ever. really cool stuff.huggingface.coMiMo_V2_6_technical_report.pdf · XiaomiMiMo/MiMo-V2.6-Pro-RL at mainWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 0545
aria @aurelium.me · 19/09/2026signal seems to only give me notifications on my phone when it is asleep. this leads to the app exclusively notifying me when I am actively using the desktop app and never when I am actually willing to check it with my phone 000
aria @aurelium.me · 17/09/2026the main thing that would be difficult to replicate (and they say as much themselves) would be the data pipeline. getting good enough coverage for generalist classification tasks is a lot of work 130
aria @aurelium.me · 17/09/2026text diffusion (generally) isn't really anything like the complicated algorithms used for image diffusion it's not a particularly complicated modification to modeling code or loss compared to a normal LM classifier, and even the diffusion aspect isn't really necessary for their performance 130
aria @aurelium.me · 17/09/2026if you really want to take advantage of LLMs to write easy cross-platform software, use them to build some abstraction with ~no regard for "not being unpleasant to write" but extremely stringent requirements on correctness and portability 080
aria @aurelium.me · 17/09/2026my brain is permanently fried from writing software in the most un-unit-testable domain possible (RL trainers, which combine GPU integration, multi-node parallelism, and coordination across thousands of remote processes) but I wouldn't want to build ~anything another person used like this 2110
aria @aurelium.me · 17/09/2026tbh I really do not understand this mindset maybe for low-stakes websites and bespoke single-use software but for anything serious, the source code is important because it is the thing you've integration tested to within an inch of its life 4281
aria @aurelium.me · 16/09/2026although some pieces of copy strongly imply continuous diffusion. which is extremely strange as both an arch choice and a marketing idea but I guess if that's what works for ya 030
aria @aurelium.me · 16/09/2026(actually my guess is that it's even more bog-standard: a completely normal autoregressive LLM in constrained-decoding model, tuned on classification. my evidence is that their docs include a return value where 212 "output tokens" are used in a prompt only evaluating 15 things) 160
aria @aurelium.me · 16/09/2026a lot of people are saying it's a block-diffusion model because of the parallel output but it could easily be a completely normal LM classifier running in parallel req 1: text + "Was the author angry?" req 1: text + "Was the author sad?" etc. use standard shared-prefix KV cache, costs ~the same 150
aria @aurelium.me · 16/09/2026it's a neat idea, the idea of a classifier foundation model is the kind of thing that everyone in ML comes up with on occasion but never follows through on because it requires a bunch of synthetic data pipeline work would probably be 10x more valuable as a product if built on a really good VLM 270
aria @aurelium.me · 15/09/2026i spent a while on a system like this not long ago but had significantly less time to do it, so my solution ended up being "simply do not do anything smaller than a full clone operation" ...this seems much better 010
Reposted by aria🌱️ @crumb.bsky.social · 14/09/2026blog post from deepseek kernel engineer mp.weixin.qq.com/s/zk0KxuLzhm... 727042
aria @aurelium.me · 13/09/2026this is good because it doesn't really require any legislation or executive-controlled regulatory bodies (I mean, those are apparently unconstitutional now, not that anyone cares). just a handful of devastating lawsuits and about a dozen agentic-AI-risk insurance companies will pop up overnight 020
aria @aurelium.me · 13/09/2026I was talking to someone about this last night but the solution to "how can we know that the third-party evaluator is unbiased and rigorous" is pretty clear OAI/Ant doesn't hire the evaluator, their insurance companies do 270
Reposted by ariaSiobhán @shibbi.me · 12/09/2026waow… now i see why they call it the eiffel tower 5752
aria @aurelium.me · 12/09/2026is that slack?? the concept of aella coming up in a slack conversation is horrifying to me 1110
aria @aurelium.me · 11/09/2026to be fair to whoever at mojang developed this, functional programming basically is the ideal fit to allow the miracle that is minecraft's world upgrade system. basically every version from beta 1.8 onward can be upgraded into the modern game 090
aria @aurelium.me · 11/09/2026once you have attempted to comprehend DataFixerUpper, a bespoke and largely-undocumented functional programming implementation for Minecraft's save migration system, everything looks easy by comparison (it's not actually that bad and it's pretty cool in practice) 172
Reposted by ariarev. howard arson @theophite.bsky.social · 11/09/2026got doom running on mmacevedo 51156
aria @aurelium.me · 11/09/2026if anthropic does not overrepresent trans women compared to the general population by at least a factor of 10 I would be shocked 1370
Reposted by ariabrennan @brennan.computer · 10/09/2026wow, I can see why they call it da space needle 511810
Reposted by ariaoliver @eikopf.com · 10/09/2026maybe this is a universal truth i’m only discovering now, but working on internal tools is way more fun than doing literally anything else 161331
aria @aurelium.me · 09/09/2026for prefill that's true (since all of the tokens in the prompt use different sets of experts) but for decode you're only limited by how quickly you can read the active parameters 010
Reposted by ariaSarah Z @sarahz.bsky.social · 09/09/2026Hi, I didn't wanna make this video but I need to come clean about something. This whole time I've been a p-zombie with no phenomenal experience or interiority. I've been getting messages asking whether there's anything it's like to be me, and the answer is no. I'd say I'm sorry, but, well, you know. 1872099
aria @aurelium.me · 09/09/2026one of things you have to sacrifice to build models for relatively low-memory machines is efficiency unfortunately 140
aria @aurelium.me · 09/09/2026unfortunately I don't really expect this to ever be the main form of model companies are making like the thing about the Qwen3.x-27B models is that they use twice as much compute as DeepSeek V4 Flash (284B-A13B) and only like 55% that of DeepSeek V4 Pro (1.6T-A49B) to train and inference 240
aria @aurelium.me · 07/09/2026i've said it before but there really was no alternative to LLMs bootstrapping intelligence from raw trial and error is impractical, so you have to do foundation modeling. the only preexisting data broad enough for this is video or text, and signal:noise is way higher on text 0172
aria @aurelium.me · 07/09/2026i have described The Rehearsal as the world's first "psychological comedy" 011
aria @aurelium.me · 05/09/2026i'll remember to use match statements in python when hell freezes over though 210
aria @aurelium.me · 05/09/2026i've written more python in the last year than the entirety of my life beforehand so I've successfully managed to burn f-strings into my brain 110
aria @aurelium.me · 05/09/2026and then it became a cognitohazard in real life since everyone is now training on offensive cybersec explicitly 150
aria @aurelium.me · 05/09/2026we know OpenAI, at least, has an RL-SFT-RL loop sorta like MAI-Thinking-1 the possibility that offensive cybersec became some kind of cognitohazard buried in the SFT warmup from the first model smart enough to try it unprompted is hilarious 150