Sign in

aria

@aurelium.me
1.8K followers 535 following 1.5K posts

research sw infra @ arcee, opinions extremely my own she/her

PostsRepliesMedia
aria @aurelium.me · 30/09/2026
keep up the good work, dot
001
aria @aurelium.me · 29/09/2026
unfortunately all of the cool things about training nondifferentiable models also make them really terrible for running on GPUs (branching, mainly) so the OOMs of per-token compute reduction you'd need to match backprop with a zeroth-order method are kind of unlikely to manifest
030
aria @aurelium.me · 29/09/2026
imo applying zeroth-order optimization to transformers is missing the point completely why would you approximate gradients via sampling (likely >50x slower) when you can compute them directly. to the extent that ZO is useful, it is for training architectures that don't work with backprop
140
aria @aurelium.me · 29/09/2026
(also, I'd be very surprised if the paper reveals that "core assumptions in optimization research are completely wrong". ZO-type methods are mathematically principled and this has been known for a while. the practical version of this in transformers exists, it's called RL)
040
aria @aurelium.me · 29/09/2026
you might think that in order to calculate the pseudograds you need to store 256x random perturbations of the weights, which would be huge. but in theory you could embed deterministic PRNG into your forward pass kernels, so all you need to store is the seeds my guess is that this is what they did
140
aria @aurelium.me · 29/09/2026
zeroth-order optimization in this context is approximating gradients via sampling. you apply a large number of random perturbations to the weights and use the effect these perturbations have on the loss to calculate pseudogradients
130
aria @aurelium.me · 29/09/2026
"population" here refers to the number of forward passes. typical backwards pass is maybe 2x the wall-clock time of the forward pass, so this is around 85x less efficient than backprop (assuming zero time losses from their weird perturbation kernels)
170
aria @aurelium.me · 28/09/2026
the chinese room is basically a verbal magic trick. surely something this mundane, a man reading a book, cannot together make up a system that "understands Chinese" even if the man does not and the book is inert without him ...except the book is 1000 OOMs larger than the observable universe
110
Reposted by aria
{🧪} +paoloricciuti.svelte @paolo.ricciuti.me · 23/09/2026
{model_name} is sooooo good, I've built a {3d_game_demo} with it and it one shotted it. {3d_game_demo_video}
2385
aria @aurelium.me · 22/09/2026
arxiv.org/pdf/2609.22978 how could I miss the most important ML event of the week... the DeepSeek sandboxing tech report!
arxiv.org
1184
aria @aurelium.me · 22/09/2026
Huawei Ascend kernels are the most insidious avenue for xrisk... thank you for saving us, dario.
Frontier LLM development (Opus 5.5 only) Opus 5.5 has classifiers similar to Fable models for a small set of capabilities related to the development of frontier LLMs, such as kernel development for certain ML accelerators. They shouldn't impact the vast majority of traditional Al or ML development, research, or general coding. These classifiers cause Claude to fall back from Opus 5.5 to Opus 5.

Note: These frontier LLM development classifiers apply only to Opus 5.5. Opus 5 doesn't fall back on frontier LLM development questions.
13510
aria @aurelium.me · 21/09/2026
huggingface.co/XiaomiMiMo/M... MiMo-V2.6 is out! and, more importantly to me, the tech report. this is the most batteries-included tech report for a modern frontier model ever. really cool stuff.
huggingface.co
MiMo_V2_6_technical_report.pdf · XiaomiMiMo/MiMo-V2.6-Pro-RL at main
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
0545
aria @aurelium.me · 19/09/2026
signal seems to only give me notifications on my phone when it is asleep. this leads to the app exclusively notifying me when I am actively using the desktop app and never when I am actually willing to check it with my phone
000
aria @aurelium.me · 17/09/2026
the main thing that would be difficult to replicate (and they say as much themselves) would be the data pipeline. getting good enough coverage for generalist classification tasks is a lot of work
130
aria @aurelium.me · 17/09/2026
text diffusion (generally) isn't really anything like the complicated algorithms used for image diffusion it's not a particularly complicated modification to modeling code or loss compared to a normal LM classifier, and even the diffusion aspect isn't really necessary for their performance
130
aria @aurelium.me · 17/09/2026
if you really want to take advantage of LLMs to write easy cross-platform software, use them to build some abstraction with ~no regard for "not being unpleasant to write" but extremely stringent requirements on correctness and portability
080
aria @aurelium.me · 17/09/2026
my brain is permanently fried from writing software in the most un-unit-testable domain possible (RL trainers, which combine GPU integration, multi-node parallelism, and coordination across thousands of remote processes) but I wouldn't want to build ~anything another person used like this
2110
aria @aurelium.me · 17/09/2026
tbh I really do not understand this mindset maybe for low-stakes websites and bespoke single-use software but for anything serious, the source code is important because it is the thing you've integration tested to within an inch of its life
4281
aria @aurelium.me · 16/09/2026
although some pieces of copy strongly imply continuous diffusion. which is extremely strange as both an arch choice and a marketing idea but I guess if that's what works for ya
We started TypeSafe because we believe that Al needs an interface software could depend on. We can't wait to see new use cases continuously diffuse through the community and economy.
030
aria @aurelium.me · 16/09/2026
(actually my guess is that it's even more bog-standard: a completely normal autoregressive LLM in constrained-decoding model, tuned on classification. my evidence is that their docs include a return value where 212 "output tokens" are used in a prompt only evaluating 15 things)
160
aria @aurelium.me · 16/09/2026
a lot of people are saying it's a block-diffusion model because of the parallel output but it could easily be a completely normal LM classifier running in parallel req 1: text + "Was the author angry?" req 1: text + "Was the author sad?" etc. use standard shared-prefix KV cache, costs ~the same
150
aria @aurelium.me · 16/09/2026
it's a neat idea, the idea of a classifier foundation model is the kind of thing that everyone in ML comes up with on occasion but never follows through on because it requires a bunch of synthetic data pipeline work would probably be 10x more valuable as a product if built on a really good VLM
270
aria @aurelium.me · 15/09/2026
i spent a while on a system like this not long ago but had significantly less time to do it, so my solution ended up being "simply do not do anything smaller than a full clone operation" ...this seems much better
010
Reposted by aria
🌱️ @crumb.bsky.social · 14/09/2026
blog post from deepseek kernel engineer mp.weixin.qq.com/s/zk0KxuLzhm...
727042
aria @aurelium.me · 13/09/2026
this is good because it doesn't really require any legislation or executive-controlled regulatory bodies (I mean, those are apparently unconstitutional now, not that anyone cares). just a handful of devastating lawsuits and about a dozen agentic-AI-risk insurance companies will pop up overnight
020
aria @aurelium.me · 13/09/2026
I was talking to someone about this last night but the solution to "how can we know that the third-party evaluator is unbiased and rigorous" is pretty clear OAI/Ant doesn't hire the evaluator, their insurance companies do
270
Reposted by aria
Siobhán @shibbi.me · 12/09/2026
waow… now i see why they call it the eiffel tower
da eiffel tower
5752
Reposted by aria
samhain ⎔ @personhood.removal.surgery · 12/09/2026
OpenAI:
Screenshot of a post by @pipkinpippa reading: “i follow a lot of odd communities and one of them is people who keep lemurs as pets

once a month someone comes along and is like ‘my lemur is suddenly viciously attacking everyone’

and every time the other lemur owners reply like”

Below is a screenshot of three Facebook comments with names and avatars censored: “Thats a lemur for you”; “That’s what owning a lemur is, they’re not pets unfortunately”; and “Now you know why lemurs are cheap.”
017917
aria @aurelium.me · 12/09/2026
is that slack?? the concept of aella coming up in a slack conversation is horrifying to me
1110
aria @aurelium.me · 12/09/2026
better than the alternative (EVE corps)
0130
aria @aurelium.me · 11/09/2026
to be fair to whoever at mojang developed this, functional programming basically is the ideal fit to allow the miracle that is minecraft's world upgrade system. basically every version from beta 1.8 onward can be upgraded into the modern game
090
aria @aurelium.me · 11/09/2026
once you have attempted to comprehend DataFixerUpper, a bespoke and largely-undocumented functional programming implementation for Minecraft's save migration system, everything looks easy by comparison (it's not actually that bad and it's pretty cool in practice)
172
Reposted by aria
rev. howard arson @theophite.bsky.social · 11/09/2026
got doom running on mmacevedo
51156
aria @aurelium.me · 11/09/2026
something like that
060
aria @aurelium.me · 11/09/2026
if anthropic does not overrepresent trans women compared to the general population by at least a factor of 10 I would be shocked
1370
Reposted by aria
brennan @brennan.computer · 10/09/2026
wow, I can see why they call it da space needle
paris, texas
511810
Reposted by aria
oliver @eikopf.com · 10/09/2026
maybe this is a universal truth i’m only discovering now, but working on internal tools is way more fun than doing literally anything else
161331
aria @aurelium.me · 09/09/2026
wdym?
010
aria @aurelium.me · 09/09/2026
for prefill that's true (since all of the tokens in the prompt use different sets of experts) but for decode you're only limited by how quickly you can read the active parameters
010
Reposted by aria
Sarah Z @sarahz.bsky.social · 09/09/2026
Hi, I didn't wanna make this video but I need to come clean about something. This whole time I've been a p-zombie with no phenomenal experience or interiority. I've been getting messages asking whether there's anything it's like to be me, and the answer is no. I'd say I'm sorry, but, well, you know.
1872099
aria @aurelium.me · 09/09/2026
one of things you have to sacrifice to build models for relatively low-memory machines is efficiency unfortunately
140
aria @aurelium.me · 09/09/2026
unfortunately I don't really expect this to ever be the main form of model companies are making like the thing about the Qwen3.x-27B models is that they use twice as much compute as DeepSeek V4 Flash (284B-A13B) and only like 55% that of DeepSeek V4 Pro (1.6T-A49B) to train and inference
240
aria @aurelium.me · 07/09/2026
i've said it before but there really was no alternative to LLMs bootstrapping intelligence from raw trial and error is impractical, so you have to do foundation modeling. the only preexisting data broad enough for this is video or text, and signal:noise is way higher on text
0172
aria @aurelium.me · 07/09/2026
i have described The Rehearsal as the world's first "psychological comedy"
011
aria @aurelium.me · 06/09/2026
how are you posting from the micro-verse i thought he shrank you
020
aria @aurelium.me · 06/09/2026
micro-bnuuy
microscope shot of a small, translucent 3D rabbit printed out of resin, labelled as being 150 micrometers. it is sitting next to a single human hair, which is around the same diameter as the entire print.
510821
aria @aurelium.me · 05/09/2026
i'll remember to use match statements in python when hell freezes over though
210
aria @aurelium.me · 05/09/2026
i've written more python in the last year than the entirety of my life beforehand so I've successfully managed to burn f-strings into my brain
110
aria @aurelium.me · 05/09/2026
and then it became a cognitohazard in real life since everyone is now training on offensive cybersec explicitly
150
aria @aurelium.me · 05/09/2026
we know OpenAI, at least, has an RL-SFT-RL loop sorta like MAI-Thinking-1 the possibility that offensive cybersec became some kind of cognitohazard buried in the SFT warmup from the first model smart enough to try it unprompted is hilarious
150