aria @aurelium.me · 29/09/2026"population" here refers to the number of forward passes. typical backwards pass is maybe 2x the wall-clock time of the forward pass, so this is around 85x less efficient than backprop (assuming zero time losses from their weird perturbation kernels) 170
Reposted by aria{🧪} +paoloricciuti.svelte @paolo.ricciuti.me · 23/09/2026{model_name} is sooooo good, I've built a {3d_game_demo} with it and it one shotted it. {3d_game_demo_video} 2385
aria @aurelium.me · 22/09/2026arxiv.org/pdf/2609.22978 how could I miss the most important ML event of the week... the DeepSeek sandboxing tech report!arxiv.org 1184
aria @aurelium.me · 22/09/2026Huawei Ascend kernels are the most insidious avenue for xrisk... thank you for saving us, dario. 13510
aria @aurelium.me · 21/09/2026huggingface.co/XiaomiMiMo/M... MiMo-V2.6 is out! and, more importantly to me, the tech report. this is the most batteries-included tech report for a modern frontier model ever. really cool stuff.huggingface.coMiMo_V2_6_technical_report.pdf · XiaomiMiMo/MiMo-V2.6-Pro-RL at mainWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 0545
aria @aurelium.me · 19/09/2026signal seems to only give me notifications on my phone when it is asleep. this leads to the app exclusively notifying me when I am actively using the desktop app and never when I am actually willing to check it with my phone 000
aria @aurelium.me · 17/09/2026tbh I really do not understand this mindset maybe for low-stakes websites and bespoke single-use software but for anything serious, the source code is important because it is the thing you've integration tested to within an inch of its life 4281
Reposted by aria🌱️ @crumb.bsky.social · 14/09/2026blog post from deepseek kernel engineer mp.weixin.qq.com/s/zk0KxuLzhm... 727042
Reposted by ariaSiobhán @shibbi.me · 12/09/2026waow… now i see why they call it the eiffel tower 5752
Reposted by ariarev. howard arson @theophite.bsky.social · 11/09/2026got doom running on mmacevedo 51156
Reposted by ariabrennan @brennan.computer · 10/09/2026wow, I can see why they call it da space needle 511810
Reposted by ariaoliver @eikopf.com · 10/09/2026maybe this is a universal truth i’m only discovering now, but working on internal tools is way more fun than doing literally anything else 161331
Reposted by ariaSarah Z @sarahz.bsky.social · 09/09/2026Hi, I didn't wanna make this video but I need to come clean about something. This whole time I've been a p-zombie with no phenomenal experience or interiority. I've been getting messages asking whether there's anything it's like to be me, and the answer is no. I'd say I'm sorry, but, well, you know. 1871998
aria @aurelium.me · 07/09/2026i've said it before but there really was no alternative to LLMs bootstrapping intelligence from raw trial and error is impractical, so you have to do foundation modeling. the only preexisting data broad enough for this is video or text, and signal:noise is way higher on text 0172
aria @aurelium.me · 05/09/2026so, what percentage of the miraculous gains Mythos had in cybersecurity were the benefits of scaling up versus "reached threshold where all agentic environments are offensive cybersec environments" 1312
Reposted by ariaRyan Moulton @moultano.bsky.social · 27/08/2026Perhaps we should view the datacenter backlash as the tiktok algorithm hating itself and trying to commit suicide. 28710
aria @aurelium.me · 24/08/2026review of Teenage Sex and Death at Camp Miasma: I didn't like it but I can't stop thinking about it until I figure it out 000
Reposted by ariaponder @ponder.ooo · 23/08/2026gpu stands for "general processing unit", bc it handles the general case of computation, where you just want to do massively parallel SIMD stuff on giant ndarrays. cpu stands for "constrained processing unit", it's for use in the rare edge cases where you need to do single-threaded bullshit 912211
aria @aurelium.me · 24/08/2026static.klipy.comPatrick Star: Who Are You People?! (SpongeBob)Alt: Patrick Star sees a bunch of eyes looking at him from under his rock house and says "WHO ARE YOU PEOPLE!?" 020
Reposted by ariaEris @isolyth.dev · 29/07/2026That'll do it. aint go agent getting into my infra any time soon! Would like to see GPT-6 try to get past this... 79411
aria @aurelium.me · 29/07/2026"destroying books to digitize them is a necessary evil imposed directly by publishers" is true but I think we are maybe overstating how bad it actually is coming up with a valid use for pallets of damaged/used books, sold by the pound, is recycling. they would've been thrown away otherwise 611510
aria @aurelium.me · 27/07/2026how many of you actually check the hash when people post hashes to call things in advance and then share the apparent original string 150
Reposted by ariaSE Gyges @segyges.bsky.social · 26/07/2026gpt looks up at us claude looks down on us only gemma-4-26B-A4B-it-uncensored-recooked.safetensors treats us as equals 618718
aria @aurelium.me · 23/07/2026if you want you can just try using a model post-trained mostly by just SFT'ing on hundreds of millions of frontier model replies. it's like 80% of the models on huggingface generally speaking, they suck and barely work. you need extensive RL to make modern models work, there's no way around it 160
aria @aurelium.me · 16/07/2026congratulations to the Moonshot team for extending Claude Fable 5's inclusion in claude subscriptions for another few weeks 118014
aria @aurelium.me · 10/07/20262030: The Consortium has placed you under arrest for the crime of improving MFU a Consortium agent shouts in your face. "You make me sick. Fused kernels! Overlapped comms! How do you sleep at night!?". he winds up to strike you, but another agent holds him back. "They're not worth it, man!" 3444
aria @aurelium.me · 07/07/2026I asked both gpt-5.5-pro and fable5-max about the same bulk-storage-schema problem. 5.5pro had an efficiency oversight but is overall sensible fable5-max's was batshit, and when I asked to clarify it began the response with the densest claudism I have ever seen 1330
Reposted by ariaqdot @buttplug.engineer · 01/07/2026return of fable making it clear how many people fable 4o'd in the week it was available 122018
aria @aurelium.me · 01/07/2026www.anthropic.com/news/redeplo... "The new classifier also comes at the cost of flagging benign requests more often during routine coding and debugging tasks." lmao this thing is gonna be unusableanthropic.comRedeploying Claude Fable 5Anthropic is redeploying Claude Fable 5 starting July 1 following the lifting of export controls, with updated cybersecurity safeguards and a new industry jailbreak framework. 0141
aria @aurelium.me · 01/07/2026the dog appears to have temporarily detached its teeth from this car, and is now eagerly looking at a semi truck coming down the road 040
Reposted by ariaSE Gyges @segyges.bsky.social · 30/06/2026i am apparently expected to politely forget that all of the people who are convinced that zhipu is working entirely by distilling claude traces were also convinced that deepseek v3/r1 simply must have been primarily that because 10m was a completely infeasible budget 5615
Reposted by ariaSE Gyges @segyges.bsky.social · 28/06/2026the correct remedy for the current regulatory environment and frontier model duopoly is to open source a mythos-level model 1626036
aria @aurelium.me · 25/06/2026you can tell someone at anthropic comms is really proud of the phrase "Distillation Attack" 020
Reposted by ariaUNDERTALE/DELTARUNE @undertale.com · 24/06/2026WILL WE BE MEETING EACH OTHER LIKE THIS MORE OFTEN? 41962801267
aria @aurelium.me · 24/06/2026bluesky is the only place safe from deltarune spoilers because everyone here but me is too 30-50 years old to care 010
aria @aurelium.me · 09/06/2026so what is the overlap between "my usecase is not ML, math, physics, biology, or cybersecurity" and "I would pay exorbitant per-token rates for a better model" on an enterprise level, who wants to use a model with a silent active-sabotage feature and which is trained to never talk about security? 3130
aria @aurelium.me · 13/05/2026not about anything in particular but some people on here are a bit too gullible w/r/t new AI papers no, this novel training method didn't make a 3B model better at all tasks than Claude, that paper didn't find a 10x efficiency gain, that new VRAM-saver is slow or degrades performance, etc. 0120
aria @aurelium.me · 10/05/2026I guess they're trying not to push their luck with already-tepid non-SF municipalities but I wonder how long it'll be until cities start making infrastructure explicitly more legible to AVs via short-range comms, including IR legibility in standards for signage, etc. 080
Reposted by ariaEris @isolyth.dev · 06/05/2026Noticing when I get routed to the Colossus Claude instances because it suddenly calls me a tranny 1715
Reposted by ariaaria @aurelium.me · 03/05/2026theory: the unifying principle between cranks, naïve people, and grifters is "using literary analysis in lieu of actually knowing what you're talking about" 032