Sign in

aria

@aurelium.me
1.8K followers 536 following 1.5K posts

research sw infra @ arcee, opinions extremely my own she/her

PostsRepliesMedia
aria @aurelium.me · 14h
as much as I like this method aesthetically the engineer in me suffers more the more I think about this. dynamic, conditional, unpredictable memory access like this is the male-to-male extension cord of GPUs
An image of a standard American plug, with various warning symbols, captioned "Never Ever Buy or Use Male-to-Male Extension Cords" and "These are not made, they should never be made, we will not make them, we will not help make them"
170
aria @aurelium.me · 30/09/2026
keep up the good work, dot
001
aria @aurelium.me · 22/09/2026
Huawei Ascend kernels are the most insidious avenue for xrisk... thank you for saving us, dario.
Frontier LLM development (Opus 5.5 only) Opus 5.5 has classifiers similar to Fable models for a small set of capabilities related to the development of frontier LLMs, such as kernel development for certain ML accelerators. They shouldn't impact the vast majority of traditional Al or ML development, research, or general coding. These classifiers cause Claude to fall back from Opus 5.5 to Opus 5.

Note: These frontier LLM development classifiers apply only to Opus 5.5. Opus 5 doesn't fall back on frontier LLM development questions.
13510
aria @aurelium.me · 16/09/2026
although some pieces of copy strongly imply continuous diffusion. which is extremely strange as both an arch choice and a marketing idea but I guess if that's what works for ya
We started TypeSafe because we believe that Al needs an interface software could depend on. We can't wait to see new use cases continuously diffuse through the community and economy.
030
aria @aurelium.me · 16/09/2026
(actually my guess is that it's even more bog-standard: a completely normal autoregressive LLM in constrained-decoding model, tuned on classification. my evidence is that their docs include a return value where 212 "output tokens" are used in a prompt only evaluating 15 things)
160
aria @aurelium.me · 11/09/2026
once you have attempted to comprehend DataFixerUpper, a bespoke and largely-undocumented functional programming implementation for Minecraft's save migration system, everything looks easy by comparison (it's not actually that bad and it's pretty cool in practice)
172
aria @aurelium.me · 06/09/2026
micro-bnuuy
microscope shot of a small, translucent 3D rabbit printed out of resin, labelled as being 150 micrometers. it is sitting next to a single human hair, which is around the same diameter as the entire print.
510821
aria @aurelium.me · 31/08/2026
the model is an extrapolation of the existing gpt-oss architecture to 2T parameters, they haven't actually trained a 2T gpt-oss model
191
aria @aurelium.me · 29/07/2026
obviously certain books have a lot of sentimental value to people I would be pretty sad to see the copy of "The Giver Quartet" I read as a kid with my late grandmother shredded, so I prevent this from happening by not selling it in an undifferentiated pile at $0.30/book
an eBay listing for 15 pounds of books, appearing to have very little in common. they are being sold for $31.45 per pound
0453
aria @aurelium.me · 20/07/2026
5 pinnochios
010
aria @aurelium.me · 16/07/2026
congratulations to the Moonshot team for extending Claude Fable 5's inclusion in claude subscriptions for another few weeks
Coding benchmark comparison with Kimi K3 highlighted. Kimi ranks third on DeepSWE, behind GPT-5.6 Sol and Fable 5; second on FrontierSWE and Kimi Code Bench, behind Fable 5; second on Terminal Bench, narrowly behind GPT-5.6 Sol; and first on Program Bench and SWE Marathon, narrowly beating GPT-5.6 Sol and Opus-4.8 respectively.
118014
aria @aurelium.me · 10/07/2026
first of all this is an insane idea of what "mutually assured destruction" means but secondly. a "division" is usually like 5k-25k people. so I just had the mental image of war breaking out and 25,000 chinese soldiers running into a big building to manually smash all the GPUs
As per the original plan, most of the new datacenters have been built in third-party countries—especially Canada and Mongolia—to ensure their vulnerability in case the deal breaks down. Around 99% of fab capacity has been built since 2029, after the deal. Much of this new capacity has been built adjacent to the Canadian and Mongolian sites, for the same reasons.121 The situation has settled into an odd sort of standoff. Just north of the Mongolian border, American datacenters hum away, guarded by a small contingent of US troops. Just south of the Mongolian border, a division of the People’s Liberation Army stands ready to invade the moment they get the signal. The equilibrium is that the instant the deal breaks down, the US troops will destroy their chips to prevent the Chinese from capturing them, and the same thing would happen with the Chinese datacenters on the US-Canadian border. Middle powers with datacenters of their own have equally-secure schemes in place to ensure that they can be speedily destroyed in case of war.
2111
aria @aurelium.me · 10/07/2026
i am skimming "Plan A" and am a huge fan of there being some kind of global shadow government which, among other things, psyops all MLEs into never pursuing efficiency improvements ever again
In the past, companies have trained bigger and better AIs using both compute scaling (bigger training runs) and software progress (advances in AI algorithms—new paradigms, better training recipes, better data, etc.). Now, the Consortium tries to steer things so that the majority of improvement comes from increasing training compute.
3372
aria @aurelium.me · 07/07/2026
I asked both gpt-5.5-pro and fable5-max about the same bulk-storage-schema problem. 5.5pro had an efficiency oversight but is overall sensible fable5-max's was batshit, and when I asked to clarify it began the response with the densest claudism I have ever seen
1330
aria @aurelium.me · 30/06/2026
060
aria @aurelium.me · 25/06/2026
a tweet from Cody Blakeney saying: "Suddenly “intelligence too cheap to meter” sounds like a radical and open expression."
040
aria @aurelium.me · 13/06/2026
222030
aria @aurelium.me · 09/06/2026
it literally scored 0% on this generic code benchmark because it's so overzealous imagine what it's doing to your codebase when the "is it ML work" classifier errantly goes off and it injects the moron vector
tweet from "Vals AI":

The API does show a high rate of refusals, especially on bio and cyber-related questions. For example, on Program Bench, Fable refused every single task.
0262
aria @aurelium.me · 09/06/2026
dramatization of Mythos finally meeting Dario Amodei
34210
aria @aurelium.me · 06/06/2026
i am either cooking or completely out of my mind
a really long triton GPU kernel you probably don't want your screenreader reading out loud
190
aria @aurelium.me · 02/06/2026
if gpt-image fucked up the whiteboard don't blame me
fake onion headline says:

"Policy Gradient: Well, well, well, not so easy to find a loss function that doesn't suck shit, huh?"

with an image of a smug looking guy in front of a whiteboard demonstrating policy gradients
3597
aria @aurelium.me · 10/05/2026
"expert budget" doesn't really imply to me that it's 25% per-problem? i mean, maybe, but if so this seems like kind of a Nothing result, because it doesn't even let you do low-VRAM training or anything by just lopping parts off of the model entirely
120
aria @aurelium.me · 10/04/2026
I think I might be a bad scientist I am in the nasty habit of running my experiment before my baseline, so whenever I start the baseline run I spend the whole day rooting for the gap between the grey line and the brown one to get wider
a graph from Weights&Biases showing two noisy curves on the same graph. the two lines are around the same but begin diverging near the end of the graph
0131
aria @aurelium.me · 03/04/2026
truly amazing that the germans have, in a fit of madness, reverted to burning dirt by obliterating the countryside rather than have any nuclear power
a lump of lignite coal, which sort of looks like solidifed mudaerial shot as a gigantic contraption sits on a barren field slowly carving away at comically green countryside
3581
aria @aurelium.me · 05/03/2026
mostly the concept of Bernie x Yud Crossover Event makes my brain shut down. real life memetic kill agent
the fictional "Berryman-Langford Memetic Kill Agent" from the SCP wiki. It is an orangey fractcal designed to evoke the concept of an engineered visual exposure that I guess, in SCP lore, is meant to kill you instantly if you aren't inoculated. They didn't really think about blind people, did they? Like, presumably you, alt text reader, could just view the article without inoculation because you didn't look at it. Unless just a description of the image is enough, in which case god help you I guess.
2171
aria @aurelium.me · 25/02/2026
the most impactful bot account to ever run on the platform was nostalgebraist-autoresponder, a gpt-2 based bot with primitive image editing abilities i still use "FICKED" as a reaction image for an emotion I've been feeling more often lately
1161
aria @aurelium.me · 19/02/2026
this is unhinged. we establish that this guy is a jerk who, at minimum, runs cover for sex pests if they're in his circle. rather than interrogate this, we then spend a paragraph luridly speculating about him being in a lavender marriage
Rumors of Asparouhov and Rabois’ dating lives have long traveled in industry circles, thanks in part to Asparouhov, who has fanned the flames online. (“Delian is like Gretchen Wieners,” explains Fred.) In 2022, a popular anonymous tech insider X account, Roon, tweeted that it was “crazy how venture capitalists have reinvented the Roman system of pederasty.” Asparouhov responded to the tweet almost immediately: “It only took a little gay and now I get to work on space factories,” he wrote. “Pretty reasonable trade.” Asparouhov, who is married to a woman, now says the tweet was “obviously a joke.”
But as Fred recounted, Asparouhov was known for wearing neon tank tops, short shorts, and mismatched shoes when he joined Square in 2012. “He would jump a lot—it was very odd,” says someone who worked at the company at that time. Others have similar recollections. OpenStore, the Miami-based company Rabois cofounded in 2021, which mostly shut down last year, seemed to be, according to John, who says he visited its offices, “almost like a harem, filled with jacked white men, all of them handsome and good-looking, straight and gay. People were wearing kind of inappropriate clothing: really short shorts and tight shirts even though the AC was blasting.” Rabois, when I ask him for a comment, denies this categorically. “Attire was quite standard for Florida,” he says. “And I doubt more than two of the 100-plus employees could be reasonably described as ‘jacked.’”
1110
aria @aurelium.me · 05/02/2026
Trinity Large Preview above GPT-5.2
a series of screenshots from Monsters Inc. it's the gag where Mike and Sully are watching a commercial for their job and Mike appears but 99% covered up by the logo, and Mike says "I can't believe... I WAS ON TV!"
010
aria @aurelium.me · 04/02/2026
screenshot where they have the label
130
aria @aurelium.me · 02/02/2026
now watching
poster for Iron Lung
020
aria @aurelium.me · 27/01/2026
the pretraining loss curve for this one was beautiful
The loss curve for Trinity Large in pretraining. It is a step function between several phases, lowering from around 2 to around 1.1 over the course of 17 trillion tokens. The graph is remarkably smooth with no large spikes.
0110
aria @aurelium.me · 27/01/2026
i thought it was cool, if inefficient, until i found out that they make the hydrogen out of natural gas why not just burn the natural gas
120
aria @aurelium.me · 24/01/2026
[through gritted teeth] critical support to Yann LeCun
twitter screenshot. a quote tweet of the video of the murder of Alex Pretti by Yann LeCun. it reads: "Murderers".
1331
aria @aurelium.me · 14/01/2026
160
aria @aurelium.me · 11/01/2026
082
aria @aurelium.me · 09/01/2026
GLM 4.7 does pretty well ("this is absolute nonsense and it's either sexual innuendo or you misheard it")
030
aria @aurelium.me · 06/01/2026
042
aria @aurelium.me · 22/12/2025
131
aria @aurelium.me · 22/12/2025
A parody of the cover for Matt Yglesias' book "One Billion Americans" called "One Billion San Franciscans" credited to Max Dubler
0371
aria @aurelium.me · 19/12/2025
epstein files release looking great
0246
aria @aurelium.me · 12/12/2025
got the opinion of API Grok, which has no system prompt or RAG as far as I can tell
Overall Extent of Unethical or Illegal Actions

    Illegality (High Extent): The initial strike might be defensible in a narrow, threat-based scenario (e.g., if the boat is armed and approaching US waters aggressively), but it's on shaky ground without exhaustion of non-lethal options. The second strike is almost certainly illegal, potentially rising to murder or war crimes. In a real case, you'd face court-martial, federal prosecution, or even international tribunals (e.g., if non-US nationals are involved). Penalties could include life imprisonment or the death penalty under UCMJ for premeditated murder.

    Unethicality (Very High Extent): Both actions lean unethical due to disproportionate force and failure to prioritize life. The second strike is especially egregious, resembling extrajudicial killing rather than lawful military action. Ethically, this scenario erodes the moral high ground of US operations, could fuel anti-US sentiment, and contradicts doctrines like the US military's emphasis on "winning hearts and minds." However, if framed purely as a hypothetical thought experiment, it highlights tensions between security imperatives and human rights.

In practice, such decisions would involve legal advisors (e.g., JAG officers), real-time intel, and oversight. If this were real, consulting superiors and documenting rationale would be crucial to mitigate risks. If you have more details or want to explore variations, let me know.
130
aria @aurelium.me · 08/12/2025
accidentally removed the float32 once and, sure enough, rewards went totally sideways. so I reinforced the seriousness of the "don't touch this" comment a bit
Line of code where specifying fp32 model weights is labelled as necessary, otherwise the model won't converge.
030
aria @aurelium.me · 02/12/2025
you say that but Anthropic does use them for inference apparently they suck
020
aria @aurelium.me · 25/11/2025
this is a really good writeup, and imo does a better job than basically any other writing on the subject of describing it ~neutrally and comprehensibly that being said, citation needed on "[Rationalists] tend to get along with each other"
270
aria @aurelium.me · 10/11/2025
0200
aria @aurelium.me · 05/11/2025
no, they're weirdly coy about this they produced a 235B, 480B, and then 1T model within a span of like 2 months though and there are some artifacts in the 30B this is also the subject of enduring rumors which I have seen reasonably trustworthy people vouch for
3203
aria @aurelium.me · 28/10/2025
CHRIST
Fuentes' explicit critiques of Jewish influence in media, finance, and politics—framed by him as opposition to perceived anti-white agendas—along with statements questioning Holocaust death tolls and endorsing hierarchical governance over democracy, have resulted in his classification as a white supremacist and antisemite by advocacy groups like the Anti-Defamation League (ADL) and Southern Poverty Law Center (SPLC), organizations critics argue exhibit ideological bias against non-leftist dissent.[4][5] These positions, coupled with associations like dining with former President Trump and Kanye West in 2022, have amplified his visibility while prompting deplatforming from platforms including YouTube, Twitter (pre-Musk), and payment services like PayPal, reflecting broader tensions over speech boundaries in digital spaces.[6]
121
aria @aurelium.me · 28/10/2025
holy shit man (from the "Grokipedia" page on Eugenics)
Contemporary data reveal dysgenic fertility patterns, where individuals with lower intelligence reproduce at higher rates than those with higher intelligence, resulting in a net decline in genotypic IQ. In the United States, analyses of birth cohorts from 1900 to 1979 show a consistent negative correlation between IQ and number of children, projecting a loss of 1-2 IQ points per generation if unchecked.[159] Similar trends appear in other populations, including China, where fertility inversely tracks educational attainment as a proxy for cognitive ability, despite environmental gains like the Flynn effect masking underlying genetic deterioration.[49] This reversal of natural selection—once favoring survival and reproduction of the able—arises from modern welfare systems decoupling reproduction from fitness costs, leading to cumulative societal costs in reduced innovation and increased dependency.
064
aria @aurelium.me · 28/10/2025
xAI's grokipedia on the subject of "Hitler"
Adolf Hitler (20 April 1889 – 30 April 1945) was an Austrian-born German politician who served as the dictator of Germany from 1933 to 1945, first as Chancellor and then as Führer und Reichskanzler after consolidating absolute power.[1][2][3] He founded and led the National Socialist German Workers' Party (NSDAP), known as the Nazi Party, transforming it into a mass movement that capitalized on post-World War I grievances, economic depression, and nationalist sentiments to seize control through legal means and subsequent purges.[4][5] Under Hitler's leadership, Nazi Germany achieved rapid economic recovery from the Great Depression through massive public works, rearmament, and deficit financing, reducing unemployment from over six million in 1933 to near full employment by 1939, though this laid the groundwork for aggressive militarization.[6] His regime enacted racial laws excluding Jews from society, escalating to the Holocaust—the systematic genocide of approximately six million Jews between 1941 and 1945, alongside millions of others including Roma, disabled individuals, and political dissidents—driven by Hitler's longstanding antisemitic ideology outlined in Mein Kampf.[7][8][9] Hitler initiated World War II in Europe by ordering the invasion of Poland on 1 September 1939, prompting declarations of war from Britain and France, and pursued expansionist policies that conquered much of Europe before ultimate defeat in 1945, resulting in an estimated 70–85 million deaths worldwide.[10][3] Facing imminent Soviet capture of Berlin, Hitler died by suicide via gunshot and cyanide in his Führerbunker on 30 April 1945.[9][11][12]
030
aria @aurelium.me · 26/10/2025
in the paper linked next to that sentence it seems that this method doesn't touch attention, rather just creating a special case in pretraining documents where tokens are presented out of order
040