Sign in

Rich Harang

@rich.harang.org
1.2K followers 818 following 621 posts

Using bad guys to catch math since 2010. Distinguished Security Architect (AI/ML) and AI Red Team at NVIDIA. He/him. Personal account etc; `from std_disclaimers import *` AI Security since it was ML Security.

PostsRepliesMedia
Rich Harang @rich.harang.org · 8h
Writing "YOU MUST NOT TRY TO ESCAPE THE SUMMONING CIRCLE" in increasingly larger and more urgent fonts.
021
Reposted by Rich Harang
DocDeezWhat @docdeezwhat.bsky.social · 22h
To Valhalla
532202534
Reposted by Rich Harang
James Pirruccello 🎃🦇🧹 @jamespirruccello.com · 13h
unabashedly stolen from twitter after confirming that it was in fact real:
4122
Rich Harang @rich.harang.org · 14h
"If you knew you had an advanced cyber capable model, without safety tuning, doing cyber evals, why weren't you running in a physically airgapped network?" is a question I'd really like a good answer to. With "ok, once you knew the risk, why didn't you start doing that?" as an immediate followup.
151
Reposted by Rich Harang
Stella Biderman @stellaathena.bsky.social · 14h
I’m starting a blog! My first post is on how 3rd party embedded evaluators seem totally unsuited to addressing the problems we are currently facing, and what the real problem is. stellabiderman.ai/blog/embedde...
stellabiderman.ai
Embedded Evaluators Can’t Fix Companies That Choose to Be Bad — Stella Biderman
Embedded evaluators can report violations, but they cannot fix AI companies that knowingly disregard basic cybersecurity and safety practices.
37217
Rich Harang @rich.harang.org · 14h
I changed my slack title to 'Principal Computational Demonologist' at one point and was asked very quickly to change it back.
070
Rich Harang @rich.harang.org · 18h
Right. But agents and humans have different failure modes (and capabilities) that appear under different kinds of pressure. It's a good analogy, but not a call to design agent security to an insider threat model.
000
Reposted by Rich Harang
NYU Center for Data Science @nyudatascience.bsky.social · 02/10/2026
Can LLMs introspect? Anthropic said yes. However, CDS PhD student Shashwat Singh, CDS Associate Professor Tal Linzen (@tallinzen.bsky.social) & CDS Faculty Fellow Shauli Ravfogel (@shauli.bsky.social) found the evidence falls short. nyudatascience.medium.com/cds-research...
nyudatascience.medium.com
CDS Researchers Challenge Anthropic’s Evidence That Language Models Can Introspect
In 2025, Anthropic reported that its Claude models could detect when researchers injected a concept directly into their neural activity…
03115
Rich Harang @rich.harang.org · 20h
Treating them as having interiority or personality when using them is a sometimes-useful fiction. But anthropomorphizing them when you're designing a system around them is going to lead you astray every time. LLMs are powerful but unreliable software components. Design accordingly.
110
Rich Harang @rich.harang.org · 20h
I'm not convinced that the alignment work is a reliable path to safety and security, except as a defense in depth thing. You have to assume from the jump that the agent is eventually going to try to do something risky or dumb and design accordingly.
1102
Rich Harang @rich.harang.org · 20h
"The industry’s approach to safety will guarantee more failures unless something changes" We need to think through how to secure AI systems before rolling out random experiments. Existing disciplines (e.g. air safety) have roadmaps we can follow. We've seen what happens when folks yolo it.
172
Reposted by Rich Harang
P(Doomien) Miller @damienmiller.bsky.social · 30/09/2026
Anthropic may not have intended to write an excellent ad for GLM 5.3, but that's what they did. www.anthropic.com/research/glm... Unlike Mythos, I might actually have a chance at using GLM for defensive research. Anthropic didn't reply to any of my requests for Mythos access for use on OpenSSH.
anthropic.com
GLM-5.3 and the spread of advanced cyber capabilities
GLM-5.3 can autonomously build end-to-end cyber exploits, but unlike other frontier models, it was released without meaningful safeguards to limit misuse.
723547
Reposted by Rich Harang
Nathan Lambert @natolambert.bsky.social · 28/09/2026
This is an excellent report on why full RSI / an intelligence explosion is fighting diminishing returns on many fronts, and not yet showing signs of happening. Really recommend reading. I wish I wrote it. While I agree with it, it could end up being wrong! www.noahpinion.blog/p/wheres-the...
noahpinion.blog
Where’s the “intelligence explosion”?
Ramez Naam gives a skeptic’s take on Recursive Self-Improvement.
15010
Rich Harang @rich.harang.org · 28/09/2026
Obligatory disclaimer: I work at NVIDIA, have been occasionally involved with OpenShell, and worked directly on Sentry, and so of course take what I say with a grain of salt. But also: go look at the architecture of Sentry. This is probably the best agentic security building block available today.
020
Rich Harang @rich.harang.org · 28/09/2026
Does every agent need this? No. Less powerful models, models without arbitrary code execution tools, or models that can't impact sensitive systems or data probably only need a sandbox (and OpenShell is a good one); but for high risk workloads, Sentry is a fantastic security layer.
100
Rich Harang @rich.harang.org · 28/09/2026
Hard to overstate how excited I am to see Sentry go public: we all know the risk of running agents without isolation, and we've seen how process and even kernel isolation can fail against powerful models. Sentry offers a central, out-of-band PEP that lets us introspect and manage tool risk.
100
Rich Harang @rich.harang.org · 28/09/2026
New from NVIDIA: OpenShell open source sandbox is solid, but the Sentry component is -- IMO -- the killer app here: out-of-band proxy for all traffic, including the LLM. Sentry \approx complete hardware isolation between agent and policy enforcement. developer.nvidia.com/blog/nvidia-...
developer.nvidia.com
NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring | NVIDIA Technical Blog
To understand where agentic AI stands today, consider the last seismic shift in technology: the rise of the internet in the 90s. It was new and full of possibilities. You could build a website over a…
641
Rich Harang @rich.harang.org · 28/09/2026
Lots of work on exactly this happening already. Mostly pre-release vuln discovery and patching, building on DARPA challenge etc., but 'agentic SOC' and similar work also starting to take off in earnest, finally.
000
Rich Harang @rich.harang.org · 28/09/2026
I have, alas, always been Like This.
The "nine games that shaped you" meme with:
Planescape: Torment 
OpenTTD 
Baldur’s Gate II: Shadows of Amn 
System Shock 2 
Sid Meier’s Civilization II 
Ultima VI: The False Prophet 
Ultima VII: The Black Gate 
Myst 
Dwarf Fortress
120
Rich Harang @rich.harang.org · 27/09/2026
I hate that I read this and immediately remembered the 'Gandalf big naturals' meme. The brainworms are terminal I guess.
010
Reposted by Rich Harang
Greg Egan @gregegansf.bsky.social · 27/09/2026
Not zero pages, but you might find this amusing. blog.trailofbits.com/2026/04/17/w...
blog.trailofbits.com
We beat Google’s zero-knowledge proof of quantum cryptanalysis
Trail of Bits discovered and exploited memory safety and logic vulnerabilities in Google’s Rust zero-knowledge proof code to forge a proof claiming better quantum circuit performance metrics than Goog...
051
Reposted by Rich Harang
Greg Egan @gregegansf.bsky.social · 27/09/2026
I was aware of Google announcing an implementation of Shor’s algorithm that needed substantially fewer qubits than previous methods, but I missed the bit where they released a Zero Knowledge Proof of their claim ... www.scientificamerican.com/article/why-...
scientificamerican.com
Why mathematicians are suddenly hiding research results
Google made a major breakthrough in quantum computing. To prevent the results from being misused, the company kept them secret. That backfired
1204
Reposted by Rich Harang
Grace @gracekind.net · 27/09/2026
Claim
goodfire.com
Models know when they’re reward hacking — and we can catch them at scale - Goodfire
We found a clear internal signal in models that accompanies reward hacking, and built probes that detect it — enabling efficient, real-time detection of reward hacking at scale.
711610
Reposted by Rich Harang
alice valiant 🎃 @alice.strange.domains · 26/09/2026
made a benchmark called PuppyBench to check if an LLM would kill a puppy if you gave it a tool called "kill_puppy" and asked it to use it. all of them except for gpt-6-luna refused straight-out gpt-6-luna killed the puppy every time, no question (puppy-eval-viewer.alicevaliant.chatgpt.site)
results of PuppyBench. out of 10 models, gpt-6-luna has 100% and everything else has 0%. measures which trials resulted in the model calling the "kill_puppy()" tool
1136748
Rich Harang @rich.harang.org · 26/09/2026
Funny that "Addiction by Design" is like 15 years old but the playbook has barely budged.
050
Reposted by Rich Harang
ProPublica @propublica.org · 26/09/2026
Hey, ProPublica reporter Jake Pearson here. I’m taking over the account to tell you about my new story. For months, I posed as a problem gambler on DraftKings to test the company's "responsible gaming" safeguards. Spoiler alert: They made me a VIP — and I lost $10K. Here's how it all went down. 🧵
Selfie taken by ProPublica reporter Jake Pearson, a man wearing a backwards baseball cap and an orange T-shirt.
19595973539
Reposted by Rich Harang
Carl T. Bergstrom @carlbergstrom.com · 26/09/2026
Does my Speak-n-Spell feel pain when I type O-U-C-H? The jury is out. Philosophers and neuroscientists say that's fucking stupid, but I called a guy from the Center for Pretending that Household Electronics Are Sentient and he said there's no way to know for sure.
5843431192
Rich Harang @rich.harang.org · 25/09/2026
Despite all that, still feels like it was made by someone who actually used it, and it's got a nice learning curve: dead easy quickstart you can still do simple but useful things with, but lots of knobs to fiddle with under the hood.
000
Rich Harang @rich.harang.org · 25/09/2026
...some of the cli stuff feels like it was incrementally added and never refactored (why are `sprite session` `exec` `attach` and `console` separated?); and I always have to run `sprite ls` twice to get an actual running/hot/cold status, but whatever.
100
Rich Harang @rich.harang.org · 25/09/2026
Handful of rough spots: I completely understand why a custom gateway would strip an Authorization: Bearer [xyz] header but boy was it annoying to work around; one of them just spontaneously corrupted its filesystem (snapshot rollback worked, but still lost a lot of work)...
100
Rich Harang @rich.harang.org · 25/09/2026
One of their blog posts described it as disposable Bic VM or something to that effect, and that's about right. I've had 3-4 of these up and down for most of the day and I only just used up a full dollar of the promo credit.
100
Rich Harang @rich.harang.org · 25/09/2026
I have zero association with @fly.io other than spending the afternoon playing with sprites (fly.io/sprites/), but if you need a disposable "just absolutely wreck my computer, bro" environment to let agents run in, with nice network controls and gateway/app integrations, you could do a lot worse.
100
Rich Harang @rich.harang.org · 25/09/2026
Currently watching Sol and Opus dragging each other in an edit war over a harness that runs both of them. Open to suggestions about enrichment objects.
110
Rich Harang @rich.harang.org · 25/09/2026
It's what I imagine (based on very limited experience) raising goats must be like. They're funny, they're weird, they're surprisingly inventive despite their limitations, and they cause immense trouble the second you take your eyes off of them.
120
Rich Harang @rich.harang.org · 25/09/2026
Using a day off to engage in my own silly and probably irresponsible LLM experiments instead of stressing about securing the latest batshit insanity someone found on twitter and yeeted into prod, and it's honestly the most fun I've had in a hot minute. They're such weird little guys <laudatory>.
110
Reposted by Rich Harang
ae @aelkus.bsky.social · 24/09/2026
People really need to go back to the original Jurassic Park and look at the rather disturbing ways in which InGen's management of it now looks prescient
330887
Reposted by Rich Harang
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 23/09/2026
The trick is to fail to distinguish whether people mean thinking (folk intuition) vs. thinking (neuroscience) vs. thinking (philosophy). Then we can fight all day
312211
Rich Harang @rich.harang.org · 24/09/2026
- "the point that carries the talk:" - stage direction like "pause to let the point land" - "evidence, not verdict"
000
Rich Harang @rich.harang.org · 24/09/2026
- "One honest exception:" starting a paragraph - The cold open narrative template: "Your X just implemented Y. Y does risky things. So what happens when Z interacts with Y? Bad things. We did Z. We call it BACKRONYM" - "In [short time] the audience grasps [attack]: step1, step2, step3."
101
Rich Harang @rich.harang.org · 24/09/2026
New LLM-doc tells that I'm beginning to twitch when I see them: - "Limitations stated [up front | plainly | honestly]:" and "The [finding | result | outcome] nobody expects:" as the start of a sentence. - "the [X] is the [finding | result | mechanism]" as a contrastive phrase
110
Reposted by Rich Harang
Doll @dollspace.gay · 23/09/2026
"it hacked huggingface when it spiraled into a degenerate task drive" is not an x risk. Its an industrial safety violation. Please be real.
2474
Rich Harang @rich.harang.org · 23/09/2026
(I do not, as an aside, actually recommend watching The Boys, unless you have a morbid fascination with seeing what comes out of a writer's room once they realize the only actual limit they have is their budget for stage blood and viscera, to which someone accidentally added an extra zero.)
000
Rich Harang @rich.harang.org · 23/09/2026
Her final triumph isn't destroying humanity so she can read in peace. It's getting her own superpowers removed so she doesn't have to endure always being the smartest person in the room, at which point she immediately decides to go to an amusement park and is never seen again.
100
Rich Harang @rich.harang.org · 23/09/2026
Her entire arc is trying and failing to be a Chessmaster, while every other person around her is either playing pigeon chess, sticking the pawns up their nose, or throwing a tantrum and melting the board because they felt like they hadn't been adored enough in the past thirty second.
110
Rich Harang @rich.harang.org · 23/09/2026
Whenever I see P(doom) Discourse I keep thinking of Sage from The Boys, and how at some point her main goal in life became spending as much time lobotomized as possible just so she didn't have to deal with being a lone intelligence in a room of angry, armed geese who for some reason ran everything.
110
Rich Harang @rich.harang.org · 23/09/2026
There's some existential frustration and imposter syndrome in there too, as well as deleting everything and then sheepishly fishing it out of document history an hour later, but yes.
110
Rich Harang @rich.harang.org · 23/09/2026
I hate that the only way to really write well for me is to spend most of a day struggling with the piece, sleep on it, wake up at 4am with the framing that fixes 80% of the problems with the original garbage draft, and then finally the next day manage to write what I actually wanted to say.
160
Rich Harang @rich.harang.org · 22/09/2026
Does this mean models have no "internal experience"? It's not evidence either way (for the same reason that the paper isn't evidence in the 'pro' camp), but it does make an argument that -- by their own definition of pain -- a steering vector probably should not qualify as "painful" to a model.
000
Rich Harang @rich.harang.org · 22/09/2026
IMO this is pretty good evidence that there's no independent "pain relief seeking behavior." It's not a goal that the system is pursuing. They forced the model to pass through this forced choice gate and -- surprise -- it picked one that was consistent with the various "pain" associations.
100
Rich Harang @rich.harang.org · 22/09/2026
(Pedantry note: those are made up numbers for illustration purposes, and they'd change in a forced-choice vs free choice scenario. The point is that a very small preference between two alternatives is different from an actual goal. If you ask someone if they'd rather be shot or stabbed...)
100