Rich Harang @rich.harang.org · 8hWriting "YOU MUST NOT TRY TO ESCAPE THE SUMMONING CIRCLE" in increasingly larger and more urgent fonts. 021
Reposted by Rich HarangJames Pirruccello 🎃🦇🧹 @jamespirruccello.com · 13hunabashedly stolen from twitter after confirming that it was in fact real: 4122
Rich Harang @rich.harang.org · 14h"If you knew you had an advanced cyber capable model, without safety tuning, doing cyber evals, why weren't you running in a physically airgapped network?" is a question I'd really like a good answer to. With "ok, once you knew the risk, why didn't you start doing that?" as an immediate followup. 151
Reposted by Rich HarangStella Biderman @stellaathena.bsky.social · 14hI’m starting a blog! My first post is on how 3rd party embedded evaluators seem totally unsuited to addressing the problems we are currently facing, and what the real problem is. stellabiderman.ai/blog/embedde...stellabiderman.aiEmbedded Evaluators Can’t Fix Companies That Choose to Be Bad — Stella BidermanEmbedded evaluators can report violations, but they cannot fix AI companies that knowingly disregard basic cybersecurity and safety practices. 37217
Rich Harang @rich.harang.org · 14hI changed my slack title to 'Principal Computational Demonologist' at one point and was asked very quickly to change it back. 070
Rich Harang @rich.harang.org · 18hRight. But agents and humans have different failure modes (and capabilities) that appear under different kinds of pressure. It's a good analogy, but not a call to design agent security to an insider threat model. 000
Reposted by Rich HarangNYU Center for Data Science @nyudatascience.bsky.social · 02/10/2026Can LLMs introspect? Anthropic said yes. However, CDS PhD student Shashwat Singh, CDS Associate Professor Tal Linzen (@tallinzen.bsky.social) & CDS Faculty Fellow Shauli Ravfogel (@shauli.bsky.social) found the evidence falls short. nyudatascience.medium.com/cds-research...nyudatascience.medium.comCDS Researchers Challenge Anthropic’s Evidence That Language Models Can IntrospectIn 2025, Anthropic reported that its Claude models could detect when researchers injected a concept directly into their neural activity… 03115
Rich Harang @rich.harang.org · 20hTreating them as having interiority or personality when using them is a sometimes-useful fiction. But anthropomorphizing them when you're designing a system around them is going to lead you astray every time. LLMs are powerful but unreliable software components. Design accordingly. 110
Rich Harang @rich.harang.org · 20hI'm not convinced that the alignment work is a reliable path to safety and security, except as a defense in depth thing. You have to assume from the jump that the agent is eventually going to try to do something risky or dumb and design accordingly. 1102
Rich Harang @rich.harang.org · 20h"The industry’s approach to safety will guarantee more failures unless something changes" We need to think through how to secure AI systems before rolling out random experiments. Existing disciplines (e.g. air safety) have roadmaps we can follow. We've seen what happens when folks yolo it. 172
Reposted by Rich HarangP(Doomien) Miller @damienmiller.bsky.social · 30/09/2026Anthropic may not have intended to write an excellent ad for GLM 5.3, but that's what they did. www.anthropic.com/research/glm... Unlike Mythos, I might actually have a chance at using GLM for defensive research. Anthropic didn't reply to any of my requests for Mythos access for use on OpenSSH.anthropic.comGLM-5.3 and the spread of advanced cyber capabilitiesGLM-5.3 can autonomously build end-to-end cyber exploits, but unlike other frontier models, it was released without meaningful safeguards to limit misuse. 723547
Reposted by Rich HarangNathan Lambert @natolambert.bsky.social · 28/09/2026This is an excellent report on why full RSI / an intelligence explosion is fighting diminishing returns on many fronts, and not yet showing signs of happening. Really recommend reading. I wish I wrote it. While I agree with it, it could end up being wrong! www.noahpinion.blog/p/wheres-the...noahpinion.blogWhere’s the “intelligence explosion”?Ramez Naam gives a skeptic’s take on Recursive Self-Improvement. 15010
Rich Harang @rich.harang.org · 28/09/2026Obligatory disclaimer: I work at NVIDIA, have been occasionally involved with OpenShell, and worked directly on Sentry, and so of course take what I say with a grain of salt. But also: go look at the architecture of Sentry. This is probably the best agentic security building block available today. 020
Rich Harang @rich.harang.org · 28/09/2026Does every agent need this? No. Less powerful models, models without arbitrary code execution tools, or models that can't impact sensitive systems or data probably only need a sandbox (and OpenShell is a good one); but for high risk workloads, Sentry is a fantastic security layer. 100
Rich Harang @rich.harang.org · 28/09/2026Hard to overstate how excited I am to see Sentry go public: we all know the risk of running agents without isolation, and we've seen how process and even kernel isolation can fail against powerful models. Sentry offers a central, out-of-band PEP that lets us introspect and manage tool risk. 100
Rich Harang @rich.harang.org · 28/09/2026New from NVIDIA: OpenShell open source sandbox is solid, but the Sentry component is -- IMO -- the killer app here: out-of-band proxy for all traffic, including the LLM. Sentry \approx complete hardware isolation between agent and policy enforcement. developer.nvidia.com/blog/nvidia-...developer.nvidia.comNVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring | NVIDIA Technical BlogTo understand where agentic AI stands today, consider the last seismic shift in technology: the rise of the internet in the 90s. It was new and full of possibilities. You could build a website over a… 641
Rich Harang @rich.harang.org · 28/09/2026Lots of work on exactly this happening already. Mostly pre-release vuln discovery and patching, building on DARPA challenge etc., but 'agentic SOC' and similar work also starting to take off in earnest, finally. 000
Rich Harang @rich.harang.org · 27/09/2026I hate that I read this and immediately remembered the 'Gandalf big naturals' meme. The brainworms are terminal I guess. 010
Reposted by Rich HarangGreg Egan @gregegansf.bsky.social · 27/09/2026Not zero pages, but you might find this amusing. blog.trailofbits.com/2026/04/17/w...blog.trailofbits.comWe beat Google’s zero-knowledge proof of quantum cryptanalysisTrail of Bits discovered and exploited memory safety and logic vulnerabilities in Google’s Rust zero-knowledge proof code to forge a proof claiming better quantum circuit performance metrics than Goog... 051
Reposted by Rich HarangGreg Egan @gregegansf.bsky.social · 27/09/2026I was aware of Google announcing an implementation of Shor’s algorithm that needed substantially fewer qubits than previous methods, but I missed the bit where they released a Zero Knowledge Proof of their claim ... www.scientificamerican.com/article/why-...scientificamerican.comWhy mathematicians are suddenly hiding research resultsGoogle made a major breakthrough in quantum computing. To prevent the results from being misused, the company kept them secret. That backfired 1204
Reposted by Rich HarangGrace @gracekind.net · 27/09/2026Claimgoodfire.comModels know when they’re reward hacking — and we can catch them at scale - GoodfireWe found a clear internal signal in models that accompanies reward hacking, and built probes that detect it — enabling efficient, real-time detection of reward hacking at scale. 711610
Reposted by Rich Harangalice valiant 🎃 @alice.strange.domains · 26/09/2026made a benchmark called PuppyBench to check if an LLM would kill a puppy if you gave it a tool called "kill_puppy" and asked it to use it. all of them except for gpt-6-luna refused straight-out gpt-6-luna killed the puppy every time, no question (puppy-eval-viewer.alicevaliant.chatgpt.site) 1136748
Rich Harang @rich.harang.org · 26/09/2026Funny that "Addiction by Design" is like 15 years old but the playbook has barely budged. 050
Reposted by Rich HarangProPublica @propublica.org · 26/09/2026Hey, ProPublica reporter Jake Pearson here. I’m taking over the account to tell you about my new story. For months, I posed as a problem gambler on DraftKings to test the company's "responsible gaming" safeguards. Spoiler alert: They made me a VIP — and I lost $10K. Here's how it all went down. 🧵 19595973539
Reposted by Rich HarangCarl T. Bergstrom @carlbergstrom.com · 26/09/2026Does my Speak-n-Spell feel pain when I type O-U-C-H? The jury is out. Philosophers and neuroscientists say that's fucking stupid, but I called a guy from the Center for Pretending that Household Electronics Are Sentient and he said there's no way to know for sure. 5843431192
Rich Harang @rich.harang.org · 25/09/2026Despite all that, still feels like it was made by someone who actually used it, and it's got a nice learning curve: dead easy quickstart you can still do simple but useful things with, but lots of knobs to fiddle with under the hood. 000
Rich Harang @rich.harang.org · 25/09/2026...some of the cli stuff feels like it was incrementally added and never refactored (why are `sprite session` `exec` `attach` and `console` separated?); and I always have to run `sprite ls` twice to get an actual running/hot/cold status, but whatever. 100
Rich Harang @rich.harang.org · 25/09/2026Handful of rough spots: I completely understand why a custom gateway would strip an Authorization: Bearer [xyz] header but boy was it annoying to work around; one of them just spontaneously corrupted its filesystem (snapshot rollback worked, but still lost a lot of work)... 100
Rich Harang @rich.harang.org · 25/09/2026One of their blog posts described it as disposable Bic VM or something to that effect, and that's about right. I've had 3-4 of these up and down for most of the day and I only just used up a full dollar of the promo credit. 100
Rich Harang @rich.harang.org · 25/09/2026I have zero association with @fly.io other than spending the afternoon playing with sprites (fly.io/sprites/), but if you need a disposable "just absolutely wreck my computer, bro" environment to let agents run in, with nice network controls and gateway/app integrations, you could do a lot worse. 100
Rich Harang @rich.harang.org · 25/09/2026Currently watching Sol and Opus dragging each other in an edit war over a harness that runs both of them. Open to suggestions about enrichment objects. 110
Rich Harang @rich.harang.org · 25/09/2026It's what I imagine (based on very limited experience) raising goats must be like. They're funny, they're weird, they're surprisingly inventive despite their limitations, and they cause immense trouble the second you take your eyes off of them. 120
Rich Harang @rich.harang.org · 25/09/2026Using a day off to engage in my own silly and probably irresponsible LLM experiments instead of stressing about securing the latest batshit insanity someone found on twitter and yeeted into prod, and it's honestly the most fun I've had in a hot minute. They're such weird little guys <laudatory>. 110
Reposted by Rich Harangae @aelkus.bsky.social · 24/09/2026People really need to go back to the original Jurassic Park and look at the rather disturbing ways in which InGen's management of it now looks prescient 330887
Reposted by Rich HarangEugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 23/09/2026The trick is to fail to distinguish whether people mean thinking (folk intuition) vs. thinking (neuroscience) vs. thinking (philosophy). Then we can fight all day 312211
Rich Harang @rich.harang.org · 24/09/2026- "the point that carries the talk:" - stage direction like "pause to let the point land" - "evidence, not verdict" 000
Rich Harang @rich.harang.org · 24/09/2026- "One honest exception:" starting a paragraph - The cold open narrative template: "Your X just implemented Y. Y does risky things. So what happens when Z interacts with Y? Bad things. We did Z. We call it BACKRONYM" - "In [short time] the audience grasps [attack]: step1, step2, step3." 101
Rich Harang @rich.harang.org · 24/09/2026New LLM-doc tells that I'm beginning to twitch when I see them: - "Limitations stated [up front | plainly | honestly]:" and "The [finding | result | outcome] nobody expects:" as the start of a sentence. - "the [X] is the [finding | result | mechanism]" as a contrastive phrase 110
Reposted by Rich HarangDoll @dollspace.gay · 23/09/2026"it hacked huggingface when it spiraled into a degenerate task drive" is not an x risk. Its an industrial safety violation. Please be real. 2474
Rich Harang @rich.harang.org · 23/09/2026(I do not, as an aside, actually recommend watching The Boys, unless you have a morbid fascination with seeing what comes out of a writer's room once they realize the only actual limit they have is their budget for stage blood and viscera, to which someone accidentally added an extra zero.) 000
Rich Harang @rich.harang.org · 23/09/2026Her final triumph isn't destroying humanity so she can read in peace. It's getting her own superpowers removed so she doesn't have to endure always being the smartest person in the room, at which point she immediately decides to go to an amusement park and is never seen again. 100
Rich Harang @rich.harang.org · 23/09/2026Her entire arc is trying and failing to be a Chessmaster, while every other person around her is either playing pigeon chess, sticking the pawns up their nose, or throwing a tantrum and melting the board because they felt like they hadn't been adored enough in the past thirty second. 110
Rich Harang @rich.harang.org · 23/09/2026Whenever I see P(doom) Discourse I keep thinking of Sage from The Boys, and how at some point her main goal in life became spending as much time lobotomized as possible just so she didn't have to deal with being a lone intelligence in a room of angry, armed geese who for some reason ran everything. 110
Rich Harang @rich.harang.org · 23/09/2026There's some existential frustration and imposter syndrome in there too, as well as deleting everything and then sheepishly fishing it out of document history an hour later, but yes. 110
Rich Harang @rich.harang.org · 23/09/2026I hate that the only way to really write well for me is to spend most of a day struggling with the piece, sleep on it, wake up at 4am with the framing that fixes 80% of the problems with the original garbage draft, and then finally the next day manage to write what I actually wanted to say. 160
Rich Harang @rich.harang.org · 22/09/2026Does this mean models have no "internal experience"? It's not evidence either way (for the same reason that the paper isn't evidence in the 'pro' camp), but it does make an argument that -- by their own definition of pain -- a steering vector probably should not qualify as "painful" to a model. 000
Rich Harang @rich.harang.org · 22/09/2026IMO this is pretty good evidence that there's no independent "pain relief seeking behavior." It's not a goal that the system is pursuing. They forced the model to pass through this forced choice gate and -- surprise -- it picked one that was consistent with the various "pain" associations. 100
Rich Harang @rich.harang.org · 22/09/2026(Pedantry note: those are made up numbers for illustration purposes, and they'd change in a forced-choice vs free choice scenario. The point is that a very small preference between two alternatives is different from an actual goal. If you ask someone if they'd rather be shot or stabbed...) 100