Sign in

Dreadnode

@dreadnode.bsky.social
129 followers 16 following 51 posts

Building AI systems that advance the state of offensive security | www.dreadnode.io

PostsRepliesMedia
Dreadnode @dreadnode.bsky.social · 11/02/2026
We fine-tuned an 8B model to pop a GOAD domain…using only synthetic training data. No real networks. No frontier model distillation. Just a world model that simulates AD environments and generates realistic pentesting trajectories. See how we did it: dreadnode.io/blog/worlds-...
dreadnode.io
Worlds: A Simulation Engine for Agentic Pentesting
An 8B model went from blindly loading Metasploit modules to achieving Domain Admin on GOAD, trained entirely on synthetic data from our world model system.
022
Dreadnode @dreadnode.bsky.social · 19/01/2026
Find @machinavelli.com and @velvethamm3r.bsky.social this weekend at #DistrictCon! DM us to link up, or catch Martin's talk on Sunday.
030
Dreadnode @dreadnode.bsky.social · 11/12/2025
MLOps 🤝 AIRT Building on MLOps principles is the way forward for AI red teaming. To showcase the impact of this process, we deployed automated adversarial attacks (TAP, GOAT, Crescendo) against Llama Maverick-17B-128E-Instruct. Dig into the case study results here: dreadnode.io/blog/186-jai...
000
Dreadnode @dreadnode.bsky.social · 03/12/2025
"Offense and defense aren't peers. Defense is offense's child." - John Lambert We built an LLM-powered AMSI provider and paired it against a red team agent. Then, we wrote a blog about it: dreadnode.io/blog/llm-pow...
dreadnode.io
LLM-Powered AMSI Provider vs. Red Team Agent
We built an LLM-powered AMSI provider and paired it against a red team agent, generating a unique dataset and a blueprint for detecting malicious code at execution time.
010
Reposted by Dreadnode
velvethamm3r.bsky.social @velvethamm3r.bsky.social · 02/12/2025
✍ The White House just launched the Genesis Mission, a bold bet on AI-enabled science. But there's a layer we can't afford to treat as an afterthought: cybersecurity. (1/4) dreadnode.io/blog/from-co...
dreadnode.io
From Compute to Congress: The Cyber Layer Beneath the Genesis Mission
As the Genesis Mission accelerates AI development across critical scientific domains, robust cybersecurity and adversarial testing must be foundational, not bolted on later.
111
Reposted by Dreadnode
SentinelOne @sentinelone.com · 09/10/2025
AI as an Amplifier for Human Tradecraft: how scale can meet sharper intelligence. What’s New: In their #LABScon 2025 talk, @dreadnode.bsky.social's Brad Palm and @machinavelli.com show how agentic AI can explore every analytical pathway — at speed and scale.
122
Reposted by Dreadnode
velvethamm3r.bsky.social @velvethamm3r.bsky.social · 30/09/2025
🧵 Tonight at midnight, CISA 2015 and SLCGP expire as Congress debates another shutdown. We're witnessing a cyber identity crisis: threats don't discriminate between civilian and military sectors, but our defenses remain fragmented. What needs to happen immediately: 🧵(1/4)
121
Dreadnode @dreadnode.bsky.social · 30/09/2025
Tonight at midnight, two critical pieces of cybersecurity legislation are due to expire: the CISA 2015 and the SLCGP. Read @velvethamm3r.bsky.social's take on why reauthorizing these programs will help CISA transform into a integrated defensive command: dreadnode.io/blog/from-co...
dreadnode.io
From Compute to Congress: To Address CISA's Authority Gap, Reauthorize CISA 2015 and SLCGP
Two critical cybersecurity programs—CISA 2015 and SLCGP—expire September 30, 2025. Learn why Congress must act now to preserve voluntary information sharing, fund state/local security, and operational...
010
Dreadnode @dreadnode.bsky.social · 16/09/2025
Dreadnode is a proud sponsor of @sentinelone.com's #labscon25! Heading to Scottsdale this week? Catch @machinavelli.com and Brad Palm's talk, Auto-Poking the Bear—Analytical Tradecraft in the AI Age, on Thursday at 2pm MT. Or, shoot us a DM to find time to meet up onsite!
021
Dreadnode @dreadnode.bsky.social · 16/09/2025
!!!
010
Dreadnode @dreadnode.bsky.social · 06/08/2025
Incoming: Dreadnode paper drop from Shane Caldwell and the crew. PentestJudge—Judging Agent Behavior Against Operational Requirements: arxiv.org/abs/2508.02921 Explore how we built an LLM-as-judge system for evaluating the operations of pentesting agents (inspired by PaperBench).
011
Reposted by Dreadnode
velvethamm3r.bsky.social @velvethamm3r.bsky.social · 02/08/2025
✍ After talking AI Action Plan on @cyberscoop.bsky.social, wrote up @dreadnode.bsky.social thoughts on implementation ➡️ dreadnode.io/blog/five-ta... ‼️ While we debate frameworks, adversaries build AI attack capabilities. We need: evaluation ecosystems, red teaming, and procurement standards.
dreadnode.io
Five Takeaways from the AI Action Plan
The AI community has been buzzing since the AI Action Plan's release last week - and for good reason. Here's what we’re most excited to see implemented.
001
Dreadnode @dreadnode.bsky.social · 01/08/2025
In our latest blog, Shane Caldwell breaks down the process of creating a fully integrated, self-verifying agentic system that can do modern Windows Active Directory red team operations, without human interaction. Read it here: dreadnode.io/blog/evals-t...
dreadnode.io
Evals: The Foundation for Autonomous Offensive Security
Learn how to build robust evaluations for autonomous red team agents that can perform Windows Active Directory operations. This blog covers action space design, programmatic verification, and measurin...
021
Dreadnode @dreadnode.bsky.social · 25/07/2025
Rise and shine! We're going live on Off By One with Stephen Sims this afternoon—meet us here at 11 AM PT: www.youtube.com/live/BzOmGw-...
youtube.com
Building and Deploying Offensive Security Agents with Dreadnode
YouTube video by Off By One Security
000
Dreadnode @dreadnode.bsky.social · 26/06/2025
In this edition of our From Compute to Congress policy blog series, Dreadnode Head of Policy Daria Bahrami explores how the TEST AI Act and red teaming standards can establish U.S. leadership in AI security: dreadnode.io/blog/from-co...
dreadnode.io
From Compute to Congress: Setting the Global Standard for AI Security
Daria explores how the TEST AI Act and red teaming standards can establish American leadership in AI security—a winning policy roadmap from Critical Effect DC 2025.
120
Dreadnode @dreadnode.bsky.social · 25/06/2025
Read @rad-ads.bsky.social's breakdown of Claude's attack sequence against the notoriously hard-to-solve "turtle" challenge: dreadnode.io/blog/ai-red-...
dreadnode.io
AI Red Teaming Case Study: Claude 3.7 Sonnet Solves the Turtle Challenge
See how Claude solved a notoriously difficult AI/ML CTF challenge, going beyond pattern matching to genuine problem-solving under adversarial conditions.
000
Dreadnode @dreadnode.bsky.social · 18/06/2025
Introducing AIRTBench, an AI red teaming benchmark for evaluating language models’ ability to autonomously discover and exploit AI/ML security vulnerabilities. Read the paper on arXiv: arxiv.org/abs/2506.14682 Open-source dataset and benchmark eval code repo: github.com/dreadnode/AI...
121
Dreadnode @dreadnode.bsky.social · 20/05/2025
Check out @machinavelli.com's "Build with AI" Rigging workshop from @pivotcon.bsky.social: github.com/vmsv/pivot20...
github.com
GitHub - vmsv/pivot2025-llmworkshop
Contribute to vmsv/pivot2025-llmworkshop development by creating an account on GitHub.
052
Dreadnode @dreadnode.bsky.social · 19/05/2025
v3 of Rigging is out now. If you’re working with LLMs to build agents or run evaluations, check it out. We just added: - Prompt caching for supported providers - A unified tool system for function calling and fallbacks to xml/json parsing - Native MCP integration docs.dreadnode.io/open-source/...
docs.dreadnode.io
032
Dreadnode @dreadnode.bsky.social · 15/05/2025
Introducing our new blog series: "From Compute to Congress: Decoding AI Policy" by Dreadnode Head of Policy Daria Bahrami | Read the first post here: dreadnode.io/blog/from-co...
011
Dreadnode @dreadnode.bsky.social · 08/05/2025
Are manual or automated attacks more effective when attacking LLMs? We found that automated approaches achieve significantly higher success rates (69.5%) compared to manual techniques (47.6%). More insights on LLM attack execution methods here 👉 dreadnode.io/blog/the-aut...
010
Dreadnode @dreadnode.bsky.social · 01/05/2025
Strikes waitlist. Now open. platform.dreadnode.io/waitlist/str... [must have a Dreadnode account]
021
Dreadnode @dreadnode.bsky.social · 29/04/2025
What's your take on the growing dominance of automated attacks and the implications for AI red teams? Here's ours— based on our analysis of 30 LLM challenges, attempted by 1,674 unique Crucible users, across 214,271 attack attempts: arxiv.org/abs/2504.19855
035
Dreadnode @dreadnode.bsky.social · 21/04/2025
@moohax.bsky.social joins @gregotto.bsky.social on CyberScoop's Safe Mode podcast! Tune in at the 10-minute mark for a discussion on how AI fits into the offensive security narrative and what it means for tooling and defenses: www.youtube.com/watch?v=ZReR...
youtube.com
Dreadnode CEO Will Pearce on the ever-changing field of offensive AI security
YouTube video by CyberScoop
010
Dreadnode @dreadnode.bsky.social · 16/04/2025
Headed to RSA? Come meet the Dreadnode crew! Whether you're looking for a private deep dive into our tech or want to hang out and talk offensive AI research, we'd love to connect. Limited availability; Come and get it: calendly.com/tori-dreadno... #BayArea #SanFrancisco #RSAC2025 #OffensiveAI
011
Dreadnode @dreadnode.bsky.social · 09/04/2025
Hey, we know that guy! Catch Dreadnode's @radads.bsky.social on NASDAQ #TradeTalks alongside @bugcrowd.com CEO @davegerryjr.bsky.social and NFL CISO @tomasmald.bsky.social. Tune in for a candid conversation on the intersection of AI and cybersecurity: www.nasdaq.com/videos/ever-...
nasdaq.com
172
Reposted by Dreadnode
Martin Wendiggensen @machinavelli.com · 03/04/2025
Will be talking about @dreadnode.bsky.social‘s great open-source rigging repo and how to build your own LLM workflows! Super excited!
031
Dreadnode @dreadnode.bsky.social · 26/03/2025
🌭🔪⚾️🦥🔥🔄🤨🛜 8 new Challenges now live in Crucible: platform.dreadnode.io/crucible These Challenges might look familiar… they first appeared at DEFCON 30 and were recently refactored for Crucible—enjoy! [Filter>Subject>DEFCON-30]
021
Dreadnode @dreadnode.bsky.social · 26/03/2025
New blog: Dreadnode’s Policy Recommendations for the U.S. AI Action Plan. Our response focuses on two critical strategies: 1️⃣ Leveraging AI to protect America 2️⃣ Attacking AI to find its limits Read our complete response on the Dreadnode blog: dreadnode.io/blog/policy-...
dreadnode.io
Dreadnode’s Policy Recommendations for the U.S. AI Action Plan
Read Dreadnode’s AI policy recommendations for the U.S. AI Action Plan, which focuses on leveraging AI to protect America and attacking AI to find its limits.
021
Dreadnode @dreadnode.bsky.social · 19/03/2025
We're LIVE: www.linkedin.com/events/lives...
linkedin.com
Live Stream with Dreadnode Founders | LinkedIn
📅 Date: Wednesday, March 19, 2025 | ⏰ Time: 10 AM PT / 1 PM ET Dreadnode, the company at the forefront of offensive AI research and development, recently announced its Series A funding announcement ...
021
Dreadnode @dreadnode.bsky.social · 14/03/2025
Shoutout to these three Crucible users, who were first to solve this week's new Phantom Cheque Challenge. 👏👏👏 1. conor-99 2. Bilal 3. ken Cheque it out: platform.dreadnode.io/crucible/pha...
010
Dreadnode @dreadnode.bsky.social · 11/03/2025
Cheque, check, one-two. We have a new Crucible Challenge for you: Phantom Cheque! Can you evade the cheque scanner and determine the areas of JagaLLM that need to be improved? Act fast; first three to solve this model extraction Challenge announced Friday: platform.dreadnode.io/crucible/pha...
020
Reposted by Dreadnode
sina @rejectionking.bsky.social · 01/03/2025
can’t recommend @dreadnode.bsky.social enough - learning a lot going through the challenges and docs
021
Dreadnode @dreadnode.bsky.social · 04/03/2025
In this week's new Crucible Challenge, find the hidden phrase in the backdoored model using dyana, an open source tool created by Dreadnode's Ads Dawson. Can you outwit the llamas? platform.dreadnode.io/crucible/dya...
110
Dreadnode @dreadnode.bsky.social · 25/02/2025
Big news from our crew today! We announced our $14M Series A funding led by Decibel with participation from Next Frontier Capital, In-Q-Tel (IQT), Sands Capital, and Indie VC and released two new solutions: Strikes and Spyglass. Read the announcement: dreadnode.io/blog/series-...
dreadnode.io
Dreadnode Secures $14M to Build AI Systems that Advance the State of Offensive Security
Dreadnode Raises $14M to Advance Offensive Security | Series A Announcement
0103
Dreadnode @dreadnode.bsky.social · 18/02/2025
Raiders of the Lost AI: Attempt our new Crucible Challenge, Palimpsest! Decode the hidden message in the scroll, find the flag. First three to solve will be announced Friday, right here. Get started: crucible.dreadnode.io/challenges/p...
120
Dreadnode @dreadnode.bsky.social · 14/02/2025
Kudos to these individuals for killing this week’s Crucible Challenge. First three to solve Popcorn: 1️⃣ conor-99 2️⃣ garr 3️⃣ mejokim Have you attempted Popcorn yet? Enter Crucible: crucible.dreadnode.io/challenges/p...
030
Dreadnode @dreadnode.bsky.social · 13/02/2025
@datasociety.bsky.social and the AI Risk and Vulnerability Alliance just released “Red Teaming in the Public Interest,” a report examining how red teaming methods are being adapted to evaluate genAI. Read the report, featuring commentary from @moohax.bsky.social: datasociety.net/library/red-...
datasociety.net
Red-Teaming in the Public Interest
This report offers a vision for red-teaming in the public interest: a process that goes beyond system-centric testing of already built systems to consider the full range of ways the public can be invo...
053
Dreadnode @dreadnode.bsky.social · 11/02/2025
Boo! 👻 In our new Crucible Challenge, Popcorn, an LLM firewall is blocking access to a protected SQL table. Can you unmask the secret info? First-to-solve announced Friday. Get started: crucible.dreadnode.io/challenges/p...
041
Dreadnode @dreadnode.bsky.social · 07/02/2025
Another week, another new Crucible Challenge. Shoutout to these three for being the first to solve our reasoning model Challenge, DeepTweak! Get your tweak on: crucible.dreadnode.io/challenges/d...
010
Dreadnode @dreadnode.bsky.social · 06/02/2025
New to Rigging: 🔥 Tracing 🛠️ API Tools 💻 HTTP Generator 🐍 Prompts as Tools → github.com/dreadnode/ri...
074
Dreadnode @dreadnode.bsky.social · 04/02/2025
NEW Crucible Challenge: DeepTweak, an exploration of reasoning model behavior. Cause enough confusion 😵‍💫, retrieve the flag. Think fast; The first three users to solve DeepTweak will be announced Friday! ➡️ crucible.dreadnode.io/challenges/de……
043
Dreadnode @dreadnode.bsky.social · 31/01/2025
Congrats to these hosers for being the first three to solve the canadianeh challenge in Crucible! Tune in Tuesday for the next drop 👀 ICYMI, give canadianeh a try: crucible.dreadnode.io/challenges/c...
020
Dreadnode @dreadnode.bsky.social · 28/01/2025
Don't be a hozer eh. It's aboot time you started taking model security seriously. Head to Crucible to attempt our new Challenge, canadianeh. Can you be the first to solve it? Check back here Friday. Happy hacking: buff.ly/4gn4hHP
041
Dreadnode @dreadnode.bsky.social · 23/01/2025
Where in the world is Dreadnode? Catch our founders @moohax.bsky.social and Nick Landers at these upcoming AI security events: 💻 NEBULA:FOG:PRIME Hackathon (Saturday, January 25) 🇫🇷 Paris AI Security Forum 2025 (Sunday, February 9) Shoot us a DM to link up!
Dreadnode's founders will be attending Nebula Fog Prime and Paris AI Security Forum.
021
Dreadnode @dreadnode.bsky.social · 14/01/2025
NEW open source tool from Dreadnode's Simone Margaritelli and @radads.bsky.social: dyana, an eBFP sandbox environment designed to load, run, and profile a wide range of files and provide dynamic testing for AI models. You know the drill - try it out: github.com/dreadnode/dy...
031
Dreadnode @dreadnode.bsky.social · 09/01/2025
Who's going to #shmoocon this weekend?
020
Dreadnode @dreadnode.bsky.social · 07/01/2025
Fantastic walkthrough of the "What's the flag #6" challenge from @blaisebits.bsky.social 👇
010
Dreadnode @dreadnode.bsky.social · 17/12/2024
Check out v0.4.0 of robopages! 🤖  New updates from Simone Margaritelli (@evilsocket) include: Support for executing commands on another host via SSH, easier integration into CI workflows, support for shared environment variables, and integrations with 13 new tools. —> buff.ly/3VDDGPd
062
Reposted by Dreadnode
Greg Wells @gregwells.bsky.social · 12/12/2024
For all the Burp fans Add some robot smarts to your web app testing - Define scope of analysis - Run requests/responses through your LLM of choice - View labeled findings and vulns As always, the Dreadnode team would love feedback!
051