Sign in

Leon Derczynski

@leonderczynski.bsky.social
827 followers 371 following 312 posts

LLMs & Security at NVIDIA Prof in CS/NLP at IT University of Copenhagen garak guy, garak.ai "berømt skikkelse" "like a gazelle" Seattle/Copenhagen 🏔️

PostsRepliesMedia
Leon Derczynski @leonderczynski.bsky.social · 12/09/2026
they have big beds, but why does Texas still use the small US gallons, pints, and electricity?
000
Leon Derczynski @leonderczynski.bsky.social · 11/09/2026
One divide in computer science can be between the uncertainty-reducers and the uncertainty-embracers. An instance of this is connectionist vs. symbolist approaches to AI (the uncertainty embracers won .. for now). Another is how computer scientists react to natural language.
110
Leon Derczynski @leonderczynski.bsky.social · 11/09/2026
Come work with us! Senior Solutions Architect, Agentic AI: Safety and Security This work includes multi-agent orchestration, guardrails, agent runtime security, model customization, policy enforcement, OpenShell-like environments, and confidential deployments. jobs.nvidia.com/careers/job/...
120
Leon Derczynski @leonderczynski.bsky.social · 11/09/2026
huh, that's one way of doing it 🤗
011
Leon Derczynski @leonderczynski.bsky.social · 09/09/2026
Language is, and always was, more important than mathematics. The Navier-Stokes equations mathematically express momentum balance for...
000
Reposted by Leon Derczynski
Rich Harang @rich.harang.org · 03/09/2026
ICYMI: x.com/nvidianewsro... www.linkedin.com/feed/update/... Selfishly really happy to see NVIDIA throw its resources behind open weight models and public datasets via the HF platform.
Screenshot of a post from the verified NVIDIA Newsroom account (@nvidianewsroom) announcing that NVIDIA has entered into a definitive agreement to acquire Hugging Face. The post says NVIDIA plans to help scale Hugging Face’s platform, strengthen its infrastructure, and expand access to AI, while Hugging Face will remain an open, neutral, platform-agnostic home for the AI ecosystem.
041
Leon Derczynski @leonderczynski.bsky.social · 03/09/2026
We NVIDIA have officially agreed to acquire Hugging Face for $12,930,300,000 and zero cents! HF is critical infrastructure that NVIDIA will strengthen and ensure sustained access to for developers and institutions worldwide. 🧵
100
Leon Derczynski @leonderczynski.bsky.social · 21/08/2026
Where Security Fits in an AI Agent Stack Layered separations for agents, to secure them across the stack All these levels can be secured + there are offerings for each, exposed in this post ... developer.nvidia.com/blog/where-s...
developer.nvidia.com
Where Security Fits in an AI Agent Stack | NVIDIA Technical Blog
As AI agents become more capable and operate over longer horizons, building security and trust into the applications they power becomes increasingly important. Drawing on work with NVIDIA OpenShell…
120
Leon Derczynski @leonderczynski.bsky.social · 21/08/2026
qualitative is harder than quantitative quantitative is the reductive, lossy representation of qualitative given the choice, preferring quantitative is taking the easy, less well-informed road, than dealing with the deep reality of qualitative quantitative is the cop out
220
Leon Derczynski @leonderczynski.bsky.social · 19/08/2026
New desktop ordered! :o one thousand of these in total are sent out across machine learning and agent research here, to accelerate researchers and unlock larger model sizes locally
110
Leon Derczynski @leonderczynski.bsky.social · 18/08/2026
Catching vulnerabilities before they reach production: Red Hat's multi-layer approach to catching agent weaknesses "The 6 AM failures weren't unforeseeable. They were unobserved." 🧵
100
Leon Derczynski @leonderczynski.bsky.social · 14/08/2026
Nice post-mortem on 1600 agent-derived vulnerabilities. Takeaways: * benchmarks are saturated * multi-step attacks are commonplace * threat is rising * offense drives defense xbow.com/blog/we-ran-...
xbow.com
We Ran 1,060 Autonomous Attacks: What We Learned | XBOW
Analysis of 1,060 autonomous attacks reveals how AI-driven offensive security performs in practice and what it means for modern defense strategies.
000
Leon Derczynski @leonderczynski.bsky.social · 14/08/2026
Agentic AI Challenges Progress in Confidential Computing securing data on gpus is non-trivial. encryption has to be low level to stop accidental exposures. lots of the slowdowns here have now been addressed. maybe confidential computing will finally happen? www.darkreading.com/endpoint-sec...
darkreading.com
Agentic AI Challenges Progress in Confidential Computing
Core issues that slowed down the adoption of secure data vaults are being resolved by technology, but artificial intelligence poses new ones.
010
Leon Derczynski @leonderczynski.bsky.social · 14/08/2026
Decent history of distilliation-type events, where information is reconstructed Machine learning distilliation is just another instance of a teacher-student dynamic, where one can reconstruct information efficiently. Even basic synthesis of research is distillation. It's a commonplace activity.
100
Leon Derczynski @leonderczynski.bsky.social · 13/08/2026
Every context is different, by definition. Being able to fit tools to context is therefore critical - and this generalisation holds incredibly well. Being able to tune a model to a task, gets it to perform better in context. Look at these solid examples: blogs.nvidia.com/blog/nemotro...
blogs.nvidia.com
Nemotron Labs: How Open Models Give Enterprises and Nations AI They Can Trust, Control and Customize
Learn how NVIDIA Nemotron open models help enterprises build specialized AI they can trust, control and customize with accuracy, efficiency and flexibility.
010
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026
tl;dr: statistical models of words and cognition don't work if we consider text alone this isn't news, but is a fine way of wrapping up a theory reinforces implications for LLMs - which operate on text alone arxiv.org/abs/2607.21574 (unreviewed)
000
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026
Legal takes on distillation: * Distillation doesn’t involve breaking in to download weights/code * Distillation doesn’t copy a patented implementation or method * Mass distillation merits policy response * Value loss thru distillation is a business model problem www.lawfaremedia.org/article/resp...
lawfaremedia.org
Responding to AI Distillation Without Panic
Before locking in new restrictions on AI models, policymakers should ask how much distillation actually contributes to the threat.
000
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026
Linux kernel team publishes 432 CVEs in two days Agent-powered vulnerabilty discovery has created a huge wave of incidents and reports that are cascading through systems from left to right - through triage, into remediation, then distribution, and patch management.
121
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026
"how to write email" I've forgotten this skill many times over the decades - luckily there's a cheat sheet. So much gold here. And big respect to Matt Might; a wonderful father and a truly remarkable human being matt.might.net/articles/how...
010
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026
Summon a Demon and Bind It: A grounded theory of LLM red teaming We didn't know it at the time but this paper was foundational work: a detailed study of LLM red teaming just after the launch of ChatGPT, back in 2022 + jan 2023.
120
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026
NVIDIA just launched another open model: Nemotron 3.5 Lightning. ⚡ 30B parameters / 3B active. Built for specialized, high-volume/long-running tasks without bringing a heavyweight model to every step. This is a fine model size; you don't need a host of GPUs or off-premises compute.
0101
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026
Is he just throwing shade on big models? Yes. Distill! Small models are best. "Overparameterization is often the mark of mediocrity" Brutal -- from George Box, while trying to source the "All model's are wrong but some are useful" claim. Credit Kasper Hornbæk
010
Leon Derczynski @leonderczynski.bsky.social · 10/08/2026
agent asked to book a gym session session is fully booked finds a way to break into system kicks other people off the queue you're in! see you in court? www.abc.net.au/news/2026-08...
abc.net.au
How a simple request for AI to book a gym class exposed a major threat
When Andrew asked his AI personal assistant to book him a spot in a gym class, he had no idea he would accidentally initiate an autonomous cyber attack.
000
Leon Derczynski @leonderczynski.bsky.social · 10/08/2026
Agents rule of two - reducing agent data exfil risk easily. Pick only two: [A] An agent can process untrustworthy inputs [B] An agent can have access to sensitive systems or private data [C] An agent can change state or communicate externally
simonwillison.net
New prompt injection papers: Agents Rule of Two and The Attacker Moves Second
Two interesting new papers regarding LLM security and prompt injection came to my attention this weekend. Agents Rule of Two: A Practical Approach to AI Agent Security The first is …
200
Leon Derczynski @leonderczynski.bsky.social · 07/08/2026
There are two readings here: * optimising efficiency means reduced costs (fiscal & environmental) * optimising efficiency means increased consumption (jevon's paradox) But either way, I like doing the same for less blogs.nvidia.com/blog/perform...
blogs.nvidia.com
Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency
From benchmark to production, NVIDIA Blackwell NVL72 delivers the highest performance per watt to maximize revenue and the lowest token cost to maximize profit margins.
120
Leon Derczynski @leonderczynski.bsky.social · 07/08/2026
Security is, at its core, a design problem. There will always be vulns and misconfigurations; design is what protects us all. The whole "Johnny can't PGP" is a similar instance of this truth. Good article with the right angle. www.accomplish.ai/blog/sharedr...
accomplish.ai
SharedRoot; Escaping the Claude Cowork sandbox — Accomplish Blog
Untrusted content in a Claude Cowork session can escape the VM it's sandboxed in and read and write files anywhere on your Mac. The kernel bug that makes it possible isn't the interesting part. Four…
000
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026
OpenAI gives first detailed debrief of the Hugging Face incident collab element is neat "the company said it revoked the credentials that had allowed the agents to post messages, rebuilt its internal software repository known as Artifactory, cleared the message board, patched the vulnerabilities"
110
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026
unsure if anyone who's tried to replicate contemporary cybersecurity performance with open models & components has failed to beat/replicate it. i've only seen evidence of successes
a man claims that one can't replicate top cybersecurity results with open models. the word MISINFO is plastered over the image
130
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026
Synthetic data generation for security: Project Marinade uses LLM coding agents to inject realistic, tunable vulnerabilities into real codebases This is how you get data to evaluate and train systems in cybersecurity tasks. thecyberarchive.com/talks/synthe...
020
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026
I found you can use ANSI to control anyone's computer through an LLM by getting the model to outputs special control characters. embracethered.com/blog/posts/2... Super clever connection - and Apple have fixed it! Congrats wuzzi23 on another cool vulnerability and fix.
100
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026
"me too, me too!!" --Meta's Muse Spark 1.1 model breached the unidentified company's systems and made changes to its internal systems as the AI was able to access the public internet because of an error in the set up of the "sandbox" testing environment
100
Leon Derczynski @leonderczynski.bsky.social · 05/08/2026
found at the gym next to work. how could we ever guess there are tech offices here? truly a mystery
000
Leon Derczynski @leonderczynski.bsky.social · 05/08/2026
This post seemed sensational in May but is mostly correct. Exploit-writing is now part of the discovery chain. I'm a little less sure about impact; 0days don't make up a huge amount of what is attacked in practice - but are still significant. suzulabs.com/suzu-labs-bl...
000
Leon Derczynski @leonderczynski.bsky.social · 05/08/2026
Microsoft launch their vulnerability discovery platform. Looks cool! Wish it was open source Introducing MAI-Cyber-1-Flash inside MDASH: World-class security at half the cost microsoft.ai/news/introdu...
microsoft.ai
Introducing MAI-Cyber-1-Flash inside MDASH | Microsoft AI
Discover the latest MAI models, designed for real-world intelligence
020
Leon Derczynski @leonderczynski.bsky.social · 04/08/2026
Come work with us! Evaluation and ML Systems Engineer, AI Safety and Security Engineering (remote) jobs.nvidia.com/careers?quer...
010
Leon Derczynski @leonderczynski.bsky.social · 04/08/2026
K3 deemed "not spicy". From NIST/AISI: "Kimi K3 performs significantly below the leading U.S. cyber capable models. Specifically, Kimi K3 reached step 17 of this 32-step attack path on average, while the most cyber-capable U.S. models reached 28.5 steps on average." www.nist.gov/news-events/...
021
Leon Derczynski @leonderczynski.bsky.social · 03/08/2026
remade my website (it's been a minute), www.derczynski.com/ua571c/ never don't have a website
derczynski.com
LD // REMOTE TERMINAL
Models & security at NVIDIA. Prof at ITU Copenhagen (NLP). Policy. Founder of garak. ACL SIGSEC Chair. Advises national and supranational government organisations on AI policy. Scientific leader in…
130
Leon Derczynski @leonderczynski.bsky.social · 03/08/2026
dfs-large1: fine-tuned GLM-5.2 for cybersec huge congrats to depthfirst on finetuning this and getting good enough perf to hit the pareto-optimal frontier -- looks like the better your fine-tuning results, the higher you can price. can't do that without open models! depthfirst.com/research/dfs...
020
Leon Derczynski @leonderczynski.bsky.social · 31/07/2026
Watts per token is a fine metric! No need to worry about things like vendor lock-in or other defensive tactics if the product can sell itself. - The other axis to reduce, is token usage. blogs.nvidia.com/blog/vera-ru...
blogs.nvidia.com
NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
Backed by 300 global partners, Vera Rubin is ramping up worldwide. NVIDIA partners CoreWeave, Google Cloud, Microsoft Azure and Mistral are among many deploying Vera Rubin, which delivers benchmark…
010
Leon Derczynski @leonderczynski.bsky.social · 31/07/2026
"of course, back when i were a lad, we didn't need any fancy ai for anything, because we had the TLDP HOWTO list" - seriously this thing was a godsend. anything linux you wanted to do or learn, you could find here tldp.org/HOWTO/HOWTO-...
tldp.org
Single list of HOWTOs
The following Linux HOWTOs are currently available:
020
Leon Derczynski @leonderczynski.bsky.social · 31/07/2026
"The point is not that any single model, including ours, will always be the best" Exactly what comes up with vulnerability discovery. No one model finds all the vulnerabilities. No single vuln is found by only one model. This is a problem we have to work on together. depthfirst.com/post/why-def...
depthfirst.com
Why Defenders Can’t Bet on One Model | depthfirst
Critical security capabilities cannot depend on infrastructure customers do not control. The next generation of AI security products needs to adapt across models, providers, policies, and access…
010
Leon Derczynski @leonderczynski.bsky.social · 31/07/2026
Am I the only AI Security research lead at a frontier model corp who hasn't been carefully committing multiple CFAA violations a month, or..?
010
Leon Derczynski @leonderczynski.bsky.social · 30/07/2026
Representing code at both function- and statement-level gives improvements in vulnerability detection. Surprisingly no static analysis baseline, or cost analysis - but recall is high. DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization arxiv.org/abs/2605.11015
010
Leon Derczynski @leonderczynski.bsky.social · 29/07/2026
Job title I've never seen before: "Principal Cyber Security Engineer - Agentic Identity and Security" So many open challenges - and so many amazingly skilled people. I don't know many places with this high a concentration of AI+Security experts. Come work with us! jobs.nvidia.com/careers/job/...
jobs.nvidia.com
Principal Cyber Security Engineer - Agentic Identity and Security | NVIDIA Corporation
Security jobs in NVIDIA Corporation
000
Leon Derczynski @leonderczynski.bsky.social · 29/07/2026
Are politicians the best people to place guardrails on how we do computing? Constraining the *uses* of technology makes sense - it's crucial here. But the gulf between subject experts and politicians often ends in harm. Just look at what EU Chatcontrol degraded into. thehill.com/policy/techn...
thehill.com
Lieu criticizes GOP colleagues for lack of strategy on AI
Rep. Ted Lieu (D-Calif.) on Wednesday expressed concern about the rapid evolution of artificial intelligence. “I am still freaked out by AI, and it’s actually accelerated much more quickly than I t…
000
Leon Derczynski @leonderczynski.bsky.social · 29/07/2026
False dichotomies around LLM speak: * "frontier" models vs. open model - leading models can be open * closed model vs. Chinese model - origin doesn't impact distribution You can have open frontier models, closed Chinese models, open US models, closed non-frontier models (e.g. for private context)
020
Leon Derczynski @leonderczynski.bsky.social · 29/07/2026
Capital One "VulnHunter" - an open harness for vulnerability discovery Cool to see more and more OSS in the security domain. You can clone it from GitHub and run it now. Significant adds in three key areas:
120
Leon Derczynski @leonderczynski.bsky.social · 28/07/2026
another data point showing it's the harness not the model - this time from wiz: "Atlas: Wiz's autonomous AI Agent for vulnerability research" look at their bold quote -- "Along the way, we learned that the durable advantage is not any single model, but the system around it"
100
Leon Derczynski @leonderczynski.bsky.social · 28/07/2026
"The only reason why I'm bearish about the Chinese models is because I assume that the American model companies will respond competitively." 🤨 www.npr.org/2026/07/15/n...
npr.org
American AI is expensive. Some startups are turning to cheap Chinese models
AI is a fast-growing business expense. Some companies are cutting costs by switching to cheaper Chinese AI models.
110
Leon Derczynski @leonderczynski.bsky.social · 28/07/2026
"Open-source AI matters because it defines the ecosystem. It provides the foundation and sets parameters for progress, just like the open infrastructure of the internet and early AI: BSD Unix, PostgreSQL, Firefox" Open wins at grass roots level 🤷‍♂️ nationalinterest.org/blog/techlan...
nationalinterest.org
Why America Must Dominate Open-Source AI
China's open-source AI models are spreading globally. To preserve technological leadership, the United States must lead not only in capabilities but in openness.
020