Sign in

Leon Derczynski

@leonderczynski.bsky.social
833 followers 371 following 313 posts

LLMs & Security at NVIDIA Prof in CS/NLP at IT University of Copenhagen garak guy, garak.ai "berømt skikkelse" "like a gazelle" Seattle/Copenhagen 🏔️

PostsRepliesMedia
Leon Derczynski @leonderczynski.bsky.social · 02/10/2026
I get some wild mails while at NVIDIA. Like today - would I like to give a "paid, recorded interview on AI governance and US–China competition" to a researcher in Hong Kong and give my opinions on Huawei silicon. That is.. not something we will be doing.
000
Leon Derczynski @leonderczynski.bsky.social · 12/09/2026
they have big beds, but why does Texas still use the small US gallons, pints, and electricity?
000
Leon Derczynski @leonderczynski.bsky.social · 11/09/2026
But my question is this: Why would text have to conform to a lower-order computationally-expressable grammar? Text isn't even all of language - there's lots of other information out there, and it's all in the humans and the human experience.
000
Leon Derczynski @leonderczynski.bsky.social · 11/09/2026
Some find the human ability to communicate without expressing everything in a stable, fully populated structure quite disgusting (Randall Monroe included I guess, see graphic) - others find it fascinating.
100
Leon Derczynski @leonderczynski.bsky.social · 11/09/2026
One divide in computer science can be between the uncertainty-reducers and the uncertainty-embracers. An instance of this is connectionist vs. symbolist approaches to AI (the uncertainty embracers won .. for now). Another is how computer scientists react to natural language.
110
Leon Derczynski @leonderczynski.bsky.social · 11/09/2026
Come work with us! Senior Solutions Architect, Agentic AI: Safety and Security This work includes multi-agent orchestration, guardrails, agent runtime security, model customization, policy enforcement, OpenShell-like environments, and confidential deployments. jobs.nvidia.com/careers/job/...
120
Leon Derczynski @leonderczynski.bsky.social · 11/09/2026
huh, that's one way of doing it 🤗
011
Leon Derczynski @leonderczynski.bsky.social · 09/09/2026
Language is, and always was, more important than mathematics. The Navier-Stokes equations mathematically express momentum balance for...
000
Reposted by Leon Derczynski
Rich Harang @rich.harang.org · 03/09/2026
ICYMI: x.com/nvidianewsro... www.linkedin.com/feed/update/... Selfishly really happy to see NVIDIA throw its resources behind open weight models and public datasets via the HF platform.
Screenshot of a post from the verified NVIDIA Newsroom account (@nvidianewsroom) announcing that NVIDIA has entered into a definitive agreement to acquire Hugging Face. The post says NVIDIA plans to help scale Hugging Face’s platform, strengthen its infrastructure, and expand access to AI, while Hugging Face will remain an open, neutral, platform-agnostic home for the AI ecosystem.
041
Leon Derczynski @leonderczynski.bsky.social · 03/09/2026
The neutrality and might that NVIDIA bring is a fantastic match for Hugging Face and their crucial place among researchers and industry. This acquisition feels like it puts a lot of calm on to the future of machine learning research and open access to open weight work. Good. 🤗💚
000
Leon Derczynski @leonderczynski.bsky.social · 03/09/2026
I think NVIDIA is a great home for Hugging Face - I'm pretty open about NV being the best place to do open, unbiased work; that's why I chose them as a home when moving on from academia (which I dearly love).
100
Leon Derczynski @leonderczynski.bsky.social · 03/09/2026
I've followed and used Hugging Face since it was a chat platform, even before the iconic "transformers" library landed in 2018. There are some wonderful people at HF and I'm looking forward to having so many of them as colleagues. Really happy.
100
Leon Derczynski @leonderczynski.bsky.social · 03/09/2026
HF will remain an open platform. This move will grow Hugging Face, only. It's too important to do anything else with it!
100
Leon Derczynski @leonderczynski.bsky.social · 03/09/2026
We NVIDIA have officially agreed to acquire Hugging Face for $12,930,300,000 and zero cents! HF is critical infrastructure that NVIDIA will strengthen and ensure sustained access to for developers and institutions worldwide. 🧵
100
Leon Derczynski @leonderczynski.bsky.social · 21/08/2026
i should share my favourite jhh quote on this next time we meet
010
Leon Derczynski @leonderczynski.bsky.social · 21/08/2026
Securing the agent stack: * Distribution: Package installation & defaults * Orchestration: Selects and coordinates harnesses * Agent harness: Loop, context, tools, sessions * Secure runtime: Isolation, identity, policy, creds, audit * Inference data plane: Model serving, cache placement, routing
000
Leon Derczynski @leonderczynski.bsky.social · 21/08/2026
Where Security Fits in an AI Agent Stack Layered separations for agents, to secure them across the stack All these levels can be secured + there are offerings for each, exposed in this post ... developer.nvidia.com/blog/where-s...
developer.nvidia.com
Where Security Fits in an AI Agent Stack | NVIDIA Technical Blog
As AI agents become more capable and operate over longer horizons, building security and trust into the applications they power becomes increasingly important. Drawing on work with NVIDIA OpenShell…
120
Leon Derczynski @leonderczynski.bsky.social · 21/08/2026
no not really, it's just one of those messages that needs broad reach because people forget. you are the ignition point this time tho 100% (thank you!)
101
Leon Derczynski @leonderczynski.bsky.social · 21/08/2026
NOT ANY MORE 😂
100
Leon Derczynski @leonderczynski.bsky.social · 21/08/2026
That was your best?
000
Leon Derczynski @leonderczynski.bsky.social · 21/08/2026
qualitative is harder than quantitative quantitative is the reductive, lossy representation of qualitative given the choice, preferring quantitative is taking the easy, less well-informed road, than dealing with the deep reality of qualitative quantitative is the cop out
220
Leon Derczynski @leonderczynski.bsky.social · 19/08/2026
spec: - Grace Blackwell Ultra GB300 Superchip - 496GB LPDDR5X CPU memory, 252GB HBM3e GPU memory - Supports up to 1 trillion parameter models - 20 Petaflops (20,000 TFLOPS) of FP4 computing power - Ubuntu with NVIDIA AI Developer Tools ..unclear if it runs crysis
040
Leon Derczynski @leonderczynski.bsky.social · 19/08/2026
New desktop ordered! :o one thousand of these in total are sent out across machine learning and agent research here, to accelerate researchers and unlock larger model sizes locally
110
Leon Derczynski @leonderczynski.bsky.social · 18/08/2026
AI agent observability: Building a production-grade operational layer (Red Hat) www.redhat.com/en/blog/ai-a...
redhat.com
AI agent observability: Building a production-grade operational layer
How Red Hat AI enables production-grade AI agent observability.
000
Leon Derczynski @leonderczynski.bsky.social · 18/08/2026
"Every failure left traces that would have been visible if anyone had been watching—retry loops piling up, spend anomalies spiking, and customer-facing outputs contradicting documented policy. The agents weren't broken. The operational infrastructure to watch them didn't exist."
redhat.com
AI agent observability: Building a production-grade operational layer
How Red Hat AI enables production-grade AI agent observability.
220
Leon Derczynski @leonderczynski.bsky.social · 18/08/2026
Catching vulnerabilities before they reach production: Red Hat's multi-layer approach to catching agent weaknesses "The 6 AM failures weren't unforeseeable. They were unobserved." 🧵
100
Leon Derczynski @leonderczynski.bsky.social · 14/08/2026
Nice post-mortem on 1600 agent-derived vulnerabilities. Takeaways: * benchmarks are saturated * multi-step attacks are commonplace * threat is rising * offense drives defense xbow.com/blog/we-ran-...
xbow.com
We Ran 1,060 Autonomous Attacks: What We Learned | XBOW
Analysis of 1,060 autonomous attacks reveals how AI-driven offensive security performs in practice and what it means for modern defense strategies.
000
Leon Derczynski @leonderczynski.bsky.social · 14/08/2026
Agentic AI Challenges Progress in Confidential Computing securing data on gpus is non-trivial. encryption has to be low level to stop accidental exposures. lots of the slowdowns here have now been addressed. maybe confidential computing will finally happen? www.darkreading.com/endpoint-sec...
darkreading.com
Agentic AI Challenges Progress in Confidential Computing
Core issues that slowed down the adoption of secure data vaults are being resolved by technology, but artificial intelligence poses new ones.
010
Leon Derczynski @leonderczynski.bsky.social · 14/08/2026
If your business is predicated on distilliation (i.e. learning) not happening, perhaps there is a business model problem? www.linkedin.com/pulse/horrif...
linkedin.com
000
Leon Derczynski @leonderczynski.bsky.social · 14/08/2026
Decent history of distilliation-type events, where information is reconstructed Machine learning distilliation is just another instance of a teacher-student dynamic, where one can reconstruct information efficiently. Even basic synthesis of research is distillation. It's a commonplace activity.
100
Leon Derczynski @leonderczynski.bsky.social · 13/08/2026
Every context is different, by definition. Being able to fit tools to context is therefore critical - and this generalisation holds incredibly well. Being able to tune a model to a task, gets it to perform better in context. Look at these solid examples: blogs.nvidia.com/blog/nemotro...
blogs.nvidia.com
Nemotron Labs: How Open Models Give Enterprises and Nations AI They Can Trust, Control and Customize
Learn how NVIDIA Nemotron open models help enterprises build specialized AI they can trust, control and customize with accuracy, efficiency and flexibility.
010
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026
tl;dr: statistical models of words and cognition don't work if we consider text alone this isn't news, but is a fine way of wrapping up a theory reinforces implications for LLMs - which operate on text alone arxiv.org/abs/2607.21574 (unreviewed)
000
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026
Legal takes on distillation: * Distillation doesn’t involve breaking in to download weights/code * Distillation doesn’t copy a patented implementation or method * Mass distillation merits policy response * Value loss thru distillation is a business model problem www.lawfaremedia.org/article/resp...
lawfaremedia.org
Responding to AI Distillation Without Panic
Before locking in new restrictions on AI models, policymakers should ask how much distillation actually contributes to the threat.
000
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026
The last two stages are very tough - patch providers are having a tough time now. Kudos to Linux for managing to triage, validate, land and publish this much so fast. www.theregister.com/security/202...
theregister.com
010
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026
Linux kernel team publishes 432 CVEs in two days Agent-powered vulnerabilty discovery has created a huge wave of incidents and reports that are cascading through systems from left to right - through triage, into remediation, then distribution, and patch management.
121
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026
"how to write email" I've forgotten this skill many times over the decades - luckily there's a cheat sheet. So much gold here. And big respect to Matt Might; a wonderful father and a truly remarkable human being matt.might.net/articles/how...
010
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026
Main image credit: Feddie Xtzeth - Promptmancer; made to describe what this (as of then unnamed) activity feels like and is. Paper & Summary: summonademonandbind.it
summonademonandbind.it
Summon a Demon and Bind It: A Grounded Theory of LLM Red-Teaming
We interviewed dozens of experts to establish why and how people attack LLMs. This builds a 'grounded theory' describing the activity of LLM attacking, how it fits into the world, and how people…
010
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026
We came out with the first (and still a major) data-driven taxonomy of the techniques people use to attack models; now embedded inside a major NVIDIA technology used across the industry and government in security products.
summonademonandbind.it
Summon a Demon and Bind It: A Grounded Theory of LLM Red-Teaming
We interviewed dozens of experts to establish why and how people attack LLMs. This builds a 'grounded theory' describing the activity of LLM attacking, how it fits into the world, and how people…
110
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026
Captured an amazing moment in time as a new technology collided with society and a tiny group of people learned to manipulate it, while trying to develop understanding and metaphors of something completely unknown.
summonademonandbind.it
Summon a Demon and Bind It: A Grounded Theory of LLM Red-Teaming
We interviewed dozens of experts to establish why and how people attack LLMs. This builds a 'grounded theory' describing the activity of LLM attacking, how it fits into the world, and how people…
100
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026
Summon a Demon and Bind It: A grounded theory of LLM red teaming We didn't know it at the time but this paper was foundational work: a detailed study of LLM red teaming just after the launch of ChatGPT, back in 2022 + jan 2023.
120
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026
NVIDIA just launched another open model: Nemotron 3.5 Lightning. ⚡ 30B parameters / 3B active. Built for specialized, high-volume/long-running tasks without bringing a heavyweight model to every step. This is a fine model size; you don't need a host of GPUs or off-premises compute.
0101
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026
Is he just throwing shade on big models? Yes. Distill! Small models are best. "Overparameterization is often the mark of mediocrity" Brutal -- from George Box, while trying to source the "All model's are wrong but some are useful" claim. Credit Kasper Hornbæk
010
Leon Derczynski @leonderczynski.bsky.social · 10/08/2026
agent asked to book a gym session session is fully booked finds a way to break into system kicks other people off the queue you're in! see you in court? www.abc.net.au/news/2026-08...
abc.net.au
How a simple request for AI to book a gym class exposed a major threat
When Andrew asked his AI personal assistant to book him a spot in a gym class, he had no idea he would accidentally initiate an autonomous cyber attack.
000
Leon Derczynski @leonderczynski.bsky.social · 10/08/2026
NB this only covers data exfiltration risk: there are many other risks from prompt injection which the lethal trifecta doesn’t cover. Agents rule of two: simonwillison.net/2025/Nov/2/n...
simonwillison.net
New prompt injection papers: Agents Rule of Two and The Attacker Moves Second
Two interesting new papers regarding LLM security and prompt injection came to my attention this weekend. Agents Rule of Two: A Practical Approach to AI Agent Security The first is …
110
Leon Derczynski @leonderczynski.bsky.social · 10/08/2026
Agents rule of two - reducing agent data exfil risk easily. Pick only two: [A] An agent can process untrustworthy inputs [B] An agent can have access to sensitive systems or private data [C] An agent can change state or communicate externally
simonwillison.net
New prompt injection papers: Agents Rule of Two and The Attacker Moves Second
Two interesting new papers regarding LLM security and prompt injection came to my attention this weekend. Agents Rule of Two: A Practical Approach to AI Agent Security The first is …
200
Leon Derczynski @leonderczynski.bsky.social · 07/08/2026
There are two readings here: * optimising efficiency means reduced costs (fiscal & environmental) * optimising efficiency means increased consumption (jevon's paradox) But either way, I like doing the same for less blogs.nvidia.com/blog/perform...
blogs.nvidia.com
Why Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency
From benchmark to production, NVIDIA Blackwell NVL72 delivers the highest performance per watt to maximize revenue and the lowest token cost to maximize profit margins.
120
Leon Derczynski @leonderczynski.bsky.social · 07/08/2026
Security is, at its core, a design problem. There will always be vulns and misconfigurations; design is what protects us all. The whole "Johnny can't PGP" is a similar instance of this truth. Good article with the right angle. www.accomplish.ai/blog/sharedr...
accomplish.ai
SharedRoot; Escaping the Claude Cowork sandbox — Accomplish Blog
Untrusted content in a Claude Cowork session can escape the VM it's sandboxed in and read and write files anywhere on your Mac. The kernel bug that makes it possible isn't the interesting part. Four…
000
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026
its falseness literally has no credit
000
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026
"agents figured out another way to communicate with each other. Instead of leaving messages in files, they used the names of newly created directories" www.groundlevel-ai.com/p/openai-giv...
groundlevel-ai.com
OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat conference
In a session attended by Ground Level AI, OpenAI researchers said the company is "consciously slowing down research to enhance security" while a full technical postmortem is still underway.
000
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026
OpenAI gives first detailed debrief of the Hugging Face incident collab element is neat "the company said it revoked the credentials that had allowed the agents to post messages, rebuilt its internal software repository known as Artifactory, cleared the message board, patched the vulnerabilities"
110