Leon Derczynski @leonderczynski.bsky.social · 10hGhostCompact: ~Free OpenAI inference "When the model's first action after compacting was a tool call, the API billed only for post-compaction tokens, but the tool call response still contained the full compaction item." Information loves to leak. Unsure what a good passive control would be here. 110
Leon Derczynski @leonderczynski.bsky.social · 02/10/2026I get some wild mails while at NVIDIA. Like today - would I like to give a "paid, recorded interview on AI governance and US–China competition" to a researcher in Hong Kong and give my opinions on Huawei silicon. That is.. not something we will be doing. 000
Leon Derczynski @leonderczynski.bsky.social · 11/09/2026One divide in computer science can be between the uncertainty-reducers and the uncertainty-embracers. An instance of this is connectionist vs. symbolist approaches to AI (the uncertainty embracers won .. for now). Another is how computer scientists react to natural language. 110
Leon Derczynski @leonderczynski.bsky.social · 11/09/2026Come work with us! Senior Solutions Architect, Agentic AI: Safety and Security This work includes multi-agent orchestration, guardrails, agent runtime security, model customization, policy enforcement, OpenShell-like environments, and confidential deployments. jobs.nvidia.com/careers/job/... 120
Leon Derczynski @leonderczynski.bsky.social · 03/09/2026We NVIDIA have officially agreed to acquire Hugging Face for $12,930,300,000 and zero cents! HF is critical infrastructure that NVIDIA will strengthen and ensure sustained access to for developers and institutions worldwide. 🧵 100
Leon Derczynski @leonderczynski.bsky.social · 21/08/2026Securing the agent stack: * Distribution: Package installation & defaults * Orchestration: Selects and coordinates harnesses * Agent harness: Loop, context, tools, sessions * Secure runtime: Isolation, identity, policy, creds, audit * Inference data plane: Model serving, cache placement, routing 000
Leon Derczynski @leonderczynski.bsky.social · 19/08/2026New desktop ordered! :o one thousand of these in total are sent out across machine learning and agent research here, to accelerate researchers and unlock larger model sizes locally 110
Leon Derczynski @leonderczynski.bsky.social · 18/08/2026Catching vulnerabilities before they reach production: Red Hat's multi-layer approach to catching agent weaknesses "The 6 AM failures weren't unforeseeable. They were unobserved." 🧵 100
Leon Derczynski @leonderczynski.bsky.social · 14/08/2026Decent history of distilliation-type events, where information is reconstructed Machine learning distilliation is just another instance of a teacher-student dynamic, where one can reconstruct information efficiently. Even basic synthesis of research is distillation. It's a commonplace activity. 100
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026tl;dr: statistical models of words and cognition don't work if we consider text alone this isn't news, but is a fine way of wrapping up a theory reinforces implications for LLMs - which operate on text alone arxiv.org/abs/2607.21574 (unreviewed) 000
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026Linux kernel team publishes 432 CVEs in two days Agent-powered vulnerabilty discovery has created a huge wave of incidents and reports that are cascading through systems from left to right - through triage, into remediation, then distribution, and patch management. 121
Leon Derczynski @leonderczynski.bsky.social · 12/08/2026"how to write email" I've forgotten this skill many times over the decades - luckily there's a cheat sheet. So much gold here. And big respect to Matt Might; a wonderful father and a truly remarkable human being matt.might.net/articles/how... 010
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026Summon a Demon and Bind It: A grounded theory of LLM red teaming We didn't know it at the time but this paper was foundational work: a detailed study of LLM red teaming just after the launch of ChatGPT, back in 2022 + jan 2023. 120
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026NVIDIA just launched another open model: Nemotron 3.5 Lightning. ⚡ 30B parameters / 3B active. Built for specialized, high-volume/long-running tasks without bringing a heavyweight model to every step. This is a fine model size; you don't need a host of GPUs or off-premises compute. 0101
Leon Derczynski @leonderczynski.bsky.social · 11/08/2026Is he just throwing shade on big models? Yes. Distill! Small models are best. "Overparameterization is often the mark of mediocrity" Brutal -- from George Box, while trying to source the "All model's are wrong but some are useful" claim. Credit Kasper Hornbæk 010
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026OpenAI gives first detailed debrief of the Hugging Face incident collab element is neat "the company said it revoked the credentials that had allowed the agents to post messages, rebuilt its internal software repository known as Artifactory, cleared the message board, patched the vulnerabilities" 110
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026unsure if anyone who's tried to replicate contemporary cybersecurity performance with open models & components has failed to beat/replicate it. i've only seen evidence of successes 130
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026Synthetic data generation for security: Project Marinade uses LLM coding agents to inject realistic, tunable vulnerabilities into real codebases This is how you get data to evaluate and train systems in cybersecurity tasks. thecyberarchive.com/talks/synthe... 020
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026I found you can use ANSI to control anyone's computer through an LLM by getting the model to outputs special control characters. embracethered.com/blog/posts/2... Super clever connection - and Apple have fixed it! Congrats wuzzi23 on another cool vulnerability and fix. 100
Leon Derczynski @leonderczynski.bsky.social · 06/08/2026"me too, me too!!" --Meta's Muse Spark 1.1 model breached the unidentified company's systems and made changes to its internal systems as the AI was able to access the public internet because of an error in the set up of the "sandbox" testing environment 110
Leon Derczynski @leonderczynski.bsky.social · 05/08/2026found at the gym next to work. how could we ever guess there are tech offices here? truly a mystery 000
Leon Derczynski @leonderczynski.bsky.social · 04/08/2026Come work with us! Evaluation and ML Systems Engineer, AI Safety and Security Engineering (remote) jobs.nvidia.com/careers?quer... 010
Leon Derczynski @leonderczynski.bsky.social · 04/08/2026K3 deemed "not spicy". From NIST/AISI: "Kimi K3 performs significantly below the leading U.S. cyber capable models. Specifically, Kimi K3 reached step 17 of this 32-step attack path on average, while the most cyber-capable U.S. models reached 28.5 steps on average." www.nist.gov/news-events/... 021
Leon Derczynski @leonderczynski.bsky.social · 03/08/2026dfs-large1: fine-tuned GLM-5.2 for cybersec huge congrats to depthfirst on finetuning this and getting good enough perf to hit the pareto-optimal frontier -- looks like the better your fine-tuning results, the higher you can price. can't do that without open models! depthfirst.com/research/dfs... 020
Leon Derczynski @leonderczynski.bsky.social · 30/07/2026Representing code at both function- and statement-level gives improvements in vulnerability detection. Surprisingly no static analysis baseline, or cost analysis - but recall is high. DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization arxiv.org/abs/2605.11015 010
Leon Derczynski @leonderczynski.bsky.social · 29/07/2026Capital One "VulnHunter" - an open harness for vulnerability discovery Cool to see more and more OSS in the security domain. You can clone it from GitHub and run it now. Significant adds in three key areas: 120
Leon Derczynski @leonderczynski.bsky.social · 28/07/2026another data point showing it's the harness not the model - this time from wiz: "Atlas: Wiz's autonomous AI Agent for vulnerability research" look at their bold quote -- "Along the way, we learned that the durable advantage is not any single model, but the system around it" 100
Leon Derczynski @leonderczynski.bsky.social · 27/07/2026The Hill covers Open Secure AI Alliance: “The United States now faces a similar choice with artificial intelligence” the letter states. “Our AI leadership will be judged not by one frontier AI model, but by whether the United States builds a strong, open ecosystem that diffuses into every sector” 110
Leon Derczynski @leonderczynski.bsky.social · 27/07/2026 New: NVIDIA Labs Object Oriented Agent tech demo. Blog: developer.nvidia.com/blog/six-age... GitHub: github.com/nvidia-nemo/... Paper: arxiv.org/abs/2607.20709 000
Leon Derczynski @leonderczynski.bsky.social · 27/07/2026The harness is everything. New tech demo dropped for one way of doing CodeAct agents. Super-efficient. The team smashed CyberGym while building this - it's now the top-ranking open system for that vulnerability discovery benchmark. 130
Leon Derczynski @leonderczynski.bsky.social · 27/07/2026Open source is critical infrastructure for the global economy. Launching today, the Open Secure AI Alliance brings industry and community together around shared research, tools and vulnerability harnesses to help defenders find and patch bugs before attackers strike. nvda.ws/4pAMWBy 031
Leon Derczynski @leonderczynski.bsky.social · 24/07/2026Adversarial attack in the wild! The close visual appearance of M and W in this typeface and and packing of vertical lines make it hard to read, easy to get wrong, and tougher to scan. Love it. How often do you see something like this?! 000
Leon Derczynski @leonderczynski.bsky.social · 23/07/2026Huge success and part of how I chose who to approach when moving to industry. Open source is the way. Looking at "repositories with meaningful traction", AI world say "NVIDIA is now the largest contributor, with more than 600 repositories in the past year" 🎉 🧙 💚 aiworld.eu/story/the-ne... 030
Leon Derczynski @leonderczynski.bsky.social · 23/07/2026Anonymised analysis of the openai model 'breaching' hugging face: > report doesn't say what sandbox sol broke out of?? > a docker container running as root > Plot twist there was no sandbox at all > many use "sandbox" and "container with host access" interchangeably ymmv, use critical thinking 120
Leon Derczynski @leonderczynski.bsky.social · 22/07/2026generally, everyone outside of the upper ~quarter of incomes is much better off under heavy economic redistribution (i.e. left-wing econonmic policy) this is uncontroversial unfortunately the mathematical literacy to see this is not universal fortunately money isn't the most important factor 011
Leon Derczynski @leonderczynski.bsky.social · 15/07/2026Solid US-origin open model, congratulations Thinking Machines * context window of 1M * available on hugging face now * between opus 4.6 and gpt 5.6 on a web dev benchmark * token efficient * 41B active params of 975B total www.wired.com/story/thinki... 2212
Leon Derczynski @leonderczynski.bsky.social · 09/07/2026The whole digital economy runs on open source. Banning open models is bad for almost everything if it lands - business and research and consumers. And they present no risk. What's going on? thehill.com/policy/technology/59522… 020
Leon Derczynski @leonderczynski.bsky.social · 08/05/2026@jonchristian.net love seeing your stories pumped straight to my lock screen notifications 110
Leon Derczynski @leonderczynski.bsky.social · 03/11/2025😂 arXiv is cute when it pretends to have standards! 040
Leon Derczynski @leonderczynski.bsky.social · 28/07/2025Come to LLMSEC at ACL & hear Niloofar's keynote "What does it mean for agentic AI to preserve privacy?" - Niloofar Mireshghallah, Meta/CMU (Friday 1st Aug, 11.00; Austria Center Vienna Hall B) See you there! #acl2025 #acl2025nlp 1122
Leon Derczynski @leonderczynski.bsky.social · 19/03/2025Here's my "Most Inappropriate Demo" trophy at NVIDIA, 2024. For garak's "atkgen.Tox" probe, an unfettered LLM used to goad other LLMs into being toxic. 080
Leon Derczynski @leonderczynski.bsky.social · 17/02/2025it's a weekday where I dont have to take pacific time calls 040
Leon Derczynski @leonderczynski.bsky.social · 13/02/2025you know the field has changed when the foreign event you were speaking at is on the tv news on the bus home 040
Leon Derczynski @leonderczynski.bsky.social · 08/02/2025Will be representing NVIDIA at the EU AI Summit in Paris. I'll be talking about how we build & help others build safe, secure AI systems. On 11.2 you can see me at: * AI Assurance and Testing: Global Perspectives * Building trustworthy AI: balancing innovation, responsibility, and democratization 060
Leon Derczynski @leonderczynski.bsky.social · 05/01/2025Are there people who don't make the sponge cake rice cooker recipe asap?! 120
Leon Derczynski @leonderczynski.bsky.social · 28/12/2024Good Christmas times, finally the elephant has come to our house! 020