Liran Tal @lirantal.com · 4hmind blowing how the variance of different fixtures in the Snyk VulnBench 2.0 benchmark data shows up for the same security agent harness, and same model (OpenAI's GPT-5.6 Sol) all the data is open on vulnbench.com 000
Liran Tal @lirantal.com · 7has I was wrapping up the Snyk VulnBench 2.0 publication, here's a quick glimpse into the various security agents and their cost to run the evals codex security agent, claude code security review and vercel's deepsec all completed and heads up, costs do not correlate with findings accuracy :) 000
Liran Tal @lirantal.com · 13hback in my day when we "hacked the fbi" we were on the run but I guess when AI agents are doing it then it's all fun and games, no repercussions eh :) 000
Liran Tal @lirantal.com · 16hthis is the new "github is down so let's take a coffee break" yes or yes 130
Liran Tal @lirantal.com · 29/09/2026Great episode from Gergely Orosz interviewing Matt Pocock and touching on topics I really like around developer education, course building, the right focus in times of AI. The podcast goes into other software engineering depths and the grill-me skill but more than anything else if you're in DevRel 001
Liran Tal @lirantal.com · 29/09/2026to whomever needs to hear this, all of you folks automating GTM, outreach, marketing and growth with AI agents - if I see an unsolicited mail in my inbox that is AI generated and 100% of them are, I am immediately pushing the "Spam & block" button. I've already done this maybe like 20-30 times in 010
Liran Tal @lirantal.com · 28/09/2026whoever you are at apple who decided that the default Finder folder is "Recent" and not "Downloads" you should resign immediately as a public service to humanity, go to Finder -> Settings and change it so that you get back 5 minutes each year 010
Liran Tal @lirantal.com · 28/09/2026if I had asked developers what they wanted they'd say better scanners what they actually need though is agent driven security 010
Liran Tal @lirantal.com · 28/09/2026kinda feel bad I'm doing a software update while at the lounge's wifi 010
Liran Tal @lirantal.com · 28/09/2026Get up to speed on all things security agents and harnesses at AI Security Summit in SF on October 15th 🤖 🔥 020
Liran Tal @lirantal.com · 23/09/2026btw hope you know that Snyk always had a pretty generous free tier for dependency scanning, code security, and other capabilities we're extending that into credits based consumption for Enterprise too 000
Liran Tal @lirantal.com · 23/09/2026hmmm, didn't Anthropic folks say that this model switch in between session is no longer an issue or did I miss something? that's from the Claude Code app 010
Liran Tal @lirantal.com · 23/09/2026have you heard of Jev from TypeSafe AI by now? super fun infra to build with for doing agentic work but, it's also extremely handy on its own merit of classification and labeling I created a CLI tool that downloads all songs for any given artist and then classifies the lyrics across theme, mood, 010
Liran Tal @lirantal.com · 23/09/2026the first rule of the new npm is that the first version is always excluded of build provenance 010
Liran Tal @lirantal.com · 22/09/2026the npm registry now performs some level of package validation and evaluation before pushing the latest tag to the latest version published I'm not entirely versed on how deep and thorough the security audit here but better than nothing... 220
Liran Tal @lirantal.com · 22/09/2026so I guess this is the new default for agent coded CLIs these past 6 months huh 010
Liran Tal @lirantal.com · 22/09/2026have you ever wondered what'w the mood of one of Madonna's song? well, now you can, with Jev from TypeSafe AI I used Jev to classify any artist (all their songs / discography) and now we get labels for mood, lyrical complexity, overall theme and some other stats this has been very fun building 000
Liran Tal @lirantal.com · 22/09/2026looks like varlock implemented a bunch of DX improvements around env secrets, very smooth and helpful nice work Theo Ephraim and varlock team 000
Liran Tal @lirantal.com · 21/09/2026is the env vars secrets space heating up finally?? varlock vs infisical 010
Liran Tal @lirantal.com · 21/09/2026what's up with the latest Codex update, it's just stuck like this for a bit 010
Liran Tal @lirantal.com · 21/09/2026Jev app onboarding experience is very cool and on-brand Diogo and the team cooked well there 010
Liran Tal @lirantal.com · 21/09/2026new academic paper (Politecnico di Torino, arXiv) just benchmarked the whole field of agent-skill security scanners against real skills.sh skills Snyk Agent Scan is one of only 3 scanners integrated into skills.sh's own audit pipeline (alongside Socket + Gen Agent Trust Hub) — and had the broadest 120
Liran Tal @lirantal.com · 18/09/2026To everyone who attended my talk today at #AGNTCon + #MCPCon Europe in Amsterdam (yay, windmills!) - I uploaded my deck as a PDF to the session on sched, so you're welcome to browse through and re-visit the topics Should be freely available to everyone else too: 010
Liran Tal @lirantal.com · 17/09/2026wrote a book on this 2 years ago maybe you want to read or give it to your agent 000
Liran Tal @lirantal.com · 17/09/2026lol, if only it was that easy, right or... maybe it is? 😯 ask me :) 010
Liran Tal @lirantal.com · 17/09/2026I like how Ezra points out generation step vulnerabilities vs prevention altogether. Recommended read: snyk.io/blog/is-prev... 000
Liran Tal @lirantal.com · 17/09/2026reminder to use the boxdown CLI to configure your ChatGPT Codex or Claude Code to use local isolated container environments for agentic work (magically sets everythig up for ya!) yes, it's open source 🎉 p.s. Cursor is supported too if you're a fan of it 020
Liran Tal @lirantal.com · 16/09/2026Snyk VulnBench headline numbers How well do other models compare with an F1 agreement score with Snyk Code security findings from a SAST scan and their error rate (Opus 4.7 Max was surprising!) 010
Liran Tal @lirantal.com · 16/09/2026geeky but the terminal is back and I was having fun building up the Snyk VulnBench harness for benchmark purposes (now prefer the open source Harbor framework / CLI) learned so much from building the harness though totally recommend 000
Liran Tal @lirantal.com · 16/09/2026early back in May when I ran the Snyk VulnBench benchmark I compared various models to a Snyk F1 reference score (can they match what Snyk Code is reporting as security vulnerabilities) here's what they reported bonus: error bars 010
Liran Tal @lirantal.com · 16/09/2026Following is how anti-trojan-source CLI detects cases of potentially harmful characters, identified from the Glassworm attack: 000
Liran Tal @lirantal.com · 15/09/2026lol what this is Sonnet 5 on Medium so yes fine not Astra or Fable but come'on always always always verify and validate, I cannot stress this enough 010
Liran Tal @lirantal.com · 15/09/2026if you haven't been working with an AI agent (whether Claude Cowork or otherwise) as your main driver for a second brain at work you're falling behind 010
Liran Tal @lirantal.com · 14/09/2026autoapprove is cool but doing it without a malicious package slipping in is cooler 😉 000
Liran Tal @lirantal.com · 14/09/2026wip for running the benchmarks on VulnBench 2.0 which is a new set of fixture data, larger apps codebase, varied language ecosystem... pretty interesting how harness + model are very much a pair, for example Codex Security agent with Terra on xhigh just isn't scoring high enough 010
Liran Tal @lirantal.com · 14/09/2026wip for running the benchmarks on VulnBench 2.0 which is a new set of fixture data, larger apps codebase, varied language ecosystem... pretty interesting how harness + model are very much a pair, for example Codex Security agent with Terra on xhigh just isn't scoring high enough 110
Liran Tal @lirantal.com · 11/09/2026We have been working on the OWASP MCP Security Taxonomy - an open, vendor-neutral framework designed to create a common language for MCP security risks, weaknesses, attack patterns, controls, detections, and test cases: github.com/OWASP/MCP-Ta... Go check it out and give feedbackgithub.comGitHub - OWASP/MCP-Taxonomy: OWASP MCP TaxonomyOWASP MCP Taxonomy. Contribute to OWASP/MCP-Taxonomy development by creating an account on GitHub. 2116
Liran Tal @lirantal.com · 11/09/2026do we have this chart with updated models? would appreciate if someone has the prompt injection resistance from model cards handy to share 000
Liran Tal @lirantal.com · 11/09/2026this no longer works I think but how fun it is that you can just send text to trigger a denial of service in downstream AI agents? 000
Liran Tal @lirantal.com · 10/09/2026Snyk partnered with AI Engineer and other great companies and labs to put together the AI Security Summit in San Francisco on October 15th I'll be there. Are you coming? Hit me up and check out the event: aisecuritysummit.com 000
Liran Tal @lirantal.com · 10/09/2026running the Codex Security agent harness through a code base and btw just its Threat Modeling scope of work exceeded $2 in cost for a typical Golang application using Sol on High reasoning mode 000
Liran Tal @lirantal.com · 09/09/2026lol this is so good www.youtube.com/watch?v=7xOU...youtube.com'Darth Vader' makes the case for Flock cameras at city council meetingA man dressed as Star Wars villain Darth Vader spoke at a San Diego city council meeting on 19 August to voice his support of Flock cameras in the citySubscr... 010