Reposted by Torr Vision Group OxfordAlasdair Paren @alasdair-p.bsky.social · 04/09/2025www.scientificamerican.com/article/hack... New article by Deni Bechard at Scientific America covering our work on hijacking Multimodal computer agents published on Arxiv earlier this year. A massive effort by Lukas Aichberger, supported by myself Yarin Gal, Philip Torr, FREng, FRS & Adel Bibi 011
Reposted by Torr Vision Group OxfordFazl Barez @fbarez.bsky.social · 01/07/2025Excited to share our paper: "Chain-of-Thought Is Not Explainability"! We unpack a critical misconception in AI: models explaining their steps (CoT) aren't necessarily revealing their true reasoning. Spoiler: the transparency can be an illusion. (1/9) 🧵 28531
Reposted by Torr Vision Group OxfordProfessor Philip Torr, Oxford University @philiptorr.bsky.social · 10/06/2025Youtuber Sabine Hossenfelder has picked up on our paper on how to fool AI agents, you can watch her desrcibing our work here: www.youtube.com/watch?v=KY7_....youtube.comAI is becoming dangerous. Are we ready?YouTube video by Sabine Hossenfelder 011
Reposted by Torr Vision Group OxfordC Emde @cemde.bsky.social · 04/04/2025🚨 New paper alert: Our recent work on LLM safety has been accepted to ICLR 2025 🇸🇬 We propose a new framework for LLMs safety. 🧵 (1/7) #LLM #AISafety #ICLR2025 #Certification #AdversarialRobustness #NLP #Shhhhhh #DomainCertification #AImedia.tenor.coma man in a suit and tie is sitting at a desk in front of a computer screen that says founder of the office .ALT: a man in a suit and tie is sitting at a desk in front of a computer screen that says founder of the office . 121
Reposted by Torr Vision Group OxfordProfessor Philip Torr, Oxford University @philiptorr.bsky.social · 03/04/2025Excited to be working with the UK Excited to be work with the UK AI Security Institute, even more important than normal in these turbulent times. www.linkedin.com/posts/philip...linkedin.comStrengthening AI Resilience | AISI Work | Philip TorrExcited to be work with the UK AI Security Institute, even more important than normal in these turbulent times. 021
Reposted by Torr Vision Group OxfordLukas Aichberger @aichberger.bsky.social · 18/03/2025⚠️ Beware: Your AI assistant could be hijacked just by encountering a malicious image online! Our latest research exposes critical security risks in AI assistants. An attacker can hijack them by simply posting an image on social media and waiting for it to be captured. [1/6] 🧵 188
Reposted by Torr Vision Group OxfordFazl Barez @fbarez.bsky.social · 10/01/2025 🚨 New Paper Alert: Open Problem in Machine Unlearning for AI Safety 🚨 Can AI truly "forget"? While unlearning promises data removal, controlling emergent capabilities is a inherent challenge. Here's why it matters: 👇 Paper: arxiv.org/pdf/2501.04952 1/8 1256
Reposted by Torr Vision Group OxfordFrancesco Pinto @ Neurips2024 @frapintoml.bsky.social · 09/12/2024🧵 [1/3] Heading to #Vancouver 🇨🇦 tomorrow to present our latest work in @OxfordTVG #UniversityOfOxford at #NeurIPS2024 🧠: - 💥 Improving on #StylizedImageNet, use #IllusionBench: can you see the cat 🐈⬛ Hidden in Plain Sight in the picture 🖼️? Paper: arxiv.org/abs/2411.06287 131