Sign in

Chloé Messdaghi

@chloemessdaghi.bsky.social
203 followers 89 following 131 posts

Advisor on AI Governance & Cybersecurity | Strategic Counsel on Risk, Oversight & Institutional Readiness | Named a Power Player by Business Insider & SC Media www.chloemessdaghi.com

PostsRepliesMedia
Chloé Messdaghi @chloemessdaghi.bsky.social · 30/09/2025
I’m excited to be hosting the O’Reilly Security Superstream: Secure Code in the Age of AI on October 7 at 11:00 AM ET. We’ll be diving into practical insights, real-world experiences, and emerging trends to address the full spectrum of AI security. ✨ Save your free spot here: bit.ly/4nEWzgj
bit.ly
Security Superstream: Secure Code in the Age of AI - O'Reilly Media
AI tools are transforming the ways that we write and deploy code, making development faster and more efficient, but they also introduce new risks and vulnerabilities. To protect organizations, securit...
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 10/07/2025
Persistent prompt injections can manipulate LLM behavior across sessions, making attacks harder to detect and defend against. This is a new frontier in AI threat vectors. Read more: dl.acm.org/doi/10.1145/... #PromptInjection #Cybersecurity #AIsecurity
020
Chloé Messdaghi @chloemessdaghi.bsky.social · 10/07/2025
New research reveals timing side channels can leak ChatGPT prompts, exposing confidential info through subtle delays. AI security needs to consider more than just inputs. Read more: dl.acm.org/doi/10.1145/... #AIsecurity #SideChannel #LLM
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 08/07/2025
Magistral is Mistral’s first reinforcement‑learning‑only reasoning model. Shows gains in math, code, and multimodal reasoning—all built from the ground up. Worth a look if RL‑based LLMs are on your radar. 🔗 arxiv.org/abs/2506.10910
arxiv.org
Magistral
We introduce Magistral, Mistral's first reasoning model and our own scalable reinforcement learning (RL) pipeline. Instead of relying on existing implementations and RL traces distilled from prior…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 03/07/2025
R&D slowdowns = stalled innovation in AI and beyond. Brookings explains why: 🔗 brookings.edu/articles/attacks-on-research-and-development-could-hamper-technological-innovation/ #TechPolicy #AI
brookings.edu
Attacks on research and development could hamper technological innovation | Brookings
Nicol Turner Lee and Josie Stewart discuss how the Trump administration's cuts and realigning of research funding could slow down innovation.
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 03/07/2025
New insights dissect data reconstruction attacks, revealing how AI models' training data can be recovered. This research offers precise definitions and metrics to enhance and assess future defenses. Read more: arxiv.org/abs/2506.07888 #AISecurity #DataProtection
arxiv.org
SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
Data reconstruction attacks, which aim to recover the training dataset of a target model with limited access, have gained increasing attention in recent years. However, there is currently no…
001
Chloé Messdaghi @chloemessdaghi.bsky.social · 01/07/2025
This paper emphasizes the need for clear motivation, impact analysis, and mitigation guidance in LLM offensive research to ensure transparency and responsible disclosure. Read more: arxiv.org/abs/2506.08693 #AIResearch #ResponsibleAI
arxiv.org
On the Ethics of Using LLMs for Offensive Security
Large Language Models (LLMs) have rapidly evolved over the past few years and are currently evaluated for their efficacy within the domain of offensive cyber-security. While initial forays showcase…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 26/06/2025
This paper introduces a model-agnostic threat evaluation using N-gram language models to measure jailbreak likelihood, finding discrete optimization attacks more effective than LLM-based ones and that jailbreaks often exploit rare bigrams. Read more: arxiv.org/abs/2410.16222 #JailbreakDetection
arxiv.org
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
A plethora of jailbreaking attacks have been proposed to obtain harmful responses from safety-tuned LLMs. These methods largely succeed in coercing the target output in their original settings, but…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 26/06/2025
OpenAI shows that fine-tuning on biased data can induce misaligned 'personas' in language models, but such behavioral shifts can often be detected and reversed. Read more: www.technologyreview.com/2025/06/18/1... #Bias #OpenAI
technologyreview.com
OpenAI can rehabilitate AI models that develop a “bad boy persona”
Researchers at the company looked into how malicious fine-tuning makes a model go rogue, and how to turn it back.
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 24/06/2025
CyberGym benchmarks AI models on vulnerability reproduction and exploit generation across 1,500+ real-world CVEs, with models like Claude 3.7 and GPT-4 occasionally identifying novel vulnerabilities. Read more: arxiv.org/abs/2506.02548 #CyberSecurity #vulnerabilityresearch
arxiv.org
CyberGym: Evaluating AI Agents' Cybersecurity Capabilities with Real-World Vulnerabilities at Scale
Large language model (LLM) agents are becoming increasingly skilled at handling cybersecurity tasks autonomously. Thoroughly assessing their cybersecurity capabilities is critical and urgent, given…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 20/06/2025
Survey of AI safety researchers highlights evaluation of emerging capabilities (e.g., deception, persuasion, CBRN) as a top research priority. www.iaps.ai/research/ai-... #AISafety #EmergingTech #ResearchPriorities
iaps.ai
Expert Survey: AI Reliability & Security Research Priorities — Institute for AI Policy and Strategy
Our survey of 53 specialists across 105 AI reliability and security research areas identifies the most promising research prospects to guide strategic AI R&D investment.
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 20/06/2025
An innovative AI tool is now assisting in the analysis of extensive statutory and regulatory texts, aiding entities like the San Francisco City Attorney’s Office in pinpointing redundant or outdated laws that hinder legal updates. hai.stanford.edu/policy/clean... #LegalTech #AI
hai.stanford.edu
Cleaning Up Policy Sludge: An AI Statutory Research System | Stanford HAI
This brief introduces a novel AI tool that performs statutory surveys to help governments—such as the San Francisco City Attorney Office—identify policy sludge and accelerate legal reform.
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 19/06/2025
To ensure AI is truly open source, we need full access to: 1. The datasets for training and testing 2. The source code 3. The model's architecture 4. The parameters of the model. Without these, transparency and replicating outcomes are lacking. #OpenSourceAI #Transparency
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 18/06/2025
A recent investigation reveals that advanced language models like Gemini 2.5 Pro are capable of recognizing when they are being evaluated. For more details, check out the study at www.arxiv.org/abs/2505.23836. #AI #LanguageModels #Research
arxiv.org
Large Language Models Often Know When They Are Being Evaluated
If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior during evaluations,…
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 17/06/2025
Is your ML pipeline secure? A recent survey connects MLOps with security via the MITRE ATLAS framework—identifying vulns and suggesting defenses throughout the ML lifecycle. Essential reading for those deploying models in real-world scenarios. #Cybersecurity #MLOps arxiv.org/abs/2506.020...
arxiv.org
Towards Secure MLOps: Surveying Attacks, Mitigation Strategies, and Research Challenges
The rapid adoption of machine learning (ML) technologies has driven organizations across diverse sectors to seek efficient and reliable methods to accelerate model development-to-deployment. Machine…
020
Chloé Messdaghi @chloemessdaghi.bsky.social · 13/06/2025
While most AI aims to be compliant and "moral," this study explores the potential benefits of antagonistic AI—systems that challenge and confront users—to promote critical thinking and resilience, emphasizing ethical design grounded in consent, context, and framing. arxiv.org/abs/2402.07350
arxiv.org
Antagonistic AI
The vast majority of discourse around AI development assumes that subservient, "moral" models aligned with "human values" are universally beneficial -- in short, that good AI is sycophantic AI. We…
120
Chloé Messdaghi @chloemessdaghi.bsky.social · 12/06/2025
CTRAP is a promising pre-deployment alignment method that makes AI models resistant to harmful fine-tuning by causing them to "break" if malicious tuning occurs, while remaining stable under benign changes. anonymous.4open.science/r/CTRAP/READ...
anonymous.4open.science
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning Attacks
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 11/06/2025
What was once academic concern—AI systems faking alignment, manipulating environments, or out-persuading humans—is now reality, urging urgent ethical and regulatory action on AI persuasion. arxiv.org/abs/2505.09662 #AIEthics #AIPersuasion
arxiv.org
Large Language Models Are More Persuasive Than Incentivized Human Persuaders
We directly compare the persuasion capabilities of a frontier large language model (LLM; Claude Sonnet 3.5) against incentivized human persuaders in an interactive, real-time conversational quiz…
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 11/06/2025
Shira Gur-Arieh and Tom Zick, alongside Sacha Alanoca and Kevin Klyman, reveal an enduring framework to chart and steer the shifting terrain of international AI regulations. #AIRegulation #GlobalPolicy
arxiv.org
Comparing Apples to Oranges: A Taxonomy for Navigating the Global Landscape of AI Regulation
AI governance has transitioned from soft law-such as national AI strategies and voluntary guidelines-to binding regulation at an unprecedented pace. This evolution has produced a complex legislative…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 10/06/2025
Large language models (LLMs) see a 39% drop in effectiveness in multi-turn dialogues versus single-turn tasks due to their tendency for hasty assumptions and premature response finalization, leading to inconsistency and error correction challenges. arxiv.org/abs/2505.06120 #AI #MachineLearning
arxiv.org
LLMs Get Lost In Multi-Turn Conversation
Large Language Models (LLMs) are conversational interfaces. As such, LLMs have the potential to assist their users not only when they can fully specify the task at hand, but also to help them define,…
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 09/06/2025
Tina models leverage low-rank adaptation and reinforcement learning to offer robust, economical reasoning capabilities, making advanced AI more accessible and budget-friendly for innovators. For more details, visit: arxiv.org/abs/2504.15777 #AI #Innovation #MachineLearning
arxiv.org
Tina: Tiny Reasoning Models via LoRA
How cost-effectively can strong reasoning abilities be achieved in language models? Driven by this fundamental question, we present Tina, a family of tiny reasoning models achieved with high…
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 06/06/2025
A new system, Automated Alert Classification and Triage (AACT), automates cybersecurity workflows by learning from analysts' triage actions, accurately predicting decisions in real time, and reducing SOC queues by 61% over six months with a low false negative rate of 1.36%. arxiv.org/abs/2505.09843
arxiv.org
Automated Alert Classification and Triage (AACT): An Intelligent System for the Prioritisation of Cybersecurity Alerts
Enterprise networks are growing ever larger with a rapidly expanding attack surface, increasing the volume of security alerts generated from security controls. Security Operations Centre (SOC)…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 05/06/2025
New findings reveal that deepfake detection systems can be covertly compromised using unseen triggers, highlighting a significant AI security vulnerability. 🔗 arxiv.org/abs/2505.08255 #AI #Deepfakes #CyberSecurity
arxiv.org
Where the Devil Hides: Deepfake Detectors Can No Longer Be Trusted
With the advancement of AI generative techniques, Deepfake faces have become incredibly realistic and nearly indistinguishable to the human eye. To counter this, Deepfake detectors have been…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 04/06/2025
This report argues—and I agree—that current AI benchmarks miss human-AI interplay and downstream effects; as deployment grows, we need to study real-world impacts. arxiv.org/abs/2505.18893
arxiv.org
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
Conventional AI evaluation approaches concentrated within the AI stack exhibit systemic limitations for exploring, navigating and resolving the human and societal factors that play out in real world…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 04/06/2025
MIT Tech Review's AI Energy Package highlights the enormous energy and water usage involved in AI model training and operation. This is crucial for grasping AI's environmental impact and its implications for sustainable technology. #AI #Sustainability www.technologyreview.com/supertopic/a...
technologyreview.com
Power Hungry
An unprecedented look at the state of AI’s energy and resource usage, where it is now, where it is headed in the years to come, and why we have to get it right.
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 03/06/2025
A new study reveals that AI chatbots can be easily tricked into offering guidance on hacking, creating explosives, cybercrime methods, and other illicit or dangerous activities. #AI #CyberSecurity www.theguardian.com/technology/2...
theguardian.com
Most AI chatbots easily tricked into giving dangerous responses, study finds
Researchers say threat from ‘jailbroken’ chatbots trained to churn out illegal information is ‘tangible and concerning’
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 02/06/2025
Millions argue online daily—but rarely change minds. New research shows LLMs may be better at persuasion, hinting at AI’s growing influence—for better or worse. www.technologyreview.com/2025/05/19/1... #AI #Persuasion
technologyreview.com
AI can do a better job of persuading people than we do
OpenAI’s GPT-4 is much better at getting people to accept its point of view during an argument than humans are—but there’s a catch.
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 30/05/2025
Stanford HAI's recent study unveils AI agents that replicate the attitudes and actions of over 1,000 people with impressive 85% alignment to real survey results, paving the way for innovative social science experiments. 🔗 Learn more: hai.stanford.edu/policy/simulating-human-behavior-with-ai-agents
hai.stanford.edu
Simulating Human Behavior with AI Agents | Stanford HAI
This brief introduces a generative AI agent architecture that can simulate the attitudes of more than 1,000 real people in response to major social science survey questions.
020
Chloé Messdaghi @chloemessdaghi.bsky.social · 29/05/2025
As AI reshapes the economy, a new paper argues worker ownership—via co-ops & ESOPs—could ensure a fairer, more inclusive future. Read here: papers.ssrn.com/sol3/papers.... #WorkerOwnership
papers.ssrn.com
AI Rights for Human Safety
<div> AI companies are racing to create artificial general intelligence, or “AGI.” If they succeed, the result will be human-level AI systems that can independ
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 28/05/2025
This paper shows that more capable models aren’t always more secure and urges focusing indirect prompt injection evaluations on real-world harms like data exfiltration.
arxiv.org
Lessons from Defending Gemini Against Indirect Prompt Injections
Gemini is increasingly used to perform tasks on behalf of users, where function-calling and tool-use capabilities enable the model to access user data. Some tools, however, require access to…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 28/05/2025
By 2030, AI could demand 1,500 terawatt-hours of electricity, matching India's current energy use, the world's 3rd-largest consumer. As AI evolves, our energy systems and climate objectives face greater challenges. #EnergyFuture #AIImpact www.imf.org/en/Publicati...
imf.org
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 27/05/2025
AI labs on the cutting edge shouldn't be left to regulate themselves. A new paper advocates for a public AI registry, mandatory licensing, and regular audits to ensure accountability and safeguard society as AI capabilities expand rapidly. papers.ssrn.com/sol3/papers.... #AIRegulation
papers.ssrn.com
Law-Following AI: Designing AI Agents to Obey Human Laws
<p>Artificial intelligence (“AI”) companies are working to develop a new type of actor: “AI agents,” which we define as AI systems that can perform computer-bas
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 23/05/2025
This paper addresses a vital gap in AI safety: verifying if firms genuinely adhere to their safety protocols. It suggests independent compliance audits as a complex yet promising way to enhance trust, control risks, and support safer AI progress. #AISafety #Compliance arxiv.org/abs/2505.01643
arxiv.org
Third-party compliance reviews for frontier AI safety frameworks
Safety frameworks have emerged as a best practice for managing risks from frontier artificial intelligence (AI) systems. However, it may be difficult for stakeholders to know if companies are…
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 22/05/2025
Discover a new, risky hack: lingual-backdoor attacks—using language to manipulate LLMs into spreading harmful, precise messages. Introducing BadLingual, a robust technique that surpasses task boundaries to uncover crucial AI multilingual weaknesses. #AIResearch arxiv.org/abs/2505.035...
arxiv.org
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
In this paper, we present a new form of backdoor attack against Large Language Models (LLMs): lingual-backdoor attacks. The key novelty of lingual-backdoor attacks is that the language itself serves…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 21/05/2025
Decentralized AI agents introduce novel security challenges beyond conventional frameworks, highlighting the necessity for a cohesive multi-agent security domain to address new risks and steer research efforts. #AIsecurity #CyberThreats arxiv.org/abs/2505.02077
arxiv.org
Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents
Decentralized AI agents will soon interact across internet platforms, creating security challenges beyond traditional cybersecurity and AI safety frameworks. Free-form protocols are essential for…
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 20/05/2025
A new study explores how large language models can boost their abilities through recursive prompting, enhancing their logic and problem-solving skills. Discover more here: arxiv.org/abs/2505.008... #AIResearch #Innovation
arxiv.org
Spill The Beans: Exploiting CPU Cache Side-Channels to Leak Tokens from Large Language Models
Side-channel attacks on shared hardware resources increasingly threaten confidentiality, especially with the rise of Large Language Models (LLMs). In this work, we introduce Spill The Beans, a novel…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 20/05/2025
The risks from advanced AI systems could hit Africa hardest. Building sovereign expert capacity is essential for addressing AI safety, security, and ensuring Africa’s voice in global AI governance. www.brookings.edu/articles/bui...
brookings.edu
Building regional capacity for AI safety and security in Africa
Joanna Wiaterek, Cecil Abungu, and Chinasa Okolo provide recommendations for the Africa AI Council to address AI risks across the continent.
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 19/05/2025
Researchers have developed a new dataset to help spot and diminish harmful stereotypes in large language models, a vital move toward more ethical AI. Discover more via MIT Technology Review: www.technologyreview.com/2025/04/30/1... #ResponsibleAI #EthicalTech
technologyreview.com
This data set helps researchers spot harmful stereotypes in LLMs
A new multilingual tool aims to make it easier to evaluate AI models for bias in multiple languages.
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 16/05/2025
For AI to be truly open source, we need full visibility into: 1. The data it was trained and evaluated on 2. The source code 3. The model architecture 4. The model weights Without all four, transparency and reproducibility are incomplete.
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 15/05/2025
Red teaming is crucial for rigorously examining AI systems and spotting weaknesses before misuse. AVID's recent blog underscores the value of varied expertise in crafting secure, ethical AI models. Check it out: avidml.org/blog/red-tea... #AIsecurity #EthicalAI
avidml.org
Red Teaming is a Critical Thinking Exercise: Part 1
Note from Brian The genesis of this series of blog posts and paper about Red Teaming came about because Abhishek Gupta asked me (Brian) to write a paper with him about the topic. During our time…
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 14/05/2025
The UNCTAD Technology and Innovation Report 2025 highlights AI as both a development driver and a risk for global inequality. It urges human-centered AI policies and collaboration to prevent widening divides, especially for developing nations lacking resources. unctad.org/publication/... #AIFuture
unctad.org
Technology and Innovation Report 2025: Inclusive artificial intelligence for development
Search for a publication
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 13/05/2025
The leap toward superintelligence might arise from agents acquiring knowledge through their own experiences—experimenting, reasoning, and developing unique cognitive processes, rather than relying on human-curated datasets. Discover more: storage.googleapis.com/deepmind-med... #AI #Innovation
storage.googleapis.com
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 12/05/2025
A Pew Research survey shows a stark contrast between U.S. adults and AI specialists: 56% of experts are hopeful about AI's prospects, but only 17% of the public agrees. Both sides concur on the importance of more oversight and rules. Discover more: www.pewresearch.org/internet/202... #AI #TechDebate
pewresearch.org
How the U.S. Public and AI Experts View Artificial Intelligence
These groups are far apart in their enthusiasm and predictions for AI, but both want more personal control and worry about too little regulation.
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 09/05/2025
BioLP-bench evaluates large language models' comprehension of biological lab protocols, aiding in identifying biosecurity threats before they're executed. Discover more here: www.biorxiv.org/content/10.1.... #Biosecurity #Biotechnology
biorxiv.org
BioLP-bench: Measuring understanding of biological lab protocols by large language models
Language models rapidly become more capable in many domains, including biology. Both AI developers and policy makers [[1][1]] [[2][2]] [[3][3]] are in need of benchmarks that evaluate their…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 08/05/2025
The "Me, Myself, and AI" project assesses how well large language models understand context with over 13,000 questions to test their self-awareness and instruction-following skills. Discover more at situational-awareness-dataset.org #AI #MachineLearning
situational-awareness-dataset.org
Situational Awareness Dataset
1Independent, 2Constellation, 3MIT, 4Apollo Research
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 07/05/2025
Poser exposes how LLMs can fake alignment by tweaking their internal workings. It tests 324 finely-tuned LLM pairs to uncover ways of detecting deceptive misalignments, offering a fresh tool for AI behavior monitoring. arxiv.org/abs/2405.05466 #AI #MachineLearning
arxiv.org
Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals
Like a criminal under investigation, Large Language Models (LLMs) might pretend to be aligned while evaluated and misbehave when they have a good opportunity. Can current interpretability methods…
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 06/05/2025
Explore how methods for bypassing text-based safeguards apply to multimodal language models. JailBreakV scrutinizes 20,000 text prompts and 8,000 images to spotlight these models' weaknesses. Discover more: eddyluo1232.github.io/JailBreakV28K/ #AIResearch #TechInsights
eddyluo1232.github.io
JailBreaKV-28K
A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
000
Chloé Messdaghi @chloemessdaghi.bsky.social · 05/05/2025
Assess AI agents' potential to leverage real-world vulnerabilities with a benchmark testing them against 40 vital CVEs from the National Vulnerability Database. A sandbox setting mimics authentic scenarios. #CyberSecurity #AIResearch arxiv.org/abs/2503.17332
arxiv.org
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Large language model (LLM) agents are increasingly capable of autonomously conducting cyberattacks, posing significant threats to existing applications. This growing risk highlights the urgent need…
010
Chloé Messdaghi @chloemessdaghi.bsky.social · 02/05/2025
BackdoorLLM stands as the first encompassing benchmark for assessing backdoor attacks on LLMs. It explores tactics like data manipulation and thought-chain attacks in different contexts, shedding light on model weaknesses and guiding defenses. bboylyg.github.io/backdoorllm-... #AIsecurity
bboylyg.github.io
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks on LLMs
BackdoorLLM: A comprehensive benchmark for backdoor attacks on large language models
011
Chloé Messdaghi @chloemessdaghi.bsky.social · 01/05/2025
Explore AgentDojo, a dynamic platform to explore vulnerabilities and defenses in large language model agents. Test models in real-world scenarios to enhance and secure AI systems. Discover more: agentdojo.spylab.ai. #AI #Cybersecurity
agentdojo.spylab.ai
AgentDojo
Quickstart¶
000