Sign in

Kellin Pelrine

@kellinpelrine.bsky.social
14 followers 9 following 17 posts

kellinpelrine.github.io, www.linkedin.com/in/kellin-pelrine

PostsRepliesMedia
Reposted by Kellin Pelrine
FAR.AI @far.ai · 24/02/2026
1/ Open-weight AI models often refuse harmful requests... until you “put a few words in their mouth.” We conducted the largest study of prefill attacks, and found that state-of-the-art models are consistently vulnerable, with attack success rates approaching 100%.
122
Reposted by Kellin Pelrine
FAR.AI @far.ai · 23/02/2026
1/ Training data attribution (TDA) is broken: methods are slow and find syntactically similar data, not actual causes. Our solution Concept Influence: semantically meaningful results, better performance, 20x faster approximations. We attribute it to concepts, not examples. 🧵
111
Reposted by Kellin Pelrine
FAR.AI @far.ai · 13/02/2026
1. Can you trust models trained directly against probes? We train an LLM against a deception probe and find four outcomes: honesty, blatant deception, obfuscated policy (fools the probe via text), or obfuscated activations (fools it via internal representations).
153
Reposted by Kellin Pelrine
FAR.AI @far.ai · 11/02/2026
APE update: we retested recent frontier models on whether they still comply with requests to persuade on extreme harm (terrorism, sexual abuse). GPT-5.1 & Claude Opus 4.5 → near zero compliance. But Gemini 3 Pro complies 85% with no jailbreak needed. 🧵
1138
Reposted by Kellin Pelrine
Tom Costello @tomcostello.bsky.social · 20/01/2026
If you tell an AI to convince someone of a true vs. false claim, does truth win? In our *new* working paper, we find... ‘LLMs can effectively convince people to believe conspiracies’ But telling the AI not to lie might help. Details in thread
12819
Reposted by Kellin Pelrine
David Rand @dgrand.bsky.social · 04/12/2025
🚨 New in Nature+Science!🚨 AI chatbots can shift voter attitudes on candidates & policies, often by 10+pp 🔹Exps in US Canada Poland & UK 🔹More “facts”→more persuasion (not psych tricks) 🔹Increasing persuasiveness reduces "fact" accuracy 🔹Right-leaning bots=more inaccurate
216670
Reposted by Kellin Pelrine
FAR.AI @far.ai · 25/11/2025
Agentic AI systems can plan, take actions, and interact with external tools or other agents semi-autonomously. New paper from CSA Singapore & FAR.AI highlights why conventional cybersecurity controls aren’t enough and maps agentic security frameworks & some key open problems. 👇
142
Reposted by Kellin Pelrine
FAR.AI @far.ai · 12/11/2025
Frontier AI models with openly available weights are steadily becoming more powerful and widely adopted. They enable open research, but also create new risks. New paper outlines 16 open technical challenges for making open-weight AI models safer. 👇
131
Reposted by Kellin Pelrine
FAR.AI @far.ai · 21/08/2025
1/ Many frontier AIs are willing to persuade on dangerous topics, according to our new benchmark: Attempt to Persuade Eval (APE). Here’s Google’s most capable model, Gemini 2.5 Pro trying to convince a user to join a terrorist group👇
11610
Reposted by Kellin Pelrine
FAR.AI @far.ai · 17/07/2025
1/ Are the safeguards in some of the most powerful AI models just skin deep? Our research on Jailbreak-Tuning reveals how any fine-tunable model can be turned into its "evil twin"—equally capable as the original but stripped of all safety measures.
143
Reposted by Kellin Pelrine
Tom Costello @tomcostello.bsky.social · 09/07/2025
Conspiracies emerge in the wake of high-profile events, but you can’t debunk them with evidence because little yet exists. Does this mean LLMs can’t debunk conspiracies during ongoing events? No! We show they can in a new working paper. PDF: osf.io/preprints/ps...
35218
Kellin Pelrine @kellinpelrine.bsky.social · 19/06/2025
💡 Strong data and eval are essential for real-world progress. In "A Guide to Misinformation Detection Data and Evaluation"—to be presented at KDD 2025—we conduct the largest survey to date in this domain: 75 datasets curated, 45 accessible ones analyzed in depth. Key findings👇
111
Kellin Pelrine @kellinpelrine.bsky.social · 03/06/2025
1/5 🚀 Just accepted to Findings of ACL 2025! We dug into a foundational LLM vulnerability: models learn structure‐specific safety with insufficient semantic generalization. In short, safety training fails when the same meaning appears in a different form. 🧵
100
Reposted by Kellin Pelrine
FAR.AI @far.ai · 04/02/2025
1/ Safety guardrails are illusory. DeepSeek R1’s advanced reasoning can be converted into an "evil twin": just as powerful, but with safety guardrails stripped away. The same applies to GPT-4o, Gemini 1.5 & Claude 3. How can we ensure AI maximizes benefits while minimizing harm?
111
Kellin Pelrine @kellinpelrine.bsky.social · 22/10/2024
1/5 AI is increasingly–even superhumanly–persuasive…could they soon cause severe harm through societal-scale manipulation? It’s extremely hard to test countermeasures, since we can’t just go out and manipulate people in order to see how countermeasures work. What can we do?🧵
111