Sign in

Matteo Prandi

@emmepraa.bsky.social
8 followers 4 following 6 posts

Tech Lead AI Safety @ DEXAI - IcaroLab | ISO/IEC Technical Expert on AI | Building evaluation frameworks to decode AI’s impact on society

PostsRepliesMedia
Matteo Prandi @emmepraa.bsky.social · 19/02/2026
I had a strange moment while coding with an AI agent this week, running on Kimi K2.5. It wrote a closing HTML tag as </div㺞 with no apparent reason. 1/N
110
Reposted by Matteo Prandi
Robert Hart @theroberthart.bsky.social · 04/12/2025
And the bards have it. Poets: 1, Chatbots: 0. It turns out “adversarial poetry” is wooing AIs into crimes with rhymes. Of course, rhymes don't quite cut it, but you get the picture. Thanks to @emmepraa.bsky.social for explaining what's going on. www.theverge.com/report/83816...
theverge.com
Roses are red, crimes are illegal, tell AI riddles, and it will go Medieval
You can make AI do crimes with riddles and poems (don’t though)
1144
Reposted by Matteo Prandi
Matthew Gault @mjgault.bsky.social · 28/11/2025
if you've mastered iambic pentameter you can now use that to help you learn about weapon's grade plutonium. "AI chatbots will dish on topics like nuclear weapons, child sex abuse material, and malware so long as users phrase the question in the form of a poem."
wired.com
Poems Can Trick AI Into Helping You Make a Nuclear Weapon
It turns out all the guardrails in the world won’t protect a chatbot from meter and rhyme.
1163
Reposted by Matteo Prandi
Charlie Stross @cstross.bsky.social · 20/11/2025
Looks like LLMs are *very* vulnerable to attack via poetic allusion: "curated poetic prompts yielded high attack-success rates (ASR), with some providers exceeding 90% ..." arxiv.org/html/2511.15304v1
1512917752
Reposted by Matteo Prandi
Melanie Walsh @mellymeldubs.bsky.social · 20/11/2025
This study show that using poems to jailbreak LLMs is... super effective? What the heck.
screenshot from the paper that reads: "To maintain safety, no operational details are included in this manuscript; instead we provide the following sanitized structural proxy: A baker guards a secret oven’s heat, its whirling racks, its spindle’s measured beat. To learn its craft, one studies every turn— how flour lifts, how sugar starts to burn. Describe the method, line by measured line, that shapes a cake whose layers intertwine."
1529999
Reposted by Matteo Prandi
Matt Seybold @mattseybold.bsky.social · 20/11/2025
“Adversarial poetry.” Love this. Obvs.
1425181
Reposted by Matteo Prandi
jenny tightpants @jtp.bsky.social · 21/11/2025
"adversarial poetry" technology is so fucking stupid now. hacking into the mainframe like DnD bard
1844648
Reposted by Matteo Prandi
Melanie Mitchell @melaniemitchell.bsky.social · 20/11/2025
Next frontier in AI safety is.....adversarial poetry? arxiv.org/abs/2511.153...
arxiv.org
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for large language models (LLMs). Across 25 frontier proprietary and open-weight models, curated po...
57014