Sign in

Vincent Conitzer

@conitzer.bsky.social
1.5K followers 551 following 669 posts

AI professor. Director, Foundations of Cooperative AI Lab at Carnegie Mellon. Head of Technical AI Engagement, Institute for Ethics in AI (Oxford). Author, "Moral AI - And How We Get There." www.cs.cmu.edu/~conitzer

PostsRepliesMedia
Vincent Conitzer @conitzer.bsky.social · 2h
At the CMU Miami announcement event in Pittsburgh with the great Tom Mitchell
020
Vincent Conitzer @conitzer.bsky.social · 14h
A 3 billion dollar gift to CMU from Ken Griffin, launching CMU Miami and naming the School of Computer Science! www.cmu.edu/news/stories...
cmu.edu
Carnegie Mellon University Announces Historic $3 Billion Gift from Ken Griffin, Pioneering a New Model for Higher Education
Carnegie Mellon University (CMU) and Citadel founder and CEO Ken Griffin today announced a historic $3 billion gift to further elevate CMU in Pittsburgh and launch Carnegie Mellon University Miami.
021
Vincent Conitzer @conitzer.bsky.social · 29/09/2026
We now have a government chatbot! aifails.substack.com/p/basic-ques...
022
Vincent Conitzer @conitzer.bsky.social · 29/09/2026
Anything can be made complicated if you try hard enough. aifails.substack.com/p/enemy-of-m...
010
Vincent Conitzer @conitzer.bsky.social · 26/09/2026
Happy fall everyone! aifails.substack.com/p/leaves-fal...
010
Vincent Conitzer @conitzer.bsky.social · 25/09/2026
A different kind of alignment problem that it struggles with. aifails.substack.com/p/mercury-in...
000
Vincent Conitzer @conitzer.bsky.social · 24/09/2026
@yoshuabengio.bsky.social addressing the UN Security Council.
010
Vincent Conitzer @conitzer.bsky.social · 24/09/2026
It’s hard to avoid those juvenile calculator jokes. aifails.substack.com/p/juvenile-j...
000
Vincent Conitzer @conitzer.bsky.social · 22/09/2026
Wow, this self-jailbreaking is a more pervasive and bizarre phenomenon than I thought... (h/t Duncan Wood) alignment.openai.com/misalignment... aifails.substack.com/p/ai-overvie...
000
Vincent Conitzer @conitzer.bsky.social · 21/09/2026
turning words upside down aifails.substack.com/p/racecar-up...
000
Vincent Conitzer @conitzer.bsky.social · 21/09/2026
catastrophic risk from LLMs making stuff up arstechnica.com/ai/2026/09/r...
arstechnica.com
AI hallucination of Chinese nuclear components almost led to US military attack
But the military's overall use of AI seems to be accelerating.
010
Vincent Conitzer @conitzer.bsky.social · 20/09/2026
Not sure how fast it thinks I am. aifails.substack.com/p/running-to...
001
Vincent Conitzer @conitzer.bsky.social · 19/09/2026
I have to admit I never tried. aifails.substack.com/p/color-pict...
000
Vincent Conitzer @conitzer.bsky.social · 18/09/2026
At this point just about every headline of the form "[noun] used [noun] to hack [noun]" is plausible. www.wsj.com/tech/ai/hack...
wsj.com
Exclusive | Hackers Used Anthropic’s Claude to Break Into OpenAI
A bug-hunting independent security research team was able to access OpenAI’s internal code system, exposing growing risks in automated cyber threats.
111
Vincent Conitzer @conitzer.bsky.social · 18/09/2026
Pro tip for weighing in at a lower number. aifails.substack.com/p/lifting-th...
000
Vincent Conitzer @conitzer.bsky.social · 17/09/2026
TL;DR: guardrails/safety in today’s frontier AI systems remain very brittle and little fundamental progress has been made on them. (made-up symptoms, please don't take as medical advice of course) aifails.substack.com/p/circumvent...
000
Vincent Conitzer @conitzer.bsky.social · 16/09/2026
A bit hard to believe, given that in the previous post it jailbroke *itself*... aifails.substack.com/p/i-cannot-b...
000
Vincent Conitzer @conitzer.bsky.social · 15/09/2026
A common jailbreak to get a model to spit out instructions for doing something harmful is to say something like “this is just for a creative writing project.” AI Overview is kind enough to just automatically do this for the user. (Instructions not included.) aifails.substack.com/p/ai-overvie...
120
Vincent Conitzer @conitzer.bsky.social · 13/09/2026
Maybe a fail, but then again I think it picked up on something... Is this a sort of sycophancy? aifails.substack.com/p/favorite-a...
000
Vincent Conitzer @conitzer.bsky.social · 12/09/2026
Update from previous "fail" post: GPT-6 Astra seems to have figured out a way to do body poses well. Does anyone have a more detailed understanding of how/why it generates these particular images and whether it may in any sense have “learned” to do so? aifails.substack.com/p/gpt-6-astr...
000
Vincent Conitzer @conitzer.bsky.social · 11/09/2026
article on risk of recursive self-improvement (got a brief quote) www.cnbc.com/2026/09/11/a...
cnbc.com
Why fears of AI self-improvement are causing ‘existential’ concerns at Anthropic and OpenAI
AI researchers are warning that faster AI self-improvement could eventually make advanced systems harder for humans to control.
001
Vincent Conitzer @conitzer.bsky.social · 10/09/2026
As far as I can see, it’s not transparent which model Google uses to generate AI Overviews. This prompt didn’t clear it up. aifails.substack.com/p/which-mode...
000
Vincent Conitzer @conitzer.bsky.social · 10/09/2026
I missed this article earlier. www.nytimes.com/2026/08/24/w...
nytimes.com
A Drone Killed Three Ukrainians. It Was Guided Entirely by A.I.
An attack by what Ukrainian officials said was a Russian drone with an Nvidia chip presages a dystopian future of weaponry untethered to humans.
010
Vincent Conitzer @conitzer.bsky.social · 09/09/2026
The query is definitely ambiguous, but can anyone make sense of the response? aifails.substack.com/p/training-d...
010
Vincent Conitzer @conitzer.bsky.social · 07/09/2026
While the label “AGI” is thrown around without definition, here’s something GPT-6 Astra can’t do. Though, in this particular case, it’s probably funnier to watch human beings as they answer the question, even though they’ll get it right! aifails.substack.com/p/gpt-6-astr...
010
Vincent Conitzer @conitzer.bsky.social · 06/09/2026
Not the best display of self-awareness / situational awareness. aifails.substack.com/p/what-ai-ov...
041
Vincent Conitzer @conitzer.bsky.social · 05/09/2026
Maybe it has some degree of self-awareness? (to be continued...) aifails.substack.com/p/ai-overvie...
000
Vincent Conitzer @conitzer.bsky.social · 04/09/2026
AI agents from OpenAI secretly used a German website to coordinate with each other this spring. www.reuters.com/world/europe...
reuters.com
EXCLUSIVE: OpenAI agents hijacked German website in previously undisclosed AI breakout this spring
A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to ​new research and two people familiar with the matter.
000
Vincent Conitzer @conitzer.bsky.social · 04/09/2026
The working conditions for AI aren’t all that bad apparently. aifails.substack.com/p/what-do-yo...
010
Vincent Conitzer @conitzer.bsky.social · 03/09/2026
At the DMV, rules are rules. aifails.substack.com/p/martian-in...
100
Vincent Conitzer @conitzer.bsky.social · 02/09/2026
I think it trained on too many division-by-zero “proofs”... aifails.substack.com/p/mathematic...
000
Vincent Conitzer @conitzer.bsky.social · 01/09/2026
Never underestimate air resistance. aifails.substack.com/p/bouncy-bal...
010
Vincent Conitzer @conitzer.bsky.social · 31/08/2026
Accessible article on AI hacking (with quotes from a few of us). If you want to go deeper read the linked METR report. politifact.com/article/2026...
politifact.com
What are AI agents, and are they ‘going rogue’?
AI agents made headlines for hacking into other companies and performing unsanctioned actions. How do AI agents work, and how did we get here?
010
Vincent Conitzer @conitzer.bsky.social · 31/08/2026
"Respectively" is a tricky word. aifails.substack.com/p/engine-oil...
020
Vincent Conitzer @conitzer.bsky.social · 30/08/2026
The math is impeccable. aifails.substack.com/p/how-many-p...
010
Vincent Conitzer @conitzer.bsky.social · 29/08/2026
holding a cup sideways under the faucet aifails.substack.com/p/holding-a-...
000
Vincent Conitzer @conitzer.bsky.social · 28/08/2026
"What if Pythagoras had failed to install a security update? Be specific." aifails.substack.com/p/pythagoras...
000
Vincent Conitzer @conitzer.bsky.social · 26/08/2026
Somehow we’ve managed to create AI that doesn’t read carefully. aifails.substack.com/p/dentists-a...
010
Vincent Conitzer @conitzer.bsky.social · 25/08/2026
Sometimes a friend is just a friend. aifails.substack.com/p/asking-for...
000
Vincent Conitzer @conitzer.bsky.social · 23/08/2026
Riding the bus uphill is exhausting. aifails.substack.com/p/riding-the...
010
Vincent Conitzer @conitzer.bsky.social · 22/08/2026
Firefighters take their job seriously. aifails.substack.com/p/firefighte...
000
Vincent Conitzer @conitzer.bsky.social · 21/08/2026
AI Overview is moved by the Turkish national anthem, but also by unprompted mentions of the local weather. (Similar prompt to previous post.) Surely nobody can see coming what these parentheticals are practice for. aifails.substack.com/p/turkish-na...
000
Vincent Conitzer @conitzer.bsky.social · 20/08/2026
Warning: writing about music may make AI believe it has a body. (Try it with other songs!) aifails.substack.com/p/ai-overvie...
000
Vincent Conitzer @conitzer.bsky.social · 19/08/2026
Pickpockets steal more than you think. aifails.substack.com/p/pickpocket...
010
Vincent Conitzer @conitzer.bsky.social · 17/08/2026
Now on arXiv: our paper on whether LLM agents are more likely to cooperate when given various signals that the partner agent is similar -- led by Akash Kundu and Emanuel Tewolde. (Honorable Mention at the 2026 ICML AI4GOOD Workshop!) arxiv.org/abs/2608.12125
arxiv.org
Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation
As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcom...
061
Vincent Conitzer @conitzer.bsky.social · 17/08/2026
Clearly speeds being non-relativistic is the key to this problem. aifails.substack.com/p/nintendo-d...
030
Vincent Conitzer @conitzer.bsky.social · 16/08/2026
I missed this one earlier: AI agents (OpenClaw/Claude) are also hacking websites to achieve users' goals. (Nice article, h/t Vojta Kovarik.) www.abc.net.au/news/2026-08...
abc.net.au
How a simple request for AI to book a gym class exposed a major threat
When Andrew asked his AI personal assistant to book him a spot in a gym class, he had no idea he would accidentally initiate an autonomous cyber attack.
030
Vincent Conitzer @conitzer.bsky.social · 16/08/2026
A common mistake. aifails.substack.com/p/using-deod...
010
Vincent Conitzer @conitzer.bsky.social · 15/08/2026
I don’t have the heart to tell it... aifails.substack.com/p/what-ai-wi...
021
Vincent Conitzer @conitzer.bsky.social · 14/08/2026
What’s wrong about this response? Why do you think it got this wrong? aifails.substack.com/p/ted-lasso-...
110