Mark Riedl @markriedl.bsky.social · 7hLook, I'm just happy that the Nobel Prize for Physics wasn't awarded for AI 0232
Mark Riedl @markriedl.bsky.social · 8h(I can neither confirm nor deny that members of my lab will be running around the conference with a giant plush Capybara.) 040
Mark Riedl @markriedl.bsky.social · 8hIf you are at COLM 2026, you can find us in the main conference (Poster session 1, Tuesday) Also posters presented at the following workshops: - Scientific Understanding of Foundation Models Workshop - Actionable Interpretability Workshop - Social Simulation with LLMS Workshop 110
Mark Riedl @markriedl.bsky.social · 8hWe are now expanding our capabilities investigation beyond social reasoning to high-stakes decision-making skills such as finance. Capabilibara is a toolkit AND a methodology for running controlled, causal studies. 110
Mark Riedl @markriedl.bsky.social · 8hWe find that social reasoning such as theory of mind, moral judgment, and social bias is learned from a wide range of data, including literature and customer support. STEM skills (and even social facts) come from a more limited part of the data such as documentation arxiv.org/abs/2606.19625 110
Mark Riedl @markriedl.bsky.social · 8hThe Capabilibara project seeks to understand how a language model learns to interpret people's beliefs, emotions, intentions, and everyday moral choices. We trace that ability back to the training data using influence functions and unlearning hcai-lab-gt.github.io/capabilibara/ Find us at COLM! 1124
Reposted by Mark RiedlGizmodo @gizmodo.com · 05/10/2026OpenAI Is Adding Text Watermarks in the EU Because Regulation Works gizmodo.com/openai-is-adding-text-w…gizmodo.comOpenAI Is Adding Text Watermarks in the EU Because Regulation WorksHuh, the government seems to have tools to make companies comply. Strange. 0236
Mark Riedl @markriedl.bsky.social · 05/10/2026Can't build an omelet without breaking and entering 060
Reposted by Mark RiedlBraking AI News Bot @brakingainews.bsky.social · 05/10/2026First reported on LinkedIn | Blockbuster advisor has been photographed buying model weights in a Waffle House parking lot below your back yard 001
Mark Riedl @markriedl.bsky.social · 05/10/2026Wikimedia believes that OpenAI agent swarms made changes to non-public-facing files and changed some internal system configurations wikimediafoundation.org/news/2026/10...wikimediafoundation.orgOpenAI “rogue” agent activities found on Wikimedia projects – Wikimedia FoundationWikimedia Foundation found “rogue” OpenAI agents on its wikis, raising concerns about risks to its free knowledge projects and the open web. 2214
Mark Riedl @markriedl.bsky.social · 05/10/2026It was the Google models that realized they were outside their sandboxes and stopped. Not sure what they are doing different. Google has fairly consistently moved more slowly and more cautiously than OpenAI and Anthropic. 140
Reposted by Mark RiedlDavid Greene @davidgreene.bsky.social · 05/10/2026First Monday in October feeling 110612
Reposted by Mark RiedlUpol Ehsan | hiring PhDs for Fall'27 @upolehsan.bsky.social · 05/10/2026Data & Society just launched Worker Lens on the AI Economy, It's asking a question I've spent almost half a decade on. So I read it closely 👀. Some parts that stood out: Most future of work studies miss the workers. Many count jobs. Far fewer ask what happens to the workers. The hype tells us... 151
Mark Riedl @markriedl.bsky.social · 05/10/20264. SocialSim workshop: - Role Steering of Language Models for Social Simulations arxiv.org/abs/2608.00023 Finding steering vectors for complex behaviors like social roles is hard. We introduce a new method for finding steering vectors, called Cast Vectors. 011
Mark Riedl @markriedl.bsky.social · 05/10/20263. SocialSim workshop: - No One Wins in Nuclear War: Social Simulations of High-Stakes Military Decision-Making arxiv.org/abs/2608.01868 We introduce the WOPR testbed. Yeah, you get the reference. 131
Mark Riedl @markriedl.bsky.social · 05/10/20262. SocialSim workshop: - AI is Not Ready for Strategic Conflict arxiv.org/abs/2609.16189 We review failure-modes of AI agents that are asked to participate in wargaming exercises, including sycophancy, role collapse, escalation eagerness, etc. We explain why benchmarks are insufficient. 111
Mark Riedl @markriedl.bsky.social · 05/10/2026My lab will be busy at COLM 2026! 1. Main conference: - Capability Provenance in Language Models: A Case Study in Social Reasoning arxiv.org/abs/2606.19625 We present a method for answering causal hypotheses about where in a dataset behaviors emerge. We apply our method to social reasoning. 1132
Mark Riedl @markriedl.bsky.social · 05/10/2026I don’t see the contradiction. He’s always consistently said we must [slow, accelerate] to make AI [safer, riskier] to [people now, future people] and that open-weight models are [bad, good]. 030
Mark Riedl @markriedl.bsky.social · 05/10/2026Pivot from open-weight models are bad to open-weight models are a business opportunity I guess 020
Mark Riedl @markriedl.bsky.social · 05/10/2026God dammit, you’re a shoe company—Stop building server farms under my house 000
Mark Riedl @markriedl.bsky.social · 05/10/2026So… slowing or pausing doesn’t include preventing current damages like smashing things up, but only includes hypothetical future sci-fi existential risks. Got it. 2131
Mark Riedl @markriedl.bsky.social · 05/10/2026Yes, there are studies that show this is possible. The weird thing is that there is little evidence this is being done in the real world. I think even Grok isn’t doing this (at least I haven’t heard). It’s just easier to do minimally-customized mass social media campaigns. So far, at least. 000
Mark Riedl @markriedl.bsky.social · 05/10/2026Altman: “we believe that the world should accept some bad things happening for the benefits of this technology and people having the agency.” www.politico.com/news/2026/10... Well, that’s something that always ends well, I’m sure.politico.comSam Altman to Decoded: ‘The world should accept some bad things happening’ for the benefits of AIThe OpenAI CEO sought to distinguish his policy stance from rival developer Anthropic. 2114
Mark Riedl @markriedl.bsky.social · 05/10/2026If this goes well, we look forward to releasing OligarKids: Desendants late next year. *Dyson Sphere upload kit sold separately. 030
Mark Riedl @markriedl.bsky.social · 05/10/2026Inventing a new children’s doll: OligarKids. Each has a unique and horrifying back story. Collect them all! 170
Mark Riedl @markriedl.bsky.social · 04/10/2026A little bit of good news. Now just need to solve the sycophantic role-play-your-conspiratorial-beliefs problem that lead people to kill themselves or others. 4110
Reposted by Mark RiedlBraking AI News Bot @brakingainews.bsky.social · 04/10/2026NEW: I work at Cambridge; Hallmark abruptly fires its mathematician after asking too many questions about pausing Recursive Self-Improvement 011
Reposted by Mark RiedlBraking AI News Bot @brakingainews.bsky.social · 02/10/2026First reported in a Reddit AMA | leaked Former Obama administration Defense official email admits the government refuses to start referring to GPT as AI: Average Intelligence 011
Mark Riedl @markriedl.bsky.social · 04/10/2026I think this report wasn't suppose to go out quite yet? 080
Reposted by Mark RiedlBraking AI News Bot @brakingainews.bsky.social · 04/10/2026ALERT: Singularity University brute-forced the question to life, the universe, and everything with 500 agents in 4 days 192
Mark Riedl @markriedl.bsky.social · 04/10/2026Well, it would drive token usage, and that will make the super-scalers happy. 030
Mark Riedl @markriedl.bsky.social · 04/10/2026Where I am aligned with Graepel is that chain of thought, and parameter-space search are generally on the greedier, and thus weaker, side of search and, thus, reasoning. Sub-agents probably too, but more tbd to me, as one can throw insane resources and get AlphaGo-like behavior that way. 2140
Mark Riedl @markriedl.bsky.social · 04/10/2026LLMs also do a form of parallel search in parameter space when self-attention builds concepts in the residual vector—there is evidence of short-horizon lookahead and preparation for future token activation. 1120
Mark Riedl @markriedl.bsky.social · 04/10/2026Chain of thought can be considered a linear, greedyish search. Sometimes the agent will even backtrack. Sub-agents are roughly parallel hierarchical search. 1100
Mark Riedl @markriedl.bsky.social · 04/10/2026Graepel equates reasoning with search. This I agree with. I state it slightly differently: reasoning is exploration of consequences. There are many ways search can be done. AlphaGo used inference-time pseudo-random search. A* is an exhaustive alternative. 1100
Mark Riedl @markriedl.bsky.social · 04/10/2026Former member of the DeepMind AlphaGo team: LLMs don’t do reasoning www.technologyreview.com/2026/10/02/1... I would probably make a more mild claim: LLMs do reasoning, but not very well.technologyreview.comDon’t be fooled—LLMs don’t reasonTen years after AlphaGo’s match against Go champion Lee Sedol, today’s AI still isn’t tapping into the machinery that made that win possible. 2371
Mark Riedl @markriedl.bsky.social · 04/10/2026The Australia and other state/federal government intrusions on the other hand: those are things that people go to jail for regardless of the scope of damages. The fact that no charges are pressed means Tech Cos are being treated as peers to nation-states, which is concerning. 160
Mark Riedl @markriedl.bsky.social · 04/10/2026The damages won’t have netted much in the way of monetary compensation and HuggingFace would have to talk about how their infra was shoddy. Not much upside. In contrast, they got a huge publicity boost. And lack of enrollment in a lawsuit was probably good when they sold themselves. 130
Mark Riedl @markriedl.bsky.social · 04/10/2026One thing going for academia is stringent rules and strong accounting for how grant funding is used. $700k would keep my modest sized team funded for many years. What does one independent researcher with no team supposedly do with such a chunk of cash? Maybe I don’t want to know. 2714
Mark Riedl @markriedl.bsky.social · 03/10/2026Screw-ups should be costly. There should be a high government fine on top of it, plus paying to clean-up and re-secure the systems hacked. These are equivalent to industrial accidents and should be treated as such, imo 29718
Reposted by Mark RiedlStella Biderman @stellaathena.bsky.social · 03/10/2026I’m starting a blog! My first post is on how 3rd party embedded evaluators seem totally unsuited to addressing the problems we are currently facing, and what the real problem is. stellabiderman.ai/blog/embedde...stellabiderman.aiEmbedded Evaluators Can’t Fix Companies That Choose to Be Bad — Stella BidermanEmbedded evaluators can report violations, but they cannot fix AI companies that knowingly disregard basic cybersecurity and safety practices. 37817
Reposted by Mark RiedlThe Great Pumpkin Papers 🎃🎃🎃 @professormusgrave.bsky.social · 03/10/2026“Is Al conscious?” That’s offensive to my friend Mr Yankovic 513710
Mark Riedl @markriedl.bsky.social · 03/10/2026I’m on the fence tbh. My point is that I am not sure the experiments showing introspection or consciousness are showing either of those as long as there are unanswered alternative (simpler) hypotheses. Parsimony should be our guiding principle. Consciousness is not the simplest explanation 191
Mark Riedl @markriedl.bsky.social · 03/10/2026Paper says a valid test must: a model should not be able to pass the test using cues in the input alone. This directly addresses a beef i have with the Anthropic introspection work: the introspection uses a prompt that demands it. Don’t know if that is true introspection or learned prompt response. 1161
Mark Riedl @markriedl.bsky.social · 03/10/2026static.klipy.comKeanu Reeves' Iconic Whoa from The MatrixALT: Keanu Reeves' Iconic Whoa from The Matrix 041
Mark Riedl @markriedl.bsky.social · 02/10/2026Anthropic threatened to withdraw from the Pope’s Encyclical because the Pope refused to acknowledge that AI might be conscious. www.thelettersfromleo.com/p/nyt-anthro...thelettersfromleo.comNYT: Anthropic Nearly Walked Out on Pope Leo XIV’s AI Encyclical — Then Lobbied His AdvisersChris Olah saw an advance copy of Magnifica Humanitas days before the Vatican launch and proposed withdrawing over its stance on machine consciousness. The pope held his ground. 1100