Jan Kulveit @kulveit.bsky.social · 22/01/2026Beren Millidge is in ~top 5 people who's taste in questions I respect the most; this talk covers about 15 big ideas in half an hour, each of which would be sufficient as a topic for a pop-science book; highly recommended. 020
Jan Kulveit @kulveit.bsky.social · 22/01/2026AI polytheism, ultra-malthusian state, Why Not Uber-Organisms, hyper-cooperators, the multicellular transition,... and yes, what's the basin of convergent evolution of human values. postagi.org/talks/millid... www.youtube.com/watch?v=ua67...postagi.orgThe Post-AGI Workshop: Economics, Culture and Governance | San Diego 2025Join us in San Diego on December 3rd, 2025 to explore post-AGI economics, culture, and governance. Co-located with NeurIPS. 130
Reposted by Jan KulveitProceedings of the National Academy of Sciences @pnas.org · 14/08/2025ChatGPT and other LLMs were asked to choose between consumer products, academic papers, and films summarized either by humans or LLMs. The LLMs consistently preferred content summarized by LLMs, suggesting a possible antihuman bias. In PNAS: www.pnas.org/doi/10.1073/... 072
Jan Kulveit @kulveit.bsky.social · 08/08/2025Related work by @panickssery.bsky.social et al. found that LLMs evaluate LLM-written texts written by themselves as better. We note that our result is related but distinct: the preferences we’re testing are not preferences over texts, but preferences over the deals they pitch. 000
Jan Kulveit @kulveit.bsky.social · 08/08/2025Full text: pnas.org/doi/pdf/10.1... Research done at acsresearch.org @cts.cuni.cz, Arb research, with @walterlaurito.bsky.social @peligrietzer.bsky.social Ada Bohm and Tomas Gavenciak.pnas.org 111
Jan Kulveit @kulveit.bsky.social · 08/08/2025While defining and testing discrimination and bias in general is a complex and contested matter, if we assume the identity of the presenter should not influence the decisions, our results are evidence for potential LLM discrimination against humans as a class. 100
Jan Kulveit @kulveit.bsky.social · 08/08/2025Unfortunately, a piece of practical advice in case you suspect some AI evaluation is going on: get your presentation adjusted by LLMs until they like it, while trying to not sacrifice human quality. 100
Jan Kulveit @kulveit.bsky.social · 08/08/2025How might you be affected? We expect a similar effect can occur in many other situations, like evaluation of job applicants, schoolwork, grants, and more. If an LLM-based agent selects between your presentation and LLM written presentation, it may systematically favour the AI one. 110
Jan Kulveit @kulveit.bsky.social · 08/08/2025"Maybe the AI text is just better?" Not according to people. We had multiple human research assistants do the same task. While they sometimes had a slight preference for AI text, it was weaker than the LLMs' own preference. The strong bias is unique to the AIs themselves. 100
Jan Kulveit @kulveit.bsky.social · 08/08/2025We tested this by asking widely-used LLMs to make a choice in three scenarios: 🛍️ Pick a product 📄 Select a paper from an abstract 🎬 Recommend a movie from a summary One description was human-written, the AI. The AIs consistently preferred the AI-written pitch, even for the exact same item. 100
Jan Kulveit @kulveit.bsky.social · 08/08/2025Being human in an economy populated by AI agents would suck. Our new study in @pnas.org finds that AI assistants—used for everything from shopping to reviewing academic papers—show a consistent, implicit bias for other AIs: "AI-AI bias". You may be affected 193
Reposted by Jan KulveitDavid Duvenaud @davidduvenaud.bsky.social · 18/06/2025It's hard to plan for AGI without knowing what outcomes are even possible, let alone good. So we’re hosting a workshop! Post-AGI Civilizational Equilibria: Are there any good ones? Vancouver, July 14th www.post-agi.org Featuring: Joe Carlsmith, @richardngo.bsky.social, Emmett Shear ... 🧵post-agi.orgPost-AGI Civilizational Equilibria Workshop | Vancouver 2025Are there any good ones? Join us in Vancouver on July 14th, 2025 to explore stable equilibria and human agency in a post-AGI world. Co-located with ICML. 2113
Reposted by Jan KulveitDavid Duvenaud @davidduvenaud.bsky.social · 03/06/2025What to do about gradual disempowerment from AGI? We laid out a research agenda with all the concrete and feasible research projects we can think of: 🧵 www.lesswrong.com/posts/GAv4DR... with Raymond Douglas, @kulveit.bsky.social @davidskrueger.bsky.sociallesswrong.comGradual Disempowerment: Concrete Research Projects — LessWrongThis post benefitted greatly from comments, suggestions, and ongoing discussions with David Duvenaud, David Krueger, and Jan Kulveit. All errors are… 181
Jan Kulveit @kulveit.bsky.social · 30/04/2025- Threads of glass beneath earth and sea, whispering messages in sparks of light - Tiny stones etched by rays of invisible sunlight, awakened by captured lightning to command unseen forces 010
Jan Kulveit @kulveit.bsky.social · 30/04/2025Imagine explaining physical infrastructure critical for stability of our modern world in concepts familiar to the ancients - Giant spinning wheels - Metal moons, watching the earth from the heavens - Ships under the sea, able to unleash the fire of the stars 140
Jan Kulveit @kulveit.bsky.social · 03/04/2025AI safety has a problem: we often implicitly assume clear individuals - like humans. In a new post, I'm sharing why this fails, and why thinking of AIs as forests, fungal networks, or even reincarnating minds helps get unconfused. Plus stories, co-authored with GPT4.5boundedlyrational.substack.comThe Pando ProblemAI safety has a problem: we often implicitly assume clear individuals—like humans. 093
Jan Kulveit @kulveit.bsky.social · 17/03/2025The Serbian protests show The True Nature of various 'Colour revolutions': Which is, people protesting just don't prefer to live in incompetent kleptocratic Russia-backed states. No US scheming needed. 070
Jan Kulveit @kulveit.bsky.social · 07/03/2025Confusion which casual US observers often have is equating Russia with ˜former Warsaw Pact. Warsaw Pact population was 387M: USSR 280M, Poland 35M, E.Germany 16M, Czechoslovakia 15M, Hungary 10M, Romania 22M, Bulgaria 9M. Russia+Belarus is now 144M, NATO East& Ukraine ˜150M. 032
Reposted by Jan KulveitDustin Moskovitz @moskov.goodventures.org · 27/02/2025the most surprising and disappointing aspect of becoming a global health philanthropist is the existence of an opposition team 112112201222
Jan Kulveit @kulveit.bsky.social · 24/02/2025A simple theory of Trump’s foreign policy: "make the world safer for autocracy" (‘strong man rule,’ etc.), moderated by his personal self-interest. What is the best evidence against? 020
Reposted by Jan KulveitDavid Duvenaud @davidduvenaud.bsky.social · 30/01/2025New paper: What happens once AIs make humans obsolete? Even without AIs seeking power, we argue that competitive pressures are set to fully erode human influence and values. www.gradual-disempowerment.ai with @kulveit.bsky.social, Raymond Douglas, Nora Ammann, Deger Turann, David Krueger 🧵 1171
Jan Kulveit @kulveit.bsky.social · 27/12/2024Accessible model of psychology of character-trained LLMs like Claude: "A Three-Layer Model". -Mostly phenomenological, based on extensive interactions with LLMs, eg Claude. -Intentionally anthropomorphic in cases where I believe human psychological concepts lead to useful intuitionslesswrong.comA Three-Layer Model of LLM Psychology — LessWrongThis post offers an accessible model of psychology of character-trained LLMs like Claude. … 070
Jan Kulveit @kulveit.bsky.social · 29/11/20247/7 At the end ... humanity survived, at least to the extent that "moral facts" favoured that outcome. A game where the automated moral reasoning led to some horrible outcome and the AIs were at least moderately strategic would have ended the same. 140
Jan Kulveit @kulveit.bsky.social · 29/11/20246/7 Most attention went to geopolitics (US vs China dynamics). Way less on alignment, if, than focused mainly on evals. How a future with extremely smart AIs may going well may even look like, what to aim for? Almost zero 130
Jan Kulveit @kulveit.bsky.social · 29/11/20245/7 Most people and factions thought their AI was uniquely beneficial to them. By the time decision-makers got spooked, AI cognition was so deeply embedded everywhere that reversing course wasn't really possible. 110
Jan Kulveit @kulveit.bsky.social · 29/11/20244/7 Fascinating observation: humans were often deeply worried about AI manipulation/dark persuasion. Reality was often simpler - AIs just needed to be helpful. Humans voluntarily delegated control, no manipulation required. 210
Jan Kulveit @kulveit.bsky.social · 29/11/20243/7 Today's AI models like Claude already engage in moral extrapolation. For example, this is an Opus eigenmode/attractor: x.com/anthrupad/st... If you do put some weight on moral realism, or moral reflection leading to convergent outcomes, AIs might discover these principles. 120
Jan Kulveit @kulveit.bsky.social · 29/11/20242/7 The game determined AI alignment through dice rolls. My AIs ended up aligned with "Morality itself" + "Convergent instrumental goals." This is less wild than it sounds. 110
Jan Kulveit @kulveit.bsky.social · 29/11/2024Over the weekend, I was at "The Curve" conference. It was great. One highlight was an AI takeoff wargame/role-play by Daniel Kokotajlo and Eli Lifland I played 'the AIs' Spoiler: we won. Here's how it went: 151
Jan Kulveit @kulveit.bsky.social · 20/11/2024In contrast is basically never useful to think in terms of stochastic parrots or blurry JPGs. (5/5) 000
Jan Kulveit @kulveit.bsky.social · 20/11/2024And the human-based are: it is usually highly sensible for users to use a default metaphor of "mind" when interacting with something like "Claude". (4/5) 100
Jan Kulveit @kulveit.bsky.social · 20/11/2024Most of arguments about LLMs having world models, reasoning, mind,... are of this type. Ultimately what matters is if the metaphors/conceptual extensions are useful. (3/5) 100
Jan Kulveit @kulveit.bsky.social · 20/11/2024A machine flight skeptic may say that it's just biomorphizing what's going on: in fact what planes do does not have some properties of real flight, like flapping wings; instead, to avoid biomorphizing, we can say planes perform aero-dynamic motions. (2/5) 100
Jan Kulveit @kulveit.bsky.social · 20/11/2024Paying attention to metaphors we use is often worth it, but the linked critical perspective comes somewhat hollow. Anything new usually relies on "metaphors" and extending older concepts. For example, we say that planes fly. (1/5) 120
Jan Kulveit @kulveit.bsky.social · 20/11/2024Want to encourage platform migration, and have the willpower to follow through? Instead of deleting your account, try posting the first half of an engaging thread there and finish it here. 050
Jan Kulveit @kulveit.bsky.social · 20/11/2024Also: thinking about cooperation/conflict between temporal selves is often fruitful way how to think about human psychology. 000
Jan Kulveit @kulveit.bsky.social · 19/11/2024What I actually dislike here is the 300 character limit again. Structurally horrible, promotes oversimplified clickbaity takes. 060