Sign in

Andreas Kirsch

@blackhc.bsky.social
4.9K followers 2K following 224 posts

My opinions only here. 👨‍🔬 RS DeepMind Past: 👨‍🔬 R Midjourney 1y 🧑‍🎓 DPhil AIMS Uni of Oxford 4.5y 🧙‍♂️ RE DeepMind 1y 📺 SWE Google 3y 🎓 TUM 👤 @nwspk

PostsRepliesMedia
Reposted by Andreas Kirsch
United Tech & Allied Workers @utaw.tech · 24/09/2026
Everyone’s talking about AI, but no one’s actually asking workers how it’s affecting us. So we’ve done it ourselves. Today we are launching the AI Workers’ Inquiry, a new study led by workers themselves, into the way AI is changing work. Read it at go.utaw.tech/ai
go.utaw.tech
AI Workers' Inquiry 2026 | Tech Workers Inquiry
How AI is changing work in the tech sector, from the workers who build, deploy, manage, evaluate and use this technology. A Tech Workers' Inquiry from UTAW.
33935
Andreas Kirsch @blackhc.bsky.social · 16/09/2026
A while ago I had Claude do some automated research on a toy continual learning problem, then forgot to post about it. Can a small network learn new digits without losing the old ones? A few interesting results on evaluation, loss scaling and synthetic replay:
Horizontal bar chart of final Split MNIST accuracy across all ten digits after sequentially learning five digit pairs with one MLP. Cross-entropy with Adam: 19.6%; cross-entropy with SGD: 19.4%; task-only BCE: 58.1%; task-only BCE times two: 51.2%; all-class BCE with prototype regularization: 90.2%; the same with four epochs per task: 91.1%; joint training: 95.2%. Means over three seeds. All BCE configurations use SGD. The prototype step changes from summed current-pair BCE to summed all-class BCE and adds synthetic replay. Default is two epochs per task; joint training uses ten epochs. Standard deviations are rounded to one decimal place.
120
Andreas Kirsch @blackhc.bsky.social · 10/09/2026
Three regimes shape Bayesian loss curves: • Misspecification → loss floor • Model complexity → eventual 1/n excess • Prior fit + strength → early/intermediate shape Diffuse priors and prior-data conflict can share a starting loss but learn differently x.com/BlackHC/sta...
031
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
Educational thread: When we scale up models, we pick the winning recipe at a proxy scale. This means it can still lose at the target scale An explainer on Bayesian model selection and scaling laws: 3 criteria for "best", when they disagree, and what that means for your evals
Landscape card titled “Bayesian Model Selection & Scaling Laws.” Three schematic loss curves cross four times as dataset size grows, with a dashed marker at one dataset size and the closing question: which of these is the best model? Lower loss is better; there are no winner bands on this variant.
130
Andreas Kirsch @blackhc.bsky.social · 28/07/2026
Another happy read that we should ponder and reflect on: Forethought's "AI-Enabled Coups". It's a careful, refreshingly unhysterical paper on how a small group (or literally one person) could use advanced AI to seize a state. Ofc: in personal capacity, not on behalf of Google
Scholarly print-style card on graph paper. Eyebrow: MARGINALIA 01 · on "AI-Enabled Coups" · Forethought Research · April 2025 · FRIDAY READING. Headline: "Every tyrant has needed other people. Until now." — "Until now." underlined in red. Tally: 3 risk factors · 4 military coup paths · 3 classes of mitigation. Below, the paper's central claim compressed as a correction: "A coup needs battalions of soldiers willing to go along" struck through in red; beneath, in green: "A coup may soon need one person with exclusive access to advanced AI." Note: Davidson, Finnveden & Hadshar argue that singularly loyal AI workforces, secretly loyal models, and exclusive capability access could make coups feasible even in established democracies. Footer credits the paper and @blackhc
351
Andreas Kirsch @blackhc.bsky.social · 16/07/2026
The Atlantic @TheAtlantic says generative AI is "an engineering disaster." I had Claude fact-check all 23 checkable claims against primary sources: 7 check out · 7 need context · 9 don't hold Verdict: the economics hold up. The computer science doesn't (I checked it too) 🧵
Fact-check summary card on a graph-paper background. Eyebrow: "Corrigendum No. 1 — re: Generative AI Is an Engineering Disaster, The Atlantic, 14 Jul 2026." 

Large serif headline: "The economics check out. The computer science doesn't," with "computer science" in red and underlined. 

A tally box scores 23 claims: 7 check out (green), 7 need context (amber), 9 don't hold up (red). 

Below, the key finding: the central chart's x-axis label "MORE TOKENS" is struck through in red and corrected to "TOKENS PER SECOND PER REQUEST" in green, captioned "A speed dial, not a volume meter. The whole argument rests on it." 

Faint red rising curves decorate the right edge. 

Footer: independent fact-check, 23 claims, compiled 15 July 2026, @blackhc.
170
Andreas Kirsch @blackhc.bsky.social · 14/07/2026
I work at Google DeepMind. This won't make me popular. But it's all public reporting: 2014: DeepMind reportedly sold to Google on conditions: no military use, independent oversight 2026: a Pentagon contract for "any lawful government purpose" Not one safeguard survived intact
Collage titled "Trust is not Governance — an essay from inside Google DeepMind, written in personal capacity." 

A 2014 memorandum, "Conditions of the Acquisition," lists: military applications of DeepMind technology banned; deployment decisions before an independent ethics board (as reported in Mallaby's The Infinity Machine). 

Red threads lead to a 2026 U.S. Department of Defense agreement for classified networks reading "any lawful government purpose," with safety settings and filters adjusted at the government's request and no contractor veto (reported terms, The Information, Apr. 2026). 

Below: a 2018 AI Principles strip ("no weapons, no surveillance") stamped DROPPED 2025, and a Project Mario 2016–2021 tag stamped ABANDONED.
171273435
Andreas Kirsch @blackhc.bsky.social · 01/07/2026
A serious essay by me with many personal thoughts: utaw.tech/news/trust-i... It's on us as Google DeepMind employees to demand real governance, and our union's (@utaw.tech) recognition push is the most realistic path to get there before it's too late
utaw.tech
UTAW: Trust is not Governance
DeepMind has bet that a strong safety culture and good leadership built on trust are sufficient to withstand outside pressure. The bet has failed.
1185
Andreas Kirsch @blackhc.bsky.social · 29/05/2026
Vibe-improved a small useful tool to render markdown & html directly from GitHub URLs, so you don't have to setup GitHub Pages etc E.g. `mdrenderer․github․io/?https꞉//github․com/mdrenderer/mdrenderer․github․com/blob/master/readme․md` (All thanks to Claude Code)
230
Andreas Kirsch @blackhc.bsky.social · 18/03/2026
A while back, Andrej Karpathy said the app store will be replaced by generated, disposable software," and Amjad Masad predicted that the value of all application software will go to zero I think this "ephemeral software hypothesis" is wrong, though, and I want to explain why:
3144
Reposted by Andreas Kirsch
Daniel Hugenroth @lambda.bsky.social · 09/06/2025
We launched CoverDrop 🎉 providing sources with a secure and anonymous way to talk to journalists. Having started five years ago as a PhD research project, this now ships within the Guardian app to millions of users—all of which provide cover traffic. Paper, code, and more info: www.coverdrop.org
coverdrop.org
CoverDrop: Blowing the Whistle Through A News App
15920
Reposted by Andreas Kirsch
Ted Underwood @tedunderwood.com · 11/06/2025
This is going to be big news in my field. While we wait for the dataset, the stuff about post-processing makes interesting reading (if you're me)
47213
Andreas Kirsch @blackhc.bsky.social · 11/06/2025
What's your favorite Veo video?
100
Andreas Kirsch @blackhc.bsky.social · 09/06/2025
I'm late to review the "Illusion of Thinking" paper, so let me collect some of the best threads by and critical takes by @scaling01 in one place and sprinkle some of my own thoughts in as well. The paper is rather critical of reasoning LLMs (LRMs): x.com/MFarajtabar...
2234
Reposted by Andreas Kirsch
Daniel Litt @littmath.bsky.social · 20/05/2025
If the last time you tried to use an LLM for math was ~4 or 5 months ago it’s worth firing up Gemini 2.5 (which you can try for free) or ChatGPT o3 and getting a sense of how rapidly things have progressed.
11465
Andreas Kirsch @blackhc.bsky.social · 17/05/2025
I want to share my latest (very short) blog post: "Active Learning vs. Data Filtering: Selection vs. Rejection." What is the fundamental difference between active learning and data filtering? Well, obviously, the difference is that: 1/11
14011
Reposted by Andreas Kirsch
Marc Lanctot @sharky6000.bsky.social · 28/04/2025
Hive (and all of its expansions) has been added to OpenSpiel! 🎉🤩🐝🐜🕷️🐞🦟🪲 From Gen42: "Hive is an award-winning board game with a difference. There is no board. The pieces are added to the playing area thus creating the board. As more and more pieces are added the game becomes a fight to ... 🧵1/5
1143
Reposted by Andreas Kirsch
ICLR Conference @iclr-conf.bsky.social · 23/04/2025
📢📢 Junior researchers attending #ICLR2025, be sure to check out the mentoring chat sessions! More info here: blog.iclr.cc/2025/04/23/i... You can find all the sessions on the ICLR.cc schedule!
0205
Andreas Kirsch @blackhc.bsky.social · 23/04/2025
I want to share a blog post on our paper "All Models are Wrong, Some are Useful: Model Selection with Limited Labels" which we will present at AISTATS 2025 next week With @pokanovic.bsky.social‬, Jannes Kasper, @thoefler.bsky.social, @arkrause.bsky.social, and @nmervegurel.bsky.social
1125
Reposted by Andreas Kirsch
Andreas Kirsch @blackhc.bsky.social · 05/04/2025
I want to reshare @brandfonbrener.bsky.social's @NeurIPSConf 2024 paper on CoLoR-Filter: A simple yet powerful method for selecting high-quality data for language model pre-training! With @hlzhang109.bsky.social @schwarzjn.bsky.social @shamkakade.bsky.social
2188
Andreas Kirsch @blackhc.bsky.social · 05/04/2025
I want to reshare @brandfonbrener.bsky.social's @NeurIPSConf 2024 paper on CoLoR-Filter: A simple yet powerful method for selecting high-quality data for language model pre-training! With @hlzhang109.bsky.social @schwarzjn.bsky.social @shamkakade.bsky.social
2188
Reposted by Andreas Kirsch
Chris Terry-Enescu @cjterry.bsky.social · 28/02/2025
The Ukrainian government has a list of places where you can donate to the war effort here. I personally just donated $100: war.ukraine.ua/donate/ Slava Ukraini.
war.ukraine.ua
Donate to Ukraine’s defenders
The National Bank of Ukraine has decided to open a special fundraising account to support the Armed Forces of Ukraine.
8519321254
Reposted by Andreas Kirsch
Daniel Hugenroth @lambda.bsky.social · 29/01/2025
I am quite excited that our brand-new module "P79: Cryptography and Protocol Engineering" has its first lecture today! @martin.kleppmann.com and I designed the course to bridge the gap between mathematical ideas and the challenge of implementing secure cryptography in the real world. @cst.cam.ac.uk
4608
Reposted by Andreas Kirsch
pokanovic.bsky.social @pokanovic.bsky.social · 22/01/2025
Check out MODEL SELECTOR, a framework for label-efficient selection of pretrained classifiers. We reduce the labeling cost by up to 94.15% to identify the best model.
132
Reposted by Andreas Kirsch
Andreas Kirsch @blackhc.bsky.social · 07/01/2025
Ever wondered why presenting more facts can sometimes *worsen* disagreements, even among rational people? 🤔 It turns out, Bayesian reasoning has some surprising answers - no cognitive biases needed! Let's explore this fascinating paradox quickly ☺️
824182
Reposted by Andreas Kirsch
Gautam Kamath @gautamkamath.com · 09/01/2025
TMLR is now on Bluesky: be sure to follow @tmlrorg.bsky.social!
14110
Andreas Kirsch @blackhc.bsky.social · 07/01/2025
Ever wondered why presenting more facts can sometimes *worsen* disagreements, even among rational people? 🤔 It turns out, Bayesian reasoning has some surprising answers - no cognitive biases needed! Let's explore this fascinating paradox quickly ☺️
824182
Andreas Kirsch @blackhc.bsky.social · 02/01/2025
I didn't talk about it but I also made heavy use of Claude 3.5 and also o1 and Gemini when creating my lecture series on info theory and active learning in 3.5 weeks: bsky.app/profile/bla...
bsky.app
Andreas Kirsch (@blackhc.bsky.social)
The slides for my lectures on (Bayesian) Active Learning, Information Theory, and Uncertainty are online now 🥳 They cover quite a bit from basic information theory to some recent papers: blackhc.github.io/balitu/ and I'll try to add proper course notes over time 🤗
2130
Andreas Kirsch @blackhc.bsky.social · 17/12/2024
The slides for my lectures on (Bayesian) Active Learning, Information Theory, and Uncertainty are online now 🥳 They cover quite a bit from basic information theory to some recent papers: blackhc.github.io/balitu/ and I'll try to add proper course notes over time 🤗
317628
Reposted by Andreas Kirsch
Clem Delangue 🤗 @clem.hf.co · 16/12/2024
Just 10 days after o1's public debut, we’re thrilled to unveil the open-source version of the technique behind its success: scaling test-time compute By giving models more "time to think," Llama 1B outperforms Llama 8B in math—beating a model 8x its size. The full recipe is open-source!
48319
Andreas Kirsch @blackhc.bsky.social · 15/12/2024
Thanks for following me here! 🫶 I went through my notifications to follow people if they are in ML research, doing PhDs, etc, to have a nice feed focused on ML. Apologies to anyone I have missed! You can unfollow and refollow me to give me a new notification (I suppose)! Plz update your profiles 🙏
1161
Reposted by Andreas Kirsch
Sebastian Ober @sebastianober.bsky.social · 14/12/2024
Excited to be presenting my work, "Big batch Bayesian active learning by considering predictive probabilities" at the Bayesian Decision Making & Uncertainty (BDU) Workshop @neuripsconf.bsky.social, as both a lightning talk and a poster!https://openreview.net/pdf?id=VikX9euujU (1/3)
2314
Reposted by Andreas Kirsch
Alexander Terenin @avt.im · 14/12/2024
The NeurIPS Workshop on Bayesian Decision-making and Uncertainty has started - our first talk is by @mvdw.bsky.social! Join us at East Meeting Room 8, 15, or online!
1416
Reposted by Andreas Kirsch
Nicolas Beltran-Velez @velezbeltran.bsky.social · 12/12/2024
Hello! We will be presenting Estimating the Hallucination Rate of Generative AI at NeurIPS. Come if you'd like to chat about epistemic uncertainty for In-Context Learning, or uncertainty more generally. :) Location: East Exhibit Hall A-C #2703 Time: Friday @ 4:30 Paper: arxiv.org/abs/2406.07457
0234
Andreas Kirsch @blackhc.bsky.social · 13/12/2024
Thanks for the support! I hope I am doing the right thing here. But having written pseudo-code for both approaches now, it does seem like the only addition is a `.flatten()` 🙃 openreview.net/forum?id=0NM...
openreview.net
Not All Tokens Are What You Need for Pretraining
Previous language model pre-training methods have uniformly applied a next-token prediction loss to all training tokens. Challenging this norm, we posit that ''Not all tokens in a corpus are...
0163
Reposted by Andreas Kirsch
Rachel Lawrence @rachel-law.bsky.social · 12/12/2024
📢 We’re recruiting Machine Intelligence PhD interns at MSR Cambridge (UK) 🌍💻 Check it out here — or connect with us at #neurips2024 this week! jobs.careers.microsoft.com/global/en/jo...
jobs.careers.microsoft.com
Search Jobs | Microsoft Careers
163
Andreas Kirsch @blackhc.bsky.social · 12/12/2024
Visualization of epistmic uncertainty via MI as the gap between average conditional entropy and predictive entropy: blackhc.github.io/balitu/lectu...
0130
Reposted by Andreas Kirsch
Sylvain Wallez @swallez.com · 11/12/2024
I'm not a Python developer, and often battle with environments and dependencies when I have to use it. This comprehensive introduction to the uv package manager makes me less hesitant to use Python! www.saaspegasus.com/guides/uv-de...
saaspegasus.com
uv: An In-Depth Guide to Python's Fast and Ambitious New Package Manager
A comprehensive guide on why and how to start using uv—the package manager (and much more) that's taken the Python world by storm.
34415
Reposted by Andreas Kirsch
Tom Andersson 🌍 @tom-andersson.bsky.social · 10/12/2024
So excited to share our Google DeepMind team's new Nature paper on GenCast, an ML-based probabilistic weather forecasting model: www.nature.com/articles/s41... It represents a substantial step forward in how we predict weather and assess the risk of extreme events. 🌪️🧵
nature.com
Probabilistic weather forecasting with machine learning - Nature
GenCast, a probabilistic weather model using artificial intelligence for weather forecasting, has greater skill and speed than the top operational medium-range weather forecast in the world and provid...
210816
Reposted by Andreas Kirsch
Hamel Husain @hamel.bsky.social · 07/12/2024
Are you frustrated by how GitHub renders Jupyter notebooks? I have public service that renders GitHub notebooks with Quarto nbsanity.com It now works with gists!
76714
Andreas Kirsch @blackhc.bsky.social · 08/12/2024
Super grateful to be teaching at AIMS in South Africa for a few weeks on Bayesian Active Learning 😇🥳 aims.ac.za
1110
Reposted by Andreas Kirsch
Nathan Lambert @natolambert.bsky.social · 05/12/2024
This is one I've wanted to do for a while: ask why RL has been continually underestimated in the last 2 years. Interviewing Finbarr Timbers on the "We are So Back" Era of Reinforcement Learning Interconnects interview #11. An overview on the past, present, and future of RL. buff.ly/3Vqrbqj
interconnects.ai
Interviewing Finbarr Timbers on the "We are So Back" Era of Reinforcement Learning
Listen now | Interconnects interview #11. An overview on the past, present, and future of RL.
2588
Andreas Kirsch @blackhc.bsky.social · 04/12/2024
Mandatory ELBO derivation in any lecture series: (I think I finally understand the unnecessarily confusing derivation in the VAE paper 😅)
1230
Reposted by Andreas Kirsch
Laura @lauraruis.bsky.social · 20/11/2024
How do LLMs learn to reason from data? Are they ~retrieving the answers from parametric knowledge🦜? In our new preprint, we look at the pretraining data and find evidence against this: Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢 🧵⬇️
36850139
Andreas Kirsch @blackhc.bsky.social · 01/12/2024
5m bluesky dataset up now. gotta applaud someone doing the right thing 👏👏👏
291
Reposted by Andreas Kirsch
Epoch AI @epochai.bsky.social · 29/11/2024
Could you have trained GPT-4 on 2012 GPUs? Today, we’re releasing an interactive distributed training simulator that allows you to answer this question, among many others!
173
Andreas Kirsch @blackhc.bsky.social · 29/11/2024
I just want to point out that all my favorite art in the last year has been generated by Midjourney. Midjourney is doing more for art literacy than most modern art museums because you have to learn about artists, styles, and aesthetics to know what to prompt for 😊
050
Reposted by Andreas Kirsch
Ferenc Huszár @inference.vc · 25/11/2024
Welcome to the Crazy Rich Bayesian Starter Pack, folk who are/were vaguely into Bayesian reasoning but - with a few exceptions - don't shun the non-Bayesian. go.bsky.app/JYH5Z6M
237413
Andreas Kirsch @blackhc.bsky.social · 28/11/2024
Has anyone shared huggingface.co/datasets/alp... as a torrent yet? Happy to support that effort
huggingface.co
alpindale/two-million-bluesky-posts · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1162