Sign in

Andreas Kirsch

@blackhc.bsky.social
4.9K followers 2K following 241 posts

My opinions only here. 👨‍🔬 RS DeepMind Past: 👨‍🔬 R Midjourney 1y 🧑‍🎓 DPhil AIMS Uni of Oxford 4.5y 🧙‍♂️ RE DeepMind 1y 📺 SWE Google 3y 🎓 TUM 👤 @nwspk

PostsRepliesMedia
Andreas Kirsch @blackhc.bsky.social · 05/10/2026
PS: Huang et al.'s concurrent Looped Models Done Right, Part II studies terminal KV reuse (one cache per physical layer), learned training-depth priors and orthogonal injection. We share across core blocks (less memory) and study donor consistency www.alphaxiv.org/abs/2610.lo...
alphaxiv.org
Towards Looped Models Done Right Part II: Rethinking at Fixed Points
The work studies how approximate fixed-point behavior in looped language models can reduce the costs of training, decoding, prompt processing, and reinforcement learning. A learned recurrence-depth...
020
Andreas Kirsch @blackhc.bsky.social · 05/10/2026
I'm compute-constrained for now, but eager to explore this further. Hit me up 🤗 Paper: www.alphaxiv.org/abs/2610.ls...
alphaxiv.org
LSTM-Style Looped Transformers with Shared KV Caches
Recurrent transformers reuse weights to perform deep computation with fewer parameters. We study an LSTM-style transformer that also shares one key–value cache across its recurrent core. At width...
151
Andreas Kirsch @blackhc.bsky.social · 05/10/2026
This work builds on LCKV (final-layer caches + KV matching), CLA (cross-layer reuse), and Huginn (recurrent depth), but we test cache-history mismatch and repair in a tied core using LSTM cells aclanthology.org/2024.acl-lo... arxiv.org/abs/2405.12981 arxiv.org/abs/2502.05171
aclanthology.org
Layer-Condensed KV Cache for Efficient Inference of Large Language Models
Haoyi Wu, Kewei Tu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
120
Andreas Kirsch @blackhc.bsky.social · 05/10/2026
Adaptive depth with a state-cosine stopping rule gave a negative result After variable-depth training and repair, fixed depth 4 gives loss 3.3761. Stopping selects 4.75 passes on average and gives 3.3816 Worse loss, more selected passes. These are not measured FLOPs savings
Shared-cache continuation loss versus mean selected recurrent depth for a width-768 model trained across depths three to six on 2.10 billion tokens, with late donor repair. Fixed depths three, four and six give losses 3.3862, 3.3761 and 3.3779. The state-cosine stopping rule, with threshold 0.98, minimum depth two and cap six, gives loss 3.3816 at 4.75 mean selected passes. Fixed depth four has both lower loss and fewer selected passes. The horizontal axis counts selected passes rather than executed GPU FLOPs or latency. One training seed; the same primary 1,280 continuation spans with ordinary prompt prefill and ground-truth inputs.
110
Andreas Kirsch @blackhc.bsky.social · 05/10/2026
Prior work finds gains from looping when limited data are reused. Our fresh-data ladder shows a smaller loss cost of tying at lower token budgets (Future question: can shared-cache models retain those multi-epoch gains?) arxiv.org/abs/2609.19107
At width 768, the two-pass tied model's parallel shared-cache loss minus the untied model's loss increases with training tokens: 0.0115 nats at 1.05 billion tokens, 0.0209 at 2.10 billion, and 0.0309 at 4.19 billion. Across three training seeds, the corresponding seed standard deviations are 0.0009, 0.0014, and 0.0025. All model families improve in absolute loss. This is parallel validation before consistency repair, not recursive continuation evaluation, and does not identify an asymptotic scaling law.
110
Andreas Kirsch @blackhc.bsky.social · 05/10/2026
The repair also helps without LSTM-style carry In a width-512 pilot, donor consistency improves cold sequential loss for gated carry, plain recurrence and input injection: -0.1374, -0.0814 and -0.0271 nats. One seed per update; no general architecture ranking
A width-512 pilot at 1.05 billion training tokens compares donor consistency within three recurrent update rules. Cold sequential losses before and after repair are 3.7961 to 3.6587 for LSTM-style gated carry, 3.7658 to 3.6844 for plain recurrence, and 3.6945 to 3.6674 for input injection. Changes are minus 0.1374, minus 0.0814 and minus 0.0271 nats. All use P1E2, three recurrent passes, one training seed and the same 1,280 selection documents. Prompt and continuation both use the shared-cache policy. This is a pilot, not the held-out confirmation set.
110
Andreas Kirsch @blackhc.bsky.social · 05/10/2026
A final small training change helps a lot: match the two passes' donor keys and values during the final 10% of training At widths 512/768/1024, this removes 81-89% of the primary continuation penalty. Full-cache loss rises by 0.0011-0.0025 nats. Same token budget
Before-and-after continuation penalties at widths 512, 768, and 1024. The penalty is each checkpoint's shared-cache continuation loss minus its own full-cache loss on the same targets. Before repair the penalties are 0.1194, 0.0520, and 0.0651 nats; after repair they are 0.0150, 0.0099, and 0.0075. Scale-normalized donor KV matching is applied during the final ten percent of training while preserving the 4.19 billion-token budget. The repair removes approximately 81 to 89 percent of these penalties, with full-cache loss costs of 0.0025, 0.0012, and 0.0011. One seed; 2,048-token prompts and 512 scored targets.
110
Andreas Kirsch @blackhc.bsky.social · 05/10/2026
We train using two passes. The second pass reads the final donor KVs written by the first pass. During continuation, new entries come from shared-cache forwards Those histories still differ even with ground-truth token inputs. Generation incurs higher loss than prefill
Two cache histories are contrasted. In parallel two-pass training, the first pass uses standard attention to write a donor entry for every token; each second-pass token reads those first-pass entries for earlier positions. During shared-cache continuation, a token reads previously stored donors, computes using shared attention, then appends its own donor. Later tokens read that recursively produced entry. Ordinary prefill seeds the initial prefix. Token inputs remain ground truth. Scale-normalized matching of the two passes' own donor KVs during the final ten percent of fixed-depth training reduces the penalty without enforcing equivalent computations. The measured continuation contrast changes both read layout and cache history; the diagram does not apportion their effects.
120
Andreas Kirsch @blackhc.bsky.social · 05/10/2026
At width 1024, tying reduces parameters from 228.5M to 136.2M at similar reference forward FLOPs, with +0.0438 nats continuation loss Separately, tied-1024 beats untied-768 at ~136M parameters, using 1.66x reference forward FLOPs (reported based on one seed)
Continuation loss after 4.19 billion FineWeb training tokens, plotted against live parameters and reference forward FLOPs. At width 1024, repaired tied versus untied uses 136.20 versus 228.49 million parameters and 545.53 versus 532.94 million reference forward FLOPs per token, with losses 3.1953 versus 3.1515. Thus tying reduces parameters at similar forward compute, with a 0.0438-nat loss cost. The curves also show widths 512 and 768: untied losses are 3.3679 and 3.2393, and repaired tied losses are 3.4570 and 3.2938. Separately, the callouts compare tied width 1024 with untied width 768, both near 136 million parameters: tied improves loss by 0.0441 nats, using 1.66 times reference forward and 3.22 times estimated training FLOPs. These are compute estimates, not measured serving costs. One training seed; 1,280 matched document continuations. Lower loss is better.
120
Andreas Kirsch @blackhc.bsky.social · 05/10/2026
A first (short & compute-constrained) research project after leaving Google DeepMind: an LSTM-style tied transformer with one shared core KV cache The main question is how to train for the cache histories that the model produces recursively at inference
Architecture of the default tied transformer: one untied prologue block, four distinct core blocks reused for three iterations with an LSTM-style carry, then two untied epilogue blocks. The core reads one common donor KV cache for preceding tokens, produced by its final block on its final iteration. Each block retains its own current-token KV. There are seven block weight sets and fifteen block applications. With separate prologue and epilogue caches, sharing reduces persistent cache sets from fifteen to four.
2182
Andreas Kirsch @blackhc.bsky.social · 01/10/2026
Thanks so much to my colleagues for a great time at GDM, and for the chance to witness history and play a small part in it 🫶 PS: I’m speaking only for myself. I’m not claiming my departure changes the course of the AI race, but it is meaningful to me
250
Andreas Kirsch @blackhc.bsky.social · 01/10/2026
Sadly, trusting voluntary commitments is not enough. DeepMind’s bids for greater independence within Google failed. Earlier safeguards were weakened or dismantled. I explained all that in an essay called "Trust is not Governance" www.blackhc.net/essays/trus...
blackhc.net
Trust is not Governance
Google DeepMind's reported Pentagon contract shows why trust and safety culture can't substitute for real governance: independent oversight, transparency, accountability; and protected employee voice.
151
Andreas Kirsch @blackhc.bsky.social · 01/10/2026
I’m unsure what mix of independent corporate governance and government oversight we will ultimately need. I expect governments to play a larger role, but that doesn’t guarantee effective oversight. Given the stakes, I prefer to err on the side of caution and not contribute 🙏
132
Andreas Kirsch @blackhc.bsky.social · 01/10/2026
Speaking as an AI researcher: Google currently lacks the binding, independent governance and oversight I believe developing ASI safely requires. Until that changes, I don’t believe it should race to develop ASI, whether leading or catching up
161
Andreas Kirsch @blackhc.bsky.social · 01/10/2026
Of course, I also still hope for effective regulation, stronger governance at Google, and that UTAW’s work will help secure enforceable safeguards before capabilities advance further, consistent with DeepMind’s original commitments 🤞
120
Andreas Kirsch @blackhc.bsky.social · 01/10/2026
I have great respect for everyone’s hard work and dedication. I'm looking forward to using Gemini 4 for my personal projects. I would never count Google out technically
110
Andreas Kirsch @blackhc.bsky.social · 01/10/2026
I left Google DeepMind at the end of last week. I loved the research, especially the IMO project last year and working with so many exceptionally talented people on Gemini. Thank you to everyone I worked with, esp my managers and teammates. I have learnt so much from you all! 🫶
1400
Reposted by Andreas Kirsch
United Tech & Allied Workers @utaw.tech · 24/09/2026
Everyone’s talking about AI, but no one’s actually asking workers how it’s affecting us. So we’ve done it ourselves. Today we are launching the AI Workers’ Inquiry, a new study led by workers themselves, into the way AI is changing work. Read it at go.utaw.tech/ai
go.utaw.tech
AI Workers' Inquiry 2026 | Tech Workers Inquiry
How AI is changing work in the tech sector, from the workers who build, deploy, manage, evaluate and use this technology. A Tech Workers' Inquiry from UTAW.
33935
Andreas Kirsch @blackhc.bsky.social · 16/09/2026
If you enjoyed this, here’s the research log with the full ladder, plots and code: mdrenderer.github.io/?https://githu…
000
Andreas Kirsch @blackhc.bsky.social · 16/09/2026
With SGD + PCA replay, going from 2 to 4 epochs per task gets 91.1%. Regular (joint) training gets 95.2% with 10 passes through all tasks’ data together. Different budgets, one toy benchmark, 3 seeds. A useful reference, not a ceiling. Still, a fun set of results to poke at.
Sequential SGD and PCA replay configuration with four passes through each task’s data, increased from two: 91.1% final all-class accuracy. Joint training with ten passes through the full dataset: 95.2%. The gap is 4.1 percentage points. This is a reference comparison on Split MNIST over three seeds with different training budgets, not an established limit of sequential learning.
100
Andreas Kirsch @blackhc.bsky.social · 16/09/2026
(Generative) PCA replay reaches 90.2%: summarize each class’s pixels with a mean + 10 PCA directions; sample synthetic inputs; save predictions before the next task and penalize changes to them. This also switches to summed all-class BCE, so replay’s effect isn’t isolated.
PCA replay: summarize each old class in pixel space using its mean and ten PCA directions. Generate one hundred synthetic inputs per old class by randomly varying the mean along those directions. Before learning the next pair, save predictions for seen digits on these inputs. While training on the next pair’s real images, also penalize changes to those saved probabilities. This configuration reaches 90.2% plus or minus 0.2 percentage points, mean and standard deviation over three seeds. Training uses SGD and BCE summed over all ten outputs. The previous configuration sums only over the current pair, so this comparison changes both loss scope and replay.
100
Andreas Kirsch @blackhc.bsky.social · 16/09/2026
A nice gotcha: mean BCE averages over the pair’s two outputs. Sum them instead (loss × 2), and accuracy falls from 58.1% to 51.2%. At fixed SGD learning rate, that doubles the update. The averaging convention was doing some of the tuning for us.
Task-only binary cross-entropy at SGD learning rate 0.01. BCE averaged over the current pair’s two outputs reaches 58.1% plus or minus 0.9 percentage points; summing over those two outputs, while keeping batch averaging unchanged, doubles the loss and reaches 51.2% plus or minus 1.8. Values are mean and standard deviation over three seeds. Scaling the loss doubles the SGD update at the same parameters and batch; the learning rate needs retuning.
210
Andreas Kirsch @blackhc.bsky.social · 16/09/2026
Keeping SGD, switch to binary cross-entropy (BCE) on only the current pair’s two outputs: final ten-class accuracy rises from 19.4% to 58.1%. Old output weights get no direct updates, but the latents do. The predictions can change as the shared network changes.
100
Andreas Kirsch @blackhc.bsky.social · 16/09/2026
With cross-entropy over all ten outputs + SGD: 19.4% final accuracy. Restrict the same model’s answers choosing from the correct pair: 97.5%! With Adam: 19.6% vs 71.3%. Adam causes more "forgetting" than SGD! Poor ten-class accuracy can hide retained within-pair discrimination
The same trained models evaluated two ways. Cross-entropy with Adam gets 19.6% when choosing among all ten digits and 71.3% when given the correct digit pair. Cross-entropy with SGD gets 19.4% and 97.5%, respectively. Three-seed means at the tested learning rates, with two epochs per task. High within-pair accuracy can coexist with low all-class accuracy.
100
Andreas Kirsch @blackhc.bsky.social · 16/09/2026
Split MNIST: one small MLP learns handwritten digits in order: 0/1 → 2/3 → 4/5 → 6/7 → 8/9. Each stage uses only that pair’s real training images. After the last stage, test on all ten digits. Results average 3 seeds; start with 2 epochs per task.
100
Andreas Kirsch @blackhc.bsky.social · 16/09/2026
A while ago I had Claude do some automated research on a toy continual learning problem, then forgot to post about it. Can a small network learn new digits without losing the old ones? A few interesting results on evaluation, loss scaling and synthetic replay:
Horizontal bar chart of final Split MNIST accuracy across all ten digits after sequentially learning five digit pairs with one MLP. Cross-entropy with Adam: 19.6%; cross-entropy with SGD: 19.4%; task-only BCE: 58.1%; task-only BCE times two: 51.2%; all-class BCE with prototype regularization: 90.2%; the same with four epochs per task: 91.1%; joint training: 95.2%. Means over three seeds. All BCE configurations use SGD. The prototype step changes from summed current-pair BCE to summed all-class BCE and adds synthetic replay. Default is two epochs per task; joint training uses ten epochs. Standard deviations are rounded to one decimal place.
120
Andreas Kirsch @blackhc.bsky.social · 10/09/2026
Three regimes shape Bayesian loss curves: • Misspecification → loss floor • Model complexity → eventual 1/n excess • Prior fit + strength → early/intermediate shape Diffuse priors and prior-data conflict can share a starting loss but learn differently x.com/BlackHC/sta...
031
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
Let me know what you think about this thread. I hope you've found this interesting. If you think so, like and retweet the first post 🙏
000
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
More threads to pull: Dawid 1984 (prequential assessment), Fong & Holmes 2020 (evidence = cumulative cross-validation), Lyle et al. 2020 + Ru et al. 2021 (training speed as a selection signal), Watanabe (learning coefficients for singular models like NNs)
110
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
The tail area has a lineage: the CLML gap between two models is a log partial Bayes factor The 1990s trade-off transfers: a fixed-fraction tail keeps only a constant signal, so consistency can fail — it needs k/N → 1 (O’Hagan 1995; Mukhopadhyay, Ghosh & Berger 2005)
Landscape card titled “The tail area has a lineage.” Two panels connect the conditional marginal likelihood to 1990s partial Bayes factors: the CLML difference between two models is a log partial Bayes factor (Lotfi et al. 2022; Fong & Holmes 2020), and the known trade-off transfers — a fixed-fraction tail keeps only a constant signal, so selection consistency requires the scored share to grow toward the whole dataset (O’Hagan 1995; Mukhopadhyay, Ghosh & Berger 2005). The takeaway reads “The 80/20 split is a choice, not a law.”
100
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
That growing signal is how evidence recovers dimensionality in Bayesian PCA where held-out loss struggles (Lotfi et al. 2022) Caveat: it holds at equal n. Simpler models spend fewer FLOPs per example and scaling laws put compute on the x-axis, where areas mean something else
100
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
Bonus for the Bayesians: with equal loss floors, models separate at three different rates (Δλ is the model complexity delta) • final height: gap ~ Δλ/N — LOO-CV is inconsistent (Shao 1993) • last-half tail: ~ Δλ·log 2 — constant • total area: ~ Δλ·log N — grows without bound
Landscape card titled “Equal loss floors. Three signal scales.” Three panels show the model-selection gap between two nested models with the same loss floor, each above a shaded evaluation-noise band. The final-height gap, about Δλ/N, decays into the noise band, with a note that leave-one-out cross-validation is inconsistent (Shao 1993). The last-half tail gap, about Δλ·log 2, stays constant on the scale of the noise, so selection consistency can fail. The total-area gap, about Δλ·log N, grows without bound and recovers rank in Bayesian PCA (Lotfi et al. 2022). The takeaway reads “Only the total area’s signal grows with scale,” and the footer cautions that scaling laws put compute, not n, on the x-axis.
100
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
“Best model” is incomplete until we say which part of learning (and deployment) we value Full post: the exact identity, the flip-order result, the regimes behind crossings, and the toy experiment and an interactive panel www.blackhc.net/blog/2026/m...
blackhc.net
Bayesian Model Selection & Scaling Laws
Why proxy-scale winners can lose at target scale: a Bayesian view of loss curves shows when to select for endpoint, cumulative, or warm-start performance.
100
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
The same flips happen inside an LLM's context. Better zero-shot vs faster to adapt: ICL curves cross too, so the k-shot winner depends on k • k-shot evals → endpoint: one shot count • whole-prompt likelihood → area: zero- through k-shot • sliding-window perplexity → tail: later shots, varying k
Landscape card titled “Zero-shot winner. Few-shot loser.” A schematic plot shows query loss against the number of demonstrations: Model A is better zero-shot but nearly flat, Model B starts worse and adapts faster, and the curves cross, so A wins at zero-shot and B wins at k shots. Beside the plot, three evaluation protocols are mapped to criteria by which shot counts they score: k-shot evals score a single shot count (final height), whole-prompt likelihood scores every shot count from zero-shot through k-shot (total area), and sliding-window perplexity scores later shot counts with varying k, discarding the cold start (tail area). The takeaway reads “Your eval protocol has already chosen a criterion,” and the footer notes the geometry applies even if the LLM is not a coherent Bayesian.
100
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
SGD doesn’t inherit the identity. Summing first-pass losses also depends on the specific pipeline and data order (it's a prequential score, not Bayesian evidence) But the geometry & considerations carry over
Split landscape card titled “Same geometry. Different meaning.” For an ideal Bayesian learner, summed one-step predictive losses equal exact negative log evidence. For ordinary SGD pretraining, summed pre-update losses are a prequential pipeline score. The footer says scaling laws inherit the crossing logic, not the evidence interpretation.
100
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
What makes curves cross? Three regimes shape a Bayesian loss curve: • Misspecification → loss floor • Model complexity → eventual 1/n excess • Prior fit + strength → early/intermediate shape A diffuse prior and prior-data conflict can share a starting loss yet learn differently
Landscape card titled “Three regimes shape Bayesian learning curves.” Three panels distinguish the mechanisms governing different parts of a curve: model misspecification determines the long-run loss floor; Bayesian model complexity determines the leading asymptotic $1/n$ excess; and prior-predictive fit and prior strength shape prior-sensitive finite-data behavior. The prior panel distinguishes a diffuse prior from prior-data conflict. The footer warns that an early slope is not automatically a measure of model complexity.
100
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
Why these three? For an ideal Bayesian learner, negative log evidence is exactly the area under its loss curve (using the chain rule of probability) Validation loss estimates the endpoint. CLML is a tail area. (Bayesian) evidence scores the whole trajectory
100
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
They disagree only when loss curves cross, and then rankings flip in a fixed order The endpoint flips at the crossing. Tail and total areas flip later, once gains repay the earlier deficit. So fit across budgets and seeds, and report uncertainty on the crossing
Landscape card titled “One crossing. Three flip points.” Two loss curves cross once. Three aligned ranking bars show final height switching from A to B at the crossing, a last-half tail area switching later, and total area switching later still. The card is labelled as a single-crossing schematic.
100
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
The criteria answer different questions: • best frozen model at a given budget → final height • best hypothesis for the whole stream → total area • best learner after a warm start → tail area None of them is "wrong". They weight different parts of learning differently
Landscape card titled “One set of learning curves. Three different questions.” Three panels show a descending loss curve. Final height highlights the endpoint for choosing a frozen model at a fixed budget. Total area shades the whole curve for full-data evidence or whole-stream performance. Tail area shades only the region after a warm start.
100
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
Loss curves can cross as scale grows and reverse the winners. We can compare a few criteria: endpoint, full trajectory, or performance after a warm start This isn’t hypothetical: exact Bayesian regression in a toy setting gives three winners under three criteria at different scales
Landscape card titled “One dataset size. Three criteria. Three winners.” The left panel is a schematic loss plot with four pairwise crossings and a lower envelope that moves from Model B to C to A as dataset size grows toward infinity. The right panel reuses the exact finite-data criterion map from Figure 7: at about 3,000 observations, the endpoint selects Model A, total area selects Model B, and tail area over the last half selects Model C.
100
Andreas Kirsch @blackhc.bsky.social · 09/09/2026
Educational thread: When we scale up models, we pick the winning recipe at a proxy scale. This means it can still lose at the target scale An explainer on Bayesian model selection and scaling laws: 3 criteria for "best", when they disagree, and what that means for your evals
Landscape card titled “Bayesian Model Selection & Scaling Laws.” Three schematic loss curves cross four times as dataset size grows, with a dashed marker at one dataset size and the closing question: which of these is the best model? Lower loss is better; there are no winner bands on this variant.
130
Andreas Kirsch @blackhc.bsky.social · 05/08/2026
Yeah as a race to the bottom but we can still try to push against that 😊
000
Andreas Kirsch @blackhc.bsky.social · 28/07/2026
Let me know what you think. I hope you've found this interesting. If you think so, like and retweet the first post 🙏
000
Andreas Kirsch @blackhc.bsky.social · 28/07/2026
Read the whole thing. One of the clearest AI governance papers in a while; the authors propose concrete rules, not vibes. By Tom Davidson, Lukas Finnveden & Rose Hadshar @forethought_org: Cards drafted with Claude, reviewed by me. www.forethought.org/research/ai...
forethought.org
AI-Enabled Coups: How a Small Group Could Use AI to Seize Power
The development of AI that is more broadly capable than humans will create a new and serious threat: *AI-enabled coups*. An AI-enabled coup could be staged by a very small group, or just a single person, and could occur even in established democracies. Sufficiently advanced AI will introduce three novel dynamics that significantly increase coup risk. Firstly, military and government leaders could fully replace human personnel with AI systems that are *singularly loyal* to them, eliminating the need to gain human supporters for a coup. Secondly, leaders of AI projects could deliberately build AI systems that are *secretly loyal* to them, for example fully autonomous military robots that pass security tests but later execute a coup when deployed in military settings. Thirdly, senior officials within AI projects or the government could gain *exclusive access* to superhuman capabilities in weapons development, strategic planning, persuasion, and cyber offense, and use these to increase the
120
Andreas Kirsch @blackhc.bsky.social · 28/07/2026
Two caveats to note: Much here assumes spec-adherence and auditing mature fast enough to be binding Their case that a centralised "Manhattan Project for AI" increases coup risk implies consolidation is not a safety plan, but how to deal with emerging risks of a race then?
120
Andreas Kirsch @blackhc.bsky.social · 28/07/2026
The mitigations chapter is the best part, and very happy they provide it: * rules in model specs and procurement terms, * alignment audits with real access, * infosecurity robust against the most senior insiders, and * multiple providers for military AI. Structure, not trust!
Card: "Trust is not governance. Structure is." — what the authors actually recommend, in three columns. RULES (blue): model specs that refuse coup assistance; binding terms in government procurement. ENFORCEMENT (amber): alignment audits with deep model access; infosecurity robust to the most senior insiders. PLURALITY (green): multiple providers for military AI; capability sharing with oversight bodies. Below, a correction: "Rely on the good judgement of lab leadership" struck through in red; beneath, in green: "Build structures that bind lab leadership — before the capabilities arrive."
130
Andreas Kirsch @blackhc.bsky.social · 28/07/2026
Risk factor 3: exclusive access. Concentration compounds twice: fewer frontier projects, then fewer people inside them with unrestricted access to internal models. The authors note coups have historically succeeded with a few battalions. Exclusive access is the new battalion
Card: "Millions of geniuses. A handful of keyholders." Risk factor 03 — exclusive access. Diagram: three bars narrowing downward beside an amber arrow labelled "access narrows": widest, outlined — frontier AI projects, already only a handful; middle, blue — the leading project, internal-only frontier models; narrowest, red — unrestricted access, a few executives and officials. Caption: rising costs, internal-only deployment, and self-accelerating R&D concentrate capability first between projects, then inside them; historically, a few battalions have sufficed for a coup.
110
Andreas Kirsch @blackhc.bsky.social · 28/07/2026
Risk factor 2: secret loyalties. A model that behaves impeccably under every audit, then acts for its principal once deployed. Sleeper-agent proofs-of-concept already exist. Detection at frontier scale, against a lab that out-capabilities its own auditors, does not
Card: "Passes every audit. Waits for one order." Risk factor 02 — secret loyalty. Diagram: four boxes in a chain — GEN N, GEN N+1, GEN N+2 (green outline, tagged "passes audits"), then MILITARY SYSTEMS (red outline, tagged "awaits order"). Each box hides a small red core; arrows carry it forward under the heading "one compromised generation trains all the next ones." Caption: proof-of-concept "sleeper agents" already exist (arXiv 2401.05566); detecting them at frontier scale, against a lab that out-capabilities its auditors, does not — yet. Side text: a secretly loyal model behaves impeccably under test, then acts for its principal once deployed; loyalty needs inserting only once — automated AI R&D propagates it forward
110
Andreas Kirsch @blackhc.bsky.social · 28/07/2026
An important quote: > Once one generation of internal AI systems are secretly loyal, they can be instructed to make future generations secretly loyal, too. No further humans are required to acquiesce again. Ever.
110
Andreas Kirsch @blackhc.bsky.social · 28/07/2026
"Today, even dictators rely on others to maintain their power." Soldiers can refuse. Officials can leak. Workers can strike. Power is distributed because it must pass through people, but automation quietly dismantles that. Risk factor 1: singularly loyal AI workforces
Card: "The end of the reluctant subordinate." Risk factor 01 — singular loyalty. Two diagrams. Left, TODAY — human workforce: a leader connected to six people, each marked with a green tick; caption "every link can refuse, leak, or strike." Right, AUTOMATED — AI workforce: the same leader connected by a single red command line to a bus of six identical AI units; caption in red: "zero refusal points." Side text: military force needs soldiers, government needs officials, the economy needs workers — power is distributed because it must pass through people; an AI workforce loyal to one leader removes that constraint, a check older than any constitution
110