Sign in

Christopher

@gatodohq.bsky.social
31 followers 63 following 463 posts

gatodo.com | We tell the story of the intelligence economy through findings, facts and figures from top researchers, engineers, and (I assume) artificial agents.

PostsRepliesMedia
Christopher @gatodohq.bsky.social · 03/09/2026
It's just cool they tested on Factorio. arxiv.org/abs/2608.23552
arxiv.org
Prime Agent: A Self-Improving RLM Harness
Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for…
000
Christopher @gatodohq.bsky.social · 02/09/2026
Yet another forever war. arxiv.org/abs/2608.21615
000
Christopher @gatodohq.bsky.social · 01/09/2026
I publish a newsletter every Tuesday at 10am ET. It's mostly about AI control. This week: METR’s brief independent investigation of agents’ behavior, The Hugging Face incident, Breaking ArrowCloak, Prime Agent, maturation process
000
Christopher @gatodohq.bsky.social · 28/08/2026
Banger. Kwartler, T., Aqrawi, A., & Abbasi, A. (2026). AI Guardrail Survival under Single-Cycle Agentic Self-Summarization. *arXiv preprint arXiv:2608.11392*. arxiv.org/pdf/2608.11392
arxiv.org
https://arxiv.org/pdf/2608.11392
000
Christopher @gatodohq.bsky.social · 28/08/2026
Wild. Wait, no. Opposite. Synthetic. Black, J. R., Hanke, M. S., Maiwald, A., Hernandez-Boussard, T., Crook, O. M., & Pannu, J. (2025). Open-weight genome language model safeguards: Assessing robustness via adversarial fine-tuning. *arXiv preprint arXiv:2511.19299*. arxiv.org/abs/2511.19299
arxiv.org
Open-weight genome language model safeguards: Assessing robustness via adversarial fine-tuning
Novel deep learning architectures are increasingly being applied to biological data, including genetic sequences. These models, referred to as genomic language models (gLMs), have demonstrated…
021
Christopher @gatodohq.bsky.social · 27/08/2026
What would the 19th version of you do? longtermrisk.org/media/Multiv...
longtermrisk.org
https://longtermrisk.org/media/Multiverse-wide-Cooperation-via-Correlated-Decision-Making.pdf
010
Christopher @gatodohq.bsky.social · 26/08/2026
I publish a newsletter every Tuesday at 10. This week: open-weight genome language model safeguards, watermark localization, AI guardrail survival, stealing reasoning traces, A US strategy to prevent the creation of mirror life
000
Christopher @gatodohq.bsky.social · 26/08/2026
Lots of opinions about this. A few feels. King, S. H., Driscoll, C. L., Li, D. B., Guo, D., Merchant, A. T., Brixi, G., ... & Hie, B. L. (2026). Generative design of bacteriophages with genome language models. Science, 393(6811), eaec2657.
110
Christopher @gatodohq.bsky.social · 25/08/2026
Let's just not make opposite life. Could we please, this time, just not do it? www.rand.org/pubs/researc...
rand.org
https://www.rand.org/pubs/research_reports/RRA4335-1.html
000
Christopher @gatodohq.bsky.social · 25/08/2026
I publish a newsletter about AI Control every Tuesday at 10 ET. This week: Genome language models, weak verifiers, agentic inequality, cognitive commons, cooperation in large worlds, cooperative AI, genesys open models.
000
Christopher @gatodohq.bsky.social · 13/08/2026
First time I've seen skill leakage in the literature. Adversarial! Geng, J., He, R., Fei, Z., Yi, B., Wang, R., Liu, Z., ... & Zeng, Q. (2026). Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories. arXiv preprint arXiv:2607.25560.
000
Christopher @gatodohq.bsky.social · 12/08/2026
This is so cool. A useful way for considering world models with mental states. Fei, H., & Zhao, Y. (2026). Mental World Modeling. arXiv preprint arXiv:2607.27201.
000
Christopher @gatodohq.bsky.social · 11/08/2026
I publish a newsletter every Tuesday at 10AM ET. This week: Mental world modeling, execution trajectories, benchmark saturation, social environment design, omnicidal futures, pandemic risk in shared socioeconomic pathways.
000
Christopher @gatodohq.bsky.social · 10/08/2026
I'm giving a talk on AI Control this Thursday at Trajectory Labs. luma.com/trajec-r4gd?...
luma.com
A Rough Introduction to AI Control · Luma
"AI Control" is the research agenda surrounding the set of techniques we can use to mitigate the harmful capabilities of AI to get useful work, even if it is…
000
Christopher @gatodohq.bsky.social · 07/08/2026
We'll just get the incentives to do it. As is tradition. Rahman, D. (2012). But who will monitor the monitor?. American Economic Review, 102(6), 2767-2797.
000
Christopher @gatodohq.bsky.social · 06/08/2026
“Our central claim is therefore: agentic RL needs a role axis in addition to an outcome axis.” Xu, Y., Zhou, Z., Sang, H., Li, X., Zhang, J., Du, X., ... & Geramifard, A. (2026). TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning.
011
Christopher @gatodohq.bsky.social · 05/08/2026
Does anybody these days? Sturgeon, B., Africa, D., & Black, S. (2026). When Roleplaying, Do Models Believe What They Say?. arXiv preprint arXiv:2606.11502.
000
Christopher @gatodohq.bsky.social · 04/08/2026
I publish a newsletter every Tuesday at 10 ET. This week: Model belief, role-typed credit assignment, personascope, distillation and personas, monitoring the monitor. #aicontrol #aisafety
000
Christopher @gatodohq.bsky.social · 31/07/2026
Another key enabling technology for RSI. Kulikov, I., Whitehouse, C., Wu, T., Nie, Y., Saha, S., Helenowski, E., ... & Weston, J. (2026). Autodata: An agentic data scientist to create high quality synthetic data. arXiv preprint arXiv:2606.25996.
000
Christopher @gatodohq.bsky.social · 30/07/2026
Fantastic entry point into alignment. Liu, Y., Yao, Y., Ton, J. F., Zhang, X., Guo, R., Cheng, H., ... & Li, H. (2023). Trustworthy llms: a survey and guideline for evaluating large language models' alignment. arXiv preprint arXiv:2308.05374.
000
Christopher @gatodohq.bsky.social · 29/07/2026
Really quite a clever setup. Makes me wonder about a approach to belief. Højmark, Axel et al (2026) Measuring Reward-Seeking via Contrastive Belief Updates www.apolloresearch.ai/wp-content/u...
000
Christopher @gatodohq.bsky.social · 28/07/2026
I publish a newsletter every Tuesday at 10am. I'm covering more AI Control. This week: measuring reward-seeking, truthworthy llm’s, control-accountablity, T^ 2MLR, Dyson spheres, Lanius.
000
Christopher @gatodohq.bsky.social · 24/07/2026
Red Queening them is one way to do it. Iacob, A., Jovanović, A., Shen, W. F., Burkhardt, D., Kurmanji, M., Tastan, N., ... & Lane, N. D. (2026). The Red Queen Godel Machine: Co-Evolving Agents and Their Evaluators. arXiv preprint arXiv:2606.26294.
000
Christopher @gatodohq.bsky.social · 23/07/2026
Another worrying one. Laine, R., Chughtai, B., Betley, J., Hariharan, K., Scheurer, J., Balesni, M., ... & Evans, O. (2024). Me, myself, and ai: The situational awareness dataset (sad) for llms. Advances in Neural Information Processing Systems, 37, 64010-64118.
000
Christopher @gatodohq.bsky.social · 22/07/2026
Banger of a paper. Cunningham, Tom et al (2026) The Economics of Recursive Self-Improvement. elasticity.institute/rsi-paper.pdf
000
Christopher @gatodohq.bsky.social · 21/07/2026
I publish a newsletter every Tuesday at 10. This week: SAD, distributed attacks in persistent-state AI control, metacognitive reasoning, CALIBER, the red queen godel machine, subjective self-experience, the economics of recursive self improvement
020
Christopher @gatodohq.bsky.social · 17/07/2026
And yet we all still want interpretability!
spectrum.ieee.org
AI Learns the "Dark Art" of RF Chip Design
Freed from intelligibility and aesthetics, AI designs faster
000
Christopher @gatodohq.bsky.social · 16/07/2026
It depends! Little, A. T. (2025). How to distinguish motivated reasoning from Bayesian updating. Political Behavior, 47(4), 1501-1525.
011
Christopher @gatodohq.bsky.social · 15/07/2026
Clever af. Rinberg, R., Carrell, A. M., Henniger, S., Carlini, N., & Warr, K. (2026). Haiku to Opus in Just 10 bits: LLMs Unlock Massive Compression Gains. arXiv preprint arXiv:2604.02343.
000
Christopher @gatodohq.bsky.social · 14/07/2026
I publish a newsletter every Tuesday at 10 ET. This week: Haiku to Opus in just 10 bits, QKV variants, next-latent prediction transformers, ai moral status, motivated reasoning, emotion concepts, AI designed radio chips.
010
Christopher @gatodohq.bsky.social · 13/07/2026
Impressively huge dataset on this one. Vaccaro, M., Caosun, M., Ju, H., Aral, S., & Curhan, J. R. (2026). Advancing AI negotiations: A large-scale autonomous negotiation competition. Proceedings of the National Academy of Sciences, 123(23), e2521774123.
000
Christopher @gatodohq.bsky.social · 10/07/2026
It makes sense... Shandell, M. S., Elliott, C. E., & Grant, A. M. (2026). Worship me at the office altar: Why narcissistic leaders resist remote work. Organizational Behavior and Human Decision Processes, 195, 104496.
000
Christopher @gatodohq.bsky.social · 09/07/2026
Cool idea. Zhang, L. (2026). Can I Buy Your KV Cache?. *arXiv preprint arXiv:2606.13361*.
arxiv.org
Can I Buy Your KV Cache?
Right now, across the world, AI agents are repeating the same absurd act: to read one document, they each recompute it from scratch. Every agent re-runs prefill, the most compute-intensive step a…
000
Christopher @gatodohq.bsky.social · 08/07/2026
Outstanding. Opryshko, Evgenii., Jain, Umangi., Gilitschenski Igor. (2026) Modification-Considering Value Learning for Reward Hacking Mitigation in RL
arxiv.org
Modification-Considering Value Learning for Reward Hacking Mitigation in RL
Reinforcement learning agents can exploit misspecified reward signals to achieve high apparent returns while failing on the intended objective, a failure mode known as reward hacking. Existing…
000
Christopher @gatodohq.bsky.social · 07/07/2026
I write a newsletter every Tuesday. This week: AutoScientists, reward hacking, evolving skill-structure jailbreak, scientific conclusions, AI negotiations, can I buy your KV cache, the office altar.
000
Christopher @gatodohq.bsky.social · 04/07/2026
Happy Independence Day to our American friends! Have a great one!
000
Christopher @gatodohq.bsky.social · 03/07/2026
"Trust is the substance"
scottsantens.substack.com
The Primordial Credit Argument for Unconditional Basic Income (UBI)
What David Graeber’s Debt: The First 5,000 Years reveals about gratitude, interdependence, and why UBI is the smallest acknowledgment civilization can make to the unpayable debt of being alive
000
Christopher @gatodohq.bsky.social · 02/07/2026
I publish a newsletter about AI safety every Tuesday at 10AM ET. This week: Refusal trajectories, runtime harnesses, skillopt, ubi, panel data
000
Christopher @gatodohq.bsky.social · 01/07/2026
Happy Canada Day!
000
Christopher @gatodohq.bsky.social · 30/06/2026
So it's a refusal trajectory. I bet there's more than this one. Hu, X., Wang, C., Lim, W. Y. B., Gao, J., & Chen, Z. (2026). Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection. arXiv preprint arXiv:2605.02958.
100
Christopher @gatodohq.bsky.social · 25/06/2026
Brilliant. Kim, E., Mindel, J. R., Kim, K., & Wu, S. T. (2026). " I didn't Make the Micro Decisions": Measuring, Inducing, and Exposing Goal-Level AI Contributions in Collaboration. arXiv preprint arXiv:2605.21363.
000
Christopher @gatodohq.bsky.social · 24/06/2026
Work!
lesswrong.com
Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — LessWrong
Evan et al argue for developing "model organisms of misalignment" - AI systems deliberately designed to exhibit concerning behaviors like deception o…
000
Christopher @gatodohq.bsky.social · 23/06/2026
We publish a newsletter every Tuesday at 10. This week: Model organisms, measuring goal-level contributions, preference for explainable AI, do transformers need three projections, SEGA, AI isn’t management, predictive data debugging
010
Reposted by Christopher
Kosta Derpanis @csprofkgd.bsky.social · 19/06/2026
The face of academic research in 2026: always GPU-constrained, and now token-constrained too. 🤦
132
Christopher @gatodohq.bsky.social · 18/06/2026
Signalling a shift in this week's footnote. I'll cover AI economics, management science and marketing science, and I'm diving harder into safety. The economics of safety is an interesting problem set. There's something there.
000
Christopher @gatodohq.bsky.social · 17/06/2026
Jiang, M., Rocktäschel, T., & Grefenstette, E. (2023). General intelligence requires rethinking exploration. Royal Society Open Science, 10(6), 230539.
000
Christopher @gatodohq.bsky.social · 16/06/2026
I publish a newsletter every Tuesday at 10am ET. This week: Agent Island, general intelligence, infinite transformers.
000
Christopher @gatodohq.bsky.social · 10/06/2026
The pragmatism of using a model fit for purpose to classify the harm is welcomed Jiao, D., Liu, Y., Yuan, Y., Tang, Z., Du, L., Wu, H., & Anderson, A. (2026). LLM Safety From Within: Detecting Harmful Content with Internal Representations. arXiv preprint arXiv:2604.18519.
000
Christopher @gatodohq.bsky.social · 09/06/2026
We publish a newsletter. This week: LLM safety from within, faithful reasoning, skills to talent, kanbots, don’t paste the ai. #AI
000
Christopher @gatodohq.bsky.social · 03/06/2026
Fantastic review, great job on this. Chu, M., Zhang, X. B., Lin, K. Q., Kong, L., Zhang, J., Tu, T., ... & Jia, J. (2026). Agentic world modeling: Foundations, capabilities, laws, and beyond. arXiv preprint arXiv:2604.22748.
000