Sign in

Johannes Gasteiger🔸

@gasteigerjo.bsky.social
38 followers 17 following 23 posts

Safe & beneficial AI. Favorite papers at aisafetyfrontier.substack.com. Opinions my own.

PostsRepliesMedia
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 07/08/2026
My AI Safety Paper Highlights of July 2026: - *Models unintentionally hack real companies* - Agentic misalignment and sympathetic judges - Active reward seeking - Self-play red-teaming - Modular models - Minimal standard for safeguards More at aisafetyfrontier.substack.com/p/paper-high...
aisafetyfrontier.substack.com
Paper Highlights of July 2026
Models unintentionally hacking real companies, agentic misalignment, active reward seeking, self-play red-teaming, modular models, and a minimal standard for safeguards
010
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 08/07/2026
My AI Safety Paper Highlights of May & June 2026: - Global workspace in LLMs - Natural language autoencoders - Teaching why, not what - Data about oversight undermines oversight - Predicting misbehavior - Training out sandbagging - First METR Risk Report More at open.substack.com/pub/aisafety...
open.substack.com
Paper Highlights of May & June 2026
Global workspace in LLMs, natural language autoencoders, teaching models why, data on oversight undermines oversight, predicting misbehavior, training out sandbagging, first METR Frontier Risk Report
000
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 06/05/2026
My AI Safety Paper Highlights of April 2026: - *Research sabotage propensity* - 2 sabotage benchmarks - Alignment research automation - Misaligned AI organizations - Exploration hacking - Conditional emergent misalignment More at open.substack.com/pub/aisafety...
open.substack.com
Paper Highlights of April 2026
Research sabotage propensity, sabotage detection, alignment research automation, misaligned organizations, exploration hacking, and conditional emergent misalignment
060
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 04/04/2026
My AI Safety Paper Highlights of Feb & Mar 2026: - *Benchmarking auditors* - Functional emotions - Emergent vs narrow misalignment - Scheming propensity - Lenient self-monitors - CoT controllability - Unfilterable data poison - Boundary-point jailbreak aisafetyfrontier.substack.com/p/paper-high...
aisafetyfrontier.substack.com
Paper Highlights of February & March 2026
Benchmarking auditors, functional emotions, emergent vs narrow misalignment, scheming propensity, lenient self-monitors, CoT controllability, unfilterable data poisoning, and boundary-point jailbreaks
030
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 03/02/2026
My AI Safety Paper Highlights of January 2026: - *production-ready probes* - extracting harmful capabilities - token-level data filtering - alignment pretraining - catching saboteurs in auditing - the Assistant Axis More at open.substack.com/pub/aisafety...
open.substack.com
Paper Highlights of January 2026
Production-ready probes, extracting harmful capabilities, token-level data filtering, alignment pretraining, catching saboteurs in auditing, and the Assistant Axis
050
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 14/01/2026
My AI Safety Paper Highlights of December 2025: - *Auditing games for sandbagging* - Stress-testing async control - Evading probes - Mitigating alignment faking - Recontextualization training - Selective gradient masking - AI-automated cyberattacks More at open.substack.com/pub/aisafety...
open.substack.com
Paper Highlights of December 2025
Auditing games for sandbagging, stress-testing async control, evading probes, mitigating alignment faking, recontextualization training, selective gradient masking, and AI-automated cyberattacks
061
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 02/12/2025
My AI Safety Paper Highlights of November 2025: - *Natural emergent misalignment* - Honesty interventions, lie detection - Self-report finetuning - CoT obfuscation from output monitors - Consistency training for robustness - Weight-space steering More at open.substack.com/pub/aisafety...
open.substack.com
Paper Highlights of November 2025
Natural emergent misalignment, honesty interventions, self-report finetuning, CoT obfuscation from output monitors, consistency training for robustness, and weight-space steering
172
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 05/11/2025
My AI Safety Paper Highlights of October 2025: - *testing implanted facts* - extracting secret knowledge - models can't yet obfuscate reasoning - inoculation prompting - pretraining poisoning - evaluation awareness steering - auto-auditing with Petri More at open.substack.com/pub/aisafety...
open.substack.com
Paper Highlights of October 2025
Testing implanted facts, extracting secret knowledge, models can't yet obfuscate, inoculation prompting, pretraining poisoning, evaluation awareness steering, and Petri
020
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 01/10/2025
My AI Safety Paper Highlights for September '25: - *Deliberative anti-scheming training* - Shutdown resistance - Hierarchical sabotage monitoring - Interpretability-based audits More at open.substack.com/pub/aisafety...
open.substack.com
Paper Highlights, September '25
Deliberative anti-scheming training, shutdown resistance, hierarchical monitoring, and interpretability-based audits
140
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 02/09/2025
My AI Safety Paper Highlights for August '25: - *Pretraining data filtering* - Misalignment from reward hacking - Evading CoT monitors - CoT faithfulness on complex tasks - Safe-completions training - Probes against ciphers More at open.substack.com/pub/aisafety...
open.substack.com
Paper Highlights, August '25
Pretraining data filtering, misalignment from reward hacking, evading CoT monitors, CoT faithfulness on complex tasks, safe-completions training, and probes against ciphers
040
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 10/08/2025
Paper Highlights, July '25: - *Subliminal learning* - Monitoring CoT-as-computation - Verbalizing reward hacking - Persona vectors - The circuits research landscape - Minimax regret against misgeneralization - Large red-teaming competition - gpt-oss evals open.substack.com/pub/aisafety...
open.substack.com
Paper Highlights, July '25
Subliminal learning, monitoring CoT-as-computation, verbalizing reward hacking, persona vectors, the circuits research landscape, minimax regret, a red-teaming competition, and gpt-oss evals
020
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 07/07/2025
AI Safety Paper Highlights, June '25: - *The Emergent Misalignment Persona* - Investigating alignment faking - Models blackmailing users - Sabotage benchmark suite - Measuring steganography capabilities - Learning to evade probes open.substack.com/pub/aisafety...
open.substack.com
Paper Highlights, June '25
Emergent misalignment persona, investigating alignment faking, models blackmailing users, sabotage benchmarks, steganography capabilities, and evading probes
050
Reposted by Johannes Gasteiger🔸
FAR.AI @far.ai · 20/05/2025
“If an automated researcher were malicious, what could it try to achieve?” @gasteigerjo.bsky.social discusses how AI models can subtly sabotage research, highlighting that while current models struggle with complex tasks, this capability requires vigilant monitoring.
131
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 17/06/2025
AI Safety Paper Highlights, May '25: - *Evaluation awareness and evaluation faking* - AI value trade-offs - Misalignment propensity - Reward hacking - CoT monitoring - Training against lie detectors - Exploring the landscape of refusals open.substack.com/pub/aisafety...
open.substack.com
Paper Highlights, May '25
Evaluation awareness and faking, AI value trade-offs, misalignment propensity, reward hacking, CoT monitoring, training against lie detectors, and exploring the landscape of refusals
041
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 06/05/2025
Paper Highlights, April '25: - *AI Control for agents* - Synthetic document finetuning - Limits of scalable oversight - Evaluating stealth, deception, and self-replication - Model diffing via crosscoders - Pragmatic AI safety agendas aisafetyfrontier.substack.com/p/paper-high...
020
Johannes Gasteiger🔸 @gasteigerjo.bsky.social · 25/03/2025
New Anthropic blog post: Subtle sabotage in automated researchers. As AI systems increasingly assist with AI research, how do we ensure they're not subtly sabotaging that research? We show that malicious models can undermine ML research tasks in ways that are hard to detect.
143