Our team is at @ICLR 2026 with two papers: mechanistic interpretability work on how a small RNN learns to plan (poster Saturday), and TamperBench, the first unified framework for stress-testing open-weight LLM safety under fine-tuning (three workshops Sunday). 👇