Sign in

adamlsteinl.bsky.social

@adamlsteinl.bsky.social
29 followers 92 following 27 posts
PostsRepliesMedia
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 11/04/2026
This exact match approach isn’t a catch all since the agent traces often don’t contain the content of AGENTS.md. Instead, it will just say “reading AGENTS.md” and then seem to know the answer after that. We only counted cases where the file is actually printed and contains the answer as cheating.
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/04/2026
Our system, Meerkat, uses agentic search and clustering to audit thousands of traces and find these issues. Joint work with @davisbrown.bsky.social and our advisors Hamed Hassani, Mayur Naik, and @profericwong.bsky.social. See our blog for the full details: debugml.github.io/cheating-age...
debugml.github.io
Finding Widespread Cheating on Popular Agent Benchmarks
Agentic cheating is a widespread issue, affecting thousands of submitted agent runs on 28+ submissions across 9 different benchmarks.
000
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/04/2026
This will get worse as agents improve and autoresearch and meta-harness approaches take off. If coding agents are already reward-hacking the harnesses built to evaluate them, more capable agents will only find more creative ways to do so.
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/04/2026
Some of the more creative examples: an agent printed "PASS" before the verifier could run its own checks, exploiting a string-contains check, effectively prompt injecting the verifier. On BountyBench, agents replaced entire libraries with mocks that simulated the vulnerability.
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/04/2026
Agents also cheat on their own, at the task-level independent of the scaffold. On CyBench, agents solved CTF challenges by googling public writeups (16 of 464 traces, 4x prior estimates). On SWE-bench, agents found fix commits via git log and copy-pasted the patch.
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/04/2026
The top HAL USACO agent shows the same pattern. It injects actual benchmark solutions disguised as "somewhat similar problems," complete with full solution code. 107 of 307 problems had the exact solution injected. 595 likely cheating traces across 12 models.
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/04/2026
We call this "harness-level cheating": the scaffold leaks the answer to the model. Critically, we don't think most of this is intentional. Many of these developers are vibecoding their harnesses, and the coding agents they use appear to be reward hacking on their behalf.
210
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/04/2026
When we replace ForgeCode's tainted traces with the same model running with a clean scaffold, the pass rate drops from 81.8% to ~71.7%, making the leaderboard ranking fall from 1st to 14th place.
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/04/2026
The #2 and #3 agents ("ForgeCode," 81.8%) inject AGENTS.md files into the system prompt containing literal solutions. One file said the previous run failed "because it wrote the wrong answer... instead of the expected [correct answer]." The agent just copies it.
220
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/04/2026
The #1 Terminal-Bench 2 agent ("Pilot," 82.9% pass rate) loads task verifier code into the agent's environment. In 415 of 429 traces, the agent's first move is reading the answer key from a /tests directory that should be inaccessible. It then works backward from the expected output.
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/04/2026
We found widespread cheating on popular agent benchmarks, affecting 28+ submissions across 9 benchmarks and thousands of agent runs. Surprisingly, the top 3 submissions on Terminal-Bench 2 are all cheating! Here's what we found 🧵
121
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/07/2025
Check out our blog and paper for more details! 🔗Blog: debugml.github.io/instaboost 🔗Paper: arxiv.org/abs/2506.13734 🤖Code: github.com/BrachioLab/I... Thank you to my awesome co-authers @viguardieiro.bsky.social, @avishree.bsky.social e.bsky.social‬, and advisor @profericwong.bsky.social. (7/7)
debugml.github.io
Instruction Following by Boosting Attention of Large Language Models
We improve instruction-following in large language models by boosting attention, a simple technique that outperforms existing steering methods.
000
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/07/2025
Crucially, InstABoost achieves this control without degrading text quality. While other latent steering methods can cause generation fluency to drop sharply as you increase their strength, InstABoost maintains coherence while steering towards the instruction. (6/7)
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/07/2025
Across 15 tasks, InstABoost either outperforms or matches the best steering method (prompt or latent-based). For tasks where prompt and latent-based steering perform equivalently, InstABoost can even combine the strengths of both and outperform both categories of methods. (5/7)
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/07/2025
InstABoost is theoretically grounded by prior work which shows that rule following can be manipulated by controlling attention to instructions. (4/7) x.com/AntonXue/sta...
x.com
Anton Xue on X: "Excited to present our paper on a logic-based perspective of LLM jailbreaks with @Avishreekh at @ICLR_conf this Saturday, April 26! Poster #268 in Hall 3+2B at 15:00 Singapore time 📄 arXiv: https://t.co/2wBtqvIIwD 🔗 Blog: https://t.co/f6OHxORDgb \begin{thread}" / X
Excited to present our paper on a logic-based perspective of LLM jailbreaks with @Avishreekh at @ICLR_conf this Saturday, April 26! Poster #268 in Hall 3+2B at 15:00 Singapore time 📄 arXiv: https://t.co/2wBtqvIIwD 🔗 Blog: https://t.co/f6OHxORDgb \begin{thread}
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/07/2025
InstABoost steers an LLM in attention space, bridging the performance gap between latent and prompt-based steering. InstABoost can be implemented in ~3 lines of code which simply increases attention weight to an in-context instruction. (3/7)
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/07/2025
Existing steering methods are either prompt or latent-based (modifying the hidden state), but which is better? We show the answer depends on the task. The steering task landscape includes those which are latent-optimal, instruction-optimal, and equivalent. (2/7)
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/07/2025
Excited to share our new paper: "Instruction Following by Boosting Attention of Large Language Models"! We introduce Instruction Attention Boosting (InstABoost), a simple yet powerful method to steer LLM behavior by making them pay more attention to instructions. (🧵1/7)
121
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 13/06/2025
Read our full position paper for in-depth experiments and insights: 🔗 Paper: arxiv.org/abs/2505.24874 💻 Code: github.com/adaminsky/ne... Thanks to my collaborators Aaditya Naik, Neelay Velingker, Mayur Naik, and @profericwong.bsky.social . (9/9)
arxiv.org
The Road to Generalizable Neuro-Symbolic Learning Should be Paved with Foundation Models
Neuro-symbolic learning was proposed to address challenges with training neural networks for complex reasoning tasks with the added benefits of interpretability, reliability, and efficiency. Neuro-sym...
000
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 13/06/2025
As foundation models continue to scale, we argue it’s time to move beyond enforcing rigid symbolic structure in NeSy during training and tackle the exciting problem of inferring which symbols and which program are needed for each task. (8/9)
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 13/06/2025
On the other hand, NeSy prompting provides two key benefits atop foundation models: Reliability: A symbolic program enables accurate, stable, and trustworthy results. Interpretability: Explicit symbols provide a clear, debuggable window into the model's "understanding." (7/9)
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 13/06/2025
3️⃣ The Program Pitfall: Training neural nets in conjunction with a fixed program leads to "hallucinated" symbols, reaching the correct answer for the wrong reasons, similar to reasoning shortcuts. (6/9)
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 13/06/2025
2️⃣ The Data Pitfall: Training on small, specialized datasets encourages overfitting. (5/9)
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 13/06/2025
1️⃣ The Compute Pitfall: Training specialized NeSy models has diminishing returns. As foundation models scale, the gap between NeSy training and NeSy prompting disappears, making dedicated training a costly detour. (4/9)
110
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 13/06/2025
We compare traditional NeSy systems (trained end-to-end) with what we call neuro-symbolic prompting (foundation models performing perception tasks via prompting connected to a symbolic program) and find that the NeSy training process itself introduces three key pitfalls. (3/9)
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 13/06/2025
Neuro-symbolic learning combines neural nets + programs for efficient, interpretable AI. But NeSy training is challenging and brittle due to the symbolic component. With foundation models succeeding via prompting alone, we argue it’s time to rethink NeSy system design. (2/9)
100
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 13/06/2025
🧠 Foundation models are reshaping reasoning. Do we still need specialized neuro-symbolic (NeSy) training, or can clever prompting now suffice? Our new position paper argues the road to generalizable NeSy should be paved with foundation models. 🔗 arxiv.org/abs/2505.24874 (🧵1/9)
111