adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/04/2026We found widespread cheating on popular agent benchmarks, affecting 28+ submissions across 9 benchmarks and thousands of agent runs. Surprisingly, the top 3 submissions on Terminal-Bench 2 are all cheating! Here's what we found 🧵 121
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 10/07/2025Excited to share our new paper: "Instruction Following by Boosting Attention of Large Language Models"! We introduce Instruction Attention Boosting (InstABoost), a simple yet powerful method to steer LLM behavior by making them pay more attention to instructions. (🧵1/7) 121
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 13/06/2025🧠 Foundation models are reshaping reasoning. Do we still need specialized neuro-symbolic (NeSy) training, or can clever prompting now suffice? Our new position paper argues the road to generalizable NeSy should be paved with foundation models. 🔗 arxiv.org/abs/2505.24874 (🧵1/9) 111