Sign in

Eric Wong

@profericwong.bsky.social
557 followers 73 following 4 posts

Assistant professor at University of Pennsylvania. Machine learning, optimization, robustness & interpretability. Home page: www.cis.upenn.edu/~exwong Lab page: brachiolab.github.io Research blog: debugml.github.io

PostsRepliesMedia
Reposted by Eric Wong
Hadas Orgad @hadasorgad.bsky.social · 23/02/2026
Our ICML 2025 workshop on Actionable Interpretability drew massive interest. But the same questions kept coming up: What does "actionable" mean? Is it achievable? How? We're ready to answer. 🧵
12410
Eric Wong @profericwong.bsky.social · 04/11/2025
What do certified guarantees look like in the age of large language models and long reasoning chains? Look for us at EMNLP to find out!
010
Eric Wong @profericwong.bsky.social · 17/07/2025
If you're at ICML, in about 15 minutes, Weiqiu & I will be at our poster on sum-of-parts models: for faithful attributions and cosmology discovery. Stop by to say hi! East Exhibition Hall A-B #E-1208 Thu 17 Jul 11 a.m. - 1:30 p.m. PDT debugml.github.io/sum-of-parts/ #ICML @youweiqiu.bsky.social
debugml.github.io
Sum-of-Parts Models: Faithful Attributions for Groups of Features
Overcoming fundamental barriers in feature attribution methods with grouped attributions
021
Eric Wong @profericwong.bsky.social · 10/07/2025
LLM ignoring instructions? Make it listen with InstABoost. ✅ Simple: Steer your model in 5 lines of code ✅ Effective: Outperforms latent steering & prompt-only methods ✅ Grounded: Based on our mechanistic theory on rule-following (LogicBreaks) Blog: debugml.github.io/instaboost
debugml.github.io
Instruction Following by Boosting Attention of Large Language Models
We improve instruction-following in large language models by boosting attention, a simple technique that outperforms existing steering methods.
011
Reposted by Eric Wong
adamlsteinl.bsky.social @adamlsteinl.bsky.social · 13/06/2025
🧠 Foundation models are reshaping reasoning. Do we still need specialized neuro-symbolic (NeSy) training, or can clever prompting now suffice? Our new position paper argues the road to generalizable NeSy should be paved with foundation models. 🔗 arxiv.org/abs/2505.24874 (🧵1/9)
111