Sign in

Elias Stengel-Eskin

@esteng.bsky.social
2K followers 682 following 62 posts

Postdoc @UNC working on NLP, AI, and computational linguistics. Formerly PhD student @JHU and undergrad @McGill esteng.github.io

PostsRepliesMedia
Reposted by Elias Stengel-Eskin
Jaemin Cho @jmincho.bsky.social · 20/05/2025
Some personal updates: - I've completed my PhD at @unccs.bsky.social! 🎓 - Starting Fall 2026, I'll be joining the CS dept. at Johns Hopkins University @jhucompsci.bsky.social as an Assistant Professor 💙 - Currently exploring options for my gap year (Aug 2025 - Jul 2026), so feel free to reach out! 🔎
3305
Reposted by Elias Stengel-Eskin
Valentina Pyatkin @valentinapy.bsky.social · 12/05/2025
📢 The SoLaR workshop will be collocated with COLM! @colmweb.org SoLaR is a collaborative forum for researchers working on responsible development, deployment and use of language models. We welcome both technical and sociotechnical submissions, deadline July 5th!
1166
Reposted by Elias Stengel-Eskin
Vaidehi Patil @vaidehipatil.bsky.social · 07/05/2025
🚨 Introducing our @tmlrorg.bsky.social paper “Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation” We present UnLOK-VQA, a benchmark to evaluate unlearning in vision-and-language models, where both images and text may encode sensitive or private information.
1128
Elias Stengel-Eskin @esteng.bsky.social · 06/05/2025
Thank you!
000
Elias Stengel-Eskin @esteng.bsky.social · 06/05/2025
Thanks @niranjanb.bsky.social!
010
Reposted by Elias Stengel-Eskin
Mohit Bansal @mohitbansal.bsky.social · 05/05/2025
🔥 BIG CONGRATS to Elias (and UT Austin)! Really proud of you -- it has been a complete pleasure to work with Elias and see him grow into a strong PI on *all* axes 🤗 Make sure to apply for your PhD with him -- he is an amazing advisor and person! 💙
1124
Elias Stengel-Eskin @esteng.bsky.social · 06/05/2025
Thank you @mohitbansal.bsky.social -- I have learned so much from your mentorship (and benefitted greatly from your job market guidance), and consider myself extremely fortunate to have found such a fantastic lab and postdoc advisor!
110
Elias Stengel-Eskin @esteng.bsky.social · 05/05/2025
Thanks @kmahowald.bsky.social looking forward to collaborating!
020
Elias Stengel-Eskin @esteng.bsky.social · 05/05/2025
Thanks @rdesh26.bsky.social ❤️
000
Elias Stengel-Eskin @esteng.bsky.social · 05/05/2025
And of course thank you to the amazing students/collaborators from @unccs.bsky.social and @jhuclsp.bsky.social 🙏
020
Elias Stengel-Eskin @esteng.bsky.social · 05/05/2025
A huge shoutout to my mentors who have supported and shaped my research! Esp. grateful to my postdoc advisor @mohitbansal.bsky.social for helping me grow along the whole spectrum of PI skills, and my PhD advisor @vandurme.bsky.social for shaping my trajectory as a researcher
120
Elias Stengel-Eskin @esteng.bsky.social · 05/05/2025
Looking forward to continuing to develop AI agents that interact/communicate with people, each other, and the multimodal world. I’ll be recruiting PhD students for Fall 2026 across a range of connected topics (details: esteng.github.io) and plan on recruiting interns for Fall 2025 as well.
esteng.github.io
Elias Stengel-Eskin
Postdoctoral Research Associate, UNC Chapel Hill
130
Elias Stengel-Eskin @esteng.bsky.social · 05/05/2025
Extremely excited to announce that I will be joining @utaustin.bsky.social Computer Science in August 2025 as an Assistant Professor! 🎉
UT Austin campus
5449
Elias Stengel-Eskin @esteng.bsky.social · 30/04/2025
🌵 I'm going to be presenting PBT at #NAACL2025 today at 2PM! Come by poster session 2 if you want to hear about: -- balancing positive and negative persuasion -- improving LLM teamwork/debate -- training models on simulated dialogues With @mohitbansal.bsky.social and @peterbhase.bsky.social
0103
Reposted by Elias Stengel-Eskin
Justin Chih-Yao Chen @cyjustinchen.bsky.social · 30/04/2025
I will be presenting ✨Reverse Thinking Makes LLMs Stronger Reasoners✨at #NAACL2025! In this work, we show - Improvements across 12 datasets - Outperforms SFT with 10x more data - Strong generalization to OOD datasets 📅4/30 2:00-3:30 Hall 3 Let's chat about LLM reasoning and its future directions!
153
Elias Stengel-Eskin @esteng.bsky.social · 29/04/2025
Links: 1⃣ arxiv.org/abs/2410.14596 2⃣ arxiv.org/abs/2503.15272 3⃣ arxiv.org/abs/2409.07394 With awesome collaborators @mohitbansal.bsky.social, @peterbhase.bsky.social, David Wan, @cyjustinchen.bsky.social, Han Wang, @archiki.bsky.social
arxiv.org
Teaching Models to Balance Resisting and Accepting Persuasion
Large language models (LLMs) are susceptible to persuasion, which can pose risks when models are faced with an adversarial interlocutor. We take a first step towards defending models against persuasion while also arguing that defense against adversarial (i.e. negative) persuasion is only half of the equation: models should also be able to accept beneficial (i.e. positive) persuasion to improve their answers. We show that optimizing models for only one side results in poor performance on the other. In order to balance positive and negative persuasion, we introduce Persuasion-Training (or PBT), which leverages multi-agent recursive dialogue trees to create data and trains models via preference optimization to accept persuasion when appropriate. PBT allows us to use data generated from dialogues between smaller 7-8B models for training much larger 70B models. Moreover, PBT consistently improves resistance to misinformation and resilience to being challenged while also resulting in the best overall performance on holistic data containing both positive and negative persuasion. Crucially, we show that PBT models are better teammates in multi-agent debates across two domains (trivia and commonsense QA). We find that without PBT, pairs of stronger and weaker models have unstable performance, with the order in which the models present their answers determining whether the team obtains the stronger or weaker model's performance. PBT leads to better and more stable results and less order dependence, with the stronger model consistently pulling the weaker one up.
010
Elias Stengel-Eskin @esteng.bsky.social · 29/04/2025
📆 04/30 2PM: Teaching Models to Balance Resisting and Accepting Persuasion 📆 05/01 2PM: MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration 📆 05/02 11AM: AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric Knowledge
110
Elias Stengel-Eskin @esteng.bsky.social · 29/04/2025
✈️ Heading to #NAACL2025 to present 3 main conf. papers, covering training LLMs to balance accepting and rejecting persuasion, multi-agent refinement for more faithful generation, and adaptively addressing varying knowledge conflict. Reach out if you want to chat!
1155
Elias Stengel-Eskin @esteng.bsky.social · 24/04/2025
Kudos to Atin Pothiraj on leading this project, with @jmincho.bsky.social and @mohitbansal.bsky.social Code: github.com/atinpothiraj... @hf.co Dataset: huggingface.co/datasets/ati... Paper: arxiv.org/abs/2504.15485
github.com
GitHub - atinpothiraj/CAPTURe
Contribute to atinpothiraj/CAPTURe development by creating an account on GitHub.
010
Elias Stengel-Eskin @esteng.bsky.social · 24/04/2025
By testing VLMs’ spatial reasoning under occlusion, CAPTURe highlights an unexpected weakness. We analyze this weakness by providing the model with additional information: ➡️ Providing object coordinates as text improves performance substantially. ➡️ Providing diffusion-based inpainting also helps.
100
Elias Stengel-Eskin @esteng.bsky.social · 24/04/2025
Interestingly, model error increases with respect to the number of occluded dots, suggesting that task performance is correlated with the level of occlusion. Additionally, model performance depends on pattern type (the shape in which the objects are arranged).
100
Elias Stengel-Eskin @esteng.bsky.social · 24/04/2025
We evaluate 4 strong VLMs (GPT-4o, InternVL2, Molmo, and Qwen2VL) on CAPTURe. Models generally struggle with multiple aspects of the task (occluded and unoccluded) Crucially, every model performs worse in the occluded setting but we find that humans can perform the task easily even with occlusion.
100
Elias Stengel-Eskin @esteng.bsky.social · 24/04/2025
We release 2 splits: ➡️ CAPTURe-real contains real-world images and tests the ability of models to perform amodal counting in naturalistic contexts. ➡️ CAPTURe-synthetic allows us to analyze specific factors by controlling different variables like color, shape, and number of objects.
100
Elias Stengel-Eskin @esteng.bsky.social · 24/04/2025
CAPTURe = Counting Amodally Through Unseen Regions, which requires a model to count objects arranged in a pattern by inferring how the pattern continues behind an occluder (an object that blocks parts of the scene). This needs pattern recognition + counting, making it a good testbed for VLMs!
100
Elias Stengel-Eskin @esteng.bsky.social · 24/04/2025
Check out 🚨CAPTURe🚨 -- a new benchmark testing spatial reasoning by making VLMs count objects under occlusion. SOTA VLMs (GPT-4o, Qwen2-VL, Intern-VL2) have high error rates on CAPTURe (but humans have low error ✅) and models struggle to reason about occluded objects. arxiv.org/abs/2504.15485 🧵👇
164
Reposted by Elias Stengel-Eskin
Archiki Prasad @archiki.bsky.social · 18/04/2025
🚨Real-world retrieval is messy: queries are ambiguous or docs conflict & have incorrect/irrelevant info. How can we jointly address these problems? ➡️RAMDocs: challenging dataset w/ ambiguity, misinformation & noise ➡️MADAM-RAG: multi-agent framework, debates & aggregates evidence across sources 🧵⬇️
3167
Reposted by Elias Stengel-Eskin
hanqix.bsky.social @hanqix.bsky.social · 16/04/2025
Excited to share my first paper as first author: "Task-Circuit Quantization" 🎉 I led this work to explore how interpretability insights can drive smarter model compression. Big thank you to @esteng.bsky.social, Yi-Lin Sung, and @mohitbansal.bsky.social for mentorship and collaboration. More to come
052
Reposted by Elias Stengel-Eskin
Zaid Khan @codezakh.bsky.social · 15/04/2025
What if we could transform advanced math problems into abstract programs that can generate endless, verifiable problem variants? Presenting EFAGen, which automatically transforms static advanced math problems into their corresponding executable functional abstractions (EFAs). 🧵👇
1165
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
Had an awesome time mentoring @hanqix.bsky.social who led this work, with Yi-Lin Sung and @mohitbansal.bsky.social @unccs.bsky.social 💻Code: github.com/The-Inscruta... 📄Paper: arxiv.org/abs/2504.07389
arxiv.org
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
Post-training quantization (PTQ) reduces a model's memory footprint by mapping full precision weights into low bit weights without costly retraining, but can degrade its downstream performance especia...
030
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
To simulate a realistic use case involving generation, we evaluate on Spider for text-to-sql. Quantization methods struggle with preserving performance on generative tasks. We show that TaCQ is the only method to achieve non-zero performance in 2-bits for Llama-3-8B-Instruct.
110
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
We also show that TaCQ generalizes to larger models, recovering 87.93% of the 16-bit Qwen2.5-32B-Instruct model’s performance at 2-bit quantization. We also include results showing TaCQ performs well on Qwen2.5-7B-Instruct.
110
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
✅TaCQ also outperforms even without conditioning on specific tasks, gaining 7.12% in 2-bit and 2.88% in 3-bit. 💡Conditioning creates consistent 10%+ gains in low bits for many quantization methods.
110
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
📊Evaluations show that TaCQ improves accuracy on average by 14.74% in 2-bit and 1-2% in 3-bit when compared to baselines using the same conditioning dataset and lower bits per weight. This holds true across multiple MMLU topics/tasks and GSM8K.
110
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
2️⃣ Magnitude Sharpened Gradient (MSG): Drawing from input attribution, MSG estimates the overall importance of a weight by simulating the impact of removing it from the network entirely to stabilize QAL’s predictions.
110
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
1️⃣ Quantization Aware Localization (QAL): Drawing from circuit discovery & model editing, QAL contrasts unquantized model weights w/ a uniformly-quantized model to estimate the expected change due to quantization and uses gradient information to predict the resulting impact on task performance.
110
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
✨Let’s introduce Task Circuit Quantization (TaCQ) Our saliency metric is composed of two parts: 1️⃣ Quantization Aware Localization (QAL) 2️⃣ Magnitude Sharpened Gradient (MSG):
110
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
Insight: We can efficiently simulate the impact of quantizing individual weights on task performance by conditioning on task data and determining which weights/circuits are most important.
110
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
Motivation ▪️Not all weights are equally important ▪️Weight importance is strongly affected by quantization dynamics ▪️LLMs are often used for specific downstream tasks where preserving task-specific information is crucial ▪️Weight importance varies based on task
110
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
🚨Announcing TaCQ 🚨 a new mixed-precision quantization method that identifies critical weights to preserve. We integrate key ideas from circuit discovery, model editing, and input attribution to improve low-bit quant., w/ 96% 16-bit acc. at 3.1 avg bits (~6x compression) 📃 arxiv.org/abs/2504.07389
1157
Elias Stengel-Eskin @esteng.bsky.social · 27/03/2025
congratulations @archiki.bsky.social! very well-deserved 🎉
110
Reposted by Elias Stengel-Eskin
Archiki Prasad @archiki.bsky.social · 27/03/2025
🥳🥳 Honored and grateful to be awarded the 2025 Apple Scholars in AI/ML PhD Fellowship! ✨ Huge shoutout to my advisor @mohitbansal.bsky.social, & many thanks to my lab mates @unccs.bsky.social , past collaborators + internship advisors for their support ☺️🙏 machinelearning.apple.com/updates/appl...
1153
Reposted by Elias Stengel-Eskin
Shoubin Yu @shoubin.bsky.social · 19/03/2025
Introducing VEGGIE 🥦—a unified, end-to-end, and versatile instructional video generative model. VEGGIE supports 8 skills, from object addition/removal/changing, and stylization to concept grounding/reasoning. It exceeds SoTA and shows 0-shot multimodal instructional & in-context video editing.
154
Elias Stengel-Eskin @esteng.bsky.social · 25/02/2025
🚨UPCORE is our new method for balancing unlearning/forgetting with maintaining model performance. Best part is it works by selecting a coreset from the data rather than changing the model, so it is compatible with any unlearning method, with consistent gains for 3 methods + 2 tasks!
042
Reposted by Elias Stengel-Eskin
Vaidehi Patil @vaidehipatil.bsky.social · 25/02/2025
🚨 Introducing UPCORE, to balance deleting info from LLMs with keeping their other capabilities intact. UPCORE selects a coreset of forget data, leading to a better trade-off across 2 datasets and 3 unlearning methods. 🧵👇
2115
Reposted by Elias Stengel-Eskin
Mohit Bansal @mohitbansal.bsky.social · 05/02/2025
🚨 Check out "UTGen & UTDebug" for learning to automatically generate unit tests (i.e., discovering inputs which break your code) and then applying them to debug code with LLMs, with strong gains (>12% pass@1) across multiple models/datasets! (see details in 🧵👇) 1/4
174
Elias Stengel-Eskin @esteng.bsky.social · 04/02/2025
🚨 Excited to announce UTGen and UTDebug, where we first learn to generate unit tests and then apply them to debugging generated code with LLMs, with strong gains (+12% pass@1) on LLM-based debugging across multiple models/datasets via inf.-time scaling and cross-validation+backtracking! 🧵👇
085
Reposted by Elias Stengel-Eskin
Archiki Prasad @archiki.bsky.social · 04/02/2025
🚨 Excited to share: "Learning to Generate Unit Tests for Automated Debugging" 🚨 which introduces ✨UTGen and UTDebug✨ for teaching LLMs to generate unit tests (UTs) and debugging code from generated tests. UTGen+UTDebug yields large gains in debugging (+12% pass@1) & addresses 3 key questions: 🧵👇
1187
Reposted by Elias Stengel-Eskin
Mohit Bansal @mohitbansal.bsky.social · 27/01/2025
🎉 Congrats to the awesome students, postdocs, & collaborators for this exciting batch of #ICLR2025 and #NAACL2025 accepted papers (FYI some are on the academic/industry job market and a great catch 🙂), on diverse, important topics such as: -- adaptive data generation environments/policies ... 🧵
1189
Elias Stengel-Eskin @esteng.bsky.social · 23/01/2025
Lots more analysis/results in the paper: Work done with @peterbhase.bsky.social @mohitbansal.bsky.social at @unccs.bsky.social Code: github.com/esteng/persu... Paper: arxiv.org/abs/2410.14596 4/4
github.com
GitHub - esteng/persuasion_balanced_training
Contribute to esteng/persuasion_balanced_training development by creating an account on GitHub.
000
Elias Stengel-Eskin @esteng.bsky.social · 23/01/2025
PBT also makes models better teammates: When pairing 2 non-PBT LLMs in a multi-agent debate, we observe order-dependence. Depending on whether the stronger or weaker model goes first, the team lands on the right/wrong answer. PBT reduces this & improves team performance. 3/4
100