Sign in

Archiki Prasad

@archiki.bsky.social
668 followers 844 following 24 posts

Ph.D. Student at UNC NLP | Apple Scholar in AI/ML Ph.D. Fellowship | Prev: FAIR at Meta, AI2, Adobe (Intern) | Interests: #NLP, #ML | archiki.github.io

PostsRepliesMedia
Reposted by Archiki Prasad
Elias Stengel-Eskin @esteng.bsky.social · 05/05/2025
Extremely excited to announce that I will be joining @utaustin.bsky.social Computer Science in August 2025 as an Assistant Professor! 🎉
UT Austin campus
5449
Reposted by Archiki Prasad
Elias Stengel-Eskin @esteng.bsky.social · 30/04/2025
🌵 I'm going to be presenting PBT at #NAACL2025 today at 2PM! Come by poster session 2 if you want to hear about: -- balancing positive and negative persuasion -- improving LLM teamwork/debate -- training models on simulated dialogues With @mohitbansal.bsky.social and @peterbhase.bsky.social
0103
Reposted by Archiki Prasad
Elias Stengel-Eskin @esteng.bsky.social · 29/04/2025
✈️ Heading to #NAACL2025 to present 3 main conf. papers, covering training LLMs to balance accepting and rejecting persuasion, multi-agent refinement for more faithful generation, and adaptively addressing varying knowledge conflict. Reach out if you want to chat!
1155
Reposted by Archiki Prasad
Elias Stengel-Eskin @esteng.bsky.social · 24/04/2025
Check out 🚨CAPTURe🚨 -- a new benchmark testing spatial reasoning by making VLMs count objects under occlusion. SOTA VLMs (GPT-4o, Qwen2-VL, Intern-VL2) have high error rates on CAPTURe (but humans have low error ✅) and models struggle to reason about occluded objects. arxiv.org/abs/2504.15485 🧵👇
164
Reposted by Archiki Prasad
Mohit Bansal @mohitbansal.bsky.social · 21/04/2025
In Singapore for #ICLR2025 this week to present papers + keynotes 👇, and looking forward to seeing everyone -- happy to chat about research, or faculty+postdoc+phd positions, or simply hanging out (feel free to ping)! 🙂 Also meet our awesome students/postdocs/collaborators presenting their work.
1194
Archiki Prasad @archiki.bsky.social · 18/04/2025
Thanks to my coauthors: @hwang98.bsky.social @esteng.bsky.social @mohitbansal.bsky.social for the fun collaboration! @unccs.bsky.social Paper: arxiv.org/abs/2504.13079 Data and Code: github.com/HanNight/RAM...
arxiv.org
Retrieval-Augmented Generation with Conflicting Evidence
Large language model (LLM) agents are increasingly employing retrieval-augmented generation (RAG) to improve the factuality of their responses. However, in practice, these systems often need to handle...
010
Archiki Prasad @archiki.bsky.social · 18/04/2025
Can RAG systems handle imbalanced evidence or increasing misinformation? ➡️ As document support becomes imbalanced, baselines ignore under-supported correct answers but MADAM-RAG maintains stable performance ➡️ As misinformation 📈, baselines degrade sharply (−46%) but MADAM-RAG remains more robust
000
Archiki Prasad @archiki.bsky.social · 18/04/2025
How important are multi-round debate and aggregation in MADAM-RAG? Increasing debate rounds in MADAM-RAG improves performance by allowing agents to refine answers via debate. Aggregator provides even greater gains, especially in early rounds, aligning conflicting views & suppressing misinfo.
100
Archiki Prasad @archiki.bsky.social · 18/04/2025
We evaluate on 3 datasets: FaithEval (suppression of misinformation), AmbigDocs (disambiguation across sources), RAMDocs (our dataset w/ different types of conflict). MADAM-RAG consistently outperforms concatenated-prompt and Astute RAG baselines across all three datasets and model backbones.
100
Archiki Prasad @archiki.bsky.social · 18/04/2025
We propose MADAM-RAG, a structured, multi-agent framework designed to handle inter-doc conflicts, misinformation, & noise in retrieved content, comprising: 1️⃣ Independent LLM agents - generate intermediate response conditioned on a single doc 2️⃣ Centralized aggregator 3️⃣ Iterative multi-round debate
100
Archiki Prasad @archiki.bsky.social · 18/04/2025
📂RAMDocs is designed to reflect the complexities of real-world retrieval. It includes: ➡️ Ambiguous queries w/ multiple valid ans. ➡️ Imbalanced document support (some answers backed by many sources, others by fewer) ➡️ Docs w/ misinformation (plausible but wrong claims) or noisy/irrelevant content
100
Archiki Prasad @archiki.bsky.social · 18/04/2025
🚨Real-world retrieval is messy: queries are ambiguous or docs conflict & have incorrect/irrelevant info. How can we jointly address these problems? ➡️RAMDocs: challenging dataset w/ ambiguity, misinformation & noise ➡️MADAM-RAG: multi-agent framework, debates & aggregates evidence across sources 🧵⬇️
3167
Reposted by Archiki Prasad
Zaid Khan @codezakh.bsky.social · 15/04/2025
What if we could transform advanced math problems into abstract programs that can generate endless, verifiable problem variants? Presenting EFAGen, which automatically transforms static advanced math problems into their corresponding executable functional abstractions (EFAs). 🧵👇
1165
Reposted by Archiki Prasad
Elias Stengel-Eskin @esteng.bsky.social · 12/04/2025
🚨Announcing TaCQ 🚨 a new mixed-precision quantization method that identifies critical weights to preserve. We integrate key ideas from circuit discovery, model editing, and input attribution to improve low-bit quant., w/ 96% 16-bit acc. at 3.1 avg bits (~6x compression) 📃 arxiv.org/abs/2504.07389
1157
Reposted by Archiki Prasad
UNC-Chapel Hill Computer Science @unccs.bsky.social · 27/03/2025
🎉 A big congratulations to @archiki.bsky.social (advised by Prof. @mohitbansal.bsky.social) for the being awarded the 2025 Apple Scholars in AI/ML PhD Fellowship!", we are proud of you! 👏
021
Archiki Prasad @archiki.bsky.social · 27/03/2025
Thanks Elias!!
010
Archiki Prasad @archiki.bsky.social · 27/03/2025
Thanks Jaemin, learned so much from you as well!
000
Reposted by Archiki Prasad
Mohit Bansal @mohitbansal.bsky.social · 27/03/2025
🎉🎉 Big congrats to @archiki.bsky.social on being awarded the @Apple AI/ML PhD Fellowship, for her extensive contributions in evaluating+improving reasoning in language/reward models and their applications to new domains (ReCEval, RepARe, System-1.x, ADaPT, ReGAL, ScPO, UTGen, GrIPS)! #ProudAdvisor
131
Archiki Prasad @archiki.bsky.social · 27/03/2025
🥳🥳 Honored and grateful to be awarded the 2025 Apple Scholars in AI/ML PhD Fellowship! ✨ Huge shoutout to my advisor @mohitbansal.bsky.social, & many thanks to my lab mates @unccs.bsky.social , past collaborators + internship advisors for their support ☺️🙏 machinelearning.apple.com/updates/appl...
1153
Reposted by Archiki Prasad
Shoubin Yu @shoubin.bsky.social · 19/03/2025
Introducing VEGGIE 🥦—a unified, end-to-end, and versatile instructional video generative model. VEGGIE supports 8 skills, from object addition/removal/changing, and stylization to concept grounding/reasoning. It exceeds SoTA and shows 0-shot multimodal instructional & in-context video editing.
154
Reposted by Archiki Prasad
Mohit Bansal @mohitbansal.bsky.social · 05/02/2025
🚨 Check out "UTGen & UTDebug" for learning to automatically generate unit tests (i.e., discovering inputs which break your code) and then applying them to debug code with LLMs, with strong gains (>12% pass@1) across multiple models/datasets! (see details in 🧵👇) 1/4
174
Archiki Prasad @archiki.bsky.social · 04/02/2025
Thanks to my amazing co-authors @esteng.bsky.social (co-lead), @cyjustinchen.bsky.social, @codezakh.bsky.social, @mohitbansal.bsky.social @unccs.bsky.social Paper: arxiv.org/abs/2502.01619 Code+Datasets: github.com/archiki/UTGe...
arxiv.org
Learning to Generate Unit Tests for Automated Debugging
Unit tests (UTs) play an instrumental role in assessing code correctness as well as providing feedback to a large language model (LLM) as it iteratively debugs faulty code, motivating automated test g...
010
Archiki Prasad @archiki.bsky.social · 04/02/2025
Lastly, we show that both test-time scaling and backtracking are crucial for UTDebug, and scaling the number of generated UTs also consistently improves code accuracy.
100
Archiki Prasad @archiki.bsky.social · 04/02/2025
Combining UTGen with UTDebug 🤝 we consistently outperform no UT feedback, randomly sampling UTs, and prompting targeted UTs across 3 models & datasets. For partially correct code with subtle errors (our MBPP+Fix hard split) debugging with UTGen improves over baselines by >12.35% on Qwen 2.5!
110
Archiki Prasad @archiki.bsky.social · 04/02/2025
RQ3: We also propose ✨UTDebug ✨ with two key modifications: 1⃣Test-time scaling (self-consistency over multiple samples) for increasing output acc. 2⃣Validation & Backtracking: Generating multiple UTs to perform validation, accept edits only when the overall pass rate increases & backtrack otherwise
100
Archiki Prasad @archiki.bsky.social · 04/02/2025
We verify the tradeoff for 0-shot LLMs: when using UTs generated w/o conditioning on the code, the outputs acc is high -- but these UTs are not effective at revealing errors (especially for MBPP+Fix Hard). When prompted to generate failing UTs, we get high attack rate but less accurate UT outputs.
100
Archiki Prasad @archiki.bsky.social · 04/02/2025
On three metrics: attack rate, output acc, and acc + attack (measuring both) we benchmark several open-source 7-8B LLMs. We find that UTGen models balance output acc and attack rate and result in 7.59% more failing/error-revealing unit tests with correct outputs on Qwen-2.5.
100
Archiki Prasad @archiki.bsky.social · 04/02/2025
✨UTGen ✨ bootstraps training data from code generation datasets to train unit test generators. Given coding problems and their solutions, we 1⃣ perturb the code to simulate errors, 2⃣ find challenging UT inputs, 3⃣ generate CoT rationales deducing the correct UT output for challenging UT inputs.
100
Archiki Prasad @archiki.bsky.social · 04/02/2025
We find zero-shot LLMs exhibit a tradeoff between output accuracy and attack rate: UTs whose outputs can be predicted correctly are too easy (not failing) and it's hard to predict outputs of challenging inputs. So how do we do both, i.e., have a high attack rate and output acc (RQ2)? A: ✨UTGen ✨
100
Archiki Prasad @archiki.bsky.social · 04/02/2025
RQ1: What are desirable properties of UT generators? 1⃣ The generator should correctly predict the output for a UT input to a task (output acc 📈) 2⃣ It should uncover errors by generating failing UT inputs, i.e., given incorrect code generate challenging inputs that raise errors (attack rate 📈)
100
Archiki Prasad @archiki.bsky.social · 04/02/2025
1⃣ What are desirable properties of unit test generators? 2⃣ How good are models at 0-shot unit test generation (spoiler alert, they are not great) ... so how do we improve LLMs' UT generation abilities? 3⃣ How can we use potentially noisy feedback from generated tests for debugging?
100
Archiki Prasad @archiki.bsky.social · 04/02/2025
🚨 Excited to share: "Learning to Generate Unit Tests for Automated Debugging" 🚨 which introduces ✨UTGen and UTDebug✨ for teaching LLMs to generate unit tests (UTs) and debugging code from generated tests. UTGen+UTDebug yields large gains in debugging (+12% pass@1) & addresses 3 key questions: 🧵👇
1187
Reposted by Archiki Prasad
Mohit Bansal @mohitbansal.bsky.social · 27/01/2025
🎉 Congrats to the awesome students, postdocs, & collaborators for this exciting batch of #ICLR2025 and #NAACL2025 accepted papers (FYI some are on the academic/industry job market and a great catch 🙂), on diverse, important topics such as: -- adaptive data generation environments/policies ... 🧵
1189
Reposted by Archiki Prasad
Mohit Bansal @mohitbansal.bsky.social · 15/01/2025
Deeply honored & humbled to have received the Presidential #PECASE Award by the @WhiteHouse and @POTUS office! 🙏 Most importantly, very grateful to my amazing mentors, students, postdocs, collaborators, and friends+family for making this possible, and for making the journey worthwhile + beautiful 💙
5438
Archiki Prasad @archiki.bsky.social · 28/12/2024
✨ Collaborating with our amazing postdocs in our lab over the past year has been a great learning experience, with lots of fun + exciting research in LLM agents, reasoning, & multimodality! Check out the new postdoc openings and become a part of the vibrant research @unccs.bsky.social !⬇️
081
Reposted by Archiki Prasad
Mohit Bansal @mohitbansal.bsky.social · 23/12/2024
🚨 We have postdoc openings at UNC 🙂 Exciting+diverse NLP/CV/ML topics**, freedom to create research agenda, competitive funding, very strong students, mentorship for grant writing, collabs w/ many faculty+universities+companies, superb quality of life/weather. Please apply + help spread the word 🙏
13715
Reposted by Archiki Prasad
Jaemin Cho @jmincho.bsky.social · 07/12/2024
🚨 I’m on the academic job market! j-min.io I work on ✨Multimodal AI✨, advancing reasoning in understanding & generation by: 1⃣ Making it scalable 2⃣ Making it faithful 3⃣ Evaluating + refining it Completing my PhD at UNC (w/ @mohitbansal.bsky.social). Happy to connect (will be at #NeurIPS2024)! 👇🧵
23010
Archiki Prasad @archiki.bsky.social · 05/12/2024
I've truly enjoyed ✨ all of our collaborations ✨ over the past year. I particularly admire his thoughtful ideas, dedication to seeing them through, and his mentorship of junior students to do the same. I'm excited to see research from his lab as a professor and an advisor! 😄
131
Reposted by Archiki Prasad
Elias Stengel-Eskin @esteng.bsky.social · 05/12/2024
🚨 I am on the faculty job market this year 🚨 I will be presenting at #NeurIPS2024 and am happy to chat in-person or digitally! I work on developing AI agents that can collaborate and communicate robustly with us and each other. More at: esteng.github.io and in thread below 🧵👇
24714
Reposted by Archiki Prasad
Mohit Bansal @mohitbansal.bsky.social · 03/12/2024
Looking forward to giving this Distinguished Lecture at StonyBrook next week & meeting the several awesome NLP + CV folks there - thanks Niranjan‬ + all for the kind invitation 🙂 PS. Excited to give a new talk on "Planning Agents for Collaborative Reasoning and Multimodal Generation" ➡️➡️ 🧵👇
1238
Reposted by Archiki Prasad
Justin Chih-Yao Chen @cyjustinchen.bsky.social · 02/12/2024
🚨 Reverse Thinking Makes LLMs Stronger Reasoners We can often reason from a problem to a solution and also in reverse to enhance our overall reasoning. RevThink shows that LLMs can also benefit from reverse thinking 👉 13.53% gains + sample efficiency + strong generalization (on 4 OOD datasets)!
11911
Archiki Prasad @archiki.bsky.social · 24/11/2024
🙋🏽‍♀️🙋🏽‍♀️Thanks!
000
Reposted by Archiki Prasad
UNC-Chapel Hill Computer Science @unccs.bsky.social · 21/11/2024
Congratulations to #UNC CS student David Wan for winning the prestigious 2024 Google PhD Fellowship in NLP. 🎉🥳 A very well-deserved honor for his impactful work on factual and faithful text+multimodal generation with Prof. @mohitbansal.bsky.social and UNC NLP group! ▶️ blog.google/technology/r...
094