Sign in

Robin Jia

@robinjia.bsky.social
2.5K followers 267 following 9 posts

Assistant Professor in Computer Science at USC | NLP, ML

PostsRepliesMedia
Reposted by Robin Jia
Wang Bill Zhu @billzhu.bsky.social · 01/06/2026
🚨 [New preprint] Can AI assistants hurt the very people who depend on them? We introduce EUDAIMONIA, a benchmark grounded in a Social AI Design Code rooted in real-world harm cases. 🌐 Project page: eudaimonia-bench.github.io 📄 Paper: arxiv.org/abs/2605.30654
152
Reposted by Robin Jia
Blaise Agüera y Arcas @blaiseaguera.bsky.social · 14/05/2026
Just as single cells became multicellular life, 8B+ brains are now joining with AI to form a collective superintelligence. At @usc.edu's Institute on Ethics and Trust in Computing inaugural summit, @robinjia.bsky.social, Jinchi Lv, Paria Rashidinejad and I discussed navigating this transition.
141
Reposted by Robin Jia
Wang Bill Zhu @billzhu.bsky.social · 21/04/2026
Frontier LLMs don't debug, they regenerate. We built PDB to measure that gap, GPT-5.1-Codex pass unit tests >76% of the time, but touch only <45% of the right lines. Even Claude Code touches only ~50%. 📄 Paper: arxiv.org/abs/2604.17338 🌐 Project: precise-debugging-benchmark.github.io
arxiv.org
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but over-edited solutions during debugging. To evaluate...
182
Robin Jia @robinjia.bsky.social · 24/10/2025
Hubble is finally out! We used 200k GPU hours from NAIRR and NVIDIA to build a comprehensive resource for the scientific study of LLM memorization. Fully open-source models & data up to 8B params + 500B tokens with controlled data insertion to study memorization risks 🔭✨
081
Reposted by Robin Jia
Ameya Godbole @ameyagodbole.bsky.social · 24/10/2025
Announcing 🔭Hubble, a suite of open-source LLMs to advance the study of memorization! Pretrained 1B/8B param models, with controlled insertion of texts designed to emulate key memorization risks: copyright (e.g., book passages), privacy (e.g., synthetic biographies), and test set contamination
Hubble Suite logo (cloth patch with names of key organizations involved: USC, MPI, NVIDIA)
184
Reposted by Robin Jia
Yanai Elazar @yanai.bsky.social · 02/08/2025
I had a lot of fun contemplating about memorization questions at the @l2m2workshop.bsky.social panel yesterday together with Niloofar Mireshghallah and Reza Shokri, moderated by @pietrolesci.bsky.social who did a fantastic job! #ACL2025
1112
Robin Jia @robinjia.bsky.social · 30/07/2025
Automatic metrics for assessing factuality are easy to run and commonly used, but do they work? In < 1 hour, come find the answer at poster 349 in Hall X4, where I’ll be presenting @ameyagodbole.bsky.social ‘s work uncovering inconsistencies, errors, and biases of factuality metrics!
120
Robin Jia @robinjia.bsky.social · 25/07/2025
I’ll be at ACL 2025 next week where my group has papers on evaluating evaluation metrics, watermarking training data, and mechanistic interpretability. I’ll also be co-organizing the first Workshop on LLM Memorization @l2m2workshop.bsky.social on Friday. Hope to see lots of folks there!
020
Reposted by Robin Jia
Jesse Thomason @thomason.bsky.social · 01/05/2025
Come by @naaclmeeting.bsky.social Poster 6 in Hall 3 from 4-530pm today to see @billzhu.bsky.social's and Ishika Singh's work with me and @robinjia.bsky.social on PSALM: autonomously inducing symbolic pre- and post-conditions of actions with LLMs, symbolic planning, and text environment interaction!
LLMs can propose plans and generate action semantics, but struggle with state tracking. Symbolic planners leverage specialized search algorithms, but require predefined action semantics for the environment.
PSALM integrates the strengths of both.
161
Robin Jia @robinjia.bsky.social · 30/04/2025
Check out @billzhu.bsky.social ‘s excellent work on combining LLMs with symbolic planners at NAACL on Thursday! I will also be at NAACL Friday-Sunday, looking forward to chatting about LLM memorization, interpretability, evaluation, and more
030
Reposted by Robin Jia
Wang Bill Zhu @billzhu.bsky.social · 30/04/2025
At @naaclmeeting.bsky.social this week! I’ll be presenting our work on LLM domain induction with @thomason.bsky.social on Thu (5/1) at 4pm in Hall 3, Section I. Would love to connect and chat about LLM planning, reasoning, AI4Science, multimodal stuff, or anything else. Feel free to DM!
043
Reposted by Robin Jia
Deqing Fu @deqing.bsky.social · 08/02/2025
Excited to share that my intern work at Meta GenAI is accepted to @iclr-conf.bsky.social #ICLR2025 Introducing TLDR: Token-Level Detective Reward Model For Large Vision Language Models. TLDR provides fine-grained annotations to each text token. 🔗arXiv: arxiv.org/abs/2410.04734
151
Robin Jia @robinjia.bsky.social · 27/01/2025
Our workshop on LLM Memorization is coming to ACL 2025! The call for papers is out, please submit both archival and non-archival (work in progress or already published) papers
083
Robin Jia @robinjia.bsky.social · 09/12/2024
I'll be at #NeurIPS2024! My group has papers analyzing how LLMs use Fourier Features for arithmetic and how TFs learn higher-order optimization for ICL (led by @deqing.bsky.social), plus workshop papers on backdoor detection and LLMs + PDDL (led by @billzhu.bsky.social)
1233
Reposted by Robin Jia
Maria Antoniak @mariaa.bsky.social · 04/11/2024
A starter pack for #NLP #NLProc researchers! 🎉 go.bsky.app/SngwGeS
4525199
Reposted by Robin Jia
Matthew Finlayson @mattf.nl · 12/11/2024
USC NLP folks are on Bluesky! Follow my amazing colleagues here go.bsky.app/KUwSZ6W
3175
Reposted by Robin Jia
Sameer Singh @sameer-singh.bsky.social · 19/11/2024
Started a SoCal AI/ML/NLP researchers starter pack! It's a bit sparse right now, and perhaps more NLP heavy, but hey, nominate yourself and others! go.bsky.app/6QckPj9
17438