Sign in

Prithviraj "Raj" Ammanabrolu

@rajammanabrolu.bsky.social
4.2K followers 245 following 174 posts

AI, RL, NLP, Games Asst Prof at UCSD Lab: pearls.ucsd.edu Personal: prithvirajva.com

PostsRepliesMedia
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 05/05/2026
Working on it!
120
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 05/05/2026
No
030
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 01/05/2026
My lab and collaborators had 4 papers on everything from multi objective alignment, reasoning during mid training, multimodal synthetic data, and generating RL tasks accepted to #ICML2026! Come hang out with us in Seoul and we can talk about the exciting follow-ups! Papers at pearls.ucsd.edu
pearls.ucsd.edu
PEARLS Lab @ UCSD | About
Lab website for the PEARLS lab @ University of California, San Diego CSE
260
Reposted by Prithviraj "Raj" Ammanabrolu
Mark Riedl @markriedl.bsky.social · 03/12/2025
My former MS student Chris Cui (now PhD student with @rajammanabrolu.bsky.social)motivates Text Adventure Games as testbeds for reasoning. Provides a new benchmark suite of text games. Observes that Zork still kicks LLM’s butts despite training on walkthroughs arxiv.org/abs/2504.14128
arxiv.org
TALES: Text Adventure Learning Environment Suite
Reasoning is an essential skill to enable Large Language Models (LLMs) to interact with the world. As tasks become more complex, they demand increasingly sophisticated and diverse reasoning capabiliti...
0236
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 24/11/2025
Ziyi Zhang, Shengqi Li (on PhD app market!) for multi agent D&D sim creation + RL openreview.net/pdf?id=3Op7k... Me mostly if you want boba and beach recs (or NVIDIA full time RS roles I guess, I'm hiring a few ppl there but not UCSD)
openreview.net
010
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 24/11/2025
Jenny Shen for pluralistic alignment + human feedback arxiv.org/abs/2510.01167 Chris Cui for text sims + RL + scalable oversight of reasoning models arxiv.org/abs/2504.14128 Lucas (on industry market) for how reasoning emerges from mid training/SFT to RL lucasdino.github.io/assets/files...
130
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 24/11/2025
Bosung Kim for all things VLA, embodied AI, and long context memory arxiv.org/abs/2505.16928 Ruiyi Wang for multi turn agentic RL and all the RL infra in and outs arxiv.org/abs/2510.01132
110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 24/11/2025
My entire PEARLS Lab, and many NVIDIA colleagues, will be at #neurips2025 to chat about their latest. Some papers accepted to the conf are already outdated so just reach out to. Thread 🧵
110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 12/11/2025
Yay congrats, Mark! Well deserved! It's def a required reading for all things comp storytelling (and creativity!)
010
Reposted by Prithviraj "Raj" Ammanabrolu
Mark Riedl @markriedl.bsky.social · 12/11/2025
I am extremely honored and humbled to have been awarded a Test-of-Time award for my 2005 paper "From Linear Story Generation to Branching Story Graphs" with R. Michael Young
5725
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 03/11/2025
I've done a few versions of this talk but this is the first that's been recorded publicly, thanks to IVADO Montreal A good overview of things my lab has been up to in the last year or so at least in balancing safety/capabilities of (embodied) AI Agents www.youtube.com/watch?v=S-kV...
youtube.com
Navigating the Safety-Capability Spectrum when Teaching Agents with Feedback -Prithviraj Ammanabrolu
YouTube video by IVADO
081
Reposted by Prithviraj "Raj" Ammanabrolu
Ruiyi Wang @ruiyiwang.bsky.social · 26/10/2025
🔥Excited to share our new work: "A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning"! We study what actually works for agentic multi-turn RL with varying 🌎Environment, 🤖Policy, and ⭐Reward. We conduct various ablations and empirical analysis on 🧩TextWorld, 🧙ALFWorld, and 🧑‍💻SWE-Gym.
192
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 01/10/2025
My students will be presenting IRPO arxiv.org/abs/2504.15477 and a paper on Personalized RLHF arxiv.org/abs/2504.07070 on Wed onwards
arxiv.org
In-context Ranking Preference Optimization
Recent developments in Direct Preference Optimization (DPO) allow large language models (LLMs) to function as implicit ranking models by maximizing the margin between preferred and non-preferred respo...
010
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 01/10/2025
My students will be presenting papers next Wed/Thursday so be sure to check those out too
130
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 01/10/2025
I'll be at #CoLM2025 and the IVADO agents workshop right before in Montreal. My students will be presenting two papers in the main conf. I'll also do a ws keynote where I'll talk about some of our latest. Come by and say hi next week!
180
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 29/06/2025
I'm probably mostly going to stop posting on this site. There's close to no engagement and it's not worth the effort to cross post for the amount of time that takes. Find me elsewhere / email me
450
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 19/06/2025
I recently left Mosaic/Databricks Research. It's been a ride building out the RL team from <4 ppl to 20+ across two companies & acquisition +figuring out RL as a Service in prod. Mosaic had insane talent density Some "relaxation" while I put out Prof fires for a smol bit then new adventures!
070
Reposted by Prithviraj "Raj" Ammanabrolu
Mark Riedl @markriedl.bsky.social · 17/06/2025
If you work in the intersection of NLP and games/narrative, then this workshop is for you! wordplay-workshop.github.io/cfp/ Organized by the amazing @laramartin.net and @rajammanabrolu.bsky.social (among others)
wordplay-workshop.github.io
/call_for_papers
Official website for the Wordplay Workshop at EMNLP 2025. Exploring interactive narratives, text-adventure games, and AI agents in language-based environments. Join us in Suzhou, China, November 5th-9...
0105
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 16/06/2025
The thing that feels so off about the core tech world is that every convo is very transactional. Maybe true elsewhere too. "Oh you're an expert in RL, can you answer questions about my new startup?" Every single (Bay) party. No I do not want to consult. I just wanna hang out.
190
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 10/06/2025
Of all the labeling startups out there to acquihire, this was... an interesting choice. Says a lot actually
100
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 09/06/2025
. @bosungkim.bsky.social will be at #CVPR2025 in Nashville this week to present this and just generally talk about scaling memory for embodied agents! Catch her at the poster sessions and also the Foundation Models meets Embodied Agents Workshop on Wed
020
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 09/06/2025
Yes AI for edu is a thing but almost all vanilla LLMs just railroad students into answers. Complete cognitive offload is not useful for improving learning outcomes
020
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 09/06/2025
I've heard this personally from multiple PMs at AI companies. Students are one of the biggest demographics and they need to "break in" and have even more usage to improve their metrics. Classic corporate economic incentives
174
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 06/06/2025
Tis the era of bringing back every AI benchmark ever but this time by the LLM people and for the LLMs
021
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 05/06/2025
Had a fun little visit to Cambridge LTL where I talked about a bunch of my lab's latest papers including some still not public with the key takeaway that "RL can absolutely learn new things and is not just resurfacing knowledge" talks.cam.ac.uk/show/archive...
talks.cam.ac.uk
talks.cam : Language Technology Lab Seminars
041
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 03/06/2025
That's fair, I guess I should rephrase to "regardless of a possible common prior, it's nearly impossible for different providers to have the same representations pop out of their post trained LLM"
100
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 03/06/2025
The moral of the story here is basically that who is making your LLM really matters. Internal use cases critical to their businesses will always influence data distributions and everything downstream of that. This is in contrast to things like Platonic Representation Hypothesis
150
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 03/06/2025
Interesting tidbit from UCSD's Victor Shih on a podcast talking about Chinese AGI efforts is that Deepseek is good at Chinese govt doc understanding cause that's what affects stock prices most and DS is a hedge fund. www.youtube.com/watch?v=b1Te...
youtu.be
Xi Jinping’s paranoid approach to AGI, debt crisis, & Politburo politics — Victor Shih
YouTube video by Dwarkesh Patel
150
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 02/06/2025
The top two scores are 332 not 322 but other than the typo the rest of this list seems legit and consistent across multiple sources www.indiatvnews.com/education/hi... x.com/RejaullahmdM...
indiatvnews.com
JEE Advanced 2025 topper list out, Rajit Gupta secures AIR 1 with 332 marks: Check complete list of toppers
IIT Kanpur has released the topper list for JEE Advanced 2025. This year, Rajit Gupta achieved AIR 1 by scoring 332 out of 360 marks. Devdutta Majhi is the leading female candidate with AIR 16. Scroll...
000
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 02/06/2025
Looks like Gemini gets AIR 6 in #JEE2025 with a score of 323 Only 5 highschoolers in all India do better than an LLM in the single most important exam of their to get into the IITs The legacy edu selection systems are now worse than useless
140
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 01/06/2025
I get prepping for worst case scenarios but a lot of AI Safety debates I somehow end up these days in boil down to "assume you have Machine God in a box, now tell me how to align it" I could rant for hours but seriously y'all this isn't productive
180
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 24/05/2025
Here for the afternoon shift!
010
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025
Paper: arxiv.org/abs/2505.16928 Website/code/data: pearls-lab.github.io/infini-thor/ Led by @bosungkim.bsky.social who has done a fantastic job on this in the last bit. Full stack from Unity gamedev to Big Model Scaler. Watch out for her in the embodied agent space!
arxiv.org
Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning
We introduce $\infty$-THOR, a new framework for long-horizon embodied tasks that advances long-context understanding in embodied AI. $\infty$-THOR provides: (1) a generation framework for synthesizing...
020
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025
But even then agents only perform well up to 130k after which perf sharply decreases and all architectures and additional context extension methods we modified fail after ~400k. None make it to the 3m context sample test set we use let alone infinite. Lots of space for progress!
110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025
And find that when accounting for hardware constraints, only a specific combo of interleaved VLA with a mix of context parallelism + some extension with high pre training context window size works well. We detail the exact architecture and describe potential improvements
100
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025
The first is a static eval called Needle(s) in the Embodied Haystack, which is like QA asking agents to post hoc analyze their trajectories putting many needles together We then go Beyond this with interactive RL style evals to see how well models interact with a changing env
110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025
So we first extended AI2's THOR embodied sim to continue generating meaningful tasks that are effectively infinite in length. Eg things you do at context len 7k can get reused at 3m We do a thorough analysis of many types of architectures x training methods on two new evals
110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025
The combo of Big Models, embodied agents, and long context has much potential. But it's very unclear what works and what doesn't. Most robotics sims don't have high cognitive complexity+long contexts and other types of sims don't require physical understanding
100
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025
"Foundation" models for embodied agents are all the rage but how to actually do complex looong context reasoning? Can we scale Beyond Needle(s) in the (Embodied) Haystack? ∞-THOR is an infinite len sim framework + guide on (new) architectures/training methods for VLA models
1102
Reposted by Prithviraj "Raj" Ammanabrolu
Mark Riedl @markriedl.bsky.social · 19/05/2025
I'm presenting today on AI Agents vs Agency Law at the 12th Governance of Emerging Technologies and Science Conference events.asucollegeoflaw.com/gets/ If agency law were to be applied to AI Agents (mostly in ecommerce settings), where does current AI align with the laws and where does it not?
2134
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 15/05/2025
I like the Ultra Scale Playbook from Huggingface and give it to my MS/first year PhD students to read as a prereq huggingface.co/spaces/nanot... Is there a "RLSys" version of this on scaling RL+LLM training? If not + there's OSS community interest, I'll prob write one?
080
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 14/05/2025
that's what's hot
150
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 13/05/2025
We know people don't read tho maybe I see the confusion cause RL with verifiable rewards... is all RL before learned rewards
010
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 10/05/2025
We now have a whole YouTube video explaining our MINDcraft paper, check it out! youtu.be/MeEcxh9St24
youtu.be
Mindcraft Research Paper!
YouTube video by Emergent Garden
1113
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 08/05/2025
If... circumstances ... were different maybe I'd enjoy the other stuff like fund raising more too but right now
040
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 08/05/2025
The part of the Prof job I'm enjoying by far the most right now is teaching actually
170
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 07/05/2025
Working on technical approaches to human AI collab in the mid term will help us focus on how to make sure systems stay under human control implicitly and explicitly. This is also why I continue to maintain an academic affiliation, companies are simply not incentivized to do this
030
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 07/05/2025
The bar to be considered even "competent" will get very high and those off loading excess brain power will get automated away sooner. Skillsets required in the workforce have changed and the rate of human employment will depend in the near term on how quickly universities adapt
100
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 07/05/2025
This is reasonably written and echoes many of my own fears. The upside of AI is too huge to pass up but also there's a high chance that the vast majority of humanity is on track to becoming economically obsolete without really any transition plan www.theguardian.com/books/2025/m...
110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 07/05/2025
What's with these arguments over whether X or Y or whatever was the first LLM RL library? These all came in the last 3 months We wrote multiturn RL4LMs like 3+ years ago github.com/allenai/RL4LMs There were other simple versions even before. ML ppl approaching goldfish memory
github.com
GitHub - allenai/RL4LMs: A modular RL library to fine-tune language models to human preferences
A modular RL library to fine-tune language models to human preferences - allenai/RL4LMs
082