Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 01/05/2026My lab and collaborators had 4 papers on everything from multi objective alignment, reasoning during mid training, multimodal synthetic data, and generating RL tasks accepted to #ICML2026! Come hang out with us in Seoul and we can talk about the exciting follow-ups! Papers at pearls.ucsd.edupearls.ucsd.eduPEARLS Lab @ UCSD | AboutLab website for the PEARLS lab @ University of California, San Diego CSE 260
Reposted by Prithviraj "Raj" AmmanabroluMark Riedl @markriedl.bsky.social · 03/12/2025My former MS student Chris Cui (now PhD student with @rajammanabrolu.bsky.social)motivates Text Adventure Games as testbeds for reasoning. Provides a new benchmark suite of text games. Observes that Zork still kicks LLM’s butts despite training on walkthroughs arxiv.org/abs/2504.14128arxiv.orgTALES: Text Adventure Learning Environment SuiteReasoning is an essential skill to enable Large Language Models (LLMs) to interact with the world. As tasks become more complex, they demand increasingly sophisticated and diverse reasoning capabiliti... 0236
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 24/11/2025Ziyi Zhang, Shengqi Li (on PhD app market!) for multi agent D&D sim creation + RL openreview.net/pdf?id=3Op7k... Me mostly if you want boba and beach recs (or NVIDIA full time RS roles I guess, I'm hiring a few ppl there but not UCSD)openreview.net 010
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 24/11/2025Jenny Shen for pluralistic alignment + human feedback arxiv.org/abs/2510.01167 Chris Cui for text sims + RL + scalable oversight of reasoning models arxiv.org/abs/2504.14128 Lucas (on industry market) for how reasoning emerges from mid training/SFT to RL lucasdino.github.io/assets/files... 130
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 24/11/2025Bosung Kim for all things VLA, embodied AI, and long context memory arxiv.org/abs/2505.16928 Ruiyi Wang for multi turn agentic RL and all the RL infra in and outs arxiv.org/abs/2510.01132 110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 24/11/2025My entire PEARLS Lab, and many NVIDIA colleagues, will be at #neurips2025 to chat about their latest. Some papers accepted to the conf are already outdated so just reach out to. Thread 🧵 110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 12/11/2025Yay congrats, Mark! Well deserved! It's def a required reading for all things comp storytelling (and creativity!) 010
Reposted by Prithviraj "Raj" AmmanabroluMark Riedl @markriedl.bsky.social · 12/11/2025I am extremely honored and humbled to have been awarded a Test-of-Time award for my 2005 paper "From Linear Story Generation to Branching Story Graphs" with R. Michael Young 5725
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 03/11/2025I've done a few versions of this talk but this is the first that's been recorded publicly, thanks to IVADO Montreal A good overview of things my lab has been up to in the last year or so at least in balancing safety/capabilities of (embodied) AI Agents www.youtube.com/watch?v=S-kV...youtube.comNavigating the Safety-Capability Spectrum when Teaching Agents with Feedback -Prithviraj AmmanabroluYouTube video by IVADO 081
Reposted by Prithviraj "Raj" AmmanabroluRuiyi Wang @ruiyiwang.bsky.social · 26/10/2025🔥Excited to share our new work: "A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning"! We study what actually works for agentic multi-turn RL with varying 🌎Environment, 🤖Policy, and ⭐Reward. We conduct various ablations and empirical analysis on 🧩TextWorld, 🧙ALFWorld, and 🧑💻SWE-Gym. 192
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 01/10/2025My students will be presenting IRPO arxiv.org/abs/2504.15477 and a paper on Personalized RLHF arxiv.org/abs/2504.07070 on Wed onwardsarxiv.orgIn-context Ranking Preference OptimizationRecent developments in Direct Preference Optimization (DPO) allow large language models (LLMs) to function as implicit ranking models by maximizing the margin between preferred and non-preferred respo... 010
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 01/10/2025My students will be presenting papers next Wed/Thursday so be sure to check those out too 130
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 01/10/2025I'll be at #CoLM2025 and the IVADO agents workshop right before in Montreal. My students will be presenting two papers in the main conf. I'll also do a ws keynote where I'll talk about some of our latest. Come by and say hi next week! 180
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 29/06/2025I'm probably mostly going to stop posting on this site. There's close to no engagement and it's not worth the effort to cross post for the amount of time that takes. Find me elsewhere / email me 450
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 19/06/2025I recently left Mosaic/Databricks Research. It's been a ride building out the RL team from <4 ppl to 20+ across two companies & acquisition +figuring out RL as a Service in prod. Mosaic had insane talent density Some "relaxation" while I put out Prof fires for a smol bit then new adventures! 070
Reposted by Prithviraj "Raj" AmmanabroluMark Riedl @markriedl.bsky.social · 17/06/2025If you work in the intersection of NLP and games/narrative, then this workshop is for you! wordplay-workshop.github.io/cfp/ Organized by the amazing @laramartin.net and @rajammanabrolu.bsky.social (among others)wordplay-workshop.github.io/call_for_papersOfficial website for the Wordplay Workshop at EMNLP 2025. Exploring interactive narratives, text-adventure games, and AI agents in language-based environments. Join us in Suzhou, China, November 5th-9... 0105
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 16/06/2025The thing that feels so off about the core tech world is that every convo is very transactional. Maybe true elsewhere too. "Oh you're an expert in RL, can you answer questions about my new startup?" Every single (Bay) party. No I do not want to consult. I just wanna hang out. 190
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 10/06/2025Of all the labeling startups out there to acquihire, this was... an interesting choice. Says a lot actually 100
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 09/06/2025. @bosungkim.bsky.social will be at #CVPR2025 in Nashville this week to present this and just generally talk about scaling memory for embodied agents! Catch her at the poster sessions and also the Foundation Models meets Embodied Agents Workshop on Wed 020
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 09/06/2025Yes AI for edu is a thing but almost all vanilla LLMs just railroad students into answers. Complete cognitive offload is not useful for improving learning outcomes 020
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 09/06/2025I've heard this personally from multiple PMs at AI companies. Students are one of the biggest demographics and they need to "break in" and have even more usage to improve their metrics. Classic corporate economic incentives 174
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 06/06/2025Tis the era of bringing back every AI benchmark ever but this time by the LLM people and for the LLMs 021
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 05/06/2025Had a fun little visit to Cambridge LTL where I talked about a bunch of my lab's latest papers including some still not public with the key takeaway that "RL can absolutely learn new things and is not just resurfacing knowledge" talks.cam.ac.uk/show/archive...talks.cam.ac.uktalks.cam : Language Technology Lab Seminars 041
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 03/06/2025That's fair, I guess I should rephrase to "regardless of a possible common prior, it's nearly impossible for different providers to have the same representations pop out of their post trained LLM" 100
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 03/06/2025The moral of the story here is basically that who is making your LLM really matters. Internal use cases critical to their businesses will always influence data distributions and everything downstream of that. This is in contrast to things like Platonic Representation Hypothesis 150
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 03/06/2025Interesting tidbit from UCSD's Victor Shih on a podcast talking about Chinese AGI efforts is that Deepseek is good at Chinese govt doc understanding cause that's what affects stock prices most and DS is a hedge fund. www.youtube.com/watch?v=b1Te...youtu.beXi Jinping’s paranoid approach to AGI, debt crisis, & Politburo politics — Victor ShihYouTube video by Dwarkesh Patel 150
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 02/06/2025The top two scores are 332 not 322 but other than the typo the rest of this list seems legit and consistent across multiple sources www.indiatvnews.com/education/hi... x.com/RejaullahmdM...indiatvnews.comJEE Advanced 2025 topper list out, Rajit Gupta secures AIR 1 with 332 marks: Check complete list of toppersIIT Kanpur has released the topper list for JEE Advanced 2025. This year, Rajit Gupta achieved AIR 1 by scoring 332 out of 360 marks. Devdutta Majhi is the leading female candidate with AIR 16. Scroll... 000
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 02/06/2025Looks like Gemini gets AIR 6 in #JEE2025 with a score of 323 Only 5 highschoolers in all India do better than an LLM in the single most important exam of their to get into the IITs The legacy edu selection systems are now worse than useless 140
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 01/06/2025I get prepping for worst case scenarios but a lot of AI Safety debates I somehow end up these days in boil down to "assume you have Machine God in a box, now tell me how to align it" I could rant for hours but seriously y'all this isn't productive 180
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 24/05/2025Here for the afternoon shift! 010
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025Paper: arxiv.org/abs/2505.16928 Website/code/data: pearls-lab.github.io/infini-thor/ Led by @bosungkim.bsky.social who has done a fantastic job on this in the last bit. Full stack from Unity gamedev to Big Model Scaler. Watch out for her in the embodied agent space!arxiv.orgBeyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context ReasoningWe introduce $\infty$-THOR, a new framework for long-horizon embodied tasks that advances long-context understanding in embodied AI. $\infty$-THOR provides: (1) a generation framework for synthesizing... 020
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025But even then agents only perform well up to 130k after which perf sharply decreases and all architectures and additional context extension methods we modified fail after ~400k. None make it to the 3m context sample test set we use let alone infinite. Lots of space for progress! 110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025And find that when accounting for hardware constraints, only a specific combo of interleaved VLA with a mix of context parallelism + some extension with high pre training context window size works well. We detail the exact architecture and describe potential improvements 100
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025The first is a static eval called Needle(s) in the Embodied Haystack, which is like QA asking agents to post hoc analyze their trajectories putting many needles together We then go Beyond this with interactive RL style evals to see how well models interact with a changing env 110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025So we first extended AI2's THOR embodied sim to continue generating meaningful tasks that are effectively infinite in length. Eg things you do at context len 7k can get reused at 3m We do a thorough analysis of many types of architectures x training methods on two new evals 110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025The combo of Big Models, embodied agents, and long context has much potential. But it's very unclear what works and what doesn't. Most robotics sims don't have high cognitive complexity+long contexts and other types of sims don't require physical understanding 100
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 23/05/2025"Foundation" models for embodied agents are all the rage but how to actually do complex looong context reasoning? Can we scale Beyond Needle(s) in the (Embodied) Haystack? ∞-THOR is an infinite len sim framework + guide on (new) architectures/training methods for VLA models 1102
Reposted by Prithviraj "Raj" AmmanabroluMark Riedl @markriedl.bsky.social · 19/05/2025I'm presenting today on AI Agents vs Agency Law at the 12th Governance of Emerging Technologies and Science Conference events.asucollegeoflaw.com/gets/ If agency law were to be applied to AI Agents (mostly in ecommerce settings), where does current AI align with the laws and where does it not? 2134
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 15/05/2025I like the Ultra Scale Playbook from Huggingface and give it to my MS/first year PhD students to read as a prereq huggingface.co/spaces/nanot... Is there a "RLSys" version of this on scaling RL+LLM training? If not + there's OSS community interest, I'll prob write one? 080
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 13/05/2025We know people don't read tho maybe I see the confusion cause RL with verifiable rewards... is all RL before learned rewards 010
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 10/05/2025We now have a whole YouTube video explaining our MINDcraft paper, check it out! youtu.be/MeEcxh9St24youtu.beMindcraft Research Paper!YouTube video by Emergent Garden 1113
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 08/05/2025If... circumstances ... were different maybe I'd enjoy the other stuff like fund raising more too but right now 040
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 08/05/2025The part of the Prof job I'm enjoying by far the most right now is teaching actually 170
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 07/05/2025Working on technical approaches to human AI collab in the mid term will help us focus on how to make sure systems stay under human control implicitly and explicitly. This is also why I continue to maintain an academic affiliation, companies are simply not incentivized to do this 030
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 07/05/2025The bar to be considered even "competent" will get very high and those off loading excess brain power will get automated away sooner. Skillsets required in the workforce have changed and the rate of human employment will depend in the near term on how quickly universities adapt 100
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 07/05/2025This is reasonably written and echoes many of my own fears. The upside of AI is too huge to pass up but also there's a high chance that the vast majority of humanity is on track to becoming economically obsolete without really any transition plan www.theguardian.com/books/2025/m... 110
Prithviraj "Raj" Ammanabrolu @rajammanabrolu.bsky.social · 07/05/2025What's with these arguments over whether X or Y or whatever was the first LLM RL library? These all came in the last 3 months We wrote multiturn RL4LMs like 3+ years ago github.com/allenai/RL4LMs There were other simple versions even before. ML ppl approaching goldfish memorygithub.comGitHub - allenai/RL4LMs: A modular RL library to fine-tune language models to human preferencesA modular RL library to fine-tune language models to human preferences - allenai/RL4LMs 082