Sign in

aakriti1kumar.bsky.social

@aakriti1kumar.bsky.social
37 followers 62 following 11 posts
PostsRepliesMedia
Reposted by @aakriti1kumar.bsky.social
Matt Groh @mattgroh.bsky.social · 11/02/2026
🚨 New in @natmachintell.nature.com 🚨 We collected 9000+ annotations of empathic communication in convos from experts, crowds & LLMs across 4 NLP/comms/psych frameworks LLM judgment exceeds crowds' reliability & nearly matches experts Soft skills can now be reliably measured by LLMs 🧵
2144
aakriti1kumar.bsky.social @aakriti1kumar.bsky.social · 17/06/2025
How do we reliably judge if AI companions are performing well on subjective, context-dependent, and deeply human tasks? 🤖 Excited to share the first paper from my postdoc (!!) investigating when LLMs are reliable judges - with empathic communication as a case study 🧐 🧵👇
140
aakriti1kumar.bsky.social @aakriti1kumar.bsky.social · 02/04/2025
Super cool opportunity to work with brilliant scientists and fantastic mentors @mattgroh.bsky.social and Dashun Wang 🌟🌟 Feel free to reach out!
020
Reposted by @aakriti1kumar.bsky.social
Abhishek Sharma @abhishekshar.bsky.social · 23/01/2025
Our paper: Decision-Point Guided Safe Policy Improvement We show that a simple approach to learn safe RL policies can outperform most offline RL methods. (+theoretical guarantees!) How? Just allow the state-actions that have been seen enough times! 🤯 arxiv.org/abs/2410.09361
arxiv.org
Decision-Point Guided Safe Policy Improvement
Within batch reinforcement learning, safe policy improvement (SPI) seeks to ensure that the learnt policy performs at least as well as the behavior policy that generated the dataset. The core challeng...
041