Sign in

Hao Zhu 朱昊

@zhuhao.me
575 followers 158 following 15 posts

AI researcher. Postdocing at Stanford NLP. Prev: PhD CMU LTI. Visit zhuhao.me Raising agents in the Opensocial.world

PostsRepliesMedia
Reposted by Hao Zhu 朱昊
Dirk Hovy @dirkhovy.bsky.social · 03/05/2025
We (w/ @diyiyang.bsky.social, @zhuhao.me, & Bodhisattwa Prasad Majumder) are excited to present our #NAACL25 tutorial on Social Intelligence in the Age of LLMs! It will highlight long-standing and emerging challenges of AI interacting w humans, society & the world. ⏰ May 3, 2:00pm-5:30pm Room Pecos
0156
Reposted by Hao Zhu 朱昊
Tomer Ullman @tomerullman.bsky.social · 13/03/2025
woooooo! Out in Child Development: "Learning Loopholes: The Development of Intentional Misunderstandings in Children" paper: srcd.onlinelibrary.wiley.com/doi/10.1111/... preprint-pdf: www.tomerullman.org/papers/kids_...
25313
Hao Zhu 朱昊 @zhuhao.me · 11/03/2025
This works like magic!
120
Hao Zhu 朱昊 @zhuhao.me · 07/03/2025
I have similar observations. But as a reviewer, I have to be honest that I cannot check each claim about previous papers, and these kinds of false references are often considered as minor issues (not really) comparing to novelty or empirical results.
100
Reposted by Hao Zhu 朱昊
Chris Paxton @cpaxton.bsky.social · 05/03/2025
New personal project with my friend Michael Cho: RoboPapers, a podcast where we chat with authors of cool robotics papers and post the discussion on YouTube and spotify. First one was with Duan Jiafei, who did the very cool paper SAM2Act, and it goes up Friday.
1323
Reposted by Hao Zhu 朱昊
Danny To Eun Kim @teknology.bsky.social · 05/03/2025
🚨New Breakthrough in Tip-of-the-Tongue (TOT) Retrieval Research! We address data limitations and offer a fresh evaluation method for these complex queries. Curious how TREC TOT track test queries are created? Check out this thread 🧵 and our paper 📄: arxiv.org/abs/2502.17776
arxiv.org
Tip of the Tongue Query Elicitation for Simulated Evaluation
Tip-of-the-tongue (TOT) search occurs when a user struggles to recall a specific identifier, such as a document title. While common, existing search systems often fail to effectively support TOT scena...
2188
Reposted by Hao Zhu 朱昊
Caleb Ziems @calebziems.com · 04/03/2025
EgoNormia (egonormia.org) exposes a major gap in Vision-Language Models understanding of the social world: they don't know how to behave when norms about the physical world *conflict* ⚔️ (<45% acc.) But humans are naturally quite good at this (>90% acc.) Check it out! ➡️ arxiv.org/abs/2502.20490
egonormia.org
EgoNormia: A Benchmark for Visual Frontier Models' Normative Reasoning
A large scale video dataset and a benchmark for evaluating frontier models' understanding of physical social norms through videos.
082
Hao Zhu 朱昊 @zhuhao.me · 04/03/2025
thanks to Leena Mathur and Su Li for helping with collecting robotics videos. thanks @michaelryan207.bsky.social @williamheld.com @echoshao8899.bsky.social @jyangballin.bsky.social @ellaminzhili.bsky.social @juliakruk.bsky.social @rewang.bsky.social @vidhijain.bsky.social @ybisk.me Dorsa Sadigh
020
Hao Zhu 朱昊 @zhuhao.me · 04/03/2025
Incredible collab with MohammadHossein Rezaei* (U of A) Yicheng Fu* Phil Cuvin* (U of T) @calebziems.com @yanzhe.bsky.social @diyiyang.bsky.social Upvote our paper at huggingface.co/papers/2502.... arxiv.org/abs/2502.20490
huggingface.co
Paper page - EgoNormia: Benchmarking Physical Social Norm Understanding
Join the discussion on this paper page
120
Hao Zhu 朱昊 @zhuhao.me · 04/03/2025
As always, we open source everything. Even our nicely made website: egonormia.org Please check out the leaderboard, the blog (w/Bibtex support), the code, data, as well as a data viewer.
egonormia.org
EgoNormia: A Benchmark for Visual Frontier Models' Normative Reasoning
A large scale video dataset and a benchmark for evaluating frontier models' understanding of physical social norms through videos.
131
Hao Zhu 朱昊 @zhuhao.me · 04/03/2025
We are getting closer to have agents operating in the real physical world. However, can we trust frontier models to make embodied decisions 🎮 aligned with human norms 👩‍⚖️ ? With EgoNormia, a 1.8k ego-centric video 🥽 QA benchmark, we show that this is surprisingly challenging!
1239
Reposted by Hao Zhu 朱昊
Shikhar Murty @shikharmurty.bsky.social · 06/02/2025
Want to make a browser agent for *any* domain like banking or healthcare? We propose methods for training LLMs with open-ended, unsupervised interaction on live websites: ✅ OSS SoTA on WebVoyager ✅ world's smallest high-performing web-agent Try it here: nnetnav.dev
192
Hao Zhu 朱昊 @zhuhao.me · 06/02/2025
Visit nnetnav.dev for more examples, code, and data.
nnetnav.dev
NNetNav - Unsupervised Browser Agents
Unsupervised learning for browser automation.
020
Hao Zhu 朱昊 @zhuhao.me · 06/02/2025
The key insight is that LLMs are good at understanding whether a traj is doing something reasonable and that guides efficient exploration and gives accurate labels. Be warned that deploying exploration algorithms in the real world has consequences -- monitor your agents closely.
120
Hao Zhu 朱昊 @zhuhao.me · 06/02/2025
Ever dreamed of AI agents learning through interacting with the open world unsupervisedly? Our latest preprint introduces NNetNav-Live which collects training data through exploration on real websites and hindsight labeling, which produces a SOTA OSS agent.
142
Hao Zhu 朱昊 @zhuhao.me · 10/12/2024
Our awesome team: @ellaminzhili.bsky.social @williamheld.com @michaelryan207.bsky.social Kunat Pipatanakul, Potsawee Manakul @diyiyang.bsky.social
080
Hao Zhu 朱昊 @zhuhao.me · 10/12/2024
My first bluesky post will be for my first project as a postdoc at Stanford. Talk Arena is our first step towards building audio LMs into interactive agents. Try it out and let me know what you think. talkarena.org
talkarena.org
Talk Arena
Interactive evaluation for audio models
2194
Reposted by Hao Zhu 朱昊
Will Held @williamheld.com · 10/12/2024
With an increasing number of Large *Audio* Models 🔊, which one do users like the most? Introducing talkarena.org — an open platform where users speak to LAMs and receive text responses. Through open interaction, we focus on rankings based on user preferences rather than static benchmarks. 🧵 (1/5)
Talk Arena: Interactive Evaluation of Large Audio Models
3318
Hao Zhu 朱昊 @zhuhao.me · 22/11/2024
matplotlib with customization. I can share the code with you
110
Hao Zhu 朱昊 @zhuhao.me · 22/11/2024
110
Hao Zhu 朱昊 @zhuhao.me · 21/11/2024
Would really appreciate it if I can be included. I build social intelligence models/agents that can cooperate with humans.
110
Hao Zhu 朱昊 @zhuhao.me · 19/11/2024
🙋‍♂️
010