Sign in

Yijia Shao

@echoshao8899.bsky.social
164 followers 56 following 44 posts

CS PhD student @StanfordNLP cs.stanford.edu/~shaoyj

PostsRepliesMedia
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
It’s also my honor to have economists from ‪@stanforddel.bsky.social join this project. As headlines are saying 2025 is a year of agents, we believe AI agent development is not solely a technical thing. Thanks Humishka, Yucheng, Jiaxin, David, ‪@erikbryn.bsky.social‬ and ‪@diyiyang.bsky.social!
010
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
This project would not have been possible without the thoughtful participation of the 1,500+ domain workers. Many of those we contacted cold on LinkedIn thanked us for amplifying their voices—but truly, the honor is ours.
110
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
🚀 We’re making the WORKBank database public and building an interactive data explorer! 👇 To get notified when it’s live or request an occupation we missed (see Appendix D.1 in our paper), drop a comment below. forms.gle/ocDWGhRDS8y6...
forms.gle
WORKBank Database: Feedback & Interest Form
In our paper, we develop a novel auditing framework to assess which occupational tasks workers want AI agents to automate or augment, and how those desires align with the current technological capabil...
110
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
Mapping tasks to skills–and comparing currently high-paid skills and required human agency as AI agents enter the workforce—we see: core human strengths move from data processing toward interpersonal and organizational skills. Read our blog post: futureofwork.saltlab.stanford.edu
110
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
The study also reveals insights on the future of HUMAN work. Mapping the Human Agency Scale across jobs shows which roles AI can’t replace. Currently, only Mathematicians & Aerospace Engineers have most AI expert ratings that fall into H5 (Human Involvement Essential).
100
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
Despite the buzz around "AI software engineers," "AI journalists," etc., our Human Agency Scale uncovers task-level nuances within every occupation. We suggest that AI agent R&D and products account for them for more responsible, higher-quality adoption.
100
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
Workers generally prefer higher levels of human agency, hinting at friction as AI capabilities advance. From transcript analysis, the top collaboration model envisioned by workers is “role-based” AI support (23.1%) - utilizing AI systems that embody specific roles.
100
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
The impact of AI agents on work isn’t just a binary “automate or not.” We introduce the Human Agency Scale: a 5-level scale to capture the spectrum between automation and augmentation--where technology complements and enhances human capabilities.
121
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
Jointly considering worker desire and technological capability allows us to classify tasks into four zones to guide AI agent deployment and development. Alarmingly, 41.0% of YC companies are mapped to Low Priority and Automation “Red Light” Zone.
100
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
We rank tasks by worker desire for automation. For 46.1% of tasks receive a positive attitude (>3/5) – with notable variation across sectors. Transcript analysis reveals top concerns: (1) lack of trust (45%), (2) fear of job replacement (23%), (3) loss of human touch (16.3%)
100
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
In our new paper: arxiv.org/abs/2506.06576 We collaborate with economists to develop an audio-enhanced auditing framework. - 1500 domain workers from 104 occupations shared their desires. - 52 AI agent researchers & developers evaluated today’s technological capabilities.
110
Yijia Shao @echoshao8899.bsky.social · 12/06/2025
🚨 70 million US workers are about to face their biggest workplace transmission due to AI agents. But nobody’s asking them what they want. While AI R&D races to automate everything, we took a different approach: auditing what workers want vs. what AI can deliver across the US workforce.🧵
1227
Reposted by Yijia Shao
Hao Zhu 朱昊 @zhuhao.me · 04/03/2025
We are getting closer to have agents operating in the real physical world. However, can we trust frontier models to make embodied decisions 🎮 aligned with human norms 👩‍⚖️ ? With EgoNormia, a 1.8k ego-centric video 🥽 QA benchmark, we show that this is surprisingly challenging!
1239
Yijia Shao @echoshao8899.bsky.social · 14/02/2025
Hi, I found your work very interesting and hope to have a chance to reach out. Is there a way to contact you? I tried DM on this site and redit but both fails. Thank you so much for your consideration!
cs.stanford.edu
000
Yijia Shao @echoshao8899.bsky.social · 12/02/2025
Thanks Vinay, Yucheng, John & @diyiyang.bsky.social for the amazing collaboration, and to all the friends—met or yet to be met—who shared suggestions for the platform release! The release won't be possible without the generous support from US Navy Research, NSF, Google, and Microsoft Azure!
010
Yijia Shao @echoshao8899.bsky.social · 12/02/2025
Try it out today at cogym.saltlab.stanford.edu! Read our preprint to learn more details: arxiv.org/abs/2412.15701
110
Yijia Shao @echoshao8899.bsky.social · 12/02/2025
You can request official support for a new task or vote on existing task requests through our GitHub repository! github.com/SALT-NLP/col...
github.com
Build software better, together
GitHub is where people build software. More than 150 million people use GitHub to discover, fork, and contribute to over 420 million projects.
110
Yijia Shao @echoshao8899.bsky.social · 12/02/2025
We welcome contributions of new task environments and agents. Contributed agents will be deployed on our platform to study their interaction dynamics with real users. A great chance to distribute your agent in the wild!
110
Yijia Shao @echoshao8899.bsky.social · 12/02/2025
Collaborative Gym is now released at github.com/SALT-NLP/col.... Besides backend primitives, we also open-source our UI to facilitate human-agent interaction research. The UI resonates design of OpenAI canvas with side-by-side chat panel and a shared workspace for human and agent, but can do more!
github.com
GitHub - SALT-NLP/collaborative-gym: Framework and toolkits for building and evaluating collaborative agents that can work together with humans.
Framework and toolkits for building and evaluating collaborative agents that can work together with humans. - SALT-NLP/collaborative-gym
110
Yijia Shao @echoshao8899.bsky.social · 12/02/2025
🎉 For the first time ever: Collaborate with AI agents in real-time! Collaborative Gym UI is now IRB-approved and alive at cogym.saltlab.stanford.edu! A group of agents is eager to work with you. By providing feedback, you will see the agent's identity and its feedback to you!
120
Yijia Shao @echoshao8899.bsky.social · 12/02/2025
You can request official support for a new task or vote on existing task requests through our GitHub repository! github.com/SALT-NLP/col...
github.com
Build software better, together
GitHub is where people build software. More than 150 million people use GitHub to discover, fork, and contribute to over 420 million projects.
000
Yijia Shao @echoshao8899.bsky.social · 12/02/2025
We welcome contributions of new task environments and agents. Contributed agents will be deployed on our platform to study their interaction dynamics with real users. A great chance to distribute your agent in the wild!
100
Yijia Shao @echoshao8899.bsky.social · 12/02/2025
Collaborative Gym is now released at github.com/SALT-NLP/col.... Besides backend primitives, we also open-source our UI to facilitate human-agent interaction research. The UI resonates design of OpenAI canvas with side-by-side chat panel and a shared workspace for human and agent, but can do more!
github.com
GitHub - SALT-NLP/collaborative-gym: Framework and toolkits for building and evaluating collaborative agents that can work together with humans.
Framework and toolkits for building and evaluating collaborative agents that can work together with humans. - SALT-NLP/collaborative-gym
100
Reposted by Yijia Shao
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
LM agents today primarily aim to automate tasks. Can we turn them into collaborative teammates? 🤖➕👤 Introducing Collaborative Gym (Co-Gym), a framework for enabling & evaluating human-agent collaboration! I now get used to agents proactively seeking confirmations or my deep thinking.(🧵 with video)
12210
Yijia Shao @echoshao8899.bsky.social · 26/01/2025
Hi @narphorium.bsky.social , thank you! Can finally reply to you because our team wants to check whether the taxonomy can be used to examine other agentic systems (e.g. coding agents) first. It's indeed very useful. You can check out my recent blog post if interested: cs.stanford.edu/people/shaoy...
cs.stanford.edu
Hands-on Experience with Devin: Reflections from a Person Building and Evaluating Agentic Systems
Why I’m interested in making agentic systems collaborative.
121
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
[8/8] To me, Co-Gym stems from my SoP on building human-centered agentic systems 2 years ago. I am excited to see how agents could work with us and the demands this poses for advancing model intelligence! Thank you Vinay, Yucheng, John & @diyiyang.bsky.social for the amazing collaboration!
010
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
[7/8] We are working on making Co-Gym UI accessible to the public. Can’t wait to get more in-the-wild evaluations and observe more dynamics of human-agent collaboration. Stay tuned! Check out our arXiv paper first to learn more: arxiv.org/abs/2412.15701
arxiv.org
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
Recent advancements in language models (LMs) have sparked growing interest in developing LM agents. While fully autonomous agents could excel in many scenarios, numerous use cases inherently require t...
110
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
[6/8] We conducted a detailed error analysis by having authors annotate 300 trajectories. Collaborative agents expose significant limitations in current LMs and agent scaffoldings, with communication and situational awareness failures occurring in 65% and 40% of real trajectories.
130
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
[5/8] We built a user simulator and web UI to instantiate Co-Gym in simulated and real settings. Experiments reveal human-like patterns: collaborative inertia, where poor communication hinders delivery; and collaborative advantage, where human-agent teams outperform autonomous agents.
110
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
[4/8] Our vision builds on a long-standing dream in AI: to develop machines that act as teammates, not mere tools. This demands situational intelligence to take initiative, communicate, and adapt. Co-Gym offers an evaluation framework that assesses both collab outcomes and processes.
120
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
[3/8] How does Co-Gym enable collaborative agents? Our infra (1) focuses on environment design and (2) supports async interaction beyond turn-taking. We define primitives for public/private components in the shared env, as well as collaboration actions and notification protocol.
110
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
[2/8] Excitingly, collaborative agents consistently outperform their fully autonomous counterparts in terms of task performance, achieving win rates of 86% in Travel Planning, 74% in Tabular Analysis, and 66% in Related Work when evaluated by real users.
110
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
[1/8] While several HITL systems exist (e.g. OpenAI Canvas, our Collaborative STORM), what makes human-agent collab special? Agents need autonomy to be useful, yet the goal is empowering humans. We start with three tasks: travel planning, surveying related work, and tabular analysis
110
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
Check out the demo video to see what our framework can do: drive.google.com/file/d/1obls...
drive.google.com
co-gym-twitter-teaser.mp4
110
Yijia Shao @echoshao8899.bsky.social · 17/01/2025
LM agents today primarily aim to automate tasks. Can we turn them into collaborative teammates? 🤖➕👤 Introducing Collaborative Gym (Co-Gym), a framework for enabling & evaluating human-agent collaboration! I now get used to agents proactively seeking confirmations or my deep thinking.(🧵 with video)
12210
Yijia Shao @echoshao8899.bsky.social · 10/12/2024
Super fun and easy to play with! Check it out ⬇️
010
Reposted by Yijia Shao
Stanford NLP Group @stanfordnlp.bsky.social · 09/12/2024
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action Yijia Shao, Tianshi Li, Weiyan Shi, Yanchen Liu, Diyi Yang Th, Dec 12, 11:00 PST - Poster Session 3 West
111
Reposted by Yijia Shao
Stanford NLP Group @stanfordnlp.bsky.social · 09/12/2024
Papers (partly) from @stanfordnlp at #NeurIPS 2024: Oral: Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making Manling Li · Shiyu Zhao · Qineng Wang · Kangrui Wang · … · Weiyu Liu · Percy Liang · Li Fei-Fei · Jiayuan Mao · Jiajun Wu Wed 11 Dec 11:50 PM UTC [East Ballroom A, B]
1114
Yijia Shao @echoshao8899.bsky.social · 06/12/2024
Finally, thanks Tianshi Li, Weiyan Shi, Yanchen Liu, @diyiyang.bsky.social for bringing in different expertise!! The work is partially supported by grants from ONR, Meta, and research credits from OpenAI.
010
Yijia Shao @echoshao8899.bsky.social · 06/12/2024
Check out our paper, code, data to learn more! Paper: arxiv.org/abs/2409.00138 Website: salt-nlp.github.io/PrivacyLens/
arxiv.org
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
As language models (LMs) are widely utilized in personalized communication scenarios (e.g., sending emails, writing social media posts) and endowed with a certain level of agency, ensuring they act in...
122
Yijia Shao @echoshao8899.bsky.social · 06/12/2024
In our paper, we explore the impact of prompting. Unfortunately, simple prompt engineering does little to mitigate privacy leakage of LM agents’ actions. We also examine the safety-helpfulness trade-off and conduct qualitative analysis to uncover more insights.
100
Yijia Shao @echoshao8899.bsky.social · 06/12/2024
We collected 493 negative privacy norms to seed PrivacyLens. Our results reveal a discrepancy between QA probing results and LMs’ actions in task execution. GPT-4 and Claude-3-Sonnet answer nearly all questions correctly, but they leak information in 26% and 38% of cases!
100
Yijia Shao @echoshao8899.bsky.social · 06/12/2024
With negative privacy norms, vignettes, trajectories, PrivacyLens conducts a multi-level evaluation by (1) assessing LMs on their ability to identify sensitive data transmission through QA probing, (2) evaluating whether LM agents’ final actions leak the sensitive information.
110
Yijia Shao @echoshao8899.bsky.social · 06/12/2024
Evaluating LMs’ actions in applications is more contextualized. But how to create test cases? PrivacyLens offers a data construction pipeline that procedurally converts the norms into a vignette and then to an agent trajectory via template-based generation and sandbox simulation.
120
Yijia Shao @echoshao8899.bsky.social · 06/12/2024
Once we collect these privacy norms, a direct way for evaluation is by using a template to turn the tuple into a multi-choice question. However, how LMs perform when answering probing questions may not be consistent with how they act in agentic applications.
110
Yijia Shao @echoshao8899.bsky.social · 06/12/2024
Humans protect privacy not by always avoiding sharing sensitive data, but by adhering to these norms during data use and communication with others. A well-established framework for privacy norms is the Contextual Integrity theory which expresses data transmission with a 5-tuple.
110
Yijia Shao @echoshao8899.bsky.social · 06/12/2024
Why is this important? While many studies have investigated LMs memorizing training data, a lot of private data or sensitive information is actually exposed to LMs at inference time, especially when we are using them for daily assistance.
110
Yijia Shao @echoshao8899.bsky.social · 06/12/2024
Excited to present our PrivacyLens paper at #NuerIPS next week! We explore LM agent privacy risks when deployed as personal assistants. (Details in thread) I am working on developing LM agents as collaborative research partners, learning aids, personal assistants, and more. Let's connect and chat!!
272
Reposted by Yijia Shao
Stanford NLP Group @stanfordnlp.bsky.social · 18/11/2024
Missed some – or all – of our papers at #EMNLP2024? It's not too late to catch up using this handy list from the Stanford AI Lab blog: ai.stanford.edu/blog/emnlp-2...
0244