Reposted by John (Yueh-Han) ChenMaksym Andriushchenko @maksym-andr.bsky.social · 19/06/2025🚨Excited to release OS-Harm! 🚨 The safety of computer use agents has been largely overlooked. We created a new safety benchmark based on OSWorld for measuring 3 broad categories of harm: 1. deliberate user misuse, 2. prompt injections, 3. model misbehavior. 132
Reposted by John (Yueh-Han) ChenNYU Center for Data Science @nyudatascience.bsky.social · 29/08/2025Frontier AI systems failed to reliably flag safety risks related to more than 40% of common safety facts tested in the SAGE‑Eval benchmark by Yueh-Han (John) Cheni, @guydav.bsky.social, and @brendenlake.bsky.social. nyudatascience.medium.com/even-the-top...nyudatascience.medium.comEven the Top LLM Failed to Reliably Flag Some Risks Related to 40% of Safety FactsCDS’ SAGE‑Eval shows top‑performing AI models failed at least 42% of safety warnings in novel scenarios. 022
Reposted by John (Yueh-Han) ChenNYU Center for Data Science @nyudatascience.bsky.social · 30/05/2025CDS PhD student @vishakhpk.bsky.social, with co-authors @johnchen6.bsky.social, Jane Pan, Valerie Chen, and CDS Associate Professor @hhexiy.bsky.social, has published new research on the trade-off between originality and quality in LLM outputs. Read more: nyudatascience.medium.com/in-ai-genera...nyudatascience.medium.comIn AI-Generated Content, A Trade-Off Between Quality and OriginalityNew research from CDS researchers maps the trade-off between originality and quality in LLM outputs. 122
Reposted by John (Yueh-Han) ChenGuy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 30/05/2025Fantastic new work by @johnchen6.bsky.social (with @brendenlake.bsky.social and me trying not to cause too much trouble). We study systematic generalization in a safety setting and find LLMs struggle to consistently respond safely when we vary how we ask naive questions. More analyses in the paper! 0103
Reposted by John (Yueh-Han) ChenBrenden Lake @brendenlake.bsky.social · 29/05/2025Failures of systematic generalization in LLMs can lead to real-world safety issues. New paper by @johnchen6.bsky.social and @guydav.bsky.social, arxiv.org/abs/2505.21828 052
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Do LLMs show systematic generalization of safety facts to novel scenarios? Introducing our work SAGE-Eval, a benchmark consisting of 100+ safety facts and 10k+ scenarios to test this! - Claude-3.7-Sonnet passes only 57% of facts evaluated - o1 and o3-mini passed <45%! 🧵 130
Reposted by John (Yueh-Han) ChenVishakh Padmakumar @vishakhpk.bsky.social · 29/04/2025What does it mean for #LLM output to be novel? In work w/ johnchen6.bsky.social, Jane Pan, Valerie Chen and He He, we argue it needs to be both original and high quality. While prompting tricks trade one for the other, better models (scaling/post-training) can shift the novelty frontier 🧵 274
Reposted by John (Yueh-Han) ChenAi2 @ai2.bsky.social · 26/03/2025Meet Ai2 Paper Finder, an LLM-powered literature search system. Searching for relevant work is a multi-step process that requires iteration. Paper Finder mimics this workflow — and helps researchers find more papers than ever 🔍 611723