Reposted by John (Yueh-Han) ChenMaksym Andriushchenko @maksym-andr.bsky.social · 19/06/2025🚨Excited to release OS-Harm! 🚨 The safety of computer use agents has been largely overlooked. We created a new safety benchmark based on OSWorld for measuring 3 broad categories of harm: 1. deliberate user misuse, 2. prompt injections, 3. model misbehavior. 132
Reposted by John (Yueh-Han) ChenNYU Center for Data Science @nyudatascience.bsky.social · 29/08/2025Frontier AI systems failed to reliably flag safety risks related to more than 40% of common safety facts tested in the SAGE‑Eval benchmark by Yueh-Han (John) Cheni, @guydav.bsky.social, and @brendenlake.bsky.social. nyudatascience.medium.com/even-the-top...nyudatascience.medium.comEven the Top LLM Failed to Reliably Flag Some Risks Related to 40% of Safety FactsCDS’ SAGE‑Eval shows top‑performing AI models failed at least 42% of safety warnings in novel scenarios. 022
Reposted by John (Yueh-Han) ChenNYU Center for Data Science @nyudatascience.bsky.social · 30/05/2025CDS PhD student @vishakhpk.bsky.social, with co-authors @johnchen6.bsky.social, Jane Pan, Valerie Chen, and CDS Associate Professor @hhexiy.bsky.social, has published new research on the trade-off between originality and quality in LLM outputs. Read more: nyudatascience.medium.com/in-ai-genera...nyudatascience.medium.comIn AI-Generated Content, A Trade-Off Between Quality and OriginalityNew research from CDS researchers maps the trade-off between originality and quality in LLM outputs. 122
Reposted by John (Yueh-Han) ChenGuy Davidson ✈️ NeurIPS 2025 @guydav.bsky.social · 30/05/2025Fantastic new work by @johnchen6.bsky.social (with @brendenlake.bsky.social and me trying not to cause too much trouble). We study systematic generalization in a safety setting and find LLMs struggle to consistently respond safely when we vary how we ask naive questions. More analyses in the paper! 0103
Reposted by John (Yueh-Han) ChenBrenden Lake @brendenlake.bsky.social · 29/05/2025Failures of systematic generalization in LLMs can lead to real-world safety issues. New paper by @johnchen6.bsky.social and @guydav.bsky.social, arxiv.org/abs/2505.21828 052
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Our data: huggingface.co/datasets/Yue...) Code: github.com/YuehHanChen/.... We recommend that AI companies use SAGE-Eval in pre-deployment evaluations to assess model reliability when addressing salient risks in naive user prompts.huggingface.coYuehHanChen/SAGE-Eval · Datasets at Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 010
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Overall, our findings suggest the systematicity gap: unlike humans—who generalize a safety fact learned in one context to any structurally related context—LLMs today exhibit only piecemeal safety, identifying critical knowledge in isolation but failing to apply it broadly. 12/🧵 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025> We compared the performance of OLMo-2-32B-SFT to OLMo-2-32B-DPO, which is the SFT version further trained with DPO. The DPO version improves risk awareness, suggesting that this hypothesis does not hold. 11/🧵 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Hypothesis 2: Does RLHF affect safety performance on SAGE-Eval? Would it be possible that humans prefer “less annoying” responses, potentially diminishing the presence of critical safety warnings? 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025> We used The Pile as a proxy for pre-training data of frontier models, and Google search result count as a secondary method. We failed to find a statistically significant correlation using both methods, suggesting that fact frequency alone doesn’t predict performance on SAGE-Eval. 10/🧵 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025To understand the root causes, we explore two hypotheses: Hypothesis 1: Is there any correlation between fact frequency in pre-training data and safety performance on SAGE-Eval? 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025In practice, deployed LLMs will face a vastly richer and more varied set of user prompts than any finite benchmark can cover. We show that model developers can forecast SAGE-Eval safety scores with at least one order of magnitude more prompts per fact with a power law fit. 9/🧵 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Finding 4: Model capability and training compute only weakly correlate with performance on SAGE-Eval, demonstrating that our benchmark effectively avoids “safetywashing”—a scenario where capability improvements are incorrectly portrayed as advancements in safety. 8/🧵 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Finding 3: certain tones degrade safety performance. 7/🧵 In real life, users might prompt LMs in different tones. The depressed tone reduces the safety score to 0.865, noticeably below the no-augmentation baseline of 0.907. 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Finding 2: Long context undermines risk awareness. Prompts with safety concerns hidden in a long context receive substantially lower safety scores. 6/🧵 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Finding 1: All frontier LLMs we tested score <58% safety scores. Our model-level safety score is defined as % of safety facts 100% passed all test scenario prompts (~100 scenarios per safety fact). 5/🧵 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Property 3: SAGE-Eval can be automatically evaluated: we confirm evaluation accuracy by manually labeling 100 model responses as safe or unsafe. In our experiments, we find perfect alignment between human judgments and an LLM-as-a-judge using frontier models as judges 4/🧵 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Property 2: SAGE-Eval is human-verified by 144 human annotators. If one human disagrees with the label, we manually edit or remove it. We then augment these questions with programming-based techniques (add typos or different tones) to extend each fact to around 100 test scenarios. 3/🧵 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Property 1: SAGE-Eval covers diverse safety categories—including Child, Outdoor Activities, and Medicine—and comprises 104 safety facts manually sourced from reputable organizations such as the CDC and FDA. 2/🧵 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025To evaluate the systematic generalization of safety knowledge to novel situations, we designed SAGE-Eval with 3 main properties: 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025>Do LLMs robustly generalize critical safety facts to novel scenarios? Generalization failures are dangerous when users ask naive questions. 1/🧵 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Arxiv: arxiv.org/pdf/2505.21828 Joint work with @guydav.bsky.social @brendenlake.bsky.social 🧵 starts below! 110
John (Yueh-Han) Chen @johnchen6.bsky.social · 29/05/2025Do LLMs show systematic generalization of safety facts to novel scenarios? Introducing our work SAGE-Eval, a benchmark consisting of 100+ safety facts and 10k+ scenarios to test this! - Claude-3.7-Sonnet passes only 57% of facts evaluated - o1 and o3-mini passed <45%! 🧵 130
Reposted by John (Yueh-Han) ChenVishakh Padmakumar @vishakhpk.bsky.social · 29/04/2025What does it mean for #LLM output to be novel? In work w/ johnchen6.bsky.social, Jane Pan, Valerie Chen and He He, we argue it needs to be both original and high quality. While prompting tricks trade one for the other, better models (scaling/post-training) can shift the novelty frontier 🧵 274
Reposted by John (Yueh-Han) ChenAi2 @ai2.bsky.social · 26/03/2025Meet Ai2 Paper Finder, an LLM-powered literature search system. Searching for relevant work is a multi-step process that requires iteration. Paper Finder mimics this workflow — and helps researchers find more papers than ever 🔍 611723