Sign in

Peter Hase

@peterbhase.bsky.social
515 followers 450 following 8 posts

Visiting Scientist at Schmidt Sciences. Visiting Researcher at Stanford NLP Group Interested in AI safety and interpretability Previously: Anthropic, AI2, Google, Meta, UNC Chapel Hill

PostsRepliesMedia
Peter Hase @peterbhase.bsky.social · 14/07/2025
Overdue job update — I am now: - A Visiting Scientist at @schmidtsciences.bsky.social, supporting AI safety & interpretability - A Visiting Researcher at Stanford NLP Group, working with @cgpotts.bsky.social So grateful to keep working in this fascinating area—and to start supporting others too :)
351
Peter Hase @peterbhase.bsky.social · 19/05/2025
Are p-values missing in AI research? Bootstrapping makes model comparisons easy! Here's a new blog/colab with code for: - Bootstrapped p-values and confidence intervals - Combining variance from BOTH sample size and random seed (eg prompts) - Handling grouped test data Link ⬇️
100
Reposted by Peter Hase
Vaidehi Patil @vaidehipatil.bsky.social · 07/05/2025
🚨 Introducing our @tmlrorg.bsky.social paper “Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation” We present UnLOK-VQA, a benchmark to evaluate unlearning in vision-and-language models, where both images and text may encode sensitive or private information.
1128
Reposted by Peter Hase
Geoffrey Irving @girving.bsky.social · 05/03/2025
AISI has a new grant program for funding academic and nonprofit-affiliated research in (1) safeguards to mitigate misuse risk and (2) AI control and alignment to mitigate loss of control risk. Please apply! 🧵 www.aisi.gov.uk/grants
aisi.gov.uk
Grants | The AI Security Institute (AISI)
View AISI grants. The AI Security Institute is a directorate of the Department of Science, Innovation, and Technology that facilitates rigorous research to enable advanced AI governance.
195
Reposted by Peter Hase
Elias Stengel-Eskin @esteng.bsky.social · 23/01/2025
🎉Very excited that our work on Persuasion-Balanced Training has been accepted to #NAACL2025! We introduce a multi-agent tree-based method for teaching models to balance: 1️⃣ Accepting persuasion when it helps 2️⃣ Resisting persuasion when it hurts (e.g. misinformation) arxiv.org/abs/2410.14596 🧵 1/4
1218
Peter Hase @peterbhase.bsky.social · 11/01/2025
Anthropic Alignment Science is sharing a list of research directions we are interested in seeing more work on! Blog post below 👇
0121
Reposted by Peter Hase
Sam Bowman @sleepinyourhat.bsky.social · 18/12/2024
New work from my team at Anthropic in collaboration with Redwood Research. I think this is plausibly the most important AGI safety result of the year. Cross-posting the thread below:
Title card: Alignment Faking in Large Language Models by Greenblatt et al.
512629
Reposted by Peter Hase
ACL 2027 @aclmeeting.bsky.social · 16/12/2024
We invite nominations to join the ACL2025 PC as reviewer or area chair(AC). Review process through ARR Feb cycle. Tentative timeline: Review 1-20 Mar 2025, Rebuttal is 26-31 Mar 2025. ACs must be available throughout the Feb cycle. Nominations by 20 Dec 2024: shorturl.at/TaUh9 #NLProc #ACL2025NLP
forms.gle
Volunteer to join ACL 2025 Programme Committee
Use this form to express your interest in joining the ACL 2025 programme committee as a reviewer or area chair (AC). The review period is 1st to 20th of March 2025. ACs need to be available for variou...
01112
Peter Hase @peterbhase.bsky.social · 10/12/2024
Recruiting reviewers + ACs for ACL 2025 in Interpretability and Analysis of NLP Models - DM me if you are interested in emergency reviewer/AC roles for March 18th to 26th - Self-nominate for positions here (review period is March 1 through March 20): docs.google.com/forms/d/e/1F...
docs.google.com
Volunteer to join ACL 2025 Programme Committee
Use this form to express your interest in joining the ACL 2025 programme committee as a reviewer or area chair (AC). The review period is 1st to 20th of March 2025. ACs need to be available for variou...
2115
Reposted by Peter Hase
Serena Booth @reniebird.bsky.social · 07/12/2024
I'm hiring PhD students at Brown CS! If you're interested in human-robot or AI interaction, human-centered reinforcement learning, and/or AI policy, please apply. Or get your students to apply :). Deadline Dec 15, link: cs.brown.edu/degrees/doct.... Research statement on my website, www.slbooth.com
cs.brown.edu
Brown CS: Applying To Our Doctoral Program
We expect strong results from our applicants in the following:
02814
Peter Hase @peterbhase.bsky.social · 06/12/2024
I will be at NeurIPS next week! You can find me at the Wed 11am poster session, Hall A-C #4503, talking about linguistic calibration of LLMs via multi-agent communication games.
1100