Sign in

Valentina Pyatkin

@valentinapy.bsky.social
5.9K followers 585 following 77 posts

Postdoc in AI at the Allen Institute for AI & the University of Washington. 🌐 valentinapy.github.io

PostsRepliesMedia
Reposted by Valentina Pyatkin
Computational Linguistics @ UZH @cl-uzh.bsky.social · 17/04/2026
🗓️ SwissText 2026 keynote speakers announced & registration open! We are delighted to welcome Prof. Dr. Alexandra Birch and Dr. Valentina Pyatkin as our keynote speakers. 📋 Register here: ema.uzh.ch/RHK4W Early-bird rates available throughout April, with additional student discounts. #NLProc 1/3
ema.uzh.ch
SwissText 2026
10. Juni 2026 | UZH Campus Oerlikon
153
Reposted by Valentina Pyatkin
Yanai Elazar @yanai.bsky.social · 03/02/2026
Excited to have the Big Picture workshop back for another iteration this year at @aclmeeting.bsky.social Submit your big picture ideas, consolidation work, phd thesis distillation, etc. by March 5th! www.bigpictureworkshop.com w/ Allyson Ettinger, @norakassner.bsky.social, @sebruder.bsky.social
093
Reposted by Valentina Pyatkin
Yanai Elazar @yanai.bsky.social · 29/01/2026
🚨 New Study 🚨 @arxiv.bsky.social has recently decided to prohibit any 'position' paper from being submitted to its CS servers. Why? Because of the "AI slop", and allegedly higher ratios of LLM-generated content in review papers, compared to non-review papers.
2299
Reposted by Valentina Pyatkin
Front Conference Zurich @frontconference.bsky.social · 17/01/2026
Front Conference Zurich is coming up soon! On Friday, February 27, an amazing group of speakers will explore how AI is reshaping the way we work, from creativity and product design to engineering and collaboration 🤩 Our lineup: frontconference.com/schedule 🎟️ Your ticket: frontconference.com/tickets
Looking forward to learning new things in 2026?
We’ve got you covered with 17 amazing talks exploring how AI reshapes the way we work!

Get your conference pass
380.-
available until January 31
123
Reposted by Valentina Pyatkin
Ai2 @ai2.bsky.social · 02/12/2025
We're at #NeurIPS2025 with papers, posters, workshops, fireside chats, & talks across the conference. Come learn about our latest research + see live demos!
182
Reposted by Valentina Pyatkin
Simon Willison @simonwillison.net · 23/11/2025
Olmo 3 is notable as a "fully open" LLM - all of the training data is published, plus complete details on how the training process was run. I tried out the 32B thinking model and the 7B instruct models, + thoughts on why transparent training data is so important simonwillison.net/2025/Nov/22/...
simonwillison.net
Olmo 3 is a fully open LLM
Olmo is the LLM series from Ai2—the Allen institute for AI. Unlike most open weight models these are notable for including the full training data, training process and checkpoints along …
219134
Valentina Pyatkin @valentinapy.bsky.social · 20/11/2025
Olmo 3 is out! 🤩 I am particularly excited about Olmo 3 models' precise instruction following abilities and their good generalization performance on IFBench! Lucky to have been a part of the Olmo journey for three iterations already.
0243
Reposted by Valentina Pyatkin
ETH Zurich @ethz.ch · 31/10/2025
Happy Halloween!
0162
Reposted by Valentina Pyatkin
Paul Röttger @paul-rottger.bsky.social · 29/10/2025
There’s plenty of evidence for political bias in LLMs, but very few evals reflect realistic LLM use cases — which is where bias actually matters. IssueBench, our attempt to fix this, is accepted at TACL, and I will be at #EMNLP2025 next week to talk about it! New results 🧵
13111
Valentina Pyatkin @valentinapy.bsky.social · 27/10/2025
I will be giving a talk at @eth-ai-center.bsky.social next week, on RLVR for verifiable instruction following, generalization, and reasoning! 📢 Join if you are in Zurich and interested in hearing about IFBench and our latest Olmo and Tülu works at @ai2.bsky.social
060
Reposted by Valentina Pyatkin
Kanishka Misra @kanishka.bsky.social · 16/10/2025
"Although I hate leafy vegetables, I prefer daxes to blickets." Can you tell if daxes are leafy vegetables? LM's can't seem to! 📷 We investigate if LMs capture these inferences from connectives when they cannot rely on world knowledge. New paper w/ Daniel, Will, @jessyjli.bsky.social
Title page of the paper: WUGNECTIVES: Novel Entity Inferences of Language Models from Discourse Connectives, with two figures at the bottom

Left: Our figure 1 -- comparing previous work, which usually predicted the connective given the arguments (grounded in the world); our work flips this premise by getting models to use their knowledge of connectives to predict something about the world.

Right: Our main results across 7 types of connective senses. Models are especially bad at Concession connectives.
2329
Valentina Pyatkin @valentinapy.bsky.social · 10/10/2025
💡We kicked off the SoLaR workshop at #COLM2025 with a great opinion talk by @michelleding.bsky.social & Jo Gasior Kavishe (joint work with @victorojewale.bsky.social and @geomblog.bsky.social ) on "Testing LLMs in a sandbox isn't responsible. Focusing on community use and needs is."
1154
Reposted by Valentina Pyatkin
Michelle L. Ding @michelleding.bsky.social · 09/10/2025
Hi #COLM2025! 🇨🇦 I will be presenting a talk on the importance of community-driven LLM evaluations based on an opinion abstract I wrote with Jo Kavishe, @victorojewale.bsky.social and @geomblog.bsky.social tomorrow at 9:30am in 524b for solar-colm.github.io Hope to see you there!
solar-colm.github.io
Third Workshop on Socially Responsible Language Modelling Research (SoLaR) 2025
COLM 2025 in-person Workshop, October 10th at the Palais des Congrès in Montreal, Canada
196
Valentina Pyatkin @valentinapy.bsky.social · 20/09/2025
Now accepted to #neurips25 datasets & benchmarks! See you in San Diego! 🥳
080
Reposted by Valentina Pyatkin
Women in AI Research - WiAIR @wiair.bsky.social · 19/09/2025
🚀 Can open science beat closed AI? Tülu 3 makes a powerful case. In our new #WiAIRpodcast, we speak with Valentina Pyatkin (@valentinapy.bsky.social) of @ai2.bsky.social and the University of Washington about a fully open post-training recipe—models, data, code, evals, and infra. #WomenInAI 1/8🧵
151
Reposted by Valentina Pyatkin
Women in AI Research - WiAIR @wiair.bsky.social · 12/09/2025
"𝐋𝐋𝐌 𝐏𝐨𝐬𝐭-𝐭𝐫𝐚𝐢𝐧𝐢𝐧𝐠: 𝐎𝐩𝐞𝐧 𝐒𝐜𝐢𝐞𝐧𝐜𝐞 𝐓𝐡𝐚𝐭 𝐏𝐨𝐰𝐞𝐫𝐬 𝐏𝐫𝐨𝐠𝐫𝐞𝐬𝐬 " 🎙️ On Sept 17, the #WiAIRpodcast speaks with @valentinapy.bsky.social (@ai2.bsky.social & University of Washington) about open science, post-training, mentorship, and visibility #WiAIR #NLProc
061
Reposted by Valentina Pyatkin
Ai2 @ai2.bsky.social · 14/08/2025
With fresh support of $75M from NSF and $77M from NVIDIA, we’re set to scale our open model ecosystem, bolster the infrastructure behind it, and fast‑track reproducible AI research to unlock the next wave of scientific discovery. 💡
1457
Valentina Pyatkin @valentinapy.bsky.social · 10/08/2025
On my way to Oxford: Looking forward to speaking at OxML 2025
070
Valentina Pyatkin @valentinapy.bsky.social · 08/08/2025
🔈For the SoLaR workshop @COLM_conf we are soliciting opinion abstracts to encourage new perspectives and opinions on responsible language modeling, 1-2 of which will be selected to be presented at the workshop. Please use the google form below to submit your opinion abstract ⬇️
184
Reposted by Valentina Pyatkin
Yanai Elazar @yanai.bsky.social · 02/08/2025
I had a lot of fun contemplating about memorization questions at the @l2m2workshop.bsky.social panel yesterday together with Niloofar Mireshghallah and Reza Shokri, moderated by @pietrolesci.bsky.social who did a fantastic job! #ACL2025
1112
Reposted by Valentina Pyatkin
Akhila Yerukola @akhilayerukola.bsky.social · 26/07/2025
I'll be at #ACL2025🇦🇹!! Would love to chat about all things pragmatics 🧠, redefining "helpfulness"🤔 and enabling better cross-cultural capabilities 🗺️ 🫶 Presenting our work on culturally offensive nonverbal gestures 👇 🕛Wed @ Poster Session 4 📍Hall 4/5, 11:00-12:30
041
Valentina Pyatkin @valentinapy.bsky.social · 18/07/2025
🔥tokenization panel!
070
Valentina Pyatkin @valentinapy.bsky.social · 18/07/2025
why is vancouver sushi so good? 🤤 (vancouver food in general actually)
390
Reposted by Valentina Pyatkin
Ai2 @ai2.bsky.social · 14/07/2025
This week is #ICML in Vancouver, and a number of our researchers are participating. Here's the full list of Ai2's conference engagements—we look forward to connecting with fellow attendees. 👋
032
Valentina Pyatkin @valentinapy.bsky.social · 11/07/2025
I'll be at ICML in Vancouver next week! #ICML2025 You can find me at the following: - giving an invited talk at the "Models of Human Feedback for AI Alignment" workshop - giving an invited talk at the "AI for Math" workshop I'll also present these two papers ⤵️
192
Valentina Pyatkin @valentinapy.bsky.social · 06/07/2025
In Geneva🇨🇭to attend the International Open-Source LLM Builders Summit and present OLMo and Tülu!
0100
Valentina Pyatkin @valentinapy.bsky.social · 03/07/2025
💡Beyond math/code, instruction following with verifiable constraints is suitable to be learned with RLVR. But the set of constraints and verifier functions is limited and most models overfit on IFEval. We introduce IFBench to measure model generalization to unseen constraints.
1295
Reposted by Valentina Pyatkin
Nathan Lambert @natolambert.bsky.social · 03/07/2025
plus, some fun RL experiments
141
Reposted by Valentina Pyatkin
Nathan Lambert @natolambert.bsky.social · 03/07/2025
This new benchmark created by @valentinapy.bsky.social should be the new default replacing IFEval. Some of the best frontier models get <50% and it comes with separate training prompts so people don’t effectively train on test. Wild gap from o3 > Gemini 2.5 pro of like 30 points.
2173
Reposted by Valentina Pyatkin
Ai2 @ai2.bsky.social · 03/07/2025
Introducing IFBench, a benchmark to measure how well AI models follow new, challenging, and diverse verifiable instructions. Top models like Gemini 2.5 Pro or Claude 4 Sonnet are only able to score up to 50%, presenting an open frontier for post-training. 🧵
1171
Reposted by Valentina Pyatkin
Yanai Elazar @yanai.bsky.social · 01/07/2025
Check out our take on Chain-of-Thought. I really like this paper as a survey on the current literature on what CoT is, but more importantly on what it's not. It also serves as a cautionary tale to the (apparently quite common) misuse of CoT as an interpretable method.
1134
Reposted by Valentina Pyatkin
Sharon Levy @sharonlevy.bsky.social · 24/06/2025
🚨Submission deadline extended to June 27th AoE!🚨 Our reviewer interest form is also open! See below for more details👇
051
Valentina Pyatkin @valentinapy.bsky.social · 17/06/2025
Interested in shaping the progress of responsible AI and meeting leading researchers in the field? SoLaR@COLM 2025 is looking for paper submissions and reviewers! 🤖 ML track: algorithms, math, computation 📚 Socio-technical track: policy, ethics, human participant research
181
Reposted by Valentina Pyatkin
Cathy Wu @cathywu.bsky.social · 27/11/2024
I collected some folk knowledge for RL and stuck them in my lecture slides a couple weeks back: web.mit.edu/6.7920/www/l... See Appendix B... sorry, I know, appendix of a lecture slide deck is not the best for discovery. Suggestions very welcome.
web.mit.edu
311418
Reposted by Valentina Pyatkin
Yoav Goldberg @yoavgo.bsky.social · 08/06/2025
i created a gist with some non default llm courses gist.github.com/yoavg/95bbc5...
gist.github.com
llm-materials-2025.md
GitHub Gist: instantly share code, notes, and snippets.
3494
Reposted by Valentina Pyatkin
Saumya Malik @saumyamalik.bsky.social · 03/06/2025
I’m thrilled to share RewardBench 2 📊— We created a new multi-domain reward model evaluation that is substantially harder than RewardBench, we trained and released 70 reward models, and we gained insights about reward modeling benchmarks and downstream performance!
2226
Reposted by Valentina Pyatkin
Ai2 @ai2.bsky.social · 02/06/2025
RewardBench 2 is here! We took a long time to learn from our first reward model evaluation tool to make one that is substantially harder and more correlated with both downstream RLHF and inference-time scaling.
The RewardBench 2 Leaderboard on HuggingFace.
1208
Reposted by Valentina Pyatkin
Robert Hawkins @rdhawkins.bsky.social · 28/05/2025
Happy to announce the first workshop on Pragmatic Reasoning in Language Models — PragLM @ COLM 2025! 🎉 How do LLMs engage in pragmatic reasoning, and what core pragmatic capacities remain beyond their reach? 🌐 sites.google.com/berkeley.edu/praglm/ 📅 Submit by June 23rd
sites.google.com
PragLM @ COLM '25
IMPORTANT DATES
13918
Reposted by Valentina Pyatkin
Conference on Language Modeling @colmweb.org · 27/05/2025
Our discussion period just started. Authors, please read our instructions carefully. We require responses by June 2. But, what you really want to hear about is stats .... right? -> 🧵
2175
Reposted by Valentina Pyatkin
Nathan Lambert @natolambert.bsky.social · 13/05/2025
Is very classic that most people don't know the Tulu 3 paper coined the term RLVR
3122
Valentina Pyatkin @valentinapy.bsky.social · 12/05/2025
📢 The SoLaR workshop will be collocated with COLM! @colmweb.org SoLaR is a collaborative forum for researchers working on responsible development, deployment and use of language models. We welcome both technical and sociotechnical submissions, deadline July 5th!
1166
Reposted by Valentina Pyatkin
Andrew Lampinen @lampinen.bsky.social · 02/05/2025
How do language models generalize from information they learn in-context vs. via finetuning? In arxiv.org/abs/2505.00661 we show that in-context learning can generalize more flexibly, illustrating key differences in the inductive biases of these modes of learning — and ways to improve finetuning. 1/
arxiv.org
47722
Valentina Pyatkin @valentinapy.bsky.social · 02/05/2025
Accepted to #ICML2025! 🥳 See you in Vancouver this summer!
0123
Valentina Pyatkin @valentinapy.bsky.social · 01/05/2025
Accepted to #ICML2025 🤩
050
Reposted by Valentina Pyatkin
Melanie Mitchell @melaniemitchell.bsky.social · 01/05/2025
Very interesting oral history -- interviews with some top NLP folks on the effects of GenAI on their field: www.quantamagazine.org/when-chatgpt...
quantamagazine.org
When ChatGPT Broke an Entire Field: An Oral History | Quanta Magazine
Researchers in “natural language processing” tried to tame human language. Then came the transformer.
013346
Reposted by Valentina Pyatkin
Abhilasha Ravichander @lasha.bsky.social · 30/04/2025
✈️ I'm in New Mexico for #NAACL2025! Would love to meet and chat about training data, factuality, transparency, robustness, doing a PhD in AI🤖, or anything else. I love to make friends, please say hi if you see me! 🌵🌞 And check out our work at NAACL👇
1141
Reposted by Valentina Pyatkin
Nathan Lambert @natolambert.bsky.social · 29/04/2025
Heading to NAACL? With "verification being the key to AI" you should go to the poster session Friday, 9-10:30am to chat with my star colleagues @valentinapy.bsky.social + @jacobcares.bsky.social about RewardBench (and really RewardBench 2, evaluation, and reward models in post-training).
0142
Reposted by Valentina Pyatkin
François Fleuret @francois.fleuret.org · 28/04/2025
I asked "on the other platform" what were the most important improvements to the original 2017 transformer. That was quite popular and here is a synthesis of the responses:
420643
Reposted by Valentina Pyatkin
Jacob Morrison @jacobcares.bsky.social · 28/04/2025
Valentina and I will be presenting RewardBench at NAACL! Come say hi at the poster session on Friday and we can chat about reward models, staying up for 30 hours straight to rapidly reset from Singapore time, and more 🏜️
053
Valentina Pyatkin @valentinapy.bsky.social · 27/04/2025
I'll be at #NAACL2025: 🖇️To present my paper "Superlatives in Context", showing how the interpretation of superlatives is very context dependent and often implicit, and how LLMs handle such semantic underspecification 🖇️And we will present RewardBench on Friday Reach out if you want to chat!
1285