Sign in

Kate Sanders

@kesnet50.bsky.social
825 followers 396 following 23 posts

Researcher at Microsoft Copilot Tuning. Cal alum, Ph.D. @ JHU CLSP. #NLProc katesanders9.github.io

PostsRepliesMedia
Reposted by Kate Sanders
Marco @mcognetta.bsky.social · 30/06/2026
2397
Reposted by Kate Sanders
samuel mehr @mehr.nz · 15/06/2026
you're telling me a desk rejected this paper
1759
Reposted by Kate Sanders
MAGMaR Workshop @magmar-workshop.bsky.social · 17/02/2026
This year's shared task allows you to submit for the retrieval track, generation track, or full RAG track on a challenging new collection of unedited ("raw") videos. Research Papers (Apr. 1) Shared Task (Apr. 20)
011
Reposted by Kate Sanders
MAGMaR Workshop @magmar-workshop.bsky.social · 17/02/2026
📹 + 🧠 + 📝 = 🔥 First call for MAGMaR 2026, the 2nd workshop on multimodal augmented generation via multimodal retrieval! If #RAG isn't hard enough for you, try multilingually and multimodally. Collocated with @aclmeeting in San Diego in July. nlp.jhu.edu/magmar/
nlp.jhu.edu
MAGMaR Workshop
MAGMaR
112
Kate Sanders @kesnet50.bsky.social · 16/01/2026
I will be at AAAI 2026 in Singapore next week! ✈️ I'm looking forward to seeing everyone's cool projects and discussing reasoning, post-training, and multimodality. Please reach out if you will be there and would like to connect.
130
Reposted by Kate Sanders
Mayank Mehta @mayankmehta.bsky.social · 31/12/2025
Bye bye 2025, a divisive year, with many divisors: 3, 5, 9, 15, 25, 27, 45, 75, 81, 135, 225, 405, 675. Happy 2026 = 2*1013 Just two primes Cheers!
1216
Kate Sanders @kesnet50.bsky.social · 15/12/2025
Thinking about my favorite amp today 😔❤️
Eleven, eleven, eleven, eleven..
010
Reposted by Kate Sanders
Krithika Ramesh @stolenpyjak.bsky.social · 07/11/2025
🚀 SynthTextEval, our open-source toolkit for generating and evaluating synthetic text data for high-stakes domains, will be featured at EMNLP 2025 as a system demonstration! GitHub: github.com/kr-ramesh/sy... Paper 📝: aclanthology.org/2025.emnlp-d... #EMNLP2025 #EMNLP #SyntheticData
github.com
GitHub - kr-ramesh/synthtexteval: SynthTextEval: A Toolkit for Generating and Evaluating Synthetic Data Across Domains (EMNLP 2025 System Demonstration)
SynthTextEval: A Toolkit for Generating and Evaluating Synthetic Data Across Domains (EMNLP 2025 System Demonstration) - kr-ramesh/synthtexteval
1133
Reposted by Kate Sanders
Naomi Saphra @nsaphra.bsky.social · 16/10/2025
In honor of some new people coming from AI twitter, I finally updated my post to recommend For You over Discover.
2335
Reposted by Kate Sanders
Mark Riedl @markriedl.bsky.social · 11/10/2025
A company that believed it was in the verge of AGI or ASI wouldn’t capitulate to the government because it wouldn’t care about government contracts. They would soon BE the economy and the government would soon be capitulating to them.
1204
Reposted by Kate Sanders
Cameron Ellis @camerontellis.bsky.social · 09/10/2025
Astronaut meme: "Wait, it's all perception?" "Always has been"
1346
Reposted by Kate Sanders
Conference on Language Modeling @colmweb.org · 22/09/2025
Keynote spotlight #4: the second day of COLM will close with @ghadfield.bsky.social from JHU talking about human society alignment, and lessons for AI alignment
082
Reposted by Kate Sanders
Cornell Tech @cornelltech.bsky.social · 19/09/2025
Congratulations to Alane Suhr '22, a #CornellTech Ph.D. #alumni advised by associate professor Yoav Artzi, for receiving the prestigious 2022 @aaai.org / @acmsigai.bsky.social Doctoral Dissertation Award! Read more about the award here: aaai.org/about-aaai/a... @yoavartzi.com
aaai.org
AAAI/ACM SIGAI Doctoral Dissertation Award - AAAI
The AAAI/ACM SIGAI Doctoral Dissertation Award recognizes and encourages superior research and writing by doctoral candidates in AI.
082
Reposted by Kate Sanders
Conrad Hackett @conradhackett.bsky.social · 15/09/2025
Time for the world to install a gigawatt of solar power capacity 2004: A year 2010: ~ a month 2015: ~ a week Now: A day ourworldindata.org/data-insight... 🧪
Line chart showing that there's been a rapid escalation in how quickly the world installs a gigawatt of solar power capacity.
332434891
Reposted by Kate Sanders
terra firma, terra eterna 🚱🌉 @kavi.bsky.social · 31/08/2025
🚨 Urban Stats 28.0.0 🚨 The mapper is now completely redesigned by me and @spudwaffle.bsky.social, allowing for much prettier looking maps and way more customization alongside significantly more options for geographies! See below for some of the examples of the maps you can create!
2357
Reposted by Kate Sanders
Yoav Goldberg @yoavgo.bsky.social · 27/08/2025
When reading AI reasoning text (aka CoT), we (humans) form a narrative about the underlying computation process, which we take as a transparent explanation of model behavior. But what if our narratives are wrong? We measure that and find it usually is. Now on arXiv: arxiv.org/abs/2508.16599
arxiv.org
Humans Perceive Wrong Narratives from AI Reasoning Texts
A new generation of AI models generates step-by-step reasoning text before producing an answer. This text appears to offer a human-readable window into their computation process, and is increasingly r...
48623
Reposted by Kate Sanders
Daniel Khashabi @danielkhashabi.bsky.social · 26/08/2025
Paper: arxiv.org/pdf/2505.22037 🔗 Project page: aka.ms/jailbreak-d... 📊 Dataset: huggingface.co/datasets/ja...
huggingface.co
jackzhang/JBDistill-Bench · Datasets at Hugging Face
111
Reposted by Kate Sanders
Daniel Khashabi @danielkhashabi.bsky.social · 26/08/2025
So, what's the future of AI safety benchmarks? Jack's solution is "renewable benchmarks" that allows us to refresh and expand benchmarks with a single click!! x.com/jackjingyuz...
111
Reposted by Kate Sanders
Rachel Flood Heaton @rachelfloodheaton.bsky.social · 22/08/2025
In our forthcoming paper, John Hummel and I ask what it would mean for a neural computing architecture such as a brain to implement a symbol system, and the related question of what makes it difficult for them to do so, with an eye toward the differences between humans, animals, and ANNs.
arxiv.org
From Basic Affordances to Symbolic Thought: A Computational Phylogenesis of Biological Intelligence
What is it about human brains that allows us to reason symbolically whereas most other animals cannot? There is evidence that dynamic binding, the ability to combine neurons into groups on the fly, is...
13713
Reposted by Kate Sanders
Shahab Bakhtiari @shahabbakht.bsky.social · 03/08/2025
This paper is making the rounds: arxiv.org/abs/2506.21734 A tiny (27M) brain-inspired model trained just on 1000 samples outperforming o3-mini-high on reasoning tasks. #MLSky 🧠🤖
412825
Reposted by Kate Sanders
Parth Nobel @ptnobel.bsky.social · 16/07/2025
Interested in large-scale GPU optimization? Interested in how modern neural networks are being deployed to solve classical optimization problems? Writing a paper on these topics? Submit to the ScaleOPT workshop at NeurIPS! www.cvxgrp.org/scaleopt/#su...
cvxgrp.org
ScaleOPT
092
Reposted by Kate Sanders
Arya McCarthy @aryamccarthy.bsky.social · 27/07/2025
I'm recruiting MLEs @ #ACL2025! Reach out if you know folks interested in legal NLP, structured prediction, and full-time at a startup environment in NYC I'll also always chat about: • population-level inference on corpora • broad-coverage semantics • which café has the best Sachertorte in Vienna
043
Reposted by Kate Sanders
Jordan Boyd-Graber @boydgraber.bsky.social · 28/07/2025
My students and I are presenting three papers on Monday at #ACL2025 and this thread will recap them (including their videos).
172
Reposted by Kate Sanders
Tu Thanh Ha @tuthanhha.bsky.social · 26/07/2025
“Wikipedia is this economic anomaly. In many ways, it’s sort of magical that people will just volunteer without explicit economic incentives to create artifacts that are meant to share knowledge with everyone in the world”
843019489
Kate Sanders @kesnet50.bsky.social · 26/07/2025
Taking off for Vienna #ACL2025! 🇦🇹 Excited to talk with people about transparent reasoning, multimodality, and fact verification. Stop by our multimodal RAG workshop on Friday 🔥🔥🔥 Please reach out if you want to grab coffee!
020
Reposted by Kate Sanders
Maria Antoniak @mariaa.bsky.social · 17/07/2025
The #ACL2025 #ACL2025NLP feed is up and running! It matches both hashtags and any posts from or mentions of @aclmeeting.bsky.social Pin it to your home 📌 and enjoy! bsky.app/profile/did:...
24814
Reposted by Kate Sanders
terra firma, terra eterna 🚱🌉 @kavi.bsky.social · 25/07/2025
Juxtastat DAU update! Crazy how we've been >1000 every day for over a year now! Thank you all for all your support, and make sure to keep spreading the word!
0162
Reposted by Kate Sanders
ACL 2027 @aclmeeting.bsky.social · 22/07/2025
🥳 🎉 ❤️ The ACL 2025 Proceedings are live on the ACL Anthology 🥰 ! We’re thrilled to pre-celebrate the incredible research 📚 ✨ that will be presented starting Monday next week in Vienna 🇦🇹 ! Start exploring 👉 aclanthology.org/events/acl-2... #NLProc #ACL2025NLP #ACLAnthology
aclanthology.org
Annual Meeting of the Association for Computational Linguistics (2025) - ACL Anthology
pdf bibProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)Wanxiang Che | Joyce Nabende | Ekaterina Shutova | Mohammad Taher Pilehvar
05719
Reposted by Kate Sanders
ℳatt @matttomic.bsky.social · 10/07/2025
This New Yorker piece is the most hopeful I've felt about the world in a long time. I had no idea solar was booming like this. And if you live in the same world as me, dominated by oil & gas guys maintaining that solar and wind are inefficient gimmicks, you might not've known some of this either.
It took from the invention of the photovoltaic solar cell, in 1954, until 2022 for the world to install a terawatt of solar power; the second terawatt came just two years later, and the third will arrive either later this year or early next.
That’s because people are now putting up a gigawatt’s worth of solar panels, the rough equivalent of the power generated by one coal-fired plant, every fifteen hours. Solar power is now growing faster than any power source in history, and it is closely followed by wind power—which is really another form of energy from the sun, since it is differential heating of the earth that produces the wind that turns the turbines.
Last year, ninety-six per cent of the global demand for new electricity was met by renewables, and in the United States ninety-three per cent of new generating capacity came from solar, wind, and an ever-increasing variety of batteries to store that power.
14791403
Reposted by Kate Sanders
Niyati Bafna @niyatibafna.bsky.social · 04/07/2025
🔈When LLMs solve tasks with a mid-to-low resource input or target language, their output quality is poor. We know that. But can we put our finger on what breaks inside the LLM? We introduce the 💥 translation barrier hypothesis 💥 for failed multilingual generation with LLMs. arxiv.org/abs/2506.22724
2267
Reposted by Kate Sanders
Naomi Saphra @nsaphra.bsky.social · 26/04/2025
I wrote something up for AI people who want to get into bluesky and either couldn't assemble an exciting feed or gave up doomscrolling when their Following feed switched to talking politics 24/7.
nsaphra.net
The AI Researcher's Guide to a Non-Boring Bluesky Feed | Naomi Saphra
How to migrate to bsky without a boring feed.
2335994
Reposted by Kate Sanders
Leland McInnes @lelandmcinnes.bsky.social · 22/06/2025
Explore Wikipedia through a data map. Pages are grouped by semantic similarity, for topic clusters. Hover to see details, zoom to explore more fine-grained topics, click to go to a page. Search by page name to find interesting starting points for exploration. lmcinnes.github.io/datamapplot_...
711648
Reposted by Kate Sanders
arxiv cs.CL @arxiv-cs-cl.bsky.social · 17/06/2025
William Walden, Kathryn Ricci, Miriam Wanner, Zhengping Jiang, Chandler May, Rongkun Zhou, Benjamin Van Durme How Grounded is Wikipedia? A Study on Structured Evidential Support arxiv.org/abs/2506.12637
042
Reposted by Kate Sanders
Kaiser Sun @kaiserwholearns.bsky.social · 16/06/2025
What happens when an LLM is asked to use information that contradicts its knowledge? We explore knowledge conflict in a new preprint📑 TLDR: Performance drops, and this could affect the overall performance of LLMs in model-based evaluation.📑🧵⬇️ 1/8 #NLProc #LLM #AIResearch
arxiv.org
What Is Seen Cannot Be Unseen: The Disruptive Effect of Knowledge Conflict on Large Language Models
Large language models frequently rely on both contextual input and parametric knowledge to perform tasks. However, these sources can come into conflict, especially when retrieved documents contradict…
131
Reposted by Kate Sanders
JHU Computer Science @jhucompsci.bsky.social · 10/06/2025
Learn about the groundbreaking work being presented by JHU researchers at @cvprconference.bsky.social’s #CVPR2025! Check out the full list here www.cs.jhu.edu/news/johns-h... or browse through the thread below! 🧵 (1/14)
cs.jhu.edu
Johns Hopkins researchers to highlight work at CVPR 2025
Researchers from the Department of Computer Science will present groundbreaking work at the IEEE/CVF Computer Vision and Pattern Recognition Conference, the premier annual computer vision event.
121
Reposted by Kate Sanders
Niyati Bafna @niyatibafna.bsky.social · 07/06/2025
We know that speech LID systems flunk on accented speech. But why? And what can we do about it? 🤔 Our work arxiv.org/abs/2506.00628 (Interspeech '25) finds that *accent-language confusion* is an important culprit, ties it to the length of feature that the model relies on, and proposes a fix.
163
Kate Sanders @kesnet50.bsky.social · 12/05/2025
Excited to announce that I'm working on a project at AWS in New York this summer! Reach out if you're in the area and want to grab coffee 😀
image of nyc
040
Reposted by Kate Sanders
Dr. Damien P. Williams, dread portent down from a mountain cave @wolven.blacksky.app · 29/04/2025
@npr.org for every minute spent talking to a non-autistic person about autistic people's needs, you should be giving at least 2x that amount to actually autistic people— & not just in written stories, but on-air. Apply said rubric to any other group under attack right now for their lived realities.
312385529
Kate Sanders @kesnet50.bsky.social · 29/04/2025
We had a ton of fun last summer hacking away on this problem at the SCALE 2024 summer workshop. This summer, we're bringing it to Vienna as an ACL 2025 shared task!
010
Reposted by Kate Sanders
MAGMaR Workshop @magmar-workshop.bsky.social · 29/04/2025
🚨 IT'S HERE! 🚨 The Eval Leaderboard is now LIVE! 🏆💻 Our video retrieval collection stumps most pre-trained models. See if you can build a better system! eval.ai/web/challeng...
eval.ai
EvalAI: Evaluating state of the art in AI
EvalAI is an open-source web platform for organizing and participating in challenges to push the state of the art on AI tasks.
011
Reposted by Kate Sanders
Nathan Lambert @natolambert.bsky.social · 16/04/2025
First draft online version of The RLHF Book is DONE. Recently I've been creating the advanced discussion chapters on everything from Constitutional AI to evaluation and character training, but I also sneak in consistent improvements to the RL specific chapter. rlhfbook.com
212218
Reposted by Kate Sanders
Nicole Rust @nicolecrust.bsky.social · 15/04/2025
THIS IS ONE NEURON! Jaw dropping. It distributes a particular neurotransmitter (norepinephrine) across the mouse brain; it’s a locus coeruleus neuron. @jeremiahycohen.bsky.social and colleagues at the @alleninstitute.bsky.social are using new biotech to see things never seen before.
Jeremiah Cohen standing in front of a giant green neuron.
1032967
Reposted by Kate Sanders
Will Smith @willsmithvision.bsky.social · 01/04/2025
So, we wrote a neural net library entirely in LaTeX...
38415
Kate Sanders @kesnet50.bsky.social · 07/04/2025
🚨 New preprint on transparent, tree-adaptive grounded reasoning! We introduce Bonsai, a versatile reasoning system that generates interpretable, grounded, and uncertainty-aware reasoning traces while enabling state-of-the-art performance on text and video benchmarks.
230
Reposted by Kate Sanders
MAGMaR Workshop @magmar-workshop.bsky.social · 04/04/2025
🚨 Deadline extension! 🚨 The MAGMaR paper submission date has been extended to May 1, 2025 (archival and non-archival). Show off your multimodal + RAG projects!
101
Reposted by Kate Sanders
Benno Krojer @bennokrojer.bsky.social · 04/04/2025
Really enjoyed diving deep into this paper today: arxiv.org/abs/2406.16860 It's so systematic and a treasure trove full of insights
arxiv.org
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
We introduce Cambrian-1, a family of multimodal LLMs (MLLMs) designed with a vision-centric approach. While stronger language models can enhance multimodal capabilities, the design choices for vision ...
041
Reposted by Kate Sanders
Sung Kim @sungkim.bsky.social · 04/04/2025
ByteDance's Recitation over Reasoning The cutting-edge LLMs unanimously exhibits extremely severe recitation behavior; by changing one phrase in the condition, top models such as OpenAI-o1 and DeepSeek-R1 can suffer 60% performance loss on elementary school-level arithmetic and reasoning problems.
15311
Reposted by Kate Sanders
Siva Reddy @sivareddyg.bsky.social · 01/04/2025
Introducing the DeepSeek-R1 Thoughtology -- the most comprehensive study of R1 reasoning chains/thoughts ✨. Probably everything you need to know about R1 thoughts. If we missed something, please let us know.
0164
Reposted by Kate Sanders
Mor Geva @megamor2.bsky.social · 31/03/2025
🎉 Our Actionable Interpretability workshop has been accepted to #ICML2025! 🎉 > Follow @actinterp.bsky.social > Website actionable-interpretability.github.io @talhaklay.bsky.social @anja.re @mariusmosbach.bsky.social @sarah-nlp.bsky.social @iftenney.bsky.social Paper submission deadline: May 9th!
34116
Reposted by Kate Sanders
Abhilasha Ravichander @lasha.bsky.social · 21/03/2025
Want to know what training data has been memorized by models like GPT-4? We propose information-guided probes, a method to uncover memorization evidence in *completely black-box* models, without requiring access to 🙅‍♀️ Model weights 🙅‍♀️ Training data 🙅‍♀️ Token probabilities 🧵 (1/5)
arxiv.org
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
High-quality training data has proven crucial for developing performant large language models (LLMs). However, commercial LLM providers disclose few, if any, details about the data used for training. ...
49627