Sign in

Sathvik

@sathvik.bsky.social
287 followers 280 following 63 posts

computational psycholinguistics @ umd he/him

PostsRepliesMedia
Reposted by Sathvik
Mal Shah @compositiomality.bsky.social · 30/09/2026
Accepted to Linguistic Inquiry! I argue on syntactic grounds “some of the apples” is really “some apples of the apples” with NP-ellipsis. This explains which quantifiers and which modifiers (adjectives, prepositional phrases, relative clauses) can’t appear. Preprint: ling.auf.net/lingbuzz/009...
ling.auf.net
44311
Reposted by Sathvik
Aaron Steven White @aaronstevenwhite.io · 03/09/2026
Really excited about this project. You can now access effectively all of UniMorph, Universal Dependencies, Universal Decompositional Semantics, MegaAttitude, and various other linguistic resources in a unified format. And there's a lot more coming down the pike in the near future.
2111
Reposted by Sathvik
Olivia Guest · Ολίβια Γκεστ @olivia.science · 08/08/2026
They keep trying to make "correlation is cognition" happen. Through asserting: 1️⃣ models with only correlationist powers (data and correlations over that data) work well & 2️⃣ our own cognitive powers are no different to these machines, which can only rehash the past. @andreaeyleen.eurosky.social 1/
1322568
Reposted by Sathvik
Iris van Rooij 💭 @irisvanrooij.bsky.social · 05/05/2026
Now reading 📖 Cruz, Nicole (2026). Illusions of Understanding from Outsourcing Thinking to LLMs. Computational Brain & Behavior doi.org/10.1007/s421... 1/🧵
doi.org
Illusions of Understanding from Outsourcing Thinking to LLMs - Computational Brain & Behavior
Computational Brain & Behavior - Some illusions of understanding are an inevitable part of the research process, while others can be avoided or overcome by careful critical thinking and...
212347
Reposted by Sathvik
Timnit Gebru @timnitgebru.blacksky.app · 26/07/2026
Now is a good reminder that machine learning is a rebranding of statistical learning. Reinforcement learning is something I first came across in control theory. Many ML formulations are from information theory, a much more theoretically grounded discipline that's been around for decades.
8737209
Reposted by Sathvik
CJ @virmalised.us · 23/07/2026
😈
3322
Reposted by Sathvik
CJ @virmalised.us · 19/07/2026
You should not have to rely on commercial, closed-sourced, closed-data systems that have been used to hobble more transparent algorithms in order to obtain life-saving information from the web bsky.app/profile/virm...
1113
Reposted by Sathvik
Dr Abeba Birhane @abeba.blacksky.app · 16/12/2025
I wrote this brief talk on why “augmenting diversity” with LLMs is empirically unsubstantiable, conceptually flawed, and epistemically harmful and a nice surprise to see the organisers have made it public synthetic-data-workshop.github.io/papers/13.pdf
title: Cheap science, real harm: the cost of replacing human
participation with synthetic data

author: Abeba Birhane

abstract: Driven by the goals of augmenting diversity, increasing speed, reducing cost, the
use of synthetic data as a replacement for human participants is gaining traction
in AI research and product development. This talk critically examines the claim
that synthetic data can “augment diversity,” arguing that this notion is empirically
unsubstantiated, conceptually flawed, and epistemically harmful. While speed and
cost-efficiency may be achievable, they often come at the expense of rigour, insight,
and robust science. Drawing on research from dataset audits, model evaluations,
Black feminist scholarship, and complexity science, I argue that replacing human
participants with synthetic data risks producing both real-world and epistemic
harms at worst and superficial knowledge and cheap science at best
20849267
Reposted by Sathvik
Cory Shain @coryshain.bsky.social · 13/07/2026
Word predictability effects are LINEAR! And logarithmic. In a new preprint led by lab alum Stephanie Cho (w/ Ryan Buggy & Adrian Staub), we find clear evidence that the two patterns coexist. 1/
Title and abstract for "Preactivation and probabilistic inference coexist during sentence comprehension"
182
Reposted by Sathvik
Olivia Guest · Ολίβια Γκεστ @olivia.science · 02/12/2025
New preprint! @marentierra.bsky.social @irisvanrooij.bsky.social & I have been working on what CAIL means to showcase & propagate the idea of thinking very differently to tech industry norms on "artificial intelligence" Towards Critical Artificial Intelligence Literacies doi.org/10.5281/zeno... 1/
Figure 1: The important dimensions of CAILs across research and education; clockwise from 12 o’clock: Con-
ceptual Clarity is the idea that terms should refer. Critical Thinking is deep engagement with the relationships
between statements about the world. Decoloniality is the process of de-centring and addressing dominant
harmful views and practices. Respecting Expertise is the epistemic compact between professionals and society.
Slow Science is a disposition towards preferring psychologically, techno-socially, and epistemically healthy
practices. The lines between dimensions represent how they are interwoven both directly and indirectly.
13380145
Sathvik @sathvik.bsky.social · 03/07/2026
LM surprisal outperforms surprisal over human cloze responses as a predictor of reading times. Why is this the case? In our ACL paper, we standardize how to treat cloze surprisal, manipulate LM probabilities to reflect biases in the cloze task, and discuss how LMs and cloze data can be combined.
170
Sathvik @sathvik.bsky.social · 22/05/2026
What (if anything) can LLMs tell us about human language processing? I discuss this question and how psycholinguistics is moving forward with @colinphillips.bsky.social , to appear in BBS as a commentary on Futrell & Mahowald's article about LLMs & linguistics. arxiv.org/abs/2604.09466
Title & Abstract:
Across the Levels of Analysis: Explaining Predictive Processing in Humans Requires More Than Machine-Estimated Probabilities
Sathvik Nair & Colin Phillips
Commentary on Futrell, R. & Mahowald, K. (in press). How Linguistics Learned to Stop Worrying and Love the Language Models. Behavioral and Brain Sciences. http://doi.org/10.1017/S0140525X2510112X
Abstract
Under the lens of Marr’s levels of analysis, we critique and extend the authors’ two points about language models (LMs) and language processing: first, predicting upcoming linguistic information based on context is key to language processing, and second, that many advances in psycholinguistics would be impossible without LLMs. We also outline directions combining LLMs’ strengths with psycholinguistic models.
0173
Sathvik @sathvik.bsky.social · 26/03/2026
I'll be presenting some work comparing how humans and LMs make predictions at #HSP2026 this week, please reach out if you'd like to meet! Tomorrow, I'll have a poster on work with @byungdoh.bsky.social comparing cloze & LM surprisals as predictors of reading time hsp2026.org/abstracts/su...
hsp2026.org
192
Reposted by Sathvik
Iris van Rooij 💭 @irisvanrooij.bsky.social · 06/01/2026
✨ Updated preprint ✨ Iris van Rooij & Olivia Guest (2026). Combining Psychology with Artificial Intelligence: What Could Possibly Go Wrong? PsyArXiv osf.io/preprints/psyarxiv/aue4m_v2 @olivia.science Our aim is to make these ideas accessible for a.o. psych students. Hope we succeeded 🙂
Figure 1
Illustration of why AI systems cannot realistically scale to human cognition within the foreseeable future: (b) Human cognitive capacities (such as reasoning, communication, problem solving, learning, concept formation, planning etc.) can handle unbounded situations across many domains, ranging from simple to complex. (a) Engineers create AI systems using machine learning from human data. (d) In an attempt to approximate human cognition a lot of data is consumed. (c) Making AI systems that approximate human cognition is intractable (van Rooij, Guest, et al., 2024), i.e., the required resources (e.g. time, data) grows prohibitively fast as input domains get more complex, leading to diminishing returns. (a) Any existing AI system is
created in limited time (hours, months or years, not millennia or eons). Therefore, existing AI systems cannot realistically have the domain-general cognitive capacities that humans have. [Made with elements from freepik.com.]
720585
Sathvik @sathvik.bsky.social · 21/12/2025
bigram of the year: "surprisal brainrot"
040
Reposted by Sathvik
UMD Language Science Center @umd-lsc.bsky.social · 22/10/2025
We captured so many great moments from Language Science Day, thanks to Andrea Zukowski! We wish we could share them all here, but you can see the full gallery on our Flickr page. Click here to check them out: flickr.com/photos/umd-l...
Yi Ting Huang shares remarks with a large crowd of Language Science Day attendeesShevaun Lewis presents to a room of people in front of a projector screen that reads "More Than One Brain: Studying Conversation"Language Science Day panelist sit at the front of the room sharing their expertiseGraduate students and faculty share their poster presentations at the Language Science Day poster session
051
Sathvik @sathvik.bsky.social · 01/08/2025
If any friends are at Cog Sci, I’ll be in SF tomorrow! Let me know if you’d like to meet!
010
Reposted by Sathvik
Sam Gershman @gershbrain.bsky.social · 09/07/2025
The sycophantic tone of ChatGPT always sounded familiar, and then I recognized where I'd heard it before: author response letters to reviewer comments. "You're exactly right, that's a great point!" "Thank you so much for this insight!" Also how it always agrees even when it contradicts itself.
518722
Reposted by Sathvik
Lindia Tjuatja @lindiatjuatja.bsky.social · 09/06/2025
When it comes to text prediction, where does one LM outperform another? If you've ever worked on LM evals, you know this question is a lot more complex than it seems. In our new #acl2025 paper, we developed a method to find fine-grained differences between LMs: 🧵1/9
27020
Reposted by Sathvik
Iris van Rooij 💭 @irisvanrooij.bsky.social · 14/05/2025
NEW paper! 💭🖥️ “Combining Psychology with Artificial Intelligence: What could possibly go wrong?” — Brief review paper by @olivia.science & myself, highlighting traps to avoid when combining Psych with AI, and why this is so important. Check out our proposed way forward! 🌟💡 osf.io/preprints/ps...
Table 1
Typology of traps, how they can be avoided, and what goes wrong if not avoided. Note that all traps in a sense constitute category errors (Ryle & Tanney, 2009) and the success-to-truth inference (Guest & Martin, 2023) is an important driver in most, if not all, of the traps.
15346104
Reposted by Sathvik
Aniello De Santo @anids.bsky.social · 03/05/2025
A bit late but since I really like this paper, a bit of self-advertising! I am presenting at CMCL today work showing that metrics measuring how a Minimalist Grammar parser modulates memory usage can help us model Self-paced reading data for SRC/ORC contrasts: aclanthology.org/2025.cmcl-1.5/
aclanthology.org
Capturing Online SRC/ORC Effort with Memory Measures from a Minimalist Parser
Aniello De Santo. Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics. 2025.
4286
Reposted by Sathvik
Ben Lipkin @benlipkin.bsky.social · 10/04/2025
New preprint on controlled generation from LMs! I'll be presenting at NENLP tomorrow 12:50-2:00pm Longer thread coming soon :)
1209
Reposted by Sathvik
Kanishka Misra @kanishka.bsky.social · 02/04/2025
another day another minicons update (potentially a significant one for psycholinguists?) "Word" scoring is now a thing! You just have to supply your own splitting function! pip install -U minicons for merriment
from minicons import scorer
from nltk.tokenize import TweetTokenizer

lm = scorer.IncrementalLMScorer("gpt2")

# your own tokenizer function that returns a list of words
# given some sentence input
word_tokenizer = TweetTokenizer().tokenize

# word scoring
lm.word_score_tokenized(
    ["I was a matron in France", "I was a mat in France"], 
    bos_token=True, # needed for GPT-2/Pythia and NOT needed for others
    tokenize_function=word_tokenizer,
    bow_correction=True, # Oh and Schuler correction
    surprisal=True,
    base_two=True
)

'''
First word = -log_2 P(word | <beginning of text>)

[[('I', 6.1522440910339355),
  ('was', 4.033324718475342),
  ('a', 4.879510402679443),
  ('matron', 17.611848831176758),
  ('in', 2.5804288387298584),
  ('France', 9.036953926086426)],
 [('I', 6.1522440910339355),
  ('was', 4.033324718475342),
  ('a', 4.879510402679443),
  ('mat', 19.385351181030273),
  ('in', 6.76780366897583),
  ('France', 10.574726104736328)]]
'''
3217
Sathvik @sathvik.bsky.social · 25/03/2025
I’ll also be presenting a talk based on this work Friday afternoon at HSP. Very excited to share it with a psycholinguistics-focused audience!
071
Sathvik @sathvik.bsky.social · 25/03/2025
I’ll be at #HSP2025! I’m presenting a poster in session 4 on how semantic factors might affect timing data from a speeded cloze task (w @virmalised.us, Philip Resnik, and @colinphillips.bsky.social) hsp2025.github.io/abstracts/19...
hsp2025.github.io
0103
Reposted by Sathvik
Sebastián @smancha.bsky.social · 22/03/2025
I’ll be presenting a poster at HSP 2025 in about a week. It’s on memory for pronominal clitic placement in Spanish, come stop by and say hi if you can!
032
Reposted by Sathvik
Iris van Rooij 💭 @irisvanrooij.bsky.social · 11/01/2025
🎬🎥🍿 Video of my keynote at MathPsych2024 now available online www.youtube.com/watch?v=WrwN... #CogSci #CriticalAI #AIhype #AGI #PsychSci #PhilSci 🧪
youtube.com
Iris van Rooij keynote at MathPsych/ICCM 2024
YouTube video by Society for Mathematical Psychology
411134
Reposted by Sathvik
Bertram Højer @brtrm.bsky.social · 04/12/2024
What do YOU mean by "intelligence", and does ChatGPT fit your definition? We collected the major criteria used in CogSci and other fields, and designed a survey to find out! Access link: www.survey-xact.dk/collect Code: 4S7V-SN4M-S536 Time: 5-10 mins
bertramhojer.github.io
Perspectives on Intelligence: Community Survey
Research survey exploring how NLP/ML/CogSci researchers define and use the concept of intelligence.
23213
Reposted by Sathvik
Hal Daumé III @haldaume3.bsky.social · 10/12/2024
starter pack for the Computational Linguistics and Information Processing group at the University of Maryland - get all your NLP and data science here! go.bsky.app/V9qWjEi
12912
Reposted by Sathvik
Leonie Weissweiler @weissweiler.bsky.social · 26/11/2024
@kanishka.bsky.social and I have made a starter pack for researchers working broadly on linguistic interpretability and LLMs! go.bsky.app/F9qzAUn Please message me or comment on this post if you've noticed someone who we forgot or would like to be added yourself!
10379
Reposted by Sathvik
Thenmozhi Soundararajan/Dalit Diva ✨️is querying✨️ @dalitdiva.bsky.social · 23/11/2024
"Hey everyone! 👋 I’ve created a starter pack of South Asian artists, authors, academics, activists, and orgs. I’ll keep it updated—DM me or reply if you or someone you know should be added! ✨" go.bsky.app/GGd6dxU
4015161
Sathvik @sathvik.bsky.social · 11/11/2024
I’ll be presenting two posters on (psycho)linguistically motivated perspectives on LM generalization at #EMNLP2024! 1. Sensitivity to Argument Roles - Session 2 & #BlackBoxNLP 2. Learning & Filler-Gap Dependencies - #CoNLL Excited to chat with other folks interested in compling x cogsci! papers⬇️
171
Sathvik @sathvik.bsky.social · 06/06/2024
5-gram of the day: "language models from computational linguistics"
000
Sathvik @sathvik.bsky.social · 16/04/2024
Today I learned that I may not have a successful psycholinguistics career because I got a B in databases.
220
Sathvik @sathvik.bsky.social · 30/03/2024
Panicked after seeing AGI on my tax form
020
Reposted by Sathvik
Nnedi Okorafor, PhD @nnedi.bsky.social · 04/03/2024
This sums it up perfectly. It’s not a conversation.
“There is no ethical way to use the major AI image generators. All of them are trained on stolen images, and all of them are built for the purpose of deskilling, disempowering and replacing real human artists.”
158173067432
Sathvik @sathvik.bsky.social · 22/02/2024
yelled about lexicalism in my NLU seminar do i get a prize
010
Reposted by Sathvik
Vicki @vickiboykis.com · 31/12/2023
LLMs are so weird because one side is people with five PhDs who have been studying neuron activations for the past three decades and on the other side is someone called leetm5n with an anime avatar just casually releasing increasingly better performing fine tunes of mistral
2385
Reposted by Sathvik
Laura Gwilliams @lauragwilliams.bsky.social · 13/12/2023
We’re excited about our first paper looking at speech encoding in single neurons across the depth of human cortex. Out today in @nature! www.nature.com/articles/s41... [1/6]
nature.com
Large-scale single-neuron speech sound encoding across the depth of human cortex - Nature
High-density single-neuron recordings show diverse tuning for acoustic and phonetic features across layers in human auditory speech cortex.
99649
Sathvik @sathvik.bsky.social · 02/11/2023
Honored my paper was accepted to Findings of #EMNLP2023! Many psycholinguistics studies use LLMs to estimate the probability of words in context. But LLMs process statistically derived subword tokens, while human processing doesn't. Does this matter? (w/Philip Resnik) 🧵 arxiv.org/abs/2310.17774
1224
Reposted by Sathvik
The Spark Society @sparksociety.bsky.social · 29/10/2023
It's great to see the SPARK Society growing on this platform. Many are interested in supporting the principles upheld by SPARK but are not members. Consider membership. Membership is for ALL SCIENTISTS who are allied in supporting scholars from diverse backgrounds in Cognitive Psychology.
03228
Reposted by Sathvik
Stephan Meylan @stephan-meylan.bsky.social · 26/10/2023
How do adults understand children’s early, highly variable speech? Our new paper in Nature Human Behavior (www.nature.com/articles/s41...) provides evidence that adults’ interpretations depend quite strongly on language expectations—what they think children are likely to say. 1/
Figure showing the difference in performance between pour  best model with rich expectations (90% agreement with annotators) and a baseline model that only uses phonetics (42%). We label the difference between the model results as the contribution of language expectations.
57130
Reposted by Sathvik
Jennifer Hu @jennhu.bsky.social · 24/10/2023
To researchers doing LLM evaluation: prompting is *not a substitute* for direct probability measurements. Check out the camera-ready version of our work, to appear at EMNLP 2023! (w/ @rplevy.bsky.social) Paper: arxiv.org/abs/2305.13264 Original thread: twitter.com/_jennhu/stat...
1729166
Reposted by Sathvik
Xinchi Yu @xinchiyu.bsky.social · 22/10/2023
Do check out these fun posters from our University of Maryland Linguistics “delegation” to SNL! :) #SNL2023 linguistics.umd.edu/news/marylan...
1103
Reposted by Sathvik
folukeifejola @folukeifejola.bsky.social · 12/10/2023
This form was created to facilitate the sharing of invitation codes for "Global South" scholars who may not have the same network privileges as scholars in the "Global North". Scholars who are geographically located in the "Global South" will be given priority. docs.google.com/forms/d/e/1F...
docs.google.com
BlueSky Invites for "Global South" scholars form
This form was created to facilitate the sharing of invitation codes for Blue Sky which do not seem to be reaching "Global South" scholars who may not have the same network privileges as scholars in th...
49487463
Sathvik @sathvik.bsky.social · 11/10/2023
this grant writing process has made me realize i focus on questions that ask "to what extent" way too much
150
Reposted by Sathvik
David Bamman @dbamman.bsky.social · 09/10/2023
This might be flying under the radar (so please RT!), but the US Copyright Office is soliciting comments for its decisions on training ML/NLP/AI systems on copyrighted material (even *non*-generative AI). They need to hear from researchers, so please comment! Deadline Oct 30. www.copyright.gov/ai/
copyright.gov
Copyright and Artificial Intelligence | U.S. Copyright Office
Copyright law
2271392
Reposted by Sathvik
Roger Levy @rplevy.bsky.social · 05/10/2023
Yi Ting Huang and I have an opening for a postdoctoral researcher on an NSF-funded project, "Syntactic processing across socioeconomic status: Linking input to comprehension". Apply by Nov 15; start Jul 1, 2024 with flexibility. Please disseminate widely! ejobs.umd.edu/postings/114...
ejobs.umd.edu
Post Doctoral Associate
Dr. Yi Ting Huang of the University of Maryland (UMD) and Dr. Roger Levy of the Massachusetts Institute of Technology (MIT) are seeking to hire a post-doctoral researcher to work on a collaborative in...
12921
Sathvik @sathvik.bsky.social · 04/10/2023
i realized that my GRFP proposal is baaaasically my idea of what my dissertation will look like??? we're not ready for this
020
Reposted by Sathvik
Timo B. Roettger @timoroettger.bsky.social · 22/09/2023
Mouse tracking for reading? 🤔 "We show that MoTR data correlate well with previously-collected eye tracking data" psyarxiv.com/4ryvs
1115