Sign in

Rabiraj Banerjee

@rabirajb.bsky.social
96 followers 273 following 16 posts

PhDing on Interpretable NLP + CSS @gesis.org Prev: Masters Student + Researcher at @ubuffalo.bsky.social and Sr. Data Scientist at Coursera

PostsRepliesMedia
Reposted by Rabiraj Banerjee
Stella Biderman @stellaathena.bsky.social · 03/10/2026
I’m starting a blog! My first post is on how 3rd party embedded evaluators seem totally unsuited to addressing the problems we are currently facing, and what the real problem is. stellabiderman.ai/blog/embedde...
stellabiderman.ai
Embedded Evaluators Can’t Fix Companies That Choose to Be Bad — Stella Biderman
Embedded evaluators can report violations, but they cannot fix AI companies that knowingly disregard basic cybersecurity and safety practices.
37817
Reposted by Rabiraj Banerjee
Gabriella Lapesa @gabriellalapesa.bsky.social · 25/06/2026
Postdoc position at CSS@GESIS and CAIS, on a project on disinformation led by great colleagues gesis.jobs.personio.de/job/2679794?...
gesis.jobs.personio.de
Postdoctoral Researcher in Computational Social Science: Research Infrastructures and Synthetic Social Media Data (CSS-119) | Jobs at GESIS – Leibniz-Institut für Sozialwissenschaften
GESIS is one of the world's leading social science infrastructure facilities, supporting researchers at all levels of their research projects with expertise and infrastructure services. We help ensure...
013
Reposted by Rabiraj Banerjee
Johannes Breuer @johannesbreuer.com · 25/06/2026
For our new project "SynDIKAT: Synthetic disinformation data, collaboration infrastructure, and an analysis toolbox for social media research" @dede1989.bsky.social and I are looking for a postdoc for a shared position between @gesis.org and @cais-research.bsky.social: rrr.is/postdocsyndi...
gesis.jobs.personio.de
Postdoc in Computational Social Science: Research Infrastructures und Synthetic Social Media Data (CSS-119) | Jobs bei GESIS – Leibniz-Institut für Sozialwissenschaften
GESIS ist eine der weltweit führenden Infrastruktureinrichtungen für die Sozialwissenschaften und steht Forscher*innen mit Expertise und Infrastrukturangeboten auf allen Ebenen ihrer Forschungsprojekt...
21113
Reposted by Rabiraj Banerjee
David Bau @davidbau.bsky.social · 20/04/2026
2026 is a whirlwind year for AI. Underlying it all is the greatest scientific mystery of our age. How does a neural network think? I talked w Oliver Whang in NYTimes Magazine, on how AI interpretability is a tangle of structure waiting to be unraveled: www.nytimes.com/2026/04/15/...
2132
Reposted by Rabiraj Banerjee
icwsm.bsky.social @icwsm.bsky.social · 15/04/2026
🚨 Big news! #ICWSM '27 is heading to Edinburgh, Scotland 🏰 More details on dates and venue coming soon ✨ ❗How submissions work: • May 15, 2026: accept (→ ICWSM '27), R&R → Sept '26 • Sept 15, 2026: accept (→ ICWSM '27), R&R → Jan '27 • Jan 15, 2027: accept (→ ICWSM '27), R&R → May '27 (→ ICWSM '28)
0189
Reposted by Rabiraj Banerjee
Kenny Joseph @kennyjoseph.bsky.social · 20/03/2026
I have been thinking for the last few years about ideologies and how they emerge in text. This paper, with @davidlazer.bsky.social and Kim Williams, reflects some of those thoughts, and how I think we can improve and expand on how we operationalize ideology in discourse. arxiv.org/abs/2603.18945
arxiv.org
A conceptual framework for ideology beyond the left and right
NLP+CSS work has operationalized ideology almost exclusively on a left/right partisan axis. This approach obscures the fact that people hold interpretations of many different complex and more specific...
0206
Reposted by Rabiraj Banerjee
ACL Anthology @aclanthology.org · 20/03/2026
Do you need a weekend read? The proceedings of EACL 2026 and co-located workshops are now online! @eaclmeeting.bsky.social aclanthology.org/events/eacl-...
aclanthology.org
19th Conference of the European Chapter of the Association for Computational Linguistics - ACL Anthology
053
Reposted by Rabiraj Banerjee
Maria Antoniak @mariaa.bsky.social · 10/03/2026
I'm lecturing about the "History of NLP" this week. What should I include? Any favorite anecdotes, images, people, methods? Slides, books, papers, or talks for inspiration or grounding? I've been maintaining a small collection here: www.are.na/maria-antoni...
are.na
🗄 history of NLP and the ACL | Are.na
267715
Reposted by Rabiraj Banerjee
Shubhendu Trivedi @shubhendu.bsky.social · 09/03/2026
The new conformal prediction book now seems to be final after a bunch of updates: arxiv.org/abs/2411.118...
arxiv.org
Theoretical Foundations of Conformal Prediction
This book is about conformal prediction and related inferential techniques that build on permutation tests and exchangeability. These techniques are useful in a diverse array of tasks, including hypot...
1113
Reposted by Rabiraj Banerjee
Naomi Saphra @nsaphra.bsky.social · 24/02/2026
why do science? it won,t make the model Bigger
4474
Reposted by Rabiraj Banerjee
Naomi Saphra @nsaphra.bsky.social · 23/02/2026
This has a very cool result on in-context learned classification tasks, where they disentangle representational quality (how well-separated concept labels are) and readout alignment (how good it is at reading out its own inner labels). Adding demo examples helps through readout, not representations!
arxiv.org
The Geometry of Prompting: Unveiling Distinct Mechanisms of Task Adaptation in Language Models
Decoder-only language models have the ability to dynamically switch between various computational tasks based on input prompts. Despite many successful applications of prompting, there is very limited...
1365
Reposted by Rabiraj Banerjee
Nikhil Garg @nkgarg.bsky.social · 23/02/2026
yeah! In this paper in ICML25, we found both directions of this -- in LLM as judge, using a bigger/more accurate model inflates accuracies bc of correlated errors, and using a worse model deflates them for the reason in the above paper arxiv.org/abs/2506.07962
Screenshot from the paper with a figure showing 15 scatterplots in a grid., evaluating LLM-as-judge on HELM. In each plot, one model is used as the judge. Each dot is another model; the y-axis is the
accuracy inflation (compared to ground truth) of using the given model as the judge, and the x-axis is the model’s true accuracy. A
vertical red line corresponds to the true accuracy of the judge. Each judge tends to inflate the accuracy of models that are less accurate
than itself, especially models from the same provider or family
0103
Rabiraj Banerjee @rabirajb.bsky.social · 20/02/2026
we need to formulate a new name for such people 😤😤😤
000
Reposted by Rabiraj Banerjee
Andrew Saxe @saxelab.bsky.social · 03/02/2026
Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures? Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham arxiv.org/abs/2512.20607
716042
Rabiraj Banerjee @rabirajb.bsky.social · 02/02/2026
God of War Ragnarok , Black Myth Wukong(very hard), Witcher 3
010
Reposted by Rabiraj Banerjee
Negar Foroutan @negarforoutan.bsky.social · 15/12/2025
1/ 🌍 How does mixing data from hundreds of languages affect LLM training? In our new paper "Revisiting Multilingual Data Mixtures in Language Model Pretraining" we revisit core assumptions about multilinguality using 1.1B-3B models trained on up to 400 languages. 🧵👇
196
Reposted by Rabiraj Banerjee
Naomi Saphra @nsaphra.bsky.social · 12/12/2025
So @lchoshen.bsky.social posted a thread on X about how different training runs tend to converge, and I just had to argue with him. Training variation is fascinating, and I think we've kinda cracked it!
x.com
2173
Reposted by Rabiraj Banerjee
Kyle Lo @kylelo.bsky.social · 20/11/2025
we released Olmo 3! lot of exciting stuff but wanna focus on: 🐟Olmo 3 32B Base, the best fully-open base model to-date, near Qwen 2.5 & Gemma 3 on diverse evals 🐠Olmo 3 32B Think, first fully-open reasoning model approaching Qwen 3 levels 🐡12 training datasets corresp to different staged training
1417
Rabiraj Banerjee @rabirajb.bsky.social · 17/11/2025
Induction through Compression I personally loved the relationship between ICL and Komogorov Complexity that this paper proposed arxiv.org/pdf/2410.14086
arxiv.org
020
Reposted by Rabiraj Banerjee
taka-yamakoshi.bsky.social @taka-yamakoshi.bsky.social · 07/11/2025
I’m excited to share our Findings of EMNLP paper w/ @cocoscilab.bsky.social , @rtommccoy.bsky.social, and @rdhawkins.bsky.social ! Language models, unlike humans, require large amounts of data, which suggests the need for an inductive bias. But what kind of inductive biases do we need?
175
Reposted by Rabiraj Banerjee
Subbarao Kambhampati (కంభంపాటి సుబ్బారావు) @rao2z.bsky.social · 14/09/2025
In the year since LRMs ("reasoning models") hit the scene, we have been trying to understand, analyze and demystify them.. Here are our efforts to date--conveniently all in one place..👇 www.linkedin.com/posts/subbar...
linkedin.com
In the year since LRMs ("reasoning models") hit the scene, we have been trying to understand, analyze and demystify them.. Here are our efforts to date--conveniently all in one… | Subbarao K...
In the year since LRMs ("reasoning models") hit the scene, we have been trying to understand, analyze and demystify them.. Here are our efforts to date--conveniently all in one place.. (𝗙𝗶𝗿𝘀𝘁..) 𝗘𝘃𝗮𝗹...
051
Reposted by Rabiraj Banerjee
Sung Kim @sungkim.bsky.social · 12/09/2025
A Survey of Reinforcement Learning for Large Reasoning Models Five sections: - Foundational Components - Foundational Problems - Training Resources - Applications - Future Directions
1183
Reposted by Rabiraj Banerjee
Yoav Goldberg @yoavgo.bsky.social · 27/08/2025
When reading AI reasoning text (aka CoT), we (humans) form a narrative about the underlying computation process, which we take as a transparent explanation of model behavior. But what if our narratives are wrong? We measure that and find it usually is. Now on arXiv: arxiv.org/abs/2508.16599
arxiv.org
Humans Perceive Wrong Narratives from AI Reasoning Texts
A new generation of AI models generates step-by-step reasoning text before producing an answer. This text appears to offer a human-readable window into their computation process, and is increasingly r...
48623
Reposted by Rabiraj Banerjee
Johan Ugander @jugander.bsky.social · 24/08/2025
Great interview with @stevenstrogatz.com with a lot of discussion of research advising. Parts reminded me of @eegilbert.org and @informor.bsky.social's (excellent) guides to PhD mentorship, with a big focus on ideation. Eric's: docs.google.com/document/d/1... Mor's: s.tech.cornell.edu/phd-syllabus/
docs.google.com
Syllabus for Eric's PhD students
Table of Contents Author 3 License 3 Purpose of this document 4 Acknowledgements 4 Perspective on the PhD 5 How long is it? 5 Funding 6 What will I (your advisor) get out of it? 6 What kinds of pr...
0497
Reposted by Rabiraj Banerjee
Nikhil Garg @nkgarg.bsky.social · 20/08/2025
We try to avoid self-promoting too much, but we (with @sjgreenwood.bsky.social) built a personalized feed with posts about papers from your network. Many people say it's the closest they can get to old academic twitter, and I hope you enjoy it and share with others too! bsky.app/profile/pape...
1195
Reposted by Rabiraj Banerjee
Jacob Eisenstein is at CoLM 🌉 @jacobeisenstein.bsky.social · 19/08/2025
061
Reposted by Rabiraj Banerjee
Margaret Mitchell @mmitchell.bsky.social · 13/08/2025
🤖 But wait! There's more! You can check out @shiraamitchell.bsky.social 's most recent update on the details of Calibration, posted yesterday! statmodeling.stat.columbia.edu/2025/08/12/s...
statmodeling.stat.columbia.edu
Survey Statistics: 2nd helpings of the 2nd flavor of calibration | Statistical Modeling, Causal Inference, and Social Science
073
Reposted by Rabiraj Banerjee
brendan o’connor @brenocon.bsky.social · 28/07/2025
#acl2025 anyone get a good quote of phil resnik's last comment? context: (some?all?) panelists & him agree the field needs more deep, careful research on smaller models to do better science. everyone is frustrated with impossibility of large-scale pretraining experiments
171
Rabiraj Banerjee @rabirajb.bsky.social · 24/07/2025
@kennyjoseph.bsky.social , Kenny check this thread out
020
Rabiraj Banerjee @rabirajb.bsky.social · 24/07/2025
aclanthology.org/2023.emnlp-m..., for Active Learning I really liked this paper, uses LLMs as annotator for knowledge distillation for small LMs
aclanthology.org
010
Reposted by Rabiraj Banerjee
Maria Antoniak @mariaa.bsky.social · 23/07/2025
What are your favorite recent papers on using LMs for annotation (especially in a loop with human annotators), synthetic data for task-specific prediction, active learning, and similar? Looking for practical methods for settings where human annotations are costly. A few examples in thread ↴
137923
Rabiraj Banerjee @rabirajb.bsky.social · 21/07/2025
This is so mean !!!
000
Reposted by Rabiraj Banerjee
Kenny Joseph @kennyjoseph.bsky.social · 17/07/2025
UB's new Department of AI and Society is hiring faculty across ranks (Assistant, Associate, Full Professor). We’re looking for transdisciplinary scholars interested in building AI by society, for society. Start dates begin Fall 2025. More info: www.ubjobs.buffalo.edu/postings/57734
ubjobs.buffalo.edu
Assistant, Associate or Full Professor, AI & Society
The Department of AI and Society (AIS) at the University at Buffalo (UB) invites candidates to apply for multiple positions as Assistant Professor, Associate Professor, or Full Professor. The new AIS ...
0119
Reposted by Rabiraj Banerjee
Alexander Hoyle @alexanderhoyle.bsky.social · 16/07/2025
Took me a second, but I knew I'd seen something related to this recently: arxiv.org/abs/2505.120...
arxiv.org
Do different prompting methods yield a common task representation in language models?
Demonstrations and instructions are two primary approaches for prompting language models to perform in-context learning (ICL) tasks. Do identical tasks elicited in different ways result in similar rep...
1102
Reposted by Rabiraj Banerjee
Yanai Elazar @yanai.bsky.social · 01/07/2025
Check out our take on Chain-of-Thought. I really like this paper as a survey on the current literature on what CoT is, but more importantly on what it's not. It also serves as a cautionary tale to the (apparently quite common) misuse of CoT as an interpretable method.
1134
Rabiraj Banerjee @rabirajb.bsky.social · 01/07/2025
These are just battle scars of doing Data Science and ML Engineering in the industry!!
000
Reposted by Rabiraj Banerjee
Mark Riedl @markriedl.bsky.social · 01/07/2025
Hi everyone. I'm excited to announce that I will be organizing a 2nd Summit on Responsible Computing, AI, and Society rcais.github.io October 27-29, 2025. We will explore the future of computing for health, sustainability, human-centered AI, and policy. Please consider submitting a 1-page abstract
Screenshot of website
The Georgia Tech School of Interactive Computing is hosting the 2025 Summit on Responsible Computing, AI, and Society, October 27-29, 2025.

Overview

The Summit on Responsible Computing, AI, and Society aims to explore the future of computing for health, sustainability, human-centered AI, and policy. The summit will bring together luminary researchers in computing for health, sustainability, human-centered AI, and tech policy to lay out the frontiers of these critical fields, and to plot out how they must evolve.
1174
Rabiraj Banerjee @rabirajb.bsky.social · 27/06/2025
Send him a note of appreciation 😊
000
Rabiraj Banerjee @rabirajb.bsky.social · 26/06/2025
To @pcarragher.bsky.social @lleibm.bsky.social , @jmendelsohn2.bsky.social , Evan and Catherine, and others for some really fruitful and nice convos, hope to see you all soon.
020
Rabiraj Banerjee @rabirajb.bsky.social · 26/06/2025
A huge shoutout to the organizing team, and to the web chair @andersgiovanni.com for updating the schedule in such an easy to follow manner, hope you get some well deserved rest (as an ex web chair I know the pain)
220
Rabiraj Banerjee @rabirajb.bsky.social · 26/06/2025
So ICWSM concluded today and it was a blast, was a great honor to attend @icwsm.bsky.social at Copenhagen and present my work with @kennyjoseph.bsky.social and other colleagues. The paper link is here : ojs.aaai.org/index.php/IC...,
ojs.aaai.org
View of Measuring Dimensions of Self-Presentation in Twitter Bios and their Links to Misinformation Sharing
192
Reposted by Rabiraj Banerjee
Shubhendu Trivedi @shubhendu.bsky.social · 26/06/2025
This is great! The idea is somewhat obvious (good!), and I'm sure many have toyed with the connection to learning-to-rank. However, no work had developed it. This should be relevant for constructing valid PIs from just preferential feedback. openreview.net/pdf?id=ENJd3...
021
Rabiraj Banerjee @rabirajb.bsky.social · 26/06/2025
Me we were in the same session :) (Session 8)
100
Reposted by Rabiraj Banerjee
Naomi Saphra @nsaphra.bsky.social · 24/06/2025
🚨 New preprint! 🚨 Phase transitions! We love to see them during LM training. Syntactic attention structure, induction heads, grokking; they seem to suggest the model has learned a discrete, interpretable concept. Unfortunately, they’re pretty rare—or are they?
35410
Reposted by Rabiraj Banerjee
Hanna Wallach @hannawallach.bsky.social · 16/06/2025
Generative language systems are everywhere, and many of them stereotype, demean, or erase particular social groups.
192
Reposted by Rabiraj Banerjee
Hanna Wallach @hannawallach.bsky.social · 15/06/2025
Alright, people, let's be honest: GenAI systems are everywhere, and figuring out whether they're any good is a total mess. Should we use them? Where? How? Do they need a total overhaul? (1/6)
13311
Reposted by Rabiraj Banerjee
Paloma ✨ @itspaloma.bsky.social · 10/06/2025
🧵 1/ Las redes están llenas de odio. ¿Puede la inteligencia artificial ayudarnos a detectarlo… ⚖️ sin discriminar, 🚫 sin reforzar estereotipos, 🔁 y sin aprender a odiar? Esa es la gran pregunta de mi tesis. 👇 Te lo cuento en este #HiloTesis @crueuniversidades.bsky.social @filarramendi.bsky.social
1149
Reposted by Rabiraj Banerjee
Ahmad Beirami @abeirami.bsky.social · 27/05/2025
As we go through a lot of excitement about RL recently with lots of cool work/results, here is a reminder that RL with a reverse KL-regularizer to the base model cannot learn any new skills that were not already present in the base model. It can only amplify the weak skills. 👇
281
Reposted by Rabiraj Banerjee
Nathan Lambert @natolambert.bsky.social · 14/05/2025
My path into AI The sort of small wins that accumulate into a real career in AI. When I started grad school AI prof's didn't have space for me in their group and when I ended I had no papers at NeurIPS/ICLR/ICML, yet the process can still work. www.interconnects.ai/p/my-path-in...
interconnects.ai
My path into AI
How I got here. Building a career brick by brick over 8 years.
1296
Reposted by Rabiraj Banerjee
Shomir Wilson @shomir.bsky.social · 12/05/2025
I posted this on LinkedIn too and it has over 600 reactions there, with the caveat that I don't know how many are from bots.
001