Sign in

Avijit Ghosh

@evijit.io
3.1K followers 734 following 380 posts

Lead Technical AI Policy Researcher at Hugging Face @hf.co 🤗. Current focus: Responsible AI, AI for Science, and @eval-eval.bsky.social‬!

PostsRepliesMedia
Avijit Ghosh @evijit.io · 17/02/2026
This has been a massive community project, and we need you all to participate! See more: evalevalai.com/projects/eve...
073
Avijit Ghosh @evijit.io · 28/11/2025
Brisket for thanksgiving >>>> Turkey for thanksgiving
040
Avijit Ghosh @evijit.io · 13/11/2025
Incredible work done with literally the smartest and most passionate researchers I am lucky to work with. Paper co-led with @ankareuel.bsky.social and Jenny Chim, and other co-authors!
100
Avijit Ghosh @evijit.io · 13/11/2025
This only strengthens our position that good-quality, independent third-party evaluations are paramount for AI safety.
111
Avijit Ghosh @evijit.io · 13/11/2025
First-party reports are less transparent or lower quality. We conducted interviews with eval practitioners and found that companies have laid off or reassigned teams dedicated to documentation & social impact evals, or they are being told to focus more on capability reporting.
110
Avijit Ghosh @evijit.io · 13/11/2025
This is true even at the provider level. We find for e.g., that Google used to do a lot more reporting about their model evaluations in 2022 and 2023 but they reduced reporting in the Gemini era, and same can be seen for Meta over successive Llama versions.
110
Avijit Ghosh @evijit.io · 13/11/2025
We find that model developers have become less transparent about their eval results over time. For instance Env Cost reporting in first party reports (release docs, model cards, system cards) has drastically declined over time. Less than 15% mention labor or the environment!
100
Avijit Ghosh @evijit.io · 13/11/2025
Extremely thrilled to talk about our new paper: "Who Evaluates AI’s Social Impacts? Mapping Coverage And Gaps In First And Third Party Evaluations". This is the first big project output from the @eval-eval.bsky.social coalition! Thread below:
1207
Avijit Ghosh @evijit.io · 12/10/2025
Trying to start a new hobby and the internet is useless. Maybe AI will finally kill unstructured information retrieval for good and then we will be forced to call or visit friends for help again
162
Avijit Ghosh @evijit.io · 06/10/2025
We are launching Hugging Science: A global community addressing these barriers through: ✅ Collaborative challenges targeting upstream problems ✅ Cross-disciplinary education ✅ Recognition for data & infrastructure work ✅ Community-owned infrastructure All links follow 🤗
100
Avijit Ghosh @evijit.io · 06/10/2025
AI for scientific discovery is a social problem: In our new position paper, @cgeorgiaw.bsky.social and I show that culture, incentives, and coordination are the main obstacles to progress, and we are launching the Hugging Science Initiative to address this!
152
Avijit Ghosh @evijit.io · 28/09/2025
So fascinating (not really) to me that company execs and tier 1 AI conferences have gone in completely opposite directions as it relates to AI usage. Surely the best minds actually developing AI models know something about overreliance, productivity, and quality? Surely?
3286
Avijit Ghosh @evijit.io · 06/09/2025
How does Claude have the same response? This is sus
0171
Avijit Ghosh @evijit.io · 31/08/2025
These official ones are hideous oh god
120
Avijit Ghosh @evijit.io · 17/08/2025
I genuinely want to know the thought process here. Is each model iteration a new being? Is Claude 4.1 its own legal entity deserving of model welfare different from 4.0? Or is it like one human updating their world knowledge and becoming smarter? Was the very first trained Claude the robot embryo?
390
Avijit Ghosh @evijit.io · 08/08/2025
The product decision to discontinue older versions of ChatGPT and the comments on Reddit around that decision reminded me once again of discussions around “robot death”, which is real insofar as people’s feelings and emotions are real.
110
Avijit Ghosh @evijit.io · 15/07/2025
[New] Husbandposting! And yes we had a croquembouche for dessert because I saw it on masterchef once and I’ve always wanted that ❤️
2120
Avijit Ghosh @evijit.io · 15/07/2025
Who are the most prolific contributors? Research institutions lead: AI2 (Allen Institute) emerges as one of the most active contributors, alongside significant activity from IBM, NVIDIA, and international organizations. The open source ecosystem spans far beyond Big Tech!
110
Avijit Ghosh @evijit.io · 15/07/2025
Let's also talk about datasets: - Most downloaded datasets are evaluation benchmarks (MMLU, Squad, GLUE) - Universities and research institutions dominate foundational data - Domain-specific datasets thrive in finance, healthcare, robotics, and science - Open datasets power most AI development!
110
Avijit Ghosh @evijit.io · 15/07/2025
Looking at a single model's stats often does not tell the full story of its usefulness. The Qwen, Llama, and Gemma models have led to a universe of derivative models on the hub, all made by the community. This is the beauty of open source!
110
Avijit Ghosh @evijit.io · 15/07/2025
Legacy models like Clip, GPT-2, BERT, etc. remain among the most downloaded models despite being years old, showing that modern chat interfaces represent just one slice of AI applications! The ecosystem is much more diverse than frontier model discussions suggest.
176
Avijit Ghosh @evijit.io · 15/07/2025
Small models consistently outperform large variants in downloads, even within the same model family. This suggests practical deployment considerations often matter more than maximum capability. The community is building for real-world use, not just benchmarks.
162
Avijit Ghosh @evijit.io · 01/07/2025
Generally a big fan of LED frame stages, I absolutely loved the Eurovision main stage this year
000
Avijit Ghosh @evijit.io · 01/07/2025
I still think the set design of the Evita revival at the American Rep theater at Harvard was the most stunning interpretation of all time - I hope this concept makes it to Broadway at some point 🤩
100
Avijit Ghosh @evijit.io · 17/06/2025
Generative AI often renders the user invisible in their limited worldview. Please sign up for a short interactive workshop on AI, Misrepresentation and Mental Health, at both @facct.bsky.social in Athens, and Alt-FAccT in NYC! Limited space, so hurry! Sign up here! tinyurl.com/ai-mirrors
041
Avijit Ghosh @evijit.io · 09/06/2025
Living downtown and literally 2 blocks from the Opera House is certainly a clutch because I’m always late to things. Catch Roméo et Juliette playing in Boston it was great 😍
010
Avijit Ghosh @evijit.io · 08/06/2025
How do you feel about compulsory vibe workplaces:
200
Avijit Ghosh @evijit.io · 15/04/2025
…people had a wide range of emotional responses to things like whether or not the outputs sounded a little robotic, what the privacy and security implications were, and how it made them think deeper about technological hurdles they face on a day to day basis (like using Alexa).
110
Avijit Ghosh @evijit.io · 15/04/2025
Thrilled to share that our paper: "It's not a representation of me": Examining Accent Bias and Digital Exclusion in Synthetic AI Voice Services - has been accepted at @facct.bsky.social 2025! - with @shiramichel.bsky.social , Sufi Kaur, Sarah Gilespie, Jeffrey Gleason and Dr. Christo Wilson.
2157
Avijit Ghosh @evijit.io · 10/04/2025
🚨 New Article: Empowering Public Organizations: Preparing Your Data for the AI Era, with @yjernite.bsky.social Let’s discuss how public organizations can unlock the full potential of their data in the age of AI.
131
Avijit Ghosh @evijit.io · 21/03/2025
On March 14, Hugging Face (@hf.co) submitted our response to the White House Office of Science and Technology Policy's request for information on the AI Action Plan. Here are the highlights:
151
Avijit Ghosh @evijit.io · 12/03/2025
H/T Anjali Singh for sending this pic. Every single day our paper seems more and more like a warning sign ;-; Automating away human oversight via cognitive laziness is in no one's best interests (other than maybe the model developer -- and even then what about liabilities?)
141
Avijit Ghosh @evijit.io · 26/02/2025
I'll be at the AAAI Conference in Philadelphia this week, where I am part of two accepted papers: 🧵
161
Avijit Ghosh @evijit.io · 25/02/2025
2. Quoted on a @businessinsider.com piece by Effie Webb on AI Agents and Job Boards. I push back against conflating autonomy with agency, and point out that agents with human oversight will augment, not replace, human workers.
110
Avijit Ghosh @evijit.io · 25/02/2025
I had a couple of press mentions this week 🧵: 1. I spoke to Shraddha Goled for Tech Circle on the harms of Openwashing and better open standards! We also talked about @hf.co's Open R1 Project.
130
Avijit Ghosh @evijit.io · 23/02/2025
New work: Protecting Human Cognition in the Age of AI - with Anjali Singh, Karan Taneja, and Klara Guan. We claim that overreliance on GenAI models disrupt traditional learning pathways. We suggest best practices for better teaching, testing, and learning tools to restore these paths:
23010
Avijit Ghosh @evijit.io · 14/02/2025
Last Friday, I spoke on a panel at the MIT Sloan AI Conference. I discussed the broken AI Harm reporting landscape, the importance of evals, safe harbors, structured disclosures, and our proposed Coordinated Flaws Disclosure framework as a path forward. Great questions and thanks for having me!
071
Avijit Ghosh @evijit.io · 11/02/2025
Contrary to popular belief, unlike my colleagues I’m not in Paris for the AI Action Summit because I’m in India for my cousin’s wedding and taking (literally) 500 pictures per day
040
Avijit Ghosh @evijit.io · 07/02/2025
I’ll be speaking about Coordinated Disclosures at a panel at the @mitsloan.bsky.social AI Conference tomorrow morning! See you there :)
140
Avijit Ghosh @evijit.io · 06/02/2025
In fact, when we look at the entire risk-benefit landscape, we find that at the highest level of agency (full autonomy), most values break down and risk > benefit:
100
Avijit Ghosh @evijit.io · 06/02/2025
And we posit that Agents have values, that express themselves differently based on the level of agency. For example, tool interoperability may be high in high agency, but consistency may be low:
100
Avijit Ghosh @evijit.io · 06/02/2025
We also break down AI Agents into levels of Agency:
100
Avijit Ghosh @evijit.io · 06/02/2025
Okay, so -- what is an AI Agent? We looked at a million different definitions...
100
Avijit Ghosh @evijit.io · 13/01/2025
Funnily, the best table we discovered for this was a blog from Caterpillar, the construction company:
130
Avijit Ghosh @evijit.io · 13/01/2025
Behind the scenes! Now that the cat is out of the proverbial bag, I want to reflect on a few fun things that I came across while doing recon for the blog:
140
Avijit Ghosh @evijit.io · 03/01/2025
ICYMI, I was happy to contribute to this IST report on Navigating AI Compliance via tracing failure patterns from history. The report examines past compliance failures in other industries and how we can apply the lessons learned to AI Governance. Read: securityandtechnology.org/virtual-libr...
020
Avijit Ghosh @evijit.io · 30/12/2024
I got to be such a nerd at the NASA Kennedy Space Center at Cape Canaveral yesterday! Sooo cool 🤩. Ian finally took me there after 3 years of begging but it was so worth it
040
Avijit Ghosh @evijit.io · 27/12/2024
As 2025 shapes up to be another eventful year of AI governance, I am thinking about @runefeather.bsky.social and I's Dual Governance paper again...
140
Avijit Ghosh @evijit.io · 16/12/2024
And so long Vancouver, you were gorgeous 🥺
010
Avijit Ghosh @evijit.io · 16/12/2024
This has been quite the dazzling experience! Couldn’t have done it without dream team!
120