Sign in

Avijit Ghosh

@evijit.io
3.1K followers 734 following 380 posts

Lead Technical AI Policy Researcher at Hugging Face @hf.co 🤗. Current focus: Responsible AI, AI for Science, and @eval-eval.bsky.social‬!

PostsRepliesMedia
Reposted by Avijit Ghosh
Margaret Mitchell @mmitchell.bsky.social · 21/09/2026
A bit from me in @wired.com today wrt our research on AI agents with @evijit.io and Samir Passi. This is a really detailed article, thanks for writing it @thiccreese.bsky.social! www.wired.com/story/metas-...
wired.com
Meta's Muse Is Better at Surveilling Than Helping Me
The Muse app continues Meta’s trend of opting users into data collection for AI training. It also nudges you to share your bank account, email, and passport information.
22411
Avijit Ghosh @evijit.io · 23/07/2026
ICYMI, here’s the full bill text: www.congress.gov/bill/119th-c...
congress.gov
010
Avijit Ghosh @evijit.io · 23/07/2026
We have long championed this cause, and I had the pleasure of consulting on H.R.9333 - AI Flaw Reporting and Security Enhancement Act. If passed, this bipartisan bill would authorize NIST to administer a centralized AI Flaw and Incident reporting system, the need for which is becoming clearer!
110
Avijit Ghosh @evijit.io · 23/07/2026
In light of everything happening this week, I wanted to reiterate that AI Flaws and Incidents should be reported in a responsible, coordinated manner, and crucially, this should be administered by a competent body as opposed to voluntary self governance, akin to analogous systems in cyber (CVE).
131
Avijit Ghosh @evijit.io · 01/06/2026
Lots to unpack here. We started with 105 annotations. Please submit pull requests for more that we may have missed! Work with Ian Reynolds, @yjernite.bsky.social, and @mmitchell.bsky.social. huggingface.co/spaces/socie...
huggingface.co
The Annotated Encyclical - a Hugging Face Space by society-ethics
Magnifica Humanitas, annotated with AI-ethics research
030
Avijit Ghosh @evijit.io · 01/06/2026
Weekend mini project! Since commentary on AI is inherently interdisciplinary, we connected the observations in the The Pope's encyclical, with decades of scholarship in Responsible AI and Ethics research, and created an interactive space with these annotations!
230
Avijit Ghosh @evijit.io · 25/05/2026
I miss pre 2022 AI when there wasn’t so much noise in the field. Almost no separation between my job and personal life (a bunch of which is attributable to FOMO). Can’t even go on vacation without hearing about AI a few tables across at dinner. Even the Pope is involved
5452
Avijit Ghosh @evijit.io · 11/05/2026
Huh so I’m not the only one who noticed this! To add to this list, the Devil Wears Prada 2 also has a plot point about AI x Creatives. www.thewrap.com/creative-con...
thewrap.com
TV Writers Worry AI Will Replace Them. Now They’re Putting Those Anxieties on Screen
The creatives behind “Hacks,” “The Comeback,” “The Pitt” and “Matlock” tell TheWrap about tackling the emerging technology in their shows.
030
Avijit Ghosh @evijit.io · 11/05/2026
I would say something like Claude design, alongside traditional purpose built AI like those for scientific research, are ones that buck this trend.
110
Avijit Ghosh @evijit.io · 11/05/2026
I wrote this paper in November (!!!) things move slower in the conference world haha. That being said, the limitations section (and the alternative futures section) does cover this. “Agentic” systems as is currently popular are still LLMs with tools.
120
Avijit Ghosh @evijit.io · 11/05/2026
Should all our resources go towards building chatbots? What if we built systems that actually give people meaningful agency? Finally out as an accepted @facct.bsky.social paper! Joint work with my past collaborators Sourojit, Pranav and Sanjana. huggingface.co/papers/2605....
huggingface.co
Paper page - What if AI systems weren't chatbots?
Join the discussion on this paper page
34314
Avijit Ghosh @evijit.io · 17/02/2026
This has been a massive community project, and we need you all to participate! See more: evalevalai.com/projects/eve...
073
Avijit Ghosh @evijit.io · 01/02/2026
This 100%
010
Avijit Ghosh @evijit.io · 17/12/2025
This has to be rage bait. Did we not see the South Park episode where ChatGPT suggested a business idea to convert fries to salad? (And I tried to prompt myself too)
050
Avijit Ghosh @evijit.io · 28/11/2025
Brisket for thanksgiving >>>> Turkey for thanksgiving
040
Reposted by Avijit Ghosh
Shayne Longpre @shaynelongpre.bsky.social · 26/11/2025
Who is winning the open AI race? Our new study Economies of Open Intelligence maps @hf.co 851k models' downloads 2020→2025. 1) Power rebalance: US tech ↓; China + community ↑ 2) Models size & efficient ↑ (MoE, quant, multimodal) 3) Intermediary layers ↑ (adapters/quantizers) 4) Transparency ↓ /🧵
293
Avijit Ghosh @evijit.io · 25/11/2025
I used to love the word “key” until AI models decided to love it and now I cringe at “key takeaways” in text material :(
020
Avijit Ghosh @evijit.io · 20/11/2025
It’s that time of the year again! I’ll be at @neuripsconf.bsky.social this year too :) If you’re interested in Responsible AI, AI Evals ( @eval-eval.bsky.social ) or AI4Science (Hugging Science), say hi!
141
Reposted by Avijit Ghosh
EvalEval Coalition @eval-eval.bsky.social · 13/11/2025
🚨 AI keeps scaling, but social impact evaluations aren’t–and the data proves it 🚨 Our new paper, 📎“Who Evaluates AI’s Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations,” analyzes hundreds of evaluation reports and reveals major blind spots ‼️🧵 (1/7)
1113
Avijit Ghosh @evijit.io · 13/11/2025
A (very incomplete) frontend of Eval Cards can be found here: evalcards.evalevalai.com, and we are now collecting eval datasets (to show in eval cards) on github: github.com/evaleval/eve... If you want to help see eval cards come alive, get in touch!
evalcards.evalevalai.com
AI Evaluation Dashboard
Professional AI system evaluation and assessment tool
000
Avijit Ghosh @evijit.io · 13/11/2025
Finally, what's next from here? Almost every developer we spoke to said that what we need is a standardized way of reporting, aggregating and comparing all the evals done by both 1st and 3rd parties for a model. This is actually our next project: Eval Cards!
100
Avijit Ghosh @evijit.io · 13/11/2025
Incredible work done with literally the smartest and most passionate researchers I am lucky to work with. Paper co-led with @ankareuel.bsky.social and Jenny Chim, and other co-authors!
100
Avijit Ghosh @evijit.io · 13/11/2025
Read the detailed results here: arxiv.org/abs/2511.05613 We also release the code, and the full annotated dataset on Hugging Face (link in paper).
arxiv.org
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
Foundation models are increasingly central to high-stakes AI systems, and governance frameworks now depend on evaluations to assess their risks and capabilities. Although general capability evaluation...
110
Avijit Ghosh @evijit.io · 13/11/2025
This only strengthens our position that good-quality, independent third-party evaluations are paramount for AI safety.
111
Avijit Ghosh @evijit.io · 13/11/2025
First-party reports are less transparent or lower quality. We conducted interviews with eval practitioners and found that companies have laid off or reassigned teams dedicated to documentation & social impact evals, or they are being told to focus more on capability reporting.
110
Avijit Ghosh @evijit.io · 13/11/2025
This is true even at the provider level. We find for e.g., that Google used to do a lot more reporting about their model evaluations in 2022 and 2023 but they reduced reporting in the Gemini era, and same can be seen for Meta over successive Llama versions.
110
Avijit Ghosh @evijit.io · 13/11/2025
We find that model developers have become less transparent about their eval results over time. For instance Env Cost reporting in first party reports (release docs, model cards, system cards) has drastically declined over time. Less than 15% mention labor or the environment!
100
Avijit Ghosh @evijit.io · 13/11/2025
We take a look at the entire eval landscape, specifically social impact evals across 7 dimensions: Bias & Harm, Sensitive Content, Performance Disparity, Env. Costs & Emissions, Privacy & Data, Financial Costs, and Moderation Labor. Who is reporting these evals?
100
Avijit Ghosh @evijit.io · 13/11/2025
Extremely thrilled to talk about our new paper: "Who Evaluates AI’s Social Impacts? Mapping Coverage And Gaps In First And Third Party Evaluations". This is the first big project output from the @eval-eval.bsky.social coalition! Thread below:
1207
Avijit Ghosh @evijit.io · 11/11/2025
… this looks like the Nature font oh no
120
Avijit Ghosh @evijit.io · 06/11/2025
We have a call for posters out! Please submit your extended abstracts, it should be quick and easy. And just like last year, provocative work is especially encouraged as it makes for such interesting conversation 😈
032
Avijit Ghosh @evijit.io · 02/11/2025
This. Copyright is a tool for protection but it’s not everything. In fact, there’s research showing that it is possible to create competitive language models using public domain data only. The proliferation of copyright respecting models would not solve the labor impact policy problem.
1242
Avijit Ghosh @evijit.io · 01/11/2025
Going to San Diego for Neurips? We at @eval-eval.bsky.social , along with the UK AISI, are hosting a closed door state of evals workshop at @ucsandiego.bsky.social on Dec 8th. Request to join below! :) evaleval.github.io/events/works...
evaleval.github.io
2025 Workshop on Evaluating AI in Practice
EvalEval, UK AI Security Institute (AISI), and UC San Diego (UCSD) are excited to announce the upcoming Evaluating AI in Practice workshop, happening on December 8, 2025, in San Diego, California.
000
Avijit Ghosh @evijit.io · 01/11/2025
The thing about non survey papers is that they can still be problematic/fake science etc, and arxiv needs a long overdue + moderated comments section
010
Avijit Ghosh @evijit.io · 28/10/2025
Datasets are the backbone of AI for Science, and we want to support scientific data natively on Hugging Face. The amazing @lhoestq.hf.co started a discussion on GH for this! Please engage (better still, submit a PR) so we can start supporting your 🫵 dataset: github.com/huggingface/...
github.com
Support scientific data formats · Issue #7804 · huggingface/datasets
List of formats and libraries we can use to load the data in datasets: DICOMs: pydicom NIfTIs: nibabel WFDB: wfdb cc @zaRizk7 for viz Feel free to comment / suggest other formats and libs you'd lik...
020
Avijit Ghosh @evijit.io · 24/10/2025
Yes! The Science/Tech/Cyber committee is doing really good work too. Well intentioned folks there trying to actually engage with researchers and industry folks. Love MA
020
Avijit Ghosh @evijit.io · 24/10/2025
Random off the cuff observation about American AI: LLM folks seem to be concentrated in SF, but AI4Science folks seem to be concentrated in Boston. Meaning as the former gets oversaturated and the latter is only getting started, I expect Boston to be the next big AI epicenter! 💪
120
Reposted by Avijit Ghosh
EvalEval Coalition @eval-eval.bsky.social · 24/10/2025
🌟 Weekly AI Evaluation Spotlight 🌟 🤖 Did you know malicious actors can exploit trust in AI leaderboards to promote poisoned models in the community? This week's paper 📜"Exploiting Leaderboards for Large-Scale Distribution of Malicious Models" by @iamgroot42.bsky.social explores this!
152
Avijit Ghosh @evijit.io · 20/10/2025
Oof
000
Avijit Ghosh @evijit.io · 20/10/2025
I have started requesting that panel moderators provide a disclaimer at panels I am on that not all my opinions are provided by my employer. HF ppl largely believe in democratization of AI and open source, but we actually have intense healthy debates internally on edge topics! It's great :)
030
Avijit Ghosh @evijit.io · 20/10/2025
+1000. I miss life pre-AI hype when the discourse around AI was more scientific and people used to attribute papers and opinions to scientists instead of to their companies. Not all orgs block research papers and sanity check their papers via legal teams, and HF, especially so, is very distributed.
130
Avijit Ghosh @evijit.io · 20/10/2025
Huh, so interesting re: art therapy! Re: The turning off adult content, this is already what Google does (SafeSearch on, off, or blurred, off by default). I do think it gives back agency to adult users without shaming sexual content from a puritan perspective.
020
Avijit Ghosh @evijit.io · 20/10/2025
This doesn’t quite answer what I’m asking. Currently there’s nothing preventing people from going to AO3, Literotica, etc. should those be banned too? What is it about porn specifically that seems to be the problem (as opposed to harms of personification/emotional attachment)
100
Avijit Ghosh @evijit.io · 20/10/2025
THANK YOU! That's the phrasing I was looking for and yes 100% want to hear from sex work experts on this!
020
Avijit Ghosh @evijit.io · 20/10/2025
Text by itself maybe less so but text in concert with visuals yes I feel
000
Avijit Ghosh @evijit.io · 20/10/2025
Fair, it also more specifically speaks to the need for better verification of the truth. I saw a demo recently where a system was developed to generate a fake nytimes article, complete with branding images, text, and sometimes generated images in the article which all looked believable.
110
Avijit Ghosh @evijit.io · 20/10/2025
But it isn’t clear to me that deepfakes will be allowed or that they’re even talking about non-text material. My read of the tweet was erotica as in fan fiction stories etc, but maybe people are reading this differently
200
Avijit Ghosh @evijit.io · 20/10/2025
Obviously, I don't have full context on what OAI will define as acceptable erotica and might have more refined views then. Still, it is a double standard to have universal content control on a generative model and not on, say, web search.
000
Avijit Ghosh @evijit.io · 20/10/2025
Hey, so can someone tell me why ChatGPT generating erotica is bad, any more so than it generating anything else? Obviously anything non-consensual or age-inappropriate is bad, but I don't see why some researchers in my timeline are up in arms about it, while Grok already does this.
7102
Avijit Ghosh @evijit.io · 17/10/2025
We're starting a weekly paper spotlight series! Come engage with the posts and let's improve evals together! :) First up: Do Large Language Model Benchmarks Test Reliability?
0111