Sign in

Shayne Longpre

@shaynelongpre.bsky.social
4.4K followers 337 following 103 posts

MTS @ Anthropic. 🇨🇦 Prev: MIT, Google, Apple, Stanford. Interests: AI/ML/NLP, Data-centric AI, transparency & societal impact

PostsRepliesMedia
Shayne Longpre @shaynelongpre.bsky.social · 18/08/2026
(New Release!) How are people really using AI? Today, researchers at MIT, Stanford, + 12 academic institutions are launching the AI Observatory: public infra for measuring how people actually use AI assistants in the wild. 🔗 ai-observatory.org 📜 www.dataprovenance.org/ai_observato...
ai-observatory.org
The AI Observatory — interactive companion
A public measure of real-world AI use. Aggregating seven sources of real AI conversations, 24,521 conversations, 92,493 user–assistant exchange pairs, and a 145-feature taxonomy for the NeurIPS 2026 p...
2132
Reposted by Shayne Longpre
Eileen Guo (is on leave for rest of 2026) @eileenguo.bsky.social · 18/08/2026
I wrote about how, beyond what AI companies tell us, we still know very about real people's AI use--and the AI Observatory project from @shaynelongpre.bsky.social @ankareuel.bsky.social + team trying to change this: www.technologyreview.com/2026/08/18/1...
technologyreview.com
We still don’t know how people are really using AI
But a new study shows that work use cases make up less of the picture than AI companies claim.
2225
Reposted by Shayne Longpre
Simons Institute for the Theory of Computing @simonsinstitute.bsky.social · 24/05/2026
1/2 We've heard of the curse of dimensionality. But what about the curse of multilinguality? "If you try to pack more languages into the same size model, they are each going to degrade," said @shaynelongpre.bsky.social of @mit.edu at the Simons Institute.
172
Shayne Longpre @shaynelongpre.bsky.social · 09/12/2025
Excited to publish this piece!
061
Reposted by Shayne Longpre
A. Feder Cooper @afedercooper.bsky.social · 05/12/2025
I’ll be hanging out at our poster on membership inference, but in the same slot Brian Lester will present our work on “The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text” (poster 102)! [arxiv.org/abs/2506.05209]
152
Reposted by Shayne Longpre
Justin Hendrix @justinhendrix.bsky.social · 26/11/2025
"China has overtaken the US in the global market for 'open' artificial intelligence models, gaining a crucial edge over how the powerful technology is used around the world."
ft.com
China leapfrogs US in global market for ‘open’ AI models
Beijing-backed technology gains ground as American giants hold fast to ‘closed’ AI strategies
1246
Shayne Longpre @shaynelongpre.bsky.social · 26/11/2025
Who is winning the open AI race? Our new study Economies of Open Intelligence maps @hf.co 851k models' downloads 2020→2025. 1) Power rebalance: US tech ↓; China + community ↑ 2) Models size & efficient ↑ (MoE, quant, multimodal) 3) Intermediary layers ↑ (adapters/quantizers) 4) Transparency ↓ /🧵
293
Reposted by Shayne Longpre
Raphaël Merx @rapha.dev · 30/10/2025
This is some legit really impressive work!!
031
Shayne Longpre @shaynelongpre.bsky.social · 28/10/2025
📢Thrilled to introduce ATLAS 🗺️: the largest multilingual scaling study to-date—we ran 774 exps (10M-8B params, 400+ languages) to answer: 🌍 Is scaling diff by lang? 🧙‍♂️ Can we model the curse of multilinguality? ⚖️ Pretrain vs finetune from checkpoint? 🔀 X-lingual transfer scores across langs? 1/🧵
1181
Reposted by Shayne Longpre
Dustin Wright @dustinbwright.com · 13/10/2025
Which, whose, and how much knowledge do LLMs represent? I'm excited to share our preprint answering these questions: "Epistemic Diversity and Knowledge Collapse in Large Language Models" 📄Paper: arxiv.org/pdf/2510.04226 💻Code: github.com/dwright37/ll... 1/10
29528
Shayne Longpre @shaynelongpre.bsky.social · 06/05/2025
Delighted to see BigGen Bench paper receive the 🏆best paper award 🏆at NAACL 2025! BigGen Bench introduces fine-grained, scalable, & human-aligned evaluations: 📈 77 hard, diverse tasks 🛠️ 765 exs w/ ex-specific rubrics 📋 More human-aligned than previous rubrics 🌍 10 languages, by native speakers 1/
1213
Reposted by Shayne Longpre
Sara Hooker @sarahooker.bsky.social · 30/04/2025
It is critical for scientific integrity that we trust our measure of progress. The @lmarena.bsky.social has become the go-to evaluation for AI progress. Our release today demonstrates the difficulty in maintaining fair evaluations on the Arena, despite best intentions.
19569
Reposted by Shayne Longpre
Knight First Amendment Institute @knightcolumbia.org · 30/04/2025
How should regulatory proposals adapt to the prevalence of general-purpose AI when the global geopolitical order is being reconfigured? @atoosakz.bsky.social, Deirdre K. Mulligan, @randomwalker.bsky.social, @alondra.bsky.social, & @shaynelongpre.bsky.social weigh in: youtu.be/cRsbjGFPJaM?...
youtu.be
Day 1 Opening Remarks and Panel 1: Regulating AI in Democratic Upheaval (AI & Democratic Freedoms)
YouTube video by Knight First Amendment Institute
042
Shayne Longpre @shaynelongpre.bsky.social · 22/04/2025
🛬 in Singapore for #ICLR2025! DM me to catch up—but only if you have a local food/bar/event rec!
030
Shayne Longpre @shaynelongpre.bsky.social · 14/04/2025
Thrilled our global data ecosystem audit was accepted to #ICLR2025! Empirically, it shows: 1️⃣ Soaring synthetic text data: ~10M tokens (pre-2018) to 100B+ (2024). 2️⃣ YouTube is now 70%+ of speech/video data but could block third-party collection. 3️⃣ <0.2% of data from Africa/South America. 1/
1124
Reposted by Shayne Longpre
Knight First Amendment Institute @knightcolumbia.org · 11/04/2025
📍EVENT: Day 2 of our “AI and Democracy” symposium will be kicking off shortly. Programming will begin with welcome remarks from George Deodatis @columbiaseas.bsky.social at 9:30am ET. #AIDemocraticFreedoms Watch the full event on our livestream here: www.youtube.com/watch?v=X1gj...
youtube.com
Artificial Intelligence and Democratic Freedoms (Day 2)
YouTube video by Knight First Amendment Institute
1134
Reposted by Shayne Longpre
Marzieh Fadaee @mziizm.bsky.social · 10/04/2025
Very excited to release Kaleidoscope—a multilingual, multimodal evaluation set for VLMs, built as part of our open-science initiative! 🌍 18 languages (high-, mid-, low-) 📚 21k questions (55% require image understanding) 🧪 STEM, social science, reasoning, and practical skills
1104
Reposted by Shayne Longpre
Knight First Amendment Institute @knightcolumbia.org · 10/04/2025
Panel 1: Regulating AI in a Time of Democratic Upheaval starts in approximately 5 minutes. Panelists: @atoosakz.bsky.social, @randomwalker.bsky.social, @alondra.bsky.social, and Deirdre K. Mulligan. Moderator: @shaynelongpre.bsky.social. #AIDemocraticFreedoms
142
Shayne Longpre @shaynelongpre.bsky.social · 09/04/2025
This week, @stanfordhai.bsky.social released the 2025 AI Index. It’s well worth reading to understand the evolving ecosystem of AI. Some highlights that stood out to me: 1/
141
Shayne Longpre @shaynelongpre.bsky.social · 01/04/2025
Excited to speak at the workshop on Technical AI Governance in Vancouver this summer! #ICML2025
082
Reposted by Shayne Longpre
Andy Tseng @andytseng.bsky.social · 17/03/2025
#AI is evolving fast, and so are its flaws. A fresh approach to finding and reporting AI bugs is long overdue. Great initiative by @shaynelongpre.bsky.social and team, transparency and accountability in AI development are essential! #AISafety #ResponsibleAI #AIEthics #MIT
wired.com
Researchers Propose a Better Way to Report Dangerous AI Flaws
After identifying major flaws in popular AI models, researchers are pushing for a new system to identify and report bugs.
162
Reposted by Shayne Longpre
WIRED @wired.com · 13/03/2025
After identifying major flaws in popular AI models, researchers are pushing for a new system to identify and report bugs.
wrd.cm
Researchers Propose a Better Way to Report Dangerous AI Flaws
After identifying major flaws in popular AI models, researchers are pushing for a new system to identify and report bugs.
610317
Shayne Longpre @shaynelongpre.bsky.social · 13/03/2025
Thank you @willknight.bsky.social for excellent coverage of our new proposal! www.wired.com/story/ai-res...
wired.com
Researchers Propose a Better Way to Report Dangerous AI Flaws
After identifying major flaws in popular AI models, researchers are pushing for a new system to identify and report bugs.
010
Shayne Longpre @shaynelongpre.bsky.social · 13/03/2025
What are 3 concrete steps that can improve AI safety in 2025? 🤖⚠️ Our new paper, “In House Evaluation is Not Enough” has 3 calls-to-actions to empower evaluators: 1️⃣ Standardized AI flaw reports 2️⃣ AI flaw disclosure programs + safe harbors. 3️⃣ A coordination center for transferable AI flaws. 1/🧵
1118
Reposted by Shayne Longpre
Andy Sellars @sellars.bsky.social · 13/03/2025
Very glad to join this paper organized by @shaynelongpre.bsky.social. Here's the paper itself: crfm.stanford.edu/2025/03/13/t...
wired.com
Researchers Propose a Better Way to Report Dangerous AI Flaws
After identifying major flaws in popular AI models, researchers are pushing for a new system to identify and report bugs.
041
Reposted by Shayne Longpre
Artificial Intelligence Institute @ekaya.bsky.social · 04/03/2025
Bringing transparency to the data used to train artificial intelligence mitsloan.mit.edu/ideas-made-t...
mitsloan.mit.edu
Bringing transparency to the data used to train artificial intelligence | MIT Sloan
Using the wrong datasets to train AI models can result in legal risks, bias, or lower-quality models. The Data Provenance Initiative’s tool can help.
073
Shayne Longpre @shaynelongpre.bsky.social · 26/02/2025
Thrilled to be at #AAAI2025 for our tutorial, “AI Data Transparency: The Past, Present, and Beyond.” We’re presenting the state of transparency, tooling, and policy, from the Foundation Model Transparency Index, Factsheets, the the EU AI Act to new frameworks like @MLCommons’ Croissant. 1/
1154
Reposted by Shayne Longpre
klaudia jaźwińska @klaudia.bsky.social · 21/02/2025
Really excellent explainer by @shaynelongpre.bsky.social‬ that clearly lays out what's at stake in the "AI crawler wars"
How we stand to lose out

As this cat-and-mouse game accelerates, big players tend to outlast little ones.  Large websites and publishers will defend their content in court or negotiate contracts. And massive tech companies can afford to license large data sets or create powerful crawlers to circumvent restrictions. But small creators, such as visual artists, YouTube educators, or bloggers, may feel they have only two options: hide their content behind logins and paywalls, or take it offline entirely. For real users, this is making it harder to access news articles, see content from their favorite creators, and navigate the web without hitting logins, subscription demands, and captchas each step of the way.

Perhaps more concerning is the way large, exclusive contracts with AI companies are subdividing the web. Each deal raises the website’s incentive to remain exclusive and block anyone else from accessing the data—competitor or not. This will likely lead to further concentration of power in the hands of fewer AI developers and data publishers. A future where only large companies can license or crawl critical web data would suppress competition and fail to serve real users or many of the copyright holders.
073
Reposted by Shayne Longpre
Alex Abdo @alexabdo.bsky.social · 20/02/2025
Re: the FTC and the platforms We *should* be concerned about platform power over speech, but it isn’t censorship. As the Supreme Court said last year, the companies’ editorial decisions to moderate content are protected by the First Amendment. 1/
273
Reposted by Shayne Longpre
Renee DiResta @noupside.bsky.social · 20/02/2025
Let’s see if the algorithm and data remains transparent.
Musk to Revise Community Notes Amid Bias Concerns
Last updated 26 minutes ago
Elon Musk has announced intentions to revise the Community Notes feature on X, previously praised as a tool for unbiased fact-checking, due to concerns over manipulation by governments and legacy media.
His critique focused particularly on a note regarding Ukrainian President Volodymyr Zelensky's approval ratings, sparking a debate on whether this move is an attempt to maintain the feature's integrity or to align it with Musk's personal views.
This story is a summary of posts on X and may evolve over time. Grok can make mistakes, verify its outputs.
59114
Shayne Longpre @shaynelongpre.bsky.social · 19/02/2025
I compiled a list of resources for understanding AI copyright challenges (US-centric). 📚 ➡️ why is copyright an issue for AI? ➡️ what is fair use? ➡️ why are memorization and generation important? ➡️ how does it impact the AI data supply / web crawling? 🧵
140
Reposted by Shayne Longpre
Nick Diakopoulos @ndiakopoulos.bsky.social · 13/02/2025
Great point by @shaynelongpre.bsky.social on the AI crawler wars: "Unless we can nurture an ecosystem with different rules for different data uses, we may end up with strict borders across the web, exacting a price on openness and transparency." www.technologyreview.com/2025/02/11/1...
technologyreview.com
AI crawler wars threaten to make the web more closed for everyone
There’s an accelerating cat-and-mouse game between web publishers and AI crawlers, and we all stand to lose.
031
Reposted by Shayne Longpre
Stanford HAI @stanfordhai.bsky.social · 13/02/2025
Our latest brief highlights the importance of independent evaluation for AI safety and accountability, addressing barriers and proposing safe harbors to protect third-party research. @klyman.bsky.social @shaynelongpre.bsky.social @sayash.bsky.social @peterhenderson.bsky.social
192
Reposted by Shayne Longpre
dame @dame.is · 12/02/2025
Appreciated this thoughtful op-ed about web crawlers and AI training… the biological metaphors are good as well as “invasive species”. by @shaynelongpre.bsky.social Link: www.technologyreview.com/2025/02/11/1...
Put simply, following this path will shrink the biodiversity of the web. Crawlers from academic researchers, journalists, and non-AI applications may increasingly be denied open access. Unless we can nurture an ecosystem with different rules for different data uses, we may end up with strict borders across the web, exacting a price on openness and transparency.
061
Reposted by Shayne Longpre
Matteo Nebbiai @matteonebbiai.bsky.social · 12/02/2025
"A future where only large companies can license or crawl critical web data would suppress competition and fail to serve real users or many of the copyright holders." Excellent @shaynelongpre.bsky.social www-technologyreview-com.cdn.ampproject.org/c/s/www.tech...
www-technologyreview-com.cdn.ampproject.org
AI crawler wars threaten to make the web more closed for everyone
043
Reposted by Shayne Longpre
Kyle Lo @ COLM2026 @kylelo.bsky.social · 12/02/2025
(not a lawyer) reading more into judge's opinion on westlaw ross case, some details that should be noted: 1. ross isn't Generative AI. it's a retrieval system where users type in a question & it returns relevant legal docs. there's no generated answer. see bsky.app/profile/dbam...
143
Reposted by Shayne Longpre
Greg Leppert @leppert.me · 12/02/2025
Great op-ed from @shaynelongpre.bsky.social on the effects AI — as a technology and as a market — is having on the web. www.technologyreview.com/2025/02/11/1...
technologyreview.com
AI crawler wars threaten to make the web more closed for everyone
There’s an accelerating cat-and-mouse game between web publishers and AI crawlers, and we all stand to lose.
031
Reposted by Shayne Longpre
Knight First Amendment Institute @knightcolumbia.org · 12/02/2025
“What I see here is media organizations that have the power to fight back against Trump but aren’t. ... This is a moment in which we need these organizations to hold the powerful to account, hold Trump to account,” @jameeljaffer.bsky.social @democracynow.org: www.democracynow.org/2025/2/11/ja...
democracynow.org
Media Outlets Cave to Trump’s Threats as FCC Launches New Probes
We look at the Trump administration’s escalating attacks on press freedom, and how the media has responded with bended knee in some cases, with Jameel Jaffer, director of the Knight First Amendment In...
2198
Shayne Longpre @shaynelongpre.bsky.social · 12/02/2025
I wrote a spicy piece on "AI crawler wars"🐞 in @technologyreview.com (my first op-ed)! While we’re busy watching copyright lawsuits & the EU AI Act, there’s a quieter battle over data access that affects websites, everyday users, and the open web. 🔗 www.technologyreview.com/2025/02/11/1... 1/
technologyreview.com
AI crawler wars threaten to make the web more closed for everyone
There’s an accelerating cat-and-mouse game between web publishers and AI crawlers, and we all stand to lose.
1145
Shayne Longpre @shaynelongpre.bsky.social · 10/02/2025
1/ Last week, we published the International AI Safety Report—supported by 30 nations plus the OECD, UN, and EU. Over 100 independent experts contributed. I’m thankful to play a small writing role, focusing on “Risks of Copyright.” 🔗 bit.ly/40Vm7Mu
bit.ly
131
Shayne Longpre @shaynelongpre.bsky.social · 03/02/2025
Our updated Responsible Foundation Model Development Cheatsheet (250+ tools & resources) is now officially accepted to @tmlrorg.bsky.social (TMLR) 2025! It covers: - data sourcing, - documentation, - environmental impact, - risk eval - model release & licensing - ++
272
Reposted by Shayne Longpre
Eryk Salvaggio @eryk.bsky.social · 02/02/2025
It feels tactical to say this but if you like a text someone has written here consider sharing it. Likes on Bsky signal your approval, but they don’t bump it or amplify it because there’s no algorithm. *Sharing* independent writer’s work helps us find and grow an audience!
33922
Shayne Longpre @shaynelongpre.bsky.social · 01/02/2025
🪶 Some thoughts on DeepSeek, OpenAI, and the copyright battles: This isn’t the first time OpenAI has accused a Chinese company of breaking its Terms and training on ChatGPT outputs. Dec 2023: They suspended ByteDance’s accounts. 1/
191
Shayne Longpre @shaynelongpre.bsky.social · 29/01/2025
Will AI eat itself? I appeared in BBC Radio 4’s The Artificial Human, alongside NYU Prof. Julia Kempe, to discuss: • The explosion of AI-generated content • Synthetic data → model collapse or helpful? • The data economy & the vital role of humans www.bbc.co.uk/sounds/play/...
bbc.co.uk
The Artificial Human - Will AI Eat Itself? - BBC Sounds
Aleks Krotoski and Kevin Fong uncover what happens when you feed an AI its own output.
110
Reposted by Shayne Longpre
Dr Abeba Birhane @abeba.blacksky.app · 29/01/2025
apply here to join the accountability lab (@aial.ie)! aial.ie/pages/hiring/ currently accepting applications for phd and postdoc positions
aial.ie
Hiring
Post-doctoral research fellow x2 The AIAL is seeking two full-time post-doctoral fellows to work with Dr. Abeba Birhane and other lab members in 1) policy translation and 2) AI evaluation, both under ...
23917
Reposted by Shayne Longpre
Alex Abdo @alexabdo.bsky.social · 21/01/2025
I wrote up a quick reaction (at @justsecurity.org) to Trump’s executive order on free speech. www.justsecurity.org/106608/free-...
justsecurity.org
A Free Speech View on the “Free Speech” Executive Order
A Trump administration executive order purporting to curtail “jawboning” raises concerns among free-speech advocates.
32010
Reposted by Shayne Longpre
Knight First Amendment Institute @knightcolumbia.org · 21/01/2025
Trump’s #TikTok Executive Order aims to consolidate his power over the digital public sphere. Our statement from @ramyakrishnan.bsky.social below. knightcolumbia.org/content/trum...
0216
Shayne Longpre @shaynelongpre.bsky.social · 21/01/2025
Excited our new report will appear at #AAAI 2025 The #DEFCON2024 @aivillage.bsky.social Generative Red Team 2 (GRT2) Case Study spanned: ⚔️495 hackers, against AI2’s Olmo + WildGuard 🐞200 model flaw reports 💰$7k+ paid bounties Check it out! 🔗 arxiv.org/pdf/2410.12104
arxiv.org
050
Reposted by Shayne Longpre
Giada Pistilli @giadapistilli.com · 19/12/2024
Just shared my thoughts with @technologyreview.com on AI's data problem. With most training data coming from English-language sources, we're building AI systems that perpetuate a Western-centric worldview. Data diversity isn't a checkbox: it determines how AI systems represent daily life globally.
1215