Sign in

Stefan Baack

@sbaack.com
286 followers 590 following 21 posts

Senior researcher studying data governance and AI training data. Mastodon: @tootbaack@infosec.exchange he/him

PostsRepliesMedia
Reposted by Stefan Baack
Christo Buschek @handle.invalid · 18/05/2026
Read more at arxiv.org/abs/2605.14164. Great collaboration with @sbaack.com and @matybohacek.bsky.social . Best co-authors anyone can wish for.
arxiv.org
Unsteady Metrics and Benchmarking Cultures of AI Model Builders
The primary way to establish and compare competencies in foundation and generative AI models has shifted from peer-reviewed literature to press releases and company blog posts, where model builders hi...
081
Reposted by Stefan Baack
Christo Buschek @handle.invalid · 18/05/2026
Benchmarks as narrative devices by AI builders! Our new paper maps 231 benchmarks highlighted in 139 model releases by 11 model builders in 2025. I'm excited to present “Unsteady Metrics and Benchmarking Cultures of AI Model Builders” next month at #FAccT2026. Paper at arxiv.org/abs/2605.14164
The abstract of the "Unsteady Metrics and Benchmarking Cultures of AI Model Builders" paper, highlighting some findings.
172
Stefan Baack @sbaack.com · 22/01/2026
Some #generativeAI developers love to destroy the foundations of the tech they build. #WIkipedia is one of the most valuable sources of genAI training data. Undermining it is not just attacking a great common resource. It's also completely self-destructive arstechnica.com/ai/2026/01/n...
arstechnica.com
Wikipedia volunteers spent years cataloging AI tells. Now there's a plugin to avoid them.
The web's best guide to spotting AI writing has become a manual for hiding it.
010
Reposted by Stefan Baack
Alex Reisner @alexreisner.bsky.social · 04/11/2025
A little-known nonprofit has been lying to news publishers while funneling millions of paywalled articles to tech companies for AI training. Read my investigation in The Atlantic. www.theatlantic.com/technology/2...
theatlantic.com
The Nonprofit Doing the AI Industry’s Dirty Work
The web archive Common Crawl has been quietly funneling paywalled articles to AI companies—and lying to publishers about it.
12111
Stefan Baack @sbaack.com · 02/10/2025
Check in if you're interested in my thoughts about what open source AI should aspire to be in relation to proprietary AI
042
Stefan Baack @sbaack.com · 16/07/2025
"The update is yet another signal that payment processors...are currently the ultimate arbiter of what kind of content can be made easily available online, or not."
020
Stefan Baack @sbaack.com · 30/06/2025
The key questions we always should ask when people talk about AI: What is being automated and why? @alexhanna.bsky.social @weizenbauminstitut.bsky.social
0165
Stefan Baack @sbaack.com · 30/06/2025
"AI is a labor disciplining device" @alexhanna.bsky.social
010
Reposted by Stefan Baack
Daniel Drepper @danieldrepper.bsky.social · 11/05/2025
“The reporter is a man of critical value. No amount of money or effort spent in fitting the right men for this work could possibly be wasted, for the health of society depends upon the quality of the information it receives.” — Walter Lippmann [a century later, I’d swap “man” for “person” though]
041
Reposted by Stefan Baack
Christo Buschek @handle.invalid · 09/12/2024
New Release! Most AI deepfakes aren't political. 90% of deepfakes are non-consensual intimate imagery. 99% of victims are women. Max Hoppensted, @rechercheur.bsky.social, @romanhoefner.bsky.social, and I uncover a deepfake community and the business behind undress apps www.spiegel.de/netzwelt/web...
spiegel.de
(S+) Deepfake-Pornos: Das perfide Geschäft mit gefälschten Sexvideos
Tausende Frauen werden Opfer von gefakten Pornos, in denen ihr Gesicht zu sehen ist. Betroffen sind minderjährige Mädchen, Prominente, Politikerinnen. Dahinter stecken skrupellose Geschäftsleute. Der ...
23322
Stefan Baack @sbaack.com · 08/04/2025
"brainstorming and iteration is...a crucial everyday part of game development...and is not a problem to be solved...I have had many discussions with other game developers who interact with AI engineers and savants who believe our industry pipelines need 'fixing' by them and them alone"
012
Reposted by Stefan Baack
FragDenStaat @fragdenstaat.de · 26/03/2025
Die Union will das Informationsfreiheitsgesetz abschaffen. @arnesemsrott.bsky.social: „Öffentliche Kontrolle &Transparenz sind der Union offenbar ein Dorn im Auge. Sie will unbehelligt durchregieren. Rechte der Öffentlichkeit stören dabei offenbar." Pressemitteilung: fragdenstaat.de/newsletter/a...
fragdenstaat.de
Union will Informationsfreiheitsgesetz abschaffen: Frontalangriff auf Transparenz und Demokratie - FragDenStaat
Das Portal für Informationsfreiheit für Bürger, Initiativen und Vereine. Stellen Sie eine IFG-Anfrage nach Behördendokumenten, die für Sie und Ihr Engagement wichtig sind! Informieren Sie sich über In...
10376149
Reposted by Stefan Baack
👻Notorischer User dieser Plattform🎃 @bildoperationen.bsky.social · 09/02/2025
«By moving fast and breaking things, DOGE forces a collapse of the system where unanswered questions are met with technological solutions. Shifting the conversation to the technical is a way of locking policymakers and the public out of decisions and shifting that power to the code they write.»
03810
Reposted by Stefan Baack
404 Media @404media.co · 05/02/2025
You can’t post your way out of fascism Authoritarians and tech CEOs now share the same goal: to keep us locked in an eternal doomscroll instead of organizing against them 🔗 www.404media.co/you-cant-pos...
404media.co
You Can’t Post Your Way Out of Fascism
Authoritarians and tech CEOs now share the same goal: to keep us locked in an eternal doomscroll instead of organizing against them, Janus Rose writes.
11661462605
Reposted by Stefan Baack
Auschwitz Memorial @auschwitzmemorial.bsky.social · 27/01/2025
Auschwitz was at the end of a long process. It did not start from gas chambers. This hatred was gradually developed by humans. From ideas, words, stereotypes & prejudice through legal exclusion, dehumanization & escalating violence... to systematic and industrial murder. Auschwitz took time.
A bird's-eye view of a former Auschwitz II-Birkenau camp showing a wide dirt pathway flanked by parallel rows of barbed-wire fences. Groups of visitors walk along the path, surrounded by the remnants of brick structures and barracks, now reduced to foundations. Green grass contrasts with the somber history of the site, as the path leads toward a guard tower in the distance.
10505284822435
Reposted by Stefan Baack
Eryk Salvaggio @eryk.bsky.social · 06/12/2024
“AI is fake and sucks” vs “AI is real and dangerous” is a Twitter argument. In reality I think the debate also has a lot of “AI is real but not for how you’re using it,” to “AI is fake and that is dangerous,” to “things are happening to real people because of AI hype and that should stop.”
320433
Stefan Baack @sbaack.com · 03/12/2024
My reading for this week, delivered to me by the great @aschrock.bsky.social themself! Thank you, looking forward to reading :-)
141
Reposted by Stefan Baack
Critical Digital Media @digital.therourke.net · 03/12/2024
Labelers training AI say they're overworked, underpaid and exploited by big American tech companies
cbsnews.com
Labelers training AI say they're overworked, underpaid and exploited by big American tech companies
Digital workers in Kenya had to sift through horrific online content to train AI, but say they were underpaid, overworked, and got inadequate mental health support. So they're fighting back.
1125
Reposted by Stefan Baack
Daniel Drepper @danieldrepper.bsky.social · 03/12/2024
Dieser Report gibt Hoffnung! Immer mehr neue, ambitionierte Medien haben sich in Deutschland und Europa gegründet. Medien mit dem Ziel, die Öffentlichkeit hochwertig zu informieren. @netzwerkrecherche.org hat für den „Journalism Value Report“ 174 Medien in 31 Ländern befragt und kann zeigen:
13717
Reposted by Stefan Baack
klaudia jaźwińska @klaudia.bsky.social · 27/11/2024
I have a new piece out with @aisvarya17.bsky.social in @columjournreview.bsky.social in which we test how OpenAI's new search feature surfaces and attributes news content. Our findings were not promising for news publishers (1/9) www.cjr.org/tow_center/h...
cjr.org
How ChatGPT (Mis)represents Publisher Content
ChatGPT search — which is positioned as a competitor to search engines like Google and Bing — launched with a press release from OpenAI touting claims that the company had “collaborated extensively wi...
817584
Reposted by Stefan Baack
Thandi Smith @thandis.bsky.social · 20/11/2024
“Without facts, you can’t have truth, and without truth, you can’t have trust”. - Maria Ressa, 2021 Nobel Peace Prize
022
Reposted by Stefan Baack
MwahahahahahadScientist @mads100tist.bsky.social · 14/11/2024
The Onion should buy Elsevier next
5753541576
Stefan Baack @sbaack.com · 21/02/2024
It ended well though. He got the job, and still has it. We met recently 😅
010
Stefan Baack @sbaack.com · 21/02/2024
I still remember when a friend asked for advice about getting a job I intended to apply for
120
Stefan Baack @sbaack.com · 06/02/2024
Long term, there should be less reliance on sources like Common Crawl and a bigger emphasis on training generative AI on datasets created and curated by people in equitable and transparent ways (10/10)
020
Stefan Baack @sbaack.com · 06/02/2024
A key issue is that filtered Common Crawl versions are not updated after their original publication to take feedback and criticism into account. Therefore, we need dedicated intermediaries tasked with filtering Common Crawl in transparent and accountable ways that are continuously updated (9/10)
110
Stefan Baack @sbaack.com · 06/02/2024
AI builders should put more effort into filtering Common Crawl, establish industry standards and best practices for end-user products to reduce potential harms when using Common Crawl or similar sources for training data (8/10)
120
Stefan Baack @sbaack.com · 06/02/2024
Both Common Crawl and AI builders can help making generative AI less harmful. Common Crawl should highlight the limitations and biases of its data, be more transparent and inclusive about its governance, and enforce more transparency by requiring AI builders to attribute using Common Crawl (7/10)
120
Stefan Baack @sbaack.com · 06/02/2024
Due to Common Crawl’s deliberate lack of curation, AI builders need to filter it with care, but such care is often lacking. Popular filtered versions like C4 are especially problematic as the filtering techniques used to create them are simplistic and leave lots of harmful content untouched (6/10)
120
Stefan Baack @sbaack.com · 06/02/2024
In addition, relevant domains like Facebook and the New York Times block Common Crawl from crawling most (or all) of their pages. These blocks are increasing, creating new biases in the crawled data www.wired.com/story/most-n... (5/10)
wired.com
Most Top News Sites Block AI Bots. Right-Wing Media Welcomes Them
Nearly 90 percent of top news outlets like 'The New York Times' now block AI data collection bots from OpenAI and others. Leading right-wing outlets like NewsMax and Breitbart mostly permit them.
120
Stefan Baack @sbaack.com · 06/02/2024
Common Crawl archive is massive, but far from being a “copy of the internet.” Its crawls are automated to prioritize pages on domains that are frequently linked to, making digitally marginalized communities less likely to be included. Moreover, most captured content is English (4/10)
120
Stefan Baack @sbaack.com · 06/02/2024
Using Common Crawl's data does not easily align with trustworthy and responsible AI development because Common Crawl deliberately does not curate its data. It doesn't remove hate speech, for example, because it wants its data to be useful for researchers studying hate speech (3/10)
140
Stefan Baack @sbaack.com · 06/02/2024
Common Crawl is created by a small nonprofit of the same name founded in 2007. Its mission is to level the playing field for technology development by giving free access to data that only companies like Google used to have. Proving data for AI training has never been its primary goal (2/10)
120
Stefan Baack @sbaack.com · 06/02/2024
Most #generativeAI models were trained on Common Crawl, a massive archive of web crawl data. Yet most people never heard of it. My new research studies Common Crawl in-depth and highlights its influence on LLM research and development foundation.mozilla.org/en/research/... (1/10)
foundation.mozilla.org
Training Data for the Price of a Sandwich: Common Crawl’s Impact on Generative AI
Mozilla research finds that Common Crawl's outsized role in the generative AI boom has improved transparency and competition, but is also contributing to biased and opaque generative AI models.
12716
Stefan Baack @sbaack.com · 24/01/2024
Same problem. In my case a conference forced me to do it. I want a world where people can only submit .txt files to conferences and it's the organizers' job to make stuff look coherent
010
Reposted by Stefan Baack
Dr Abeba Birhane @abeba.blacksky.app · 21/11/2023
me and @zeerak.bsky.social wrote about if decolonial AI is at all possible www.elgaronline.com/edcollchap/book…
elgaronline.com
23424
Stefan Baack @sbaack.com · 10/08/2023
Generative AI is shaped by the values, practices and objectives of people. I wrote an explainer showing how to help demystifying the tech and focus on questions of accountability foundation.mozilla.org/en/blog/the-…
041