Sign in

maxschulze.bsky.social

@maxschulze.bsky.social
16 followers 249 following 1 posts
PostsRepliesMedia
Reposted by @maxschulze.bsky.social
Mushtaq Bilal, PhD @mushtaqbilalphd.bsky.social · 07/02/2025
Meta illegaly downloaded 80+ terabytes of books from LibGen, Anna's Archive, and Z-library to train their AI models. In 2010, Aaron Swartz downloaded only 70 GBs of articles from JSTOR (0.0875% of Meta). Faced $1 million in fine and 35 years in jail. Took his own life in 2013.
“Torrenting from a corporate laptop doesn’t feel right”: Meta emails unsealedA photo of Aaron Swartz (1986-2013) when he was 19.
Last month, Meta admitted to torrenting a controversial large dataset known as LibGen, which includes tens of millions of pirated books. But details around the torrenting were murky until yesterday, when Meta's unredacted emails were made public for the first time. The new evidence showed that Meta torrented "at least 81.7 terabytes of data across multiple shadow libraries through the site Anna’s Archive, including at least 35.7 terabytes of data from Z-Library and LibGen," the authors' court filing said. And "Meta also previously torrented 80.6 terabytes of data from LibGen."
5174783989
Reposted by @maxschulze.bsky.social
Ketan Joshi @ketanjoshi.co · 28/12/2024
OpenAI doesn't report a single number for its climate impacts - not anywhere, in any way, shape or form. It sounds kind of obvious but the fact that a company with such incredible electricity hunger isn't disclosing basic information about its energy consumption and associated emissions is v bad..
9157861531
Reposted by @maxschulze.bsky.social
Megan Carpentier @megancarpentier.bsky.social · 05/12/2024
This is one of the most well-written explanations of certain kinds of AI errors I’ve seen to date. (The whole article is great too.) www.theverge.com/c/24300623/a...
When you pose the question to a language model, what you are really asking is, “Given the statistical distribution of words in the vast public corpus of text, what are the words most likely to follow the sequence ‘what country is to the south of Rwanda?’” Even if the system responds with the word “Burundi,” this is a different sort of assertion with a different relationship to reality than the human’s answer, and to say the AI “knows” or “believes” Burundi to be south of Rwanda is a category mistake that will lead to errors and confusion.
502124782
Reposted by @maxschulze.bsky.social
Princeton University Press @princetonupress.bsky.social · 02/12/2024
"The way we think about technology is shaped by the tech companies themselves," says @marietjeschaake.bsky.social to Andrew Anthony in an interview for @observeruk.bsky.social. Read the full discussion: www.theguardian.com/technology/2...
theguardian.com
AI expert Marietje Schaake: ‘The way we think about technology is shaped by the tech companies themselves’
The Dutch policy director and former MEP on the unprecedented reach of big tech, the need for confident governments, and why the election of Trump changes everything
03216