Sign in

sebnagel.bsky.social

@sebnagel.bsky.social
68 followers 32 following 0 posts
PostsRepliesMedia
Reposted by @sebnagel.bsky.social
Common Crawl Foundation @commoncrawl.bsky.social · 27/05/2026
RSVP and join speakers @very-laurie.bsky.social and @pjox.bsky.social from the Common Crawl Foundation and Kostis Saitas Zarkias and Robert Pugh from Mozilla Data Collective for a truly hands-on session. Thursday, June 4th 6 PM CEST | 12 PM ET | 9 AM PDT Register via Zoom: zoom.us/meeting/regi...
zoom.us
Welcome! You are invited to join a meeting: Text Language Identification (LID) with CommonCrawl and Mozilla Data Collective. After registering, you will receive a confirmation email about joining the ...
Welcome! You are invited to join a meeting: Text Language Identification (LID) with CommonCrawl and Mozilla Data Collective. After registering, you will receive a confirmation email about joining the ...
042
Reposted by @sebnagel.bsky.social
Common Crawl Foundation @commoncrawl.bsky.social · 01/04/2026
Common Crawl is switching to reporting dataset sizes in nibbles. As an organisation dedicated to data preservation, we feel it would be remiss to allow this underrepresented unit to fall out of use. Our latest crawl now exceeds 689 tebibbles. commoncrawl.org/blog/announc...
commoncrawl.org
Common Crawl - Blog - Announcing a Change to Common Crawl Dataset Size Reporting
Common Crawl is switching to reporting dataset sizes in nibbles. As an organisation dedicated to data preservation, we feel it would be remiss to allow this underrepresented unit to fall out of use. O...
051
Reposted by @sebnagel.bsky.social
Common Crawl Foundation @commoncrawl.bsky.social · 10/02/2026
Language identification still proves to be a challenging task, especially for web data. In collaboration with @mlcommons.org @eleutherai.bsky.social @jhu.edu and 97 community members, we created CommonLID, a new benchmark for LangID for 100+ languages!
Examples of mislabeled web text by existing LangID systems. A full text version is available on the blog post below.Examples of mislabeled web text by existing LangID systems. A full text version is available on the blog post below.
1105
Reposted by @sebnagel.bsky.social
eleutherai.bsky.social @eleutherai.bsky.social · 13/02/2026
Announcing our latest paper: CommonLID In collaboration with @commoncrawl.bsky.social @mlcommons.org @jhu.edu we built a LID benchmark on actual Common Crawl text covering 109 languages. Existing evaluations overestimate how well LangID works on web data. arxiv.org/abs/2601.18026
arxiv.org
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data of...
12212
Reposted by @sebnagel.bsky.social
Common Crawl Foundation @commoncrawl.bsky.social · 02/10/2025
Common Crawl’s Web Languages initiative has had many contributions since its introduction. We’re calling for native speakers of certain languages to review language contributions, to ensure that links we’re adding to our seed crawl are of good quality. commoncrawl.org/blog/web-lan...
commoncrawl.org
Common Crawl - Blog - Web Languages Needing Review by Native Speakers
Common Crawl’s Web Languages initiative has had many contributions since its introduction. We’re calling for native speakers of certain languages to review language contributions, to ensure that links...
024
Reposted by @sebnagel.bsky.social
Common Crawl Foundation @commoncrawl.bsky.social · 21/07/2025
"MOIC will also partner with Common Crawl, one of the largest free and open repositories of web crawled data. MOIC will fund work at Common Crawl, leveraging native speakers to annotate and seed European language data in the publicly available Common Crawl data set."
blogs.microsoft.com
Unlocking data to advance European commerce and culture - Microsoft On the Issues
Microsoft launches 2 initiatives to open Europe’s languages and culture, building on AI, cloud, and digital sovereignty commitments.
031
Reposted by @sebnagel.bsky.social
Common Crawl Foundation @commoncrawl.bsky.social · 21/07/2025
The Common Crawl Foundation, MLCommons, EleutherAI, and John Hopkins' Center for Language and Speech Processing have the pleasure of inviting you to register for the 1st shared task on Language Identification for web data. commoncrawl.org/blog/wmdqs-s...
commoncrawl.org
Common Crawl - Blog - WMDQS Shared Task on Language Identification
The Common Crawl Foundation, MLCommons, EleutherAI, and John Hopkins' Center for Language and Speech Processing have the pleasure of inviting you to register for the 1st shared task on Language Identi...
065
Reposted by @sebnagel.bsky.social
Apache Software Foundation (The ASF) @apache.org · 31/07/2025
Apache Nutch 1.21 is now available for download: buff.ly/juTZlwE Nutch is a well matured, production ready #web crawler. #opensource
032
Reposted by @sebnagel.bsky.social
Apache Software Foundation (The ASF) @apache.org · 03/06/2025
[NEWS] The Apache Software Foundation Announces Two New Top-Level Projects buff.ly/09CyVxC
news.apache.org
The Apache Software Foundation Announces Two New Top-Level Projects - The Apache Software Foundation Blog
Newest TLPs enable seamless management for data and AI workloads Wilmington, DE –  June 3, 2025 – The Apache Software Foundation (ASF), the global home of open source software the world relies on,…
042
Reposted by @sebnagel.bsky.social
Mia Nahrgang @mianahrgang.bsky.social · 28/05/2025
🚨🚀 Looking for a comparative dataset on social media platforms? We’re excited to launch COMPARE! This is a collaborative effort by @nilsweidmann.bsky.social , @friederikeq.bsky.social , @sebnagel.bsky.social , @yannistheocharis.bsky.social & Molly Roberts. 🧵⤵️ (1/5)
1134
Reposted by @sebnagel.bsky.social
Common Crawl Foundation @commoncrawl.bsky.social · 29/05/2025
Call for papers! We are organising the 1st Workshop on Multilingual Data Quality Signals with @mlcommons.org and @eleutherai.bsky.social, held in tandem with @colmweb.org. Submit your research on multilingual data quality! Submission deadline is 23 June, more info: wmdqs.org
wmdqs.org
1st Workshop on Multilingual Data Quality Signals
098
Reposted by @sebnagel.bsky.social
Universität Konstanz @uni-konstanz.de · 22/05/2025
The #UniKonstanz Cluster of Excellence "The Politics of Inequality @excinequality.bsky.social will continue to receive funding in the context of the German #ExcellenceStrategy @dfg.de @wissenschaftsrat.de. Details: t1p.de/ouhyj
27016
Reposted by @sebnagel.bsky.social
Common Crawl Foundation @commoncrawl.bsky.social · 11/12/2024
commoncrawl.org
Common Crawl - Blog - Expanding the Language and Cultural Coverage of Common Crawl
We aim to enhance linguistic diversity in our dataset by inviting community contributions of non-English URLs and collaborating with MLCommons on a Language Identification campaign.
062