Sign in

Common Crawl Foundation

@commoncrawl.bsky.social
420 followers 61 following 135 posts

Common Crawl is a non-profit foundation dedicated to the Open Web.

PostsRepliesMedia
Common Crawl Foundation @commoncrawl.bsky.social · 27/05/2026
Under-represented languages deserve better tools! On June 4th, The Common Crawl Foundation and Mozilla Data Collective will host a webinar to test language identification for the languages you care about.
An image containing the logos of Common Crawl and Mozilla Data Collective. It cites the information of the webinar already mentioned in the post.
120
Common Crawl Foundation @commoncrawl.bsky.social · 10/02/2026
Current benchmarks over-estimate LangID performance on web data. In our evaluations, we show top existing models have < 80% F1, even when limiting to languages the models explicitly support.
A table showing the results of our evaluation of existing LangID systems across 6 different datasets. Full text of the table is available on the paper linked below.
130
Common Crawl Foundation @commoncrawl.bsky.social · 10/02/2026
Language identification still proves to be a challenging task, especially for web data. In collaboration with @mlcommons.org @eleutherai.bsky.social @jhu.edu and 97 community members, we created CommonLID, a new benchmark for LangID for 100+ languages!
Examples of mislabeled web text by existing LangID systems. A full text version is available on the blog post below.Examples of mislabeled web text by existing LangID systems. A full text version is available on the blog post below.
1105
Common Crawl Foundation @commoncrawl.bsky.social · 06/11/2025
Common Crawl celebrates World Digital Preservation Day Nov. 6, which invites the community to unite in answering a powerful question: Why Preserve? commoncrawl.org/blog/common-...
Banner for the World Digital Preservation Day, 6th of November 2025
030