Reposted by Kirill Solovev
Here's my new blogpost on the dataset I have created alongside @ksolovev.com: we decided to solve the issue of data availability for research on news articles by processing, cleaning, tagging and indexing the whole news subset of the Common Crawl, resulting in
cs2.uni-graz.at/blog/infini-...