Sign in

Luca

@sciencialab.com
35 followers 124 following 63 posts
PostsRepliesMedia
Luca @sciencialab.com · 14h
We are pleased to announce that pdfalto, the PDF → ALTO XML converter behind GROBID, is now on PyPI 🎉 pip install pdfalto Binary bundled in the wheels, no manual setup, plus a Python API. pypi.org/project/pdf... github.com/kermitt2/...
pypi.org
pdfalto · PyPI
PDF to ALTO XML conversion — Python bindings for the pdfalto command line tool
000
Luca @sciencialab.com · 23/09/2026
At #TPDL2026 in Faro ☀️: 🔹 Today, Michael Paris presents our Common Crawl paper on what web crawlers actually see 🔹 Thu afternoon I demo the INRIA DataLake: HAL papers → structured knowledge + software mentions Come say hi! #DigitalLibraries #CommonCrawlFoundation #Grobid
011
Luca @sciencialab.com · 23/08/2026
After releasing Grobid 0.9.1 we released the companion Python client optimised for processing PDF documents at scale. What's new: - Stream PDFs directly from ZIP archives — local or S3, no unzip needed - In-memory processing, lighter on disk I/O - Flexible file selection with glob patterns 1/2
100
Luca @sciencialab.com · 07/08/2026
🚀 GROBID 0.9.1 is out! 📄 Large PDF parsing (pdfalto 0.6.2): reduced bounded memory e.g. on large docs including dissertation of 1Gb size. 🔒 Security: command-injection, zip-slip, Jackson, OpenNLP 2.5 & Jetty CVEs fixed 📊 Monitoring: Prometheus/OTLP metrics + Grafana docs, 1/3
100
Luca @sciencialab.com · 22/07/2026
GROBID just passed 5,000 GitHub stars ⭐ From a 2008 hobby project to the engine behind Semantic Scholar, OpenAlex, ResearchGate, Internet Archive Scholar, CERN, and more. 18 years of steady open-source development. No hype, just reliable PDF→structured data at scale. 1/2
100
Luca @sciencialab.com · 21/05/2026
Great experience at #LREC2026! With colleagues and friends from DFKI And Common Crawl, we presented our paper on the SciLaD 🥗 corpus and models arxiv.org/abs/2512.1... 1/4
100
Luca @sciencialab.com · 07/05/2026
1/ 📣 Three GROBID community updates in one go: 📬 New mailing list 💬 New Discord server 🎨 Logo vote is open Details below 👇 1/5
100
Luca @sciencialab.com · 22/04/2026
🧵 Excited to share that I'll be presenting SciLaD at #LREC2026! SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing arXiv → arxiv.org/abs/2512.1... 1/5
arxiv.org
SciLaD: A Large-Scale, Transparent, Reproducible Dataset for...
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing...
100
Luca @sciencialab.com · 14/04/2026
I'm please to announce that GROBID 0.9.0 is out. Release notes and highlights below. 1/6
100
Luca @sciencialab.com · 31/03/2026
You found the paper. Now where's the BibTeX? 🧵 It happens every time! You track down the right reference — some older journal article, a conference paper, a preprint — scroll down to "Cite this paper", and get… a plain text citation. 1/7
100
Reposted by Luca
Common Crawl Foundation @commoncrawl.bsky.social · 02/03/2026
@thomvaughan.bsky.social did a WCAG colour contrast audit of 240 top domains using Common Crawl's February 2026 archive. You can read more about his study in the thread below
031
Reposted by Luca
Common Crawl Foundation @commoncrawl.bsky.social · 24/02/2026
We're happy to announce the release of the Web Graphs for December 2025 and January/February 2026, consisting of 288.6 million nodes and 12.4 billion edges at the host level, and 134.2 million nodes and 5.4 billion edges at the domain level. www.commoncrawl.org/blog/host--a...
commoncrawl.org
Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026
We're happy to announce the release of the Web Graphs for December 2025 and January/February 2026, consisting of 288.6 million nodes and 12.4 billion edges at the host level, and 134.2 million nodes a...
032
Luca @sciencialab.com · 08/12/2025
On the 26-27 November we held the #Grobid Camp at the Centre de #Inria Paris. The goal was to have a meeting with the major players in the French community which spaces from government institutes, to companies and large scale projects. 1/4
110
Luca @sciencialab.com · 01/11/2025
Me: Please, fix the tests! LLM Agent: OK
010
Luca @sciencialab.com · 24/08/2025
The Safari browser is like a car with one gear that claim it does not pollute...
000
Luca @sciencialab.com · 28/07/2025
Exactly! There is a common misconception that by throwing any kind of crap into a vector it will magically work. Still at the age of AI, metadata information cannot still be ignored.
010
Reposted by Luca
Eric Topol @erictopol.bsky.social · 01/07/2025
Yes. The time is now. Vaccines to treat and prevent cancer. www.jci.org/articles/vie...
432680
Luca @sciencialab.com · 18/05/2025
Grobid 0.8.2 is out! 🚀 - 🧠 New processing "flavors" for different doc types (e.g. SDO, corrections, editorials) - 🔗 Improved URL extraction - ✅ Better text extraction for paragraphs around figures and tables 🧵🔽
100
Luca @sciencialab.com · 11/05/2025
Dear @github, I wonder whether it would be possible to have a way to save certain "search parameters" inside the issues/pulls so that our work may be framed to important tasks. E.g. working on a specific milestone and wanting to know everything that is not yet done:
020
Reposted by Luca
Marian Dörk @nrchtct.bsky.social · 18/12/2013
GROBID by Patrice Lopez turns messy PDFs into well-structured text in TEI format including references- super useful! github.com/kermitt2/grobid
github.com
GitHub - kermitt2/grobid: A machine learning software for...
A machine learning software for extracting information fr...
101
Reposted by Luca
Hans de Jonge @hldejonge.bsky.social · 10/02/2025
To what extent do researchers funded by Dutch Research Council NWO and ZonMw share the research data and code underlying their publications? Today we published an analysis based on 10.000+ papers using the open source tool Grobid: www.nwo.nl/en/news/shar... All underlying data openly available!
03212
Luca @sciencialab.com · 06/05/2025
Grobid popularity is still growing, despite LM, LLM, LLLM....
000
Luca @sciencialab.com · 22/11/2024
Suggestion not asked. If you don't want advertisements anymore on Twitter, you can pay 200 EUR per year (100 EUR only reduced them by half, lol), or you can pay 0 EUR and 1/2
200
Luca @sciencialab.com · 21/11/2024
Alleluja!! www.nature.com/artic...
nature.com
Superconductivity researcher who committed misconduct exits university
Nature - The University of Rochester has confirmed that it no longer employs Ranga Dias, who was found by investigators to have committed data fabrication.
000
Luca @sciencialab.com · 21/11/2024
Here a few tips for using Bluesky github.com/JefTek/Blues...
github.com
GitHub - JefTek/BlueskyGuide: Collection of Tips & Tricks for collaborating on Bluesky
Collection of Tips & Tricks for collaborating on Bluesky - JefTek/BlueskyGuide
000