Sign in

Luca

@sciencialab.com
35 followers 124 following 63 posts
PostsRepliesMedia
Luca @sciencialab.com · 20h
We are pleased to announce that pdfalto, the PDF → ALTO XML converter behind GROBID, is now on PyPI 🎉 pip install pdfalto Binary bundled in the wheels, no manual setup, plus a Python API. pypi.org/project/pdf... github.com/kermitt2/...
pypi.org
pdfalto · PyPI
PDF to ALTO XML conversion — Python bindings for the pdfalto command line tool
000
Luca @sciencialab.com · 23/09/2026
At #TPDL2026 in Faro ☀️: 🔹 Today, Michael Paris presents our Common Crawl paper on what web crawlers actually see 🔹 Thu afternoon I demo the INRIA DataLake: HAL papers → structured knowledge + software mentions Come say hi! #DigitalLibraries #CommonCrawlFoundation #Grobid
011
Luca @sciencialab.com · 23/08/2026
Python client release note: github.com/grobidOrg... And.. Grobid 0.9.1 release note (in case you missed it) github.com/grobidOrg... 🖖 2/2
github.com
Release v0.2.0 · grobidOrg/grobid-client-python
What's Changed Add type hints and py.typed marker (PEP 561) by @lfoppiano in #112 Fixed offsets crash on Markdown transformation: removing the html.unescape and repetitive… by @Sanakhamassi in...
010
Luca @sciencialab.com · 23/08/2026
After releasing Grobid 0.9.1 we released the companion Python client optimised for processing PDF documents at scale. What's new: - Stream PDFs directly from ZIP archives — local or S3, no unzip needed - In-memory processing, lighter on disk I/O - Flexible file selection with glob patterns 1/2
100
Luca @sciencialab.com · 07/08/2026
Wanna join the community? - Announcement mailing list: grobid.netlify.app/mail - Discord: grobid.netlify.app/d... ... and don't forget to star us on Github #GROBID #NLP #TextMining #PDF #OpenScience #TDM #MachineLearning 3/3
010
Luca @sciencialab.com · 07/08/2026
🔗 Reworked author–affiliation linking mechanism, up to +8% F1 score; and more: rootless Docker (k8s/OpenShift-ready), DeLFT 0.4.6, various fixes 👉 github.com/kermitt2/... 2/3
github.com
Release 0.9.1 · grobidOrg/grobid
What's Changed Added Training web API to list ongoing trainings and interrupt a running one, releasing the per-model lock #1483 Development/model debug API exposing the first-level models for ...
110
Luca @sciencialab.com · 07/08/2026
🚀 GROBID 0.9.1 is out! 📄 Large PDF parsing (pdfalto 0.6.2): reduced bounded memory e.g. on large docs including dissertation of 1Gb size. 🔒 Security: command-injection, zip-slip, Jackson, OpenNLP 2.5 & Jetty CVEs fixed 📊 Monitoring: Prometheus/OTLP metrics + Grafana docs, 1/3
100
Luca @sciencialab.com · 22/07/2026
Thank to every contributor and user 🙏 for the support! github.com/grobidOrg... 2/2
github.com
GitHub - grobidOrg/grobid: A machine learning software for extracting information from scholarly documents
A machine learning software for extracting information from scholarly documents - grobidOrg/grobid
010
Luca @sciencialab.com · 22/07/2026
GROBID just passed 5,000 GitHub stars ⭐ From a 2008 hobby project to the engine behind Semantic Scholar, OpenAlex, ResearchGate, Internet Archive Scholar, CERN, and more. 18 years of steady open-source development. No hype, just reliable PDF→structured data at scale. 1/2
100
Luca @sciencialab.com · 21/05/2026
- Paper: arxiv.org/abs/2512.1... - Models: huggingface.co/colle... - Dataset: huggingface.co/colle... (including English cleaned text only, TEI-XML, JSON, Markdown) 4/4
arxiv.org
SciLaD: A Large-Scale, Transparent, Reproducible Dataset for...
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing...
000
Luca @sciencialab.com · 21/05/2026
Finally, I was positively surprised to see that such a large number of people using and talking about #Grobid. 3/4
100
Luca @sciencialab.com · 21/05/2026
Moreover, we met in person after working remotely for a few years. Tall people don't look that tall on video conference. 2/4
100
Luca @sciencialab.com · 21/05/2026
Great experience at #LREC2026! With colleagues and friends from DFKI And Common Crawl, we presented our paper on the SciLaD 🥗 corpus and models arxiv.org/abs/2512.1... 1/4
100
Luca @sciencialab.com · 07/05/2026
5/ ⭐ One last thing: if you find GROBID useful, please star us on GitHub — it goes a long way in helping the project grow. → github.com/grobidOrg... Come help shape what's next. 🙌 5/5
github.com
GitHub - grobidOrg/grobid: A machine learning software for extracting information from scholarly documents · GitHub
A machine learning software for extracting information from scholarly documents - grobidOrg/grobid
000
Luca @sciencialab.com · 07/05/2026
4/ 🎨 And finally — we're refreshing the GROBID logo, and you get to pick. Vote here → forms.gle/aGDNma9Qzn... 4/5
forms.gle
GROBID Logo
We are looking for selecting the GROBID logo. Something not serious and slightly stereotypical French.
100
Luca @sciencialab.com · 07/05/2026
3/ 💬 Prefer real-time chat? Join our new Discord server to hang out with users, contributors, and maintainers. Invite → discord.gg/yuEaC4tYnz 3/5
discord.gg
Grobid
Check out the Grobid community on Discord - hang out with 19 other members and enjoy free voice and text chat.
100
Luca @sciencialab.com · 07/05/2026
2/ 📬 We've launched an official mailing list (EN/FR) for announcements, discussions, and Q&A. Subscribe → groupes.renater.fr/s... 2/5
100
Luca @sciencialab.com · 07/05/2026
1/ 📣 Three GROBID community updates in one go: 📬 New mailing list 💬 New Discord server 🎨 Logo vote is open Details below 👇 1/5
100
Luca @sciencialab.com · 22/04/2026
This builds on the foundational harvesting work by Patrice Lopez & James Howison (SoftCite project), and is a collaboration with @DFKI, @HUBerlin, @CommonCrawl & Uni Mannheim. Attending LREC? Let's connect!👋 #NLP #ScientificNLP #MultilingualNLP #SciLaD #ScienciaLAB #grobid 5/5
001
Luca @sciencialab.com · 22/04/2026
To validate the data quality, we pre-trained SciLaD-M — a RoBERTa encoder trained from scratch on SciLaD. Result: performance comparable to or better than SciBERT across scientific NLP benchmarks, using only openly available data. Open data → open models. 💡 4/5
100
Luca @sciencialab.com · 22/04/2026
We also release a curated English split of 10M+ documents — filtered and deduplicated — along with the full open-source pipeline so anyone can reproduce or extend the dataset. No black boxes. Full transparency. 🔓 3/5
100
Luca @sciencialab.com · 22/04/2026
Scientific NLP needs large, open, high-quality data. SciLaD is our answer: a corpus of Open Access publications harvested from Unpaywall, converted into structured TEI XML using open-source tools. 35M+ multilingual publications across languages and scientific domains. 🌍 2/5
100
Luca @sciencialab.com · 22/04/2026
🧵 Excited to share that I'll be presenting SciLaD at #LREC2026! SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing arXiv → arxiv.org/abs/2512.1... 1/5
arxiv.org
SciLaD: A Large-Scale, Transparent, Reproducible Dataset for...
SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing...
100
Luca @sciencialab.com · 14/04/2026
5/ Infrastructure: JDK 21, Gradle 9, TensorFlow 2.17 (Python 3.10–3.11), pdfalto 0.6.0, wapiti 1.5.1, virtualenv/conda support for DeLFT. 6/ Full release notes → github.com/kermitt2/... #GROBID #OpenSource #NLP #ScholarlyInfrastructure 6/6
github.com
Release 0.9.0 · grobidOrg/grobid
What's Changed Added Conflict of interest and author contributions statement extraction in header and segmentation models #1319 Extract figures, tables and equations from back/annex sections #...
020
Luca @sciencialab.com · 14/04/2026
4/ New pluggable NLP engines: Lingua for language identification, Blingfire for sentence segmentation — both available as drop-in alternatives to the existing defaults. 5/6
100
Luca @sciencialab.com · 14/04/2026
3/ Revised Crossref consolidation, developed in collaboration with the Crossref team — improved rate limit handling, better error recovery, and more robust reference matching. 4/6
110
Luca @sciencialab.com · 14/04/2026
Multi-architecture Docker images (amd64 + arm64) are now available, enabling native deployment on Apple Silicon and ARM cloud instances. 3/6
100
Luca @sciencialab.com · 14/04/2026
1/ New extraction coverage: conflict of interest & author contribution statements; figures, tables and equations from back/annex sections; URLs from PDF annotations; ORCID identifiers fetched via Crossref when absent from the source document. 2/ Native Linux ARM64 support. 2/6
110
Luca @sciencialab.com · 14/04/2026
I'm please to announce that GROBID 0.9.0 is out. Release notes and highlights below. 1/6
100
Luca @sciencialab.com · 31/03/2026
If you've ever lost time reformatting citations by hand, this one's for you. sciencialab.github.i... 7/7
000
Luca @sciencialab.com · 31/03/2026
Results are automatically extracted, so always give the output a quick review before dropping it into your .bib file — especially on older or inconsistently formatted references. 6/7
100
Luca @sciencialab.com · 31/03/2026
Under the hood it's powered by #Grobid, a battle-tested machine learning library for extracting structured data from scientific documents. The same technology used to process millions of PDFs at scale — now doing one job really well, in one click. 5/7
100
Luca @sciencialab.com · 31/03/2026
No tab-switching, no copy-pasting, no reformatting. github.com/user-atta... 4/7
100
Luca @sciencialab.com · 31/03/2026
But the real magic is the bookmarklet. Drag it once to your bookmarks bar. Then, on any webpage — a journal site, a preprint server, a bibliography — highlight the citation, click the bookmarklet, and the BibTeX is already in your clipboard. 3/7
100
Luca @sciencialab.com · 31/03/2026
We built a simple tool to fix exactly this: text2bibtex. The app is simple: paste any plain-text citation and get a clean, ready-to-use BibTeX entry back. github.com/user-atta... 2/7
100
Luca @sciencialab.com · 31/03/2026
You found the paper. Now where's the BibTeX? 🧵 It happens every time! You track down the right reference — some older journal article, a conference paper, a preprint — scroll down to "Cite this paper", and get… a plain text citation. 1/7
100
Reposted by Luca
Common Crawl Foundation @commoncrawl.bsky.social · 02/03/2026
@thomvaughan.bsky.social did a WCAG colour contrast audit of 240 top domains using Common Crawl's February 2026 archive. You can read more about his study in the thread below
031
Reposted by Luca
Common Crawl Foundation @commoncrawl.bsky.social · 24/02/2026
We're happy to announce the release of the Web Graphs for December 2025 and January/February 2026, consisting of 288.6 million nodes and 12.4 billion edges at the host level, and 134.2 million nodes and 5.4 billion edges at the domain level. www.commoncrawl.org/blog/host--a...
commoncrawl.org
Common Crawl - Blog - Host- and Domain-Level Web Graphs December 2025 and January/February 2026
We're happy to announce the release of the Web Graphs for December 2025 and January/February 2026, consisting of 288.6 million nodes and 12.4 billion edges at the host level, and 134.2 million nodes a...
032
Luca @sciencialab.com · 08/12/2025
We tried to carry on without the sharp mind of Patrice Lopez, who is currently immersed in new and exciting challenges. And on a lighter note, we might finally have a logo for Grobid! 4/4
000
Luca @sciencialab.com · 08/12/2025
It also enabled us to better coordinate our roadmap, integrating the benefits of advanced Large Language Models while preserving our commitment to enabling processing on standard consumer-grade hardware. 3/4
100
Luca @sciencialab.com · 08/12/2025
The meeting was extremely productive and helped us consolidate our community-driven vision for Grobid’s development in the coming years. 2/4
100
Luca @sciencialab.com · 08/12/2025
On the 26-27 November we held the #Grobid Camp at the Centre de #Inria Paris. The goal was to have a meeting with the major players in the French community which spaces from government institutes, to companies and large scale projects. 1/4
110
Luca @sciencialab.com · 08/12/2025
The meeting was extremely productive and helped us consolidate our community-driven vision for Grobid’s development in the coming years. 2/4
000
Luca @sciencialab.com · 01/11/2025
Me: Please, fix the tests! LLM Agent: OK
010
Luca @sciencialab.com · 24/08/2025
The Safari browser is like a car with one gear that claim it does not pollute...
000
Luca @sciencialab.com · 28/07/2025
Exactly! There is a common misconception that by throwing any kind of crap into a vector it will magically work. Still at the age of AI, metadata information cannot still be ignored.
010
Reposted by Luca
Eric Topol @erictopol.bsky.social · 01/07/2025
Yes. The time is now. Vaccines to treat and prevent cancer. www.jci.org/articles/vie...
432680
Luca @sciencialab.com · 18/05/2025
Your feedback will help us improve Grobid! 🌟 Feel free to share your thoughts, star us on GitHub, and let’s keep building! 💬🚀
000
Luca @sciencialab.com · 18/05/2025
Next up, we're focusing on supporting more platforms (Linux ARM), improving figures and tables extraction, enhancing CJK language support, and providing better handling for more document types like theses, reports, and more. 🔽
100
Luca @sciencialab.com · 18/05/2025
- 🔤 Improved recognition of non-standard fonts - 🛠️ Various bug fixes and security vulnerabilities addressed github.com/kermitt2/... 🔽
github.com
Release 0.8.2 · kermitt2/grobid
What's Changed Added New model specialization/variants (flavors) mechanism #1151 Specialization/variant process for a lightweight processing that covers other types of scientific articles that...
110