Sign in

IDI

@institutional.org
297 followers 2 following 32 posts

A research center at Harvard working to strengthen society’s connection to knowledge by advancing our access to and understanding of the data that shapes AI.

PostsRepliesMedia
IDI @institutional.org · 08/09/2026
The Institutional Data Initiative at Harvard @institutional.org works with knowledge institutions to increase the availability, diversity, and responsible use of training data for AI. Reach out and join us.
010
IDI @institutional.org · 08/09/2026
Together, these projects establish Institutional Books as an ongoing resource that can grow through continued research and community use. The open-source pipelines can be adapted to new collections, models, languages, and questions: institutional.org/posts/extend...
institutional.org
Extending Institutional Books with enriched text and 22M extracted images | IDI
The Institutional Data Initiative at the Harvard Law School Library expands the ~1M-book Institutional Books dataset with paragraph-level text analysis and 22M extracted images, enabling new forms of ...
143
IDI @institutional.org · 08/09/2026
Visual Elements extends the dataset to images embedded in scanned pages, producing 22.6M visual elements – including illustrations, photographs, charts, musical notation, and decorative materials – linked to their source volumes and pages.
120
IDI @institutional.org · 08/09/2026
Enriched Text contains 217B tokens from 983K @harvardlibrary.bsky.social volumes, organized into 1.39B annotated paragraphs across 250 languages. Researchers can use the dataset to analyze language, repeated passages, line breaks, and paragraph predictability.
140
IDI @institutional.org · 08/09/2026
Institutional Books is expanding. The Institutional Data Initiative @institutional.org at @harvardlawadmin.bsky.social Library has added paragraph-level text analysis and 22M extracted images to our work on ∼1M books, opening new possibilities for research, experimentation, and discovery. 🧵
1288
IDI @institutional.org · 08/09/2026
The Institutional Data Initiative at Harvard (@instdin) works with knowledge institutions to increase the availability, diversity, and responsible use of training data for AI. Reach out and join us.
000
IDI @institutional.org · 08/09/2026
Together, these projects establish Institutional Books as an ongoing resource that can grow through continued research and community use. The open-source pipelines can be adapted to new collections, models, languages, and questions: institutional.org/posts/extend...
institutional.org
Extending Institutional Books with enriched text and 22M extracted images | IDI
The Institutional Data Initiative at the Harvard Law School Library expands the ~1M-book Institutional Books dataset with paragraph-level text analysis and 22M extracted images, enabling new forms of ...
100
IDI @institutional.org · 08/09/2026
Visual Elements extends the dataset to images embedded in scanned pages, producing 22.6M visual elements – including illustrations, photographs, charts, musical notation, and decorative materials – linked to their source volumes and pages.
100
IDI @institutional.org · 08/09/2026
Enriched Text contains 217B tokens from 983K @HarvardLibrary volumes, organized into 1.39B annotated paragraphs across 250 languages. Researchers can use the dataset to analyze language, repeated passages, line breaks, and paragraph predictability.
100
IDI @institutional.org · 02/09/2026
The Institutional Data Initiative at the @hls.harvard.edu Library works with knowledge institutions to increase the availability, diversity, and responsible use of training data for AI. Reach out and join us.
000
IDI @institutional.org · 02/09/2026
The dataset supports more than keyword search: you can follow reading order, identify people and places, and examine how newspaper design and advertising changed across time. Institutional Newspapers includes the dataset, models, and an open-source processing pipeline:
institutional.org
Making 1.47 Million Historical Newspaper Pages Available as High-Quality, Granular Data | IDI
The Institutional Data Initiative at the Harvard Law School Library and Boston Public Library release an open dataset and processing pipeline that make historical newspapers more accessible for resear...
120
IDI @institutional.org · 02/09/2026
Instead of treating each page as one image, our pipeline divides it into individual crops. Those crops are processed as articles, advertisements, photographs, cartoons, headings, mastheads, or empty space, with OCR tailored to each segment.
100
IDI @institutional.org · 02/09/2026
Institutional Newspapers contains 1.47 million pages from the Boston Public Library’s collections, representing 16.3 billion tokens and a multilingual record of Boston’s communities across 73 languages.
100
IDI @institutional.org · 02/09/2026
When a newspaper becomes a wall of OCR text, columns collapse, headlines mix with ads, and reading order disappears. The Institutional Data Initiative (@institutional.org) and Boston Public Library (@bpl.boston.gov) built a pipeline and dataset to make historical newspapers more accessible.🧵
institutional.org
Making 1.47 Million Historical Newspaper Pages Available as High-Quality, Granular Data | IDI
The Institutional Data Initiative at the Harvard Law School Library and Boston Public Library release an open dataset and processing pipeline that make historical newspapers more accessible for resear...
132
IDI @institutional.org · 02/09/2026
The Institutional Data Initiative at the @hls.harvard.edu Library works with knowledge institutions to increase the availability, diversity, and responsible use of training data for AI. Reach out and join us.
000
IDI @institutional.org · 02/09/2026
The dataset supports more than keyword search: you can follow reading order, identify people and places, and examine how newspaper design and advertising changed across time. Institutional Newspapers includes the dataset, models, and an open-source processing pipeline:
institutional.org
Making 1.47 Million Historical Newspaper Pages Available as High-Quality, Granular Data | IDI
The Institutional Data Initiative at the Harvard Law School Library and Boston Public Library release an open dataset and processing pipeline that make historical newspapers more accessible for resear...
110
IDI @institutional.org · 02/09/2026
Instead of treating each page as one image, our pipeline divides it into individual crops. Those crops are processed as articles, advertisements, photographs, cartoons, headings, mastheads, or empty space, with OCR tailored to each segment.
100
IDI @institutional.org · 02/09/2026
Institutional Newspapers contains 1.47 million pages from the Boston Public Library’s collections, representing 16.3 billion tokens and a multilingual record of Boston’s communities across 73 languages.
100
IDI @institutional.org · 02/09/2026
The Institutional Data Initiative at the @hls.harvard.edu Library works with knowledge institutions to increase the availability, diversity, and responsible use of training data for AI. Reach out and join us.
000
IDI @institutional.org · 02/09/2026
The dataset supports more than keyword search: you can follow reading order, identify people and places, and examine how newspaper design and advertising changed across time. Institutional Newspapers includes the dataset, models, and an open-source processing pipeline:
100
IDI @institutional.org · 02/09/2026
Instead of treating each page as one image, our pipeline divides it into individual crops. Those crops are processed as articles, advertisements, photographs, cartoons, headings, mastheads, or empty space, with OCR tailored to each segment.
100
IDI @institutional.org · 02/09/2026
Institutional Newspapers contains 1.47 million pages from the Boston Public Library’s collections, representing 16.3 billion tokens and a multilingual record of Boston’s communities across 73 languages.
100
Reposted by IDI
Greg Leppert @leppert.me · 13/08/2026
AI enables new forms of research, but only where information is computationally accessible. Today, @institutional.org and @umich.edu Library are releasing a structured dataset that makes 186 years of U-M’s Board of Regents Proceedings newly accessible. institutional.org/posts/instit...
institutional.org
Institutional Proceedings: 186 years of university decisions. | IDI
IDI and the University of Michigan Library have collaborated to release two centuries of the university’s governance decisions as a structured dataset that can be searched, analyzed, and reused.
111
Reposted by IDI
Greg Leppert @leppert.me · 20/11/2025
Even if you're not a partner library, you might be curious about what it's like to work with GRIN. Our technical report has a wealth of details. arxiv.org/abs/2511.11447
arxiv.org
GRIN Transfer: A production-ready tool for libraries to retrieve digital copies from Google Books
Publicly launched in 2004, the Google Books project has scanned tens of millions of items in partnership with libraries around the world. As part of this project, Google created the Google Return Inte...
001
Reposted by IDI
Greg Leppert @leppert.me · 20/11/2025
We're also sharing the pipeline we developed for Institutional Books that seamlessly dedupes, classifies, and enhances the data once GRIN Transfer brings it down. www.institutional.org/tools
institutional.org
Institutional Books | Institutional Data Initiative
Institutional Books 1.0 is our first release of public domain books. This set was originally digitized through Harvard Library’s participation in the Google Books project..
101
Reposted by IDI
Greg Leppert @leppert.me · 20/11/2025
When libraries join Google Books, Google not only scans their books, it also makes a wealth of image, OCR, & metadata available to them via the Google Return Interface (GRIN). But working with GRIN can be challenging, so we're releasing a tool to make it easier. www.institutional.org/posts/grin-t...
institutional.org
Announcing the release of GRIN Transfer
GRIN Transfer, an open source tool that allows Google Books partner libraries to more easily access their Google Books collection.
153
IDI @institutional.org · 18/09/2025
Join us tomorrow at 10AM EST: tinyurl.com/y3ye6cz6
tinyurl.com
Welcome! You are invited to join a meeting: IDI Talk with Michele Dolfi & Peter Staar on SmolDocling. After registering, you will receive a confirmation email about joining the meeting.
Please join us for a talk with Michele Dolfi & Peter Staar from IBM Research in Zurich to discuss their work on SmolDocling, an “ultra-compact” model for diverse OCR tasks.
000
IDI @institutional.org · 09/09/2025
Register to join the talk virtually: harvard.zoom.us/meeting/regi...
harvard.zoom.us
Welcome! You are invited to join a meeting: IDI Talk with Michele Dolfi & Peter Staar on SmolDocling. After registering, you will receive a confirmation email about joining the meeting.
Please join us for a talk with Michele Dolfi & Peter Staar from IBM Research in Zurich to discuss their work on SmolDocling, an “ultra-compact” model for diverse OCR tasks.
000
IDI @institutional.org · 09/09/2025
Can a small visual language model read documents as effectively as models 27 times its size? Next Friday, IDI will host Michele Dolfi and Peter Staar from IBM Research Zurich to discuss their work on SmolDocling, an “ultra-compact” model for diverse OCR tasks.
100
Reposted by IDI
Greg Leppert @leppert.me · 20/06/2025
This Monday, @institutionaldatainitiative.org will host Petr Knoth to share his experience leading CORE ("The world’s largest collection of open access research papers") as the rise of AI brings new meaning, and challenges, to stewarding knowledge repositories. Join us virtually via the link below.
222
Reposted by IDI
Greg Leppert @leppert.me · 16/06/2025
Cohosted by @institutionaldatainitiative.org and The Berkman Klein Center. harvard.zoom.us/webinar/regi...
harvard.zoom.us
Welcome! You are invited to join a webinar: Open AI Development. After registering, you will receive a confirmation email about joining the webinar.
For AI to truly benefit society, it must be built on foundations of transparency, fairness, and accountability—starting with the most foundational building block that powers it: data. Not long ago, ...
001
IDI @institutional.org · 12/06/2025
We hope Institutional Books will be the beginning of a process that makes millions more books accessible to the public for a variety of uses. We welcome feedback as we continue to expand this dataset, refine its contents, and sharpen our process. www.institutionaldatainitiative.org/institutiona...
institutionaldatainitiative.org
Institutional Books | Institutional Data Initiative
Institutional Books 1.0 is our first release of public domain books. This set was originally digitized through Harvard Library’s participation in the Google Books project..
091
IDI @institutional.org · 12/06/2025
We look forward to growing Institutional Books through community. We welcome collaboration from researchers and model makers as we: - Evaluate the dataset’s impact on model outputs - Continuing to refine our OCR pipelines View the dataset on Hugging Face: huggingface.co/datasets/ins...
huggingface.co
institutional/institutional-books-1.0 · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1162
IDI @institutional.org · 12/06/2025
As part of our refinement work, we supplemented the original OCR-extracted text with a post-processed version that utilizes line detection to reassemble the text according to the line type.
120
IDI @institutional.org · 12/06/2025
We included extensive volume-level metadata with both original and generated components, such as results from text-level language detection.
130
IDI @institutional.org · 12/06/2025
We analyzed the dataset’s coverage across time, topic, and language and found: - 40% of English text + long tail of 254 languages - 20 clear topical tranches - Largely published in the 19th and 20th centuries Technical report here: arxiv.org/abs/2506.08300
281
IDI @institutional.org · 12/06/2025
Today we released Institutional Books 1.0, a 242B token dataset from Harvard Library's collections, refined for accuracy and usability. 🧵
97529
Reposted by IDI
Greg Leppert @leppert.me · 14/04/2025
The @institutionaldatainitiative.org is proud to support The New Commons challenge. $100k grants along with mentorship. Let's get impactful data into the AI ecosystem.
065
Reposted by IDI
Greg Leppert @leppert.me · 12/03/2025
As the @institutionaldatainitiative.org expands its mission, we’re announcing a collaboration with @bpl.boston.gov to develop AI-driven tools capable of accelerating new digitization at libraries across the world, starting at the Boston Public Library. institutionaldatainitiative.org/posts/using-...
institutionaldatainitiative.org
Using AI to Accelerate Digitization at Boston Public Librarys
Today, as part of our mission expansion, we’re announcing a collaboration with BPL to develop AI-driven tools capable of accelerating new digitization of large collections at libraries across the worl...
11710
Reposted by IDI
Greg Leppert @leppert.me · 05/03/2025
I'm pleased to announce we're expanding our mission at the @institutionaldatainitiative.org with an open call for institutional collaborators, new digitization at Harvard Law School Library, and additional support to advance this work. institutionaldatainitiative.org/posts/open-c...
institutionaldatainitiative.org
Expanding Our Mission: An Open Call for Collaborators
Today, we’re pleased to announce an open call for institutional collaborators as new support expands the research capacity of the Institutional Data Initiative.
1119
IDI @institutional.org · 13/12/2024
Hello world. institutionaldatainitiative.org/hello-world....
institutionaldatainitiative.org
How Knowledge Institutions Can Build a Promethean Moment
Why we’re launching the Institutional Data Initiative to work with libraries, government agencies, and other knowledge institutions to develop data collections and best practices for artificial intell...
073