Sign in

IDI

@institutional.org
297 followers 2 following 32 posts

A research center at Harvard working to strengthen society’s connection to knowledge by advancing our access to and understanding of the data that shapes AI.

PostsRepliesMedia
IDI @institutional.org · 08/09/2026
The Institutional Data Initiative at Harvard @institutional.org works with knowledge institutions to increase the availability, diversity, and responsible use of training data for AI. Reach out and join us.
010
IDI @institutional.org · 08/09/2026
The Institutional Data Initiative at Harvard (@instdin) works with knowledge institutions to increase the availability, diversity, and responsible use of training data for AI. Reach out and join us.
000
IDI @institutional.org · 09/09/2025
Can a small visual language model read documents as effectively as models 27 times its size? Next Friday, IDI will host Michele Dolfi and Peter Staar from IBM Research Zurich to discuss their work on SmolDocling, an “ultra-compact” model for diverse OCR tasks.
100
IDI @institutional.org · 12/06/2025
As part of our refinement work, we supplemented the original OCR-extracted text with a post-processed version that utilizes line detection to reassemble the text according to the line type.
120
IDI @institutional.org · 12/06/2025
We included extensive volume-level metadata with both original and generated components, such as results from text-level language detection.
130
IDI @institutional.org · 12/06/2025
We analyzed the dataset’s coverage across time, topic, and language and found: - 40% of English text + long tail of 254 languages - 20 clear topical tranches - Largely published in the 19th and 20th centuries Technical report here: arxiv.org/abs/2506.08300
281
IDI @institutional.org · 12/06/2025
Today we released Institutional Books 1.0, a 242B token dataset from Harvard Library's collections, refined for accuracy and usability. 🧵
97529