Sign in

Greg Leppert

@leppert.me
1.2K followers 141 following 46 posts

Working on AI and access to knowledge at Harvard. Executive Director of the Institutional Data Initiative; Chief Technologist of the Berkman Klein Center.

PostsRepliesMedia
Greg Leppert @leppert.me · 08/09/2026
The Institutional Data Initiative continues our work to increase the computational accessibility of the ~1M volumes in Insitutional Books. Today’s releases add access to 22M photos and illustrations as well as 1.9B annotated paragraphs. Dig in.
030
Greg Leppert @leppert.me · 02/09/2026
Thanks. We've deleted and reposted: bsky.app/profile/inst...
000
Greg Leppert @leppert.me · 02/09/2026
Immensely proud of this collaboration with @bpl.boston.gov. A dataset of 1.47M newspaper pages and a new processing pipeline that breaks pages down to their component parts, increasing OCR accuracy and computational addressability. The historical record has never been more searchable.
032
Greg Leppert @leppert.me · 13/08/2026
Find the dataset, as well as a SKILL file for agentic access, here: huggingface.co/collections/...
huggingface.co
Institutional Proceedings - a institutional Collection
A growing corpus of decision-making records, parsed and optimized for computational access.
010
Greg Leppert @leppert.me · 13/08/2026
This project is the first of many that @institutional.org is undertaking to help libraries and other knowledge institutions enable computational access to their collections. If you represent a library and would like to work together, we'd love to hear from you.
100
Greg Leppert @leppert.me · 13/08/2026
Across 186yrs, we identified leadership changes, fundraising progress, and other strategic plans including the creation of race and gender policies following the Civil Rights Act as well as a 1921 $1M drive for a “women’s building,” launched after women were barred from entering the Michigan Union.
110
Greg Leppert @leppert.me · 13/08/2026
Organizational decision-making is fertile ground for AI-assisted research, but board minutes often bury key decisions in narrative prose. Using reasoning models, we worked with U-M Library to extract & structure these decisions so researchers can use computational tools to study the evolution of U-M
100
Greg Leppert @leppert.me · 13/08/2026
AI enables new forms of research, but only where information is computationally accessible. Today, @institutional.org and @umich.edu Library are releasing a structured dataset that makes 186 years of U-M’s Board of Regents Proceedings newly accessible. institutional.org/posts/instit...
institutional.org
Institutional Proceedings: 186 years of university decisions. | IDI
IDI and the University of Michigan Library have collaborated to release two centuries of the university’s governance decisions as a structured dataset that can be searched, analyzed, and reused.
111
Greg Leppert @leppert.me · 07/12/2025
🤪
010
Greg Leppert @leppert.me · 07/12/2025
AI is autotune (and Beat Detective) for culture writ large. www.nytimes.com/2025/12/03/m...
nytimes.com
Why Does A.I. Write Like … That?
110
Greg Leppert @leppert.me · 20/11/2025
Even if you're not a partner library, you might be curious about what it's like to work with GRIN. Our technical report has a wealth of details. arxiv.org/abs/2511.11447
arxiv.org
GRIN Transfer: A production-ready tool for libraries to retrieve digital copies from Google Books
Publicly launched in 2004, the Google Books project has scanned tens of millions of items in partnership with libraries around the world. As part of this project, Google created the Google Return Inte...
001
Greg Leppert @leppert.me · 20/11/2025
We're also sharing the pipeline we developed for Institutional Books that seamlessly dedupes, classifies, and enhances the data once GRIN Transfer brings it down. www.institutional.org/tools
institutional.org
Institutional Books | Institutional Data Initiative
Institutional Books 1.0 is our first release of public domain books. This set was originally digitized through Harvard Library’s participation in the Google Books project..
101
Greg Leppert @leppert.me · 20/11/2025
That's why we built GRIN Transfer: a tool for downloading collections, big or small. GRIN Transfer handles request batching, failure recovery, and data aggregation so that libraries can focus on using the data rather than simply gaining access to it. www.institutional.org/posts/grin-t...
institutional.org
Announcing the release of GRIN Transfer
GRIN Transfer, an open source tool that allows Google Books partner libraries to more easily access their Google Books collection.
100
Greg Leppert @leppert.me · 20/11/2025
We learned this lesson over the months it took to download 1M of Harvard Library's books for our Institutional Books release. As a result, many libraries have yet to take full advantage of the wonderful resources GRIN provides.
100
Greg Leppert @leppert.me · 20/11/2025
When libraries join Google Books, Google not only scans their books, it also makes a wealth of image, OCR, & metadata available to them via the Google Return Interface (GRIN). But working with GRIN can be challenging, so we're releasing a tool to make it easier. www.institutional.org/posts/grin-t...
institutional.org
Announcing the release of GRIN Transfer
GRIN Transfer, an open source tool that allows Google Books partner libraries to more easily access their Google Books collection.
153
Greg Leppert @leppert.me · 25/07/2025
Data is everything.
000
Greg Leppert @leppert.me · 24/07/2025
Amanda Watson, @institutionaldatainitiative.org's Library Chair and leader of Harvard Law School Library, spoke about the importance of publishing library collections as data to guide the future of AI. hls.harvard.edu/today/food-f...
hls.harvard.edu
Food for (AI) thought and the library initiative improving AI’s digital diet - Harvard Law School
Amanda Watson of the Harvard Law School Library says the release of Harvard’s digitized collection is only the beginning of collaborations between libraries and tech firms.
020
Reposted by Greg Leppert
Ben Lee @bcgl.bsky.social · 02/07/2025
With @yh-huang.bsky.social, I'm excited to share our Digital Collections Explorer, an open-source, multimodal viewer for digital collections! Users can search with both natural language inputs and reverse image search. Paper: arxiv.org/abs/2507.00961 Public demo: digital-collections-explorer.com
arxiv.org
Digital Collections Explorer: An Open-Source, Multimodal Viewer for Searching Digital Collections
We present Digital Collections Explorer, a web-based, open-source exploratory search platform that leverages CLIP (Contrastive Language-Image Pre-training) for enhanced visual discovery of digital col...
27426
Greg Leppert @leppert.me · 23/06/2025
This starts in an hour.
000
Greg Leppert @leppert.me · 20/06/2025
June 23rd at 12:45pm ET. RSVP here: harvard.zoom.us/meeting/regi...
harvard.zoom.us
Welcome! You are invited to join a meeting: IDI Talk with Petr Knoth (CORE). After registering, you will receive a confirmation email about joining the meeting.
Welcome! You are invited to join a meeting: IDI Talk with Petr Knoth (CORE). After registering, you will receive a confirmation email about joining the meeting.
010
Greg Leppert @leppert.me · 20/06/2025
This Monday, @institutionaldatainitiative.org will host Petr Knoth to share his experience leading CORE ("The world’s largest collection of open access research papers") as the rise of AI brings new meaning, and challenges, to stewarding knowledge repositories. Join us virtually via the link below.
222
Greg Leppert @leppert.me · 16/06/2025
Cohosted by @institutionaldatainitiative.org and The Berkman Klein Center. harvard.zoom.us/webinar/regi...
harvard.zoom.us
Welcome! You are invited to join a webinar: Open AI Development. After registering, you will receive a confirmation email about joining the webinar.
For AI to truly benefit society, it must be built on foundations of transparency, fairness, and accountability—starting with the most foundational building block that powers it: data. Not long ago, ...
001
Greg Leppert @leppert.me · 16/06/2025
Tomorrow, it's our pleasure to host @ayahbdeir.bsky.social to talk about the power of data in building an AI ecosystem that's open, transparent, and fair. 11am ET on June 17th. Register at the link below to attend virtually.
100
Greg Leppert @leppert.me · 14/04/2025
The @institutionaldatainitiative.org is proud to support The New Commons challenge. $100k grants along with mentorship. Let's get impactful data into the AI ecosystem.
065
Reposted by Greg Leppert
Free Law Project ⚖ @free.law · 21/03/2025
To start the weekend, we've got a brand new experience for case law on CourtListener. It has better typography, more features and metadata, five million scanned decisions from @harvardlil.bsky.social, and a lot more. Read all about it and let us know what you think: free.law/2025/03/21/c...
free.law
A Faster, Smarter, Unified Case Law Experience
A redesigned case law modernizes the reading experience with enhanced layout and typography, more advanced features, better speed, and more.
1278
Greg Leppert @leppert.me · 12/03/2025
The @institutionaldatainitiative.org at Harvard works with knowledge institutions to increase the availability, diversity, and responsible use of training data for AI. Reach out and join us.
010
Greg Leppert @leppert.me · 12/03/2025
Our goal is to develop methods and tools that can support expert staff at libraries everywhere, increasing the breadth of materials that can be digitized and the speed at which they’re made accessible to the public. Learn more at BPL: www.bpl.org/news/boston-...
bpl.org
Boston Public Library Expands Access to Collections Through AI-Enhanced Digitization
BOSTON, MA – March 12, 2025 - The Boston Public Library (BPL) is launching a large-scale digitization project to unlock hundreds of thousands…
121
Greg Leppert @leppert.me · 12/03/2025
Together, we’ll research opportunities to generate machine-readable representations of items, add searchable metadata, and begin the structuring of entire collections—all at the moment each item leaves the imaging station.
100
Greg Leppert @leppert.me · 12/03/2025
IDI and BPL are working to change this by collaborating at the outset of a large digitization project, exploring how AI might complement human expertise and strengthen the process in its earliest stages.
100
Greg Leppert @leppert.me · 12/03/2025
BPL is embarking on a new initiative to digitize hundreds of thousands of historic items. Conventional approaches to this scale lead to an impossible choice: sacrifice depth for breadth or drastically limit what gets digitized. AI tools can help, but they’re relegated to the end of the process.
110
Greg Leppert @leppert.me · 12/03/2025
As the @institutionaldatainitiative.org expands its mission, we’re announcing a collaboration with @bpl.boston.gov to develop AI-driven tools capable of accelerating new digitization at libraries across the world, starting at the Boston Public Library. institutionaldatainitiative.org/posts/using-...
institutionaldatainitiative.org
Using AI to Accelerate Digitization at Boston Public Librarys
Today, as part of our mission expansion, we’re announcing a collaboration with BPL to develop AI-driven tools capable of accelerating new digitization of large collections at libraries across the worl...
11710
Greg Leppert @leppert.me · 05/03/2025
With our digitization at Harvard Law School Library, we'll work to increase access to unique collections, such as the Supreme Court Records and Briefs that are critical to understanding decision-making at the highest U.S. court yet remain largely inaccessible.
000
Greg Leppert @leppert.me · 05/03/2025
If you're part of a library, university, or other knowledge institution and interested in working with a team of data scientists to refine and publish your data, we'd love to chat. And if you're a data scientist or community builder interested in working with institutions, we're hiring.
100
Greg Leppert @leppert.me · 05/03/2025
IDI is building a collection of large, impactful, and widely available datasets to increase AI’s accessibility and diversity while reaffirming institutions as stewards of knowledge.
100
Greg Leppert @leppert.me · 05/03/2025
I'm pleased to announce we're expanding our mission at the @institutionaldatainitiative.org with an open call for institutional collaborators, new digitization at Harvard Law School Library, and additional support to advance this work. institutionaldatainitiative.org/posts/open-c...
institutionaldatainitiative.org
Expanding Our Mission: An Open Call for Collaborators
Today, we’re pleased to announce an open call for institutional collaborators as new support expands the research capacity of the Institutional Data Initiative.
1119
Greg Leppert @leppert.me · 12/02/2025
Great op-ed from @shaynelongpre.bsky.social on the effects AI — as a technology and as a market — is having on the web. www.technologyreview.com/2025/02/11/1...
technologyreview.com
AI crawler wars threaten to make the web more closed for everyone
There’s an accelerating cat-and-mouse game between web publishers and AI crawlers, and we all stand to lose.
031
Greg Leppert @leppert.me · 29/01/2025
In 15mins (1pm ET), I'll be giving a talk about @institutionaldatainitiative.org and our quest to build the most boring dataset in the world. Tune in here: lu.ma/iqkqvcus
lu.ma
Greg Leppert, Harvard | The Most Boring Dataset in the World · Luma
Foresight Institute’s Intelligent Cooperation Group The Most Boring Dataset in the World Abstract: Data is a critical raw material in the construction of AI.…
040
Greg Leppert @leppert.me · 21/12/2024
The key is involving the institutions themselves in the conversation. Their missions are as much a reflection of the cultures they help to preserve as the data itself, and integrating them is critical as we look for new models to foster and interact with knowledge.
090
Greg Leppert @leppert.me · 21/12/2024
@karaswisher.bsky.social and @ylecun.bsky.social discuss our @institutionaldatainitiative.org in the latest Pivot. It's essential that we establish equitable and sustainable models for data access and stewardship in the age of AI, and IDI is working on exactly that podcasts.apple.com/us/podcast/m...
podcasts.apple.com
Meta's Chief AI Scientist Yann LeCun Makes the Case for Open Source | On With Kara Swisher
Podcast Episode · Pivot · 12/21/2024 · 55m
1102
Greg Leppert @leppert.me · 13/12/2024
To learn more about the Institutional Data Initiative, read our launch announcement here: institutionaldatainitiative.org/hello-world....
institutionaldatainitiative.org
How Knowledge Institutions Can Build a Promethean Moment
Why we’re launching the Institutional Data Initiative to work with libraries, government agencies, and other knowledge institutions to develop data collections and best practices for artificial intell...
050
Greg Leppert @leppert.me · 13/12/2024
Our launch is generously supported by @microsoft.com and OpenAI, and these projects are just the start. If you’re part of an institution, we’d love to hear how we can help. If you’re an AI/ML researcher, we'd love to collaborate. We're also hiring for our core team, here at Harvard. Reach out.
130
Greg Leppert @leppert.me · 13/12/2024
We're also working with Boston Public Library on an active scanning project of theirs involving millions of pages taken from public domain newspapers. We're hoping to make progress on some of the classic OCR challenges presented by newspapers, while evaluating the resulting data for model training.
100
Greg Leppert @leppert.me · 13/12/2024
Our first project is a collection of 1M public domain books, scanned at Harvard Library as part of the Google Books project. We're fortunate enough to have the support of Google in releasing this dataset and will do so in early 2025 with accompanying analysis.
100
Greg Leppert @leppert.me · 13/12/2024
Our goal is to work within the missions of institutions to help unlock and improve access to their collections, not only for AI uses but traditional patron access as well. This work can help expand the breadth of people that institutional collections and AI models are able to empower.
100
Greg Leppert @leppert.me · 13/12/2024
As a team of technologists who deeply value knowledge institutions, we designed @institutionaldatainitiative.org as a gear to help the AI and institutional communities, each moving at their own unique speeds, mesh and transfer energy between them. Energy, experience, and expertise.
100
Greg Leppert @leppert.me · 13/12/2024
Data access and integrity are two things institutions know well. They hold vast collections and think a lot about knowledge access and stewardship. Stewardship is a nice word because it conveys a sense of integrity over time, and time is something institutions excel at utilizing.
100
Greg Leppert @leppert.me · 13/12/2024
Data access plays an important role in the development of AI models. It helps define who is represented in them, who they empower, and even who can build them. Data integrity, as a result, is important as well.
100
Greg Leppert @leppert.me · 13/12/2024
Yesterday we launched the Institutional Data Initiative at Harvard Law School Library to work with libraries, government agencies, and other knowledge institutions to help refine and publish their collections as data, with an eye toward AI. 🧵 bsky.app/profile/inst...
1134