Mass Digitization Leftovers: Digitizing Public Domain Books from the Stacks
Lovett, Julia
- Publisher
- Zenodo
- Language
- en
Abstract
In the early 2000s, the heyday of mass digitization, large universities participating in the Google Books Project and other efforts aimed to digitize their entire print holdings and put them online. As a result, you can find most older published books in Google Books, HathiTrust, Internet Archive, or smaller collections (with full-text access varying based on copyright status). Given the vast scope of previous mass digitization efforts, universities have mostly moved on to digitizing unique, rare, or archival collections. The same is true for URI, where recently we have been concentrating on scanning master's theses, dissertations, and student newspapers to include in the IR. But I recently became curious about whether any non-unique titles from our regular stacks were not yet available online. It seemed unlikely that any of these would have escaped the mass digitization dragnet. To investigate this question, I researched approximately 1600 titles from the stacks—specifically, titles that were published in the US between 1924 and 1929, which had recently entered the public domain and could be made openly available. To my surprise, about 2% of them were not yet available online. They included scientific studies of animals and plants, history of science, novels, poetry, local history, and even pharmacy. We have begun scanning these and adding them to our IR (https://digitalcommons.uri.edu/pd-books/) as well as the Internet Archive. This project is a great way for URI to contribute to our collective online, shared commons of published books, and indicates that other university libraries may have a similar percentage of public domain books that are not yet digitized. However, it is also time consuming to identify these titles, and we don’t know how much usage they will generate. To my knowledge, this is a unique approach to prioritizing materials for digitization. I will discuss the pros and cons.
Full text
Mass Digitization Leftovers: Digitizing Public Domain Books from the Stacks Julia Lovett Digital Initiatives Librarian University of Rhode Island Northeast Institutional Repository Day 2025 December 4, 2025
Public Domain Books - DigitalCommons@URI -Research Question: Given the prevalence of mass digitization projects of the early 2000s, how many published books (not unique/archival materials) havenʼt yet been digitized? - Focused on recently Public Domain titles, 1924-1929 - Because we can legally scan these! - Intended Outcomes: - Finding the percentage of public domain books not yet digitized in URI collection - Digitizing these books and sharing through both our IR and Internet Archive 2 Julia Lovett, University of Rhode Island Northeast Institutional Repository Day 2025
Identifying not-yet-digitized books - Alma Analytics report spreadsheet: - Books - Published 1924-1929 - Location ʻstacksʼ (excludes special collections/archives) - Resulted in 1,942 titles 3 Julia Lovett, University of Rhode Island Northeast Institutional Repository Day 2025
Identifying not-yet-digitized books - Searched 1,675 titles online - For each title, looked at: -HathiTrust - Internet Archive - The Online Books Page https://onlinebooks.library.upenn.edu/ - Google search - 33 were not found → 2% of titles - This was unexpectedly high! - Subjects: ecology; science; local history; poetry; fiction; ships! 4 Julia Lovett, University of Rhode Island Northeast Institutional Repository Day 2025
Digitizing and uploading to IR - Scanned on Zeutschel OS12000 scanner - PDFʼs uploaded: - DigitalCommons@URI - Internet Archive -IA viewer embedded into Digital Commons record page https://digitalcommons.uri.edu/pd-books/ 5 Julia Lovett, University of Rhode Island Northeast Institutional Repository Day 2025
Outcomes so far - 19 books digitized - 841 downloads (as of 11/25/25) - Time estimates: - Search: 2 minutes per title (most were just a few seconds) equaling 56 hours of work - Scanning and uploading time 2.5 hours/book, equaling 47.5 hours of work - TOTAL: 103 hours (done a little at a time, over 2 summers) - Could other libraries have approximately 2% of their public domain holdings not yet available on Google Books, HathiTrust, Internet Archive, other sites…? And if so, could this be a valuable way to prioritize materials for digitization? Thank you! Questions? jalovet[email protected] 6 Julia Lovett, University of Rhode Island Northeast Institutional Repository Day 2025