From Annotation to Insight: Human-in-the-Loop Machine Learning for Historical Archives in HAICu WP2
Romein, C. A. (Annemieke); van Schuijlenburg, Koen; Peeters, Sebastiaan; Weber, Andreas
- Publisher
- Zenodo
- Language
- en
Abstract
This poster gives insight into ongoing research of new machine-learning architectures, which promote continual-machine learning through continual, ‘life-long’ harvesting of labels and annotations in order to enable multi-modal data mining (1). This work is carried out in the context of work pagage 2 of the large Dutch HAICu project (digital Humanities, Artificial Intelligence & Cultural heritage, 2024-2030) (2). The work package’s focus lies on layout clustering, document structure detection, contextual text linking. We research and develop scalable solutions for studying and interpreting handwritten and other multimodal collections with complex layouts by creating innovative feedback loops between volunteers-in-the-loop and machine-generated output (e.g. recurring textual and graphic patterns). Our use cases stem from three different Dutch archives: the National Archives in The Hague (NA), the Groninger Archives (GA), and the Collection Overijssel (CO). In the context of the NA, we work with the archives of the Dutch Ministry of Colonial Affairs (1850-1900), entailing ca. 4 million scans. Researchers from the NHL-Stenden University of Applied Sciences use advanced deep-learning technologies such as Laypa (3), DinoV2 vision transformers (4), and openTSNE mapping (5) to develop sophisticated latent space embeddings. These techniques transform page layouts based on structural features, going beyond conventional automatic text recognition to capture historical documents' nuanced visual and contextual complexity. Visualizing this latent space exposes clusters of handwriting styles for printed tabular structures, implicit tables on mixed layout pages, and writing density for different column widths (see first illustration). Concurrently, research at the University of Twente attempts to exploit the colonial archive’s innate preexisting structure and historical ordering principles (e.g. indices and klappers) for rapid information retrieval through keyword search. Also at the University of Twente, researchers study the Staten van Overijssel collection, comprising 60,000 pages. It documents centuries of political and social interactions in the Netherlands (1578-1795). We pay in particular attention to historical petitions, representing the 'voice of the people' and offering insights into societal concerns (see second illustration). One of the core objectives is to generate innovative automatic metadata - likely with ANNIF (6) - that will make these documents more accessible and interpretable for the general public. Central to WP2's methodology is research into robust human-in-the-loop frameworks that recognize the limitations of fully automated systems. Solutions developed within this WP rather continuously refine machine learning models by integrating academic expertise and volunteer contributions. This collaborative approach democratizes cultural heritage research and actively engages diverse stakeholders in preserving and interpreting colonial, regional and other multimodal archives. Acknowledgement: This work is supported by the Dutch Research Agenda (NWA) of the Dutch Research Council (NWO), grant: NWA.1518.22.105. References: (1) Schomaker, L.R.B. (2020). Lifelong Learning for Text Retrieval and Recognition in Historical Handwritten Document Collections. https://doi.org/10.1142/9789811203244_0012. (2) HAICu: https://www.haicu.science/. (3) Klut, S., et al. (2023). Laypa: A Novel Framework for Applying Segmentation Networks. https://doi.org/10.1145/3604951.3605520. (4) Oquab, M., et al. (2023). Dinov2: Learning robust visual features. https://doi.org/10.48550/arXiv.2304.07193. (5) https://opentsne.readthedocs.io/en/stable/# and Poličar, P.G., et al. (2024). OpenTSNE. https://doi.org/10.18637/jss.v109.i03. (6) Annif: https://annif.org/ and Suominen, O. (2019). Annif: DIY automated subject indexing. https://doi.org/10.18352/lq.10285.
Full text
NWO FUNDING: NWA.1518.22.105 From Annotation to Insight: Human-in-the-Loop Machine Learning for Historical Archives in HAICu WP2 C.A. Romein1 , K. van Schuijlenburg2, S. Peeters1 , B.J. Wolf3 , S.J.L. Weggeman3 , R. van Koert 4 , K. Dijkstra3 ,M.A. Dhali2, A. Weber12 , L.R.B. Schomaker2* 1 University of Twente; 2 University of Groningen; 3 NHL Stenden; 4 KNAW Humanities Cluster/ Digital Infrastructure *The order is based on relative contributions., starting with the largest contribution. CRedIT for this poster: Conceptualisation: CAR, AW, KvS Data curation: BW, CAR, KvS, SP Funding acquisition: AW, LS Investigation: all Methodology: BW, CAR, KvS, SP, RvK Resources: all Supervision: AW, MD Visualisation: CAR, KvS, SP Writing orig.: CAR, KvS Writing rev.: all t-SNE features of a DINO-embedding. National Archives, The Hague 4 Million Scans of Colonial Governance The archives of the Dutch Ministry of Colonial Aairs (1850-1900) contain millions of pages with diverse layouts: printed tables, handwrien annotations, mixed-format registers, and correspondence. Traditional text recognition fails to capture the structural complexity of these documents. At the University of Groningen, we extract the page of interest from images, allowing subsequent methods to focus only on relevant features. Deep learning models are continuously trained during data labelling, and computer vision algorithms are utilized to make labelling eicient. Working with NHL-Stenden University of Applied Sciences, we employ advanced deep-learning technologies, including Laypa for document layout analysis, and DINOv2 vision transformers to create sophisticated latent space embeddings for the unsupervised clustering of structural features. Visualising this latent space reveals meaningful clusters: handwriting styles, printed tabular structures, implicit tables on mixed-layout pages, and varying writing densities across dierent column configurations. Simultaneously, researchers at the University of Twente exploit the archive’s inherent historical organisation—indices and klappers—to develop rapid keyword-based information retrieval systems. Continuous Learning Through Citizen Science Central to our approach is the recognition that fully automated systems cannot capture the nuanced complexity of historical documents. We develop new frameworks where machine-learning models improve lifelong machine learning through strategic human intervention. Active Learning Strategies Algorithms identify uncertain or ambiguous cases, prioritising these for expert review. Scholars and trained volunteers provide targeted annotations precisely where models need guidance, maximising the impact of human eort. Continual Learning Architecture Rather than training a model for each scenario, we aim to incrementally learn models new models, tasks and classes without forgeing what has previously been learned. Democratising Research This collaborative approach engages diverse stakeholders—from academic researchers to heritage volunteers —in preserving and interpreting cultural archives. Human expertise guides computational power, whilst machine analysis reveals paerns invisible to manual review. Collectie Overijssel - Resolutions Centuries of Political Voice (1578-1795) The Staten van Overijssel archives document political and social interactions across more than two centuries of Dutch history. Within these ±60,000 pages, historical petitions oer particularly rich insights into societal concerns, representing authentic voices from early modern communities. A researcher at the University of Twente focuses on generating innovative automatic metadata to make these documents accessible beyond specialist circles. Using tools such as ANNIF for subject indexing, we will create semantic descriptions that connect historical content to contemporary understanding. The challenge lies not merely in transcribing text but in understanding context: recognising petitioners, identifying recurring grievances, tracing administrative responses, and mapping networks of political interaction. Human-in-the-loop methods prove essential here, as historical expertise guides the machine in distinguishing between formulaic language and substantive content. This work makes regional archives interpretable for the general public, revealing how ordinary people engaged with political institutions centuries ago. The HAICu project (digital Humanities, Artificial Intelligence & Cultural heritage, 2024-2030) develops innovative multimodal machine learning architectures for cultural heritage collections. Work Package 2 researchers lifelong machine learning frameworks that evolve through ongoing collaboration between computational models and human expertise. Rather than treating digitisation as a one-time process, we create feedback loops where volunteers and scholars continuously refine algorithms, enabling analysis of complex historical documents. Our research addresses three diverse Dutch archival collections, each presenting unique challenges in layout analysis, content extraction, and semantic interpretation. Groninger Archives Our third use case develops complementary approaches for collections at dierent stages of digitisation, ensuring our methods remain adaptable to varying institutional capacities and metadata standards. Looking Forward: 2024-2030 The HAICu WP2 project oers a new avenues of digital heritage research, demonstrating how continual learning and human-machine collaboration can unlock complex historical sources. Over the coming years, we will refine these methods and expand their application across diverse archival contexts.