DCMS x NHM AI Pilot Project - RBGE
Full text
DCMS AI Pilot Project Report Enhancing herbarium digitisation with AI-driven computer vision for image quality monitoring Natural History Museum, London Naifeng Zhang1, Sanson T. S. Poon, Arianna Salili-James, and Ben Scott Royal Botanic Garden Edinburgh Elspeth Haston, Milo Phillips, Robyn Drinkwater, and Rob Cubey 31st March 2025 1 Project 1.1 Aim The aim of this project is to develop a flexible and scalable quality control (QC) framework that assists in the digitisation and quality assurance of herbarium sheet images. This study addresses several image quality issues, such as cropping errors, focus problems, missing elements, as well as rare anomalies, by employing advanced artificial intelligence (AI)- driven object detection and classification techniques. By integrating AI-powered quality checks, the framework improves efficiency and ensures more consistent image quality. If successful, this approach will benefit the institutions involved in this study while also assisting others with limited capacity or resources for comprehensive quality assessments, helping to streamline workflows and reduce costly errors in the digitisation process. Using an ‘issue’ dataset, generated and provided by the Royal Botanic Garden Edinburgh (RBGE), we aim to ensure that the AI-assisted tool we develop is capable of detecting these issues with high accuracy while also offering detailed insights to improve decision-making. Ultimately, our objective is to build a comprehensive system that can adapt to emerging challenges in herbarium digitisation workflows across global institutions, thereby improving operational efficiency and accuracy, and supporting national and global digitisation efforts. 1.2 Background Over the past two decades, approximately 13,000 herbarium sheet images at RBGE have required re-digitisation due to quality or workflow issues, such as unsatisfactory image 1[email protected]
quality, missing content, or duplication within large-scale digitisation projects. While this figure may appear low compared to the one million specimens digitised, it vastly underrepresents the true scale of the problem, as most errors are identified and corrected through rigorous manual quality control by highly skilled digitisers. With a digitisation rate of 600–1,000 images per day and a re-digitisation cost of approximately £1.15 per specimen, even a slight increase in the error rate could result in significant financial and operational burdens. As digitisation programmes continue to expand, both the frequency of errors and the associated costs are expected to rise. This challenge is not unique to RBGE, as institutions across the country face similar constraints. The integration of AI-driven quality control could transform digitisation workflows by reducing reliance on manual inspection, minimising errors, and delivering substantial cost savings at scale. Herbarium sheets are crucial resources in botanical research, serving as permanent records of plant diversity and enabling taxonomic studies (BARAL,1992;Besnard et al., 2018), ecological research (Heberling and Isaac,2017;James et al.,2018), and conservation efforts (MacDougall et al.,1998;Nualart et al.,2017;Albani Rocchetti et al.,2021). The digitisation of these specimens allows for easier access and widespread dissemination, which significantly enhances research capabilities and data sharing (Soltis,2017;Niedzielski and Markiewicz,2023). However, converting these physical specimens into high-quality digital images faces several challenges, notably issues related to image quality. The digitisation process in herbaria has evolved from flat-bed scanners to the more widespread use of high-resolution cameras. During the image capture and processing workflows, a number of quality issues can occur, including mis-cropping, out-of-focus images, and the occurrence of unexpected bars or other anomalies. These quality issues can compromise the reliability and usability of digital archives, making quality control a critical component of the digitisation workflow. Moreover, as noted in Ahl et al. (2023), quality assessment and data cleaning are essential steps in ensuring dataset integrity, as error-free datasets are rare. Traditional quality control methods often rely on manual inspection and experiencebased evaluations, making the process time-consuming and error-prone (e.g., de la Hidalga et al.,2020;De Smedt et al.,2024;Novikov and Nachychko,2025). A substantial number of natural history images are being generated through the digitisation of UK collections. However, many institutions lack the capacity to effectively assess the quality of these images efficiently before publishing them. This limitation underscores the need for comprehensive solutions that can address a broad spectrum of quality problems efficiently. Recent advancements in computer vision, particularly in object detection (Zou et al., 2023), offer promising solutions for automating quality checks. Models from traditional machine learning methods to convolutional neural network (CNN)-based deep learning models such as ResNet (Residual Neural Network, He et al.,2016) and YOLO (You Only Look Once, Redmon et al.,2016) provide a robust foundation for developing advanced quality assurance tools. There is a pressing need to develop a flexible and scalable quality 2
control framework that can easily incorporate new detection algorithms, allowing for continuous adaptation to emerging challenges in herbarium sheet digitisation. By leveraging state-of-the-art techniques in computer vision and machine learning, this framework can significantly enhance the accuracy and efficiency of quality control processes, supporting the global effort to digitise botanical collections effectively. 2 Dataset 2.1 Raw Dataset The issue dataset, provided by RBGE, consists of digitised herbarium sheet images that represent issues encountered during their everyday herbarium sheet digitisation and quality control processes. The dataset contains a total of 766 images, including 120 standard images without known issues, while the remaining 646 images were sets with intentionally introduced errors for training purposes. This study focuses on five selected issues frequently observed in herbarium sheet digitisation: bars, focus errors, missing elements, under-cropping, and over-cropping (Figure 1). These images were produced by the digitisation team at RBGE as part of this project. They represent a broad range of taxa across different regions of the world. To facilitate structured analysis, the dataset is organised into five distinct sub-datasets based on the specific issue present in each image. Each image contains only one identified issue, ensuring no duplication across sub-datasets. The distribution of images across the sub-datasets is presented in Table 1. Note that this distribution only reflects the allocation used in this study and does not necessarily represent the underlying prevalence of digitisation issues in any collection or organisation. Table 1: Distribution of the raw dataset provided by RBGE, where each subset represents a specific type of issue. The count indicates the number of occurrences in each category. Sub-dataset Count Bars 2 Focus 300 Barcode-Missing 30 Under-cropped 73 Over-cropped 241 Standard 120 2.2 Preprocessing To incorporate various digitisation issues for training, validation, and testing across multiple AI models, we structured subsets tailored to specific objectives. The bars issue, due 3
(a) Standard Image (b) Abnormal Bars (c) Focus Errors (d) Missing Elements (e) Under-cropping (f) Over-cropping Figure 1: Examples of standard and flawed images, indicating five frequent issues we selected in this project: abnormal bars, focus errors, missing elements, under-cropping and over-cropping. to its rarity and minimal sample size, was addressed using a non-AI approach. Table 2 summarises the dataset distribution and preprocessing steps for each target issue. We begin with the raw dataset as a foundation and create specialised image sets for training the object detection model and addressing specific digitisation issues. To ensure consistency in image processing, all images are resized to a fixed width of 1080 pixels while maintaining their original aspect ratio. For object detection tasks (e.g., detecting missing barcode, colour bar, or scale bar), we incorporate all 120 standard, issue-free images along with 50 under-cropped and 50 over-cropped images, forming a dataset of 220 images. Including images with cropping issues enhances the model’s robustness in detecting objects at varying scales. This dataset is divided into training, validation, and test sets using a 6:2:2 split. 4
Table 2: Summary of datasets used for different purposes. For object detection, 60% of the dataset is used for training, and 20% is used for validation in each epoch to identify the best model. The ‘Additional Notes’ column provides details about the target issue images, except for the ‘bars’ issue. Target Dataset size Train:Test ratio Additional notes Object Detection 220 (6:2):2 100 images with cropping issue Out of Focus 220 7:3 100 least distorted images only Under-Cropping 193 8:2 73 under-cropped samples Over-Cropping 370 8:2 200 selected, 50 manually created Bars 2 N/A Not AI model, no training required The raw focus dataset consists of three sub-datasets, each containing 100 images with varying degrees of distortion. During model development, we use the 100 least distorted images alongside standard images to ensure a more balanced dataset. This dataset is split into training and test sets using a 7:3 ratio. For under-cropping detection, we utilise all 73 raw under-cropped images along with 120 standard images, forming a dataset of 193 images. A 8:2 train-test split is applied to this dataset. An analysis of the 241 available over-cropped images revealed that fewer than 20 specifically exhibited barcode over-cropping. To create a more balanced dataset, we randomly select 200 images from the over-cropping dataset and manually generate 50 additional samples by applying barcode over-cropping to 10 standard images at five random positions. This results in a dataset of 250 over-cropped images and 120 standard images, split into training and test sets with a 8:2 ratio. Due to the rarity of the bars issue and the availability of only two samples, a nonAI approach using traditional edge extraction and contour detection methods is applied. These techniques are tested directly on the two raw, non-preprocessed bar images. 3 Methods As discussed in Section 1, new issues frequently arise in the digitisation workflow. To effectively support quality checkers, tools require continuously updated to incorporate newly encountered issues while maintaining coverage of existing ones. Therefore, it is essential to develop a flexible and scalable framework that can seamlessly integrate new features. This section outlines the framework’s design and examines key digitisation issues, detailing how they were addressed and incorporated into the system. The source code for this work has been made available on Github2. 2github.com/NaturalHistoryMuseum/nhm dcms rbge qualitycontrol 5
3.1 Quality Control AI Assistant Framework For many quality issues, such as image mis-cropping and missing objects, it is difficult to directly identify these issues from raw images. Since each issue has distinct characteristics, developing a single end-to-end model to handle all of them is impractical. To address this, we designed a scalable QC AI assistant framework, as illustrated in Figure 2. Figure 2: Quality Control AI Assistant Framework. The first step in the framework involves standardising the input data. To better identify all possible issues, we implemented an object detection method to generate bounding boxes of each object, providing location and size information of each item on the herbarium sheet. These data, along with the original image, are used as input to identify various issues. The second step involves developing dedicated datasets and models for each specific issue. This modular approach enables independent and parallel training, validation, and testing of each issue detector, ensuring greater flexibility and scalability. 3.2 Object Detection To enhance the detection of issues in herbarium sheet images, we employ a robust object detection method to acquire the location, scale, and size of each item on the sheet. We use YOLO (You Only Look Once, Redmon et al.,2016), a CNN-based object identification model, renowned for its speed and efficiency. Unlike methods such as Faster R-CNN (Ren et al.,2016), which segment images into region proposals, YOLO treats object detection as a single regression problem, enabling the processing of an entire image at once. This architecture makes YOLO exceptionally fast and capable of real-time processing. This is a crucial feature when testing large volumes of images during quality control process. Despite its speed, YOLO ensures high 6
accuracy and minimal bounding box distortion, which is critical for identifying issues that relate to objects in the images. The model we use, YOLOv11 (Jocher and Qiu,2024), is the latest model that offers state-of-the-art performance in both accuracy and speed. 3.3 Cropping Error Detection Before detecting cropping issues, we first define under-cropping and over-cropping, as shown in Figure 3. Over-cropping, shown in Figure 3b, occurs when at least one edge of an image is excessively cropped, leading to the potential loss of important information. On the other hand, under-cropping, depicted in Figure 3c, happens when unnecessary space remains around the specimen on at least one edge due to insufficient cropping. (a) Correctly cropped image. (b) Over-cropped image. (c) Under-cropped image. Figure 3: Illustration of correctly cropped, over-cropped, and under-cropped images. In the over-cropped image, excessive cropping on the left results in a loss of information from the label and specimen. The under-cropped image retains unnecessary space to the right of the colour-card and ruler. 3.3.1 Over-cropping Detection To detect over-cropping issues in herbarium sheet images, we leverage the fixed and standard items present on the sheets: the colour card, ruler, and barcode. According to RBGE’s digitisation workflow, the colour card and ruler are consistently placed along the right edge of the image capture platform, while barcodes are positioned in the bottom-left corner in newer images and typically near the bottom in older ones. When over-cropping occurs, these fixed items are likely to be partially cut. However, even though all colour cards and rulers are identical across images, and barcodes share a similar appearance, we cannot rely on pixel width and height alone to determine over-cropping. Image scaling varies, with widths normalised to 1080 pixels, making direct size measurements unreliable. Instead, we use the barcode’s aspect ratio (width-to-height) as a key indicator of over-cropping. However, aspect ratio does not work for the colour card, as sometimes both its width 7
and height shrink proportionally when cropped, keeping the aspect ratio unchanged. Figure 4shows that 56.5% of over-cropped images have a colour card aspect ratio within the standard range, making it an unreliable indicator. To improve detection, we normalise the colour card’s pixel width and height by dividing them by the ruler’s height, allowing for consistent cross-image comparisons. Figure 4: Boxplot of colour card Aspect Ratio (AR) for Raw Standard and Over-Cropped Image Datasets. A significant portion of the AR values in the over-cropped images fall within the range defined by the minimum and maximum values of the standard images. For the ruler, we use the distance between its right edge and the image’s right edge as an additional feature, normalised by the ruler’s height. In total, we extract four key features for detecting over-cropping, two from colour card and two from barcode and ruler respectively. To classify this over-cropping issues, we employ a decision tree classifier, a fast and interpretable machine learning approach (Xu et al.,2019). To improve classification performance, we standardise feature trends so that smaller values consistently indicate a higher likelihood of belonging to a specific class, while larger values suggest the opposite. Specifically, we apply median normalisation to the barcode aspect ratio, defined as: D=|X−median(X)|,(1) where Xrepresents the barcode aspect ratio, and Ddenotes its absolute deviation from the median. The normalised barcode aspect ratio X′is then computed as: X′= 1 −D−min(D) max(D)−min(D),(2) ensuring that smaller values correspond to a higher likelihood of a cropping issue. Standardising feature trends in this manner simplifies the decision tree’s construction, enabling it 8
to form decision boundaries more efficiently without reconciling inconsistent trends. This results in a more compact and interpretable model, reducing over-fitting and improving generalisation to unseen data. 3.3.2 Under-cropping Detection Detecting under-cropping is more straightforward. First, we compute the minimum bounding box that encloses all detected objects in the image. Next, we extract five key features: the objects ratio, r(the area of the bounding box divided by the total image area), and the distances from each side of the bounding box to the corresponding image borders. Similar to the approach used for over-cropping detection, in order to facilitates the subsequent application of a decision tree model, we adjust the rby replacing it with 1 −rto ensure consistent trends across all five features. 3.4 Out of Focus Detection In the digitisation process, incorrect camera parameter settings and shifts in camera position can lead to the camera not focusing properly on the image. This focus error can cause the image to become blurred. To detect whether an image has out-of-focus issues, we have adopted a method combining statistics and deep learning. Firstly, we use the Laplacian operator (Bansal et al.,2016) to measure the degree of blurriness in the image. The formula for the Laplacian operator is: ∇2f=∂2f ∂x2+∂2f ∂y2,(3) where frepresents the image intensity function, and ∇2fdenotes the Laplacian of f, which captures regions of rapid intensity change to detect edges and measure blurriness. It quantifies the rate of change in pixel intensity at a specific location by calculating the gradient, thus reflecting the clarity of the image. A higher gradient indicates a clearer image. Subsequently, we fine-tuned a simple pre-trained network, ResNet18 (He et al., 2016), using the gradient map produced by the Laplacian filter as input to perform a binary classification. This helps us determine whether the image is blurred and consequently infer if there is an out-of-focus issue. 3.5 Bars Detection In the practical process of quality control, the QC team may encounter some extremely rare issues. The strange bars shown in the Figure 5are one such example. The RBGE team provided only these two images as samples. We have included this issue as part of the problems our project aims to solve, in order to demonstrate the extensibility of our assistant framework. 9
cropping models exhibit limited generalisability, as the decision tree model for detecting over-cropping and under-cropping is based on predefined priors related to the specific kind and location of colour card, ruler, and barcode, limiting applicability to different designs of herbarium sheets. To enhance the framework, the next phase should focus on several key areas. This includes gathering and generating a larger, more diverse dataset that includes a variety of quality issues, developing a more robust model for detecting over-cropping issues based on varied herbarium designs, exploring additional quality issues, and implementing corresponding detection models for improved coverage. Additionally, promoting collaboration with other institutions will foster cross-institutional knowledge exchange, enabling the sharing of quality control methodologies, error detection strategies, and AI-driven approaches, which can enhance the consistency and accuracy of digitisation processes across various organisations. By addressing these limitations and broadening the project’s scope, the next phase aims to deploy an even more powerful and adaptable quality control system that can be widely adopted across various institutions and domains. The digitisation of the UK’s remaining 20.5 million botanical specimens, beyond RBGE’s existing efforts, presents considerable financial risks due to inconsistent resourcing across institutions. Survey data from DiSSCo UK highlights (Smith et al.,2022) these challenges: only 13% of institutions employ full-time digitisers, over half rely on part-time staff, and fewer than 25% have access to advanced imaging equipment. This disparity in capacity suggests that error rates in digitisation could be significantly higher than RBGE’s exceptionally low 0.067% re-imaging rate, which has been maintained by a small team of highly skilled specialists. Without intervention, institutions with limited expertise or equipment could experience error rates 10 to 20 times higher, leading to a dramatic increase in the costs associated with rework. The AI-assisted quality estimation introduced in this pilot project presents a practical solution to mitigate these risks. If AI can reduce the error rate from an estimated 1% to the current RBGE level of 0.1%, the impact would be substantial—preventing over 180,000 unnecessary re-digitisations and saving more than £200,000 across the full 20.5 million specimens. Beyond direct cost reductions, AI would also alleviate hidden expenses, such as the significant labour required for manual error detection and correction, which is not accounted for in the £1.15 per-specimen re-digitisation cost. This challenge extends beyond RBGE, affecting institutions nationwide. Implementing AI-driven quality control has the potential to transform digitisation workflows by reducing reliance on manual inspection, minimising errors, and achieving substantial cost savings at scale. 16
References L. Ahl, L. Bellucci, P. Brewer, P.-Y. Gagnier, E. Haston, L. Livermore, S. De Smedt, H. Hardy, and H. Enghoff. Digitisation of natural history collections: criteria for prioritisation. Research Ideas and Outcomes, 9:e114548, 2023. G. Albani Rocchetti, C. G. Armstrong, T. Abeli, S. Orsenigo, C. Jasper, S. Joly, A. Bruneau, M. Zytaruk, and J. C. Vamosi. Reversing extinction trends: new uses of (old) herbarium specimens to accelerate conservation action on threatened species. New Phytologist, 230(2):433–450, 2021. R. Bansal, G. Raj, and T. Choudhury. Blur image detection using laplacian operator and open-cv. In 2016 International Conference System Modeling & Advancement in Research Trends (SMART), pages 63–67. IEEE, 2016. H. BARAL. Vital versus herbarium taxonomy: morphological differences between living and dead cells of ascomycetes, and their taxonomic implications. Mycotaxon, 44(2): 333–390, 1992. G. Besnard, M. Gaudeul, S. Lavergne, S. Muller, G. Rouhan, A. P. Sukhorukov, A. Vanderpoorten, and F. Jabbour. Herbarium-based science in the twenty-first century, 2018. A. N. de la Hidalga, P. L. Rosin, X. Sun, A. Bogaerts, N. De Meeter, S. De Smedt, M. S. van Schijndel, P. Van Wambeke, and Q. Groom. Designing an herbarium digitisation workflow with built-in image quality management. Biodiversity Data Journal, 8:e47051, 2020. S. De Smedt, A. Bogaerts, N. De Meeter, M. Dillen, H. Engledow, P. Van Wambeke, F. Leliaert, and Q. Groom. Ten lessons learned from the mass digitisation of a herbarium collection. PhytoKeys, 244:23, 2024. K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. J. M. Heberling and B. L. Isaac. Herbarium specimens as exaptations. American Journal of Botany, 104(7):963–965, 2017. S. A. James, P. S. Soltis, L. Belbin, A. D. Chapman, G. Nelson, D. L. Paul, and M. Collins. Herbarium data: Global biodiversity and societal botanical needs for novel research. Applications in plant sciences, 6(2):e1024, 2018. G. Jocher and J. Qiu. Ultralytics yolo11, 2024. URL https://github.com/ultralytics /ultralytics. 17
N. Kanopoulos, N. Vasanthavada, and R. L. Baker. Design of an image edge detection filter using the sobel operator. IEEE Journal of solid-state circuits, 23(2):358–367, 1988. W. E. Lorensen and H. E. Cline. Marching cubes: A high resolution 3d surface construction algorithm. ACM SIGGRAPH Computer Graphics, 21(4):163–169, 1987. A. S. MacDougall, J. A. Loo, S. R. Clayden, J. Goltz, and H. Hinds. Defining conservation priorities for plant taxa in southeastern new brunswick, canada using herbarium records. Biological conservation, 86(3):325–338, 1998. P. Niedzielski and J. Markiewicz. Digitalization of herbarium collections as a tool for the commercialization of scientific knowledge. Procedia Computer Science, 225:2194–2203, 2023. A. Novikov and V. Nachychko. Some notes on the digitisation workflow at the lws herbarium. ARPHA Preprints, 6:e148553, 2025. N. Nualart, N. Ib´a˜nez, I. Soriano, and J. L´opez-Pujol. Assessing the relevance of herbarium collections as tools for conservation biology. The Botanical Review, 83:303–325, 2017. F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011. J. Redmon, S. Divvala, R. Girshick, and A. Farhadi. You only look once: Unified, realtime object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 779–788, 2016. S. Ren, K. He, R. Girshick, and J. Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence, 39(6):1137–1149, 2016. V. S. Smith, H. Hardy, T. Wainwright, L. Livermore, N. Fraser, J. Horak, J. Aspinall, and M. Howe. Harnessing the power of natural science collections: a blueprint for the uk. Natural History Museum, pages 1–28, 2022. doi: https://doi.org/10.5281/zenodo.647 2239. P. S. Soltis. Digitization of herbaria enables novel research. American journal of botany, 104(9):1281–1284, 2017. F. Xu, H. Uszkoreit, Y. Du, W. Fan, D. Zhao, and J. Zhu. Explainable ai: A brief survey on history, research areas, approaches and challenges. In Natural language processing and Chinese computing: 8th cCF international conference, NLPCC 2019, dunhuang, China, October 9–14, 2019, proceedings, part II 8, pages 563–574. Springer, 2019. 18
Z. Zou, K. Chen, Z. Shi, Y. Guo, and J. Ye. Object detection in 20 years: A survey. Proceedings of the IEEE, 111(3):257–276, 2023. 19