scieee AI-readable full text Open interactive document viewer

Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks

Broustail, Danaé; Lipták, Panna; Saiz, Pablo; Le Meur, Jean-Yves

Abstract

We present the digital integration of the Tracks Images collection, a recently digitized set of photographs documenting CERN’s early history through tracks captured in Hydrogen Bubble Chambers, into the CERN Document Server. Our workflow incorporated insights from the collection’s analog envelope structure into MARCXML records using the OpenRefine framework. MARCXML records were constructed and enriched with basic metadata fields, while an image similarity algorithm was applied to identify and cluster near-duplicate scans arising from the multiplicity of negatives. The resulting clustered dataset was used to generate representative records for upload. In total, 6365 image records covering 5339 envelopes were published on the CERN Document Server, ensuring their long-term preservation and visibility. This project presents a pipeline for large-scale digitized image collection processing and lays the groundwork for future improvements in similarity benchmarking, keyword-based metadata enrichment, and expert-driven curation.

Full text

Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks August 2025 AUTHOR(S): Danaé Broustail ETH Zürich SUPERVISOR(S): Panna Lipták Pablo Saiz Jean-Yves Le Meur CERN openlab Report 2025 PROJECT SPECIFICATION Tens of thousands of images of particles tracks from the 1970s have been digitized in 2024 and stored on the CERN Data Cloud (EOS). They are grouped into folders corresponding to the envelopes in which they were stored and are named with the information printed on the back of each photo. The student will have to analyze the set of images, design a process to create albums and records, and finally implement the programs enabling the upload of all the photos into a public collection of the CERN Document Server. It will therefore give these images a new audience, as they will become available online. Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 1 CERN openlab Report 2025 ABSTRACT We present the digital integration of the Tracks Images collection, a recently digitized set of photographs documenting CERN’s early history through tracks captured in Hydrogen Bubble Chambers, into the CERN Document Server. Our workflow incorporated insights from the collection’s analog envelope structure into MARCXML records using the OpenRefine framework. MARCXML records were constructed and enriched with basic metadata fields, while an image similarity algorithm was applied to identify and cluster near-duplicate scans arising from the multiplicity of negatives. The resulting clustered dataset was used to generate representative records for upload. In total, 6365 image records covering 5339 envelopes were published on the CERN Document Server, ensuring their long-term preservation and visibility. This project presents a pipeline for large-scale digitized image collection processing and lays the groundwork for future improvements in similarity benchmarking, keyword-based metadata enrichment, and expert-driven curation. Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 2 CERN openlab Report 2025 TABLE OF CONTENTS 1 Introduction 5 1.1 The particle tracks collection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 1.2 The CERN Document Server and the MARC21 format . . . . . . . . . . . . . . 6 1.3 Goalandmotivation ................................. 7 2 Data exploration 8 2.1 Description of the collection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 2.2 Metadataexploration................................. 9 2.2.1 Available metadata for envelopes . . . . . . . . . . . . . . . . . . . . . . 9 2.2.2 Metadataforimages ............................. 10 2.2.3 Metadata insights from the archives . . . . . . . . . . . . . . . . . . . . . 10 2.3 Challengesanddecisions............................... 12 3 Metadata extraction from TIF EOS file path 14 3.1 Renamingadjustments................................ 14 3.1.1 Resolving inconsistent filenames . . . . . . . . . . . . . . . . . . . . . . . 14 3.1.2 Resolving mapping between filename to envelope ID . . . . . . . . . . . . 15 3.2 MARCfieldextraction................................ 15 4 Data transformations for MARC fields 17 4.1 Checksumextraction................................. 17 4.2 Imagesimilarity.................................... 17 4.2.1 Intuition.................................... 18 4.2.2 SimilarityAlgorithm ............................. 18 4.2.3 Generating similarity scores . . . . . . . . . . . . . . . . . . . . . . . . . 19 4.2.4 Clusteringimages............................... 19 4.2.5 Threshold determination for clustering . . . . . . . . . . . . . . . . . . . 21 4.2.6 URL references for images in the same envelope . . . . . . . . . . . . . . 22 4.3 JPEGcreation .................................... 22 4.3.1 JPGcreation ................................. 22 4.3.2 MARC Fulltext File Transfer formatting . . . . . . . . . . . . . . . . . . 23 5 Upload to the CERN Document Server 24 5.1 Finalizing the MARC XML snippets for the collection . . . . . . . . . . . . . . . 24 5.2 MovingtheTIFfiles ................................. 24 5.3 CDSUpload...................................... 25 6 Conclusion 26 6.1 Alternative similarity approaches . . . . . . . . . . . . . . . . . . . . . . . . . . 26 6.2 Correcting scanning mistakes . . . . . . . . . . . . . . . . . . . . . . . . . . . . 27 6.3 Keyword extraction for metadata enrichment . . . . . . . . . . . . . . . . . . . . 27 6.4 Metadatacuration .................................. 28 7 References 29 Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 3 CERN openlab Report 2025 8 Appendix 30 8.1 Renaming inconsistent filenames . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 8.2 Resolving mapping from filename to envelope ID . . . . . . . . . . . . . . . . . . 30 8.3 Duplicatescans.................................... 31 Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 4 CERN openlab Report 2025 (a) The compactus layout of the archives. Photo by the CERN Scientific Information Service [13] (b) Storage units (cardboard boxes, binders, shelves) in a compactus. Photo by Danaé Broustail. Figure 1: The CERN archives. 1 Introduction Images, videos, and audio recordings have been accumulated at CERN since its foundation in 1954. These analog materials were initially stored in cardboard boxes, distributed across various storage rooms at CERN. Over time, they were progressively identified and transferred to compacti in the archives (see Figure 1), where particular care was taken to preserve their physical integrity through the use of non-acidic storage materials and controlled humidity and temperature conditions [13]. Analog materials such as tapes, photo slides, and negatives carry a unique historical value: they document the story of CERN’s predecessors, and serve as a testament to the evolution of engineering and physics to fulfill CERN’s mission. Since 2016, several preservation phases have been carried out to digitize these collections, with a dual objective: to conserve unique historical data, preventing deterioration and loss of information, and to share CERN’s rich history to a wider audience. This work is embedded in the larger-scale Digital Memory project [3], conducted by the Institutional Repositories section in CERN’s Information Technology department. 1.1 The particle tracks collection The particle tracks collection, which is the main focus of this report, is one of the last analog collections to undergo digitization. Its digitization was delayed because the collection lacked detailed captions and a clear historical context. In 2024, the collection was scanned by ScanCorner India, a company specializing in digitizing analog photo formats. The original photos were stored in labeled envelopes or taped into binders, often with a handwritten date beneath each image (see Figure 3). The collection contains highly heterogeneous material spanning several decades of CERN history: color and grayscale photographs of personnel, sites, and machines; experimental apparatus diagrams; hand-drawn graphs; handwritten notes; mathematical proofs; and, at its core, the particle tracks themselves (see Figure 2). The particle tracks were produced by hydrogen bubble chamber experiments [14]. The hyCreating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 5 CERN openlab Report 2025 Figure 2: Preview of the particle tracks collection drogen bubble chamber, used in the 1950s and 1960s, was a particle detector in which collisions were photographed as bubble trails left by particles moving through superheated liquid hydrogen. This manual analysis method was later superseded by the multiwire proportional chamber, developed by Georges Charpak in 1968 [4], which marked the transition to computer-assisted processing of particle collisions. The next crucial steps are to explore, analyze, and process the images in this collection to ensure their long-term digital preservation. With the aim of making the collection accessible to the public, the CERN Document Server emerges as the natural platform for endeavor. 1.2 The CERN Document Server and the MARC21 format The CERN Document Server (CDS) [8] is CERN’s institutional repository, providing access to digitized archives and multimedia resources. It holds the CERN Archives, a vast collection of reports, articles, photos, and videos, and the Pauli Archives, which contain Wolfgang Pauli’s personal manuscripts and correspondence. Our photos were uploaded to the CERN Archives in the CERN Photolab category (see Figure 4), which is an author category for all analog photos created at CERN. On CDS, each document is represented as a record: for example, a single image corresponds to one record. A grouping between a set of records constitutes an album record. In this system, images are preserved digitally as record instances. The MARC21 bibliographic format [1], developed by the Library of Congress, is used internally by CDS to represent the metadata attached to these records. Currently, CDS is undergoing a migration to the new CDS, based on InvenioRDM, which will no longer rely on MARC21 [15]. However, since the archive collection has not yet been migrated to the new CDS, our collection was uploaded to the current CDS system. Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 6 CERN openlab Report 2025 Figure 3: Storage structure of the collection. Envelope labels are a combination of numbers and letters. They typically contained multiple negatives of the same image, either in medium format or in film format (26 x 35 mm). Photos were found in binders and their label was a date caption. Photos by Danaé Broustail. Figure 4: The CERN Photolab page on CDS. 1.3 Goal and motivation Our aim is to upload the heterogeneous, metadata-limited particle tracks collection to CDS, both to visually showcase CERN’s heritage and to enrich these digitized images with as much contextual information as possible, transforming them into contextualized digital records. Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 7 CERN openlab Report 2025 (a) The folder structure of the digitized collection, showing the digital image contents of envelope 10074. The scans of envelope 10074 are almost identical due to the presence of several negatives within the same envelope. (b) Negatives (film and medium format) stored in envelope 10074. Figure 5: Comparing the structure of the analog and digital collections, with envelope 10074. In the digital collection, each negative scan is uniquely identified by its envelope number and its number suffix in the envelopeNb_tirageNb format, where envelopeNb is the envelope number and tirageNb is the print number, also known as the tirage number. A photo of the envelope label is also provided, with a filename format of envelopeNb_Label. Photos by Danaé Broustail. 2 Data exploration In this section, we will describe the exploratory data analysis of the collection. We first explored the collection’s structure and existing metadata. We then investigated correlations between the digitized collection and its analog counterpart stored in the CERN Archives. Finally, we refined the scope of the project following these insights. 2.1 Description of the collection The photos of the collection were scanned and digitized into TIF images. The TIF format is an image file format known for its ability to store high-quality images, making it a suitable choice for archiving. This choice led to file sizes in the range of 60 MB, making it impossible to download the full dataset locally. We used the Linux Public Login User Service (LXPLUS [10]), a cluster of Linux PCs provided by the CERN IT department, to access the images. We accessed the images remotely by mounting the collection on LXPLUS from its storage directory on EOS (see Code Snippet 1). mkdir PHOTOS eosxd -ofsname=eosmedia.cern.ch:/eos/media/cds/public/www/digital-memory/media-archive/ image/ScanCorner-mirror/PHOTOS PHOTOS,→ cd PHOTOS Listing 1: Mount command for the collection dataset. Executed from the command line of my LXPLUS user space. The digitized collection was uploaded to EOS in nine batches, each named according to "PHOTOS.N" where N was a number ranging from 1 to 9. A PHOTOS batch contained album Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 8 CERN openlab Report 2025 3.1.2 Resolving mapping between filename to envelope ID We aimed to ensure that album folder names extracted from the TIF file paths matched exactly with the envelope names in the Media Description table, so that the corresponding media information could be included in the MARC fields of the image records. To check for the existence of such mappings, we used a General Refine Expression Language (GREL) cross operation on the envelope name column in both tables (see Code Snippet 2). When no match was found, the cross operation returned a blank value, which allowed us to visualize missing envelope-to-file relationships between Media Description and the List of TIF table, and vice versa. cell.cross(List of TIF V3, 'TIF album ID').cells['TIF filename'].value.join(",") Listing 2: GREL expression used to retrieve the list of files per folder name. The transformation was applied on the envelope name column of Media Description and matched to the ’TIF Album ID’ column of the List of TIF file paths table. A missing match returned a blank value, while a match produced a comma-separated list of TIF files. Through this process, we identified 30 envelope names in Media Description that could not be mapped to the List of TIF table, and 82 files in the List of TIF table with no corresponding envelope name in Media Description. We prioritized resolving the 82 missing file-to-envelope relationships, as they pointed to image files that existed in the file system, whereas unmatched albums in Media Description were of lesser concern since they did not correspond to actual files. The issue was resolved by correcting inconsistencies in envelope names within Media Description, typically by adjusting a single character or replacing a separator (see Section 8.2 for additional details). This procedure resulted in matching all 5344 albums in the TIF file path list to a media description. 3.2 MARC field extraction To construct MARC snippets from the TIF file path and Media Description tables, we created each MARC field as a separate column in the TIF file path table. Since the CDS Invenio upload command bibupload (see CDS documentation [7]) required MARC snippets to be formatted in XML, we applied the XML formatting directly within the TIF file path table by wrapping the MARC column entries in XML syntax (see XML-ized MARC features in Figure 10b). This approach simplified the subsequent XML combination and export steps. The table in Figure 10a summarizes the MARC fields created, along with the transformations applied to generate them. Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 15 CERN openlab Report 2025 (a) Table showing a list of MARC fields that were added as columns to the TIF file path table, their meaning, and an intuitive description of the MARC extraction transformations. Constant MARC fields are indicated as such with "Invariant". Transformations were implemented in GREL and Python, and can be found on the OpenRefine table history (see Digitization GitLab [2]). A complete description of checksum extraction is provided in section 4. (b) MARC fields formatted as XML, for image record 01-4-56X_001. Figure 10: Description of transformations applied to the TIF file path table and Media Description tables in 10a. Illustration of resulting XML MARC fields for one image record in 10b. MARC fields are shown on both tables as XML-{datafield}_{subfield_list}, where datafield is the MARC datafield and subfield_list is the concatenated list of subfields. Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 16 CERN openlab Report 2025 4 Data transformations for MARC fields This section describes the Python scripts executed on the initial list of TIF files to complete the creation of MARC snippets. These transformations required accessing the image source files in order to (1) collect additional metadata fields and (2) generate intermediate files and computations needed for CDS upload. In particular, we computed image checksums, deduplicated similar images to identify distinct items in the collection, and created JPG versions of the filtered images for download and display on CDS (see Figure 11 for an overview of the pipeline). Figure 11: Overview of the image transformation pipeline. The left panel shows the flow common to the image similarity script compare_images.py, the image clustering script cluster_images.py, and the JPG conversion script convert_images.py. Inputs were PHOTOS-separated TIF path .xlsx files exported from OpenRefine. Log files were generated during the execution of each script for every PHOTOS batch. The right panel shows the outputs of the transformation scripts. The script generate_md5_checksums.sh was the only standalone script, independent of the TIF path lists. 4.1 Checksum extraction Checksums were added to the electronic access MARC field (XML-8564, subfield z) to uniquely identify the digital image content. We extracted image checksums using a Bash script that recursively searched for TIF files from the root folder of the collection and applied the md5sum command to each file. The results were stored in a .csv file indexed by the TIF file path. By using the TIF file path as a joining column, we incorporated the checksums into the XML-8564_uqyz column of the main TIF file path table, thereby completing the electronic access field. 4.2 Image similarity This section explains how we addressed the challenge of having near-identical images within the same album folder of our collection. Our goal was to design an approach that could both (1) assess the extent of the similarity problem and (2) provide intermediate metrics to guide decision-making and apply thresholds for selecting representative images from clusters of similar ones. Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 17 CERN openlab Report 2025 To avoid overloading CDS, we chose to upload only a relevant subset of distinct images. This selected fraction was then used to generate reference URLs pointing to other images in the same envelope. 4.2.1 Intuition Our intuition was to compute similarity scores algorithmically between pairs of images. Since we assumed that similar content appeared only within the same album folder due to the scanning procedure, we restricted comparisons to the folder level, which contained an average of five images (see Figure 14a). We interpreted the comparison scores as edge weights in a graph linking together all images within a folder. After thresholding these edge weights, similar images appeared in the same connected component of the similarity graph, while distinct images belonged to separate components. From this album folder graph structure, representative images could be selected for each component, thereby preserving a meaningful fraction of the collection (see Figure 12). Figure 12: Illustration of the image similarity logic for album folder 10074. Analog negatives in an envelope were digitized into near-identical images within the corresponding album folder. Pairwise comparisons within the folder can be interpreted as a similarity graph. Representative images were then selected for each connected component of the graph. 4.2.2 Similarity Algorithm Since our approach required comparing over 32000 images across 5400 albums, we prioritized a lightweight similarity solution available in Python. We selected ORB (Oriented FAST and Rotated BRIEF) [6], a feature detection and description algorithm implemented in the Open Source Computer Vision library (OpenCV) library. ORB is a fusion of two methods, the FAST keypoint detector and the BRIEF descriptor, making it both efficient and accurate for image matching. In ORB, keypoints are informative pixel locations detected in an image, while descriptors are binary vectors describing the neighborhood of each keypoint. We used ORB to match the descriptors of two images, for example, descriptors des1 and des2 from images image_1 and image_2. We used a Brute Force Matcher (BFMatcher) to match the two image descriptors with the Hamming distance, which is a well-suited distance Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 18 CERN openlab Report 2025 for binary vectors. The matcher produced a list of matches, linking a descriptor des1_i in image_1 to its closest descriptor des2_j in image_2. The average distance across all matches was normalized and converted to a similarity percentage (see Code Snippet 3). # 1. Get image descriptors kp1, des1 =orb.detectAndCompute(image_1, None) kp2, des2 =orb.detectAndCompute(image_2, None) # 2. Match descriptors bf =cv2.BFMatcher(cv2.NORM_HAMMING, crossCheck=True) matches =bf.match(des1, des2) # 3. Get similarity percentage distances =[m.distance for min matches] max_distance = 256 # Max Hamming distance for ORB similarity = 100 * (1-np.mean(distances) /max_distance) Listing 3: Similarity percentage computation between two images, using ORB by OpenCV. As image similarity is a long-studied problem in computer vision, we also qualitatively tested other approaches on toy examples with artificially introduced anomalies. In particular, we compared ORB with (1) histogram-based similarity (cv2.calcHist in OpenCV) and (2) the Structural Similarity Index (SSIM) from the scikit-image library. For each method, we verified whether thresholding could effectively separate anomalous images from a folder containing similar images. The choice of ORB was motivated by its accuracy and faster runtime in such examples, compared to the other two solutions. Ultimately, we adopted ORB as a proof of concept to demonstrate how similarity scores could be used to cluster near-identical images and retain only representative ones. Our focus was on building a working similarity-driven clustering framework rather than performing an exhaustive benchmarking of algorithms. For future work on benchmarking and image similarity analysis, the pipeline’s similarity score function, defined between two images to output a similarity percentage, can be replaced and tested with alternatives. The following sections describe the implementation of the comparison and clustering routines, which were based on the ORB similarity function. 4.2.3 Generating similarity scores We created similarity scores for all pairs of images within each album folder. We ran the script compare_images.py on the original TIF path list table to obtain an output comparison table, with the filenames for each pairwise image comparison, the album folder name that they belonged to, and, finally, the similarity percentage scores (see Figure 13). 4.2.4 Clustering images We used the script cluster_images.py to identify image groups and select representative images, working from the similarity scores generated in Section 4.2.3. Connected components were identified by constructing envelope-wise graph representations. Edges were added between two images whose similarity score exceeded a similarity threshold, meaning they could be considered near-identical. We then used a depth-first search traversal of the graph to visit the neighbors of the image nodes and find connected components of the graph. A decision rule was put in place to select the lowest file size image as the representative image of a component. This was motivated by (1) the fact that larger file sizes could induce Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 19 CERN openlab Report 2025 Figure 13: OpenRefine tables for similarity scores and clustering. Left panel: comparison table for all images in the collection. Since comparisons were made at the folder level and there was an average of five images for around 5400 folders, this created comparisons in the magnitude of (5+4+3+3+1)·5400 ≈80000, which roughly corresponds to the 70000 rows in the table. Right panel: clustering table with image filenames selected from components that were grouped from the similarity scores of the comparison table. The "Component ID" column shows the one (or more) connected component ID found within an album folder. "Images per cluster" gives the original count of images in the component, which were collapsed into one representative image indicated in "filename". Bottom panel: graph visualizations of the tables for an example folder containing three envelopes. The comparison table draws edges with similarity weights between images in the same album folder. The clustering table removes edges below a similarity threshold (82%) and identifies connected components, selecting one image per component. space limitations and hinder subsequent JPG creation and upload, and (2) the fact that larger images could capture irrelevant smudges. A clustering table was used to store information about the connected components and representative images chosen for each component (see Figure 13). Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 20 CERN openlab Report 2025 4.2.5 Threshold determination for clustering (a) Histogram of original file counts. File counts are equally distributed from one to nine images per album, with a peak at seven images. (b) Histogram of post-clustering file counts, with similarity threshold 82%. File counts are concentrated at one file per album. (c) Barplot of file count reduction per album, with a 70% similarity threshold. (d) Barplot of file count reduction per album, with a 82% similarity threshold. Figure 14: Visualization of clustering-induced file count reduction. Histograms in 14a and 14b show the number of images per album on the x-axis and the album count on the y-axis. Barplots in 14c and 14d compare the original file count (pink) and reduced post-clustering file count (purple) (y-axis) per album (x-axis). The effect of the threshold value for clustering was visualized through barplots created by evaluate_clustering.py, comparing original file counts to post-clustering file counts per album folder. By analyzing these plots, we could intuitively find a satisfactory threshold range for the clustering procedure (see Figure 14d and 14c). A threshold value of 70% (see Figure 14c) produced a flat reduction profile, erasing meaningful distinctions between images and collapsing almost all album folders into a single image. A lower threshold implies a more lenient identity criterion, which tends to merge all images into the same connected component. In contrast, a range of 82 to 85% yielded a spiked reduction profile (see Figure 14d), indicating that meaningful visual differences were preserved, while near-duplicates managed to be grouped together. We manually inspected five to ten album folders, including both near-duplicates and visually distinct images, to validate the choice of the threshold. Given a reasonable threshold value of 82%, validated by several album folder examples, we observed that the clustering procedure effectively reduced a significant fraction of album folders to a single representative file (see Figure 14b). This significant shift in the file count distribution justified our image deduplication approach, as it suggested that a large part of our collection contained duplicates. Finally, by taking the 6366 filenames from the 82% thresholded clustering table and ensuring the presence of the initial 5344 album folders, we refined our original TIF file path table to Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 21 CERN openlab Report 2025 include only the representative images identified in the clustering step. Album folders containing only label files were discarded, leaving 5339 album folders. The record IDs for this selection were then recomputed by incrementing from the first available record ID. The resulting list of TIF paths will be referred to as the 82%-threshold TIF file path table. 4.2.6 URL references for images in the same envelope In addition to the existing MARC fields, we sought to enhance user navigation of the collection by including links to other images from the same envelope. This approach allowed the collection to be intuitively explored following the original envelope structure. We based this operation on the 82%-threshold TIF file path table, as it contained the images that would be accessible on CDS as records. To implement this, we created the references MARC field (XML-773), with URLs pointing to image records from the same envelope, i.e. images originating from the same album folder in the digitized collection. Each URL was formatted as a CDS search query, targeting the report number of the image (XML-037_a) and the collection name, “Tracks Images.” For a given image record, the references section aggregated XML-773 fields, each containing a single search URL to another image in the same envelope (see XML snippet 4). <datafield tag="773" ind1=""ind2=" "><subfield code="p">Related image 01-4-56X_001</subfield><subfield code="u">https://cds.cern.ch/ search?cc=Tracks+Images&p=037__a%3A%2201-4-56X_001%22&of=hd</ subfield></datafield> ,→ ,→ ,→ <datafield tag="773" ind1=""ind2=" "><subfield code="p">Related image 01-4-56X_002</subfield><subfield code="u">https://cds.cern.ch/ search?cc=Tracks+Images&p=037__a%3A%2201-4-56X_002%22&of=hd</ subfield></datafield> ,→ ,→ ,→ Listing 4: Example of a MARC XML-773_up section for image 01-4-56X_003. The two data fields in this section correspond to the two other images from the same envelope: 01-4-56X_001 and 01-4-56X_002. The subfield p is the text of the link, and the subfield u is the CDS search URL, querying the image names. "&" was used in place of "&" and "%22" in place of spaces for XML URL formatting. 4.3 JPEG creation We converted our filtered TIF image collection into JPG files to enable the download option on CDS. This entailed creating 3 sizes of JPG images specified by widths of 180, 640 and 1440 pixels, to enable various download sizes on CDS. We first created a script convert_images.py to convert the TIF images into JPG images of the 3 required widths. Then, we incorporated the JPG file paths into an FFT ("fulltext file transfer") section in the image record MARC snippets to enable their upload to CDS. 4.3.1 JPG creation The script convert_images.py iterated over the 82%-threshold TIF file path table to load TIF images, resize them to widths 180, 640, and 1440, and finally save the resized images as JPG files in a temporary storage directory. We used a public CERN Andrew File System (AFS) workspace for that purpose. The conversion to JPG handled heterogeneous image modes: the RGB mode (8-bit color mode) and the L mode (8-bit grayscale mode) could directly be resized and saved as JPG, while Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 22 CERN openlab Report 2025 the I mode (16-bit grayscale mode) first needed to undergo pixel value normalization and L conversion before JPG conversion. The files were named according to the following format: reportNb_iconwidth.jpg, where reportNb is the report number (XML-037_a) and width is the pixel width of the image, one of 180, 640, or 1440. During the conversion process, we discovered that 10848X_003.tif (record ID 2491721) was a corrupted image because it could not be converted. It was therefore removed from the final list of uploaded files, but the initial record ID ordering of the TIF path list was preserved. 4.3.2 MARC Fulltext File Transfer formatting JPG paths were then included into an FFT section in the MARC snippet for upload to CDS (see XML snippet 5). <datafield tag="FFT" ind1=""ind2=" "><subfield code="a"> /afs/cern.ch/work/d/dbrousta/public/tracks_images_jpg/10850X_003_icon1440.jpg</subfield> <subfield code="t">Image</subfield><subfield code="d">Image 10850X_003</subfield><subfield code="f">jpg;icon-1440</subfield></datafield> ,→ ,→ ,→ <datafield tag="FFT" ind1=""ind2=" "><subfield code="a"> /afs/cern.ch/work/d/dbrousta/public/tracks_images_jpg/10850X_003_icon640.jpg</subfield> <subfield code="t">Image</subfield><subfield code="d">Image 10850X_003</subfield><subfield code="f">jpg;icon-640</subfield></datafield> ,→ ,→ ,→ <datafield tag="FFT" ind1=""ind2=" "><subfield code="a"> /afs/cern.ch/work/d/dbrousta/public/tracks_images_jpg/10850X_003_icon180.jpg</subfield> <subfield code="t">Image</subfield><subfield code="d">Image 10850X_003</subfield><subfield code="f">jpg;icon-180</subfield></datafield> ,→ ,→ ,→ Listing 5: MARC FFT section for image 10850X_003 for JPG upload to CDS. There are 3 data fields for each image size. Within a data field, subfield a contains the path of the image, subfield d contains the original image name, and subfield f is the width of the image in pixels. Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 23 CERN openlab Report 2025 5 Upload to the CERN Document Server After processing the images in our collection and extracting the required MARC fields, we took care of the upload of the 6365 images in the collection, spanning 5339 envelopes. 5.1 Finalizing the MARC XML snippets for the collection The final MARC XML was produced by concatenating all field components into a single string, by adding each feature sequentially and separating them by newlines for readability. Because XML does not rely on indentation, formatting was straightforward. Blank values, which were occasionally found in date or URL reference fields, were represented as empty strings. For each image record, a MARC XML snippet was generated by applying a GREL expression to the 82%-threshold TIF path list (see Code Snippet 6). Each snippet was enclosed within <record> and </record> tags to delimit the image record. '<record>'+ forEach([ cells['XML-001'].value, cells['XML-037__a'].value, cells['XML-110__a'].value, cells['XML-245__a'].value, cells['XML-260__c'].value, cells['XML-269__c'].value, cells['XML-500__a'].value, cells['XML-595__as'].value, cells['XML-596__a'].value, cells['XML-690__a'].value, cells['XML-773__up'].value, cells['XML-8564__uqyz'].value, cells['XML-FFT'].value, cells['XML-960__a'].value, cells['XML-980__a'].value, cells['XML-999__a'].value], xml, if(isNull(xml), '', xml.strip()+'\n')).join('') + '</record>',→ Listing 6: GREL expression to create the final MARC XML snippet. MARC XML snippets were divided into six batches of 1000 records. For each batch, we generated a corresponding .xml file, enclosed within <xml> and </xml> tags, containing the MARC XML snippets. These files then served as input for the upload commands. 5.2 Moving the TIF files Before executing the upload commands, the TIF files of the collection had to first be moved to the CDS image path. To avoid overwriting any existing CDS source files, we added the tracks/tiff folder on the CDS images path, where our collection would be transferred. We created cp commands for each image record in TIF file path table in the "Move command to cds" column to move the source TIF files to the CDS images directory under the tracks/tiff/directory. These commands were formatted as [! -e "destination_path" ⌋ ] && cp "source_path" "destination_path" || echo "source_path" >> skipped_file ⌋ s.txt, with destination_path as the CDS destination path, and source_path as the source TIF file path. Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 24 CERN openlab Report 2025 8.3 Duplicate scans Finally, we identified a small number of envelopes that had been scanned twice, resulting in duplicate digitized folders. This occurred for: •Envelopes 25292X and 25929X, •Envelopes 1-4-65X and 01-4-65X. Manual inspection of the envelope label photos confirmed that 25929X was correct, and 25292X was erroneous. Accordingly, the 25292X folder was deleted. For the second case, the Media Description table entry for 1-4-65X was corrected to 01-465X. No files were deleted, but rows referring to the incorrect path were removed from the TIF path table. Creating a collection of albums on the CERN Document Server from digitized historical photos of particle tracks 31