scieee AI-readable full text Open interactive document viewer

Open Raman spectral library for biomolecule identification

Teran, Marcelo; Ruiz González, José Javier; Loza-Alvarez, Pablo; Masip, David; Merino, David

Abstract

Raman spectroscopy combined with Multivariate Curve Resolution (MCR) analysis is widely used in biomedical applications. However, assignation of biomolecules to the components extracted by MCR can be challenging due to the absence of an open Raman spectral library for biomolecules. Raman experts typically identify unmixed component spectra as biomolecules by comparing them with reference spectra from the literature. This process can be time-consuming and subject to human bias. In this work, we created an open Raman spectral database with 140 biomolecules by implementing an algorithm to digitalize the spectra plots and most relevant peaks from articles available in the literature. Additionally, we implemented two search algorithms. The first one uses the spectral linear kernel or cosine similarity on the full spectra. The second algorithm is based on peak matching, and relies on the intersection over the union of the matched peaks with a defined tolerance for peak matching. Our experimental validation showed 100 % top 10 accuracy in molecule identification (e.g. collagen) and 100 % accuracy in molecule type identification (e.g. protein) in both pure biomolecule measurements and also when replicating results from prior studies. Objectively narrowing the identification to the top 10 ranked candidates and providing type identification can significantly reduce both the time required for visual identification and the need to purchase reference component samples. We publish our spectral library as an open-source tool so it can be expanded collaboratively by the research community. It is available at: https://github.com/mteranm/rama nbiolib.

Full text

Open Raman spectral library for biomolecule identification Marcelo Ter´ an a,*,1 , Jos´ e Javier Ruiz b,1 , Pablo Loza-Alvarez b , David Masip a , David Merino a a Faculty of Computer Science and Telecommunications, Universitat Oberta de Catalunya (UOC), 08018, Barcelona, Spain b ICFO-Institut de Ciencies Fotoniques, The Barcelona Institute of Science and Technology, Castelldefels, 08860, Barcelona, Spain ARTICLE INFO Keywords: Raman spectroscopy Spectral library Biomolecules Biomedicine Database Open-source ABSTRACT Raman spectroscopy combined with Multivariate Curve Resolution (MCR) analysis is widely used in biomedical applications. However, assignation of biomolecules to the components extracted by MCR can be challenging due to the absence of an open Raman spectral library for biomolecules. Raman experts typically identify unmixed component spectra as biomolecules by comparing them with reference spectra from the literature. This process can be time-consuming and subject to human bias. In this work, we created an open Raman spectral database with 140 biomolecules by implementing an algorithm to digitalize the spectra plots and most relevant peaks from articles available in the literature. Additionally, we implemented two search algorithms. The first one uses the spectral linear kernel or cosine similarity on the full spectra. The second algorithm is based on peak matching, and relies on the intersection over the union of the matched peaks with a defined tolerance for peak matching. Our experimental validation showed 100 % top 10 accuracy in molecule identification (e.g. collagen) and 100 % accuracy in molecule type identification (e.g. protein) in both pure biomolecule measurements and also when replicating results from prior studies. Objectively narrowing the identification to the top 10 ranked candidates and providing type identification can significantly reduce both the time required for visual identification and the need to purchase reference component samples. We publish our spectral library as an open-source tool so it can be expanded collaboratively by the research community. It is available at: https://github.com/mteranm/rama nbiolib. 1. Introduction Raman spectroscopy (RS) is an optical technique that uses the phenomenon of inelastic scattering of monochromatic radiation to provide vibrational fingerprints of the different types of molecules present in the sample under study [1]. In recent years, RS has gained attention as an analytical tool in biomedical applications due to its non-invasive and label-free characteristics [2,3]. RS does not require any specific sample preparation, making it particularly advantageous for real-time and in vivo analysis of biological systems without the need for biopsies. This makes it especially valuable in clinical diagnosis, as it enables the examination of tissue composition and the detection of disease markers without invasive procedures, reducing patient discomfort and eliminating the risks associated with sample collection [4]. Although RS can be used to study the chemical composition of samples formed by the mixture of multiple components, it is usually difficult to identify each of these components, specially when measuring complex mixtures, such as biological tissue [5]. Molecules with similar chemical characteristics generate vibrational bands that overlap in the same region of the Raman spectrum. This is usually happening when multiple biomolecules of the same type are present in the mixture; for instance proteins, lipids, carbohydrates or nucleic acids. In addition, Raman scattering signal is usually weak and is often masked by other effects, such as auto-fluorescence, which can be intense in cells and tissues [6]. Therefore, complex analytical techniques are required to extract the relevant information from spectral data. To overcome the overlapping challenge that mixtures present, the different spectra variables must be considered to identify which Raman bands or peaks belong to one molecule and which to others, which justifies the necessity of implementing multivariate analysis. Specifically, Multivariate Curve Resolution (MCR) analysis has been widely used to unmix Raman spectral data [2]. It performs a matrix factorization method that decomposes a matrix of spectra mixtures into a matrix of component spectra and another matrix with their relative abundances [7]. Also, it is usually non-supervised, which means that it does not need any a priori information regarding the chemical content of the sample. * Corresponding author. E-mail address: [email protected] (M. Ter´ an). 1 These authors contributed equally to this work. Contents lists available at ScienceDirect Chemometrics and Intelligent Laboratory Systems journal homepage: www.elsevier.com/locate/chemometrics https://doi.org/10.1016/j.chemolab.2025.105476 Received 18 March 2025; Received in revised form 19 May 2025; Accepted 27 June 2025 Chemometrics and Intelligent Laboratory Systems 264 (2025) 105476 Available online 27 June 2025 0169-7439/© 2025 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license ( http://creativecommons.org/licenses/bync-nd/4.0/ ). In addition, it is possible to impose meaningful constraints that improve its interpretation, as for instance non-negativity values in the spectra components and abundances it provides. Importantly, the first step towards a biochemical interpretation of spectral data analysis consists of assigning each spectral component provided by the MCR method to its corresponding molecule. This step is usually performed by comparing the most prominent bands in the unknown spectrum with published Raman spectra references and band positions of the biological molecules expected in the measurement. Often, this task can become challenging and requires Raman spectroscopy experience due to: the high complexity of mixtures, which leads to the decomposition of components that can be assigned to multiple biomolecules; the effect of background signal; and the variations of the spectra related to the specific setup used [8]. This task can be time-consuming, subject to human bias, and allows the validation of only a limited subset of biomolecule candidates. Although reference Raman spectra of some of the main biological molecules are reviewed and shared in the bibliography [9–12], currently, there is no open Raman spectra digital database for biological molecules, nor any open search library tool that can assist researchers in the identification process. This can be a limitation for the broad usage of Raman spectroscopy in biomedical analysis. An open Raman spectral library can provide a standardized tool for identifying the Raman spectra of biomolecules, which can be expanded through collaborative efforts by research community members. Spectral digital databases and search tools are available in other areas of application of Raman spectroscopy, such as minerals and painting pigments identification [13–16], and closer to the biological field, in microbiology, although in this case spectra do not specify biomolecules [17]. Considering this, in this work we aim to create a digital database containing the Raman spectra of the main biological molecules and to provide and develop an open source tool for Raman spectral search and identification. 2. Materials and methods 2.1. Extraction of the Raman spectra biological molecules reference database We have identified and digitalized already published Raman spectra. Even though spectra files are commonly not shared in the literature, two different types of Raman information can be commonly found in publications: plots of full Raman spectra or tables with the positions of the most relevant Raman bands or peaks, in wavenumbers (cm −1 ). Due to this situation, we created two different files in the database: one with the spectra, and another with a list of Raman peaks. In the Raman plots file, the Raman shift and the intensity values of the spectrum plot are included, while in the Peak’s positions database the list of relevant peaks positions and the list of relevant peaks positions intensity are included. An additional file contains the metadata of the spectra, such as the component type (lipid or protein, for instance), the DOI of the reference, sample information, and details of spectra acquisition and processing. The full list of metadata is defined in Supplementary Data Table ST1. All files are in the CSV format. Although original data is always desirable, and despite the possible loss in data quality, we extracted the spectra plots from the images of figures from articles using a custom-developed Python script based on classical computer vision contour detection techniques, using opencv library [18,19], see Supplementary Data S2 for more details. Since the spectra acquired by different authors contained different signal-to-noise ratio, fluorescence baseline, Raman spectral range, spectral resolution and intensity values, we standardized the database spectra applying the following preprocessing. 1. Smoothing by means of a Savitsky-Golay filter. 2. Baseline removal using airPLS algorithm [20], implemented in the BaselineRemoval library [21]. 3. Spectrum range standardization from 450 cm −1 to 1800 cm −1 with a step of 1 cm −1 , using linear interpolation. 4. Min-max normalization of intensity values between [0, 1]. We obtained the peak position information from articles tables by a custom-developed text processing code, in Python, from the PDF or HTML sources of the articles. A total of 17 articles were processed to analyze 202 sets of data, and 140 different molecule components, since some of the sets were repeated in 2 or more articles [9,10,12,22–34]. Table 1 shows the summary of the database components by type and the database full information is defined in Supplementary Data Table ST1. The database contains multiple spectrum entries for 36 components, see Table 2. Among them, 11 of the duplicated entries were measured by different authors using different spectrometers and data acquisition conditions. Additionally, 25 components were obtained by the same authors using the same spectrometer equipped with lasers operating at 488 nm and 532 nm. As it will be shown later, these repeated components will be used to evaluate the searching algorithm. We have made the database available on GitHub with the aim of enabling the research community to expand it by including additional Raman spectra from biological molecules that were not considered in the scope of this article. The guidelines of the contribution process can be found in Supplementary Data S4. 2.2. Raman spectra based biological molecules search algorithms Once the databases of Raman spectra have been built, here we propose two spectral search algorithms to assist in the identification of an unknown Raman spectrum from a biological molecule. The algorithms aim to provide a ranking of the similarity of the unknown spectra to each of the database spectra, reducing the identification to a manageable list of likely candidates. They were implemented in an open-source Python library available on Github 2 and through Python package installer. Additionally, a desktop application was developed to provide a Graphical User Interface (GUI) to the core library for non-technically adept users, which can be downloaded from GitHub. 2 2.2.1. Search algorithm based on Raman spectra plot similarity The first algorithm uses the Raman spectra plots. It relies on a similarity ranking between the unknown query Raman spectrum, Su, and each Raman spectrum of the database, Sdb, using cosine similarity (CS) or spectral linear kernel (SLK) as a similarity score. In the final ranking, we deduplicate the multiple occurrences of Sdb for a single component by retaining only the Sdb with the highest similarity to Su. CS is a general similarity metric used in a wide range of applications and data types [35]. It calculates the cosine of the angle between two vectors, in this case the query and database Raman spectra, which reflects their similarity in terms of direction rather than magnitude. In terms of Raman signal, it means that it is based on the shape of the spectrum, rather than on its absolute intensity. This is particularly useful, since absolute intensities are more prone to change if different experiment conditions are used, such as different laser power, acquisition time, molecular concentration, different system components or even a different system alignment. Considering the two spectra array Su and Sdb with the same length N, the CS is defined as: CS=Su⋅Sdb  Su⋅Su √ Sdb⋅Sdb √(1) Where Su⋅Sdb is the dot product between both spectra. 2 https://github.com/mteranm/ramanbiolib. M. Ter´ an et al. Chemometrics and Intelligent Laboratory Systems 264 (2025) 105476 2 SLK is a similarity metric that was specifically designed for Raman spectra similarity calculations by Khan et al. [36]. SLK takes into consideration the intensity at each wavenumber and its neighbouring points, within a window. They designed and evaluated SLK not only considering pure molecule spectra but also mixtures. The SLK point similarity is defined as: SimSLK(xi,zi)=xi.zi+∑ j=i+W j=i−W(xi−xj)(zi−zj)(2) Where x and z are the spectra values of Su and Sdb to compare and W is the window size, for which we selected a value of W =25 cm −1 to fit with the most relevant peaks in the database. The SLK score is defined as the sum of the point-to-point similarity: SLK(x,z)= ∑ i=N i=1 SimSLK(xi,zi)(3) Analogously, to align the SLK score range with the range of the CS, [−1, 1], we propose the following similarity normalization: SLK(x,z)norm =SLK(x,z)  SLK(x,x) √  SLK(z,z) √(4) The search consists of computing both similarities between the unknown query spectrum and all spectra from the dataset. Then, the components of the library are ranked by descending matching similarity order. 2.2.2. Search algorithm based on Raman spectra peak positions matching When dealing with components in the database that are described by means of a table with Raman peak positions list, we propose a peak matching (PM) algorithm that considers a symmetrical tolerance T in the matching of the peak positions from both the unknown query spectrum and database spectra. We defined a peak matching score to compare if two sets of peak positions contain the same peak positions considering a symmetrical tolerance T (cm −1 ). The score is called the intersection over the union ratio, IUR@T, and aims to provide a score that considers which peaks are matched and which are missed in both sets. Considering a spectrum Su, represented by a set of peak positions Pu, and a spectrum Sdb, represented by a set of peak positions Pdb, we defined the IUR@T as the number of peaks positions in the intersection of the subset of peaks in Pu that are also found in Pdb , considering T, over the number of peak positions in the union of the set of Pu and Pdb: IUR@T= Pu∩Pdb|T |Pu∪Pdb|(5) Based on our experimental observations, we selected a symmetrical tolerance T =5 cm −1 to compensate for the peak position shift that can occur due to different spectral resolution and calibration between different spectrometers and the spectra pixel images used to create the database. This value was found to provide a balance between correcting for instrumental variations and preserving the differentiation of spectral bands. The proposed algorithm is defined as follows. 1. For each Pdb in the database. a. Find the peaks of Pu that are within a distance T from a peak in Pdb. b. Deduplicate the cases where a single peak in Pdb was assigned to multiple peaks in Pu. The assignation combination that maximizes the number of matched peaks is used. c. Calculate the IUR score between Pu and Pdb. 2. Deduplicate the multiple occurrences of Pdb for a single component by retaining only the Pdb with the highest matching score to Pu. 3. Rank the components by descending matching score order Since this algorithm only uses the peaks position database, it is not affected by the variations in spectra peak’s shape and intensity due to different acquisition conditions [8]. The input of the peak matching algorithm is a set of the spectrum peak positions in wavenumber units. It is necessary to first extract the most prominent peaks of the spectrum plot, which we call in this article peak extraction. In this article we performed the peak extraction using SciPy find_peaks function [37]. SciPy find_peaks finds all local maxima by comparison of neighbouring values. A peak minimum prominence threshold was defined for each spectrum to filter out the less prominent peaks that are attributed to the noise of the signal. 2.2.3. Component type identification using k-NN We used the k-nearest neighbors (k-NN) algorithm for the molecular type identification by identifying the k-nearest neighbors of the database with respect to the unknown component spectrum [38]. To identify them, the CS, SLK and IUR@T similarities were considered as distance Table 1 Summary of the database components count by biomolecule type. Lipids Proteins Saccharides Amino Acids Primary Metabolites Nucleic Acids Pigments Other Count [reference] 48 [9,11,24,25,34] 30 [10,26,27,32,33] 26 [9,12,31] 13 [9,29,30] 10 [9] 7 [9,22,23] 2 [9,28] 4 [9] Table 2 Components with duplicated spectrum in the database. Type Component References Lipids Fatty acids myristic acid [9,11] oleic acid palmitic acid stearic acid Triglycerides trilinolein trilinolenin triolein Saccharides Monosaccharides amylopectin [9,12] amylose Polysaccharides d-(+)-xylose d-(−)-fructose Proteins carbonic anhydrase [10] collagen cytochrome c elastase ferritin glutathione transferase hemoglobin horseradish peroxidase albumin lactalbumin lectin major proteinase myoglobin papain pepsin pepsinogen superoxide dismutases thaumatin triosephosphate isomerase trypsin trypsin inhibitor trypsinogen ubiquitin xylanase α -chymotrypsinogen a M. Ter´ an et al. Chemometrics and Intelligent Laboratory Systems 264 (2025) 105476 3 metrics. The type of the unknown component was determined by majority voting among the k nearest neighbors’ types, and ties were resolved based on the type mean similarity. We selected a number of k = 5 nearest neighbors. 2.2.4. Search evaluation To evaluate the search algorithms, we determined if a database component entry, which is known to be present in the spectrum, appears at top 1, 5, and 10 elements of the search result. Additionally, we evaluated if we could provide type identification for the query spectrum by defining the following experiments. ● We identified a duplicated spectrum in the database, removed it from the database, and used the search algorithm expecting to find the duplicated spectrum that is still in the database amongst the top results. ● We measured Raman spectra from biomolecules pure samples, expecting to find the correct component identification in the top results. ● We used two open datasets of mixtures of saccharides and amino acids to assess component identification when using controlled aqueous mixtures [39–42]. We compared the direct identification in mixtures against the identification after MCR application. When identifying directly in mixtures, we evaluated the top 6 and top 10 accuracy of each component present in the mixture, since the mixtures can contain up to 6 biomolecules. We analyzed how the biomolecule concentration affects the results by defining 3 ranges of concentration – low, medium and high. In the saccharides dataset, the concentrations of 30 μ l, 75 μ l, and 120 μ l were respectively marked as low, medium and high. In the amino acids dataset, we split the concentration range in 3 subranges of equal length. ●We replicated published articles that assigned a biological component to measured spectra, where spectral unmixing algorithms were applied to biomolecules mixtures spectra. 3. Results and discussion 3.1. Search evaluation using components with duplicated spectra in the database In this experiment, our goal is to evaluate how Raman spectra of the same component acquired under different acquisition conditions affect the search performance. Table 3 shows the results of the search for each component with a duplicated spectrum in the database, which are indicated in Table 2. In the case of stearic acid, measurement with the 532 nm appears at first position in PM, but at 10th and 12th for CS and SLK respectively. Fig. 1 shows that the Raman spectra of the top-ranking components in the stearic acid search - a fatty acid - are actually very Table 3 Duplicated spectrum position in search results, for each duplicated spectrum in the database. Protein data was acquired with the same spectrometer, but with different laser wavelengths. The rest, lipids and saccharides, were acquired with different spectrometers and laser wavelengths. Type Component Component position CS SLK PM (IUR) 488 nm 532 nm 488 nm 532 nm 488 nm 532 nm Proteins carbonic anhydrase 1 1 1 1 3 1 collagen 1 1 1 1 1 20 cytochrome c6 3 3 3 1 3 elastase 1 1 1 5 7 7 ferritin 29 11 25 14 22 1 glutathione transferase 1 1 1 1 2 5 hemoglobin 5 4 5 3 2 13 horseradish peroxidase 1 1 2 1 1 125 human albumin 1 8 1 2 2 1 lactalbumin 1 1 1 1 7 1 lectin 3 2 1 1 25 3 major proteinase 1 1 1 1 1 1 myoglobin 5 2 5 3 3 3 papain 1 1 1 1 1 1 pepsin 1 1 9 1 6 34 pepsinogen 5 22 2 5 1 1 superoxide dismutases 1 1 1 1 3 2 thaumatin 2 1 8 1 20 11 triosephosphate isomerase 1 4 1 13 1 2 trypsin 2 2 2 2 5 5 trypsin inhibitor 1 1 1 1 1 2 trypsinogen 1 1 1 1 4 2 ubiquitin 1 1 1 1 2 2 xylanase 2 2 1 1 4 3 α -chymotrypsinogen a 1 1 1 1 6 12 Type Component 532 nm 785 nm 532 nm 785 nm 532 nm 785 nm Lipids myristic acid 2 3 1 3 3 1 oleic acid 2 3 2 3 9 1 palmitic acid 8 4 9 11 1 1 stearic acid 10 1 12 10 1 1 Type Component 785 nm 1064 nm 785 nm 1064 nm 785 nm 1064 nm Lipids trilinolein 1 1 1 1 2 2 trilinolenin 1 1 1 1 4 4 triolein 3 3 3 2 2 7 Saccharides amylopectin 2 4 1 3 2 2 amylose 3 4 1 4 2 4 d-(+)-xylose 1 1 1 1 1 1 d-(−)-fructose 1 1 1 1 1 1 M. Ter´ an et al. Chemometrics and Intelligent Laboratory Systems 264 (2025) 105476 4 similar and share many of the peaks. All of them are lipids, including 5 fatty acids and 4 triglycerides. Triglycerides are esters derived from glycerol attached to three fatty acids, so they exhibit Raman peaks similar to single fatty acids, explaining then this top-ranking result. The main differences between them are the position of weaker peaks or lack thereof, however they do share the strongest peaks. In this case, PM performs better because all peaks are equally considered. We evaluated how different laser wavelengths affect the search performance. We considered the 25 protein components with duplicated spectra in the database that were acquired by Rygula et al. [10], see Table 2, using the same spectrometer operating at both 488 nm and 532 nm. Fig. 2 shows examples of raw spectra compared to the spectra in the database, after applying the standardization preprocessing defined in section 2.1 for ferritin, horseradish peroxidase and hemoglobin. The results exemplify how Raman spectra, and therefore the search results, can be heavily affected by the laser wavelength. These iron-containing molecules exhibit different Raman resonance modes depending on the different laser wavelengths [10]. This introduces differences in spectra in addition to a different fluorescence baselines. Moreover, component identification relies heavily on peak extraction. In this example, it is performed by the authors with criteria that can be complex. The different criteria in peak extraction can contribute to differences in the results. An example of this phenomenon can be seen in the case of Fig. 1. Comparison of query spectrum and the top 4 results for stearic acid (532 nm) search using the methods: PM (left), SLK (center) and CS (right). In the results plot, spectra are ranked by results position, with the highest similarity at the bottom. Fig. 2. Comparison of the raw spectrum (top row) and final database spectrum after baseline correction (bottom row), for the spectra of ferritin (left), horseradish peroxidase (center) and hemoglobin (right) using 488 nm (in blue) and 532 nm (in red) lasers. Spectra were acquired by Rygula et al. [10]. The vertical dashed lines represent the database peaks, for each laser, while the triangular markers represent the extracted peaks from the query spectrum. The spectrum plot and peak position information is plotted for both acquisition wavelengths for easy comparison. (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.) M. Ter´ an et al. Chemometrics and Intelligent Laboratory Systems 264 (2025) 105476 5 horseradish peroxidase measured with the 532 nm laser. We observed significant shifts comparing the peak positions we extracted from its spectrum to the published peaks of the same component measured with the 488nn laser. These shifts reduced the accuracy of the search matching, as shown in Fig. 2. This situation can explain the results obtained in the identification of these molecules. Table 4 shows the component identification accuracy in the top 1, 5 and 10 and the type identification accuracy for each search algorithm. The results show an accuracy larger than 88 % at top 10 but lower than 59 % in the top 1. The difficulty of finding a component in the top 1 can be explained by the high similarity of the multiple spectra in the database, especially within the same molecule type. This similarity makes it harder to find a specific molecule in the top 1. However, component type identification accuracy is larger than 95 %. For this, CS and SLK obtained the best results. 3.2. Search evaluation using Raman spectra from samples with isolated pure biomolecules In this experiment we aim to analyze the search performance in the case of spectral data measured in the laboratory. We acquired spectra from samples of 6 molecules that exist in the database: collagen, albumin, cytochrome c, DNA, RNA and glycogen. Spectra can be seen in the supplementary materials, Figure SF2, the samples were generated following the information in the supplementary materials, Table ST3. The spectra was acquired using a Raman microscope (inVia Renishaw, Apply Innovation, Gloucestershire, UK) using a 532 nm diode laser and a 1800 lines/mm grating, with a spectral resolution of approximately 1.8 cm −1 . Spectra standardization preprocessing described in section 2.1 was used on this data. Most of the components were measured in a phosphate-buffered saline solution (PBS). Water in PBS shows a large peak at wavenumbers higher than 1600 cm −1 [43]. To analyze the effect this peak may have on the results, spectra were analyzed in two different sets: one that included wavenumbers up to 1800 cm −1 , and the same data considered only up to 1600 cm −1 , therefore excluding this water peak. Table 5 presents the search results of the experiment. Despite differences between the spectra in the database and those measured in the laboratory, the search algorithms can identify the unknown component at top 5 in most cases. Poorest performance is observed when using CS considering full data range. This can be attributed to the overlap of the characteristic Amide I band of proteins (1600-1690 cm −1 ) with the main peak of PBS beyond 1600 cm −1 [10]. SLK performs better than the rest at top 1 and matches the performance of PM at top 5 across all cases. For DNA, glycogen and RNA the challenge to identify the component in the first position arises from the presence of highly similar components of the same type in the database, as shown in Fig. 3. Table 6 summarizes the identification accuracy at top 1, 5, 10, along with the k-NN type identification accuracy, achieving a type identification higher than 83 % in all cases. 3.3. Search evaluation using biomolecules mixtures datasets In this experiment, we aim to evaluate how the algorithms can identify components within mixtures. To do this, we first searched directly the mixture spectra, then we used MCR to extract the components of the mixture, and searched for them in the database. Fig. 4 shows the results of the identification of the mixtures directly. The results show that the identification accuracy is directly related to the concentration of the biomolecule, with top 10 accuracy between 74.7 and 93.7 for molecules with high concentrations and between 36.5 and 50.2 for molecules with low concentrations. This results in overall top 6 and top 10 accuracies of less than 66.6 and 73.7 respectively. SLK shows better performance than the rest, which can be attributed to the inclusion of mixtures analysis in its design and evaluation [36]. Since the previous results depend heavily on the concentration of the molecules in the mixture, we analyzed how the identification performs after unmixing the spectra using MCR, see Supplementary Data S6. We believe that this is a common procedure and reflects the potential real use of our database. Table 7 shows the search results for each extracted component. The worst results observed in the saccharides dataset can be explained by the molecules’ similar chemical composition (e.g maltose is composed of two glucose molecules), and difference between the database saccharides, in solid state, and in the mixture dataset, in aqueous solution [12,40]. Table 8 presents the overall accuracy results for component and type identification. Accuracy at top 6 and top 10 for the MCR case shows higher values compared with the search directly in mixtures, showing the advantages of this method. Biomolecule type identification shows worse performance compared to other experiments, as amino acids type identification is especially difficult for k-NN due to large differences in spectra between various amino acids. 3.4. Search evaluation replicating published articles assignation This experiment aims at replicating results published in the literature. We identified published works where the authors assigned an unknown spectrum to a specific molecule. It is important to note that these works also published the spectra they identified as a specific molecule. This is not very common practice, and identification of these articles has been useful as a validation of the method presented here. Feng et al. [44] performed Raman spectral imaging on skin sections and used MCR to extract the spectrum of collagen, elastin, keratin, triolein, ceramide, melanin, water and DNA. They published the numerical data of the spectra. All of these components are present in our two databases. The extracted components were identified by measuring synthetic samples of the expected components and visually identifying the relevant Raman bands in the different spectra. Marro et al. [45] deconvoluted 6 components from retina cultures, but only one was assigned to a specific biomolecule that is present in both of our databases, namely phosphatidylcholine. The unknown spectrum was identified by visual inspection, comparing the Raman bands with published spectra. The spectrum of phosphatidylcholine used as unknown spectrum was extracted from the referenced article plot figure, using the method described in section S2. Fig. 5 illustrates the differences between the unmixed unknown spectra and the database spectra. Table 9 details the results for each component analyzed in this section. Collagen and DNA are found in the top 1 position for all metrics, despite noticeable differences in peak shapes and intensities between the unmixed unknown spectra and the database spectra. In contrast, in the case of ceramide, CS and SLK methods fail to identify the unknown component within the top 10. This may be due to peak shape and intensity differences and the presence of numerous similar spectra of the same type in the database. However, the PM method produced better results by relying exclusively on peak wavenumber position information. Phosphatidylcholine results highlight the challenges of peak interference due to MCR unmixing errors, where a protein peak at 1003 cm −1 causes the methods based on the full spectra plot to misidentify Phosphatidylcholine as protein instead of lipid. Table 10 shows the accuracy of finding the component assigned in Table 4 Top 1, 5 and 10 component finding accuracy and type k-NN accuracy, in percentage, for each duplicated spectrum in the database, for each method. Method Component Type Accuracy (%) k-NN Accuracy (%) Top 1 Top 5 Top 10 k =5 CS 55.55 90.27 95.83 97.22 SLK 59.72 87.50 93.05 97.22 PM (IUR) 34.72 77.77 87.50 95.83 M. Ter´ an et al. Chemometrics and Intelligent Laboratory Systems 264 (2025) 105476 6 the articles in the top 1, 5 and 10 of the search results. PM obtained the best results in the top 5 and 10. SLK performs slightly better than PM at top 1 but worse in the rest of the cases. CS results are worse compared with the previous methods. Type identification accuracy shows a value higher than 85 % in all cases, with a 100 % accuracy for PM. Therefore, it is observed again that type identification is successful even though component identification may not be as accurate. 3.5. Discussion We present a database of Raman spectra, as well as a method to identify biomolecules based on their Raman spectra. The method was tested in four different ways: using the duplicate spectra in the database, generating new laboratory measurements of pure samples, using public mixtures datasets, and replicating published articles assignation. These tests allowed us to assess how the methods proposed handle variations in peak shape and intensity caused by different acquisition conditions and unmixing errors. Despite a top 1 accuracy below 50 % in most cases, the top 10 accuracy and type identification accuracy exceeded 90 % for at least one search method across all experiments. Type identification alone is often sufficient when studying complex mixtures in cellular and tissue samples, as molecules of the same type typically share biological functions. Additionally, providing a list of the 10 most similar spectra can significantly reduce the time required for identification by narrowing the number of candidates to consider. When identifying components in mixtures, our algorithms achieved a top 10 accuracy of 93 %, using SLK, when the target biomolecules were present at high concentrations. However, performance declined notably at medium and low concentration. The results suggest that the algorithms are more effective when identifying unknown pure biomolecules or components extracted via MCR unmixing, rather than when applied directly to complex mixtures. The CS metric achieved its best performance when the query spectrum was free of background interference and could be attributed to a single biomolecule. However, its performance dropped significantly when peak shapes and intensities were substantially altered, as CS was not specifically designed for robust spectral comparison. The SLK score demonstrated robustness against background interferences and variations in peak ratios between two measurements of the same component. Nonetheless, it showed limitations when comparing highly similar spectra when only the least prominent peaks were different, as occurs with components of the same type, such as lipids and proteins. The PM method performed better than the rest when analyzing the MCRTable 5 Search results for spectra from isolated pure biomolecules using all the algorithms. Component Type Laser Component position Max Range =1600 cm −1 Max Range =1800 cm −1 CS PM (IUR) SLK CS PM (IUR) SLK Albumin Proteins 532 2 2 1 1 1 1 Collagen Proteins 532 1 1 1 1 2 1 Cytochrome C Proteins 532 2 1 1 2 1 2 DNA Nucleic Acids 532 1 2 1 5 2 1 Glycogen Saccharides 532 4 2 4 24 2 4 RNA Nucleic Acids 532 2 4 1 26 3 1 Fig. 3. Spectra in the top 3 results in DNA search (left) and top 4 results in glycogen search (right), for the DNA and glycogen measurements (bottom). Table 6 Accuracy of all methods finding the correct component at top 1, 5 and 10 and component type using k-NN, from isolated pure biomolecules spectra. Method Max Range =1600 cm −1 Max Range =1800 cm −1 Component Type (k-NN) Component Type (k-NN) Accuracy (%) Accuracy (%) Top 1 Top 5 Top 10 k =5 Top 1 Top 5 Top 10 k =5 CS 33.33 100 100 100 33.33 66.67 66.67 50 PM (IUR) 33.33 100 100 83.33 33.33 100 100 83.33 SLK 83 100 100 100 66.67 100 100 100 M. Ter´ an et al. Chemometrics and Intelligent Laboratory Systems 264 (2025) 105476 7 decomposed output spectra from real biological samples. That is because PM can ignore the peak shape and intensity ratio variations caused by the decomposition errors, since it relies only on the peak’s position. However, the peak extraction stage is heavily affected by noise and peak position shifts, which needs to be mitigated with a refined tolerance value and peak detection method. 4. Conclusions We constructed an open Raman spectral library for biomolecule identification from the literature. We have made the database available on GitHub. We also present some algorithms for spectral data extraction that we have made openly available on GitHub. The digital database presented can be broadened by the research community including additional Raman spectra from biological molecules that were not considered in the scope of this article. We also present different fast algorithms that identify spectra in the database that may match unknown spectral data. The use of these algorithms and databases can help significantly narrow the matching candidates to the top 10 ranked spectra. The algorithm can also provide type identification, significantly reducing the time required for visual identification and the need to purchase reference component samples. The results highlight the library’s potential to ease Raman spectral analysis for biological molecule identification, enhancing wider biomedical applications. CRediT authorship contribution statement Marcelo Ter´ an: Writing – original draft, Visualization, Validation, Software, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Jos´ e Javier Ruiz: Writing – review & editing, Validation, Methodology, Investigation, Formal analysis, Conceptualization. Pablo Loza-Alvarez: Writing – review & editing, Supervision, Resources, Project administration, Funding acquisition, Conceptualization. David Masip: Writing – review & editing, Supervision, Resources, Project administration, Funding acquisition, Conceptualization. David Merino: Writing – review & editing, Supervision, Resources, Project administration, Funding acquisition, Conceptualization. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence Fig. 4. (a) Average top 6 and top 10 identification accuracy by concentration in saccharides and amino acids mixtures datasets. (b) Overall average top 6 and top 10 identification accuracy by similarity method. Table 7 Search results for spectra from saccharides and amino acids mixtures datasets after MCR extraction using all the algorithms. Dataset Component Component position CS PM (IUR) SLK Amino Acids Alanine 1 1 1 Asparagine 2 1 7 Aspartic acid 1 1 1 Glucosamine 1 15 1 Glutamic acid 1 27 1 Histidine 1 1 1 Saccharides Fructose 2 1 1 Glucose 2 5 2 Maltose 4 78 5 Sucrose 9 8 25 Table 8 Accuracy of all methods finding the correct component at top 1, 6 and 10 and component type using k-NN, from saccharides and amino acids mixtures datasets after MCR extraction. Method Component Type Accuracy (%) k-NN Accuracy (%) Top 1 Top 6 Top 10 k =5 CS 50.0 90.0 100.0 70.0 SLK 60.0 80.0 90.0 70.0 PM (IUR) 50.0 60.0 70.0 50.0 M. Ter´ an et al. Chemometrics and Intelligent Laboratory Systems 264 (2025) 105476 8 the work reported in this paper. Acknowledgements The authors acknowledge funding from Fundaci´ o CELLEX; Ministerio de Economía y Competitividad - Severo Ochoa programme for Centres of Excellence in R&D (CEX2019-000910-S); CERCA programme (999619436); Laserlab-Europe (871124); Ministerio de Ciencia e Innovaci´ on PID2021-122807OB-C31 and PID2022-138721NBI00 projects funded by MCIN/AEI/10.13039/501100011033/FEDER, UE; CARET project. The SLN facility corresponds to a “Grup reconegut” 2021 SGR 01456 Departament de Recerca i Universitats de la Generalitat de Catalunya. Appendix A. Supplementary data Supplementary data to this article can be found online at https://doi. org/10.1016/j.chemolab.2025.105476. Data availability Data available on GitHub: https://github. com/mteranm/ramanbiolib References [1] R.R. Jones, D.C. Hooper, L. Zhang, D. Wolverson, V.K. Valev, Raman techniques: fundamentals and frontiers, Nanoscale Res. Lett. 14 (2019) 231, https://doi.org/ 10.1186/s11671-019-3039-2. [2] H. Noothalapati, K. Iwasaki, T. Yamamoto, Biological and medical applications of multivariate curve resolution assisted Raman spectroscopy, Anal. Sci. 33 (2017) 15–22, https://doi.org/10.2116/analsci.33.15. [3] S. Cui, S. Zhang, S. Yue, Raman spectroscopy and imaging for cancer diagnosis, J. Healthc. Eng. 2018 (2018) e8619342, https://doi.org/10.1155/2018/8619342. [4] I.P. Santos, E.M. Barroso, T.C.B. Schut, P.J. Caspers, C.G.F. van Lanschot, D.- H. Choi, M.F. van der Kamp, R.W.H. Smits, R. van Doorn, R.M. Verdijk, V.N. Hegt, J.H. von der Thüsen, C.H.M. van Deurzen, L.B. Koppert, G.J.L.H. van Leenders, P. C. Ewing-Graham, H.C. van Doorn, C.M.F. Dirven, M.B. Busstra, J. Hardillo, A. Sewnaik, I. ten Hove, H. Mast, D.A. Monserez, C. Meeuwis, T. Nijsten, E. B. Wolvius, R.J.B. de Jong, G.J. Puppels, S. Koljenovi´ c, Raman spectroscopy for cancer detection and cancer surgery guidance: translation to the clinics, Analyst 142 (2017) 3025–3047, https://doi.org/10.1039/C7AN00957G. [5] N. Kuhar, S. Sil, T. Verma, S. Umapathy, Challenges in application of Raman spectroscopy to biology and materials, RSC Adv. 8 (2018) 25888–25908, https:// doi.org/10.1039/C8RA04491K. [6] R. teja Vulchi, V. Morgunov, R. Junjuri, T. Bocklitz, Artifacts and anomalies in raman spectroscopy: a review on origins and correction procedures, Molecules 29 (2024) 4748, https://doi.org/10.3390/molecules29194748. Fig. 5. Comparison between the component spectrum extracted in the evaluated articles (in blue) and the spectrum of the same component in the database (in red). (For interpretation of the references to colour in this figure legend, the reader is referred to the Web version of this article.) Table 9 Component matching coincidence position for spectra published in selected articles. Reference Component Type Component position CS SLK PM (IUR) Feng et al. [44] Ceramide Lipids 14 15 5 Triolein Lipids 1 1 1 DNA Nucleic Acids 1 1 1 Collagen Proteins 1 1 1 Elastin Proteins 24 1 3 Keratin Proteins 15 2 2 Marro et al. [45] Phosphatidylcholine Lipids 27 13 2 Table 10 Accuracy of all methods finding the correct component at top 1, 5 and 10 and component type using k-NN, from spectra published in selected articles. Method Component Type Accuracy (%) k-NN Accuracy (%) Top 1 Top 5 Top 10 k =5 CS 42.86 42.86 42.86 85.71 SLK 57.14 71.43 71.43 85.71 PM (IUR) 42.86 100.00 100.00 100.00 M. Ter´ an et al. Chemometrics and Intelligent Laboratory Systems 264 (2025) 105476 9