NanoToxRadar: A Multitarget Nano-QSAR Model for Predicting the Cytotoxicity Values of Multicomponent Nanoparticles
Full text
NanoToxRadar: A Multitarget Nano-QSAR Model for Predicting the Cytotoxicity Values of Multicomponent Nanoparticles Jaehyeon Park, Shahzad Rashid, Helena Copsey, Lang Tran, Alex Zabeo, Danail Hristozov, Giorgos P. Gakis, Costas Charitidis, Seokjoo Yoon, and Hyun Kil Shin* Cite This: ACS Nanosci. Au 2025, 5, 344−352 Read Online ACCESS Metrics & More Article Recommendations * sı Supporting Information ABSTRACT: Nanotechnological advances have led to the development of nanoparticles with complex structures. In this context, nano-QSAR models have been developed to assess toxicity; however, the applicability domain (AD) of such models is significantly restricted to specific types of nanoparticles (i.e., bare metal oxides, coated metals, or carbon-based nanomaterials) and target cell lines. Accordingly, NanoToxRadar, a web-based platform for predicting the toxicity of multicomponent nanoparticles (MC-NPs) toward various cell lines, was developed to extend the AD of the nano-QSAR model. The size-dependent electron-configuration fingerprint was used to represent the molecular structures of MC-NPs, and one-hot encoded cell types were used to predict toxicities toward 110 cell lines. The CatBoost regression model achieved good performance (R2Test = 0.877) and was deployed online (https://www.kitox.re.kr/nanotoxradar). The Web site takes the nanoparticle composition of the core as well as the shell, dopant, coating material, and diameter as inputs and predicts pIC50 values for 110 cell lines. KEYWORDS: Advanced nanomaterials, safe and sustainable by design, QNTR, computational nanotoxicology, nano descriptor, prediction web application The emergence of nanoparticles as key components of modern technology 1 has raised concerns about their potentially harmful effects on humans 2,3 and the environment. 4 Nanoparticle hazards, which are often attributed to their small size and large surface area, 5,6 have been continuously researched over the past few decades. 7 In addition to rigorous and time-consuming in vitro and in vivo toxicity-assessment methods, 8,9 in silico approaches, including quantitative nano structure−activity relationship (e.g., Nano-QSAR or QNAR) models have been developed to efficiently predict the toxicities of nanoparticles 10,11 as functions of their structural properties. However, these models have a limited applicability domain (AD) due to the limited amount of nanotoxicity data available for model building and the limited number of descriptors available for representing nanomaterials. 12,13 Given that advancing nanotechnology has led to increasingly complex nanoparticles, interest has shifted toward the toxicities of multicomponent nanoparticles (MC-NPs). 14 Therefore, models applicable to nanoparticles with highly complex compositions are also needed to efficiently screen the toxicities of MC-NPs. 15 Available descriptors for complex nanomaterials are calculated based on composition or size alone 16,17 except quantum mechanical (QM) and molecular dynamics (MD) descriptors. 18−21 Although QM and MD descriptors can precisely represent nanomaterials with complicated structures, both QM and MD calculations are computationally expensive. This becomes a bottleneck for screening large numbers of nanomaterials within a reasonable amount of time and for deploying a model publicly accessible to the research community, regardless of available computational resources. Moreover, many nano-QSAR models have been developed separately to target specific end points, such as cytotoxicity toward specific cell lines, 10,13,15 , 19,22,23 only fraction of data points was used in the model development. Machine learning model achieves better prediction accuracy when the model is trained using a larger volume of data. 17,24−26 Therefore, multitarget prediction is a better approach for increasing data size by integrating different target end points into the training set. 27−30 In this study, we developed a multitarget nano-QSAR model with an improved AD using a data set comprised of metalbased MC-NPs with and without surface modifications. The model was trained to predict the cytotoxicity of MC-NPs Received: April 13, 2025 Revised: August 5, 2025 Accepted: August 6, 2025 Published: August 8, 2025 Letterpubs.acs.org/nanoau © 2025 The Authors. Published by American Chemical Society 344 https://doi.org/10.1021/acsnanoscienceau.5c00035 ACS Nanosci. Au 2025, 5, 344−352 This article is licensed under CC-BY-NC-ND 4.0 Downloaded via 185.17.204.34 on November 3, 2025 at 11:59:50 (UTC). See https://pubs.acs.org/sharingguidelines for options on how to legitimately share published articles.
across 110 cell lines, demonstrating the possibility of a generalized model for prediction of MC-NPs’ cytotoxicities. The model’s AD was improved using the size-dependent electron-configuration fingerprint (SDEC FP) to represent the structures of the MC-NPs. 31 SDEC FP represents the electronic structures of the MC-NPs by estimating the electrons in each atomic orbital index. Computation of SDEC FP is significantly lower than QM and MD calculation since SDEC FP is prepared based on composition and size information.. The data set used for this study comprised 637 MC-NPs with cytotoxicity values (pIC50) measured for 110 cell lines. While SDEC FP can be used for carbon-based nanomaterials, this data set focused on various metal-based nanoparticles and even quantum dots with various surface modifications. Hence, carbon-based nanomaterials were outside the scope of this study. The model was trained using all of the cytotoxicity data by introducing cell features. The CatBoost regressor achieved best performance (R2Test: 0.88; RMSETest, 0.65). The optimal model was deployed on the web platform as “NanoToxRadar” (https://www.kitox.re.kr/nanotoxradar) for easy access by the research community. The web service requires composition (core, shell, dopants, and coating materials) and diameter to predict pIC50 values over 110 cell lines for a query NP. Therefore, unlike previous nano-QSAR models that were restricted to specific nanoparticle types and cell lines, and lacked easy accessibility, NanoToxRadar enables prediction of MC-NP cytotoxicity across 110 cell lines simultaneously with broader AD and easy web access. To ensure consistency, accuracy, and comprehensiveness, MC-NP data were collected using standardized data collection templates. The data then underwent quality control procedures to identify and correct any inconsistencies, missing information, or inaccuracies in accordance with the FAIR (findable, accessible, interoperable, reusable) principles. 32 Additional MC-NP data were obtained from the work of Oh et al., following the same data collection template (Table S1). 33 The collected MC-NP structural information included core, dopant, coating, and shell compositions, as well as diameter. Cell information included cell name, anatomical classification, source species, source tissue/organs, and cell-type category, such as primary cell, cell line, or bacteria, as described by Oh et al. 33 Further preprocessing steps were applied, which included converting cytotoxicity IC50 values (mol/L) into pIC50 (minus log base 10).Duplicate data points for identical MC-NPs were also removed if standard deviation of their pIC50 values was greater than 0.5. In cases where the standard deviation was less than 0.5, one data point was retained and labeled with the average pIC50 value. This preprocessing resulted in the removal of 284 data points, leaving 637 data points for model development. SDEC FP was calculated to represent the structures of the MC-NPs, which depend on both their composition and size. 31 Each SDEC FP has an atomic orbital index, and the value for each index is the estimated number of electrons. The SDEC FP was calculated as follows: 1) the number of atoms in the MCNP is estimated based on its composition and size. 2) the number of electrons in the MC-NP is estimated based on the number of atoms with the assumption that all electrons in MCNP are in the ground state. 3) Electrons have different spin numbers, which are distinguished by positive or negative values in the SDEC FP. As a result, each SDEC FP value has atomic orbital indices with positive or negative signs. 4) the number of electrons is then converted to a logarithmic scale. To predict cytotoxicity values across 110 cell lines, cell information was introduced into the model. Five different categories of cell information were used to prepare the cell features: cell name, anatomical classification, source species, source tissue/organ, and cell origin category. One-hot encoded vectors were prepared for each of these five categories of cell information, which included 110 cell names, 22 anatomical types, 13 source species, 34 source tissue/organs origins, and three types of cell origin. The number of values in each cell represents the feature size, as the cell information is converted into one-hot encoded vectors. The initial feature set, which comprised the SDEC FP and one-hot encoded vectors, was too large compared to the size of the data. Accordingly, the number of features was reduced to lower the possibility of model overfitting. The SDEC FP was subjected to feature compression rather than feature selection to minimize the loss of structural information. Three versions of SDEC FP were prepared: 1) a full-size SDEC FP without compression, 2) an aggregated SDEC FP obtained by adding up values for atomic orbital indices with identical energy levels, and 3) an aggregated SDEC FP that ignores the sign of the spin (i.e., without positive or negative signs). Cell features were selectively used instead of all of the one-hot encoded vectors from the information acquired from all five categories. One-hot encoded vectors were prepared for: 1) all five categories of cell information, 2) the cell name alone, 3) the cell name and source tissue/organs, and 4) the cell name and anatomical classification. MC-NP compositional complexity was visualized using the periodic table in the Periodic Trend Plotter Python library. 34 The diameters and pIC50 values of the MC-NPs were visualized to examine the correlation between diameter and pIC50. Cell diversity was examined to determine the distribution of each cytotoxicity target end point in the data set. The feature spaces of the training and test sets were compared using principal component analysis (PCA) as well as uniform manifold approximation and projection (UMAP). Machine learning algorithms, including the random forest regressor (RF), extra tree regressor (ET), extreme gradient boosting regressor (XGBoost), gradient boosting regressor, CatBoost regressor, support vector regressor (SVR), multilayer perceptron (MLP), and transformer (encoder block), were applied during model development using the PyTorch, 35 scikitlearn, 36 XGBoost, 37 and CatBoost 38 libraries in Python. Since a few cell names were associated with only one or two MC-NPs, data points with rare cell information (89 data points in total) were used exclusively in the training set and were excluded from internal and external validation. Internal validation was conducted using modified 3-fold cross validation (Figure S1). 39 This modification involved ensuring that MC-NPs with rare cell information were always included in the training data portion. This modified process was applied during hyperparameter optimization using Optuna. 40 External validation was performed using a test set that was isolated from the entire data set. Each machine learning algorithm was systematically optimized using algorithm-specific hyperparameters: CatBoost (learning_rate, depth, l2_leaf_reg, min_child_samples, border_count, thread_count), ExtraTrees (n_estimators, max_depth, min_samples_split, min_samples_leaf, bootstrap, max_features), XGB (n_estimators, max_depth, learning_rate, min_child_weight, subsample, colsample_bytree), SVR (C, epsilon, gamma), Gradient Boosting (n_estimators, learning_rate, max_depth, min_samples_split, min_samples_leaf, subsample), RandomForest (n_estimators, max_depth, min_ACS Nanoscience Au pubs.acs.org/nanoau Letter https://doi.org/10.1021/acsnanoscienceau.5c00035 ACS Nanosci. Au 2025, 5, 344−352 345
samples_split, min_samples_leaf, bootstrap, max_features), MLP (epochs, batch_size, learning_rate, weight_decay, dropout), and Transformer (d_model, nhead, num_layers, dropout, epochs, batch_size, learning_rate, weight_decay). The multitarget nano-QSAR model was deployed in a web environment at https://www.kitox.re.kr/nanotoxradar/. The Web site was implemented using the CodeIgniter framework (PHP), and the model runs in Python 3.7.9, taking SDEC FP and cell information as inputs. NanoToxRadar considers the core, shell, doping, and coating compositions, along with the doping ratio and diameter of the NPs, as essential information for calculating the SDEC FP. However, only the core and diameter are mandatory input values for making a prediction. When a user provides structural information for a query NP, NanoToxradar internally adds one-hot encoding vectors for 110 cell lines to the SDEC FP to calculate the pIC50 of all 110 cell lines. The predicted pIC50 values for each cell line are then visualized using a radar plot. Additionally, a GitHub repository (https://github.com/Jaehyeon-O-Ob/NanoToxRadar) was created to provide detailed instructions for the model. The data distribution was examined to understand the characteristics of the MC-NP data set (Figure 1). The heatmap of the cell-line distribution revealed that the data set includes 110 cell lines, among which 68 were tested for only one or two MC-NPs, accounting for 89 data points in total (Figure 1A). In the model, one-hot encoding was used to represent the cell features. The model cannot learn the appropriate weights for a cell line if its one-hot encoded vector is absent from the training set. This results in poor generalization for unseen instances of these rare cell lines. This underscores the critical importance of a strategic training-set design that ensures adequate representation of all feature categories, particularly for infrequent cell lines that might otherwise be overlooked using random sampling. The model’s predictive capability for the rare cell lines would be severely compromised without proper training data representation, leading to biased predictions that favor more common cell features. To address this, we ensured that the 89 data points containing rare cell information were always included in the training set, as random splitting could lead to these rare cell lines being completely absent from the training set. This process ensures the model can learn from rare cell information. The remaining 548 data points (total 637 data points without 89 data points of rare cell information) were then split to maintain an approximately 70:30 training-to-test ratio. This resulted in a final split of 445 data points in the training set (89 rare cell data points with 356 data points from the split data chunk) and 192 data points in the external test set. This approach maintained the desired ratio while ensuring proper representation of all 110 cell lines in the training process (Figure S1). Figure 1B shows the compositional complexity of the MC-NPs in the data set. The most common elements are oxygen, hydrogen, carbon, and sulfur, as many MC-NPs are coated with organic molecules. Zn and Cd are also common elements in the data set because they are frequently used in coatings, shells, dopants, and core materials. Figure 1C demonstrates the Figure 1. (A) Data distribution analysis for multicomponent nanoparticles in the data set. The heatmap visualizes the distribution of various cell lines across the data points. The color intensity corresponds to the frequency of cells in the data set, ranging from 1 to 75 nanoparticles. This is categorized into the top 45 cell lines (more than 2 nanoparticles) and others (with 1 nanoparticle each). (B) A periodic table highlighting the compositional complexity of the MC-NPs in the data set, with a color gradient indicating the occurrence frequency of each element. (C) A scatter plot showing the relationship between MC-NP diameter (1−700 nm) and cytotoxicity (pIC50). Higher cytotoxicity values are observed for MCNPs with diameters less than 5 nm. ACS Nanoscience Au pubs.acs.org/nanoau Letter https://doi.org/10.1021/acsnanoscienceau.5c00035 ACS Nanosci. Au 2025, 5, 344−352 346
correlation between pIC50 and diameter, showing that pIC50 increases exponentially with decreasing MC-NP size. In particular, highly cytotoxic MC-NPs are smaller than 5 nm. This size-dependent toxicity relationship demonstrates the importance of incorporating size information into our QSAR modeling approach. The clear correlation also indicates that size-related structural information is necessary for accurate cytotoxicity prediction across diverse MC-NPs, supporting our use of SDEC FPs. The pIC50 distribution was examined after the training and test sets had been split (Figure 2A), which revealed a similar pattern to that shown in Figure 1C, with values primarily distributed between 3 and 6. The range of MC-NPs in the data set (1 ∼700 nm) indicates that the model can give reliable predictions within the size range. Model overfitting is the most common issue encountered during model development. The best option is to divide the data set into three parts: a training set, a validation set, and an external test set. However, in current study, the data set was not sufficient to be divided into three; therefore, internal validation was used instead of hold-out validation approach. A model can become overfitted to the training set when the number of features is much larger than the size of the data set. 41 Therefore, we developed a model with fewer features to avoid overfitting. To determine the optimal number of features, we assessed different combinations of SDEC FP and one-hot encoded cell features using eight machine learning algorithms via the modified cross-validation. Modified 3-fold cross-validation was conducted for hyperparameter optimization, and external test set was used to validate the model and minimize risk of overfitting (Table 1). The SDEC FP is designed to be compressible without losing information by aggregating different indices. 31 Cell-information-based one-hot encoded vectors were generated based on the five categories of cell information, meaning that the cell feature size can also be reduced by selectively choosing which cell information to include for one-hot encoding. The feature space between the training and test sets was visualized using PCA and UMAP (Figures S2 and S3, respectively). The smallest feature set, which was 130 in size (composed of aggregated SDEC FP without spin and cell-name one-hot encoded vectors), showed good feature-space similarity between the training and test sets. With this feature set, the first principal component (PC1) accounted for 60.9% of the total variance, and the second principal component (PC2) captured an additional 24.8%, for a total of 85.7% of the variance (Figure 2B and 2C). Different combinations of SDEC FP and cell information features were processed by eight machine learning algorithms. To clearly present the model’s stability and performance, Table S2 provides validation results for all combinations. According to Table 1, SVR with aggregated SDEC FP and the cell-name one-hot encoded vector achieved the best prediction accuracy (R2Test = 0.89), which indicates that not all features are necessary for developing a multitarget nano-QSAR model to predict cytotoxicity values across 110 cell lines. Therefore, we developed a model with fewer features to avoid overfitting. The SVR models still exhibited good prediction accuracies close to those of the best SVR model with 130 features (R2Test = 0.882), even when the SDEC FP were further compressed. However, the SVR models yielded very similar predicted values across different cell lines for most MC-NPs (Figure S4), which indicates that SVR failed to learn discrepancies among cell information. SVR makes predictions based on support vectors (the actual data points in the training set) to define the margins within which errors are ignored. Since cell information is simply defined by one-hot encoding in the current data set, the cell information features alone can only weakly distinguish between cell lines. The weak classification power of one-hot encoding cannot provide sufficient information for capturing differences in the cytotoxicity values measured in different cells. Hence, SVR models make predictions that are heavily dependent on the SDEC FP rather than cell information, leading to similar pIC50 values being predicted regardless of cell line. The CatBoost model with 130 features (aggregated SDEC FP without spin and cell-name one-hot encoding) delivered the second-highest performance (R2Test = 0.877, Figure 3) after SVR. This implies that fewer features are sufficient for delivering good performance not only with SVR but also with other machine learning algorithms. The SDEC FP is designed to be reducible without information loss; hence, Figure 2. (A) Data-distribution analysis and feature-space visualization: training vs test sets. The histogram shows the distribution of end points (pIC50) across the training (blue) and test (red) data sets, demonstrating a balanced representation of end points. (B) A principal component analysis (PCA) visualization of the 130-feature data set, with the first two principal components explaining 85.7% of the total variance (PC1: 60.9%, PC2: 24.8%). Training (blue) and test (orange) sets show similar distribution patterns. (C) A uniform manifold approximation and projection (UMAP) plot for the 130-feature data set, revealing the underlying structure of the data in two dimensions, with consistent distribution patterns between the training (blue) and test (orange) data sets. ACS Nanoscience Au pubs.acs.org/nanoau Letter https://doi.org/10.1021/acsnanoscienceau.5c00035 ACS Nanosci. Au 2025, 5, 344−352 347
aggregated SDEC FP without spin remains capable of effectively differentiating the molecular structures of MCNPs. Five categories of cell information were available; however, one-hot encoding alone was sufficient to deliver good prediction accuracy. The CatBoost model was also capable of predicting differences in pIC50 values across different cell types (Figure S5). Thus, CatBoost with 130 features was selected as the optimal model and deployed as the web service (Figure 4). The transformer encoder block was also used to develop a model with 150 features (aggregated SDEC FP and cell-name one-hot encoding) and delivered R2test value of 0.867. The deep-learning model was not as effective as the machine learning models, presumably due to insufficient data size. Although 637 data points were considered as a large volume of data for nano-QSAR model development, this volume is not as large as the data sets in other fields where deep learning has performed well. Deep-learning models for small molecule design are often first pretrained with unlabeled data and then fine-tuned for the target end point. This approach helps overcome the limited size of data for the target end point by leveraging large databases such as Zinc 42 or PubChem. 43 However, current nanomaterial databases do not contain unlabeled nanomaterials with molecular-structural information. This lack of data hinders the use of self-supervised learning, which is important to fully leverage deep-learning algorithms for nanomaterials. Since the publication of the initial study by Fourches et al., 11 numerous nano-QSAR models have been developed; however, these models are rarely accessible to the research community. To predict toxicity of MC-NPs, QM and MD descriptors are typically used to model relationships between complex MCTable 1. Performance Metrics of Machine Learning Models with Different Feature Combinations Number of features Model RMSECV R 2 CV RMSETest R 2 Test RMSETest over end point range (%) Feature description a 314 CatBoost 0.602 ± 0.013 0.903 ± 0.005 0.623 0.886 6.49% All cell information & SDEC FP ExtraTrees 0.818 ± 0.032 0.820 ± 0.019 0.704 0.855 8.83% SVR 0.691 ± 0.038 0.871 ± 0.021 0.633 0.883 7.46% XGBoost 0.638 ± 0.073 0.889 ± 0.026 0.672 0.868 6.88% GBR 0.662 ± 0.022 0.883 ± 0.007 0.776 0.823 7.14% RandomForest 0.776 ± 0.021 0.839 ± 0.012 0.705 0.854 8.37% MLP 0.743 ± 0.009 0.852 ± 0.011 0.667 0.870 8.02% Transformer 0.717 ± 0.046 0.861 ± 0.026 0.720 0.848 7.74% 150 CatBoost 0.652 ± 0.047 0.885 ± 0.017 0.703 0.855 7.04% Cell names and aggregated SDEC FP ExtraTrees 0.885 ± 0.020 0.790 ± 0.011 0.750 0.835 9.55% SVR 0.727 ± 0.033 0.857 ± 0.020 0.617 0.888 7.85% XGBoost 0.698 ± 0.059 0.869 ± 0.021 0.691 0.860 7.53% GBR 0.698 ± 0.072 0.868 ± 0.028 0.69 0.860 7.53% RandomForest 0.833 ± 0.015 0.814 ± 0.002 0.736 0.841 8.99% MLP 0.778 ± 0.048 0.837 ± 0.025 0.760 0.831 8.39% Transformer 0.806 ± 0.049 0.824 ± 0.030 0.673 0.867 8.69% 130 CatBoost 0.691 ± 0.029 0.872 ± 0.005 0.649 0.877 7.45% Cell names and aggregated SDEC FP without spin ExtraTrees 0.987 ± 0.039 0.739 ± 0.013 0.817 0.804 10.65% SVR 0.775 ± 0.013 0.839 ± 0.012 0.635 0.882 8.36% XGBoost 0.725 ± 0.024 0.859 ± 0.006 0.697 0.858 7.83% GBR 0.706 ± 0.040 0.866 ± 0.010 0.703 0.855 7.62% RandomForest 0.859 ± 0.008 0.802 ± 0.011 0.728 0.845 9.27% MLP 0.822 ± 0.044 0.817 ± 0.029 0.744 0.838 8.87% Transformer 0.816 ± 0.022 0.821 ± 0.018 0.740 0.840 8.80% a Feature description: The full SDEC FP is 132 in size, and a one-hot encoded vector with five types of cell information is 182 in size. ACS Nanoscience Au pubs.acs.org/nanoau Letter https://doi.org/10.1021/acsnanoscienceau.5c00035 ACS Nanosci. Au 2025, 5, 344−352 348
NP structures and their toxicities. Preparing these descriptors is computationally expensive, and the molecular clusters used to represent nanomaterials are often poorly reproducible, which are major obstacles to model deployment. While the predicted values from a nano-QSAR model are accurate if the descriptors are computed correctly, QM and MD calculations are sensitive to molecular-cluster design. Slight changes in cluster design can lead to significant variations in descriptor values, potentially resulting in erroneous predictions for the target MC-NP. In particular, designing molecular clusters for MC-NPs to provide correct QM or MD descriptors is a tremendously complicated task. Consequently, many models cannot be reliably applied to MC-NPs due to limitations in descriptor preparation. In this study, we used SDEC FP, which is not computationally expensive. This allows the model to be deployed in a web environment and made publicly accessible to research communities. Furthermore, despite the simplicity of the SDEC FP, identical descriptor values are obtained for identical MC-NPs ensuring that the model’s predicted values are highly reproducible. NanoToxRadar was developed with a responsive web-design environment; thus, researchers can also use the model in mobile environments. In conclusion, we developed and deployed NanoToxRadar, a web-based platform for predicting nanotoxicity. The implemented model expands the AD of nanotoxicity prediction models by including MC-NPs and their cytotoxicity data for 110 cell lines. SDEC FP efficiently represents the molecular structures of MC-NPs. Unlike nano-QSAR models that use computationally expensive QM and MD descriptors, preparing the SDEC FP is inexpensive and easily reproducible based on the composition and size of MC-NPs. This allowed us to easily deploy the model in a web environment. One-hot cell-name encoding was used to represent different cell types in the data set. This use of cell information enabled the integration of the data set and the training of the model with a larger volume of data. By extending the representable chemical space through Figure 3. A parity plot comparing predicted and experimental values for the CatBoost regressor. The model demonstrates excellent predictive performance with R2values of 0.99 for the training data set (blue) and 0.88 for test data set (red). Figure 4. NanoToxRadar web interface and distribution of nanotoxicity predictions across cell lines. (A) The user interface for NP query, include core, shell, doping, and coating compositions with the doping ratio and diameter. (B) A radar plot depicting the pIC50 distribution. ACS Nanoscience Au pubs.acs.org/nanoau Letter https://doi.org/10.1021/acsnanoscienceau.5c00035 ACS Nanosci. Au 2025, 5, 344−352 349
SDEC FP, the model accurately predicted the multitarget cytotoxicity of MC-NPs, successfully generalizing a single model for the task. To make NanoToxRadar easily accessible to the research community, the final model was deployed in a web environment to support safe and sustainable nanomaterial design (https://www.kitox.re.kr/nanotoxradar/). The model is also available on mobile devices. ■ASSOCIATED CONTENT * sı Supporting Information The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acsnanoscienceau.5c00035. Supplementary_information, additional supporting data and preprocessed data set used for model development (PDF) Supplementary_table (XLSX) ■AUTHOR INFORMATION Corresponding Author Hyun Kil Shin −Department of Predictive Model Research, Korea Institute of Toxicology, Daejeon 34114, Republic of Korea; Human and Environmental Toxicology, University of Science and Technology (UST), Daejeon 34113, Republic of Korea; orcid.org/0000-0003-3665-0841; Email: [email protected] Authors Jaehyeon Park −Department of Predictive Model Research, Korea Institute of Toxicology, Daejeon 34114, Republic of Korea; Human and Environmental Toxicology, University of Science and Technology (UST), Daejeon 34113, Republic of Korea; orcid.org/0009-0004-7459-809X Shahzad Rashid −The Institute of Occupational Medicine (IOM), Edinburgh EH14 4AP, U.K. Helena Copsey −The Institute of Occupational Medicine (IOM), Edinburgh EH14 4AP, U.K. Lang Tran −The Institute of Occupational Medicine (IOM), Edinburgh EH14 4AP, U.K. Alex Zabeo −East European Research and Innovation Enterprise (EMERGE), 1303 Sofa, Bulgaria Danail Hristozov −East European Research and Innovation Enterprise (EMERGE), 1303 Sofa, Bulgaria Giorgos P. Gakis −Research Lab of Advanced, Composite, Nano-Materials and Nanotechnology, Materials Science and Engineering Department, School of Chemical Engineering, National Technical University of Athens, Zografos, Athens 15780, Greece; orcid.org/0000-0002-3879-7276 Costas Charitidis −Research Lab of Advanced, Composite, Nano-Materials and Nanotechnology, Materials Science and Engineering Department, School of Chemical Engineering, National Technical University of Athens, Zografos, Athens 15780, Greece; orcid.org/0000-0003-1367-7603 Seokjoo Yoon −Department of Predictive Model Research, Korea Institute of Toxicology, Daejeon 34114, Republic of Korea; Human and Environmental Toxicology, University of Science and Technology (UST), Daejeon 34113, Republic of Korea Complete contact information is available at: https://pubs.acs.org/10.1021/acsnanoscienceau.5c00035 Author Contributions The manuscript was prepared through contributions of all authors. Final version of the manuscript is approved by all authors. Jaehyeon Park: experiments, results analysis and visualization, writing original draft preparation. Shahzad Rashid, Helena Copsey, Lang Tran, Alex Zabeo, Danail Hristozov, Giorgos P. Gakis, and Costas Charitidis: data collection, data quality check, model evaluation, manuscript review, and editing. Seokjoo Yoon and Hyun Kil Shin: methodology, supervision, funding acquisition, and manuscript editing Funding This work was financially supported by the Ministry of Trade, Industry, and Energy (MOTIE) and Korea Institute for Advancement of Technology (KIAT) through the International Cooperative R&D program (P0019147), the Research Program for Agriculture Science and Technology Development (Project No. 00400007) from the National Institute of Agricultural Sciences, Rural Development Administration, as well as Korea Institute of Toxicology (KIT) Research Program (2710008763). This research was conducted in cooperation with SUNSHINE (funded by the European Union’s HORIZON 2020 program under grant agreement no. 952924) and the Explainable AI for Molecules-AiChemist project (funded by the Horizon Europe Marie SkłodowskaCurie Actions Doctoral Network under grant agreement no. 101120466). Notes The authors declare no competing financial interest. ■REFERENCES (1) Haase, A.; Klaessig, F.; Nymark, P.; Paul, K.; Greco, D. EU US Roadmap Nanoinformatics 2030, 2018. DOI: 10.5281/zenodo.1486012 (2) Bessa, M. J.; Brandao, F.; Viana, M.; Gomes, J. F.; Monfort, E.; Cassee, F. R.; Fraga, S.; Teixeira, J. P. Nanoparticle exposure and hazard in the ceramic industry: an overview of potential sources, toxicity and health effects. Environmental Research 2020,184, 109297. (3) Manigrasso, M.; Protano, C.; Astolfi, M. L.; Massimi, L.; Avino, P.; Vitali, M.; Canepari, S. Evidences of copper nanoparticle exposure in indoor environments: Long-term assessment, high-resolution field emission scanning electron microscopy evaluation, in silico respiratory dosimetry study and possible health implications. Sci. Total Environ. 2019,653, 1192−1203. (4) Blasco, J.; Corsi, I.; Matranga, V. Particles in the oceans: Implication for a safe marine environment. Marine Environmental Research 2015,111, 1−4. (5) Armstead, A. L.; Li, B. Nanotoxicity: emerging concerns regarding nanomaterial safety and occupational hard metal (WCCo) nanoparticle exposure. International journal of nanomedicine 2016, 11, 6421−6433. (6) Xuan, L.; Ju, Z.; Skonieczna, M.; Zhou, P.-K.; Huang, R. Nanoparticles-induced potential toxicity on human health: Applications, toxicity mechanisms, and evaluation models. MedComm 2023,4 (4), No. e327. (7) Thu, H. E.; Haider, M.; Khan, S.; Sohail, M.; Hussain, Z. Nanotoxicity induced by nanomaterials: A review of factors affecting nanotoxicity and possible adaptations. OpenNano 2023,14, 100190. (8) Kumar, V.; Sharma, N.; Maitra, S. S. In vitro and in vivo toxicity assessment of nanoparticles. International Nano Letters 2017,7(4), 243−256. (9) Savage, D. T.; Hilt, J. Z.; Dziubla, T. D. In Vitro Methods for Assessing Nanoparticle Toxicity. In Nanotoxicity: Methods and Protocols; Zhang, Q., Ed.; Springer: New York: New York, NY, 2019; pp 1−29. ACS Nanoscience Au pubs.acs.org/nanoau Letter https://doi.org/10.1021/acsnanoscienceau.5c00035 ACS Nanosci. Au 2025, 5, 344−352 350
(10) Puzyn, T.; Rasulev, B.; Gajewicz, A.; Hu, X.; Dasari, T. P.; Michalkova, A.; Hwang, H. M.; Toropov, A.; Leszczynska, D.; Leszczynski, J. Using nano-QSAR to predict the cytotoxicity of metal oxide nanoparticles. Nat. Nanotechnol 2011,6(3), 175−8. (11) Fourches, D.; Pu, D.; Tassa, C.; Weissleder, R.; Shaw, S. Y.; Mumper, R. J.; Tropsha, A. Quantitative Nanostructure−Activity Relationship Modeling. ACS Nano 2010,4(10), 5703−5712. (12) Cao, J.; Pan, Y.; Jiang, Y.; Qi, R.; Yuan, B.; Jia, Z.; Jiang, J.; Wang, Q. Computer-aided nanotoxicology: risk assessment of metal oxide nanoparticles via nano-QSAR. Green Chem. 2020,22 (11), 3512−3521. (13) Kar, S.; Pathakoti, K.; Tchounwou, P. B.; Leszczynska, D.; Leszczynski, J. Evaluating the cytotoxicity of a large pool of metal oxide nanoparticles to Escherichia coli: Mechanistic understanding through In Vitro and In Silico studies. Chemosphere 2021,264 (Pt1), 128428. (14) Zhang, F.; Wang, Z.; Peijnenburg, W. J. G. M.; Vijver, M. G. Review and Prospects on the Ecotoxicity of Mixtures of Nanoparticles and Hybrid Nanomaterials. Environ. Sci. Technol. 2022,56 (22), 15238−15250. (15) Mikolajczyk, A.; Gajewicz, A.; Mulkiewicz, E.; Rasulev, B.; Marchelek, M.; Diak, M.; Hirano, S.; Zaleska-Medynska, A.; Puzyn, T. Nano-QSAR modeling for ecosafe design of heterogeneous TiO2based nano-photocatalysts. Environmental Science: Nano 2018,5(5), 1150−1160. (16) Mikolajczyk, A.; Sizochenko, N.; Mulkiewicz, E.; Malankowska, A.; Rasulev, B.; Puzyn, T. A chemoinformatics approach for the characterization of hybrid nanomaterials: safer and efficient design perspective. Nanoscale 2019,11 (24), 11808−11818. (17) Gakis, G. P.; Aviziotis, I. G.; Charitidis, C. A. A structure− activity approach towards the toxicity assessment of multicomponent metal oxide nanomaterials. Nanoscale 2023,15 (40), 16432−16446. (18) Sifonte, E. P.; Castro-Smirnov, F. A.; Jimenez, A. A. S.; Diez, H. R. G.; Martínez, F. G. Quantum mechanics descriptors in a nanoQSAR model to predict metal oxide nanoparticles toxicity in human keratinous cells. J. Nanopart. Res. 2021,23 (8), 161. (19) Na, M.; Nam, S. H.; Moon, K.; Kim, J. Development of a nanoQSAR model for predicting the toxicity of nano-metal oxide mixtures to Aliivibrio fischeri. Environmental Science: Nano 2023,10 (1), 325− 337. (20) Chew, A. K.; Pedersen, J. A.; Van Lehn, R. C. Predicting the Physicochemical Properties and Biological Activities of MonolayerProtected Gold Nanoparticles Using Simulation-Derived Descriptors. ACS Nano 2022,16 (4), 6282−6292. (21) Shin, H. K.; Kim, K. Y.; Park, J. W.; No, K. T. Use of metal/ metal oxide spherical cluster and hydroxyl metal coordination complex for descriptor calculation in development of nanoparticle cytotoxicity classification model. SAR and QSAR in Environmental Research 2017,28 (11), 875−888. (22) Subramanian, N. A.; Palaniappan, A. NanoTox: Development of a Parsimonious In Silico Model for Toxicity Assessment of MetalOxide Nanoparticles Using Physicochemical Features. ACS Omega 2021,6(17), 11729−11739. (23) Cheng, K.; Pan, Y.; Yuan, B. Cytotoxicity prediction of nano metal oxides on different lung cells via Nano-QSAR. Environ. Pollut. 2024,344, 123405. (24) Kleandrova, V. V.; Luan, F.; González-Díaz, H.; Ruso, J. M.; Speck-Planche, A.; Cordeiro, M. N. D. S. Computational Tool for Risk Assessment of Nanomaterials: Novel QSTR-Perturbation Model for Simultaneous Prediction of Ecotoxicity and Cytotoxicity of Uncoated and Coated Nanoparticles under Multiple Experimental Conditions. Environ. Sci. Technol. 2014,48 (24), 14686−14694. (25) Luan, F.; Kleandrova, V. V.; González-Díaz, H.; Ruso, J. M.; Melo, A.; Speck-Planche, A.; Cordeiro, M. N. D. S. Computer-aided nanotoxicology: assessing cytotoxicity of nanoparticles under diverse experimental conditions by using a novel QSTR-perturbation approach. Nanoscale 2014,6(18), 10623−10630. (26) Gakis, G. P.; Aviziotis, I. G.; Charitidis, C. A. Metal and metal oxide nanoparticle toxicity: moving towards a more holistic structure−activity approach. Environmental Science: Nano 2023,10 (3), 761−780. (27) Halder, A. K.; Delgado, A. H. S.; Cordeiro, M. N. D. S. First multi-target QSAR model for predicting the cytotoxicity of acrylic acid-based dental monomers. Dental Materials 2022,38 (2), 333− 346. (28) Speck-Planche, A. Multicellular Target QSAR Model for Simultaneous Prediction and Design of Anti-Pancreatic Cancer Agents. ACS Omega 2019,4(2), 3122−3132. (29) Choi, J.-S.; Ha, M. K.; Trinh, T. X.; Yoon, T. H.; Byun, H.-G. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci. Rep. 2018,8 (1), 6110. (30) Speck-Planche, A.; Kleandrova, V. V. Multi-Condition QSAR Model for the Virtual Design of Chemicals with Dual Pan-Antiviral and Anti-Cytokine Storm Profiles. ACS Omega 2022,7(36), 32119− 32130. (31) Shin, H. K.; Kim, S.; Yoon, S. Use of size-dependent electron configuration fingerprint to develop general prediction models for nanomaterials. NanoImpact 2021,21, 100298. (32) Dumit, V. I.; Ammar, A.; Bakker, M. I.; Banares, M. A.; Bossa, C.; Costa, A.; Cowie, H.; Drobne, D.; Exner, T. E.; Farcal, L.; Friedrichs, S.; Furxhi, I.; Grafström, R.; Haase, A.; Himly, M.; Jeliazkova, N.; Lynch, I.; Maier, D.; Noorlander, C. W.; Shin, H. K.; Soler-Illia, G. J. A. A.; Suarez-Merino, B.; Willighagen, E.; Nymark, P. From principles to reality. FAIR implementation in the nanosafety community. Nano Today 2023,51, 101923. (33) Oh, E.; Liu, R.; Nel, A.; Gemill, K. B.; Bilal, M.; Cohen, Y.; Medintz, I. L. Meta-analysis of cellular toxicity for cadmiumcontaining quantum dots. Nat. Nanotechnol. 2016,11 (5), 479−486. (34) Rosen, A. S. Periodic Trend Plotter (version 3.1, accessed on 2023-09-15). https://github.com/Andrew-S-Rosen/periodic_trends. (35) Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L., Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 2019,32. (36) Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V. Scikit-learn: Machine learning in Python. J. Machine Learning Research 2011,12, 2825−2830. (37) Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; Association for Computing Machinery: San Francisco, California, USA, 2016; pp 785−794. (38) Prokhorenkova, L.; Gusev, G.; Vorobev, A.; Dorogush, A. V.; Gulin, A. CatBoost: unbiased boosting with categorical features. Advances in neural information processing systems 2018,31. (39) OECD (Q)SAR Assessment Framework: Guidance for the regulatory assessment of (Quantitative) Structure Activity Relationship models and predictions; OECD Publishing: Paris, 2023. (40) Akiba, T.; Sano, S.; Yanase, T.; Ohta, T.; Koyama, M. Optuna: A Next-generation Hyperparameter Optimization Framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &Data Mining; Association for Computing Machinery: Anchorage, AK, USA, 2019; pp 2623−2631. (41) Cherkasov, A.; Muratov, E. N.; Fourches, D.; Varnek, A.; Baskin, I. I.; Cronin, M.; Dearden, J.; Gramatica, P.; Martin, Y. C.; Todeschini, R.; Consonni, V.; Kuz’min, V. E.; Cramer, R.; Benigni, R.; Yang, C.; Rathman, J.; Terfloth, L.; Gasteiger, J.; Richard, A.; Tropsha, A. QSAR Modeling: Where Have You Been? Where Are You Going To? J. Med. Chem. 2014,57 (12), 4977−5010. (42) Irwin, J. J.; Tang, K. G.; Young, J.; Dandarchuluun, C.; Wong, B. R.; Khurelbaatar, M.; Moroz, Y. S.; Mayfield, J.; Sayle, R. A. ZINC20�A Free Ultralarge-Scale Chemical Database for Ligand Discovery. J. Chem. Inf. Model. 2020,60 (12), 6065−6073. (43) Kim, S.; Chen, J.; Cheng, T.; Gindulyte, A.; He, J.; He, S.; Li, Q.; Shoemaker, B. A; Thiessen, P. A; Yu, B.; Zaslavsky, L.; Zhang, J.; ACS Nanoscience Au pubs.acs.org/nanoau Letter https://doi.org/10.1021/acsnanoscienceau.5c00035 ACS Nanosci. Au 2025, 5, 344−352 351
Bolton, E. E PubChem 2025 update. Nucleic Acids Res. 2025,53 (D1), D1516−D1525. ACS Nanoscience Au pubs.acs.org/nanoau Letter https://doi.org/10.1021/acsnanoscienceau.5c00035 ACS Nanosci. Au 2025, 5, 344−352 352