Supplementary Material for "Hyperbolic Nature of Differential Expression Signatures"
Full text
1 Supplementary Material for paper: Hyperbolic Nature of Differential Expression Signatures Domonkos Pog´ any, P´ eter Antal OVERVIEW OF SUPPLEMENTARY MATERIALS This supplementary material provides additional details and extended results supporting our main study. •Section S.I expands on the theoretical connection between differential expression gene (DEG) signatures and hyperbolic geometry, complementing Section IV of the main paper: –S.I.A: Analyzing the scale-free nature of DEG signatures across different datasets and distance metrics. •Section S.II provides additional details for the dimensionality reduction experiment in Section V of the main manuscript: –S.II.A: Detailing the hyperparameter optimization process. –S.II.B: Replicating the results of the comparative analysis using the SigCom LINCS dataset. –S.II.C: Providing supplementary 2D visualizations. –S.II.D: Investigating the connection between the embeddings and the scale-free nature of DEG signatures. S.I. SCALE-FREE NATURE OF DEG SIGNATURES A. Results over Datasets and Distance Metrics This subsection provides additional analyses supporting the findings presented in Section IV of the main paper. To verify the robustness of the observed scale-free properties of differentially expressed gene (DEG) signatures, we repeated the soft thresholding analysis on the L1000FWD landmark dataset [S1] using alternative distance metrics. In addition to cosine similarity, rescaled to [0,1] as Sij = (Sij + 1)/2, we applied Euclidean and Canberra distances, rescaled as Sij = 1−Dij/max(D). Figure S1 illustrates the results, demonstrating that for both metrics, DEG signatures reached an R2above 0.9, reinforcing their scale-free nature. Additionally, to assess the generalizability of our findings across datasets, we conducted the same analysis on the full L1000FWD dataset as well as the SigCom LINCS A375, NPC, NEU, and HELA datasets [S2] (using cosine similarity). As shown in Figure S2, our observation holds for all the datasets. Fig. S1. Analysis of the scale-free nature of L1000FWD landmark DEG signatures across different distance metrics. The R2values of the power-law fit are plotted against different powers of the similarity matrices. Fig. S2. Analysis of the scale-free nature of DEG signatures across different datasets. The R2values of the power-law fit are plotted against different powers of the similarity matrices. D. Pog´ any and P. Antal are with the Department of Artificial Intelligence and Systems Engineering, Budapest University of Technology and Economics, Budapest, Hungary. E-mail: {pogany, antal}@mit.bme.hu.
2 S.II. DIMENSIONALITY REDUCTION ON DEG SIGNATURES A. Hyperparameters This subsection provides additional details for the hyperparameter optimization presented in Section V of the main paper. To ensure the robustness of our evaluations, we conducted a comprehensive hyperparameter search across multiple settings. To assess local structure preservation, we performed K-Nearest Neighbor (KNN) classification on several tasks, using 5-fold repeated cross-validation, reporting the mean and standard deviation of the F1 score for each metric. A similar repeated procedure was applied to global structure evaluations using Spearman correlation (SC) and random triplet accuracy (RTA), where we sampled 5,000 pairs and triplets per evaluation. In addition, each experiment using manifold learning methods (Uniform Manifold Approximation and Projection (UMAP) [S3]; its hyperbolic version, HUMAP; and the Poincar´ e maps (PMAP) [S4]) was repeated five times to account for their non-deterministic nature. We evaluated different distance metrics for DEG signatures, testing Euclidean, cosine, and Canberra distances, both with and without prior standardization. Figures S3 to S5 present the KNN classification results on DEG signatures across the L1000FWD landmark, L1000FWD full, and SigCom LINCS A375 datasets. Standardization did not improve performance for any metric, and cosine consistently outperformed Euclidean with the same execution time. Canberra distance, as previously reported [S5], achieved slightly better KNN classification results, particularly for cell type and time point prediction. However, it did not improve MOA classification or later global evaluation metrics. Regarding the computational cost of the KNN classification, even in the relatively small L1000FWD landmark dataset, it had 10 times longer execution time than the cosine and Euclidean distances. Moreover, unlike the other metrics, its time depends on the number of features, not just the samples. As a consequence, it took more than 12 hours to calculate the Canberra distance matrix on the larger datasets. Considering the only marginal improvement and its vast computational cost, we excluded it from further analysis. Figure S6 compares manifold learning methods using Euclidean and cosine metrics as inputs, showing that cosine consistently outperformed Euclidean across all methods and hyperparameter settings. Based on these findings, we used cosine similarity without prior standardization throughout our study for KNN classification on DEG signatures and as the input metric for dimensionality reduction. This choice aligns with standard practices in analyzing DEG signatures derived using the Characteristic Direction (CD) method [S6]. We also examined the hyperparameters of the KNN classifier, namely the number of neighbors and whether to use distanceweighted predictions. Our experiments with k=5, 10, and 50 neighbors showed no significant performance differences, leading us to select k= 5 with weighted distances for consistency. Figures S3 to S5 present the KNN classification results across multiple datasets, further supporting this choice. We tested various combinations to investigate potential interactions between the KNN classification neighbor number and the number of neighbors used for constructing KNN graphs at the initial step of manifold learning methods. As shown in Figure S7, we found no connection between the two parameters. While the performance of manifold learning methods is sensitive to the KNN graph neighbor number, the KNN classification neighbor number had little to no impact. We then optimized the hyperparameters of the manifold learning methods. The results for local and global structure preservation across multiple datasets and latent dimensions are presented in Figures S8 to S10 for UMAP, Figures S11 to S13 for HUMAP, and Figures S14 to S17 for PMAP. Additionally, we report execution times, but here, they refer to the dimensionality reduction step alone, excluding KNN classification and global evaluation time. We focused on optimizing the number of neighbors in the KNN graph for all methods, the minimum distance for UMAP and HUMAP, and the learning rate for PMAP. Other hyperparameters were also tested but showed minimal impact, so we retained their default values. For instance, sigma and gamma parameters for PAMP were kept at their default values of 1.0 and 2.0, respectively, as shown in Figure S15. For UMAP and HUMAP, we tested minimum distance values of 0.001, 0.01, 0.05, 0.1, and 0.5. We observed minimal variation in local and global metrics and no substantial effect on execution time. To balance local and global structure preservation, we selected a minimum distance of 0.1. For PMAP, we explored learning rates between 0.01 and 1.0. Learning rates below 0.1 resulted in suboptimal performance, while those above 0.1 showed no significant improvements. However, since higher learning rates led to faster convergence, we selected 1.0 for efficiency. Across all methods, the number of KNN graph neighbors was the most influential hyperparameter, determining the trade-off between local and global structure preservation. To capture both extremes, we used two configurations: k= 5 for enhanced local structure and k= 250 for improved global structure. The only exception was UMAP and HUMAP on the SigCom LINCS dataset, where k= 50 (the default setting for UMAP visualization in this dataset [S2]) yielded the best overall performance. Regarding execution time, most hyperparameters had little impact. Since all methods operate on the distance matrix of DEG signatures, they are independent of the number of features (genes) in the signatures. Still, they are affected by the number of samples (signatures). UMAP and HUMAP showed minor sensitivity to embedding dimensions, whereas PMAP required much more time in higher dimensions. As we mentioned before, there was a strong negative correlation between learning rate and the execution time of PMAP. The number of KNN graph neighbors also strongly affected the execution time for both methods. Interestingly, while UMAP and HUMAP ran faster with fewer KNN graph neighbors, the opposite is true for PMAP, which was significantly slower when using a small number of neighbors. We discovered experimentally that PMAP converges faster with large neighbor counts, reaching its early stopping criterion sooner.
3 Fig. S3. Hyperparameter search on the L1000FWD dataset for the landmark DEG signatures, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks and the execution times. Fig. S4. Hyperparameter search on the L1000FWD dataset for the full DEG signatures, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks. Fig. S5. Hyperparameter search on the SigCom LINCS A375 cell line for the DEG signatures, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks.
4 Fig. S6. KNN classification results on the L1000FWD landmark dataset across various signature distance metrics, dimensionality reduction methods, and hyperparameters. The figure shows the mean and standard deviations of F1 scores for different classification tasks. Fig. S7. KNN classification results on the L1000FWD landmark dataset across dimensionality reduction methods, dimensionality reduction graph neighbor counts (neighbors), and KNN classification neighbor counts (KNN neighbors). The figure shows the mean and standard deviations of F1 scores for different classification tasks.
5 Fig. S8. Hyperparameter search on the L1000FWD dataset for the UMAP method in 2D, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks, the global structure preservation metrics, and the execution times. Fig. S9. Hyperparameter search on the L1000FWD dataset for the UMAP method in 64D, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks, the global structure preservation metrics, and the execution times.
6 Fig. S10. Hyperparameter search on the SigCom LINCS A375 cell line for the UMAP method in 2D, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks and the global structure preservation metrics. Fig. S11. Hyperparameter search on the L1000FWD dataset for the HUMAP method in 2D, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks, the global structure preservation metrics, and the execution times.
7 Fig. S12. Hyperparameter search on the L1000FWD dataset for the HUMAP method in 64D, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks, the global structure preservation metrics, and the execution times. Fig. S13. Hyperparameter search on the SigCom LINCS A375 cell line for the HUMAP method in 2D, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks and the global structure preservation metrics.
8 Fig. S14. Hyperparameter search on the L1000FWD dataset for the PMAP method in 2D, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks, the global structure preservation metrics, and the execution times. Fig. S15. Hyperparameter search on the L1000FWD dataset for the PMAP method in 2D, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks and the global structure preservation metrics
9 Fig. S16. Hyperparameter search on the L1000FWD dataset for the PMAP method in 64D, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks, the global structure preservation metrics, and the execution times. Fig. S17. Hyperparameter search on the SigCom LINCS A375 cell line for the PMAP method in 2D, illustrating mean and standard deviations of the F1 scores for various KNN classification tasks and the global structure preservation metrics.
16 Fig. S27. Two-dimensional signature visualization on the SigCom LINCS HELA cell line, comparing different dimensionality reduction methods and neighbor counts. Embeddings are colored according to the most frequent mechanisms of action. Fig. S28. Two-dimensional signature visualization on the SigCom LINCS HELA cell line, comparing different dimensionality reduction methods and neighbor counts. Embeddings are colored according to the most frequent perturbation times.
17 Fig. S29. Two-dimensional signature visualization on the SigCom LINCS NPC cell line, comparing different dimensionality reduction methods and neighbor counts. Embeddings are colored according to the most frequent mechanisms of action. Fig. S30. Two-dimensional signature visualization on the SigCom LINCS NPC cell line, comparing different dimensionality reduction methods and neighbor counts. Embeddings are colored according to the most frequent perturbation times.
18 Fig. S31. Two-dimensional signature visualization on the SigCom LINCS NEU cell line, comparing different dimensionality reduction methods and neighbor counts. Embeddings are colored according to the most frequent mechanisms of action. Fig. S32. Two-dimensional signature visualization on the SigCom LINCS NEU cell line, comparing different dimensionality reduction methods and neighbor counts. Embeddings are colored according to the most frequent perturbation times.
19 D. Connection with the Scale-Free Nature In the following, we provide details of the connection between the scale-free nature of DEG signatures and the obtained embeddings. Specifically, we examined how well these embeddings capture the degree distribution of the DEG signatures. As outlined in Section IV of the main paper, we defined node degrees in two ways: (i) a discrete degree derived from a hard threshold of 0.65 applied to the cosine similarity matrix Sij, and (ii) a continuous weighted degree, computed as the sum of the weighted adjacency matrix Aij =Sβ ij, with β= 7 chosen to maintain a scale-free structure. We first analyzed a subset of the L1000FWD dataset containing only 2,000 signatures, the same subset used in Figure 1 of the main paper. Figure S33 visualizes the network obtained using the 0.65 threshold, where node size is proportional to discrete degree and color reflects weighted degree. To quantitatively assess the relationship between degree and embeddings, we plotted weighted degree against embedding norms and computed signed R2values for a linear fit. Results for various dimensionality reduction methods, KNN neighbor numbers, and embedding dimensionality are detailed in Figures figs. S34 to S37. We then extended this analysis to the entire dataset with 16,848 signatures. Figure S38 presents a similar visualization as before. To get consistent results, we translated the signature with the highest weighted degree to the origin, reflecting our expectation that highly connected nodes should be in the center of the embedding space. For HUMAP and PMAP, this was performed using Poincar´ e translations [S4]. Figure S39 illustrates a two-dimensional example of the Poincar´ e translation. The results of the analysis are shown in Figures figs. S40 to S44. Notably, after translation, embedding norms were equal to distances from the highest-degree node. On the small subset, we observed a strong negative correlation between PMAP embedding norm and degree with |R2|>0.8. However, this effect weakened in lower dimensions and small KNN neighbor numbers. In contrast, none of the embeddings showed a strong alignment with degree distribution on the full dataset. Methods with strong global structure preservation, such as PCA and PMAP with high KNN neighbor numbers, captured degree information slightly better than other approaches but only within the range of 0.3<|R2|<0.4. This may be because all tested methods rely on KNN graph construction, which differs topologically from the scale-free network obtained via hard or soft thresholding. These findings suggest that future research could explore manifold learning techniques specifically designed to embed the thresholded graphs into hyperbolic space. Fig. S33. Two-dimensional network visualization on the L1000FWD subset with the first 2,000 signatures. The network was constructed using a similarity threshold of 6.5. Node coordinates were obtained from different dimensionality reduction methods and neighbor counts. Node sizes are proportional to their degree, and colors represent their weighted degree.
20 Fig. S34. Relationship between weighted degree and PCA embedding norm on the L1000FWD subset with the first 2,000 signatures. Embedding norms are plotted against weighted degrees and colored by degree. R-squared values of a linear fit between weighted degrees and norms are assessed across different dimensions. Fig. S35. Relationship between weighted degree and UMAP embedding norm on the L1000FWD subset with the first 2,000 signatures. Embedding norms are plotted against weighted degrees and colored by degree. R-squared values of a linear fit between weighted degrees and norms are assessed across different dimensions and neighbor counts.
21 Fig. S36. Relationship between weighted degree and HUMAP embedding norm on the L1000FWD subset with the first 2,000 signatures. Embedding norms are plotted against weighted degrees and colored by degree. R-squared values of a linear fit between weighted degrees and norms are assessed across different dimensions and neighbor counts.
22 Fig. S37. Relationship between weighted degree and PMAP embedding norm on the L1000FWD subset with the first 2,000 signatures. Embedding norms are plotted against weighted degrees and colored by degree. R-squared values of a linear fit between weighted degrees and norms are assessed across different dimensions and neighbor counts. Fig. S38. Two-dimensional network visualization on the L1000FWD dataset. The network was constructed using a similarity threshold of 6.5. Node coordinates were obtained from different dimensionality reduction methods and neighbor counts. Node sizes are proportional to their degree, and colors represent their weighted degree.
23 Fig. S39. Hyperbolic translation operation in the 2D PMAP space. The same L1000FWD network is visualized, but to the right, the node with the highest weighted degree is shifted to the center. Fig. S40. Relationship between weighted degree and adjusted 2D FWD embedding norm on the L1000FWD dataset. Adjusted norms are plotted against weighted degrees and colored by degree. Fig. S41. Relationship between weighted degree and adjusted PCA embedding norm on the L1000FWD dataset. Adjusted norms are plotted against weighted degrees and colored by degree. R-squared values of a linear fit between weighted degrees and adjusted norms are assessed across different dimensions.
24 Fig. S42. Relationship between weighted degree and adjusted UMAP embedding norm on the L1000FWD dataset. Adjusted norms are plotted against weighted degrees and colored by degree. R-squared values of a linear fit between weighted degrees and adjusted norms are assessed across different dimensions and neighbor counts. Fig. S43. Relationship between weighted degree and adjusted HUMAP embedding norm on the L1000FWD dataset. Adjusted norms are plotted against weighted degrees and colored by degree. R-squared values of a linear fit between weighted degrees and adjusted norms are assessed across different dimensions and neighbor counts.
25 Fig. S44. Relationship between weighted degree and adjusted PMAP embedding norm on the L1000FWD dataset. Adjusted norms are plotted against weighted degrees and colored by degree. R-squared values of a linear fit between weighted degrees and adjusted norms are assessed across different dimensions and neighbor counts.