scieee AI-readable full text Open interactive document viewer

MicroSplit: Semantic Unmixing of Fluorescent Microscopy Data

Jug, Florian

Full text

MicroSplit: Semantic Unmixing of Fluorescent Microscopy Data Ashesh Ashesh1, Federico Carrara1, Igor Zubarev1, Vera Galinova1, Melisande Croft1, Melissa Pezzotti2, Daozheng Gong3, Francesca Casagrande1, Elisa Colombo1, Stefania Giussani1, Elena Restelli1, Eugenia Cammarota1, Juan Manuel Battagliotti1, Nikolai Klena1, Moises Di Sante2, Raghabendra Adhikari4, Daniel Feliciano4, Gaia Pigino1, Elena Taverna1, Oliver Harschnitz1, Nicola Maghelli1, Norbert Scherer3, Damian Edward Dalle Nogare1, Joran Deschamps1, Francesco Pasqualini2, Florian Jug1* 1Fondazione Human Technopole, V.le Rita Levi-Montalcini, 20157, Milan, Italy. 2University of Pavia, Corso Strada Nuova, 65, 27100, Pavia, Italy. 3University of Chicago, 5801 S Ellis Ave, 60637, Chicago, USA. 4HHMI/Janelia Research Campus, 19700 Helix Drive, Ashburn, 20147, VA, USA. Abstract Fluorescence microscopy, a key driver for progress in the life sciences, faces limitations due to the microscope’s optics, fluorophore chemistry, and photon exposure limits, necessitating trade-offs in imaging speed, resolution, and depth. Here, we introduce MicroSplit, a computational multiplexing technique based on deep learning that allows multiple cellular structures to be imaged in a single fluorescent channel and then unmix them by computational means, allowing faster imaging and reduced photon exposure. We show that MicroSplit efficiently separates up to four superimposed noisy structures into distinct denoised fluorescent image channels. Furthermore, using Variational Splitting Encoder-Decoder (VSE) networks, our approach can sample diverse predictions from a trained posterior of solutions. The diversity of these samples scales with the uncertainty in a given input, allowing us to estimate the true prediction errors by computing the variability between posterior samples. We demonstrate the robustness of MicroSplit networks, which are trained for each splitting task at hand, across various datasets and noise levels and show its utility to image more, to image faster, and to improve downstream analysis. We provide MicroSplit along with all associated training and evaluation datasets as open resources, enabling life scientists to immediately benefit from the potential of computational multiplexing and thus help accelerate the rate of their scientific discovery process. Keywords: deep learning, fluorescence microscopy, computational multiplexing, semantic channel unmixing, variational inference 1 Introduction Fluorescence microscopy is an essential tool for exploring structures and dynamics within cells, tissues, and organisms. Multiplexed acquisition techniques involving multiple fluorophores, each tagged to a different cellular structure, are used to acquire multichannel image data. However, overlap in fluorophore excitation spectra imposes practical limits on the number of fluorophores that can be used in a biological sample (see Figure 1a) without problems such as cross-talk or bleed-through [1]. Even if a suitable set of fluorescent labels have been chosen, the sequential nature of multiplexed acquisitions requires multiple exposures of the sample, which reduces the photon budget available for other purposes [2,3] and also limits the maximum sampling frequency in live-cell imaging scenarios. A limited photon budget can be addressed by reducing the photon exposure per acquisition, which in turn will lead to more noisy images that will be harder to analyze [3–5]. In this work, we describe a computational multiplexing approach, MicroSplit, which addresses the limitations outlined above. We propose labeling and imaging multiple biological structures in a single fluorescent 1 Fig. 1 Semantic unmixing of fluorescence microscopy data with MicroSplit. (a) Multiplexed fluorescent imaging is conducted by labeling cellular components with different fluorescent markers and imaging one after another into separate image channels. (b) We propose a method that allows multiple cellular structures to be imaged simultaneously in a single image channel. The superimposed structures are then split into separate channels using an adequately trained neural network. (c) Imaging multiple structures in one go saves precious photon budget. It is well known that the total photon budget limits the capabilities of light microscopes (left). Using content-aware denoising methods, e.g. [3,4], can help repurpose some photon budget by acquiring more noisy raw images (middle). Imaging multiple cellular components in a single fluorescent channel saves additional photon budget that can then, for example, be invested into imaging additional structures, image more gentle, or image faster (right). (d-f) Exemplary qualitative results of two-channel splitting (d), three channel splitting (e), and fourchannel splitting (f). Each panel shows the image channel containing multiple structures to be split (input), insets of the noisy target channels (Ci, i ∈ {1,2,3,4}) as they are used during training (target), and insets of the split channels as predictions by MicroSplit (prediction). Note that the predictions are noticeable less noisy than the training data itself, which is a key feature of MicroSplit’s unsupervised denoising component (see also the main text and Figure 2). channel and then employing our proposed method to split superimposed structures from within this multistructure channel computationally into separate unmixed image channels (see Figure 1b). We recently presented the underlying methodological components that can enable such computational multiplexing in theory [6,7] and have now created a practical method that (i) combines the benefits of both approaches into a single framework, (ii) enables the processing of volumetric image data by creating a highly optimized network architecture for it (see Figure 1and Table 1), and (iii) we propose a procedure to assess the calibration of a trained network and the estimation of prediction errors (see Figure 2and Section 2.3). MicroSplit is designed for easy use by microscopists and life scientists, even those who are not machine learning experts. 2 Using MicroSplit allows users to (i) simultaneously image more structures than previously possible by combining up to four structures in a single channel, or to (ii) image the same number of structures but in fewer channels and, hence, at reduced photon exposure (see Figure 1c). The saved photon budget can then be allocated towards other objectives such as faster temporal sampling, better signal-to-noise ratio in raw acquisitions, e.g. to image more gentle and avoid phototoxicity, or any combination of these possibilities. Note that each semantic unmixing task requires training a dedicated MicroSplit model. While training a universal foundation model is conceptually possible, we deliberately employ narrow, task-specific models to avoid out-of-distribution issues and to ensure robust performance on the data at hand. Our experiments show that MicroSplit is capable of separating two, three, or even four jointly imaged structures into denoised unmixed channel predictions, even if the training data are itself noisy (see Figure 1df). A MicroSplit model learns to unmix superimposed structures by learning from examples (i.e. in a supervised way). At the same time, it co-learns to denoise the data without supervision (see Figure 2a). This means that it is sufficient to use a body of noisy training data and MicroSplit will still learn to predict denoised images for each unmixed structure (images labeled “Target”1in Figures 1and 2show such noisy training data). These properties make our method straightforward to apply in practical applications, as we show in detail in Section 2. An additional and distinctive feature of the Variational Splitting Encoder-Decoder (VSE) networks we are using is their ability to generate multiple plausible solutions for a given input image [7]. When given inputs are unambiguous and thus can be split with high certainty, the generated solutions will closely resemble each other. In contrast, as the uncertainty in inputs increases, the diversity among the sampled solutions will reflect these inherent ambiguities by becoming more different from one another. We show that the networks we trained are calibrated, allowing us to estimate the otherwise hard to assess true error of predictions, even in the absence of ground truth images that are notoriously hard or even impossible to come by. Technically, this is enabled by evaluating the easy-to-compute inter-sample variability (see Section 2.3). This feature is crucial because it enables users to identify the parts of their data where MicroSplit predictions may be unreliable due to ambiguities in the fed inputs. Such regions can either be disregarded or subjected to expert review, effectively addressing a persistent challenge in AI-driven bioimage analysis–the difficulty in evaluating the accuracy of predictions [1]. We show the performance of MicroSplit on 24 2D and 5 3D semantic unmixing tasks, showing that we achieve accurate results in a wide range of noise levels and superimposed structures, consistently leading to high-quality denoised single-structure predictions. In addition to the above-mentioned use cases, we propose a way to use MicroSplit to remove unwanted structured image artifacts by unmixing them from the structures in which we are truly interested. We showcase this capability by removing spurious fluorescent puncta from a fluorescent microscopy dataset (see Figure 4and Section 2.4.2). 2 Results As introduced above, MicroSplit enables microscopists to image multiple structures that would typically be imaged in separate fluorescence channels and instead: (i) capture these structures in one channel, and then (ii) unmix the superimposed structures contained in that channel by computational means. The training procedure we propose requires supervision signals (target images) for each structure to be unmixed (see Figure 1b). In the following section, we outline three practical ways to obtain suitable training data and will use them in various experimental settings to demonstrate their effectiveness. 2.1 Training Modes and Required Training Data Traditionally, data are acquired using conventional multiplexing (see Figure 1a). Hence, we propose Training Mode I, where previously recorded multiplexed image channels are used as targets for supervised training to achieve semantic unmixing of cellular structures (see Figure 2a). We generate mixed input images by pixel-wise summation of multiplexed image channels. These input images closely resemble what can later be acquired in a single acquisition (see Figure 1b). Although this approach yields high-quality training data, some differences in noise properties and intensity variations between structures may arise, as discussed in Section C.2. Experiments using Training Mode I show consistently high semantic unmixing performance, as seen in Figures 1and 2, Table 1, and in Sections B.2 and C.2. 1In many supervised training settings, the supervision data is referred to as “ground truth”. In this manuscript we instead refer to the supervision signals as “Target” since MicroSplit can be trained exclusively on noisy data that is therefore not ground truth. 3 (a) Two-channel unmixing tasks Task Dataset Task Details 2D/3D PSNR MicroMS-SSIM C1 C2 C1 C2 I HT-H24 TM-I 3D 38.8 33.8 0.970 0.951 II HT-P23A TM-I 3D 25.1 31.4 0.747 0.932 III HT-P23B TM-I 3D 26.4 21.6 0.847 0.599 IV Pavia-P24 TM-III, high, 50:50 2D 24.3 29.9 0.682 0.839 V Pavia-P24 TM-III, high, 66:33 2D 28.2 25.6 0.780 0.696 VI Pavia-P24 TM-III, high, 84:16 2D 25.2 24.1 0.722 0.623 VII Pavia-P24 TM-III, mid, 50:50 2D 23.1 24.3 0.673 0.755 VIII Pavia-P24 TM-III, mid, 66:33 2D 24.0 22.3 0.729 0.678 IX Pavia-P24 TM-III, mid, 84:16 2D 24.3 22.4 0.735 0.659 X Pavia-P24 TM-III, low, 50:50 2D 21.9 23.0 0.612 0.684 XI Pavia-P24 TM-III, low, 66:33 2D 22.9 21.7 0.593 0.564 XII Pavia-P24 TM-III, low, 84:16 2D 23.3 22.8 0.682 0.663 XIII HT-T24 TM-III 2D 40.3 32.8 0.978 0.951 XIV HT-LIF24 TM-III 2D 32.0 32.9 0.965 0.960 XV Chicago-Sch23 TM-I, C0 vs C1 2D 41.3 42.9 0.984 0.993 XVI Chicago-Sch23 TM-I, C0 vs C2 2D 38.4 41.1 0.973 0.987 XVII Chicago-Sch23 TM-I, C0 vs C3 2D 58.2 41.7 0.998 0.993 XVIII Chicago-Sch23 TM-I, C1 vs C2 2D 43.7 44.9 0.995 0.995 XIX Chicago-Sch23 TM-I, C1 vs C3 2D 65.4 44.3 1.000 0.996 XX Chicago-Sch23 TM-I, C2 vs C3 2D 66.6 43.6 1.000 0.996 (b) Three-channel unmixing tasks Task Dataset Task Details 2D/3D PSNR MicroMS-SSIM C1 C2 C3 C1 C2 C3 XXI CBG-Z18 TM-I 3D 28.1 28.9 29.1 0.920 0.929 0.913 XXII CBG-N18 TM-I 3D 38.4 42.4 35.0 0.977 0.981 0.974 XXIII HHMI-D258bit TM-I 3D 22.5 31.3 24.3 0.840 0.768 0.793 XXIV HT-LIF24 TM-III, 2ms 2D 31.0 32.2 36.3 0.940 0.973 0.944 XXV HT-LIF24 TM-III, 3ms 2D 30.8 32.2 36.1 0.940 0.973 0.948 XXVI HT-LIF24 TM-III, 5ms 2D 32.9 34.3 37.6 0.960 0.983 0.963 XXVII HT-LIF24 TM-III, 20ms 2D 37.0 39.7 41.4 0.984 0.994 0.989 XXVIII HT-LIF24 TM-III, .5s 2D 39.5 41.6 43.1 0.991 0.995 0.995 (c) Four-channel unmixing tasks Task Dataset 2D/3D PSNR MicroMS-SSIM C1 C2 C3 C4 C1 C2 C3 C4 XXIX HT-LIF24 2D 32.1 32.5 34.2 37.6 0.964 0.959 0.983 0.957 XXX Chicago-Sch23 2D 35.6 39.4 38.8 36.2 0.956 0.986 0.980 0.942 Table 1 Quantitative performance on two-, three-, and four-channel splitting tasks. Throughout this work, we refer to specific splitting tasks by their Task-Id (first column). Some of the microscopy datasets we use (column two), give rise to multiple splitting tasks, e.g. by selecting a subset of the fluorescent channels or using different noise levels or channel mixing weights (see Section 4.1). Such Task Details are in abbreviated form given in the third column along with the training mode used for the task. For tasks with Task-Id XXIX and XXX, the training modes used were TM-IIIand TM-I, respectively. For volumetric tasks (labeled 3D in column 2D/3D), we fed a 3D image stack to MicroSplit, which in turn also predicts 3D outputs (posterior samples). We evaluate all the tasks on held-out test sets and report MicroMS-SSIM [8] and CARE-PSNR (PSNR) metrics for each unmixed channel prediction. In Table G.3 we list all standard errors for the results in the above table. Note that Task XXIII is the most difficult semantic unmixing task of visually different cellular structures we encountered so far, particularly w.r.t. the quality of predictions for the third channel. In Section 2.5 and Supplementary Section B.1.1, we elaborate on how low SNR and several other properties of the raw data are contributing to the simplicity or difficulty of semantic unmixing, and how potentially occurring problems can be mitigated. In cases where multiplexed data for Training Mode I are not available, we propose Training Mode II . Here, we assume that images of each single structure exist, but we no longer assume that all structures have been imaged in each sample. As before, here we also generate superimposed inputs by summing images showing different fluorescent structures. However, unlike before, summed-up input images no longer originate from the same sample, and any spatial correlations that might exist between these cellular structures are lost. Since a network trained with Training Mode II cannot leverage these correlations, we reasoned that Training Mode I should be at least as performant as Training Mode II . In Section C.2 we test this hypothesis and show that Training Mode II indeed comes with a slight performance drop in cases where the channels to be unmixed offer actionable spatial correlations that MicroSplit can leverage. 4 A variation of this approach, Training Mode II-b, was used to obtain the results shown in Section 2.4.2. Instead of summing uncorrelated sets of images, we used image data of superimposed structures. If those structures mix in some areas, but are visible in isolation in others, we cropped regions of interest (ROIs) showing individual structures and summed them randomly into superimposed training inputs. This training mode is best understood in the context of the results we present in Section 2.4.2. Finally, in Training Mode III , we do not create input images by summing images of individual structures but rather acquire them also at the microscope. In this mode, as in Training Mode I , all structures of interest must be individually labeled to allow multiplexed imaging of the required training targets. Additionally, we acquire an additional image channel by exciting all used labels at once and collecting the entirety of the emitted light, hence, directly imaging also the superimposed input directly at the microscope. Naturally, Training Mode III then uses this channel instead of the summation of the target channels as input to MicroSplit. The advantage of this is that the input image is also subject to realistic image noise and that the relative intensity of the different structures is realistic as well (see results for tasks with IDs IV-XIV and XXIII-XXIX). These training modes provide flexibility in the preparation of training data for MicroSplit, ensuring robust performance under a variety of experimental conditions and noise levels. We provide an overview of training modes in Supplementary Table G.3. 2.2 MicroSplit Yields High-Quality Unmixed Structures To explore the performance of MicroSplit in various biological samples, imaging modalities, and training modes (see Section 2.1), we collected a total of 10 data sets and defined a total of 30 + 6 semantic unmixing tasks, as shown in Table 1and Table C.3. A brief overview of the datasets can be found in the Online Methods (Section 4) and more details are given in Supplementary Section F. In Figure 1d-f we show qualitative results for three of these tasks. A quantitative assessment to the available ground truth, called training targets throughout this manuscript, can be found in Table 1. In the same table, we list all tasks and the achieved quality of unmixed channels in terms of CARE-PSNR (PSNR) [1] and MicroMS-SSIM [8], a variation on SSIM optimized for quantitative evaluation of microscopy image data. Throughout all tasks, we observed average PSNR and MicroMS-SSIM values of 32.53 and 0.886, respectively, which for more common tasks, such as image denoising, would typically be considered high enough to warrant downstream processing and analysis. The lowest score of all semantic unmixing experiments still shows PSNR/MicroMS-SSIM values of 21.6/0.564 (Task III-channel 2 and Task VI-channel 2, respectively), which, depending on the desired analysis to be performed, is arguably still reliable enough for downstream analysis. However, neither PSNR nor MicroMS-SSIM are sufficient to know how trustworthy the predictions of MicroSplit are, since these metrics can only be calculated when ground truth targets are available. For this reason, MicroSplit employs a variational training paradigm of its Splitting Encoder-Decoder Neural Network, as shown in Figure 2. The network architecture we use is similar to a hierarchical variational autoencoder (HVAE) [10], sometimes also referred to as a Ladder-VAE [11]. The most prominent difference in our setup is that MicroSplit is not an autoencoder, since the predictions are not meant to be the same as the given input. Details about the precise network architectures and the training procedure of MicroSplit can be found in Section 4.3, as well as in Section A.1 and in [6,7]. In the next section, we show how the variational nature of MicroSplit can be exploited to estimate prediction errors caused by ambiguous inputs being fed. 2.3 Error Estimation, Data Uncertainty, and Calibration MicroSplit exploits the variational nature of its underlying architecture to enable uncertainty quantification. Since variational networks are not simple point-predictors, but instead are capable of learning an approximate posterior of possible solutions [1], MicroSplit can efficiently sample such solutions. As for VAEs [12], more likely solutions will be sampled more frequently, suggesting that the analysis of the intersample variability can be a good surrogate for the uncertainty in the input data and, therefore, also for the trustworthiness of semantic unmixing results. We tested this hypothesis and show our findings in Figure 2b,c. More specifically, we show an input patch and the lateral context that was fed to MicroSplit, two posterior samples, the difference between them as a heatmap, and the (approximate) minimum mean squared error (MMSE) prediction, obtained 5 Fig. 2 Network Architecture, Posterior Sampling, and Calibration. (a) MicroSplit jointly learns to denoise input images (unsupervised, blue shaded area) and split superimposed structures (supervised, red shaded area). For this, we use a hierarchical network architecture, variational training (yellow shaded area), and use lateral context (LC, [6]) for better performance. Due to unsupervised denoising, using an adequate noise model [7,9], the supervised target images can be subject to noise, and the final prediction will still be noise-free. (b) Leftmost column shows an input patch containing three labeled structures in a single channel at native resolution (bottom) and two LC input patches that add additional spatial context surrounding the primary input patch (towards the top). Since the input is noisy, not all structural details in the data are visible. To account for this noise-induced data uncertainty, the variational architecture we use is capable of sampling plausible “interpretations” from a learned posterior of possible solutions. For each of the three superimposed image structures (C1-C3) we show two such posterior samples in columns two and three, and their difference as a heat map in column four. The fifth column shows an approximation of the posterior minimum mean-squared-error (MMSE), computed by pixel-wise averaging 50 posterior samples. The last three columns show the noisy target images (as used during training), an estimate of the pixelwise root mean square error (RMSE) and the calibration plots for the three channels, respectively. The calibration plots show that the uncertainty estimated solely from predictions (RMV) correlate with the true prediction error (RMSE), which means that this easy-to-calculate quantity (RMV) can be used as a surrogate for an otherwise difficult-to-assess prediction error of MicroSplit. (c) Ten posterior samples for C3, the channel showing the highest estimated RMSE values in (b). Note that puncta at low estimated RMSE value locations (cyan arrows) remain relatively unchanged and closely resemble the structure present in the noisy target at that location, while puncta at high estimated RMSE locations (red arrows) show significant variations between samples in (c). by pixelwise averaging of 50 posterior samples. We also show the target patches used to train the shown 3-channel splitting task (Task XXIII). Note also here how much lower the signal-to-noise ratio is in the training data than in the predictions of a trained MicroSplit network. Finally, we show the estimated pixelwise root mean squared error (RMSE), computed using the 50 posterior samples that we also used to generate the previously shown MMSE solution. The overall idea might be best conveyed by looking at a concrete example. Cyan and red arrows in Figure 2b,c point at image locations that might show puncta-like structures. Note that in the 10 posterior samples shown in Figure 2c, the locations pointed at with cyan arrows remain relatively unchanged and the predictions match the structure present in the target channel. However, those locations pointed at with red arrows change their appearance considerably from one posterior sample to the next. This is, in fact, also reflected in the estimated RMSE patches shown in Figure 2b, where higher estimated RMSE values appear precisely at locations where the posterior samples show elevated variability and is hence indicating that 6 Fig. 3 Segmentations on Unmixed Channels are in line with Inter-Observer Variability. We define three segmentation tasks, each on one unmixed channel predicted by MicroSplit, and compare the segmentations created by three Bioimage Analysts with each other. (a) For the three tasks at hand, we show overview images (left) and insets (right) of three two-channel splitting tasks, with the superimposed raw input image in the top row, the predicted channel to be segmented in the middle row, and a single-channel control acquisitions via regular multiplexing (target, see Figure 1a) in the bottom row. (b, c) For Task 1 and Task 2, we show segmentation results obtained by three analysts (A1-A3) obtained on the MicroSplit predictions (top row), the single-channel control acquisitions (target, middle row). The rightmost column shows overlays of the results obtained by all three analysts, with consistently segmented pixels being shown in white. The bottom row first shows the overlay of the segmentation results of a single analyst on predicted (top) vs. single channel control inputs (middle), and in the last column a box plot of all pairwise DICE scores between predicted vs. target channel inputs of the entire test set (5 images of size 1600 ×1600 for Task 1 and size 4096 ×4096 for Task 2). Note that the intra-observer variability between predicted and target segmentations (first three box-plots) are in about the same range as the inter-observer variability shown in the rightmost box-plot. 7 the input data is ambiguous in these regions. However, whether variability in posterior samples actually manifests itself in locations where predictions of MicroSplit are prone to error and, conversely, also only in these places, is not obvious. Posterior samples could consistently predict structures in places where they are not, or consistently predict the absence of structures where, in fact, such structures should be. To show that a trained MicroSplit network does not consistently hallucinate the presence or absence of structures, we propose to check how well calibrated the network is [7,13–15]. A calibrated neural network is one whose predicted probabilities or uncertainties accurately reflect the true likelihood of outcomes. In the context of a regression task like the one executed by MicroSplit, calibration means that the predicted error estimates of the model align with the actual distribution of errors in the data. We therefore estimated the true error by computing the pixel error between the MMSE prediction of MicroSplit and the ground truth target images we derived from the available training data, and plotted this “true error” against the RMSE errors we described above. As the calibration plots in Figure 2b show, the true error and RMSE scale roughly linearly with respect to each other, meaning that the estimated RMSE, which can be computed from posterior samples only, is a good estimator of the true error. Hence, calibrated MicroSplit networks offer a reliable and efficient way to estimate the uncertainty of the data and, therefore, also the magnitude of the error of its predictions. This property is a considerable advantage over non-variational approaches and will help future users to (i) filter images or image regions that lead to uncertain predictions, and (ii) provide evidence to themselves and others that unmixed images do not suffer from hallucinations. 2.4 Downstream Processing of Unmixed Data Sample preparation and microscopy data acquisition are typically only the first steps in the pursuit of scientific insight. Once images have been acquired, their content must be analyzed and quantified. In this work, we cannot cover the full breadth of image data analyses that life scientists routinely perform on microscopy data [16]. However, in many analysis pipelines, image segmentation is a key step since it determines the identity, location, shape, and relative location of biological entities of interest. In the following section, we have quantified the segmentation performance on MicroSplit predictions compared to the same analysis carried out on traditionally multiplexed image data. 2.4.1 Segmentation of Unmixed Image Data is of High Quality We show that a typical segmentation task, as it is conducted in biological research laboratories on a daily basis, leads to comparable quality results when conducted on regular multiplexed microscopy data or on unmixed images. To this end, we compared the segmentation results of three bioimage analysts who were instructed to interactively train a pixel-classifier until they reach best-possible results. We show the results in Figure 3and Figure S7, and explain the details of this experiment below. First, each analyst segmented the single-channel microscopy data acquired by regular multiplexed fluorescent microscopy (Target) and the two-channel semantic unmixing predictions of MicroSplit (Prediction) without being informed about the nature of the data. We then compared the consistency between the segmentations of each analyst on the two sets of data and the consistency of the segmentations between the analysts. We observed that the variability between segmentations on target and unmixed images lies within the range of the inter-observer variability between the three analysts. We qualitatively show the segmentation results on two experiments (Tasks) in Figure 3bc and plot the measured intra-observer variabilities (A1,A2,A3) and the inter-observer variability. Hence, in all experiments we conducted, the quality of segmentations remained unchanged when using unmixed image data compared to multiplexed images. But more interestingly, imaging two structures in a single channel can free up photon budget, which then becomes available for imaging at a higher frame rate, imaging at a higher signal-to-noise ratio, or imaging more labeled structures in additional image channels. 2.4.2 Removing Unwanted Imaging Artifacts Naturally, MicroSplit predictions are just images and any subsequent analysis can be performed with them. As a second downstream processing example, we present an innovative way to remove imaging artifacts. More specifically, we imaged the post-mitotic neuronal marker CTIP2 in sliced hPSC-derived forebrain organoids using a secondary antibody conjugated to a 555 dye. The resulting image data not only had CTIP2+ nuclear staining (wanted signal), but also puncta that accumulated because of non-specific signal (i.e. undesired artifacts). 8 Fig. 4 Removing Image Artifacts with MicroSplit. We remove unwanted puncta present in raw microscopy data using our image unmixing approach by (i) selecting a set of cropped regions of interest (ROI) in the raw data that do not contain the unwanted structures (puncta, in this case), and other set of ROIs that do only contain said puncta without the fluorescent signal of interest, followed by (ii) creating superimposed input images by mixing (summing) randomly sampled ROIs of the two collected sets. These mixed images, together with the original pairs of ROIs, are now the input and target input used to train a MicroSplit network. (a) An example of a raw input image (left), showing labeled nuclei and unwanted puncta, and the puncta-free prediction by MicroSplit. Both images show intensity histograms in cyan and red, respectively. The dashed gray plot in the predicted image corresponds to the cyan plot of the input. The y-axes are shown on a logarithmic scale to better emphasize the puncta intensities MicroSplit successfully removed. (b) For better visual comparison, we show input and prediction insets from (a) in the top two rows, respectively, and the difference images (input - prediction) in the bottom row (red, intensities removed; gray intensities unchanged; blue, intensities increased). Ideally, we could use MicroSplit to unmix the labeled nuclei (desired image content) and the undesired puncta (i.e. imaging artifacts). However, for this we would need training targets that contain only nuclei and others that contain only puncta, a requirement that the experimental setup of our colleagues does not permit. In Section 2.1, we introduce Training Mode II-b, the manual selection of image regions that show only the artifacts (puncta) or only the desired structures (the signal, in this example, nuclei). The selected image regions are then combined, just as in any Training Mode II experiment, and used to train a MicroSplit network. Since crops without puncta or crops that show only puncta are of limited size with respect to full microscopy images (the smallest crops we collected were 100 ×100 pixels), lateral contextualization (LC; see Figure 2a) could not be used and was therefore disabled during training. In total, we have manually cropped 82 content regions and 48 puncta regions from the data. This led to a total of 53.7Mpixels for training and took an analyst about eight hours of focused work. After training had converged, we evaluated the model on previously held out full-size microscopy images. A representative result is shown in Figure 4a,b. The insets in Figure 4b allow a more detailed comparison 9 References [1] Shroff, H., Testa, I., Jug, F., Manley, S.: Live-cell imaging powered by computation. Nat. Rev. Mol. Cell Biol. (2024) [2] Weigert, M., Royer, L., Jug, F., Myers, G.: Isotropic reconstruction of 3d fluorescence microscopy images using convolutional neural networks. In: Descoteaux, M., Maier-Hein, L., Franz, A., Jannin, P., Collins, D.L., Duchesne, S. (eds.) Medical Image Computing and Computer-Assisted Intervention - MICCAI 2017, pp. 126–134. Springer, Cham (2017) [3] Weigert, M., Schmidt, U., Boothe, T., M¨uller, A., Dibrov, A., Jain, A., Wilhelm, B., Schmidt, D., Broaddus, C., Culley, S., Rocha-Martins, M., Segovia-Miranda, F., Norden, C., Henriques, R., Zerial, M., Solimena, M., Rink, J., Tomancak, P., Royer, L., Jug, F., Myers, E.W.: Content-aware image restoration: pushing the limits of fluorescence microscopy. Nature Publishing Group 15(12), 1090–1097 (2018) [4] Krull, A., Buchholz, T.-O., Jug, F.: Noise2Void - learning denoising from single noisy images, pp. 2129–2137 (2019) [5] Prakash, M., Krull, A., Jug, F.: DivNoising: Diversity denoising with fully convolutional variational autoencoders. ICLR 2020 (2020) [6] Ashesh, Krull, A., Sante, M.D., Pasqualini, F.S., Jug, F.: µSplit: image decomposition for fluorescence microscopy. (2023) [7] Ashesh, A., Jug, F.: denoisplit: A method for joint microscopy image splitting and unsupervised denoising. In: Leonardis, A., Ricci, E., Roth, S., Russakovsky, O., Sattler, T., Varol, G. (eds.) Computer Vision – ECCV 2024, pp. 222–237. Springer, Cham (2025) [8] Ashesh, A., Deschamps, J., Jug, F.: MicroSSIM: Improved Structural Similarity for Comparing Microscopy Data. In: BIC Workshop, ECCV, Milan (2024). https://doi.org/10.48550/arXiv.2408.08747 [9] Prakash, M., Delbracio, M., Milanfar, P., Jug, F.: Interpretable unsupervised diversity denoising and artefact removal. (2021) [10] Vahdat, A., Kautz, J.: NVAE: A deep hierarchical variational autoencoder. In: Neural Information Processing Systems (NeurIPS) (2020) [11] Sønderby, C.K., Raiko, T., Maaløe, L., Sønderby, S.K., Winther, O.: Ladder variational autoencoders. In: Proceedings of the 30th International Conference on Neural Information Processing Systems. NIPS’16, pp. 3745–3753. Curran Associates Inc., Red Hook, NY, USA (2016) [12] Kingma, D.P., Welling, M.: An introduction to variational autoencoders (2019) arXiv:1906.02691 [cs.LG] [13] DeGroot, M.H., Fienberg, S.E.: The comparison and evaluation of forecasters. Journal of the Royal Statistical Society. Series D (The Statistician) 32(1), 12–22 (1983). Accessed 2025-01-26 [14] Niculescu-Mizil, A., Caruana, R.: Predicting good probabilities with supervised learning. In: Proceedings of the 22nd International Conference on Machine Learning. ICML ’05, pp. 625–632. Association for Computing Machinery, New York, NY, USA (2005). https://doi.org/10.1145/1102351.1102430 . https://doi.org/10.1145/1102351.1102430 [15] Levi, D., Gispan, L., Giladi, N., Fetaya, E.: Evaluating and calibrating uncertainty prediction in regression tasks. Sensors 22(15) (2022) [16] Schroeder, A.B., Dobson, E.T.A., Rueden, C.T., Tomancak, P., Jug, F., Eliceiri, K.W.: The ImageJ ecosystem: Open-source software for image visualization, processing, and analysis. Protein Sci. 30(1), 234–249 (2021) [17] Blau, Y., Michaeli, T.: The perception-distortion tradeoff. In: 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pp. 6228–6237. Computer Vision Foundation / IEEE Computer Society, ??? (2018). https://doi. org/10.1109/CVPR.2018.00652 .http://openaccess.thecvf.com/content cvpr 2018/html/Blau The PerceptionDistortion Tradeoff CVPR 2018 paper.html 16 [18] Shroff, H., Testa, I., Jug, F., Manley, S.: Live-cell imaging powered by computation. Nature Reviews Molecular Cell Biology 25(6), 443–463 (2024) https://doi.org/10.1038/s41580-024-00702-6 . Accessed 2025-11-18 [19] Zimmermann, T.: Spectral imaging and linear unmixing in light microscopy. Adv. Biochem. Eng. Biotechnol. 95, 245–265 (2005) [20] Seo, J., Sim, Y., Kim, J., Kim, H., Cho, I., Nam, H., Yoon, Y.-G., Chang, J.-B.: PICASSO allows ultramultiplexed fluorescence imaging of spatially overlapping proteins without reference spectra measurements. Nat. Commun. 13(1), 1–17 (2022) [21] Nogare, D.D., Hartley, M., Deschamps, J., Ellenberg, J., Jug, F.: Using AI in bioimage analysis to elevate the rate of scientific discovery as a community. Nat. Methods 20(7), 973–975 (2023) [22] Deschamps, J., Dalle Nogare, D., Jug, F.: Better research software tools to elevate the rate of scientific discovery or why we need to invest in research software engineering. Front Bioinform 3, 1255159 (2023) [23] Laporte, M.H., Klena, N., Hamel, V., Guichard, P.: Visualizing the native cellular organization by coupling cryofixation with expansion microscopy (Cryo-ExM). Nature Methods 19(2), 216–222 (2022) https://doi.org/ 10.1038/s41592-021-01356-4 . Accessed 2025-02-02 [24] Gong, D., Cai, C., Strahilevitz, E., Chen, J., Scherer, N.F.: Easily scalable multi-color DMD-based structured illumination microscopy. Opt. Lett. 49(1), 77–80 (2024) [25] Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P.: Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing 13(4), 600–612 (2004) https://doi.org/10.1109/TIP. 2003.819861 [26] Wang, Z., Bovik, A.C.: A universal image quality index. IEEE Signal Processing Letters 9(3), 81–84 (2002) https://doi.org/10.1109/97.995823 [27] Wang, Z., Simoncelli, E.P., Bovik, A.C.: Multiscale structural similarity for image quality assessment. In: The Thrity-Seventh Asilomar Conference on Signals, Systems and Computers, 2003, vol. 2, pp. 1398–14022 (2003). https://doi.org/10.1109/ACSSC.2003.1292216 [28] Di Sante, M., Pezzotti, M., Zimmermann, J., Enrico, A., Deschamps, J., Balmas, E., Becca, S., Reali, A., Bertero, A., Jug, F., Pasqualini, F.S.: Calipers: Cell cycle-aware live imaging for phenotyping experiments and regeneration studies. bioRxiv (2024) https://doi.org/10.1101/2024.12.19.629259 https://www.biorxiv.org/content/early/2024/12/22/2024.12.19.629259.full.pdf [29] Walsh, R.M., Luongo, R., Giacomelli, E., Ciceri, G., Rittenhouse, C., Verrillo, A., Galimberti, M., Bocchi, V.D., Wu, Y., Xu, N., Mosole, S., Muller, J., Vezzoli, E., Jungverdorben, J., Zhou, T., Barker, R.A., Cattaneo, E., Studer, L., Baggiolini, A.: Generation of human cerebral organoids with a structured outer subventricular zone. Cell Reports 43(4), 114031 (2024) https://doi.org/10.1016/j.celrep.2024.114031 . Accessed 2025-07-21 17 A Model Architecture and Training In MicroSplit, we have combined the benefits of µSplit [6] and denoiSplit [7], cast those ideas in a common learning framework, enabled direct training and prediction on volumetric data, and have extensively tested and evaluated its performance on a wide range of datasets and semantic unmixing tasks. We have also made openly available all training and prediction code and all data we used. We do this to foster rapid adoption of MicroSplit by the scientific community and to allow others to improve our approach and compare our results with relative ease. In this section, we will discuss in detail the aspects of the above-mentioned works that were integrated into MicroSplit. Subsequently, we also describe our loss function and important training hyperparameters. From µSplit, we inherit the ability to efficiently incorporate spatial context using additional inputs, called Laterally Contextualization (LC) inputs [6]. Here we feed, next to the primary input, for which the prediction will be made, additional LC inputs that help the network to better understand the image context from which the primary input is taken (see Figure 2a and S1). These successive LC inputs are larger and larger patches centered on the primary input patch, but downscaled to the same pixel dimensions as the primary input itself. Hence, LC inputs capture the spatial context around the primary input patch but do so at lower resolutions to ensure efficient learning and predictions in a reasonably sized overall network. The network architectures we proposed come in flavors that trade computational complexity and GPU consumption with the best possible prediction quality. Its most GPU-efficient variant, Lean-LC, can train on a single GPU using less than 5GB of GPU memory. If more resources are available, it is advisable to opt for setups such as Deep-LC, which show better predictive performance at increased computational cost. In Figure S1, we have briefly described the three µSplit variants. The white regions correspond to features originating from the primary input patch. As in U-Net architectures, spatial resolution halves at each successive hierarchy level through pooling operations, causing these white areas to progressively diminish in size. In µSplit (and hence also in MicroSplit), the pooled embeddings undergo zero-padding before being concatenated with feature maps from lateral contextualization (LC) inputs. These LC features are processed through dedicated ‘Input branch’ sub-networks consisting of convolutional layers with nonlinear activations, dropout, and normalization components. ‘Input branch’ does not have pooling operations, and so their feature maps, at each hierarchy level, maintain spatial dimensions identical to the primaryinput (represented by gray regions). This preservation enables merged embeddings from both pathways to retain the original spatial dimensions throughout the network hierarchy (gray squares). This is the core idea of LC MicroSplit has inherited from µSplit. In the caption of Figure S1 we provide more details regarding the differences in the three variants of µSplit. Please refer to [6] for more details. From denoiSplit, we inherit the ability to jointly perform unsupervised denoising, using suitable Noise Models. This also enables MicroSplit to sample diverse predictions from a learned approximate posterior that captures a notion of the data uncertainty, as demonstrated by our trained networks being calibrated (see Section 2.3). While also µSplit is a variational approach that is in theory capable of generating multiple predictions from its posterior, we found that denoiSplit, arguably due to its different KL-loss formulation, produces a higher diversity that is better in line with the uncertainty in the data. In Figure S1, we present the architecture of MicroSplit, which has LC inputs and Noise models, all integrated into a single setup. As mentioned above, we also enabled MicroSplit to operate directly on volumetric image data, a possibility that was absent in both µSplit and denoiSplit. A.1 Loss Function used to train MicroSplit The loss function of MicroSplit is the weighted average between the µSplit loss and denoiSplit loss, that is, lossMicroSplit =w∗lossdenoiSplit + (1 −w)∗lossµSplit (2) Unless explicitly specified, w= 0.9 is used in all experiments we conducted. This simple design also gives us the ability to switch to pure µSplit or denoiSplit mode by simply setting wto 0 or 1, respectively. To incorporate LC inputs into the denoiSplit setup, we observed the need to modify the KL loss formulation used in lossdenoiSplit. In denoiSplit, pixel-wise KL divergence is computed at every hierarchy level. Let KLidenote the pixel-wise KL divergence tensor at the ith hierarchy level. KL-loss component for this hierarchy level, kliis defined as kli=α·X j,h,w KLi[j, h, w].(3) 18 Fig. S1 Network architectures employed by MicroSplit. This figure provides a detailed breakdown of the components shown in Figure 2a. Specifically, it illustrates the encoder-decoder structure of our architecture, which is adapted from our prior work, µSplit [6]. Since this diagram focuses on the transformation of input data into predictions, it does not include the loss terms (such as the KL-divergence loss and the noise-model-based likelihood loss). As in [6], we improve the predictive capacity of MicroSplit by feeding not only an image patch (primary input), but also additional image context. To this end, we introduced lateral contextualization (LC). These LC inputs capture a larger portion of the input data but at increasingly lower pixel-resolution. This lead to three meta-architectures, Lean-LC, Regular-LC, and Deep-LC, each with increasingly higher computational and GPU memory demands but also leading to increasingly higher unmixing performance. The gray region shown in the embeddings in all three meta-architectures represents the extra spatial size in the embeddings coming solely due to LC inputs. In Lean-LC, embeddings of different hierarchy levels in the encoder benefit from LC inputs, but the decoder does not. In Regular-LC, a more GPU-consuming variation, both the encoder’s and the decoder’s embeddings use LC (i.e. show gray regions in the figure). Finally, since in Regular-LC the spatial dimensions of the embedding layers do not decrease, it enabled us to stack some additional hierarchy levels on top of the previously used ones. We labeled the resulting architecture as Deep-LC, the most performant but also most resource heavy variation. In order to get best results, we have used Deep-LC whenever possible, but it is important to note that with just two simple hyper-parameter switches, one can instantly make use of Regular-LC or Lean-LC and train on cheaper and older consumer GPUs. 19 With LC inputs, the spatial dimensions of the latent space tensors and therefore KLido not decrease and the summing operation in this formulation leads to a higher value, owing to the larger number of summands (which are all non-negative). This gives unnecessarily high weight to the KL loss with respect to the likelihood loss component. We observed that this degrades the performance. To handle this, we center-cropped KLito the shape they would assume if there were no LC inputs. Let KLcropped idenote the appropriately center-cropped version of KLi. Our modified KL loss for denoiSplit becomes kli=α·X j,h,w KLcropped i[j, h, w].(4) We encountered a very similar issue when working with volumetric data. In that case, pixel-wise KLdivergence is a 4-dimensional tensor C×Z×Hi×Wi. As we work with larger and larger Z, the summation in Equation 3would increase since the summation would be on all 4 dimensions. This again leads to giving more weight to the KL loss component against the likelihood loss, thus rendering the performance inferior. Note that this affect will become more severe when we increase the number of Zframes in the input. In other words, adding more information in the input was not beneficial. To handle this, we separately took care of the extra Zdimension by taking the average along this dimension. The resultant 3D tensor is then passed to Equation 4to compute the KL loss component. Note that there is an additional dimension of batch size which we have not mentioned in the above explanation. That is because KL-loss is computed separately for every element in the batch. A.2 Hyper-parameters used during Training We use PyTorch package for creating our training and evaluation pipelines. We use a batch_size of 32, max_epoch of 400 and learning_rate of 0.001. We use Adamax optimizer and ReduceLROnPlateau as the learning rate scheduler with lr_scheduler_patience set to 150. During training, we use 16-bit precision. We use 2 LC inputs in our Deep-LC configuration (multiscale_lowres_count = 3). Please refer to our code (https://github.com/CAREamics/MicroSplit-reproducibility) for more details. We want to state that all experiments were done using code hosted at https://github.com/juglab/MicroSplit. However, we have developed https://github.com/CAREamics/MicroSplit-reproducibility with the objective of providing user-friendly code with easier adaptability to custom datasets. B Analyzing Factors that Affect Predictive Performance B.1 Pixel-noise (pixel-wise independent noise) From our different experiments, we have found that there are two prominent factors that affect the performance on semantic unmixing tasks. The first factor is the amount of pixel-independent noise that is present in the superimposed input images and target images. More noise means inferior performance. To quantify the effect of noise, we imaged our HT-LIF24 dataset with exposure durations of 2ms, 3ms, 5ms, 20ms, and 500ms. We imaged it such that the underlying content in these sub-datasets is identical. That is, for every frame in the 2ms acquisition, we have the corresponding higher SNR frames in 3ms, 5ms, 20ms and 500ms acquisitions. We trained a three-channel semantic unmixing task separately for each exposure duration subdataset. We made predictions on the held-out test set input frames from their respective exposure duration sub-datasets. We evaluate the prediction against the target channels present in the 500ms sub-dataset. Although the results (Tasks XXIV-XXIX) in Table 1show the expected trend (semantic unmixing quality decreases when using lower SNR training data), even the shortest exposure time of 2ms still lead to unmixed predictions that are fit for downstream processing and analysis (in all cases we measured a PSNR >30.8 and MicroMS-SSIM >0.94). To explain the performance drop, supplementary figure S60 shows that reduced PSNR primarily results from the loss of high-frequency details. B.1.1 Lessons learned from the difficult Task XXIII (on HHMI-D25 data) Out of all tasks mentioned in the Table 1, the task XXIII which uses HHMI-D258bit dataset has considerable performance issues, especially in the third channel (see Supplementary Figure S57). In the next few paragraphs we will describe our approach to investigating this issue and present a working solution. We do this to provide an example for how users of MicroSplit can improve solutions that might initially not lead up to the required semantic unmixing quality. 20 Signal-to-noise (SNR) is one of the factors which plays a role in almost all deep-learning based methods and denoising approaches, and semantic unmixing is no exception. To investigate the role of SNR for task XXIII, we first denoised the raw data of HHMI-D258bit using Noise2Void [4], and then trained MicroSplit using those denoised images, calling this training task ‘Task XXXI’. Comparing the results of tasks XXIII and XXXI, see also Table C.3, it becomes apparent that prior denoising has a rather strong positive effect on the quality of the achieved semantic unmixing results (>7db PSNR improvement), suggesting that the low SNR in the original data might indeed have caused the bad performance. Since HHMI-D258bit is stored in unsigned int8, meaning that pixel values are in [0,255], there are actually only a few distinct integer values presenting most of the data. Denoising this data, besides increasing the SNR, also increases the number of unique pixel values. We hypothesized that in addition to SNR, an overly discreet nature of pixel values might also be detrimental to semantic unmixing. To test this hypothesis, we imaged the HHMI-D2516bit data subset, where the unsigned int16 format was used to store the data, increasing the pixel intensity range to [0,65535]. Doing so, we ensured that the SNR (ratio of average foreground value to average background value) is as similar as possible for both, the 8 and 16 bit versions of the HHMI-D25 data. Using the 16bit data did indeed improve the quality of predictions, also for the problematic third channel, as can be seen in Supplementary Figures S58 and S62. The quantitative metrics in Table C.3, however, do not capture the improvement we can perceive by comparing those figures. To fully validate our hypothesis regarding SNR and overly discrete pixel intensities, we imaged another data subset of HHMI-D25, namely HHMI-D2516bit,0.25, which not only uses unsigned uint16 format but also bins four pixels into one (thereby increasing SNR on the cost of spatial resolution). On this data subset we defined Task XXXVI, and indeed observe a much improved semantic unmixing performance (see Table C.3 and Figure S59). We also experimented with synthetic Gaussian and Poisson noise with HHMI-D25 dataset versions. The motivation was to start from a working setup, make the training and evaluation dataset noisy and inspect the performance degradation. For this purpose we picked HHMI-D2516bit and HHMI-D258bit,denoised. We added Gaussian noise (σ) and Poisson noise (λ). Given an image x, its noisy version can be expressed Poi(x/λ)·λ+ϵ, where ϵ∼N(0, σ) and Poi() represents the Poisson distribution with parameter λ. As it was to be expected, the performance degrades with noise (see Tasks XXXII, XXXIV, and XXXV in Supplementary Figures S62, and S61, and a quantitative comparison in Supplementary Table C.3. B.1.2 Out-of-distribution SNR . We also use the set of models trained on different HT-LIF24 sub-datasets to understand how the performance degrades with out-of-distribution inputs. For this, we evaluate the performance of MicroSplit trained on one exposure duration on the superimposed input images coming from a different exposure duration. We present the results in Figure S5. The different curves represent individual MicroSplit models trained on one specific exposure duration sub-dataset as specified in the legend. On the x-axis, we have different evaluation sub-datasets, referred to by their exposure duration. From CARE-PSNR and MicroMS-SSIM plots, one can observe that performance improves as one increases the exposure duration. Additionally, upon observing performance on 2ms and 500ms sub-datasets, one can see that in most cases, the larger the difference between the exposure duration of the training sub-dataset and evaluation sub-dataset, the higher the performance drop. For instance, if we look at CARE-PSNR plot, the two worst-performing MicroSplit models on 2ms acquisition were trained on 20ms and 500ms. And the two worst-performing MicroSplit models on 500ms acquisition were trained on 2ms and 3ms. On a different note, one can observe much less variation in SSIM and MS-SSIM plots. We discuss this aspect in Section E. B.2 Spatial Correlation The structures present in the cell are spatially correlated. For example, nuclei are typically in the central regions of cells, while the cell boundary is, by definition, on the boundary of a cell. The knowledge about the cell surface, therefore, can tell something about where the nucleus should or should not be present. We wanted to understand how important this spatial correlation is for our semantic unmixing task. For this, we worked with the Pavia-P23 dataset where we modified the input patch process during training. The default method is to pick a random location in a frame, extract patches for both structures (target channels) from that location, and sum them to create superimposed input (i.e.Training Mode I). Using Training Mode I , the spatial correlation between the imaged structures is naturally maintained. To disrupt this, we conducted experiments with the following alterations. 21 In the first alteration, we kept Training Mode I for 50% of all training patches, but created the other 50% of training patches by picking two different random locations and adding them together to create an input patch (i.e.Training Mode II ). Hence, we maintain sound spatial correlations between the structures to be unmixed in half the training data. In the second alteration, we create all training data patches according to Training Mode II , thus eliminating all spatial correlations between the structures that should be unmixed and forcing the trained network to rely fully on the structural appearance of the structures only. We report the metrics CARE-PSNR and MicroMS-SSIM in Table ST8 which shows that the absence of spatial correlation indeed results in a drop in performance by 0.3−0.5 dB PSNR. B.3 Similarity of Structures to be Unmixed Since our method relies heavily on the spatial appearance of the structures to be unmixed, we wondered how dissimilar two structures have to be to lead to good semantic unmixing results. Hence, it is best to image structures in single channels that are as dissimilar as possible. In µSplit [6], we proposed a synthetic semantic unmixing dataset based on combinations of sinusoidal curves that required longer-range structure integration. Although the structures were very simple, the network without using LC was unable to split the input and ended up only predicting the input for both target channels equally. However, to better evaluate to what extent a network can separate structures that have a similar appearance, we designed the following experiments. Using the microtubule channel of the HT-LIF24 dataset, we created a 2-channel splitting task by mixing patches from the microtubule channel with other patches from the exact same channel. We reasoned that the network will not be able to split structures that are literally taken from the same set of data (see bottom row in Figure S2). To give the network a chance, we then started to alter one of the two superimposed copies by scaling the data (uppermost four rows in Figure S2). This leads to the superposition of data that is structurally still very similar, but one copy will have a slightly larger appearance. In Figures S2 and S3 we show the results of these experiments, conducted with a range of different relative scaling between the two copies of the microtubule data. The results clearly show that the network is capable of semantic unmixing the structures in all cases, even if the scaling factor was as low as 1.125. Note also that it takes the network increasingly longer before the splitting performance starts lifting off and reaching its peak, as can be inferred from the inflection point and convergence behavior seen in the PSNR vs. training steps plot in Figure S3. B.4 Unequal Channel Intensities Another factor that plays a key role in the performance of MicroSplit is the skewness in intensity between different target channels in the superimposed input image. This effect can be caused by diverging fluorophore densities for different structures, by fluorophore bleaching, or by using inadequate laser lines and/or laser power settings for some fluorophores. Any of this can lead to acquisitions where one or more of the structures are only weakly present in the superimposed input. We hypothesized that predictions for the brighter (dominant) structure should still be of good quality but that the quality for the dim structure might be worse. To test this hypothesis, we worked with the 2-channel semantic unmixing tasks from the Pavia-P24 dataset, where we varied the laser power for the two channels we imaged from a balanced setting (50:50) to increasingly more skewed settings (i.e. 66:33 and 84:16). Note that the total laser power was kept the same. We acquired images at these three levels of skewness at three different total laser powers and resulting signalto-noise ratios, giving us a total of 9 sets of image acquisitions. We show the results of all semantic unmixing experiments in Table ST2, where ‘Skew’ denotes the asymmetry in the power distribution (‘Balanced’, 50:50; ‘Mid’, 66:33; ‘High’, 84:16). We found that, across three total laser power levels, the performance of the bright first channel generally increases as we go from Balanced to High skew. The second dim channel shows the inverse behavior. Note that since the samples we imaged at different imaging settings were not the same, each experiment is based on a different body of training and evaluation data. This causes additional fluctuations in the performance metrics we report and makes the presented results to be non-monotonic. If all 9 imaging sessions had captured the same samples, we would expect the results to show a monotonic trend. 22 Unequal Channel Intensities - Two Extreme Examples For a first example, we worked with the nucleus and tubulin channels of the Chicago-Sch23 dataset. As always, we create the superimposed input by summing the two channels. Here, the nucleus is very weak (the channels being very skewed in their relative intensity), so much so that nuclei are de facto invisible to the naked eye. However, MicroSplit was still able to unmix these channels at a reasonable quality. One can observe the superimposed input, the two targets, and the predictions of MicroSplit in Figure S8. For a second example, we inspect Task II which uses the HT-P23A dataset. Here, the two structures are microtubules and nuclei, with the latter being the weaker channel. Similarly to the first example, the weaker nucleus channel is effectively invisible in superimposed images, as can be seen in Figure S9. However, unlike in the previous example, the structure of nuclei in this dataset is less regular and more variable. Due to this and in light of the high uncertainty caused by the highly skewed channel intensities, MicroSplit MMSE predictions for the nucleus channel become rather blurry, as can be seen in Figure S9. Still, in many practical applications, i.e. every time the detailed texture of the nuclei is not important, such predictions can still be sufficient for the desired data analysis to be carried out. B.5 Sufficient Lateral Image Context Our method combines the benefits of µSplit and denoiSplit. One of the benefits of µSplit architecture we demonstrated in [6] is its ability to utilize additional surrounding spatial context for a given superimposed input patch through LC inputs and its ability to employ much deeper architectures. In [6], we showed that having sufficient spatial context helps with relatively large structures spanning hundreds of pixels. C Additional Experiments C.1 MicroSplit vs. PICASSO In this section, we compare the performance of MicroSplit with PICASSO [20] (see also Figure S4). For semantic unmixing kchannels MicroSplit needs as input a single superimposed image from the microscope whereas PICASSO needs kimages from the microscope which correspond to kspectrally overlapping fluorophores. Due to this mismatch in the data requirement, a direct comparison is not feasible. So, to compare MicroSplit with PICASSO, we generate synthetic inputs using our HT-LIF24 dataset. However, we argue that our way of generation does not degrade the performance of PICASSO but instead, it should be easier for PICASSO to predict on this data as compared to the real data. Since fluorophores can be ordered according to the wavelength of the maximum intensity in their emission spectra, we first define such an order of our structure types. Next, for generating every channel of the input for PICASSO, we define three weights and take the weighted average of the three structures using these weights. The weights are set according to the order of the structures set above. For example, for the first channel, the weight given to the second structure will be higher than the weight given to the third structure. We generate two sets of weights, one being harder than the other. In the hard case, the dominant structure type is given 1.5 times more weight than the next dominant type and 3 times more weight than the least dominant type. In the easy version, the dominant structure type is given 2.5 times more weight than the next dominant type and 5 times more than the least dominant type. Once the input channels are generated, we add Gaussian noise of σ= 500 to each channel independently. We also train MicroSplit on this data with the same Gaussian noise applied on top of the data. In Figure S4, we show the results on one random frame. In Table ST6 and Table ST7, we show the quantitative results. C.2 Training Mode I vs Training Mode II - How Important are Spatial Correlations? In this experiment, we inspect the performance drop between data acquisition types I and II. In cells, the location of different structures is often quite co-related with one another. For example, actin is mostly concentrated on the cell periphery whereas the nucleus is typically found at the center of the cell. In acquisition modes I and II, input is created by summing the crops from individual channels. In acquisition type I, since the channels are independently acquired, summing the crops of these different channels will generate an input patch where the naturally occurring co-location property cannot be preserved. In acquisition type II, since both channels are concurrently acquired, the inputs are created from summing the crops of different 23 Task Dataset Synthetic Noise PSNR MicroMS-SSIM XXIII HHMI-D258bit - 22.5 31.3 24.3 0.840 0.768 0.793 XXXI HHMI-D258bit,denoised - 37.9 35.2 38.6 0.990 0.903 0.974 XXXII HHMI-D258bit,denoised σ= 20, λ = 30 31.8 27.3 32.4 0.939 0.664 0.859 XXXIII HHMI-D2516bit - 23.2 27.9 24.8 0.772 0.849 0.779 XXXIV HHMI-D2516bit σ= 2K, λ = 5K23.2 27.9 24.8 0.827 0.861 0.778 XXXV HHMI-D2516bit σ= 4K, λ = 10K23.0 27.4 24.6 0.746 0.838 0.700 XXXVI HHMI-D2516bit,0.25 - 32.4 32.2 35.0 0.990 0.929 0.991 Table ST1 Performance of MicroSplit on sub-datasets of HHMI-25. This table presents quantitative results of MicroSplit across various tasks defined on parts of the HHMI-25 dataset. While the original Task XXIII on HHMI-D258bit does not lead to satisfying predictions, mainly for channel 3 (see Supplementary Figure S57), tasks on similar data with higher SNR (Task XXXI, see Supplementary Figure S61, row 3), increased pixel diversity (Task XXXIII, see Supplementary Figure S62, row 2), or both (Task XXXVI, see Supplementary Figure S59) demonstrate notably improved semantic unmixing performance. Tasks XXXII, XXXIV, and XXXV are identical to Tasks XXXI and XXXIII, respectively, but with added synthetic noise, simply to demonstrate how the lower SNR drops the unmixing performance achievable with MicroSplit. Channel 1 Channel 2 Skew Laser Power Laser Power High Mid Low High Mid Low Balanced 24.3 23.1 21.9 29.9 24.3 23.0 Mid 28.2 24.0 22.9 25.6 22.3 21.7 High 25.2 24.3 23.3 24.1 22.4 22.8 Table ST2 Varying the laser power and skew with Pavia-P24 dataset. Skew column denotes the relative importance given to channel 1. High skew means larger laser power allocated to Channel 1 compared to Channel 2 channels with each crop taken from the same location in the micrographs and therefore these inputs preserve the naturally occurring co-location property. In this experiment we quantify the benefit of using this co-location information. We work with the HT-T24 dataset which falls under Training Mode III . We train three models. In the first model, we create the input using Training Mode I , that is, we create the input by picking target patches from the same location and therefore maintain the spatial co-relation. In the second model, we use Training Mode II , meaning that we pick crops from random locations from the different channels and use them to create the input. This model naturally does not have access to naturally occurring spatial co-relation information in its training data. The third model is trained using Training Mode III . We evaluate all three models on the held-out test set where the inputs have spatial co-relation preserved and are not synthetic, that is, they are imaged from the microscope. In Table ST3, we find that the first model outperforms the second by 1.3 CARE-PSNR and 0.009 MicroMS-SSIM. Naturally, the model trained with Training Mode III is best and outperforms Training Mode I by 0.6 CARE-PSNR and 0.004 MicroMS-SSIM. C.3 Training Mode I vs.Training Mode IIISummed vs. Acquired Inputs Here, we quantify how much the performance degrades if, during training, the input is created by simply summing the two channels as compared to input coming directly from the microscope. We find that while there is indeed a performance drop as can be seen in Table ST3, the drop is not detrimental. This experiment shows the utility of our approach in the case when synthetic input is used for its training but for evaluation, inputs coming directly from the microscope are used. C.4 MicroSplit enables a more effective use of the available photon budget By filtering fewer photons: Traditional multiplexed imaging relies on emission filters that selectively pass photons from one fluorophore at a time in order to minimize spectral overlap. As a consequence, a substantial fraction of emitted photons is discarded, and relaxing the filters leads to bleedthrough artifacts. MicroSplit changes this trade-off because multiple structures can be imaged in the same acquisition. This allows microscopists to use substantially broader emission filters, collecting photons from several fluorophores simultaneously, without introducing the ambiguities that would otherwise arise in a multi-channel setting. To roughly quantify this advantage, we analyzed three fluorophores from the HT-LIF24 dataset (DAPI, FITC, TRITC), using their emission spectra (downloaded from fpbase.org) normalized as probability mass 24 Model PSNR SSIM Training Mode I: input = C1+C235.9 0.956 Training Mode II: input = C1+C2(shuffled) 34.6 0.947 Training Mode III: input comes from microscope 36.5 0.960 Table ST3 Performance comparison for different Acquisition Modes. We use the Sox2 vs Golgi splitting task of the HT-T24 dataset for this purpose. For Acquisition Mode I and II, real input image is not used during training. Instead, input is created by synthetically summing the two target channels. In all cases, evaluation is done on the held-out test set using the real input channel, that is, on the input which is not synthetic and comes directly from the microscope. C1 C2 (2D) Z=1 Z=5 Z=9 Z=15 (2D) Z=1 Z=5 Z=9 Z=15 CARE-PSNR 35.8 38.7 39.5 39.7 30.9 33.7 34.5 34.7 MicroSSIM 0.865 0.878 0.885 0.886 0.729 0.757 0.772 0.767 MicroS3IM 0.950 0.970 0.973 0.974 0.929 0.950 0.956 0.956 Table ST4 Performance improvement with 3D models on HT-H24 dataset. As we increase the number of z-slices fed to the model, we see the performance improve in all our metrices. Set I Set II C1C2C3C1C2C3 Input C10.50 0.33 0.17 0.625 0.25 0.125 Input C20.33 0.50 0.33 0.25 0.625 0.25 Input C30.17 0.33 0.50 0.125 0.25 0.625 Table ST5 We use the weights to mix the three channels C1,C2and C3. functions. We compared the fraction of photons transmitted by conventional multi-color filter configurations to the photons captured when using a broad, highly permissive filter suitable for MicroSplit. Across three representative scenarios, with emission filter thresholds chosen to (i) maximize photon collection, (ii) reduce bleedthrough by 25%, and (iii) reduce bleedthrough by 50%, with conventional multiplexed imaging being respectively 22%, 34%, and 55% less photons efficient than the results obtained using MicroSplit (see top half of Supplementary Figure S6). In practical terms, this means that even in bleedthrough-optimized multiplexed imaging, each channel discards a large fraction of emitted photons, whereas MicroSplit can reclaim much of this loss by aggregating photons from multiple fluorophores in a single measurement. By enabling gentler imaging due to denoising: Our method, next to performing semantic unmixing, also performs unsupervised denoising. Since denoising improves the signal-to-noise (SNR) of images it is applied to, microscopists can acquire the raw data more gentle, accepting a lower initial SNR [3]. To illustrate this on a concrete example, we assessed the similarity of a biological structure imaged at various exposure times between 2ms and 20ms with very high SNR data of the same regions of interest acquired at 500ms exposure time. More concretely, we have conducted these experiments on the 3 channel data (Nucleus, Microtubules, Kinetocore) of the HT-LIF24 dataset. As shown in Supplementary Figure S6 (bottom), even the denoised 2ms exposure micrographs lead to a considerably higher quality w.r.t. the 500ms images than even the 5ms raw data, for Channel 1 even compared to the 20ms raw acquisitions, suggesting at least a 3 to 10-fold reduction of the required photon budget and therefore enabling users of MicroSplit to image considerably more gentle to reach the same quality required for downstream-processing. D Details on Uncertainty Quantification and Calibration In this section, we detail our uncertainty estimation and calibration module. Our approach builds on the methodology of denoiSplit [7], originally inspired by [15], with a modification to the calibration process as described below. Our variational approach is inherently capable of sampling, meaning it can generate slightly different outputs for the same input each time a prediction is made. We leverage this property to produce multiple predictions for a single input, resulting in several predicted values per pixel. Using these, we calculate the standard deviation for each pixel. To ensure these pixel-wise standard deviations correspond closely to the actual prediction errors, we apply a straightforward linear scaling. Opting for 25 Sample preparation Human BJ fibroblast cells were cultured in high-glucose DMEM (Life Technologies, 10569) supplemented with 10% fetal bovine serum (Life Technologies, 26140) and penicillin-streptomycin. Cells were maintained in a humidified incubator at 37 degrees Celsius with 5% carbon dioxide. Prior to imaging, live cells were washed with PBS (Life Technologies, 15140) and trypsinized using 2.5 mL of 0.05% trypsin (Life Technologies, 25300) at 70–80% confluence. The cells were then transferred to 35 mm glass-bottom dishes (MatTek, P35G-1.5-14-C) for microscopy. The staining solution was prepared by diluting Tubulin Tracker Deep Red (Thermo Fisher Scientific, T34077) to 1x, CellMask Orange Actin Tracking Stain (Thermo Fisher Scientific, A57247) to 2x, MitoTracker Green FM (Thermo Fisher Scientific, M7514) to 100 nM, and Hoechst 34580 (Thermo Fisher Scientific, H21486) to 5µg/mL in growth media. For imaging, cells were incubated with 1mL of the staining solution for 30 minutes at 37 degree Celsius, rinsed five times with FluoroBrite DMEM (Thermo Fisher Scientific, A1896701), and subsequently imaged and analyzed in FluoroBrite DMEM. Multi-color structured illumination microscopy (SIM) super-resolution imaging A custom-built structured illumination microscope (SIM) was used for multi-color imaging1. Excitation wavelengths of 642nm (Spectra-Physics, Excelsior One 642), 532nm (Spectra-Physics, Millennia V), 488nm (Spectra-Physics, Excelsior One 488), and 405nm (Spectra-Physics, Excelsior One 405) were employed to excite Tubulin Tracker, CellMask Orange, MitoTracker, and Hoechst, respectively. The four laser beams were combined using three dichroic mirrors (Semrock, Di03-R405-t1; Semrock, Di03-R488-t1; Thorlabs, DMSP550R) and subsequently expanded by 4x. The combined beams were first diffracted into monochromatic beams by a blazed grating (Thorlabs, GR13-0605) and then recombined at the plane of a digital micromirror device (DMD; Texas Instruments, DLP9000X VIS WQXGA). The DMD-generated patterns were projected onto the sample plane through an objective lens (Nikon, SR Plan Apo, 60x, 1.27 WI). A multi-band dichroic mirror (Semrock, Di01-R405/488/532/635-25x36) and an emission filter (Semrock, FF01-446/510/581/703-25) were used to separate excitation and emission light. Additionally, a 2×beam expander was employed to further magnify the emission signal, resulting in a total magnification of 120×. Fluorescence images were captured using an sCMOS detector (Photometrics, Kinetix). For SIM super-resolution imaging, striped binary patterns with a second-order spatial frequency of 2.86µm−1at the sample plane were displayed on the DMD. Patterns at three different angles with six phase shifts each were sequentially projected, with an exposure time of 100 ms per pattern. Super-resolution SIM reconstruction was performed using FairSIM 2, an ImageJ-based open-source software. Each raw image stack of 2048 ×2048 ×18 was reconstructed into 4096 ×4096 super-resolution images, where each pixel corresponds to 27 nm. F.7 The HHMI-D25 dataset Animal Experiments: Heterozygous PhAMexcised female mice carrying the Mito-Dendra2 transgene were generated through inhouse breeding: first, PhAMexcised heterozygous males and females (strain #018397, Jackson Laboratories) were crossed to derive homozygous males, which were then bred with wild-type C57BL/6J females. Mice were housed in sound-attenuated, temperatureand humidity-controlled rooms under a 12-hour light/dark cycle, with food and water provided ad libitum. All procedures were conducted in accordance with NIH guidelines and were approved by the Institutional Animal Care and Use Committee (IACUC) at the Janelia Research Campus (Protocol #22-0229.04). Livers were collected via cardiac perfusion: first with 1×PBS to remove blood, followed by 30 mL of 4% paraformaldehyde (PFA) at a flow rate of 2.5 mL/min to minimize endothelial damage. Tissues were post-fixed in 4% PFA for 24 hours, rinsed three times in 1×PBS, and stored in 1% PFA until further processing. Immunostaining and Image Acquisition: For imaging, samples were embedded in 4.6% low-melting-point agarose and sectioned into 120 µm slices using a Leica VT 1200S vibratome. Sections were blocked in 10% fetal bovine serum with 0.5% Triton X-100 for 1 hour and incubated with primary antibodies at 4◦C for 48 hours. Mouse anti-PMP70 (MilliporeSigma, SAB4200181, 1:75) and rabbit anti-LAMP1 (Abcam, AB208943, 1:50) were used to label peroxisomes and lysosomes, respectively. Secondary antibodies — Alexa Fluor 647 goat anti-mouse (Thermo Fisher, A21235, 1:500) and Alexa Fluor 750 goat anti-rabbit (Thermo Fisher, A21039, 1:500) — were applied overnight at 32 4◦C. Additional markers included Alexa Fluor Plus 555 Phalloidin (Thermo Fisher, A30106, 1:100) and HCS LipidTox Red (Thermo Fisher, H34476, 1:100) to label actin and lipid droplets, respectively. Nuclei were stained with DAPI (Thermo Fisher, D3571, 1:1000 dilution of 1 mg/mL stock) during a 25-minute PBS wash the following day. Sections were cleared using EasyIndex (LifeCanvas Technologies, EI-500-1.52) by incubating samples first in 50% EasyIndex for 1 hour, followed by 100% EasyIndex for 3–5 hours, and mounted using Secure-Seal spacers (Thermo Fisher, 0523073). Imaging was performed on a Leica Stellaris 8 confocal microscope using a 63×/1.40 NA oil immersion objective. Images were acquired at 2048 ×2048 pixels using bidirectional scanning, 2×optical zoom, and a pinhole size of 0.5 Airy units (AU). A total depth of 10µm was captured across 50 z-sections. Fluorophores were imaged using two acquisition lines: the first included mitochondria, actin, and lysosomes, while the second included nuclei, lipid droplets, and peroxisomes. G Further Details on all Experiments (Learning Tasks) In this section, we mention specific details about individual tasks mentioned in the Table 1. Unless explicitly specified, we use the same hyperparameters for all tasks. Please refer to the code for details on the hyperparameters. G.1 Two-channel semantic unmixing tasks Task I It is created from the HT-H24 dataset. It works on 3D z-stacks. The acquisition mode used in this dataset was Training Mode I . We show one qualitative example in Supplementary Figure S18. Task II It is created from HT-P23A dataset. It works on 3D z-stacks. The acquisition mode used in this dataset was Training Mode I . We show one qualitative example in Supplementary Figure S19. Task III It is created from the HT-P23B dataset. It works on 3D z-stacks. The acquisition mode used in this dataset was Training Mode I . We show one qualitative example in Supplementary Figure S20. Task IV-XII These tasks are created from the Pavia-P24 dataset. The acquisition mode used in them was Training Mode III-a. These tasks work on 2D frames. There are nine different acquisitions within this dataset that differ in (a) the laser power distribution among the two target channels and (b) the overall SNR. SNR has three levels, namely low, mid, and high, with low having the lowest SNR and high having the highest SNR level. The SNR level fixes the total combined laser power used for both channels. Within a single SNR level, we distribute the power among the two channels in three ways, namely, 50/50,66/33, and 84/16 denoting the percentage laser power allocation to the respective channel. So, 84/16 means 84% of the total laser power was allocated to channel 1 and 16% was allocated to channel 2. We show one qualitative example per task in Supplementary Figures S27,S29,S28,S30,S31,S32,S33,S34,S35. Task XIII This task was created from the HT-T24 dataset. It works on 2D frames. The acquisition mode used in them was Training Mode III-a. We show one qualitative example in Supplementary Figure S36. Task XIV This task was created from the HT-LIF dataset. It used (Nucleus) and Microtubules as the two channels. It works on 2D frames. The acquisition mode used in them was Training Mode III-a. This task worked with the sub-dataset that had the exposure duration of 5ms. We show one qualitative example in Supplementary Figure S37. 33 Task XV-XX These tasks were created from the Chicago-Sch23 dataset. The Chicago-Sch23 dataset has four structures, and these tasks capture all possible 2-channel semantic unmixing tasks from these four structures. These are Structured Illumination Microscopy (SIM) images, and due to the computational post-processing that happens in SIM, the resulting images have very different noise characteristics. Therefore, we disabled the use of noise models for all tasks generated from the Chicago-Sch23 dataset. More specifically, in Equation 2, we set w= 0. We show one qualitative example per task in Supplementary Figures S21,S22,S23,S24,S25,S26. G.2 Three-channel semantic unmixing tasks Task XXI This task is a 3-channel semantic unmixing task generated from the CBG-Z18 dataset. This is a 3D dataset and so, we employ the 3D version of MicroSplit. We show one qualitative example in Supplementary Figure S16. Task XXII This task is a 3-channel semantic unmixing task generated from the CBG-N18 dataset. This is a 3D dataset and so, we employ the 3D version of MicroSplit. We show one qualitative example in Supplementary Figure S17. Task XXIII This task is a 3-channel 3D semantic unmixing task generated from HHMI-D25 dataset. Specifically, mitochondria, lysosomes and nuclei channels from HHMI-D258bit sub-dataset is used to create this task. This is a 3D dataset and so, we employ 3D version of MicroSplit. We show one qualitative example in Supplementary Figure S57. Tasks XXIV-XXVIII These tasks are generated from the HT-LIF24 dataset which has four structures. The three structures picked for these tasks are Nucleus, Microtubules and Kinetocore. While these tasks aim at splitting apart the above-mentioned three structures, they differ in the exposure duration of the training and evaluation dataset. The exposure duration used is mentioned in Task Details column of Table 1. We show one qualitative example per task in Supplementary Figures S13,S14,S15,S11,S12. Task XXXI, XXXII These two tasks are 3-channel 3D semantic unmixing tasks generated from the denoised version of HHMID258bit dataset, with denoising done by Noise2Void [4]. Mitochondria, lysosomes and nuclei channels from HHMI-D258bit sub-dataset are used. This is a 3D dataset and so, we employ 3D version of MicroSplit. While the task XXXI works directly on the above-mentioned data whereas task XXXII adds Gaussian and Poisson noise on top of the individual channels. Please refer to Supplementary sub-section B.1.1 on how the noise was added. On a technical node, since the task XXXI directly works on the denoised data, there is no rationale to use noise models on such data. Hence the denoiSplit loss component is completely disabled (w= 0 in Equation 1) and the Task XXXI is effectively trained with µSplit configuration. We show one qualitative example for each task in Supplementary Figure S61. Task XXXIII-XXXV These three tasks are 3-channel 3D semantic unmixing tasks generated from HHMI-D2516bit dataset. Mitochondria, lysosomes and nuclei channels from HHMI-D2516bit sub-dataset are used. This is a 3D dataset and so, we employ 3D version of MicroSplit. While the task XXXIII works directly on the above-mentioned data whereas tasks XXXIV and XXXV adds Gaussian and Poisson noise on top of the individual channels. Please refer to Supplementary sub-section B.1.1 on how the noise was added. We show one qualitative example for each task in Supplementary Figure S62. Task XXXVI This task is a 3-channel 3D semantic unmixing task generated from HHMI-D2516bit,0.25 dataset. Mitochondria, lysosomes and nuclei channels from HHMI-D2516bit,0.25 sub-dataset are used. This is a 3D dataset and 34 Training Mode Correlation between structures preserved? Input data is. . . Note Training Mode I Yes . . . mixed computationally. Existing imaging data can be used (no special acquisitions required). Training Mode II No . . . mixed computationally. Existing imaging data can be used (no special acquisitions required). Results might be of lesser quality if reliable structural correlations between channels to be unmixed do exist. Training Mode III Yes . . . imaged directly. Mixed input must directly be imaged along with the individual target channels. Can lead to best performance when imaging is done well, but requires additional microscopy work. Table ST11 Overview of Training Modes: Properties and key advantages/ disadvantages of the training modes we propose. We say that the correlation between structures is preserved when their relative positioning in the computationally mixed input maintains their biologically occurring relative positioning in the sample. so, we employ 3D version of MicroSplit. We show one qualitative example for the task in Supplementary Figure S59. G.3 Four-channel semantic unmixing tasks Task XXIX This task is generated from the HT-LIF24 dataset and uses all four structures to create this task. The exposure duration used for this task is 5ms. We show one qualitative example per task in Figure 1f. Task XXX This task is generated from the Chicago-Sch23 dataset and uses all four structures to create this task. Due to the same reason as described for Tasks XV-XX, we have disabled the Noise model for this task as well. We show one qualitative example in Supplementary Figure S10. 35 Task Idx Dataset Task Details 2D/3D PSNR MicroMS-SSIM C1 C2 C1 C2 I HT-H24 - 3D 0.09 0.55 0.973 0.956 II HT-P23A - 3D 0.53 0.45 0.016 0.006 III HT-P23B - 3D 0.29 0.39 0.005 0.019 IV Pavia-P24 high, 50/50 2D 0.17 1.1 0.010 0.009 V Pavia-P24 high, 66/33 2D 1.83 0.19 0.045 0.002 VI Pavia-P24 high, 84/16 2D 0.35 1.44 0.016 0.049 VII Pavia-P24 mid, 50/50 2D 0.26 0.55 0.005 0.013 VIII Pavia-P24 mid, 66/33 2D 0.25 0.09 0.009 0.003 IX Pavia-P24 mid,84/16 2D 0.02 0.42 0.005 0.006 X Pavia-P24 low, 50/50 2D 0.41 0.42 0.010 0.012 XI Pavia-P24 low, 66/33 2D 0.02 0.28 0.006 0.006 XII Pavia-P24 low, 84/16 2D 0.72 0.76 0.025 0.018 XIII HT-T24 - 2D 0.66 0.52 0.005 0.004 XIV HT-LIF24 - 2D 0.66 1.21 0.003 0.004 XV Chicago-Sch23 C0 vs C1 2D 1.80 0.92 0.003 0.001 XVI Chicago-Sch23 C0 vs C2 2D 1.68 1.32 0.005 0.002 XVII Chicago-Sch23 C0 vs C3 2D 1.20 0.54 0.000 0.000 XVIII Chicago-Sch23 C1 vs C2 2D 0.89 1.46 0.001 0.001 XIX Chicago-Sch23 C1 vs C3 2D 0.99 0.81 0.000 0.001 XX Chicago-Sch23 C2 vs C3 2D 1.01 0.82 0.000 0.000 Task Idx Dataset Task Details 2D/3D PSNR MicroMS-SSIM C1 C2 C3 C1 C2 C3 XXI CBG-Z18 - 3D 0.12 0.18 0.12 0.002 0.001 0.001 XXII CBG-N18 - 3D 0.18 0.21 0.06 0.000 0.000 0.000 XXIII HHMI-D258bit - 3D 0.018 0.035 0.027 0.000 0.002 0.001 XXIV HT-LIF24 2ms 2D 1.19 0.90 0.66 0.007 0.003 0.004 XXV HT-LIF24 3ms 2D 1.30 0.96 0.75 0.008 0.003 0.004 XXVI HT-LIF24 5ms 2D 1.11 0.89 0.53 0.004 0.002 0.004 XXVII HT-LIF24 20ms 2D 0.85 0.62 0.73 0.002 0.001 0.002 XXVIII HT-LIF24 500ms 2D 0.61 0.35 0.66 0.001 0.001 0.001 XXXI HHMI-D258bit,denoised - 3D 0.073 0.11 0.151 0.000 0.001 0.001 XXXII HHMI-D258bit,denoised σ= 20, λ = 30 3D 0.076 0.094 0.153 0.000 0.003 0.002 XXXIII HHMI-D2516bit - 3D 0.03 0.126 0.016 0.002 0.003 0.001 XXXIV HHMI-D2516bit σ= 2K, λ = 5K3D 0.031 0.13 0.014 0.002 0.003 0.001 XXXV HHMI-D2516bit σ= 4K, λ = 10K3D 0.033 0.133 0.016 0.002 0.004 0.001 XXXVI HHMI-D2516bit,0.25 - 3D 0.082 0.103 0.021 0.000 0.001 0.000 Task Idx Dataset 2D/3D PSNR MicroMS-SSIM C1 C2 C3 C4 C1 C2 C3 C4 XXIX HT-LIF24 2D 0.78 1.05 0.86 0.53 0.004 0.004 0.002 0.006 XXX Chicago-Sch23 2D 1.57 0.87 1.24 0.65 0.006 0.002 0.003 0.008 Table ST12 Standard error values for all entries of Table 1. 36 Fig. S2 Semantic unmixing of similar structures with MicroSplit. We show how the same structure in two target channels can still be unmixed, even if the only structural difference is a controllable scaling factor. See Supplementary Section B.3 for a detailed description of the conducted experiments. 37 Fig. S3 Semantic unmixing of similar structures with MicroSplit. Here we plot Validation PSNR vs.training steps for all experiments shown in Figure S2. Observe that, as the magnification factor gets closer to 1, it takes longer for the network to initiate the learning, and the quality of the splitting converges to an overall lower level. See Supplementary Section B.3 for a detailed description of the conducted experiments. 38 Fig. S4 Comparison with PICASSO on a 3-channel unmixing task. Please note that MicroSplit distinguishes itself from PICASSO [20] by only requiring a single superimposed image as input, rather than multiple images with different spectral overlaps. Here, we used the high-SNR data from the HT-LIF24 dataset, prepared inputs that are suitable for PICASSO, and added additional (synthetic) noise to demonstrate how different both systems deal with noisy data. In row 1, we show example images from the the three input channels used for PICASSO. In rows 2 and 3, we show prediction by PICASSO and MicroSplit, respectively. In the last row, we show the high-SNR ground truth for visual evaluation. We observe that, in the presence of noise, Picasso starts to be challenged to unmix the data, mostly if it starts having similar features (see columns 2 and 3, where PICASSO removes data around the nucleus region seen in column 1. 39 Fig. S5 Quantitative evaluations of the effect of signal-to-noise ratios (SNR) on in-distribution and out-ofdistribution unmixing results (using the HT-LIF24 dataset). We trained MicroSplit models on micrographs acquired using a range of different exposure times. Note that the underlying sample ROIs remain identical. We then used all trained models (for 2ms, 3ms, 5ms, 20ms and 500ms exposure time data, indicated in the legend) and evaluated their semantic unmixing performance on all exposure times (x-axis), respectively. The rows show unmixed results quantified using MicroSSIM, MicroMS-SSIM, CARE-PSNR, SSIM and MS-SSIM metrics. All plots utilize the legend presented in the plot in the first row, first column. Note that the plots in this figure also demonstrate that the MS-SSIM and SSIM metrics do not work well on microscopy data (while MicroSSIM and MicroMS-SSIM show better sensitivity [8]). 40 400 500 600 700 Wavelength (nm) 0.00 0.01 0.02 Optimized for photon collection (PC) Photons wrongly assigned: 22% DAPI FITC TRITC 400 500 600 700 Wavelength (nm) 0.00 0.01 0.02 Balanced (BT & PC) Photons wrongly assigned + lost: 34% DAPI FITC TRITC 400 500 600 700 Wavelength (nm) 0.00 0.01 0.02 Optimized for Bleed-through (BT) Photons wrongly assigned + lost: 55% DAPI FITC TRITC GT Noisy vs GT High-SNR Prediction vs GT High-SNR Acq. PSNR MicroMS-SSIM PSNR MicroMS-SSIM Duration C1 C2 C3 C1 C2 C3 C1 C2 C3 C1 C2 C3 2ms 23.3 25.1 26.1 0.839 0.772 0.869 31.0 32.2 36.3 0.940 0.973 0.944 3ms 23.6 26.0 27.2 0.842 0.780 0.871 30.8 32.2 36.1 0.940 0.973 0.948 5ms 24.6 28.3 29.9 0.857 0.817 0.875 32.9 34.3 37.6 0.960 0.983 0.963 20ms 30.0 35.6 38.16 0.920 0.942 0.914 37.0 39.7 41.4 0.984 0.994 0.989 Fig. S6 Photon efficient imaging with MicroSplit: a concrete example. We quantify two mechanisms through which MicroSplit reduces the required photon budget: (i)Reduced photon filtering (top):Emission spectra of the three fluorophores DAPI, FITC, and TRITC are shown together with example emission filter bands used in conventional multi-color imaging. From left to right, we illustrate three settings that increasingly trade off bleed-through (BT) against photon efficiency: a configuration that strongly suppresses BT, a balanced configuration that collects more photons at the cost of some BT, and a highly permissive configuration that captures nearly all emitted photons but suffers notable spectral overlap. MicroSplit, by contrast, allows imaging multiple structures within a single broad emission band and subsequently reassigning the collected intensities to their respective output channels leading to the indicated photon-efficiency increases. (ii)Denoising (bottom): Denoising enables repurposing of the available photon budget by acquiring lower-SNR micrographs, which MicroSplit can restore to high-SNR predictions. We compare raw data acquired at 2ms, 3ms, 5ms, and 20ms exposure times with high-SNR (500ms) reference images of the same regions in the HT-LIF24 dataset (three channels: Nucleus, Microtubules, Kinetochore). The first six data columns in the table quantify the similarity of low-exposure raw data to the 500ms reference, while the rightmost six columns show the corresponding similarities for MicroSplit predictions (identical data as in Table1). In all cases, the MicroSplit predictions exhibit higher quality than the corresponding raw inputs. Strikingly, predictions from 2ms exposures already surpass the quality of 5ms raw data for all channels, and for Channel 1 even exceed the 20ms raw data. In this example, this corresponds to at least a three-fold reduction in required photon budget per acquisition, and likely closer to an order of magnitude when averaged across the three channels. 41 Fig. S21 Qualitative Evaluation for Task XV from Chicago-Sch23 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. 48 Fig. S22 Qualitative Evaluation for Task XVI from Chicago-Sch23 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. 49 Fig. S23 Qualitative Evaluation for Task XVII from Chicago-Sch23 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. 50 Fig. S24 Qualitative Evaluation for Task XVIII from Chicago-Sch23 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. 51 Fig. S25 Qualitative Evaluation for Task XIX from Chicago-Sch23 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. 52 Fig. S26 Qualitative Evaluation for Task XX from Chicago-Sch23 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. Fig. S27 Qualitative Evaluation for Task IV from Pavia-P24 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. 53 Fig. S28 Qualitative Evaluation for Task VI from Pavia-P24 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. 54 Fig. S29 Qualitative Evaluation for Task V from Pavia-P24 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. Fig. S30 Qualitative Evaluation for Task VII from Pavia-P24 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. 55 Fig. S31 Qualitative Evaluation for Task VIII from Pavia-P24 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. Fig. S32 Qualitative Evaluation for Task IX from Pavia-P24 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. 56 Fig. S33 Qualitative Evaluation for Task X from Pavia-P24 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. Fig. S34 Qualitative Evaluation for Task XI from Pavia-P24 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. 57 Fig. S54 Calibration plot for Task I from Dataset HT-H24 Fig. S55 Calibration plot for Task XXI from Dataset CBZ-Z18 Fig. S56 Calibration plot for Task XXII from Dataset CBZ-N18 64 Fig. S57 Qualitative Evaluation for Task XXIII from HHMI-D258bit dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. Also note that the predictions for channel 3 are of rather poor quality and that you can find a description of how this problem was solved in the Supplementary Section B. Fig. S58 Qualitative Evaluation for Task XXXIII from HHMI-D2516bit dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. 65 Fig. S59 Qualitative Evaluation for Task XXXVI from HHMI-D2516bit,0.25 dataset. Note that we show the target and the prediction corresponding to the input crop which is denoted in Input panel by a white dotted rectangle. Fig. S60 Effect of SNR on Model Performance (HT-LIF24 Dataset): We evaluate how SNR influences model performance using the HT-LIF24 dataset. Different models are trained on data subsets acquired with varying exposure durations—leading to different SNRs—and their predictions are compared over a common region of interest. High-frequency details in the predictions (especially the third channel) are visibly reduced when the input SNR is lower. 66 Fig. S61 Effect of SNR on Model Performance (HHMI-D258bit Dataset): We assess the impact of SNR on a subset of the HHMI-D25 dataset. Comparing the predictions (row 2, results of Task XXIII) with the ground truth (row 1), we observe that the prediction quality, particularly for the third channel, is not good, with entire parts of the structures being put into the other channels. We then train MicroSplit using a Noise2Void [4] denoised version of the same data, leading to much improved semantic unmixing performance (row 3, Task XXXI). Finally, we re-introduced synthetic Gaussian and Poisson noise to the denoised HHMI-D25 data used in Task XXXI and retrained MicroSplit on this lower-SNR data (row 4, Task XXXII). As it was likely to be expected, this does again drop the semantic unmixing performance. 67 Fig. S62 Effect of SNR on Model Performance (HHMI-D2516bit Dataset): We assess the impact of SNR on the HHMI-D2516bit dataset. For this part of the HHMI-D25 data, the predictions (row 2, Task XXXIII) are visually more close to the target images compared to results on HHMI-D258bit dataset (Figure S61, row 2, Task XXIII). We then introduce two levels of additional Gaussian and Poisson noise to the HHMI-D2516bit data to reduce SNR and retrain MicroSplit on those noisier versions of the data (row 3 and 4, Tasks XXXIV and XXXV, respectively). As it was likely to be expected, the reduced SNR leads to a noticeable decline in the semantic unmixing performance. 68