Accepted on the jury’s recommendation for the award of the degree of Docteur ès Sciences (PhD) by SPAD Image Sensors with Embedded Intelligence Yang LIN Thesis n° 11 448 2025 Presented on 11th November 2025 Prof. C. Petersen, jury president Prof. E. Charbon, Dr C. Bruschini, thesis directors Prof. L. Ratti, examiner Dr D. Li, examiner Prof. G. De Micheli, examiner School of Engineering Advanced Quantum Architecture Lab Doctoral program in Computational and Quantitative Biology
Acknowledgements It is never easy to align one’s interests with one’s work, and I feel very fortunate to have found that alignment throughout my Ph.D. journey. This work has been both intellectually rewarding and personally fulfilling, and it would not have been possible without the guidance, collaboration, and support of many people. First and foremost, I would like to express my deepest gratitude to my advisor, Prof. Edoardo Charbon, and co-advisor, Dr. Claudio Bruschini, for their unwavering guidance, insightful advice, and constant encouragement. Their vision and mentorship have profoundly shaped the way I approach research and problem-solving. I am also grateful to my colleagues in the AQUA lab for the stimulating discussions, teamwork, and all the memorable moments we shared outside of work — especially the barbecues, which brought laughter and friendship to balance the long research days. In particular, I am grateful to Paul, Andrei, Andrada, Milo, Tommaso, Baris, and Kerim for their professional help and insightful discussions that greatly contributed to this work. I would also like to thank my officemates, Yating and Jad, for creating such a friendly and motivating environment during countless days in the office. Finally, I owe my deepest thanks to my family and friends for their unconditional love, understanding, and encouragement. Their belief in me has been my greatest source of strength and motivation throughout this journey. To all of you — thank you for being part of this experience. Neuchâtel, October 30, 2025 Yang Lin i
Abstract Single-photon avalanche diodes (SPADs) are solid-state photodetectors that can detect individual photons with picosecond timing precision, enabling powerful time-resolved imaging across scientific, industrial, and biomedical applications. Despite their unique sensitivity, conventional SPAD imaging workflows passively collect photons, transfer large volumes of raw data off-chip, and reconstruct results through offline post-processing, leading to inefficiencies in photon usage, high latency, and limited adaptability. This thesis explores the potential of embedded artificial intelligence (AI) for efficient, real-time, intelligent processing in SPAD imaging through hardware-software co-design, bringing computation directly to the sensor to process photon data in its native form. Two general frameworks are proposed, each representing a paradigm shift from the conventional process. The first framework is inspired by the power of artificial neural networks (ANNs) in computer vision. It employs recurrent neural networks (RNNs) that operate directly on timestamps of photon arrival, extracting temporal information in an event-driven manner. The RNN is trained and evaluated for fluorescence lifetime estimation, achieving high precision and robustness. Quantization and approximation techniques are explored to enable FPGA implementation. Based on this, an imaging system integrating a SPAD image sensor with an on-FPGA RNN is developed, enabling real-time fluorescence lifetime imaging and demonstrating generalizability to other time-resolved tasks. The second framework is inspired by the human visual system, employing spiking neural networks (SNNs) that operate directly on the asynchronous pulses generated by SPAD avalanche breakdown upon photon arrival, thereby enabling temporal analysis with ultra-low latency and energy-efficient computation. Two hardware-friendly SNN architectures, Transporter SNN and Reversed start-stop SNN are proposed, which transform the phase-coded spike trains into density-coded and inter-spike-interval-coded representations, enabling more efficient training and processing. Dedicated training methods are explored, and both architectures are validated through fluorescence lifetime imaging. Based on the Transporter SNN architecture, the first SPAD image sensor with on-chip spike encoder for active time-resolved imaging is developed. This thesis encompasses a full-stack imaging workflow, spanning SPAD image sensor design, FPGA implementation, software development, neural network training and evaluation, mathematical modeling, fluorescence lifetime imaging, and optical system setup. Together, these contributions establish new paradigms of intelligent SPAD imaging, where sensing and computation are deeply integrated. The proposed frameworks demonstrate significant gains in photon efficiency, processing speed, robustness, and adaptability, illustrating how embedded AI can transform SPAD systems from passive detectors into intelligent, iii
Abstract adaptive, and autonomous imaging platforms for next-generation applications. Keywords single-photon avalanche diode (SPAD), image sensor, artificial intelligence (AI), embedded AI, deep learning, neural network, spiking neural network (SNN), fluorescence lifetime imaging microscopy (FLIM), time-resolved imaging iv
Résumé Les diodes à avalanche à photon unique (SPAD) sont des photodétecteurs à semi-conducteurs capables de détecter des photons individuels avec une précision temporelle de l’ordre de la picoseconde, ouvrant la voie à une imagerie temporellement résolue de pointe dans des domaines scientifiques, industriels et biomédicaux. Malgré leur sensibilité exceptionnelle, les flux de travail conventionnels en imagerie SPAD collectent passivement les photons, transfèrent de grands volumes de données brutes hors du capteur, puis reconstruisent les résultats par post-traitement hors ligne, entraînant une utilisation inefficace des photons, une forte latence et une adaptabilité limitée. Cette thèse explore le potentiel de l’intelligence artificielle embarquée (embedded AI) pour un traitement efficace, intelligent et en temps réel en imagerie SPAD à travers une co-conception matériel-logiciel, en rapprochant le calcul directement du capteur afin de traiter les données de photons sous leur forme native. Deux cadres généraux sont proposés, chacun représentant un changement de paradigme par rapport aux approches conventionnelles. Le premier cadre s’inspire de la puissance des réseaux de neurones artificiels (ANN) en vision par ordinateur. Il exploite des réseaux de neurones récurrents (RNN) qui opèrent directement sur les horodatages d’arrivée des photons, extrayant l’information temporelle de manière événementielle. Le RNN est entraîné et évalué pour l’estimation des durées de vie de fluorescence, atteignant une grande précision et une robustesse élevée. Des techniques de quantification et d’approximation sont étudiées afin de permettre une implémentation sur FPGA. Sur cette base, un système d’imagerie intégrant un capteur SPAD et un RNN embarqué sur FPGA est développé, permettant une imagerie de fluorescence en temps réel et démontrant sa généralisabilité à d’autres tâches temporellement résolues. Le second cadre s’inspire du système visuel humain, en employant des réseaux de neurones impulsionnels (SNN) qui opèrent directement sur les impulsions asynchrones générées par l’avalanche des SPAD lors de l’arrivée d’un photon, permettant ainsi une analyse temporelle à très faible latence et à consommation énergétique réduite. Deux architectures SNN adaptées au matériel, transporter SNN et Reversed Start-Stop SNN, sont proposées. Elles transforment les trains de spikes codés en phase en représentations codées en densité et en intervalles inter-spikes, ce qui permet un entraînement et un traitement plus efficaces. Des méthodes d’entraînement dédiées sont explorées, et les deux architectures sont validées par imagerie de fluorescence. Sur la base de l’architecture Transporter SNN, le premier capteur SPAD intégrant un encodeur de spikes sur puce pour l’imagerie active temporellement résolue est développé. Cette thèse couvre un flux de travail complet en imagerie, englobant la conception de capteurs SPAD, l’implémentation FPGA, le développement logiciel, l’entraînement et l’évaluation de v
Résumé réseaux de neurones, la modélisation mathématique, l’imagerie de fluorescence et la mise en place du système optique. Ensemble, ces contributions établissent de nouveaux paradigmes pour l’imagerie SPAD intelligente, où la détection et le calcul sont profondément intégrés. Les cadres proposés démontrent des gains significatifs en efficacité photonique, vitesse de traitement, robustesse et adaptabilité, illustrant comment l’IA embarquée peut transformer les systèmes SPAD, de simples détecteurs passifs en plateformes d’imagerie intelligentes, adaptatives et autonomes pour les applications de nouvelle génération. Mots-clés diode à avalanche à photon unique (SPAD), capteur d’image, intelligence artificielle (IA), embedded AI, apprentissage profond, réseau de neurones, réseau de neurones impulsionnels (SNN), microscopie d’imagerie de durée de vie de fluorescence (FLIM), imagerie résolue en temps vi
Contents Acknowledgements i Abstract iii Résumé v List of Figures ix List of Tables xi List of Abbreviations xiii 1 Introduction 1 1.1 Background and Motivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 1.2 Objectives . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 1.3 Contribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 1.4 Thesis Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2 Fundamentals of Intelligent Time-Resolved Imaging 13 2.1 Single-Photon Avalanche Diode . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.2 Fluorescence Lifetime Imaging . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 2.3 Recurrent Neural Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 2.4 Spiking Neural Networks . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 3 Modeling for SPAD-Based FLIM 35 3.1 Modeling of SPAD . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 3.2 Fluorescence Lifetime Imaging . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 3.3 TCSPC and Time-Gated Time-Resolved Imaging . . . . . . . . . . . . . . . . . . . 42 3.4 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49 4 SPAD Imagers with Recurrent Neural Networks 51 4.1 RNN for Active Time-Resolved Imaging . . . . . . . . . . . . . . . . . . . . . . . . 52 4.2 RNN for Fluorescence Lifetime Estimation . . . . . . . . . . . . . . . . . . . . . . 57 4.3 FPGA Implementation of RNN for Real-Time Processing . . . . . . . . . . . . . . 67 4.4 Real-Time FLIM with RNN-Coupled SPAD Sensor . . . . . . . . . . . . . . . . . . 72 4.5 RNN for NIROT . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73 vii
List of Abbreviations IRF Instrument Response Function. 23–25, 43, 45, 46, 59, 63, 76 ISI Inter-Spike Interval. 32, 82, 90, 91, 102 LiDAR Light Detection and Ranging. 1–3, 7, 33, 77, 80, 90 LIF Leaky Integrate-and-Fire. x, 30, 78, 80, 82, 84, 85 LLM Large Language Model. 26 LS Least Squares. 25, 57, 63–65, 67, 68, 72, 92, 93 LSB Least-Significant Bit. 92 LSTM Long Short-Term Memory. 27, 28, 56, 60, 61, 63–65, 67, 68, 75, 76, 91 MAC Multiply-Accumulate. 56, 57, 67, 73, 78, 84 MAE Mean Absolute Error. 62, 63 MAPE Mean Absolute Percentage Error. 62, 63, 76, 91, 92 MLE Maximum Likelihood Estimation. 23, 57 MOSFET Metal-Oxide-Semiconductor Field-Effect Transistor. 15 NIROT Near-Infrared Optical Tomography. 52 PDF Probability Density Function. 41, 42, 45–47 PDP Photon Detection Probability. 13, 14, 16, 35, 71 PMF Probability Mass Function. 35, 36, 39, 40 PMT Photomultiplier Tube. 1, 13, 14, 16, 23, 39, 57 PQPR Passive Quenching and Passive Recharging. 14, 15, 84, 94–96 R6G Rhodamine 6G. 59, 65, 67, 68 RAM Random Accesss Memory. 56, 57, 71 ReLU Rectified Linear Unit. 32, 80, 91 RGC Retinal Ganglion Cell. 6, 7 RMSE Root Mean Squared Error. 62, 63 RNN Recurrent Neural Network. 6–8, 10, 11, 13, 26, 27, 52–58, 60, 61, 63–65, 67, 73–76, 86, 93, 104 RO Ring Oscillator. 85, 86 RS Reverse Start-Stop. 82, 84, 91–93 SiPM Silicon Photomultiplier. 16 SNN Spiking Neural Network. 4–7, 9–11, 13, 28, 29, 31–33, 51, 52, 77–86, 89–94, 99, 101–103, 105 SNR Signal-to-Noise Ratio. 64 SPAD Single-Photon Avalanche Diode. ix, 1–11, 13–20, 22–24, 28, 33, 35, 39, 42, 44, 49–54, 57, 67, 71–82, 84, 85, 87, 88, 90, 94–98, 100–106 SRM Spike Response Model. 28–30 STDP Spike-Timing Dependent Plasticity. 32, 106 TCSPC Time-Correlated Single-Photon Counting. ix, 8, 10, 18–23, 26, 42, 43, 45–48, 50, 54, 60, 65, 83 xiv
List of Abbreviations TDC Time-to-Digital Converter. 5, 18–20, 52, 71, 72, 74, 87, 89 ToF Time of Flight. 19 xv
1Introduction 1.1 Background and Motivation Vision is the ability to detect light, with which humans are endowed by Nature to perceive and interpret the world. Beyond the limits of the naked eye, humanity has continuously developed new techniques and built advanced devices to extend visual perception: recording the past, sensing the remote, and unveiling the unseen. From the invention of heliography in 1820s to the ubiquitous application of charge-coupled device (CCD) and complementary metal-oxide-semiconductor (CMOS) image sensors today, people have captured their vision in both still and moving images [ 1 ]. However, both human eyes and CCD or CMOS image sensors still struggle in extremely low-light conditions and fast-moving scenarios, where single-photon sensitivity and subnanosecond precision are required. Applications such as fluorescence lifetime imaging (FLIM) and light detection and ranging (LiDAR) demand precise measurement of photon arrival times. These requirements drive the need for highly sensitive and time-resolved imaging solutions. Single-Photon Avalanche Diodes Standard image sensors, including CCD and CMOS sensors, struggle under these conditions due to limited sensitivity and high noise levels [ 2 ]. While intensified CCDs (ICCDs) and electron-multiplying CCDs (EMCCDs) enhance low-light imaging by amplifying weak signals, they still suffer from excess noise, limited photon-counting capability, and high costs [ 3 , 4 ]. Photomultiplier tubes (PMTs) achieve single-photon detection but are bulky, powerhungry, and unscalable, restricting widespread adoption [ 4 ]. In the last decade, single-photon avalanche diodes (SPADs) have emerged as cost-effective single-photon detectors, which offer a compact, efficient, and scalable solution for imaging in photon-starved conditions while enabling precise time-resolved measurements [ 5 , 6 ]. Owing to its successful implementation with CMOS technology, SPADs can be manufactured massively at affordable prices, arranged in arrays for widefield imaging and combined with other integrated circuit (IC) components to realize on-chip processing [7]. 1
Chapter 1. Introduction (a) FLIM (b) LiDAR (c) Quantum ghost imaging (d) Telescope Fig. 1.1: Applications of SPAD image sensors. (a) A wide-field FLIM system with a 0.5 megapixel SPAD array is used for lifetime measurements of Convallaria and HT1080 cells [ 16 ]. (b) A one megapixel SPAD camera is used for LiDAR [ 17 ]. (c) Two SPAD sensors are integrated into the microscope for quantum ghost imaging [ 12 ]. (d) SPADs are integrated into the telescope for various astronomical tasks [18]. SPADs have revolutionized imaging and sensing across various fields, as illustrated in Fig. 1.1. In biomedical imaging, SPADs increasingly play a crucial role in FLIM, where their ability to precisely measure photon arrival times powers the study of molecular interactions, tissue diagnostics, and early disease detection [ 6 , 8 ]. In LiDAR, SPAD-based sensors facilitate highresolution depth mapping and 3D imaging, essential for autonomous vehicles, robotics, and remote sensing [ 9 , 10 ]. Quantum optics also benefit from SPADs, as their single-photon sensitivity supports photon correlation measurements, enabling fundamental physics research and new imaging modalities such as quantum ghost imaging [ 11 , 12 ]. Furthermore, in astronomy, SPADs help detect faint celestial objects by capturing individual photons from distant stars and galaxies [ 13 – 15 ]. Their compact size, CMOS compatibility, and ability to function in extreme low-light environments continue to expand their applications in scientific research and industrial imaging. As SPAD sensors have found multiple roles in scientific research and industry, it is worth noting that SPAD sensors have gradually entered the consumer market in recent years [ 19 ], as shown in Fig. 1.2. In 2013, STMicroelectronics announced SPAD-based proximity sensors that are now in most smartphones [ 20 ]. In 2020, Apple Inc. first introduced SPAD image 2
1.1 Background and Motivation (a) Apple iPad Pro (b) Apple Vision Pro (c) Sony IMX611 (d) Canon MS-500 Fig. 1.2: Recent SPAD-based products. (a) Apple first introduced a SPAD-based LiDAR system into iPad Pro for 3D imaging. (b) Apple placed the LiDAR system at the center of its mixed reality headset, Vision Pro, for 3D mapping along with two TrueDepth cameras. (c) Sony introduced a cost-effective SPAD-based LiDAR system for smartphones. (d) Canon introduced the first SPAD-based, interchangeable-lens camera for low-light imaging. sensors into consumer-grade devices (iPhone Pro and iPad Pro) as part of a LiDAR system for depth imaging [ 21 ], and further extended to the latest mixed reality headset Apple Vision Pro [ 22 ]. In 2023, Sony announced IMX611, an affordable SPAD image sensor for smartphones [ 23 ]. In the same year, Canon Inc. launched MS-500, the world’s first ultra-high-sensitivity interchangeable-lens SPAD sensor camera, aiming at long-range surveillance at night [ 24 ]. These emerging applications usher in opportunities and challenges of developing computer vision systems and algorithms for SPAD image sensors. SPAD’s single-photon sensitivity, however, is both a blessing and a curse. While it enables precise temporal resolution and ultra-low light imaging, it also results in the generation of massive amounts of raw photon event data, especially in time-correlated or histogram-based modalities. Each event typically includes a pixel address and a timestamp, causing the total data volume to scale rapidly with the increase of SPAD array size and temporal resolution. In large arrays or high-speed applications, data rates can reach gigabits per second, overwhelming standard off-chip communication interfaces and leading to severe bandwidth bottlenecks and power inefficiencies. The high-volume data from SPAD image sensors also raises significant challenges for processing, which impedes low-latency and real-time operation, and increases the demands on memory and computational resources needed for efficient data handling and analysis, especially for off-chip processing. Embedded Artificial Intelligence To address these challenges, embedded artificial intelligence (AI), or embedded intelligence, emerges as a key paradigm shift for enabling real-time processing within imaging and sensing devices [ 25 , 26 ]. Broadly situated under the umbrella of edge AI, which refers to the integration of AI at the “edge” of the network, embedded AI denotes the integration of intelligent processing capabilities directly into hardware platforms, allowing data to be processed locally rather 3
Chapter 1. Introduction than offloaded to external computers or even remote servers. More specifically, in the context of image sensors, embedded AI can be broadly divided into near-sensor AI and in-sensor AI. Near-sensor AI refers to deploying intelligence on computing hardware tightly coupled to the sensor, such as field-programmable gate arrays (FPGAs) or image signal processors, while insensor AI denotes the integration of intelligence directly within the sensor itself. The principle of embedded AI has been increasingly applied to image sensors, driving the development of smart image sensors and intelligent vision systems [ 27 – 29 ]. As AI-based architectures, such sensors and systems inherit the advanced functionalities, and by embedding intelligence at or near the sensor, they hold the potential to achieve low latency, real-time operation, improved energy efficiency, and substantial reductions in data bandwidth. Most existing efforts have concentrated on embedding intelligence into conventional CMOS or CCD image sensors, as well as into event-based cameras, whereas extending embedded intelligence to SPAD image sensors presents distinct challenges, particularly in the context of time-resolved imaging. Unlike conventional sensors that produce intensity values or changes in intensity, SPADs generate a pulse for each detected photon, resulting in a fundamentally different data format and a substantially higher data volume. Embedded intelligence for SPAD image sensors can be investigated along two principal directions. The first is motivated by the transformative capabilities of artificial neural networks (ANNs), which have profoundly influenced domains ranging from fundamental research to industrial applications over the past decade [ 30 ]. A central research question, therefore, is how this powerful computational paradigm can be effectively leveraged to achieve a deeper integration of intelligence in proximity to, or directly within, SPAD image sensors. The second direction is inspired by the human visual system, in which light is captured by cones and rods and subsequently processed through hierarchical layers of interconnected neurons that communicate via spiking activity. As spiking neural networks (SNNs) render themselves as more efficient yet powerful alternatives to ANNs, their spike-based computation naturally matches with the pulse outputs of SPADs, enabling direct interfacing for low-latency and efficient imaging. The signal processing of the aforementioned methods is demonstrated in Fig. 1.3. Artificial Neural Networks A phenomenal wave of AI has swept over science, technology, and society in the last decade, driving breakthroughs that continually reshape modern life. Its applications span computer vision, natural language processing, robotics, healthcare, and beyond, giving rise to groundbreaking models such as ChatGPT for human-like language understanding [ 31 , 32 ], AlphaFold for protein structure prediction [33], and embedded AI models in smartphones and cameras that unlock entirely new possibilities. At the core of this sweeping revolution are the ANNs, brain-inspired computational models that embrace data-driven learning. ANNs have been extensively used in computer vision, performing post-processing such as image classification and semantic segmentation. By embedding ANNs near or within the image sensor, these computer vision tasks can be performed with reduced latency and power consumption, while 4
1.1 Background and Motivation (a) CMOS/CCD camera with ANN processing (b) Event camera with SNN processing (c) SPAD image sensor with TDC and ANN processing (d) SPAD image sensor with SNN processing Fig. 1.3: Signal processing from image sensors to neural networks. (a) Photodiodes generate continuous-valued signal, which is processed by the ANN. (b) Photodiodes create continuousvalued signal. Then the temporal contrast is calculated and generates spikes indicating the polarity, which are processing by the SNN. (c) SPADs generate spikes, which are timestamped by TDCs, thus converted to continuous-valued signal. Then it is processed by the ANN. (d) SPADs generate spikes, which are directly processed by the SNN. 5
Chapter 1. Introduction also lowering data bandwidth requirements, enhancing privacy, and enabling continuous real-time analysis. While deep learning has revolutionized image processing pipelines for conventional CCD or CMOS imaging, a significant gap remains between conventional neural network models and the unique characteristics of SPAD-based imaging modalities. Although several deep learning models have been introduced to process photon arrival histograms from SPAD sensors, their computational pipelines often resemble conventional processing frameworks, limiting performance gains. Realizing the full potential of neural networks in SPAD-based imaging demands a holistic co-design across hardware, firmware, and software layers, enabling efficient data handling and model execution. In this context, embedded AI, especially those based on recurrent neural networks (RNNs), emerges as a critical component, facilitating low-latency, energy-efficient inference directly on-device. Embedding intelligence directly on or near the sensor enables real-time processing of photon event data, e.g., compressing, denoising, or even interpreting it before it leaves the chip. Instead of transmitting raw photon streams, neural networks can extract task-relevant features such as fluorescence lifetimes, depth estimates, or object classifications on the fly. This drastically reduces bandwidth requirements and energy consumption while unlocking new capabilities like adaptive sensing and autonomous decision-making. Spiking Neural Networks The human visual system, as shown in Fig. 1.4, is a remarkable source of inspiration for both image sensors and AI [ 34 , 35 ], having already laid the foundation for architectures such as convolutional neural networks (CNNs) [ 30 , 36 ]. Photoreceptors such as rods and cones detect incoming photons and convert them into graded electrical potentials. These signals are processed within the retina by several layers of interconnected neurons, including bipolar cells, horizontal cells, and amacrine cells, before reaching the retinal ganglion cells (RGCs). RGCs are the first neurons in the pathway to generate action potentials. Their spiking activity is driven primarily by changes in intensity, contrast, and motion. These action potentials, often called spikes or pulses, travel along the optic nerve to the lateral geniculate nucleus and are then relayed to the visual cortex for higher-level processing and perception. CMOS/CCD image sensors resemble the role of photoreceptors by recording absolute light intensity at every pixel, whereas the retina instead encodes only spatial and temporal contrast information. Event cameras operate more like human retina, which produces asynchronous spikes when intensity changes [ 29 ]. SPADs are closest to rods in their ability to detect single photons but with far greater temporal precision. Unlike CMOS/CCD image sensors which produces “graded eletrical potentials”, they produce “action potentials” directly. As a result, no additional structures mimicking RGCs are required, enabling a direct connection to the “brain”. This principle inspires the development of an architecture that couples an SNN as the “visual cortex” with a SPAD as the “photoreceptors” [ 37 ]. Unlike conventional ANNs, SNNs incorporate temporal dynamics through discrete spikes, more closely emulating the information 6
1.1 Background and Motivation angiography have demonstrated changes in retinal microvasculature in terms of both reduced perfusion and vessel density. These abnormalities are mainly associated with RNFL thinning (see above) 89 . Previous studies also observed changes in retinal venules, which dilate mainly due to chronic retinal hypoxia 90–92 . However, a similar effect on small vein widening was also observed as a result of an increased concentration of retinal DA 93 . If we were able to understand the retrograde or anterograde origin of the onset of morphological changes, it would be possible to target therapy specifically to these sites, thereby slowing or stopping the degradation of individual cell populations of the retina, LGN, and optic nerve. Changes in the electrophysiology of retinal cells Morphological and biochemical changes in the retina of schizophrenia patients are accompanied by alteration in the electrophysiological response of individual retinal cells to light stimulation. Abnormalities in sensitivity to certain wavelengths, frequencies, and intensities of light during stimulation were recorded by electroretinography (ERG) 65 . These abnormalities manifest as changes in amplitude and a delayed onset of the electrophysiological response to a light stimulus (latency), but also in the structure of the aand b-waves (Fig. 2) 94 . The largest ERG study of psychiatric disorders performed to-date (150 schizophrenia patients, 150 patients with bipolar disorder, and 200 healthy Fig. 2 Schema of ERG signal from retinal cells. An illustration of the retina (left) and a representative ERG comparing HCs and schizophrenia patients (right). In the dark-adapted retina, a light stimulus elicits a presynaptic response from photoreceptor cells, represented by the downward-deflecting a-wave. The subsequent postsynaptic response, mediated largely by bipolar and Müller cells, produces the b-wave. The a-wave amplitude (measured from the baseline to the trough of the a-wave) depends on the intensity of the light stimulus and the integrity of the photoreceptors. The b-wave amplitude (measured from the trough of the a-wave to the peak of the b-wave) depends on the a-wave and the integrity of signal transmission within the retina. Redrawn from Hanjin Deivasse web illustration. Fig. 1 Retinal layers. The composition of the individual layers of the retina in the area of the optic nerve. NFL nerve fiber layer, GCL ganglion cell layer, IPL inner plexiform layer, INL inner nuclear layer, OPL outer plexiform layer, ONL outer nuclear layer, ELM external limiting membrane, IS/OS rod and cone inner and outer segments, RPS retinal pigment epithelium. Redrawn from retinareference.com. P. Adámek et al. 3 Published in partnership with the Schizophrenia International Research Society Schizophrenia (2022) 27 (a) Retina (b) Human visual pathway Fig. 1.4: Human visual system. (a) The retina is primarily composed of three layers: photoreceptors (rods and cones), bipolar cells, and RGCs [ 38 ]. (b) Signals from the retina are transmitted via the optic nerve to the lateral geniculate nucleus, and subsequently relayed to the visual cortex. The image is credited to Miquel Perello Nieto, licensed under CC BY-SA 4.0. processing mechanisms of biological neural systems. Moreover, SNNs are regarded as more energy-efficient than traditional ANNs while holding the promise of achieving comparable performance in a range of tasks. SNNs are particularly well-suited to SPAD-based imaging because they share a fundamental alignment in how they represent and process information. The output of a SPAD, discrete photon detection events, naturally corresponds to the input of spiking neural networks, namely, pulses or spikes. Unlike conventional neural networks that process data in dense, synchronous frames, SNNs operate on discrete spikes, enabling asynchronous and highly efficient processing of sparse photon events. Additionally, SNNs offer inherent energy efficiency through sparse activation and local processing, which is critical for edge deployments in power-constrained environments. The spiking neurons, building blocks of spiking neural networks, can be implemented on hardware with minimal computational and memory resources, compared to conventional artificial neurons. By leveraging temporal coding and biologically inspired dynamics, SNNs can perform tasks like lifetime estimation, depth inference, or event-based classification directly on raw SPAD data, opening the door to fully neuromorphic, real-time imaging systems. SPAD Image Sensors with Embedded Intelligence Embedded intelligence, empowered by both RNNs and SNNs, is envisioned as driving two paradigm shifts in SPAD-based imaging, particularly for active, time-resolved applications such as FLIM and LiDAR. When deployed near or directly within the sensor, these neural architectures can eliminate the need for traditional histogramming and even bypass timeto-digital converters (TDCs), thereby reducing hardware complexity and enhancing system 7
Chapter 2. Fundamentals of Intelligent Time-Resolved Imaging Fig. 2.1: Three-dimensional schematic of a SPAD structure. Green denotes p-type silicon and orange denotes n-type silicon, and gray indicates metal (metal layers, vias, and contacts). The p-n junction is biased above breakdown, with the depletion region sustaining a high electric field ( E ) that enables avalanche multiplication from single-photon absorption. A guard-ring edge-termination structure is included to suppress premature edge breakdown and ensure uniform field distribution. particularly in the infrared regime. Moreover, not all generated electron-hole pairs can trigger an avalanche, further reducing PDP. SPAD has higher DCR than its competitors. Compared to the photodiodes in EMCCD, ICCD, and conventional CMOS/CCD cameras, SPADs have larger depletion volume, which gives a higher probability of thermal generation. Additionally, the high electric field may cause electron tunneling, which further contributes to the dark count. In terms of the timing jitter, SPAD and PMT generally outperform EMCCD and ICCD. Apart from these characteristics, SPADs have the advantage of CMOS-compatibility, allowing large arrays and on-chip integrations of additional circuits. 2.1.1 Pixel Circuit A pixel circuit refers to the circuitry within a pixel that links the photodiode to local transistors for functions such as biasing, storage, quenching, amplification, and readout, and is carefully designed to perform these tasks efficiently while minimizing area usage to maximize the fill factor, defined as the ratio of the photosensitive area to the total pixel area. In the context of SPADs, the essential pixel circuitry consists of the quenching and recharging circuits, which ensure the continuous operation of the detector [ 39 ]. With the possibility of passive and active operation, the pixel circuit of a SPAD is divided into four categories: passive quenching and passive recharging (PQPR), passive quenching and active recharging, active quenching and passive recharging, and active quenching and active recharging (AQAR). Among these four 14
2.1 Single-Photon Avalanche Diode VOP R (a) Vbias VOP (b) Vcascode Vbias VOP (c) Fig. 2.2: Schematic of common PQPR SPAD pixel circuit. (a) Simplest implementation using a resistor R for both quenching and recharging of the SPAD. (b) MOSFET-based implementation where the transistor provides PQPR, which works in the linear regime and acts as a variable resistor through the control of Vbias .(c) Cascode transistor configuration, where the cascode transistor works in the linear regime and further raises the excess bias of the SPAD. It is controllber by the bias voltage Vcascode. quenching and recharging methods, PQPR and AQAR are the most common ones. Passive quenching and passive recharging is the most common pixel circuit due to its simplicity, which translates into a high fill factor. The schematic of common PQPR circuits is shown in Fig. 2.2. In passive quenching, a high-value resistor is connected in series with the SPAD. When an avalanche is triggered by a photon, the current through the resistor causes a voltage drop that brings the SPAD bias below the breakdown voltage, thereby quenching the avalanche. Once the avalanche stops, the resistor slowly recharges the SPAD by restoring the bias voltage above breakdown, which is known as passive recharging. The entire quenching and recovery cycle occurs without active control circuitry, making it simple and compact. However, it introduces a relatively long and paralyzable dead time, which limits the maximum count rate and increases the chance of missing photons during high-flux operation. In practice, the high-value resistor is typically replaced by a transistor operating in the linear region, where its effective resistance can be tuned by adjusting the bias voltage. Active quenching and active recharging is an advanced method used in SPAD circuits to overcome the drawbacks of PQPR, typically achieving fast and precise control over the avalanche detection cycle [ 39 ]. In active quenching, a dedicated circuit continuously monitors the SPAD for the onset of an avalanche. Once an avalanche is detected, the circuit rapidly pulls the SPAD bias below the breakdown voltage, forcibly quenching the avalanche. Following this, the active recharging circuit promptly restores the bias voltage above breakdown, typically through a controlled current or voltage source, allowing the SPAD to resume photon detection with minimal delay. This approach is preferred when high count rates and precise timing are required, as it provides short, controllable, and non-paralyzable dead time [ 40 ]. However, the 15
Chapter 2. Fundamentals of Intelligent Time-Resolved Imaging increased complexity, area, and power consumption of the circuits make this technique more demanding in terms of design and integration, often leading to a smaller fill factor. In addition to the quenching and recharge circuitry, SPAD pixels may incorporate additional functional blocks to support imaging and time-resolved measurements. For instance, in-pixel counters can be integrated for direct photon counting, while gating circuitry enables timegated measurements. While in-pixel preprocessing is always desirable, it remains challenging due to limited pixel area and the inherent trade-off between detection performance and functionality. As a result, TDCs and processing circuits, including spike processors, are typically placed next to the SPAD array rather than within each pixel, although this approach can limit scalability [ 41 , 42 ]. The adoption of 3D stacking technology, to some extent, alleviates this constraint by enabling more efficient integration of additional circuitry. 2.1.2 Characterization Photon Detection Probability Photon detection probability (PDP) is a key performance metric for SPADs, silicon photomultipliers (SiPMs), or PMTs. PDP is the probability that a single incident photon on the detector’s active area will produce a detectable electrical signal. It includes both physical absorption and internal conversion processes [ 2 ]. In some cases, the term photon detection efficiency is used instead, which accounts for the fill factor in addition to the intrinsic detection probability. The peak PDP of SPADs typically ranges from 20% to 60%, and it strongly depends on the wavelength of incident light, temperature, and bias voltage. To measure the PDP, a monochromatic light of a given wavelength and known photon flux is used. The photon flux Φphotons can be calculated from the measured optical power Popt: Φphotons =Popt ·λ h·c(2.1) where λ is the wavelength, h is Planck’s constant, and c is the speed of light. The PDP is then given by: PDP =Ncounts/T Φphotons (2.2) when the photon count Ncounts is recorded over a time interval T. Dark Count Rate Dark count rate (DCR) is the rate at which a photodetector (such as a SPAD) generates false detection events in the absence of incident photons. The dark counts originate from thermal generation of carriers, band-to-band tunneling, trap-assisted tunneling, afterpulsing, and radiation-induced effects, with thermal generation being the most dominant source. A SPAD typically exhibits a DCR of a few to tens of thousands of counts per second (cps) at room 16
2.1 Single-Photon Avalanche Diode temperature [ 5 , 6 ], which increases exponentially with temperature. For a fair comparison between devices of different sizes, the DCR is often normalized to the detector’s active area and expressed in counts per second per square micrometer (cps/µm2). To measure the DCR, the detector must be placed in total darkness and cooled down to the desired temperature. Then the detector is running for a significantly long time T and the generated pulses are counted as Ncounts. The DCR is calculated as: DCR =Ncounts T(2.3) To normalize for pixel area, sometimes DCR is reported as: DCR Density =DCR Active Area (2.4) which helps compare different geometries or technologies fairly. Dead Time Dead time is the period after a detection event during which a SPAD is inactive and unable to detect another photon. The dead time is mainly determined by the quenching and recharging circuits. During the dead time, the avalanche must be quenched either passively or actively, and the bias voltage returns above breakdown. Any photons arriving during this time are missed, introducing a non-linearity in photon counting. Passive quenching typically results in a longer dead time compared to active quenching. This dead time significantly impacts performance under high photon flux, as many photons are missed during this period, thereby limiting the maximum achievable count rate to: fmax =1 τdead (2.5) The impact of dead time on photon counting will be elaborated in Section 3.1.2. Dead time can be measured by constructing a histogram of photon inter-arrival times, which is the time intervals between consecutive photon detections. In this histogram, dead time appears as a gap or sharp drop at short inter-arrival times. The same histogram can also be used to estimate the afterpulsing probability, which will be discussed in the following sections. Afterpulsing Afterpulsing refers to false detection events that occur after a real detection, due to the release of trapped charge carriers from deep-level traps in the SPAD’s depletion region [ 6 , 43 ]. The trapping is mainly caused by defects or impurities in the silicon. These carriers can trigger a secondary avalanche, leading to correlated noise. Notably, afterpulsing deviates from a 17
Chapter 2. Fundamentals of Intelligent Time-Resolved Imaging Poisson process and exhibits correlated behavior. To measure the afterpulsing probability (AP), a histogram similar to that used for dark count measurements is constructed. Afterpulsing shows up as an excess of counts at short delays, ranging from nanoseconds to microseconds, and decays over time. As a result, the histogram typically exhibits a bi-exponential profile, with the tail dominated by dark counts. An exponential fit is usually applied to this tail, and the difference between the histogram and the fitted curve is attributed to afterpulses. Timing Jitter Timing jitter is the uncertainty in the time measurement of a detected photon. In other words, it’s the temporal variation between the actual photon arrival time and the registered detection time. It is often quantized by the standard deviation or the full width at half-maximum (FWHM) of the detection time uncertainty. Timing jitter arises from various sources. Variations in avalanche build-up time contribute to the intrinsic jitter of the SPAD. Additionally, circuit components, such as TDC if used, introduce further jitter at the system level. Timing jitter is measured using a picosecond laser and histogramming techniques. The laser, emitting ultra-short pulses, periodically illuminates the SPAD. Time differences between the laser pulses and photon detections are recorded using TDCs or a TCSPC module. A histogram of the photon arrival times is then constructed, from which the FWHM or standard deviation is calculated to quantify the jitter. Assuming Gaussian noise, the measured timing jitter can be decomposed into several contributing components: σtotal =qσ2 SPAD +σ2 laser +σ2 TDC (2.6) 2.1.3 Operation After photon detection, the generated pulses can be processed in several ways. The simplest approach is photon counting, where the pulses are counted to form intensity images. Each pulse can also be timestamped, a method known as time-correlated single-photon counting. Alternatively, pulses can be counted within predefined time windows, or gates, a technique referred to as time gating. Photon Counting Photon counting is the most fundamental task for an image sensor. Traditional CMOS and CCD image sensors accumulate photo-generated charge over an exposure period, then convert that charge into a voltage and digitize it, an approach that introduces readout noise and limits sensitivity at low light levels. In contrast, SPADs operate in Geiger mode, directly generating a binary signal for each detected photon. The output signal is inherently digital and uniform in 18
2.1 Single-Photon Avalanche Diode amplitude, regardless of the photon’s energy or the gain of the avalanche, thus eliminating the need for analog amplification and readout noise. Notably, SPAD image sensors have been commercialized for extremely low-light imaging [24]. The counting process can be performed either on the chip, FPGA, or computer, depending on system architecture, application requirements, and available resources. On-chip counting is typically limited by area and power constraints, making the implementation of high-bitwidth counters impractical. Moreover, integrating multi-bit counters within each pixel or column often comes at the cost of other functionalities, such as time tagging or gating, due to limited silicon area and routing complexity. As for the on-FPGA counting, SPAD output pulses, either directly or through TDCs, are captured and processed in real time. Furthermore, this approach is often preferred due to its compatibility with other functionalities. Photon counting can also be performed in software on a host computer, especially in research environments or proof-of-concept systems where flexibility and post-processing capability are prioritized over real-time performance. In this architecture, SPAD outputs, typically timestamped by TDCs or TCSPC modules, are streamed to the computer, where post-processing can be performed flexibly in software. Time-Correlated Single-Photon Counting Time-correlated single-photon counting is a highly sensitive technique used to measure the arrival times of single photons with respect to a periodic reference signal [ 44 ]. It is widely used in applications requiring precise temporal information, such as FLIM, direct ToF, and photon correlation studies. In a typical TCSPC system, the sample is excited by a pulsed light source, and a single-photon detector such as a SPAD registers the arrival of individual photons emitted from the sample. The time interval between each laser pulse and the corresponding photon detection event is measured with picosecond resolution using a TDC. A TDC is a digital circuit that measures time intervals with high precision and converts them into digital codes [ 45 ]. Hardware implementations of TDCs can vary depending on performance requirements and system constraints. Two common methods for implementing high-resolution TDCs in SPAD-based systems are the delay line and ring oscillator approaches [ 46 ]. In a delay line TDC, a reference signal propagates through a chain of delay elements, and the number of activated stages at the time of detection provides a high-resolution timestamp. This method offers good linearity and resolution but consumes significant area. In contrast, a ring oscillator-based TDC samples the oscillator’s phase at the time of detection, offering a more compact and power-efficient solution suitable for columnor pixel-level integration. However, it generally suffers from higher jitter and nonlinearity. The choice between these methods involves trade-offs among resolution, area, power, and calibration complexity, depending on the system’s requirements. The placement of the TDC in SPAD imaging systems significantly impacts performance and design complexity. Per-pixel TDCs offer maximum parallelism and temporal resolution, 19
Chapter 2. Fundamentals of Intelligent Time-Resolved Imaging but they are areaand power-intensive, limiting scalability. Column-level TDCs provide a balance between resolution and resource efficiency by sharing one TDC across multiple SPADs, making them suitable for moderate-resolution, low-power designs. Off-chip TDCs, often implemented on FPGAs or ASICs, centralize timing conversion and simplify layout but suffer from limited scalability. Commercial solutions such as PicoQuant’s TCSPC modules provide high-resolution, multi-channel time tagging with picosecond accuracy and are commonly used in research setups for applications like FLIM and fluorescence correlation spectroscopy. The choice depends on the required trade-offs among timing precision, area, power, and system bandwidth. Time Gating Time gating is a technique used in SPAD imaging systems to restrict photon detection to specific time intervals, improving temporal selectivity and reducing background noise. In a time-gated SPAD, gating can be applied either by modulating the SPAD’s bias voltage, activating it only during desired time windows, or by electronically masking the SPAD output, allowing detection signals to pass through only during predefined intervals. Gating is typically synchronized with an external timing reference, such as a pulsed excitation source, and is especially effective in rejecting photons arriving outside the expected time-of-arrival window. This enhances the signal-to-noise ratio in time-resolved applications. Compared to TCSPC, time gating provides a hardware-efficient alternative for temporal filtering, often reducing data volume and system complexity. Large SPAD arrays can be made thanks to its simplicity [ 8 , 47 ]. Its simpler pixel circuit translates to high fill factor. The choice between bias gating and output gating depends on factors such as timing resolution, power constraints, and circuit complexity. Human vision is nonlinear and more sensitive to changes in darker regions than in brighter ones, and gamma correction exploits this property to make luminance appear perceptually uniform. It is a common technique in video processing, where a nonlinear curve is applied to the intensities acquired from CMOS or CCD cameras so that more detail is preserved in dark regions and the signal better matches human perception [ 48 ]. A time-gated SPAD features a native Gamma correction . By modeling the incident photon, a Poisson process, with a binary distribution, time-gated SPADs effectively realize Gamma correction by nature, mapping the real intensity Iinto the measured intensity ˆ Iwith the following rule [49]: I=−Nrep lnµ1−ˆ I Nrep ¶(2.7) where Nrep is the number of repetitions or binary acquisitions. The native Gamma correction enables high dynamic range, at the expense of loss of details in the highlights. 20
2.2 Fluorescence Lifetime Imaging S0 S2 T1 S1 Intersystem Crossing Vibrational Relaxation Internal Conversion Phosphorescence Fluorescence Absorption Fig. 2.3: Jablonski diagram showing the possible radiative and non-radiative transitions. Upon absorption of a photon, the molecule is excited from the ground state to an excited state, followed by vibrational relaxation. It subsequently returns to the ground state while emitting a photon, i.e., fluorescence. The figure is adapted from Edinburgh Instruments. 2.2 Fluorescence Lifetime Imaging 2.2.1 Fluorescence Lifetime Fluorescence is the process by which a fluorophore absorbs a photon, becomes raised to an excited state, and then returns to the ground state by emitting a photon [ 50 ]. As illustrated in Fig. 2.3, each absorption event may lead to excitation into different vibrational sublevels of S1 , followed by rapid non-radiative relaxation to the lowest vibrational level. From there, the molecule returns to the ground state by emitting a photon. This emission does not occur at a fixed time. Instead, each path introduces a different delay between absorption and emission. Statistically, these delays follow an exponential distribution, governed by the fluorescence lifetime τ , which represents the average time the molecule remains in the excited state before emitting a photon. FLIM is an imaging technique for the characterization of molecules based on the fluorescence lifetime [ 51 ]. Compared with fluorescence intensity imaging, FLIM is insensitive to excitation intensity fluctuations, variable probe concentration, and limited photobleaching. Besides, through the appropriate use of targeted fluorophores, FLIM is able to quantitatively measure the parameters of the microenvironment around fluorescent molecules, such as pH, viscosity, and ion concentrations [ 52 , 53 ]. With these advantages, FLIM has wide applications in the biological sciences, for example to monitor protein-protein interactions [ 54 ], and plays an increasing role in medical and clinical settings such as visualization of tumor margins [ 55 ], cancerous tissue detection [51, 56], and computer-assisted robotic surgery [57, 58]. TCSPC is popular among FLIM systems due to its superiority over other techniques in terms 21
Chapter 2. Fundamentals of Intelligent Time-Resolved Imaging ... Fig. 2.4: Workflow of TCSPC FLIM. The pulsed laser repeatedly illuminates the sample, as shown by the blue curve. The orange curve represents the probability of photon emission. When a photon is detected within a period (indicated by the red dot), its timestamp is recorded in the histogram. After many repetitions, the accumulated histogram is obtained for further analysis. of time resolution, dynamic range, and robustness. In TCSPC, one records the arrival time of individual photons emitted by molecules upon photoexcitation [ 44 , 59 , 60 ]. After repeated measurements, one can construct a histogram of photon arrivals, which closely matches the true response of molecules, thus enabling the extraction of FLIM, as shown Figure 4.1. The instrumentation of a typical TCPSC FLIM system features a confocal setup, including a singlephoton detector, a dedicated TCSPC module for time tagging, and a computer for lifetime estimation [ 44 , 61 ]. Such systems are mostly unsuitable for increasing clinical applications such as non-invasive monitoring, where a miniaturized and fast FLIM system is desired [ 62 ]. Additionally, the substantial volume of data produced by TCSPC imposes a significant load on data transfer, storage, and processing. A powerful computer, sometimes equipped with dedicated GPUs, is required to acquire and process TCSPC data. TCSPC requires photodetectors with picosecond time resolution and single-photon detection capability. In the last decade, SPADs have been used successfully in TCSPC systems and, with the advent of CMOS SPADs, the expansion of these detectors into high-resolution image sensors for widefield imaging has been accomplished [ 17 ]. Several reviews of the use of SPADs in biophotonics have recently appeared [6, 63, 64]. Unlike TCSPC, which records the precise arrival time of individual photons relative to an excitation pulse, time-gating groups photon detections into predefined temporal windows following the excitation. In FLIM, this approach involves sequentially opening detection gates at increasing time delays to sample different portions of the fluorescence decay curve. Each gate captures photons arriving within a specific time range, and by combining data from multiple gates, the fluorescence lifetime can be estimated for each pixel. Time-gating typically requires less complex timing circuitry than TCSPC, making it well-suited for parallelized detection with SPAD arrays. While it offers coarser temporal resolution than TCSPC, timegated FLIM is believed to enable faster, scalable acquisition and is more tolerant to high 22
2.2 Fluorescence Lifetime Imaging photon flux, making it attractive for real-time or wide-field imaging applications [8, 47, 49]. 2.2.2 Instrumentation The instrumentation for FLIM typically consists of a pulsed excitation source, a fluorescence collection optics system, a photon detector such as a SPAD or PMT, and precise timing electronics for resolving the time delay between excitation and photon emission. Depending on the implementation, systems may operate in the time domain, using techniques such as TCSPC or time-gated detection, or in the frequency domain, where phase and modulation of the fluorescence signal are measured. Time-resolved approaches often require high temporal resolution, low-noise detectors, and accurate synchronization between the excitation pulses and photon detection. The design of FLIM instrumentation must also consider factors such as photon throughput, timing jitter, detector sensitivity, and system latency to ensure reliable and high-contrast lifetime imaging, particularly in low-light or high-speed biological imaging applications. 2.2.3 Fluorescence Lifetime Estimation After acquiring time-resolved measurements with either TCSPC or time-gating, the next step is to estimate the fluorescence lifetime from the recorded decay data. This process is crucial for translating photon timing information into meaningful biochemical contrast. These range from traditional curve-fitting techniques to faster, more computationally efficient approaches like the method of moments, phasor analysis, or machine learning-based inference. Each method offers a different trade-off between robustness, precision, and suitability for real-time or photon-limited conditions. Maximum Likelihood Estimation Maximum likelihood estimation (MLE) is one of the classic methods for estimating fluorescence lifetimes from time-resolved photon detection data [ 65 , 66 ]. MLE provides an optimal framework for fitting an exponential decay model to the measured photon arrival times, particularly under low-light conditions typical of biological samples. Given the IRF and the fluorescence lifetimes τand intensities α, the target function p(t|τ,α) is represented as: p(t|τ,α)=p(t|τ1,τ2,...,τn,α1,α2,...,αn)=Zt 0IRF(s)· n X k=1 αk τk e−t−s τkds (2.8) The IRF characterizes a system’s temporal response to an instantaneous excitation. The photon arrival times are binned into a histogram {hj} over M bins with edges {b0,b1,...,bM} . 23
Chapter 2. Fundamentals of Intelligent Time-Resolved Imaging ELENa EK RLRNa RK CmU Iext (a) HH model Cm E R U Iext (b) LIF model Fig. 2.7: Neuron circuits of HH model and LIF model in the subthreshold domain. (a) Three channels are considered in the HH model, including two with variable resistors, which are excluded in (b). While the HH model provides a detailed and biophysically plausible description of neuronal dynamics, its complexity makes it computationally expensive and analytically intractable for large-scale simulations. This has motivated the development of simplified spiking neuron models that retain essential features of neural excitability and spike generation while offering greater mathematical and computational tractability. The leaky integrate-and-fire (LIF) model is one of the most widely used simplified neuron models [ 87 ]. It describes the membrane potential as a leaky integrator of input current, with a fixed threshold for spike generation. The subthreshold membrane potential dynamics is mathematically expressed as: Cm dU dt =−g(U−E)+Iext (2.19) where U is the membrane potential, Cm is the membrane capacitance, Iext is the external current, g is the conductance, and E is the reversal potential. When the membrane potential reaches the threshold of the neuron, a spike is emitted and the potential is reset. Despite its simplicity, the LIF model captures the basic mechanism of temporal integration and thresholdbased firing, making it a useful abstraction for studying spike-based computation. It is the most common spiking neuron model employed in the computing community, and will also be the main focus of this thesis. The SRM generalizes the LIF framework by representing the membrane potential as a sum of response kernels triggered by incoming spikes and the neuron’s own firing history. This approach provides more flexibility in shaping subthreshold dynamics and refractory effects, allowing for a closer approximation to experimentally observed neuronal responses without explicitly modeling biophysical processes. 30
2.4 Spiking Neural Networks t T tk−1tktk+1 R=RT 0S(t) T (a) Rate coding t ∆t tk−1tktk+1 d(t)=Rt+∆t tS(t) ∆t (b) Density coding t φ[k]T tk−1tktk+1 φ[k]=tkmod T (c) Phase coding t I[k−1] I[k] tk−1tktk+1 I[k]=tk−tk−1 (d) Inter-spike time coding Fig. 2.8: Neural coding schemes. The spike train is modeled as S ( t ) =Pkδ ( t−tk ), where tk denotes the firing time of the k-th spike. 2.4.2 Neural Coding Since it is spikes, not continuous values, that are transmitted between neurons, a natural question arises: how should these spikes be interpreted? Or conversely, how is information encoded in spike trains? Four coding schemes discussed in this thesis are introduced [ 88 ], which are illustrated in Fig. 2.8. Rate coding is one of the most widely studied neural coding schemes, in which information is represented by the firing rate of a neuron over a given time window. In this framework, the precise timing of individual spikes is neglected, and only the average number of spikes within a period is considered. A higher firing rate corresponds to a stronger signal or stimulus, while a lower rate indicates weaker activity. Rate coding is appealing due to its robustness to noise and its conceptual similarity to activation levels in traditional ANNs. It also forms the basis for many ANN-to-SNN conversion methods, where continuous-valued neuron outputs are translated into spike rates. Density coding is a neural coding scheme in which information is represented by the temporal distribution or concentration of spikes. It shares similarities with rate coding but operates on much shorter time windows, capturing rapid, moment-to-moment variations in spike activity. By focusing on local variations in spike density rather than long-term averages, this approach enables faster and more responsive encoding of dynamic inputs. Density coding thus bridges the gap between rate-based and precise spike timing codes, offering a balance between temporal resolution and robustness for both biological and neuromorphic systems. Phase coding is a neural coding strategy in which information is encoded in the timing of 31
Chapter 2. Fundamentals of Intelligent Time-Resolved Imaging spikes relative to an underlying rhythmic signal, such as a neural oscillation. Rather than relying on firing rate or spike count, phase coding leverages the precise timing of a spike within a cycle of the oscillatory rhythm to convey information. This mechanism has been observed in various brain regions, such as the hippocampus, where neurons exhibit phase precession. Phase coding allows for compact and temporally precise representations, enabling neurons to multiplex information by combining spike timing with oscillatory phase. Inter-spike interval (ISI) coding is a temporal coding scheme in which information is conveyed through the time intervals between consecutive spikes. In this framework, shorter or longer intervals can represent different stimulus features or signal intensities, enabling neurons to encode information with fine temporal precision. ISI coding has been observed in various sensory systems, such as the auditory and somatosensory pathways, where rapid discrimination and adaptation are critical. 2.4.3 Training Spike-timing dependent plasticity (STDP) is a biologically inspired learning rule that adjusts synaptic weights based on the precise timing of preand postsynaptic spikes [ 89 , 90 ]. In its classical form, STDP strengthens a synapse if a presynaptic spike precedes a postsynaptic spike within a short temporal window, and weakens it if the postsynaptic spike occurs first. This asymmetric temporal dependence captures a form of Hebbian learning, often summarized as “cells that fire together wire together”, with added temporal sensitivity that reflects causality. STDP has been experimentally observed in various brain regions and is believed to play a key role in activity-dependent development, learning, and memory formation. In spiking neural networks, STDP offers a local and online mechanism for learning, enabling the network to adapt based on temporal correlations in spiking activity without requiring global error signals or labeled data. A spiking neural network can be unfolded through time, making it possible to be trained by backpropagation through time (BPTT). However, because the LIF neuron’s firing is nondifferentiable, gradients cannot be directly propagated through spike events [ 91 , 92 ]. Various methods have been proposed to tackle this problem. ANN-to-SNN conversion is a widely used strategy to leverage the performance of conventional ANN while deploying models in the more energy-efficient and biologically inspired SNN framework [ 93 – 95 ]. The core idea is to first train a standard ANN using established backpropagation techniques and then convert it into an SNN by interpreting activations as spike rates. This typically involves mapping rectified linear unit (ReLU) activations to spike counts over a time window, adjusting weights and thresholds to preserve information flow. This approach focuses solely on firing rates, disregarding the latency or phase information of individual spikes. However, it enables deep SNNs to inherit the high performance of trained ANNs, enabling efficient processing on neuromorphic hardware. 32
2.4 Spiking Neural Networks To leverage the information of individual spikes, several methods are proposed, but they typically set strict constraints on the coding scheme or the firing behavior. For example, backpropagation through spike time assumes latency coding and limits the number of firings [ 96 , 97 ]. Recently, surrogate gradients have emerged as a powerful and versatile training method for spiking neural networks, applicable across various neural coding schemes [ 93 , 98 – 100 ]. In this approach, the non-differentiable Dirac delta function associated with spike firing is replaced by a smooth, differentiable approximation during the backward path, enabling gradient-based optimization. This allows the use of standard deep learning techniques for SNN training while preserving spike-based computation. However, the forward path still uses the true non-differentiable spike function, meaning the computed gradient does not correspond to the true gradient of the forward dynamics. As a result, surrogate-gradient training optimizes a proxy objective rather than the exact spiking behavior, which introduces bias and may lead to suboptimal solutions if the surrogate does not faithfully reflect the underlying neuronal mechanisms. 2.4.4 Applications Despite the difficulty of efficient training, SNNs have been used in several applications. ANNs are transformed into SNNs to reduce energy consumption and running time [ 101 , 102 ]. SNNs have been used for detectors such as electroencephalogram [ 103 ] and event cameras [ 104 – 106 ] for various applications. The community, however, focuses more on classification than regression [ 104 , 107 ], more on still images than sequences, and more on rate-based coding than other coding schemes, while SPAD-based applications are more of phase coding, sequential processing, and regression. Only a few studies have been performed on the SNNs for SPAD and other single-photon detectors. In [ 108 ], it was first proposed to use SNNs to process single-photon signals; the authors use SNN to process LiDAR raw data for object detection. However, it is assumed that all the spikes come within one repetition period, which is not applicable to other tasks such as FLIM. The authors in [ 109 ] implemented SNNs with SPAD sensors, but the proposed approach only works on handwritten digit recognition with simulation data as passive imaging. In [ 110 ], intensities from an active imaging setup are converted into spike trains using a Poisson encoder, which are then processed by an SNN for hand gesture recognition. The work described in Chapter 5 represents one of the first attempts to explore general frameworks using SNNs for active SPAD-based time-resolved imaging [ 19 ]. Following this work, the authors of [ 111 ] proposed a SPAD–SNN framework for LiDAR, which aggregates macropixel sums over time. The resulting values are binarized, and the multi-bit temporal sequence is used as the spike train input to the SNN. 2.4.5 Neuromorphic Chip Neuromorphic chips are specialized hardware systems designed to emulate the structure and function of the brain [ 112 , 113 ]. They draw attention from scientists from both the biology and computer science communities. From a neuroscientist’s perspective, neuromorphic chips 33
Chapter 2. Fundamentals of Intelligent Time-Resolved Imaging provide a platform to test biologically plausible models of neural dynamics at scale, enabling the simulation of spiking neurons, synaptic plasticity, and network-level behavior in real time. These chips allow for experimental validation of theories on neural coding, learning, and sensory processing under conditions that are difficult to reproduce in traditional computing environments. From a computer scientist’s viewpoint, neuromorphic chips represent a paradigm shift in computation, moving away from the von Neumann architecture toward event-driven, massively parallel systems that are inherently energy-efficient and scalable. By encoding and processing information using spikes, similar to biological neurons, neuromorphic systems can perform tasks such as perception, pattern recognition, and decision-making with minimal power consumption. This convergence of neuroscience and computer science opens the door to building intelligent machines that are not only faster and more efficient but also fundamentally different in how they learn and interact with the world. Several neuromorphic chips have been developed in recent years, each reflecting different design philosophies inspired by the brain’s architecture and dynamics. Among the most prominent is IBM’s TrueNorth, which features a massively parallel, event-driven architecture composed of one million spiking neurons and 256 million synapses, designed for ultra-lowpower pattern recognition tasks [ 114 ]. Intel’s Loihi chip advances this further by incorporating on-chip learning through programmable synaptic plasticity rules, enabling real-time adaptation and supporting a wide range of spiking neural models [ 115 , 116 ]. Another notable platform is SpiNNaker, developed by the University of Manchester, which uses ARM cores to simulate large-scale spiking networks in a biologically realistic manner, making it ideal for neuroscience research [ 117 ]. BrainScaleS, from Heidelberg University, offers a mixed-signal approach where analog circuits emulate neural dynamics at accelerated timescales, facilitating fast experiments and network exploration [ 118 , 119 ]. These chips vary in their trade-offs between biological realism, programmability, speed, and energy efficiency, but all share the common goal of enabling brain-like computation in silicon. Combined with the image sensor, neuromorphic image sensors have been developed [ 29 , 120 , 121 ]. Most existing work has concentrated on integrating neuromorphic processors into event-based cameras, enabling low-latency and energy-efficient vision processing directly at the sensor level. This paradigm has even reached commercialization, with companies such as Prophesee 2 offering event-driven vision systems. More recently, the integration of neuromorphic processors with SPAD arrays has begun to attract research interest [ 42 , 122 , 123 ]. Current studies, however, have primarily focused on passive imaging applications, such as low-light or photon-efficient imaging, while the unique potential of neuromorphic computing for time-resolved SPAD imaging remains largely unexplored. 2https://www.prophesee.ai/ 34
3Modeling for SPAD-Based FLIM Accurate modeling plays a central role in understanding and optimizing the use of SPADs for time-resolved imaging, including FLIM. It provides a way to connect experimental observations with the underlying behavior of the system, with which a structured framework to predict outcomes and guide development can be developed. This chapter is therefore devoted to building that framework, not only to support the analysis and results presented in the rest of the thesis, but also to address questions raised by the collaborators, who sought a clearer understanding of how modeling can inform the design and application of SPAD-based FLIM. 3.1 Modeling of SPAD Single-photon detection was extensively modeled and discussed in the 1970s. However, these models have, to some extent, become disconnected from the current SPAD community. This section revisits the modeling of SPADs as single-photon detectors, with a focus on photon detection and dead-time effects. 3.1.1 Photon Detection For coherent light such as laser, the arrival of photons is modeled as a Poisson process. Given λ as the expected number of arriving photons within a time window T , the number of photons N arriving in the time window follows a Poisson distribution, whose probability mass function (PMF) is: P(N=k)=λke−λ k!,k∈N(3.1) As not every photon will necessarily trigger the avalanche process, each of the N photons has a probability p∈ [0 , 1] of being detected, which can be modeled by the Bernoulli process, and p is defined as PDP. For given N=k , the number of detected photons D is the sum of a 35
Chapter 3. Modeling for SPAD-Based FLIM sequence of independent Bernoulli variables: D= N X i=1 Zi(3.2) where P(Z=1) =p=1−P(Z=0) and Dfollows a Binomial distribution D∼Binomial(k,p): P(D=d|N=k)=Ãk d!pd(1−p)k−d(3.3) The marginal distribution of detected photons D can be retrieved by integrating over all possible values of N: P(D=d)=∞ X k=d P(D=d|N=k)·P(N=k) =∞ X k=dÃk d!pd(1−p)k−d·λke−λ k! =(λp)de−λp d! (3.4) which is the PMF of a Poisson distribution with mean λp . Hence, the number of detected photons D∼Poisson ( λp ). The Poisson process of the arrival of photons is actually thinned by the Bernoulli process of avalanche triggering. To verify the theoretical derivation, a Monte Carlo simulation is carried out. Each simulation trial generates a photon count given λ and p . The number of incident photons N is generated by a Poisson distribution with λ= 1000, and the number of detected photons D is generated by a Binomial distribution with p= 0 . 3 and k=N . The experiments are performed million times. The result is shown in Fig. 3.1a. Simulation results are consistent with analytical predictions, as a Poisson distribution with a mean of 300 is perfectly fitted to the simulated detected photon distribution. The thinning effect of the Bernoulli process is illustrated in Fig. 3.1b. Poisson distributions with means of 300 and 1000 are plotted, respectively. The marginal distribution P∞ k=dP ( D=d|N=k ) ·P ( N=k ) is computed and plotted, which overlaps with Poisson(λ=300) as expected. It is worth noting that the choice between binomial and Poisson modeling depends on the nature of the photon count randomness. When the photon count is fixed, a binomial distribution is used. This case will be further discussed in the context of incoherent light in the following paragraphs. Conversely, when the photon count is random and follows a Poisson distribution, a thinned Poisson distribution is required. As a result of the Bernoulli process from an imperfect photon detection probability, both the mean and variance of the photon count are reduced from λ to λp , which raises the question of whether this conclusion applies to incoherent light as well. Examples of photon distribution is 36
3.1 Modeling of SPAD 200 225 250 275 300 325 350 375 Number of Detected Photons 0.000 0.005 0.010 0.015 0.020 Probability Poisson Fit mean=300.01, scale=1.0 Simulated Detected Photons (a) Simulated detected photons and Poisson fit 0 200 400 600 800 1000 1200 1400 Number of Detected Photons 0.000 0.005 0.010 0.015 0.020 Probability Poisson ( λ = 300) Poisson ( λ = 1000) Marginal distribution (b) Thinning of the Poisson distribution Fig. 3.1: Monte Carlo simulation of the Poisson nature of the detection of photons. (a) Photon detection, taking Poisson process and Bernoulli detection into account, is simulated. A Poisson distribution is fitted well on it. (b) Poisson distribution with λ= 1000 and λ= 300 are plotted. With a smaller λ , the distribution is “thinner” and “taller”. The marginal distribution and the Poisson distribution with λ=300 overlap with each other. shown in Fig. 3.2. The short answer is no. The photon arrivals of antibunched light, which are more equally spaced than in the case of coherent light, are governed by sub-Poissonian statistics [ 124 ]. An extreme case is when photons are equally spaced (e.g., single-photon emitters), resulting in deterministic counts over a given period. In this case, the photon count follows the binomial distribution. The mean is λp , which is the same as coherent light. The variance is λp (1 −p ), which is smaller than coherent light. While the variance of coherent light is positively proportional to p , the variance of such antibunched light reaches maximum when p=0.5. Thermal light is a typical example of bunched light, whose count of arriving photons follows a geometric distribution: P(N=k)=1 1+λµλ 1+λ¶k (3.5) The marginal distribution is: P(D=d)=∞ X k=d P(D=d|N=k)·P(N=k) =∞ X k=dÃk d!pd(1−p)k−d·1 1+λµλ 1+λ¶k =(λp)d (1+λp)d+1 (3.6) which still follows a geometric distribution, with new mean λp and new variance λp· (1 +λp ). The marginal distributions P∞ k=dP ( D=d|N=k ) ·P ( N=k ) with three different P ( N=k ) are 37
Chapter 3. Modeling for SPAD-Based FLIM Super-Poissonian p(t)=t τ2e−t τ Poissonian p(t)=1 τe−t τ Sub-Poissonian p(t)=δ(t−τ) τ Fig. 3.2: Photon distribution of Poissonian, sub-Poissonian, and super-Poissonian light. Photons are randomly generated with the given distribution with a mean of τ . The photons from super-Poissonian light tend to be “bunched” together, as the photons from subPoissonian light tend to be more evenly spaced, or “anti-bunched”. 0 200 400 600 800 1000 Number of Detected Photons 0.000 0.005 0.010 0.015 0.020 0.025 Probability Poisson distribution Binomial distribution Geometric-binomial distribution (a) Simulated detected photons 0.0 0.2 0.4 0.6 0.8 1.0 Photon detection probability 0 250 500 750 1000 1250 1500 1750 2000 Variance Poisson distribution Binomial distribution Geometric-binomial distribution (b) Variance vs photon detection probability Fig. 3.3: Monte Carlo simulation of photon detection with incident photon count following Poissonian, sub-Poissonian, and super-Poissonian distributions. (a) All three distributions are scaled with p= 0 . 3. (b) The variance of the Geometric distribution shoots sky-high as it increases quadratically with λp . The variance of the Poisson distribution is proportional to p . The variance of a single-photon emitter reaches the maximum when p=0.5. 38
3.1 Modeling of SPAD computed and plotted, given λ= 1000 and p= 0 . 3. The result is shown in Fig. 3.3a. Monte Carlo simulation is carried out, where each data point is simulated with 10 million photons. Variances are calculated with respect to p ranging from 0 to 1, given λ= 1000, as illustrated in Fig. 3.3b. The result is in accordance with the theoretical prediction. 3.1.2 Dead Time The problem of the dead time of PMTs has been extensively studied in the mid-20th century [ 125 ]. The same principle applies to the dead time of SPAD. The statistical model of dead time can be essentially simplified into two types, non-paralyzable or non-extended dead time, and paralyzable or extended dead time. When the dead time is non-paralyzable, incoming photons during dead time are simply ignored. For the paralyzable dead time, each incoming photon during dead time extends the dead time. The inter-arrival time of photons governed by Poisson statistics follows an exponential distribution: p(t;ρ)=ρe−ρt(3.7) With a non-paralyzable dead time τ, the distribution is updated: p(t;ρ,τ)= 0, 0 <t<τ ρe−ρ(t−τ),t≥τ.(3.8) The photon counts Nis written as: P(N=k)=P(Sk≤T<Sk+1)=ZT 0p∗k(s)Z∞ T−s p(t)dtds (3.9) where Sk is the arrival time of the k -th photon and p∗k ( s ) is the k -fold convolution of p ( s ). Such a process is considered as the generalization of the Poisson process, and is termed renewal process. Given Eq. 3.8. The PMF does not have an analytical form. The approximation is often made by considering a live time T−kτ , as k dead periods of length τ remove kτ seconds from the total interval T . The probability of photon arrival within a live time is assumed to follow a Poisson distribution within the reduced interval, whose PMF is: P(N=k)≈ (ρ(T−kτ))k k!e−ρ(T−kτ), 0 ≤k≤¥T τ¦ 0, otherwise. (3.10) Given λ=ρTand η=τ/T, it can be rewritten as: P(N=k)≈ (λ(1−kη))k k!e−ρ(1−kη), 0 ≤k≤j1 ηk 0, otherwise. (3.11) 39
Chapter 3. Modeling for SPAD-Based FLIM background level. And the PDF for time-gating is denoted at pTG ( t ; τ,µ,σ,B,Tg ), where Tg is the gate length. As closed-form expressions of pTCSPC ( t ; τ,µ,σ,B ) and pTG ( t ; τ,µ,σ,B,Tg ) are difficult to obtain due to convolution operations, numerical simulations are employed to compute the CRLB. Two scenarios are considered: equal photon budget and equal exposure duration. In the equal photon budget scenario, an infinite number of gate channels is assumed, meaning that every detected photon is registered across all included gate windows, as in the TCSPC. In the equal exposure duration scenario, a single gate channel is considered. Given the same total exposure time as in TCSPC, only a fraction Tg Tof the photons are registered. Numerical Simulation Given the simulation step Tstep , pTCSPC ( t ; τ,µ,σ,B ) and pTG ( t ; τ,µ,σ,B,Tg ) are wrttien as PTCSPC(nTstep;τ,µ,σ,B) and PTG(nTstep;τ,µ,σ,B,Tg). Then the FI are: ITCSPC(τ)=En·logPTCSPC(nTstep;τ+∆τ,µ,σ,B)−PTCSPC(nTstep;τ−∆τ,µ,σ,B) 2∆τ¸2 (3.30) and ITG(τ)=En·logPTG(nTstep;τ+∆τ,µ,σ,B,Tg)−PTG(nTstep;τ−∆τ,µ,σ,B,Tg) 2∆τ¸2 (3.31) CRLB’s dependence on the number of detected photons M is particularly important in photonlimited regimes such as fluorescence lifetime imaging. In general, the CRLB for estimating the fluorescence lifetime τ scales inversely with M , meaning that the estimation precision improves as more photons are collected: Var(τ)>1 M·I(τ)(3.32) The simulation is performed in Python using 64-bit floating point precision. The IRF parameters are fixed at µ= 5 ns and σ= 2 ns. The CRLB is calculated across a range of parameters: background level B from 0 to 1, fluorescence lifetime τ from 1 to 10 ns, and gate width Tg from 1 to 49 ns. The simulation time step is set to Tstep = 0 . 05 ns, and the finite difference step for derivative calculation is ∆τ=τ/100. The photon count is set to M=106·(1+B). Scenario I: Equal Photon Budget The equal photon budget scenario assumes that both TCSPC and time-gating methods register the same number of photons. In practice, time-gating can only achieve this by employing a longer exposure time due to its limited temporal sampling. This scenario is studied to evaluate how background noise and gate length influence the estimation precision of the time-gating 46
3.3 TCSPC and Time-Gated Time-Resolved Imaging 5 10 15 20 25 30 35 40 45 Gate Length (ns) 0.00.10.20.30.40.50.60.70.80.91.0 Background Noise 0.002 0.003 0.004 0.005 0.006 0.007 0.008 0.009 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09 CRLB (ns) Fig. 3.5: CRLB with different combinations of background noise level and gate length (Scenario I). The lifetime is 5 ns, and repetition period is 50 ns. method, and to compare its performance against that of TCSPC. Given a lifetime of 5 ns, the CRLB with different combinations of background noise level and gate length is calculated, shown in Fig. 3.5. As the background noise level increases, the CRLB rises accordingly across all gate lengths. Intuitively, this is expected, as higher background noise reduces the precision of the lifetime estimation. A more precise explanation can be obtained by examining the definition of the CRLB. The PDF can be expressed as a weighted sum of the signal component and the background noise component. The presence of background noise effectively “dilutes” the contribution of the signal to the overall PDF. When a small perturbation ∆τ is applied to the lifetime parameter, the resulting change in the PDF becomes less pronounced in the presence of the background noise, leading to a reduction in Fisher Information. Since the CRLB is inversely proportional to the Fisher Information, this results in an increase in the CRLB. Given a fixed background noise level, increasing the gate width Tg leads to an increase in the CRLB. This can be interpreted in several ways. Intuitively, as Tg→ 0, the method becomes equivalent to TCSPC, which achieves the minimum CRLB due to its maximal temporal resolution. In contrast, when Tg→T , where T is the total observation window, all temporal information is lost, and the CRLB diverges to infinity. Hence, there exists a transition region in which increasing Tg results in a monotonic increase in the CRLB. From a statistical standpoint, as Tg increases, the effective PDF becomes smoother due to temporal integration. As a result, a small perturbation ∆τ induces a smaller change in the PDF, reducing the FI and thereby increasing the CRLB. From an information-theoretic perspective, timegating can be viewed as a function of the underlying TCSPC measurement. According to the data processing inequality, any such transformation cannot increase the information about the parameter of interest. Therefore, time-gating always performs no better, and typically worse, than TCSPC in terms of FI. 47
Chapter 3. Modeling for SPAD-Based FLIM 5 10 15 20 25 30 35 40 45 Gate Length (ns) 12345678910 Lifetime (ns) 0.002 0.003 0.004 0.005 0.006 0.007 0.008 0.009 0.02 0.03 CRLB/lifetime Fig. 3.6: CRLB with different combinations of lifetime and gate length (Scenario I). The background noise is 50%, and the repetition period is 50 ns. Given a background noise level of 50%, the CRLB is calculated for various combinations of background noise and gate length. As expected, the CRLB increases with longer lifetimes, since the perturbation ∆τ becomes less pronounced for larger lifetime values compared to shorter ones. To reflect the relative estimation error and facilitate comparison across lifetimes, the CRLB is normalized by the corresponding lifetime. The results are presented in Fig. 3.6. Shorter lifetimes exhibit larger relative CRLBs, as they approach the resolution limit set by the simulation step size. For short lifetimes, the relative CRLB increases markedly with gate length. This suggests that the temporal averaging effect introduced by wider gates disproportionately degrades the estimation precision when the decay process is fast. In essence, broader gates smear out the temporal information needed to resolve fast decays. For lifetimes above ∼ 4 ns, the relative CRLB remains low and relatively stable across all gate lengths. This indicates that longer lifetimes are more robust to variations in gate width, as their decay characteristics are easier to resolve even with broader integration windows. Scenario II: Equal Exposure Duration The equal exposure duration scenario assumes that both TCSPC and time-gating methods are exposed for the same amount of time. Since a single gate channel is considered, time-gating registers fewer photons than TCSPC, with a ratio of Tg T . This represents a more realistic case when the maximum allowable exposure duration is limited by sample photobleaching or system operation time. Given a lifetime of 5 ns, the CRLB with different combinations of background noise level and gate length is calculated, shown in Fig. 3.7. The CRLB is weighted by the partial photon count registration qT Tg . The CRLB-background noise relationship discussed under the equal 48
3.4 Summary 5 10 15 20 25 30 35 40 45 Gate Length (ns) 0.00.10.20.30.40.50.60.70.80.91.0 Background Noise 0.004 0.005 0.006 0.007 0.008 0.009 0.02 0.03 0.04 0.05 0.06 0.07 0.08 0.09 CRLB (ns) Fig. 3.7: CRLB with different combinations of background noise level and gate length (Scenario II). The lifetime is 5 ns, and repetition period is 50 ns. photon budget assumption still holds. Hpwever, a U-shaped trend is observed along the gate length axis, where CRLB initially decreases with increasing gate width due to improved photon collection, but rises again for wide gates due to temporal averaging and overlap effects. Given a background noise level of 50%, the normalized CRLB is calculated for various combinations of background noise and gate length, as shown in Fig 3.8. A U-shaped dependence on gate length is observed across lifetimes, with optimal estimation performance occurring between 20–30 ns. Short lifetimes exhibit higher relative error, especially at wide gates, due to their sensitivity to temporal averaging. The high CRLB observed at short gate lengths is primarily due to limited photon counts, whereas the high CRLB at long gate lengths is dominated by information loss resulting from the averaging effect inherent to time gating. Fig 3.8 also reveals a weak U-shaped trend along the lifetime axis. For most gate lengths, the normalized CRLB decreases as the fluorescence lifetime increases from 1 ns to approximately 5–7 ns, then increases again for longer lifetimes. This behavior suggests that there exists an optimal range of lifetimes, typically in the mid-range, where the relative estimation error is minimized. These results indicate that estimation accuracy, when normalized by lifetime, is not uniform across decay constants and is best in a regime where the fluorescence decay is well matched to the temporal structure of the gate window. 3.4 Summary This chapter presents the modeling efforts developed in response to key questions raised throughout the PhD, both individually and in collaboration with others. It focuses on the behavior of SPAD sensors in the context of fluorescence lifetime imaging, aiming to understand the limits and trade-offs inherent in time-resolved photon detection. 49
Chapter 3. Modeling for SPAD-Based FLIM 5 10 15 20 25 30 35 40 45 Gate Length (ns) 12345678910 Lifetime (ns) 0.003 0.004 0.005 0.006 0.007 0.008 0.009 0.02 0.03 CRLB/lifetime Fig. 3.8: CRLB with different combinations of lifetime and gate length (Scenario II). The background noise is 50%, and the repetition period is 50 ns. The chapter begins with a detailed model of SPAD detection, incorporating photon arrival statistics, the probabilistic nature of SPAD detection, and the effect of dead time. It is shown that the combination of coherent photon arrival, modeled as a Poisson process, and SPAD detection, modeled as a Bernoulli process, preserves the Poisson distribution. This result is further extended to certain incoherent light sources, such as single-photon emitters and anti-bunched light, highlighting the generality and applicability of the statistical framework. Building on this, fluorescence lifetime imaging is modeled at the photon level, in contrast to the existing efforts to construct a dataset on the histogram level. It incorporates fluorescence, instrument response function, and background noise. This model serves as a basis for the training of neural networks on synthetic datasets, and enables Monte Carlo simulation and analysis of Cramér-Rao lower bounds. A significant part of the chapter compares two dominant time-resolved acquisition methods: TCSPC and time gating. This comparison emerged from both practical design questions and theoretical considerations during the project. TCSPC is modeled as an event-based, highresolution approach, while time gating is modeled as a windowed integration method with higher throughput but coarser temporal resolution. The models helped clarify how to choose the best method and parameters for a given application, depending on photon budget, timing constraints, and hardware limitations. Overall, this chapter reflects a process of iterative questioning and modeling that shaped the theoretical foundation for later system design and experimentation. It captures how SPAD physics, lifetime estimation challenges, and acquisition trade-offs were approached from first principles through to implementation-level implications. 50
4 SPAD Imagers with Recurrent Neural Networks Over the past decade, neural network models, especially CNN and more recently vision transformers [ 128 ], have revolutionized the field of computer vision and image processing. By automatically learning hierarchical representations from raw image data, these models have dramatically improved performance across a wide range of tasks. These models have been deployed across platforms, from edge devices to cloud servers, for tasks such as face recognition and instance segmentation. However, these models are primarily designed for images captured by conventional CMOS or CCD cameras, where each pixel encodes the intensity of the incident light. As a result, neural networks typically operate on static or video frames represented by intensity values. In contrast, event cameras, or neuromorphic cameras, offer a novel imaging paradigm: instead of recording intensity values directly, each pixel responds to changes in intensity [ 29 ]. The output of these cameras is event-driven, consisting of a pixel address, a timestamp, and a binary indication of intensity change. The new paradigm calls for a new way of processing, where SNNs have been extensively studied. The operation and output of SPAD imagers are different from these cameras. In the passive imaging setup, SPADs can be operated in two ways. Blessed by the single-photon sensitivity, one intuitive way is to process the photon arrivals directly using SNNs. In this case, the incoming photons are encoded as rate-coded spikes. As endeavors have been made to transform intensity-based images into Poisson-coded spike trains for efficient neuromorphic processing, these existing models can be readily adapted to process the rate-coded outputs from SPAD imagers. Alternatively, SPADs can count the number of photons within a defined time window, generating frames with intensity values. These frames are compatible with standard deep neural networks, allowing direct application of conventional models. While many established methods can be adapted for SPAD-based passive imaging, the story is notably different for SPAD-based active imaging. In an active time-resolved imaging setup, where the scene is illuminated by a pulsed source and photon arrival times are measured, SPADs can be operated in two distinct modes. In one approach, the spikes generated by incident photons are directly fed into SNNs for processing. Since the information is inherently phase-coded, this method requires specially designed 51
Chapter 4. SPAD Imagers with Recurrent Neural Networks SNNs. Alternatively, TDCs can be used to timestamp photon arrivals relative to a reference signal, producing temporal data for further analysis. Traditionally, these timestamps are histogrammed before processing. In this thesis, a RNN-based approach is proposed to process the raw timestamps directly, which is the focus of this chapter. The SNN-based method is discussed in the following chapter. As the proposed methods serve as general frameworks, fluorescence lifetime estimation, a representative time-resolved imaging task, is primarily used to evaluate their performance. Additional applications, such as near-infrared optical tomography (NIROT), are also considered. The research described in this chapter led in part to the publication in [ 129 ], with further material presented here for completeness. 4.1 RNN for Active Time-Resolved Imaging 4.1.1 Overview Witnessing the remarkable success of neural networks in conventional computer vision tasks, researchers have increasingly sought to integrate these powerful models into the workflow of SPAD time-resolved imaging. For instance, CNNs have been applied to denoise photon-count histograms, while fully connected networks have been explored for estimating fluorescence lifetimes from aggregated photon arrival distributions. However, most current efforts remain confined to post-processing in the software domain, without addressing the unique, event-driven data flow generated by SPAD imagers, and do not attempt to optimize or adapt the photon acquisition workflow itself. These approaches typically rely on constructing histograms from the timestamps collected by SPADs and TDCs, which can be represented as high-dimensional images with a width W , height H , and multiple temporal channels C , or a tensor with a shape of [ W,H,C ]. This representation is perfect input for models such as CNN [ 76 ], but they fail to exploit the sequential nature of photon arrivals that is intrinsic to SPAD-based measurements. RNNs deployed on the edge can not only unleash the power of deep neural networks, but also take advantage of their inherent ability to process sequential data in real time. RNNs can be trained end-to-end on either synthetic or experimental datasets without the need for extensive preprocessing. By incorporating diverse factors such as different noise profiles, photon count rates, and timing jitter into the training data, the resulting models can achieve robust performance across a wide range of SPAD instruments and imaging systems. Unlike conventional models that require aggregating photon events into static histograms, RNNs can directly consume the stream of photon arrival times produced by SPAD sensors, capturing temporal correlations as they unfold. This sequential processing enables adaptive denoising, lifetime estimation, or depth reconstruction on a per-photon basis, paving the way for dynamic optimization of acquisition parameters during imaging. By operating on the edge, RNNs can thus reduce data bandwidth, improve energy efficiency, and enable intelligent, autonomous 52
4.1 RNN for Active Time-Resolved Imaging SPAD imaging systems that respond instantly to changing conditions or sample characteristics, far beyond what histogram-based or frame-based methods can achieve. The principle of an RNN-based time-resolved imaging system is illustrated in Fig. 4.1, where it is compared with the traditional histogram-based approach for fluorescence lifetime imaging. Each RNN unit consists of a group of neurons that receive both the past information, represented by the hidden states hn−1 , and the current input xn . At every time step, the RNN updates the memory of past information hn and generates a prediction yn . While the histogram is being built during the acquisition, the RNN unit processes the timestamp directly and updates the prediction upon the arrival of each photon. This architecture allows the system to continuously integrate temporal information, offering a more efficient alternative to static histogram-based methods. The algorithm is shown below: Real-Time RNN Processing Require: Initial hidden state h0in memory Require: Weight matrices Whi ,Whh,Wyh and biases bh,by 1: n←1 2: while system is running do 3: Wait for new input xt 4: Read previous hidden state hn−1from memory 5: hn←tanh(Whi xn+Whhhn−1+bh) 6: Write updated hidden state htto memory 7: yn←Wyhhn+by 8: Output yn 9: n←n+1 10: end while The hidden state is central to the success of RNNs. Once trained on appropriate datasets, it captures a high-level abstraction of input features tailored to a specific application, effectively summarizing temporal information across sequences. Unlike a histogram, which merely aggregates arrival time data into a general distribution, the hidden state offers significantly higher information density. This compact representation not only preserves essential temporal patterns but also greatly reduces memory requirements, making it a more efficient and scalable solution for applications demanding real-time processing and limited hardware resources. 4.1.2 Discussion on Hardware Implementation of RNN The RNN can be implemented on or near the sensor, for example on an FPGA, to enable efficient, real-time processing and significantly reduce data transfer bandwidth. The data throughput of conventional histogram-based methods and the proposed RNN-based method is shown in Fig. 4.2. Existing imaging systems often transfer all raw timestamps to the computer for post-processing, which leads to a total data transfer load of: Vdata =N×(§log2H¨+§log2W¨+§log2T¨) (4.1) 53
Chapter 4. SPAD Imagers with Recurrent Neural Networks Traditional Methods Our Method Excitation Pulse Emission Probability Incoming Photon LS Fitting CMM ... ... ... histogram Fig. 4.1: Principles of the proposed RNN-coupled TCSPC SPAD system. In a traditional TCSPC FLIM system, the sample is excited by a laser repeatedly, and the emission photons are detected and time-tagged. A histogram is gradually built on these timestamps, from which the lifetime can be extracted after the acquisition is completed. In our proposed system, upon the receiving of a photon, the timestamp is fed into the RNN immediately. The RNN updates the hidden state accordingly and idles for the next photon. The schematic and formula of simple RNN are shown, as in Section 2.3. At timestep n , the RNN takes the current information xn and the past information hn−1 as input, then updates the memory to the current information hn and gives out a prediction yn . Whi , Whh , Whi , bh , and by are the weights and biases to be learned from training. σ ( · ) is the non-linear activation function, which is usually the hyperbolic tangent (tanh). where N is the number of incident photons, H and W are the height and width of the SPAD array, and T is the maximum timestamp. The data transfer load thus scales linearly with the photon count and logarithmically with spatial resolution and timing precision, resulting in potentially massive bandwidth requirements. In practice, this either necessitates a dedicated TCSPC module with a high-speed computer interface or results in data rate limitations that constrain system performance. This can be optimized by performing histogramming on an FPGA placed between the SPAD sensor and the computer, as the FPGA can handle the sensor’s high raw data rates, compress the data through real-time histogramming, and transmit the reduced information to the computer at a significantly lower data transfer rate. The data transfer load is thus reduced to: Vdata =H×W×T×§log2D¨(4.2) where D is the bin depth of the histogram. D ranges from ⌈N/T⌉ , representing the optimistic case where photons are evenly distributed across all bins, to N , the pessimistic case where 54
4.1 RNN for Active Time-Resolved Imaging SPAD Laser pulsed illumination emission reflection or Time-to-digital Converter ( x,y,t ) Histogramming h ( x,y,t ) N×(§log2H¨+§log2W¨+§log2T¨) Data Processing f ( x,y,c ) H×W×T×§log2D¨ Conventional Method H×W×C×E Time-to-digital Converter ( x,y,t ) Recurrent Neural Network f ( x,y,c ) N×(§log2H¨+§log2W¨+§log2T¨) H×W×C×E Our Proposed Method Fig. 4.2: Comparison of data throughput between the conventional method and the proposed RNN-based method. H is the number of rows (height). W is the number of columns (width). T is the number of timesteps per laser repetition period. N is the photon count for the whole sensor. D is the bin depth of the histogram. C is the number of features to be predicted. Eis the bit depth. all photons accumulate in a single bin. It should be noted that the argument only stands when N≫H,W,T , which is the case for most applications. While histograms require the full collection and sorting of the photons into bins according to their arrival time, the RNN updates its hidden state continuously and makes predictions on the fly, directly at the sensor/FPGA level. Therefore, only the final results need to be transferred to the computer, which greatly reduces the data transfer load. The typical data volume in this case is given by: Vdata =H×W×C×E(4.3) where Cis the number of features to be predicted and Eis the bit depth. This approach minimizes the amount of data that needs to be transmitted, lowers latency, and decreases power consumption. By extracting and compressing high-level temporal features in situ, the system achieves faster decision-making and improved scalability, which are essential for applications with stringent resource or speed constraints, such as biomedical imaging or autonomous systems. RNNs can be implemented on hardware, such as chips or FPGAs, primarily using two architectural approaches: the Harvard architecture and the dataflow architecture [ 130 ]. The Harvard architecture features separate memory systems and data paths for instructions and data, allowing simultaneous access to both. In hardware implementations of RNNs, this separation 55
Chapter 4. SPAD Imagers with Recurrent Neural Networks gate biases are initialized to 1, a common practice that encourages information retention during the early stages of training and enhances learning stability. The dataset is randomly split into training, evaluation, and test set, with the ratio of sizes being 8:1:1. The batch size is 32. Adam optimizer is used with an initial learning rate of 0.001 [ 152 ]. The learning rate decays every 5 epochs at the rate of 0.9. The whole training process takes 100 epochs. 4.2.5 Evaluation To evaluate the performance of the fluorescence lifetime estimators, three commonly used regression metrics are adopted: root mean squared error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE). RMSE and MAE are among the most commonly used metrics in regression tasks due to their simplicity, interpretability, and effectiveness. RMSE measures the square root of the average squared differences between predicted values ˆ yiand true values yi: RMSE =1 Nv u u t N X i=1 (yi−ˆ yi)2, (4.6) It penalizes larger errors more heavily, making it sensitive to outliers. MAE, in contrast, computes the average of the absolute differences: MAE =1 N N X i=1|yi−ˆ yi|, (4.7) It treats all errors linearly and provides a more interpretable measure of typical prediction error in the same units as the target variable. MAPE, on the other hand, expresses the error as a percentage of the true value: MAPE =1 N N X i=1¯¯¯¯ yi−ˆ yi yi¯¯¯¯ . (4.8) which makes it scale-independent and easier to interpret across different ranges of lifetimes. This is especially relevant in fluorescence lifetime imaging, where lifetimes can span a wide dynamic range. Unlike RMSE and MAE, which emphasize absolute error, MAPE focuses on relative error, making it more informative in applications where proportional accuracy is more critical than absolute deviation. While the above three metrics provide a practical basis for comparing model performance, they do not indicate how close the estimators are to the theoretical performance limits. To address this, the CRLB is introduced as a benchmark. By comparing empirical errors to the 62
4.2 RNN for Fluorescence Lifetime Estimation Model RMSE MAE MAPE LS Fitting 0.1642 0.1201 0.0553 CMM 0.0915 0.0642 0.0250 Simple RNN-8 0.2516 0.1979 0.0969 Simple RNN-16 0.2396 0.1798 0.0771 Simple RNN-32 0.1877 0.1415 0.0659 GRU-8 0.0957 0.0695 0.0297 GRU-16 0.0928 0.0666 0.0274 GRU-32 0.0908 0.0647 0.0261 LSTM-8 0.0981 0.0720 0.0423 LSTM-16 0.0928 0.0669 0.0277 LSTM-32 0.0916 0.0656 0.0267 Table 4.3: Performance comparison of benchmarks and RNN variants on fluorescence lifetime estimation. RNN models are trained and tested on a synthetic dataset, where the fluorescence decay model is mono-exponential, lifetime ranges from 0.2 and 5 ns, laser repetition frequency is 20 MHz, and background noise is not considered. Their performance is benchmarked against the LS fitting and CMM. CRLB, it becomes possible to assess how efficiently the model extracts information from the data and how much room remains for improvement. The CRLB is calculated with an open-source script [153]. 4.2.6 Results Analysis on Synthetic Data The analysis begins with a simplified scenario in which background noise is absent. All tests are conducted on a computer using 32-bit floating point precision. The fluorescence decay is modeled as mono-exponential, with a fixed Gaussian IRF. All estimations are performed using 1,000 randomly sampled timestamps. The results, summarized in Table 4.3, show that CMM achieves the lowest error in MAE and MAPE, while GRU-32 yields the lowest RMSE. Notably, GRU-32 and LSTM-32 exhibit performance comparable to CMM. The strong performance of CMM in this case is expected. Without background noise and with a repetition period set to ten times the longest lifetime, CMM closely approximates the maximum likelihood estimator. When comparing the three RNN variants, GRU slightly outperforms LSTM, and both significantly outperform the Simple RNN. As model size decreases, a corresponding increase in error is observed, highlighting the trade-off between complexity and accuracy. Background noise is often unavoidable in fluorescence lifetime imaging, particularly in diagnostic and clinical environments where disruption to existing workflows must be minimized 63
Chapter 4. SPAD Imagers with Recurrent Neural Networks 1% Background Noise 5% Background Noise RMSE MAE MAPE RMSE MAE MAPE LS fitting 0.1678 0.1226 0.0562 0.1883 0.1368 0.0609 CMM 0.2367 0.2168 0.1577 1.0742 1.0635 0.7799 CMM†0.1099 0.0839 0.0456 0.2476 0.2128 0.1444 LSTM-32 0.1019 0.0733 0.0304 0.1097 0.0784 0.0323 Table 4.4: Performance comparison of benchmarks and RNN variants on fluorescence lifetime estimation with background noise. Performance of LS fitting, CMM, CMM with background subtraction, and LSTM-32 is compared in the presence of 1% and 5% background noise. LSTM-32 is trained on a dataset including 0% to 10% random background noise for generalization in different scenarios. †CMM with background noise subtraction, assuming that the background noise level is known. [ 62 ]. In the FLIM system under consideration, it is estimated that at least 1% of the detected timestamps originate from background noise. To assess the robustness of each method, performance is evaluated under varying levels of background noise. For simplicity, only LSTM-32 is selected for comparison against benchmark methods. The LSTM-32 model is trained on a synthetic dataset in which 0–10% uniform background noise is randomly added to the samples. Additionally, CMM performance is shown both with and without background noise subtraction. In the corrected case, it is assumed that the number of background photons is known, a condition that is rarely met in real-time FLIM systems. Two synthetic datasets are used for evaluation, with background noise levels set to 1% (SNR = 20 dB) and 5% (SNR = 12.8 dB), respectively. The results are presented in Table 4.4. LSTM32 consistently outperforms the other methods across all metrics and noise levels. When compared with the noise-free results in Table 4.3, it is evident that performance degrades as background noise increases for all methods. However, LSTM and least-squares fitting exhibit greater robustness to noise, while CMM proves to be highly sensitive, an observation that aligns with findings from previous studies [ 51 , 153 ]. In practice, accurate performance of CMM relies on effective background noise subtraction. Cramér-Rao Lower Bound To evaluate performance against the theoretical optimum, the CRLB for lifetime estimation accuracy is computed using an open-source software package [ 153 ], based on the specified parameter settings. The variance of each lifetime estimation method is obtained through Monte Carlo simulations. For the CMM and RNN approaches, 3000 samples are used, while for the least squares method, 1000 samples are used to reduce computation time. Figure 4.3a illustrates the relationship between lifetime and the relative standard deviation of various estimators, with a fixed photon count of 1024. It can be observed that the variances 64
4.2 RNN for Fluorescence Lifetime Estimation of CMM and LSTM-32 closely approach the CRLB, indicating that both are near-optimal estimators. Given that the laser repetition period is much longer than the lifetime and background noise is excluded, it is reasonable that CMM achieves the CRLB, as it approximates a maximum likelihood estimator. In contrast, LS fitting exhibits higher variance, likely due to its assumption of Gaussian-distributed errors. Figure 4.3b shows the relationship between photon count and the relative standard deviation of various estimators, with the lifetime fixed at 2.5 ns. Consistent with the results in Figure 4.3a, the relative standard deviations of CMM and LSTM-32 closely approach the CRLB, while LS fitting performs noticeably worse. These results indicate that CMM and LSTM-32 are efficient estimators across varying photon counts, demonstrating excellent photon efficiency. They require less than half the data to achieve performance comparable to that of LS fitting. We also analyze the CRLB in the presence of background noise, as shown in Figure 4.3. Comparing Figures 4.3c and 4.3e with Figure 4.3a, we observe that the CRLB increases slightly when background noise is introduced. The relative standard deviation of LS fitting remains largely unchanged, while that of LSTM-32 increases modestly but still outperforms LS fitting. In contrast, CMM shows a pronounced increase in relative standard deviation at shorter lifetimes, indicating a strong sensitivity to background noise in this regime. This result is consistent with earlier discussions. Although background subtraction can be applied, CMM still suffers from significantly increased variance under low signal-to-noise conditions. Similarly, from Figures 4.3d and 4.3f compared to Figure 4.3b, we find that 1% background noise has minimal impact. However, with 5% background noise, CMM exhibits clear performance degradation, with its relative standard deviation approaching that of LS fitting. Analysis on Experimental Data To verify the performance of RNNs, which are purely trained on synthetic datasets, on realworld data, the RNNs are tested on experimental data along with CMM and LS fitting as benchmarks. Three samples are prepared and measured with a commercial TCSPC FLIM system, which are the R6G solution, colon tissue stained with H&E, and lifetine-encoded fluorescent beads, as described in Section 4.2.1. The commercial FLIM system is estimated to have a background noise level below 1%. To ensure robustness, the LSTM-32 model is trained on datasets with varying levels of background noise. For each data point, represented as a series of timestamps, a background noise ratio is randomly sampled from a uniform distribution between 0% and 10%. This ratio determines the proportion of timestamps attributed to background noise. In the case of R6G solution and colon tissue stained with H&E, the synthetic dataset assumes a uniform lifetime distribution between 0.2 and 5 ns. For lifetime-encoded fluorescent beads, the lifetime distribution ranges from 1 to 8 ns. Notably, the RNN-based approach leverages prior knowledge of the expected lifetime range, enabling more accurate lifetime estimation. For the LS fitting of lifetime images from colon tissue and fluorescent beads, the histograms are spatially binned by a factor of two 65
Chapter 4. SPAD Imagers with Recurrent Neural Networks =(ns) 0.5 1 1.5 2 2.5 3 3.5 4 4.5 5 <=/= 0.03 0.035 0.04 0.045 0.05 0.055 0.06 0.065 Cramer-RaoLowerBound LeastSquaresFitting CMM LSTM-32 (a) NumberofPhotons 100 200 300 400 500 600 700 800 900 1000 <=/= 0 0.05 0.1 0.15 0.2 0.25 0.3 0.35 Cramer-RaoLowerBound LeastSquaresFitting CMM LSTM-32 (b) =(ns) 0.5 1 1.5 2 2.5 3 3.5 4 4.5 5 <=/= 0.03 0.035 0.04 0.045 0.05 0.055 0.06 0.065 Cramer-RaoLowerBound LeastSquaresFitting CMM LSTM-32 (c) NumberofPhotons 100 200 300 400 500 600 700 800 900 1000 <=/= 0 0.05 0.1 0.15 0.2 0.25 0.3 0.35 Cramer-RaoLowerBound LeastSquaresFitting CMM LSTM-32 (d) =(ns) 0.5 1 1.5 2 2.5 3 3.5 4 4.5 5 <=/= 0.03 0.035 0.04 0.045 0.05 0.055 0.06 0.065 Cramer-RaoLowerBound LeastSquaresFitting CMM LSTM-32 (e) NumberofPhotons 100 200 300 400 500 600 700 800 900 1000 <=/= 0 0.05 0.1 0.15 0.2 0.25 0.3 0.35 Cramer-RaoLowerBound LeastSquaresFitting CMM LSTM-32 (f) Fig. 4.3: Comparison of methods on synthetic datasets against CRLB. (a),(c), and (e) illustrate the relative standard deviation of different estimation methods and CRLB as a function of lifetime for a fixed number of detected photons (1024), including 0, 1%, and 5% background noise, respectively. (b),(d), and (f) illustrate the relative standard deviation of different methods and CRLB as a function of the number of detected photons for a fixed lifetime of 2.5 ns, including 0, 1%, and 5% background noise, respectively. 66
4.3 FPGA Implementation of RNN for Real-Time Processing in both the horizontal and vertical directions to accelerate the fitting process and increase the photon counts per bin. The results are presented in Fig.4.4 using a rainbow scale, in which brightness represents intensity and color encodes lifetime. The corresponding lifetime histograms across all pixels are shown below. Fig.4.4a illustrates the lifetime image of the spatially uniform R6G solution. The LSTM and CMM yield comparable results, whereas the LS fitting produces a noisier outcome, primarily due to the low photon counts. Fig.4.4b presents the lifetime images of colon tissue. Both LSTM and LS fitting provide similar results, with lifetimes centered around 0.24 ns and exhibiting small variance. The LS fitting demonstrates slightly reduced variance, which may be attributed to improved photon statistics resulting from spatial binning. By contrast, the CMM produces a significantly poorer result due to the presence of background. Because the lifetime is short and the background noise is uniformly distributed over a wide range (50 ns), the background strongly influences the variance of the estimate of CMM. Fig.4.4c shows the lifetime images of fluorescent beads. All three methods yield comparable results, and well distinguish the three types of beads encoded with different lifetimes. 4.3 FPGA Implementation of RNN for Real-Time Processing After validating the effectiveness of the RNN for fluorescence lifetime estimation, the next step is to implement the model on an FPGA for integration into a SPAD time-resolved imaging system. This section discusses the model selection process, RNN quantization, high-level synthesis (HLS) implementation, and the integration of the RNN into the overall SPAD imaging pipeline. 4.3.1 Model Selection Targeting at an available Opal Kelly XEM7360-160T (FPGA Development Board with AMDXilinx Kintex 7) FPGA, which is the FPGA used by the SPAD system described in Section 4.3.4, models and parameters are compared to find the most suitable implementation. As discussed in Section. 4.1.1, the simple RNN, while lightweight in terms of parameters and computational complexity, does not perform as well as its more advanced variants, such as GRU and LSTM. GRU and LSTM exhibit comparable performance in lifetime prediction. However, GRU is preferred for hardware implementation because it omits the cell state used in LSTM and requires fewer parameters, MAC operations, and activations, making it more efficient for deployment on resource-constrained platforms. To achieve low-latency, parallelized computation, the hidden size must be minimized while preserving acceptable prediction accuracy. Given the hardware constraints of the XEM7360160T and the system’s latency requirements, hidden sizes ranging from 8 to 32 are evaluated. Additionally, various quantization techniques can reduce hardware resource demands and 67
Chapter 4. SPAD Imagers with Recurrent Neural Networks 0 50 100 150 200 250 0 50 100 150 200 250 LSTM 0 50 100 150 200 250 0 50 100 150 200 250 CMM 0 50 100 150 200 250 0 50 100 150 200 250 LS Fitting 012345 0 500 1000 Count 012345 Lifetime (ns) 0 500 1000 012345 0 250 500 400 450 500 550 Photon Counts 3.5 4.0 4.5 Lifetime (ns) (a) R6G in methonal. 0 100 200 300 400 500 0 100 200 300 400 500 LSTM 0 100 200 300 400 500 0 100 200 300 400 500 CMM 0 50 100 150 200 250 0 50 100 150 200 250 LS Fitting 012345 0 20000 Count 012345 Lifetime (ns) 0 5000 012345 0 10000 200 400 600 800 1000 Photon Counts 0.1 0.2 0.3 0.4 Lifetime (ns) (b) Colon tissue stained with H&E. 0 100 200 300 400 500 0 100 200 300 400 500 LSTM 0 100 200 300 400 500 0 100 200 300 400 500 CMM 0 50 100 150 200 250 0 50 100 150 200 250 LS Fitting 0246 0 1000 2000 Count 0246 Lifetime (ns) 0 1000 2000 0246 0 200 400 100 200 300 400 500 Photon Counts 2 4 6 8 Lifetime (ns) (c) A mixture of fluorescent beads with three different lifetimes (1.7, 2.7, and 5.5 ns). Fig. 4.4: Comparison of LSTM, CMM, and LS Fitting on experimental data. The fluorescence lifetime images are displayed using a rainbow scale, where the brightness represents photon counts and the hue represents lifetimes. 68
4.3 FPGA Implementation of RNN for Real-Time Processing allow for larger hidden sizes within acceptable limits, which will be discussed in the following section. 4.3.2 Quantization Quantization offers an effective means to reduce hardware resource usage and latency. In mainstream deep-learning frameworks such as PyTorch and TensorFlow, model weights and activations are typically stored as 32 - bit floating - point numbers. Performing computations at this precision is often inefficient and unnecessary for low-level inference. Instead, the 32-bit floating-point representations can be replaced with fixed-point formats, and the bit - width reduced as far as possible, while minimizing the loss of the model’s behavior and performance. Both PyTorch and TensorFlow offer quantization toolkits for edge devices, namely PyTorch Quantization and TensorFlow Lite. However, these toolkits rely on their respective runtime libraries, and the quantized weights cannot be easily exported for standalone deployment. To address this, a quantized GRU model is implemented using Python and an open-source fixed-point number library for evaluation. Fixed-point formats with 8-bit, 16-bit, and 32-bit precision are used to quantize weights and activations separately, with various rounding methods. The numbers of bits for integer and fractional parts are chosen based on the distribution of the parameters of the GRU. Monte Carlo simulation is performed for the quantized GRU models to predict the series of timestamps with lifetimes of 0.5, 2.5, and 4,5 ns. The results are shown in Fig. 4.5. Weight quantization reduces the memory for the parameters, while the activations are still calculated with 32-bit floating point precision. The result is shown in Fig. 4.5a, as 16-bit fixed point and 8-bit fixed point are compared against 32-bit floating point. Results indicate that weights can be quantized to 16-bit fixed-point numbers without significant accuracy degradation, and to 8-bit with a significant loss. Activation quantization is employed to further reduce the memory for intermediate result and computational complexity, as the result is shown in Fig. 4.5b. Activations can also be quantized to 32-bit and 16-bit fixed-point precision without major impact, but 8-bit quantization causes model collapse (not shown in the figure). Beyond precision, the choice of rounding method is found to strongly influence performance. Truncation, rounding, and flooring are compared in Fig. 4.5c. Truncation, often the default method, introduces greater errors, whereas convergent rounding yields behavior nearly identical to that of floating-point computation. Activation functions are approximated using piecewise linear functions. Furthermore, quantization-aware training can slightly recover accuracy, although the accuracy loss due to quantization is not significant to begin with. 69
Chapter 4. SPAD Imagers with Recurrent Neural Networks (a) Weight quantization, including 32-bit floating point, signed 32-bit fixed point with 26-bit fractional part, and signed 16-bit fixed point with 10-bit fractional part. (b) Activation quantization, including 32-bit floating point, signed 16-bit fixed point with 10-bit fractional part, and signed 8-bit fixed point with 3-bit fractional part. (c) Activation quantization, including 32-bit floating point, signed 16-bit fixed point with rounding, truncation, and flooring. Fig. 4.5: Violin plot of the prediction of the GRU with different quantization schemes. The quantized GRUs are tested on samples with lifetime of 0.5, 2.5, and 4.5 ns, respectively. 4.3.3 High-Level Synthesis Implememtation The quantized GRU model is then implemented on FPGA. The dataflow architecture is adopted to fully utilize the parallelization of the FPGA. The GRU is written in C++ and compiled to Vivado IP with Vitis High-level Synthesis (HLS). The template is used in the C++ implementation, which makes the adapation of the GRU to other precisions and topologies simple. The computation unit, which is designed to be shared among a group of pixels, is divided into two parts: a GRU core and an FCNN. The GRU updates the hidden state given the new input, and the FCNN predicts the lifetime based on the hidden state. In principle, FCNN can be run upon every update of the hidden state. However, when the real-time update is not required, it only needs to be run once at the end of the acquisition. Due to the low latency of 70
4.3 FPGA Implementation of RNN for Real-Time Processing Module Latency (ns) Interval BRAM (%) DSP (%) FF (%) LUT (%) URAM (%) GRU 1050 169 1 8 3 16 0 FCNN 12500 19991 3 ∼0 4 6 0 Table 4.5: Report from Vitis HLS on the hardware resource utilization of GRU and FCNN. The target FPGA is Opal Kelly XEM7360-K160T (Xilinx Kintex-7, XC7K160T-1). The target FPGA can accommodate 5 GRU cores in parallel at most, and each core can process ∼0.95 Mcps. the FCNN, only one FCNN module is needed, and it will be run sequentially for each pixel after integration. The parameters of the neural networks are stored in the ROM. After the synthesis, the IP blocks of GRU core and FCNN are imported into Vivado and integrated into the FPGA design. Dual-port block RAM (BRAM) IPs are used to store the hidden state for each pixel. Upon receiving a timestamp, GRU core loads hidden states from BRAM, updates the hidden states, and sends them back to BRAM. After the integration of each repetition period, the FCNN loads the hidden state from BRAM, and streams the estimated lifetime to a FIFO. The resource utilization of the GRU core and FCNN are shown in Table. 4.5. They are designed to run at a clock frequency of 160 MHz. Targeting at a period of 6.25 ns, the uncertainty is 1.69 ns. The weights are imported from PyTorch and then quantized to 16-bit fixed point. It is worth noting that the activation functions, Sigmoid function and tanh, are realized with piecewise linear approximation. Both functions are sliced into 32 equally spaced pieces, ranging from -10 to 10. 4.3.4 Integration into SPAD Time-Resolved Imaging System The target SPAD image sensor, internally codenamed Piccolo, is a 32 × 32 SPAD array equipped with 128 shared TDCs [ 154 , 155 ]. Its macrograph is shown in Fig. 4.7. Piccolo is implemented in 180 nm CMOS with a pixel pitch of 28.5 µm and a fill factor of 28%. The design features 128 shared TDCs with a resolution of 50 ps and a dynamic range of 12 bits, enabled by a ring oscillator-based architecture. PDP reaches 47.8% at 520 nm, with acceptable near-infrared sensitivity (e.g., 8.4% at 840 nm), and the median DCR is 113 cps ( ≈ 0.49 cps/ µm2 ) at room temperature. The sensor supports both timestamping and photon-counting modes with maximum throughputs of 222 Mcps and 465 Mcps, respectively, while maintaining a total power consumption of approximately 310 mW. The architecture employs event-driven readout and dynamic TDC allocation via a collision-detection bus, allowing for efficient and scalable operation. The very same FPGA, the XEM7360-160T, serves as the interface between the sensor and the computer, handling both control and communication. Four groups of computation units are implemented on the FPGA alongside the control and communication modules. The schematic of the FPGA design is shown in Figure 4.6. Each computation unit is in charge of a quarter of the SPAD array (32 × 8 pixels). The timestamps, sent to the FPGA in parallel, are serialized and distributed to four computation units based on their SPAD IDs (i.e., row and column 71
Chapter 5. SPAD Imagers with Integrated Spiking Processor R etina V isual C o r t e x S P AD Spi k ing Neu r al Ne t w o r k Fig. 5.1: Conceptual analogy between the biological visual pathway and a SPAD-SNN system. The retina and visual cortex in the brain form a hierarchical, event-driven visual processing pipeline, where light is detected by photoreceptors, passing through the optic nerve, and interpreted by downstream neural circuits in the visual cortex. Similarly, a SPAD array captures incoming photons with high temporal resolution, and the data is processed by a SNN, mimicking the neural coding and efficiency of the biological visual system. environment in real time, much like the human visual system. SNNs not only mimic their biological counterparts and exhibit biological plausibility, but also offer an elegant solution for efficient data processing. Ideally, neuron models such as LIF, one of the simplest yet effective models for simulating spiking neuron behavior, can be implemented using minimalistic hardware [ 159 , 160 ]. The LIF model is typically realized using analog or mixed-signal circuits. In the subthreshold regime, membrane potential integration can be emulated using capacitors, while the leak can be modeled with resistors. Together, they form an RC circuit that enables input integration and exponential decay of the membrane potential. When the voltage crosses a threshold, a comparator triggers a spike and resets the potential, mimicking neuronal firing. A simplified schematic is shown in Fig. 5.2. While accurately controlling large-scale capacitance and resistance is challenging, digital implementations are also appealing. Compared to artificial neurons, spiking neurons do not require MAC operations. Instead, they rely only on addition and comparison, making them more hardware-efficient. This simplicity enables compact and energy-efficient neuromorphic designs, making LIF an attractive choice for edge computing and real-time sensory processing applications. Attempts have been made to explore the integration of SPAD image sensors with SNNs, which are summarized in Table 5.1. Existing work has primarily focused on passive imaging using 78
5.1 Perception in SPAD Sensors via Spiking Neural Networks C R − + VMP S(t) I(t) Vthre Fig. 5.2: Schematic of an analog LIF model. The membrane input I ( t ) is injected into the RC circuit with a time constant RC . The membrane potential VMP is compared with the threshold Vthre . When the membrane potential exceeds the threshold, the neuron fires a spike S ( t ), which further resets the membrane potential. Literature Year Imaging Mode Coding SNN Implementation Application [109] 2023 Passive Rate On-chip digital Image classification [19] 2024 Active Phase to density/ISI Computer FLIM [122] 2024 Passive Rate On-chip digital Image classification [110] 2024 Active Rate Computer Image classification [111] 2024 Active Rate Computer LiDAR [157] 2025 Active Phase to density On-FPGA FLIM [123] 2025 Passive Rate On-chip analog Edge detection [42] 2025 Passive Rate On-chip digital Image classification Table 5.1: Recent studies on SPAD-SNN integration. Most existing studies focus on image classification in a passive SPAD-based imaging setup. SPAD sensors. Under coherent or approximately coherent illumination, photon arrivals follow a Poisson process, and the information is typically represented by the photon arrival rate, also known as the spiking rate. In other words, the incoming photon spike train is encoded using rate coding. Meanwhile, in the AI community, researchers are exploring ANN-to-SNN conversion as a means of enabling energy-efficient processing, with rate coding being the predominant approach. For computer vision tasks, image pixel intensities are converted into Poisson-distributed spike trains, which are then processed by an SNN obtained through the conversion of a pre-trained ANN using established transformation methods. While CMOS image sensors capture photon intensity, which is then artificially converted into random spike trains for SNN-based processing, SPAD sensors naturally produce spike trains, inherently preserving temporal information and even carrying cues about incoherence. As a result, existing SNN models, originally developed for conventional computer vision pipelines, can be directly applied to passive imaging with SPAD sensors. While SPADs have been used for passive imaging under low-light conditions, their primary 79
Chapter 5. SPAD Imagers with Integrated Spiking Processor applications lie in active time-resolved imaging such as FLIM and LiDAR. In these applications, the scene is actively illuminated, typically with periodic laser pulses, and the key information is encoded in the time difference between photon arrival and the laser reference, effectively creating a phase-coded temporal signal. While the spike rate still reflects the intensity of the illumination, it cannot capture time-resolved information such as fluorescence decay lifetime or depth. Therefore, existing SNN models, primarily designed for rate-coded spike trains, are not directly applicable to active time-resolved imaging, and new approaches must be developed to handle temporally encoded data. This thesis mainly focuses on active timeresolved imaging with SNN. An interesting perspective can be added by considering information encoding in biological neural networks, where rate coding and phase coding are two fundamental strategies [ 85 , 161 ]. In rate coding, information is represented by the average firing rate of a neuron over a period of time, higher stimulus intensity typically leads to higher spike rates. This form of coding is relatively robust to noise and it inspires artificial neural networks. Phase coding encodes information in the timing of spikes relative to an oscillatory cycle or reference signal, such as the phase of theta rhythms in the hippocampus [ 162 ]. This temporal precision allows for richer representations, such as encoding spatial position or stimulus features with high fidelity. The brain likely uses a combination of both strategies, depending on the context and computational demands of the task. Despite its biological plausibility and hardware efficiency, integrating SNNs with SPAD image sensors presents several challenges. One major difficulty lies in training SNNs. For simplified neuron models such as the LIF model, spiking is represented by a Dirac delta function, whose non-differentiability prevents gradient flow during BPTT. ANN-to-SNN conversion addresses this issue by leveraging the relationship between artificial and spiking neurons. For instance, a neuron with a ReLU activation can be interpreted as a rate-coded abstraction of an integrateand-fire neuron under constant input. Based on this equivalence, SNNs can be derived from pretrained ANNs, a widely adopted approach for rate-based models. However, this conversion technique is not directly applicable to phase-based models used in active time-resolved imaging, where precise timing information is essential. Another challenge lies in the hardware implementation of spiking neural networks. While a spiking neuron can be realized using an RC circuit coupled with a comparator, accurately controlling the resistance and capacitance remains difficult. Some analog implementations of spiking neurons have been demonstrated, but they typically suffer from limited precision and performance [ 113 , 160 , 163 ]. Digital implementations are more common due to their predictable behavior and ease of integration, although they tend to incur higher hardware utilization and increased latency. Real-time processing of spikes generated by SPAD is also challenging when high temporal resolution is required. 80
5.2 Problem Formulation reference detection (a) reference OR detection (b) Fig. 5.3: Intuitive input coding for single-photon detectors. (a) The SNN has two input nodes, which take the reference signal and the detection signal as input separately. (b) The SNN has only one input node, which takes the sum of the reference and detection signals as input. 5.2 Problem Formulation Active SPAD imagers naturally generate phase-coded spike trains, making them well-suited as inputs for SNNs. In contrast to event cameras or passive SPAD imagers, which primarily rely on rate-based spike encoding, phase-coded spikes pose greater challenges for SNN processing. While encoding methods and preprocessing schemes can be freely chosen in software, such flexibility is significantly constrained in hardware. Therefore, it is essential to consider hardware feasibility when selecting encoding strategies and preprocessing techniques. A straightforward approach is to feed the reference and detection signals directly into the SNN, either separately or as their sum, as illustrated in Figure 5.3. Notably, the summed input is a special case of the separate-input scheme, as the OR gate can be modeled as a spiking neuron with identical synaptic weights for both the reference and detection signals. However, this approach introduces two key challenges: ultra-long spike sequences and significant variance in spike density. In active time-resolved SPAD imaging, achieving higher temporal resolution is always desirable to improve precision, which leads to an increased number of timesteps within each repetition period. Additionally, hundreds to thousands of photons are typically required to construct a complete event image, so the total number of timesteps can easily reach the order of millions. In practice, the photon rate is kept low relative to the repetition frequency to avoid the pile-up effect. As a result, many repetition periods contain no photon detections, introducing numerous “blank” inputs that further extend the sequence length by several orders of magnitude. The number of total timesteps is given by: N=NT×NC φ, (5.1) where NT is the number of timesteps in one repetition period, NC is the number of desired 81
Chapter 5. SPAD Imagers with Integrated Spiking Processor photons, and φ is the ratio between photon rate and repetition frequency. The ultra-long sequence results in inefficiency and excessive computation during the inference and makes it impossible to train the neural network through BPTT due to the gradient exploding/diminishing problem. The probability of photon reception within a single repetition period varies across pixels and scenarios, leading to variations in the number of “blank” inputs. When using spiking neuron models with temporal dynamics such as membrane potential decay and input decay, it would be challenging for the SNN to handle inputs with such a high temporal dynamic range. In order to implement the SNN for active time-resolved SPAD imaging on hardware, it is essential to employ a hardware-feasible encoder to convert phase-coded spike trains from SPADs to denser and more informative ones, where the “blank” inputs are eliminated and the encoded information is easier to learn. Tailored training techniques are also required for supervised learning on regression tasks. 5.3 Proposed SNN Architecture This thesis proposes two SNN architectures: the Transporter SNN and the Reverse Start-Stop (RS) SNN. Special consideration is given to hardware implementations, particularly digital designs. Dedicated circuits are devised to convert phase-coded spike trains into density-coded and ISI-coded representations, respectively, enabling more efficient learning and processing. Examples of these spike trains are illustrated in Fig. 5.4, where the spikes are generated from fluorescence emission. 5.3.1 Spiking Neuron Model Spiking neuron models come in a diverse “zoo”, ranging from simple abstractions to biologically detailed simulations, each offering different trade-offs between computational efficiency and biological realism. At the most basic level, models like the perfect integrate-and-fire model and LIF model approximate spiking behavior using simple differential equations, making them popular in neuromorphic hardware due to their low computational cost. More advanced models, such as the Izhikevich and HH neurons, capture a wider range of neuronal dynamics, including bursting and adaptation, at the expense of greater complexity [ 85 ]. Other variants, like the exponential integrate-and-fire and adaptive exponential integrate-and-fire models, extend the basic LIF framework by incorporating an exponential term to better capture the rapid onset of spikes, with adaptive exponential integrate-and-fire further introducing spike-frequency adaptation to model more complex and biologically realistic firing patterns. Among various spiking neuron models, the LIF model is adopted to keep the balance of the trade-off between the simplicity and capacity of the neural network. Considering that subtraction is easier than multiplication in hardware implementation, we assume that the 82
5.3 Proposed SNN Architecture (a) Examples of encoding of Transporter SNNs. There are 256 repetition periods and 1563 timesteps for each repetition period. (b) Examples of encoding of RS SNNs. There are 128 repetition periods and 1000 timesteps for each repetition period. Fig. 5.4: Examples of the encoding of the proposed SNNs. Spike trains for fluorescence TCSPC data with different lifetimes are shown here. The LSB is 0.016 ns and no background noise is considered. The shift and FWHM of the IRF are 1.968 ns and 0.1673 ns, respectively. For better visualization, different numbers of repetition periods and timesteps are used. . membrane potential decays linearly instead of exponentially. The membrane potential U of neuron iin layer lat timestep nwith soft reset is given by: U(l) i[n]=U(l) i[n−1]−τdecay odecay −VthreS(l) i[n−1] oreset +PNl−1 j=0w(l−1) jS(l−1) j[n]oinput (5.2) where τdecay is the decay constant, Vthre is the firing threshold, Nl is the number of neurons in the layer l , w is the synapse weight, and S(l) i [ n ] is the output of neuron i in layer l , which fires 83
Chapter 5. SPAD Imagers with Integrated Spiking Processor Vbias VOP (a) − + wi ti j OUT RST IN Potential Memory τdecay Vthre - (b) Fig. 5.5: Schematic of passive-quenching SPAD and simple leaky integrate-and-fire spiking neuron model. (a) The SPAD with PQPR generates a pulse upon detection of photons, which can be directly sent to SNN for processing. (b) Binary output ti j from neuron j in layer i is multiplied by the synapitic weight wi , which can be replaced by a look-up table to avoid MAC operation. The weighted sum is integrated into the membrane potential and undergoes leakage τdecay . The resuting membrane potential is further compared to the threshold Vthre . When the membrane potential exceeds the threshold, the neuron fires and the membrane potential is reset. when the membrane potentialU(l) i[n] exceeds the threshold Vthre: S(l) i[n]=(1, 0, if U(l) i[n]≥Vthre otherwise (5.3) When the hard reset is adopted, the membrane potential U(l) i is simply reset to 0 when it exceeds the threshold. Hard reset is used for Transporter SNN, and soft reset is used for RS SNN. It is also possible to realize the decay by halving, which can be easily implemented on hardware by shifting. Thus the Eq. 5.2 can be adapted to: U(l) i[n]=1 2U(l) i[n−1] odecay −Vthre S(l) i[n−1] oreset +PNl−1 j=0w(l−1) jS(l−1) j[n]oinput (5.4) The simplified LIF model requires minimal hardware resources to implement. The necessary electronic building blocks are two adders for integration and decay, one comparator for firing detection, and one memory for membrane potential storage. The schematic is shown in Figure 5.5. The simple structure allows it to be implemented massively on the sensor [164]. The spiking neurons and networks are built with SpikingJelly 1 , an open-source deep learning 1https://github.com/fangwei123456/spikingjelly/tree/master 84
5.3 Proposed SNN Architecture Ring Oscillator ... Injection Device detection Readout Device time density Fig. 5.6: The proposed Transporter SNN model. The detected spikes are injected into an RO, whose period equals the repetition period of the laser. When the RO is synchronized with the laser, the incoming arrival times are folded into one spike sequence, which is read out after the acquisition. The density of sequence is basically proportional to the histogram, which is an ideal input for SNNs. frameworkfor SNN based on PyTorch. To realize the simplified LIF model, the perfect integrateand-fire model neuron.IFNode in SpikingJelly is used, and a bias is added along with synaptic weights, serving as a constant current into the IF model. It is worth noting that the bias can be either positive or negative. 5.3.2 Transporter SNN After Captain Montgomery Scott experienced a crash at Dyson sphere, he rigged the Transporter and left himself in the diagnostic loop indefinitely until a rescue vessel could come and “rematerialize” him (Star Trek S6E4). Inspired by this concept of “Transporter suspension”, the Transporter SNN is proposed, forming a loop that stores temporal information directly in time rather than in memory, deferring processing until rematerialization. The schematic of Transporter SNN is illustrated in Figure 5.6. The spike generated by a SPAD is injected into a ring oscillator (RO), with delay elements of, if any, a few ps. The injected spikes would start circulating indefinitely until “rematerialization” by a readout circuit. The RO is synchronized with the reference signal, keeping the same period. After running for thousands of periods until the required information is collected, the spikes are read out sequentially, where the differences among arrival times are maintained. The resulting sequence is exactly the binarized sum of all repetition periods, thus the density of the spikes through time is basically the histogram of arrival times. The number of timesteps is reduced by several orders of magnitude, significantly improving the efficiency and facilitating the training. The length of the sequence is reduced to hundreds 85
Chapter 5. SPAD Imagers with Integrated Spiking Processor SPAD Laser pulsed illumination emission reflection or Time-to-digital Converter ( x,y,t ) Histogramming h ( x,y,t ) N×(§log2H¨+§log2W¨+§log2T¨) Data Processing f ( x,y,c ) H×W×T×§log2D¨ Conventional Method H×W×C×E Spike Encoder p ( x,y,t ) = 1 h(x,y,t)>0 Spiking Neural Network f ( x,y,c ) H×W×T H×W×C×E Our Proposed Method Fig. 5.7: Comparison of data throughput between the conventional method and the proposed SNN-based method. H is the number of rows (height). W is the number of columns (width). T is the number of timesteps per laser repetition period. N is the photon count for the whole sensor. D is the bin depth of the histogram. C is the number of features to be predicted. Eis the bit depth. or thousands, which makes the training through BPTT possible. Since the delay element can only store binary information, several spikes falling into the same delay element can cause information loss and distortion. Thus one has to balance the trade-off between hardware implementation difficulty and information fidelity in practice. With such a spike encoder, the data transmission can be largely reduced compared to the conventional method, as shown in Fig. 5.7. Compared with the RNN-based method, the data transmission is further reduced from N×(§log2H¨+§log2W¨+§log2T¨) to H×W×T. Training Transporter SNN is trained with Surrogate Gradient [ 98 , 165 ]. With the folding of the time domain by the RO, millions of timesteps of the sequence are compressed into thousands to tens of thousands of timesteps, which makes it possible to train the SNN with BPTT. Surrogate Gradient smoothes the SNN and helps the gradients “flow” backward during BPTT. The firing function (Eq. 5.3) is non-differentiable at U=Vthre and has derivatives of 0s elsewhere, which makes it impossible to use BPTT directly. To overcome this, the nondifferentiable firing function is replaced by a differentiable surrogate function, e.g. Sigmoid 86
5.3 Proposed SNN Architecture and arctan functions, in the backward path. Arctan is used here for the training: S(U)=1 πarctan³π 2α(U−Vthre)´+1 2(5.5) where αis a scaling factor. Its derivative is given by: ∂S ∂U=α 2³1+¡π 2(U−Vthre)¢2´(5.6) which is everywhere defined, continuous, and non-zero. ∂S / ∂U reaches the maximum at U=Vthre. Hardware Implementation Two types of hardware implementations are envisioned for the on-chip spike encoder: the delay-line-based Transporter and the ring-based Transporter. The delay-line-based Transporter employs a delay line arranged similarly to a TDC. A pulse triggered by photon arrival propagates through the delay line and is registered upon the arrival of a reference signal. In contrast, the ring-based Transporter consists of a ring of memory units driven by a common high-speed clock. The encoded spike circulates continuously within the ring until it is read out. Fig. 5.8 and Fig. 5.9 show the schematics of two delay-line-based Transporter designs, implemented using latches and DRAM cells, respectively. In both architectures, the photon arrival signal propagates through a delay line composed of buffers. A reference signal, typically the laser synchronization signal, is connected to the gates of all NMOS transistors, which act as switches. In the latch implementation, when both the source and gate of an NMOS transistor are high, the photon is registered, setting the corresponding latch to 1. All the Qs can be readout through a data bus. In the DRAM implementation, when both NMOS transistors are switched on, a high voltage on the W/R line is stored in the capacitor. The main difference between a delay-line-based spike encoder and a TDC lies in their functionality. The spike encoder can collect and store multiple photon events, whereas the TDC is designed to capture the arrival time of a single photon. Additionally, the TDC determines the photon arrival time using a multiplexer to select a single output, while the spike encoder reads out all bits to represent multiple events. The spike encoder can also be implemented with a ring composed of memory units such as D-type flip-flop (DFF) and CCD. A DFF ring-based Transporter system is illustrated in Fig. 5.10. The spike encoder is a ring composed of DFFs, such as a ring counter. All the DFFs are driven by a common high-speed clock, which is synchronized with the laser. The frequency must guarantee that the pulse returns to the same DFF after the propagation within the ring after one laser period. The photon arrival signal from the SPAD is inserted into the ring with an 87
Chapter 5. SPAD Imagers with Integrated Spiking Processor 5.4 Transporter: A 128 × 4 SPAD Sensor with On-Chip Preprocessing for SNN Based on the idea of Transporter SNN, a proof-of-concept SPAD image sensor with an in-sensor spike encoder, Transporter, is designed using 110 nm CMOS technology. The chip has 128 × 4 SPADs, each of them coupling to a 128 DFF ring-based spike encoder. The SNN is envisioned to be deployed on the FPGA, which also controls and communicates with the Transporter. As the chip is still under manufacturing, the chip design and post-layout simulation are reported in this section. 5.4.1 Chip Design Transporter is composed of a clock module, 128 rows of 1 × 4 SPAD-Counter-Ring, and a readout module, as shown in Fig 5.15. The clock module, located on the top, provides and distributes the clock to all the rings for operation and readout. For each row of 1 × 4 SPADCounter-Ring, the 1 × 4 SPADs are located in the middle, while two PQPR circuits and rings are placed on either side. The rings can be organized into four columns, where a tri-state buffer bus goes through each of them and the data are read out by the readout module at the bottom. A screenshot of the layout is shown in Fig. 5.14, and an abstract schematic is shown in Fig 5.15. Clock Module The clock module is designed to supply either a high-speed clock to the ring, delivering exactly 128 pulses per laser period, or a low-speed clock for readout. A multiplexer is used to switch between an external clock and an internal clock. The internal clock is implemented by a 3stage NAND-based ring oscillator, which generates the high-speed clock for ring operation. A separate VDD is supplied to improve the stability of the ring oscillator. The clock is distributed to the four columns of rings with a clock buffer H-tree, and further goes to the 128 rings. Ensuring exactly 128 pulses per laser period is crucial for maintaining data integrity within the ring. Any deviation, whether an excess or shortage of pulses, would result in the loss of relative timing information between spikes. To address this, a gating mechanism is implemented to enforce the 128-pulse constraint. This mechanism utilizes a 7-bit counter, composed of a 2-bit synchronous counter and a 5-bit asynchronous counter, designed to operate at gigahertz frequencies. The schmematic is shown in Fig. 5.16. The combination of a synchronous and an asynchronous counter enables counting a GHz clock, which neither a synchronous nor an asynchronous counter alone can achieve. The 2-bit synchronous counter captures rapid changes in the two least significant bits. In combination with logic circuits, it enables fine adjustments to the timing of the gating signal, controlled by SHIFT[0] and SHIFT[1], which are specifically designed to compensate for the delay of the GATING signal. Additionally, a delay unit for the GATING signal, controlled by DELAY[0:1], is implemented to compensate for delays present in other signal paths. A clock gate masks the clock upon receiving the GATING 94
5.4 Transporter: A 128 ×4 SPAD Sensor with On-Chip Preprocessing for SNN Spike Encoder 12-bit Counter Control Unit Readout SPAD and PQPR Clock Module Fig. 5.14: Design layout of the proposed Transporter SPAD image sensor. The chip features a symmetrical design with the SPAD array positioned at the center. Flanking both sides of the SPADs are the PQPR circuits. On either side, there are two groups consisting of a control unit, counter, and spike encoder, each group responsible for one column of SPADs. A clock module is located at the top, providing the clock signal to drive the spike encoders. At the bottom, four readout circuits are arranged, each handling the output from one group of spike encoders and counters. 95
Chapter 5. SPAD Imagers with Integrated Spiking Processor 1×4 SPAD-Counter-Ring Clock Generator & Clock Tree 1×4 SPAD Array ... Counter Readout & Ring Readout 1×2 PQPR 12-bit Counter 128 D Flipflop Ring 12-bit Counter 128 D Flipflop Ring 1×2 PQPR 12-bit Counter 128 D Flipflop Ring 12-bit Counter 128 D Flipflop Ring Fig. 5.15: High-level schematic of Transporter.Transporter is composed of a clock module, 128 rows of 1×4 SPAD-Counter-Ring, and a readout module. CLK Q D Q CLK Q D Q CLK Q D Q CLK Q D Q CLK Q D Q CLK Q D Q CLK Q D Q CLK CLK SHIFT[0] SHIFT[1] GATING Fig. 5.16: Clock gating module of Transporter.It ensures that 128 cycles in one laser repetition period. The clock of interest is fed into CLK. SHIFT[0] and SHIFT[1] are designed to compensate for the delay of the GATING, which is sent when 124, 125, 126, or 127 pulses are detected. GATING is sent to a clock gate to mask the CLK. signal and keeps it masked until the next laser synchronization signal arrives, at which point it is re-enabled. SPAD and Pixel Circuits The SPAD features φ 10 µ m round active area. A PQPR circuit, composed of two NMOS transistors, is employed for each SPAD. The photon detection probability and the dead time can be tuned by adjusting the gate voltage of these transistors. Two inverters are cascaded at the output to digitize the pulse. The pixel pitch is restricted by the height of the Transporter ring, which is 25 µ m in our implementation. The horizontal pitch is set equal to the vertical 96
5.4 Transporter: A 128 ×4 SPAD Sensor with On-Chip Preprocessing for SNN pitch to maintain isotropic pixel spacing, though it can be reduced to 15 µ m. As a result, the fill factor is 12.6%. Transporter Spike Encoder and Counter The Transporter spike encoder functions as the spike encoder shown in Fig. 5.7. In this work, DFFs are adopted as the delay element. A simplified schematic is illustrated in Fig. 5.10. 128 DFFs are cascaded like a circular shift register, except that an OR gate is inserted between the two ends, with its second input connected to the SPAD output for spike injection. The DFFs are arranged into two rows, with clock buffer H-tree placed on the top and bottom. As mentioned earlier, the accurate clock timing for the DFFs is crucial to the integrity of the spikes. Considering the setup and hold time at four corners, the maximum operation frequency is 1 GHz. A 12-bit ripple counter is implemented to count the incoming photons, but its role extends beyond simple counting. In a ring-based delay structure, the number of delay elements is limited, and each element can store only one spike. When the photon count is too high, the ring becomes saturated, leading to a loss of temporal information. This saturation is common in time-resolved applications due to varying light intensities across pixels. To mitigate this, a spike gate is introduced to block further spike injection once a predefined number of photons have been collected. In this proof-of-concept implementation, only 128 DFFs are available. As a result, the control units can configure the system to operate in three modes: 128-count cutoff, 256-count cutoff, or free-running. Readout Circuits and Control Units The readout circuits and control units are designed to deliver control signals to each row and to retrieve data from both the encoder and counter. A high-speed H-buffer clock tree is implemented to distribute the clock signal uniformly across all rows with minimal skew. Regular control signals are delivered through cascaded buffers. For readout, a 1-bit tri-state buffer line is used to sequentially access the encoder output bit by bit, while a 12-bit tri-state buffer bus allows for parallel readout of the counter. Both outputs are subsequently shifted off-chip, one bit at a time, for each group. An address selector is employed to choose the appropriate row for readout. The address selector is not implemented in the common bus-decoder structure. Instead, it is a ladder-style logic array that generates 7-bit output words across 128 rows, as shown in Fig. 5.17. Each row transforms the previous row’s output using a combination of buffers and inverters, selected per bit based on the bitwise difference between the current row’s address and the previous row’s address. If a bit changes, an inverter is used; if it remains the same, a buffer is applied. This results in a cumulative XOR pattern, where each row’s output represents the XOR of the initial input and the total address transitions from row 0 to the current row. To 97
Chapter 5. SPAD Imagers with Integrated Spiking Processor 0000001100110177 ADDR_SEL[77] 0000011100111078 ADDR_SEL[78] 0000001100111179 ADDR_SEL[79] XOR[n,n−1]BinaryRow ... ... ... ... ... ... ... ... ADDR[6] ... ADDR[5] ... ADDR[4] ... ADDR[3] ... ADDR[2] ... ADDR[1] ... ADDR[0] Fig. 5.17: Address selector of Transporter.The address selector is implemented as a “ladder” of butter/inverter. determine when a row’s output matches a target condition, a 7-input NOR gate is applied to the row’s output bits. The NOR output serves as a row selection signal, which goes high only when all output bits are zero. This mechanism eliminates the need for a conventional address decoder by using address transitions and NOR-based zero detection to implicitly select rows, and it shares conceptual similarity with Gray code logic in its use of controlled bit transitions. Test Structure Two test structures are placed on each side. The counter and the spike encoder are connected to PADs for direct readout. Pad Ring A total of 204 PADs are utilized in this chip, with 51 PADs allocated per edge. Since the control units are arranged vertically, most data input/output PADs are placed on the top and bottom edges. The left and right edges are primarily reserved for test structures and power supply connections (VDD/GND). Excluding test structures, 35 PADs are dedicated to data input and output: 29 GPINs for control signals, 4 GPIOs for data readout, and 2 analog I/Os for the external clock and the ring oscillator output. Additionally, 4 PADs each are allocated for VOP , Vq , and Vcascode , supporting SPAD operation; these are evenly distributed between the top and 98
5.4 Transporter: A 128 ×4 SPAD Sensor with On-Chip Preprocessing for SNN bottom edges. Each edge contains 4 VDD and GND PADs to ensure even power distribution. An additional VDD is dedicated to the ring oscillator to enhance its stability. 5.4.2 Post-Layout Simulation The chip is still in the manufacturing process at the time of writing and will be tested in the near future. In the following subsections, the post-layout simulation results of the chip are shown, as well as the simulation of the SNN on the computer. Clock Module In the post-layout simulation, the ring oscillator generates a clock at 1.068 GHz (ss: 0.77 GHz, ff: 1.346 GHz). The simulation result of the gated clock is shown in Fig. 5.18a. One can observe the clock is disabled when the pulse count reaches 128 in every laser period (100 ns). SPAD and Pixel Circuits Given the operation voltage VOP = 25V, breakdown voltage VB= 20V, cascode voltage Vcascode = 3 . 3V, the quenching voltage Vq is swept from 100 mV to 1 V with 10 equal steps. The result is shown in Fig. 5.18b. By adjusting Vq , the dead time can be tuned from a few nanoseconds to hundreds of nanoseconds. Transporter Spike Encoder The simulation result of the test structure for Transporter ring is shown here. The collection of photons takes 2 µ s, given a ring clock of 1 GHz. The readout of the spikes takes 1 µ s, given a readout clock of 200 MHz. The laser period is 128 ns. With a simulated photon arrival with a period of 130 ns, the ring is supposed to receive the photons every two stages. The result is shown in Fig.5.18c. During the accumulation period, 13 spikes are inserted into the ring, which are read out through the tri-state buffer bus. One can observe that the relative timing information is well preserved within the ring. 5.4.3 Spiking Neural Networks for FLIM Based on the specifications of Transporter, an SNN is designed, trained, and evaluated for fluorescence lifetime estimation. The SNN consists of a single-node input layer, a 512-neuron hidden layer, and a single-node output layer. Training is performed using BPTT with surrogate gradients. Since the Transporter has a temporal resolution of ∼ 1 ns, it imposes a limitation on the achievable lifetime resolution. This limitation can be lifted by employing the delay-line-based 99
Chapter 5. SPAD Imagers with Integrated Spiking Processor 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 Time (s) 1e 7 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Voltage (V) (a) The output of the clock module with the internal ring oscillator and gating. 0.00 0.25 0.50 0.75 1.00 1.25 1.50 1.75 2.00 Time (s) 1e 7 0 1 2 3 4 5 Voltage (V) v_q=0.1 v_q=0.2 v_q=0.3 v_q=0.4 v_q=0.5 v_q=0.6 v_q=0.7 v_q=0.8 v_q=0.9 v_q=1 SPAD_ANODE V_OUT (b) The anode and output of a SPAD with different Vq . The solid lines represent the voltage at the anode of the SPAD, and the dash lines represent the voltage digitized by inverters. Different colors represent different Vq. 2.0 2.1 2.2 2.3 2.4 2.5 2.6 Time (s) 1e 6 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Voltage (V) (c) Readout of the spikes in the ring. The readout clock is 200MHz, so it takes 6.4 µ s to read out 128 DFFs. There are 13 spikes injected into the ring every two delay elements. Fig. 5.18: Post-layout simulation results of Transporter. 100
5.5 Summary (a) No background noise (b) 10% background noise Fig. 5.19: Simulation results of spiking neural networks for fluorescence lifetime estimation. The parameters correspond to the ones in the Transporter, i.e., timestep of 128 and clock period of 1 ns. The lifetime ranges from 5 to 20 ns. 256 photons are collected for each inference. architecture. For the training and evaluation dataset, it is assumed that lifetime ranges from 5 to 20 ns. For each data entry, 256 timestamps are histogrammed and binarized as input for the SNN. Scenarios with no background noise and 10% noise are considered. The result is shown in Fig. 5.19. The mean average percentage errors are 6.60% and 9.29%, respectively. 5.5 Summary This chapter explores the integration of SPAD sensors with SNNs to enable intelligent, eventdriven processing of time-resolved photon data. Motivated by the sparse and temporally encoded nature of SPAD outputs, this work investigates how SNNs can be leveraged to process photon arrival events in a biologically inspired and hardware-efficient manner. This chapter begins with the general concept of coupling SNNs to SPAD sensors, inspired by the human visual system, which efficiently processes sparse and asynchronous stimuli. Drawing parallels between biological retinal circuits and the temporal characteristics of SPAD outputs, the system architecture is designed to mimic early visual processing by encoding photon arrivals into spike trains. The distinction between passive imaging and active timeresolved imaging is emphasized, with particular attention given to the unique challenges posed by the latter. To ground the exploration, the chapter provides an in-depth discussion of the challenges involved in processing and learning from the ultra-long, sparse spike trains generated by SPAD sensors in time-resolved imaging. These sequences, characterized by high temporal resolution 101
Chapter 5. SPAD Imagers with Integrated Spiking Processor and low event density, pose significant difficulties for direct SNN processing frameworks. The chapter further formulates the core problem: how to convert asynchronous photon arrival events into meaningful neural representations in a manner that is both computationally efficient and hardware-feasible, while enabling effective inference and training. This formulation serves as the foundation for developing bio-inspired, spike-based processing pipelines tailored to SPAD-based time-resolved imaging. To address the challenges posed by the ultra-sparse and high-temporal-resolution data from SPAD sensors, this chapter proposes two SNN architectures tailored for time-resolved imaging. These architectures incorporate two hardware-feasible spike encoders that convert the raw, phase-coded spike trains from SPADs into density-coded and ISI-coded spike trains, which are more efficient for neural processing. Specific training methods are developed to adapt these architectures to the characteristics of SPAD data. The networks are trained for fluorescence lifetime imaging tasks and demonstrate performance comparable to conventional benchmarks. Building on the proposed architecture, a proof-of-concept chip, Transporter, is designed to implement the spike encoder adjacent to the SPAD, demonstrating the feasibility of hardwarebased conversion from phase-coded SPAD outputs to density-coded spike trains. This encoder enables efficient preprocessing of photon event streams, translating time-resolved data into a format suitable for downstream neural processing. Post-layout simulations confirm its ability to reliably store relative temporal information, while additional software simulations show that the encoder’s output is both sufficient and efficient for SNN-based fluorescence lifetime estimation. 102
6Conclusion and Outlook 6.1 Conclusion This thesis presents a unified approach to photon-efficient, low-latency, and intelligent timeresolved imaging by integrating innovations in SPAD-based sensing, computational modeling, and edge artificial intelligence. Through a series of contributions across theory, algorithm design, and hardware prototyping, this work demonstrates the feasibility of bringing intelligence closer to the sensor. The result is an imaging system that not only collects data but also begins to interpret it directly at the point of acquisition. To approach this novel imaging system, two distinct yet complementary directions have been explored, each rooted in a different source of inspiration. Inspired by the success of ANNs across a wide range of applications, this work investigates the potential of embedding such models directly within the SPAD imaging pipeline. By applying deep architectures, particularly recurrent neural networks, to raw photon timing data, it becomes possible to perform lifetime estimation and image reconstruction in real time. This approach enables close coupling between data acquisition and interpretation, offering improvements in both speed and robustness for SPAD-based imaging. The second direction is rooted in the structure of biological vision and the natural compatibility between SPAD data and SNNs. SPAD sensors produce sparse, asynchronous data streams, which align closely with the event-based nature of SNNs. This synergy allows for minimalistic hardware implementations that are both highly efficient and biologically plausible. This work demonstrates that a simplified, fully digital SNN can perform fluorescence lifetime estimation with precision comparable to established benchmarks. To enable this, spike encoders are introduced to bridge the SPADs and the SNN, significantly reducing data flow and easing the computational burden of implementation. Furthermore, the first SPAD image sensor with an on-chip spike encoder is developed, enabling future integration with neuromorphic processors for end-to-end, sensor-level intelligence. This direction highlights the advantages of SNNs in terms of minimal hardware complexity, low power consumption, and biological plausibility, making them well-suited for compact, efficient imaging systems. 103
Bibliography [1] Takao Kuroda. Essential principles of image sensors. CRC press, 2017. [2] Franco Zappa, Simone Tisa, Alberto Tosi, and Sergio D. Cova. Principles and features of single-photon avalanche diode arrays. Sensors and Actuators A: Physical, 140(1):103–112, October 2007. [3] David Dussault and Paul Hoess. Noise performance comparison of ICCD with CCD and EMCCD cameras. In Infrared Systems and Photoelectronic Technology, volume 5563, pages 195–204. SPIE, October 2004. [4] Matthew D Eisaman, Jingyun Fan, Alan Migdall, and Sergey V Polyakov. Invited review article: Single-photon sources and detectors. Review of scientific instruments, 82(7), 2011. [5] Danilo Bronzi, Federica Villa, Simone Tisa, Alberto Tosi, and Franco Zappa. SPAD figures of merit for photon-counting, photon-timing, and imaging applications: A review. IEEE Sensors Journal, 16(1):3–12, 2016. [6] Claudio Bruschini, Harald Homulle, Ivan Michel Antolovic, Samuel Burri, and Edoardo Charbon. Single-photon avalanche diode imagers in biophotonics: review and outlook. Light: Science and Applications, 8(1):87, 2019. [7] Darek P. Palubiak and M. Jamal Deen. CMOS SPADs: Design issues and research challenges for detectors, circuits, and arrays. IEEE Journal of Selected Topics in Quantum Electronics, 20(6):409–426, November 2014. [8] Arin C. Ulku, Claudio Bruschini, Ivan Michel Antolovic, Edoardo Charbon, Yung Kuo, Rinat Ankri, Shimon Weiss, and Xavier Michalet. A 512 × 512 SPAD image sensor with integrated gating for widefield FLIM. IEEE Journal of Selected Topics in Quantum Electronics, 25(1), 2019. [9] Cristiano Niclass, Claudio Favi, Theo Kluter, Marek Gersbach, and Edoardo Charbon. A 128 × 128 single-photon image sensor with column-level 10-bit time-to-digital converter array. IEEE Journal of Solid-State Circuits, 43(12):2977–2989, December 2008. [10] Robert H. Hadfield, Jonathan Leach, Fiona Fleming, Douglas J. Paul, Chee Hing Tan, Jo Shien Ng, Robert K. Henderson, and Gerald S. Buller. Single-photon detection for long-range imaging and sensing. Optica, 10(9):1124–1141, September 2023. [11] Carsten Pitsch, Dominik Walter, Leonardo Gasparini, Helge Bürsing, and Marc Eichhorn. 3D quantum ghost imaging. Applied Optics, 62(23):6275–6281, August 2023. 111
Bibliography [12] Audrey Eshun, Dominique Davenport, Brandon Demory, Shervin Kiannejad, Paul Mos, Yang Lin, Tiziana Bond, Michael C. Rushford, Erin E. Nuccio, Ty J. Samo, Peter K. Weber, Claudio Bruschini, Edoardo Charbon, and Ted A. Laurence. 3D quantum ghost imaging microscope. Optica, 12(7):1109–1112, July 2025. [13] G. Naletto, C. Barbieri, T. Occhipinti, I. Capraro, A. Di Paola, C. Facchinetti, E. Verroi, P. Zoccarato, G. Anzolin, M. Belluso, S. Billotta, P. Bolli, G. Bonanno, V. Da Deppo, S. Fornasier, C. Germanà, E. Giro, S. Marchi, F. Messina, C. Pernechele, F. Tamburini, M. Zaccariotto, and L. Zampieri. Iqueye, a single photon-counting photometer applied to the ESO new technology telescope. Astronomy & Astrophysics, 508(1):531–539, December 2009. [14] F. Zappa, S. Tisa, S. Cova, P. Maccagnani, R. Saletti, R. Roncella, F. Baronti, D. Bonaccini Calia, A. Silber, G. Bonanno, and M. Belluso. Photon counting arrays for astrophysics. Journal of Modern Optics, 54(2-3):163–189, January 2007. [15] Nicholas R. Shade, Gillian Kyne, Shouleh Nikzad, Edoardo Charbon, and Eric R. Fossum. Characterization and validation of next generation image sensors for space applications. In X-Ray, Optical, and Infrared Detectors for Astronomy XI, volume 13103, pages 351–362. SPIE, August 2024. [16] Vytautas Zickus, Ming-Lo Wu, Kazuhiro Morimoto, Valentin Kapitany, Areeba Fatima, Alex Turpin, Robert Insall, Jamie Whitelaw, Laura Machesky, Claudio Bruschini, Daniele Faccio, and Edoardo Charbon. Fluorescence lifetime imaging with a megapixel SPAD camera and neural network lifetime estimation. Scientific Reports, 10(1):20986, 2020. [17] Kazuhiro Morimoto. Megapixel SPAD cameras for time-resolved applications. PhD thesis, École Polytechnique Fédérale de Lausanne (EPFL), 2021. [18] Luca Zampieri, Giampiero Naletto, Cesare Barbieri, Enrico Verroi, Mauro Barbieri, G Ceribella, Maurizio D’Alessandro, G Farisato, Andrea Di Paola, and P Zoccarato. Aqueye+: a new ultrafast single photon counter for optical high time resolution astrophysics. In Photon Counting Applications 2015, volume 9504, pages 50–63. SPIE, 2015. [19] Yang Lin and Edoardo Charbon. Spiking neural networks for active time-resolved SPAD imaging. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 8147–8156, 2024. [20] STMicroelectronics. Stmicroelectronics proximity sensor solves smartphone hang-ups. https://www.st.com/content/st_com/en/about/media-center/pressitem.html/stmicroelectronics-proximity-sensor-solves-smartphone-hang-ups.html, 2/25/2013. [21] Apple Inc. Apple unveils new iPad Pro with LiDAR Scanner and trackpad support in iPadOS. https://www.apple.com/newsroom/2020/03/apple-unveils-new-ipad-prowith-lidar-scanner-and-trackpad-support-in-ipados/. [22] Apple Inc. Introducing Apple Vision Pro: Apple’s first spatial computer. https://www.apple.com/newsroom/2023/06/introducing-apple-vision-pro/. [23] Sony Semiconductor Solutions. Sony semiconductor solutions to release SPAD 112
Bibliography depth sensor for smartphones with high-accuracy, low-power distance measurement performance, powered by the industry’s highest photon detection efficiency | news releases | sony semiconductor solutions group. https://www.sonysemicon.com/en/news/2023/2023030601.html. [24] Canon Inc. Canon developing world-first ultra-high-sensitivity ILC equipped with SPAD sensor, supporting precise monitoring through clear color image capture of subjects several km away, even in darkness. https://global.canon/en/news/2023/20230403.html. [25] Raghubir Singh and Sukhpal Singh Gill. Edge AI: A survey. Internet of Things and Cyber-Physical Systems, 3:71–92, January 2023. [26] William Fabre, Karim Haroun, Vincent Lorrain, Maria Lepecq, and Gilles Sicard. From near-sensor to in-sensor: A state-of-the-art review of embedded AI vision systems. Sensors, 24(16):5446, January 2024. [27] Tzu-Hsiang Hsu, Guan-Cheng Chen, Yi-Ren Chen, Ren-Shuo Liu, Chung-Chuan Lo, Kea-Tiong Tang, Meng-Fan Chang, and Chih-Cheng Hsieh. A 0.8 V Intelligent Vision Sensor With Tiny Convolutional Neural Network and Programmable Weights Using Mixed-Mode Processing-in-Sensor Technique for Image Classification. IEEE Journal of Solid-State Circuits, 58(11):3266–3274, November 2023. [28] Zachary Ballard, Calvin Brown, Asad M. Madni, and Aydogan Ozcan. Machine learning and computation-enabled intelligent sensor design. Nature Machine Intelligence, 3(7):556–565, July 2021. [29] Guillermo Gallego, Tobi Delbruck, Garrick Orchard, Chiara Bartolozzi, Brian Taba, Andrea Censi, Stefan Leutenegger, Andrew J. Davison, Jorg Conradt, Kostas Daniilidis, and Davide Scaramuzza. Event-based vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(1):154–180, 2022. [30] Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep learning. Adaptive computation and machine learning. The MIT Press, Cambridge, Massachusetts, 2016. [31] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020. [32] OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke 113
Bibliography Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, RaphaGontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, C. J. Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang 114
Bibliography Zhuang, William Zhuk, and Barret Zoph. GPT-4 technical report, March 2024. [33] John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A. A. Kohl, Andrew J. Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David Reiman, Ellen Clancy, Michal Zielinski, Martin Steinegger, Michalina Pacholska, Tamas Berghammer, Sebastian Bodenstein, David Silver, Oriol Vinyals, Andrew W. Senior, Koray Kavukcuoglu, Pushmeet Kohli, and Demis Hassabis. Highly accurate protein structure prediction with AlphaFold. Nature, 596(7873):583–589, August 2021. [34] Kalanit Grill-Spector and Rafael Malach. The human visual cortex. Annual Review of Neuroscience, 27(1):649–677, 2004. [35] Richard H. Masland. The neuronal organization of the retina. Neuron, 76(2):266–280, October 2012. [36] Kunihiko Fukushima. Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological Cybernetics, 36(4):193–202, April 1980. [37] Samanwoy Ghosh-Dastidar and Hojjat Adeli. Spiking neural networks. International Journal of Neural Systems, 19(04):295–308, August 2009. [38] Petr Adámek, Veronika Langová, and Jiˇrí Horáˇcek. Early-stage visual perception impairment in schizophrenia, bottom-up and back again. Schizophrenia, 8(1):27, 2022. [39] Andrea Gallivanoni, Ivan Rech, and Massimo Ghioni. Progress in quenching circuits for single photon avalanche diodes. IEEE Transactions on Nuclear Science, 57(6):3815–3826, December 2010. [40] Fabio Severini, Iris Cusini, Davide Berretta, Klaus Pasquinelli, Alfonso Incoronato, and Federica Villa. SPAD pixel with sub-ns dead-time for high-count rate applications. IEEE Journal of Selected Topics in Quantum Electronics, 28(2: Optical Detectors):1–8, March 2022. [41] Man Yao, Ole Richter, Guangshe Zhao, Ning Qiao, Yannan Xing, Dingheng Wang, Tianxiang Hu, Wei Fang, Tugba Demirci, Michele De Marchi, Lei Deng, Tianyi Yan, Carsten Nielsen, Sadique Sheik, Chenxi Wu, Yonghong Tian, Bo Xu, and Guoqi Li. Spike-based dynamic computing with asynchronous sensing-computing neuromorphic chip. Nature Communications, 15(1):4464, May 2024. [42] Xu Yang, Fuming Lei, Na Tian, Cong Shi, Zhe Wang, Shuangming Yu, Runjiang Dou, Peng Feng, Nan Qi, Zhongming Wei, Jian Liu, Kaiyou Wang, Nanjian Wu, and Liyuan Liu. A 10 000-inference/s bio-inspired spiking vision chip based on an end-to-end SNN embedding image signal enhancement. IEEE Journal of Solid-State Circuits, pages 1–17, 2025. [43] Abdul Waris Ziarkash, Siddarth Koduru Joshi, Mario Stipˇcevi´c, and Rupert Ursin. Comparative study of afterpulsing behavior and models in single photon counting avalanche 115
Bibliography photo diode detectors. Scientific Reports, 8(1):5076, March 2018. [44] Wolfgang Becker. TCSPC Handbook 9th edn. Becker & Hickl GmbH, 2021. [45] Stephan Henzler. Time-to-Digital Converters, volume 29. Springer Netherlands, Dordrecht, 2010. [46] Gordon W. Roberts and Mohammad Ali-Bakhshian. A brief introduction to time-todigital and digital-to-time converters. IEEE Transactions on Circuits and Systems II: Express Briefs, 57(3):153–157, March 2010. [47] Kazuhiro Morimoto, Andrei Ardelean, Ming-Lo Wu, Arin Can Ulku, Ivan Michel Antolovic, Claudio Bruschini, and Edoardo Charbon. A megapixel time-gated SPAD image sensor for 2D and 3D imaging applications, December 2019. [48] Charles Poynton. Digital Video and HD: Algorithms and Interfaces. Elsevier, February 2012. [49] Ivan Michel Antolovic, Claudio Bruschini, and Edoardo Charbon. Dynamic range extension for photon counting arrays. Optics Express, 26(17):22234–22248, 2018. [50] Jeff W. Lichtman and José-Angel Conchello. Fluorescence microscopy. Nature methods, 2(12):910–919, 2005. [51] Rupsa Datta, Tiffany M. Heaster, Joe T. Sharick, Amani A. Gillette, and Melissa C. Skala. Fluorescence lifetime imaging microscopy: fundamentals and advances in instrumentation, analysis, and applications. Journal of Biomedical Optics, 25(7):1–43, 2020. [52] Erik B. van Munster and Theodorus W. J. Gadella. Fluorescence lifetime imaging microscopy (FLIM). In Microscopy Techniques, pages 143–175. Springer, Berlin, Heidelberg, 2005. [53] Klaus Suhling, Liisa M. Hirvonen, James A. Levitt, Pei-Hua Chung, Carolyn Tregidgo, Alix Le Marois, Dmitri A. Rusakov, Kaiyu Zheng, Simon Ameer-Beg, Simon Poland, Simao Coelho, Robert Henderson, and Nikola Krstajic. Fluorescence lifetime imaging (FLIM): Basic concepts and some recent developments. Medical Photonics, 27:3–40, 2015. [54] Horst Wallrabe and Ammasi Periasamy. Imaging protein molecules using FRET and FLIM microscopy. Current Opinion in Biotechnology, 16(1):19–27, 2005. [55] Jakob Unger, Christoph Hebisch, Jennifer E. Phipps, João L. Lagarto, Hanna Kim, Morgan A. Darrow, Richard J. Bold, and Laura Marcu. Real-time diagnosis and visualization of tumor margins in excised breast specimens using fluorescence lifetime imaging and machine learning. Biomedical Optics Express, 11(3):1216–1230, 2020. [56] Mikael T. Erkkilä, David Reichert, Johanna Gesperger, Barbara Kiesel, Thomas Roetzer, Petra A. Mercea, Wolfgang Drexler, Angelika Unterhuber, Rainer A. Leitgeb, Adelheid Woehrer, Angelika Rueck, Marco Andreana, and Georg Widhalm. Macroscopic fluorescence-lifetime imaging of NADH and protoporphyrin IX improves the detection and grading of 5-aminolevulinic acid-stained brain tumors. Scientific Reports, 10(1):20492, 2020. [57] Brent W. Weyers, Mark Marsden, Tianchen Sun, Julien Bec, Arnaud F. Bewley, Regina F. Gandour-Edwards, Michael G. Moore, D. Gregory Farwell, and Laura Marcu. Fluores116
Bibliography cence lifetime imaging for intraoperative cancer delineation in transoral robotic surgery. Translational biophotonics, 1(1-2), 2019. [58] Jennifer Phipps, Jakob Unger, Regina Gandour-Edwards, Michael G. Moore, Arnaud Bewley, D. Gregory Farwell, and Laura Marcu. Head and neck cancer evaluation via transoral robotic surgery with augmented fluorescence lifetime imaging. In Biophotonics Congress: Biomedical Optics Congress 2018 (Microscopy/Translational/Brain/OTS), page CTu2B.3, Washington, D.C., 2018. OSA. [59] W. Becker. Fluorescence lifetime imaging–techniques and applications. Journal of Microscopy, 247(2):119–136, 2012. [60] Liisa M. Hirvonen and Klaus Suhling. Wide-field TCSPC: methods and applications. Measurement Science and Technology, 28(1):012003, 2017. [61] Peter Kapusta, Michael Wahl, and Rainer Erdmann. Advanced photon counting: Applications, methods, instrumentation / volume editors, Peter Kapusta, Michael Wahl, Rainer Erdmann ; with contributions by A. Ahlrichs [and fifty others], volume 15 of Springer Series on Fluorescence, 1617-1306. Springer, Cham, 2015. [62] Alba Alfonso-Garcia, Julien Bec, Brent Weyers, Mark Marsden, Xiangnan Zhou, Cai Li, and Laura Marcu. Mesoscopic fluorescence lifetime imaging: Fundamental principles, clinical applications and future directions. Journal of Biophotonics, 14(6):e202000472, 2021. [63] Massimo Caccia, Luca Nardo, Romualdo Santoro, and Daniel Schaffhauser. Silicon photomultipliers and SPAD imagers in biophotonics: Advances and perspectives. Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment, 926:101–117, 2019. [64] Iris Cusini, Davide Berretta, Enrico Conca, Alfonso Incoronato, Francesca Madonini, Arianna Adelaide Maurina, Chiara Nonne, Simone Riccardo, and Federica Villa. Historical perspectives, state of art and research trends of SPAD arrays and their applications (part ii: SPAD arrays). Frontiers in Physics, 10:606, 2022. [65] Željko Bajzer, Terry M. Therneau, Joseph C. Sharp, and Franklin G. Prendergast. Maximum likelihood method for the analysis of time-resolved fluorescence decay curves. European Biophysics Journal, 20(5):247–262, December 1991. [66] Michael Maus, Mircea Cotlet, Johan Hofkens, Thomas Gensch, Frans C. De Schryver, J. Schaffer, and C. A. M. Seidel. An experimental comparison of the maximum likelihood estimation and nonlinear least-squares fluorescence lifetime analysis of single molecules. Analytical Chemistry, 73(9):2078–2086, May 2001. [67] Amiram Grinvald and Izchak Z. Steinberg. On the analysis of fluorescence decay kinetics by the method of least-squares. Analytical Biochemistry, 59(2):583–598, June 1974. [68] Malte Köllner and Jürgen Wolfrum. How many photons are necessary for fluorescencelifetime measurements? Chemical Physics Letters, 200(1-2):199–204, 1992. [69] Irvin Isenberg and Robert D. Dyson. The analysis of fluorescence decay by a method of moments. Biophysical Journal, 9(11):1337–1350, November 1969. 117
Bibliography [70] Irvin Isenberg, Robert D. Dyson, and Richard Hanson. Studies on the analysis of fluorescence decay data by the method of moments. Biophysical Journal, 13(10):1090–1115, October 1973. [71] David Day-Uei Li, Hongqi Yu, and Yu Chen. Fast bi-exponential fluorescence lifetime imaging analysis methods. Optics Letters, 40(3):336–339, 2015. [72] Gang Wu, Thomas Nowotny, Yongliang Zhang, Hong-Qi Yu, and David Day-Uei Li. Artificial neural network approaches for fluorescence lifetime imaging techniques. Optics Letters, 41(11):2561–2564, 2016. [73] Dong Xiao, Yu Chen, and David Day-Uei Li. One-dimensional deep learning architecture for fast fluorescence lifetime imaging. IEEE Journal of Selected Topics in Quantum Electronics, 27(4):1–10, 2021. [74] Jason T. Smith, Nathan Un, Ruoyang Yao, Nattawut Sinsuebphon, Alena Rudkouskaya, Joseph Mazurkiewicz, Margarida Barroso, Pingkun Yan, and Xavier Intes. Fluorescent lifetime imaging improved via deep learning. Novel Techniques in Microscopy, page NM3C.4, 2019. [75] Dong Xiao, Natakorn Sapermsap, Yu Chen, and David Day Uei Li. Deep learning enhanced fast fluorescence lifetime imaging with a few photons. Optica, 10(7):944–951, July 2023. [76] Sofia Kapsiani, Nino F. Läubli, Edward N. Ward, Ana Fernandez-Villegas, Bismoy Mazumder, Clemens F. Kaminski, and Gabriele S. Kaminski Schierle. Deep learning for fluorescence lifetime predictions enables high-throughput in vivo imaging. Journal of the American Chemical Society, 147(26):22609–22621, July 2025. [77] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. [78] Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3):39:1–39:45, March 2024. [79] Jeffrey L. Elman. Finding structure in time. Cognitive Science, 14(2):179–211, 1990. [80] Yoshua Bengio, Patrice Simard, and Paolo Frasconi. Learning long-term dependencies with gradient descent is difficult. IEEE Transactions on Neural Networks, 5(2):157–166, March 1994. [81] Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9(8):1735–1780, 1997. [82] Kyunghyun Cho, Bart van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio. On the properties of neural machine translation: Encoder–decoder approaches. In Dekai Wu, 118
Bibliography Marine Carpuat, Xavier Carreras, and Eva Maria Vecchi, editors, Proceedings of SSST-8, Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation, pages 103–111, Doha, Qatar, October 2014. Association for Computational Linguistics. [83] Amirhossein Tavanaei, Masoud Ghodrati, Saeed Reza Kheradpisheh, Timothée Masquelier, and Anthony Maida. Deep learning in spiking neural networks. Neural Networks, 111:47–63, March 2019. [84] Samanwoy Ghosh-Dastidar and Hojjat Adeli. Spiking neural networks. International Journal of Neural Systems, 19(04):295–308, August 2009. [85] Wulfram Gerstner, Werner M Kistler, Richard Naud, and Liam Paninski. Neuronal dynamics: From single neurons to networks and models of cognition. Cambridge University Press, 2014. [86] Alan L. Hodgkin and Andrew F. Huxley. A quantitative description of membrane current and its application to conduction and excitation in nerve. The Journal of Physiology, 117(4):500–544, 1952. [87] Larry F. Abbott. Lapicque’s introduction of the integrate-and-fire model neuron (1907). Brain Research Bulletin, 50(5):303–304, November 1999. [88] Simon Thorpe, Arnaud Delorme, and Rufin Van Rullen. Spike-based strategies for rapid processing. Neural Networks, 14(6):715–725, July 2001. [89] Guo-qiang Bi and Mu-ming Poo. Synaptic modifications in cultured hippocampal neurons: Dependence on spike timing, synaptic strength, and postsynaptic cell type. Journal of Neuroscience, 18(24):10464–10472, December 1998. [90] Natalia Caporale and Yang Dan. Spike timing–dependent plasticity: A hebbian learning rule. Annual Review of Neuroscience, 31(Volume 31, 2008):25–46, July 2008. [91] Michael Pfeiffer and Thomas Pfeil. Deep learning with spiking neurons: Opportunities and challenges. Frontiers in Neuroscience, 12:774, 2018. [92] Jason K. Eshraghian, Max Ward, Emre O. Neftci, Xinxin Wang, Gregor Lenz, Girish Dwivedi, Mohammed Bennamoun, Doo Seok Jeong, and Wei D. Lu. Training spiking neural networks using lessons from deep learning. Proceedings of the IEEE, 111(9):1016– 1054, 2023. [93] Nicolas Perez-Nieves and Dan Goodman. Sparse spiking gradient descent. Advances in Neural Information Processing Systems, 34:11795–11808, 2021. [94] Jianhao Ding, Zhaofei Yu, Yonghong Tian, and Tiejun Huang. Optimal ANN-SNN conversion for fast and accurate inference in deep spiking neural networks. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence (IJCAI-21), 2021. [95] Bodo Rueckauer, Iulia-Alexandra Lungu, Yuhuang Hu, Michael Pfeiffer, and Shih-Chii Liu. Conversion of continuous-valued deep networks to efficient event-driven networks for image classification. Frontiers in Neuroscience, 11:682, 2017. [96] Sander M. Bohte, Joost N. Kok, and Han La Poutré. Error-backpropagation in temporally encoded networks of spiking neurons. Neurocomputing, 48(1-4):17–37, 2002. [97] Olaf Booij and Hieu tat Nguyen. A gradient descent rule for spiking neurons emitting 119
Yang LIN Curriculum Vitae Advanced Quantum Architecture Lab (AQUA) École Polytechnique Fédérale de Lausanne By[email protected]
[email protected] ÍYang’s Homepage §Github ïLinkedin Google Scholar Education 2020–2025 Ph.D. in Computational and Quantitative Biology,École Polytechnique Fédérale de Lausanne, Lausanne. PhD thesis: SPAD Image Sensors with Embedded Intelligence. This research focuses on integrating recurrent and spiking neural networks within or near SPAD image sensors to enable efficient, real-time, intelligent edge processing, especially for biomedical imaging. The work encompasses SPAD image sensor design, processor architecture, FPGA implementation, software development, neural network training and evaluation, mathematical modeling, fluorescence lifetime imaging, and optical system setup. Supervisor: Prof. Edoardo Charbon and Dr. Bruschini Claudio. 2016–2020 Bachelor of Engineering in Automation,Shanghai Jiao Tong University, Shanghai. Bachelor dissertation: Subcellular localization of long noncoding RNA with interpretable deep learning. This study explores the application of natural language processing techniques to RNA sequences for predicting the subcellular localization of long noncoding RNAs. The developed prediction software is deployed on a server, providing biologists with accessible and reliable tools for analysis. Supervisors: Prof. Hong-Bin Shen and Prof. Xiaoyong Pan Publications Journal Articles 2025 Duncan P Ryan, Paul Moş, Yang Lin, Claudio Bruschini, Edoardo Charbon, and James H Werner. Time-resolved detectors for quantum ghost imaging. The European Physical Journal Plus, volume 140, page 511. Springer, 2025. 2025 Audrey Eshun, Dominique Davenport, Brandon Demory, Shervin Kiannejad, Paul Mos, Yang Lin, Tiziana Bond, Michael C Rushford, Erin E Nuccio, Ty J Samo, et al. 3d quantum ghost imaging microscope. Optica, volume 12, pages 1109–1112. Optica Publishing Group, 2025. 2024 Yang Lin, Paul Mos, Andrei Ardelean, Claudio Bruschini, and Edoardo Charbon. Coupling a recurrent neural network to SPAD TCSPC systems for real-time fluorescence lifetime imaging. Scientific Reports, volume 14, pages 1–13. Nature Publishing Group, 2024. 2024 Dominique Davenport, Audrey Eshun, Brandon Demory, Shervin Kiannejad, Paul Mos, Yang Lin, Michael Wayne, Sam Jeppson, Ashleigh Wilson, Tiziana Bond, et al. Quantum ghost imaging microscopy depth-of-field study. Optics Express, volume 32, pages 36031–36047. Optica Publishing Group, 2024. 2021 Yang Lin, Xiaoyong Pan, and Hong-Bin Shen. lnclocator 2.0: a cell-line-specific subcellular localization predictor for long non-coding RNAs with interpretable deep learning. Bioinformatics, 2021. In Conference Proceedings 2025 Yang Lin, Claudio Bruschini, and Edoardo Charbon. Transporter: A 128 × 4 SPAD imager with on-chip encoder for spiking neural network-based processing. In International Image Sensor Workshop (IISW), 2025. 127
2025 Duncan P Ryan, James H Werner, Kati A Seitz, Dean P Morales, Rebecca Holmes, Edoardo Charbon, Claudio Bruschini, Paul Mos, and Yang Lin. Moving quantum ghost imaging from demonstration to application. In Photonic Instrumentation Engineering XII, page PC1337303. SPIE, 2025. 2025 Dominique Davenport, Audrey Eshun, Brandon Demory, Shervin Kiannejad, Paul Mos, Yang Lin, Edoardo Charbon, Tiziana Bond, Mike Rushford, Claudio Bruschini, et al. 3d quantum ghost imaging microscope. In Quantum Sensing, Imaging, and Precision Metrology III, page PC133922Q. SPIE, 2025. 2024 Yang Lin and Edoardo Charbon. Spiking neural networks for active time-resolved SPAD imaging. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 8147–8156, 2024. Awards 2022 Best Talk of a Young Researcher award at the Biomedical Photonics Network (BMPN 2022) conference in Lausanne, Switzerland. Teaching Assistantship 2023-24 : MICRO-428: Metrology, Preparing course materials, answering questions and running exercises sessions, esp. on statistics. EPFL, Lausanne 2024-25 : MICRO-429: Metrology Practicals, Instructing students in the experiments on optical metrology, incl. measuring dark count rate, afterpulsing probability, and jitters. EPFL, Lausanne 128