scieee AI-readable full text Open interactive document viewer

Exploring image-based AI methods for analyzing timelapse EmbryoScope data

Meggle, Lukas

Abstract

Els avenços en les Tecnologies de Reproducció Assistida (ART) han posat de manifest la necessitat d’eines objectives i basades en dades per donar suport a la selecció d’embrions en la Fecundació In Vitro (FIV). Aquesta tesi explora mètodes d’intel·ligència artificial basats en imatges per predir la formació de blastocists de bona qualitat al dia 3, utilitzant dades temporals de temps obtingudes d’incubadores EmbryoScope. Es va recopilar i preprocessar un conjunt de dades personalitzat amb 10.274 seqüències d’embrions anotades, en col·laboració amb l’Hospital Clínic de Barcelona, per a l’entrenament dels models. Es van avaluar diversos enfocaments, incloent-hi models tradicionals d’aprenentatge automàtic i models d’aprenentatge profund, en diferents modalitats: imatges individuals, seqüències de dades temporals de temps i característiques morfocinètiques fusionades. El model amb millor rendiment —basat en EfficientNetB0 i una fusió de característiques amb mecanisme d’atenció— va assolir una precisió equilibrada del 77,37% en el conjunt de test, mentre que un model completament automatitzat basat en vídeo, amb una xarxa LSTM, va arribar a una precisió equilibrada del 75,3%. Els resultats demostren un alt rendiment dels models embrionaris, assolint resultats d’última generació en el conjunt de dades disponible i mostrant un potencial prometedor per al suport clínic real en la selecció d’embrions.

Full text

id194852   EXPLORING IMAGE-BASED AI METHODS FOR ANALYZING TIMELAPSE EMBRYOSCOPE DATA LUKAS MEGGLE Thesis supervisor DARIOGARCÍAGASULLA(BARCELONASUPERCOMPUTINGCENTER-CENTRONACIONALDE SUPERCOMPUTACION) Tutor:JAVIERBÉJARALONSO(DepartmentofComputerScience) Degree Master'sDegreeinArtificialIntelligence Master's thesis School of Engineering Universitat Rovira i Virgili (URV) Faculty of Mathematics Universitat de Barcelona (UB) Barcelona School of Informatics (FIB) Universitat Politècnica de Catalunya (UPC) - BarcelonaTech  Resum Els avenços en les Tecnologies de Reproducció Assistida (ART) han posat de manifest la necessitat d’eines objectives i basades en dades per donar suport a la selecció d’embrions en la Fecundació In Vitro (FIV). Aquesta tesi explora mètodes d’intel·ligència artificial basats en imatges per predir la formació de blastocists de bona qualitat al dia 3, utilitzant dades temporals de temps obtingudes d’incubadores EmbryoScope. Es va recopilar i preprocessar un conjunt de dades personalitzat amb 10.274 seqüències d’embrions anotades, en col·laboració amb l’Hospital Clínic de Barcelona, per a l’entrenament dels models. Es van avaluar diversos enfocaments, incloent-hi models tradicionals d’aprenentatge automàtic i models d’aprenentatge profund, en diferents modalitats: imatges individuals, seqüències de dades temporals de temps i característiques morfocinètiques fusionades. El model amb millor rendiment —basat en EfficientNetB0 i una fusió de característiques amb mecanisme d’atenció— va assolir una precisió equilibrada del 77,37% en el conjunt de test, mentre que un model completament automatitzat basat en vídeo, amb una xarxa LSTM, va arribar a una precisió equilibrada del 75,3%. Els resultats demostren un alt rendiment dels models embrionaris, assolint resultats d’última generació en el conjunt de dades disponible i mostrant un potencial prometedor per al suport clínic real en la selecció d’embrions. Resumen Los avances en las Tecnologías de Reproducción Asistida (ART) han subrayado la necesidad de herramientas objetivas y basadas en datos para apoyar la selección de embriones en la Fecundación In Vitro (FIV). Esta tesis explora métodos de inteligencia artificial basados en imágenes para predecir la formación de blastocistos de buena calidad en el día 3, utilizando datos de lapso de tiempo obtenidos de incubadoras EmbryoScope. Se recopiló y preprocesó un conjunto de datos personalizado con 10.274 secuencias de embriones anotadas, en colaboración con el Hospital Clínic de Barcelona, para el entrenamiento de los modelos. Se evaluaron múltiples enfoques, incluidos modelos tradicionales de aprendizaje automático y modelos de aprendizaje profundo, en diferentes modalidades: imágenes individuales, secuencias de lapso de tiempo y características morfocinéticas fusionadas. El modelo con mejor rendimiento —basado en EfficientNetB0 y una fusión de características con atención— logró una precisión equilibrada del 77,37% en el conjunto de prueba, mientras que un modelo completamente automatizado basado en video, que incorpora una red LSTM, alcanzó una precisión equilibrada del 75,3%. Los resultados demuestran un alto rendimiento de los modelos embrionarios, alcanzando resultados de última generación en el conjunto de datos disponible y mostrando un potencial prometedor para el apoyo clínico real en la selección de embriones. Abstract Advances in Assisted Reproductive Technology (ART) have underscored the need for objective, data-driven tools to support embryo selection in In Vitro Fertilization (IVF). This thesis explores image-based artificial intelligence methods for predicting good-quality blastocyst formation at day 3 using time-lapse data from EmbryoScope incubators. A custom dataset of 10,274 annotated embryo sequences, collected in collaboration with Hospital Clínic de Barcelona, was assembled and preprocessed for model training. Multiple approaches, including traditional machine learning and deep learning models, were evaluated across different modalities: single images, time-lapse sequences, and fused morphokinetic features. The best-performing model - based on a EfficientNetB0 and attention-based feature fusion - achieved a balanced accuracy of 77.37% on the test set, while a fully automated video-based model incorporating an LSTM network reached a balanced accuracy of 75.3%. The results demonstrate a strong performance of the embryo models, reaching state-of-the-art results on the dataset at hand and showing promising potential for real-world clinical support in embryo selection. Table of Contents List of Figures 3 List of Tables 5 Abbreviations 7 1 Introduction, Motivation and Objectives 9 2 State of the Art 11 2.1 Morphology and Morphokinetics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.2 AI Models for Embryo Assessment . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 2.2.1 Embryo Labelling, Scoring and Selection . . . . . . . . . . . . . . . . . . . . . 12 2.2.2 Blastocyst Prediction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 3 Dataset 15 3.1 Metadata .......................................... 15 3.2 EmbryoSampleSelection ................................. 17 3.3 Labeling........................................... 18 3.4 Feature Extraction and Imputation Strategy . . . . . . . . . . . . . . . . . . . . . . . 18 3.5 FrameTimeExtraction .................................. 19 3.6 FinalDatasetPreparation................................. 20 4 Methods 21 4.1 MachineLearningModel.................................. 21 4.2 DeepLearningModels................................... 22 4.2.1 SingleImageModel ................................ 22 4.2.2 EmbryoFeatureModel .............................. 26 4.2.3 EmbryoVideoModel ............................... 27 5 Results 31 5.1 MachineLearningModel.................................. 31 5.2 DeepLearningModels................................... 31 5.2.1 Single Image Model Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 5.2.2 EmbryoFeatureModel .............................. 34 5.2.3 EmbryoVideoModel ............................... 35 6 Sustainability analysis and ethical implications 37 1 PN pronuclear. 15, 18, 55 TE trophectoderm. 11, 41 UPC Universitat Politècnica de Catalunya. 10 ViTs Vision Transformers. 23 Z-score zygote score. 11 8 Chapter 1 Introduction, Motivation and Objectives In Vitro Fertilization (IVF) is one of the most widely applied Assisted Reproductive Technology (ART) for individuals and couples experiencing infertility. Since its first successful use in 1978, IVF has undergone substantial advancements, resulting in improved success rates and offering hope to those struggling with natural conception [1]. Infertility has emerged as a major public health concern in industrialized nations [ 2 ]. This trend is closely linked to the rise of unhealthy lifestyle behaviors characteristic of modern society [ 3 ]. Factors such as physical inactivity, poor dietary habits, and excess body weight are negatively associated with reproductive health, contributing to a decline in fertility potential [ 4 ]. In addition, environmental exposures, for example endocrine-disrupting chemicals, have been shown to impair both female and male reproductive function [5]. The procedure of IVF includes ovarian stimulation, egg retrieval, fertilization in a laboratory environment, embryo culture, and subsequent embryo transfer to the uterus [ 6 ]. Fertilization occurs through either standard IVF or Intracytoplasmic Sperm Injection (ICSI). In standard IVF, eggs and sperm are combined in a culture dish, allowing fertilization to occur naturally as it would in the female reproductive tract. This method relies on the sperm’s ability to penetrate the egg without medical intervention and is primarily used in cases where sperm quality is sufficient for natural fertilization. In contrast, ICSI involves the direct injection of a single sperm cell into the cytoplasm of a mature oocyte, allowing fertilization to occur even in cases of severe male-factor infertility, however, it is increasingly used for all types of infertility, accounting for now 70-80% of ART cycles globally [7]. After fertilization, embryos are cultured in incubators that maintain stable environmental conditions, including temperature, gas composition, and pH levels, which are essential for optimal embryonic development. Modern time-lapse incubators, such as the EmbryoScope (Vitrolife©, Sweden), offer continuous image acquisition without removing embryos from their controlled environment. This uninterrupted culture combined with real-time monitoring allows embryologists to evaluate both morphological characteristics and morphokinetic parameters - the timing of cell divisions and other developmental milestones. These time-lapse features have been associated with improved embryo selection, as more frequent observations yield deeper insights into developmental competence, implantation potential or chromosomal normality [8]. 9 Chapter 1. Introduction, Motivation and Objectives A key decision in ART is whether to transfer embryos at an early stage or allow further in vitro development before selection. Early-stage transfer occurs at the cleavage stage, typically on day two or three, when the embryo comprises four to eight cells. This approach is often preferred when embryo numbers are limited, as only about half of all embryos reach the blastocyst stage under extended culture conditions. In contrast, blastocyst-stage transfer, performed on day five or six, offers improved implantation potential due to enhanced embryonic development and self-selection. However, not all embryos successfully develop to the blastocyst stage, and in some cases, especially among patients with a limited number of embryos, this may result in no embryos available for transfer [9]. The selection of embryos, as well as the decision between early-stage or blastocyst-stage transfer, relies heavily on the expertise of embryologists. These decisions are based on the assessment of both morphokinetic and morphological characteristics to estimate which embryo has the highest potential for implantation or successful development to the blastocyst stage. However, this evaluation process is inherently subjective and susceptible to human bias. Variations in experience, training, and interpretation among clinicians can lead to inconsistencies in embryo selection, ultimately affecting implantation and pregnancy outcomes. In retrospective analyses, some Artifical Intelligence (AI)-based models have demonstrated promising performance, with reported accuracies exceeding those of manual evaluations by embryologists [10]. In this thesis, the predictive capabilities of various AI models for forecasting blastocyst formation from day 3 embryos are investigated. The goal is to develop a model capable of determining whether an embryo will successfully reach a good-quality blastocyst stage, thereby supporting clinical decisions on whether to proceed with early-stage or extended embryo culture and transfer. A dataset comprising 10,274 time-lapse images of embryos was assembled as the foundation for this study. An evaluation will be conducted to determine which types of data and modeling approaches offer the highest predictive power, using a range of neural networks and machine learning techniques applied to different data modalities. This work was carried out during a master’s thesis at the Barcelona Supercomputing Center (BSC) and Universitat Politècnica de Catalunya (UPC), as part of a collaborative project between the BSC and the Hospital Clínic de Barcelona. 10 Chapter 2 State of the Art 2.1 Morphology and Morphokinetics Studies have shown that both morphological - the embryo’s visual quality at fixed time points in hours post insemination (HPI) - and morphokinetic information - the timing of developmental events such as the timing of cell divisions - can help determine the probability of blastocyst formation. Wong et al. [ 11 ] reported that embryos exhibiting faster early cleavage times were significantly more likely to develop into high-quality blastocysts. This supports the hypothesis that morphokinetics carry prognostic value. However, the study was limited by a relatively small dataset. In clinical practice, embryo assessment relies heavily on morphology-based grading systems. Earlystage evaluation includes the zygote score (Z-score) - which assesses the quality of the zygote shortly after fertilization - and the day 3 embryo morphology score, both of which have been shown to correlate with embryo viability at later stages [ 12 ]. This highlights that morphology at different time points during embryo development is a strong indicator of blastocyst formation. At later stages, clinicians use other well-established morphology-based scoring systems such as the Gardner grading system and the Veeck and Zaninovic grading scheme to evaluate blastocyst quality based on expansion, inner cell mass (ICM), and trophectoderm (TE) structure. While high-scoring blastocysts typically demonstrate higher implantation potential, it is important to note that even lower-scoring embryos can still result in successful pregnancies, however with a decreased probability [13]. Figure 2.1: Selected frames from a time-lapse sequence of an embryo developing to the blastocyst stage, annotated with time in HPI and corresponding morphokinetic stages. The red marker indicates the approximate prediction time considered in this thesis, at which the likelihood of reaching a good-quality blastocyst should be assessed. For details on developmental stage annotations, see section 3.1. 11 Chapter 2. State of the Art 2.2 AI Models for Embryo Assessment As mentioned in the introduction, AI models have been applied to various tasks related to embryo assessment. These models can support experts by reducing the need for extensive manual annotation and labour and may help mitigate human bias and variability in decision-making. 2.2.1 Embryo Labelling, Scoring and Selection Different studies attempt to predict scores or labels for embryos. For example, [ 14 ] predicted morphokinetic annotations of individual frames from time-lapse sequences with an accuracy of 84.79% with a dataset of 1,309 embryos, while [ 15 ] counted blastomeres from 1 to 5 cells with a modified U-Net trained on 190 images, reporting an average accuracy of 88.2%. [ 16 ] predicted embryo quality at day 3 in four categories, outperforming embryologists with an ensemble learning model based on various Convolutional Neural Network (CNN) networks, trained on 3,601 microscopic images. The model or code for the previously mentioned studies are not published. The STORK model [ 17 ] predicts the final Veeck and Zaninovic blastocyst quality grading from a blastocyst image using a pretrained Inception-V1 CNN, trained with a dataset of 10,148 human embryos. The code is available and the model accessible though a web interface. One of the most common applications of AI in reproductive medicine is embryo selection, where models are trained to predict the viability of embryos at the blastocyst stage (day 5) for transfer. The classification task is typically based on implantation success or live birth outcomes. VerMilyea et al. [ 18 ] developed the Life Whisperer AI system using actual pregnancy outcomes as ground truth data. The system is commercially available as a decision-support tool for embryo selection and uses static two-dimensional optical microscope images of day 5 blastocysts as input. In the context of the original study and subsequent retrospective analyses, the model outperformed embryologists [ 19 ]. It was trained on 8,886 embryos collected from 11 IVF clinics. One of the most complex models, iDAScore v2.0, was developed by Lassen et al. [ 20 ]. It utilizes separate Two-Stream Inflated 3D ConvNet (I3D) video models [ 21 ] for day 2/3 and day 5+ embryos, processing 64 or 128 frames respectively from time-lapse sequences. The model predicts implantation potential without requiring manual annotations. Trained on a dataset of 181,428 embryos from 22 IVF clinics worldwide, it represents the largest dataset used for embryo model training to date. Inclusion of discarded embryos in the training data enhanced the model’s ability to differentiate between embryos suitable for transfer and those to be discarded, avoiding bias towards only high-quality embryos with known implantation outcomes. iDAScore v2.0 is available as an optional software module for EmbryoScope time-lapse incubators. Although the prediction goals of these studies differ from that of this thesis, they offer valuable reference points by demonstrating how deep learning can extract meaningful visual and temporal patterns from embryo imagery. These approaches highlight strategies for model design and data handling that are relevant to the predictive task addressed in this work. 2.2.2 Blastocyst Prediction The prediction task of this work - the formation probability for good-quality blastocysts from day 3 embryos - was previously addressed by [ 22 ]. The ensemble model, called STEM, reached 0.82 Area Under Curve (AUC), with a balanced accuracy (see formula 5) of 76.1%. The model consists of a morphological stream model, a DenseNet feature extractor using five fixed time points starting from tPNF, and a temporal stream model, counting cells in each frame and feeding this information into an Long Short-Term Memory (LSTM) network. The dataset consisted of 10,432 embryo time-lapse 12 Chapter 2. State of the Art videos. They noted that cleavage time differs significantly between ICSI and conventional IVF embryos, as previously mentioned in the introduction. Therefore, they employed a tPNF recognition network to align the time-lapse sequences to this reference point. An interesting study by [ 23 ] introduced a novel approach for assessing blastocyst formation based on cytoplasmic particle movement during the first 44 hours of embryo development. Using time-lapse imaging combined with Particle Image Velocimetry, the study extracted embryo movement vectors by analyzing pixel-level changes between consecutive frames. These motion-derived features were used as input to various machine learning models, including LSTM neural networks and k-Nearest Neighbors, to predict the likelihood of blastocyst development. It is particularly notable since no information about the embryo’s actual appearance or morphokinetic time points was used. The input data consisted solely of features derived from cytoplasmic movement, such as summed pixel-wise motion vectors, without incorporating appearance-based or morphokinetic annotations. However, the dataset was small (230 embryos) and based on a highly selective scenario - sibling embryos where only one developed to blastocyst - limiting its real-world applicability. Another study [ 24 ] reported a balanced accuracy of 66% for predicting if an embryo develops into a blastocyst, using a Xception CNN. The input is a single image at 70 HPI sampled from a dataset of 3,469 embryo time lapse images. The state-of-the-art model was proposed by [ 25 ] with their Adaptive Key Frame Selection (AdaKFS) approach, which combines an LSTM and a policy network. The study used a dataset of 3,300 human embryo time-lapse videos, with each embryo labeled for blastocyst formation outcome based on the Gardner scoring system. From each video, 32 frames were uniformly sampled and passed through a ResNet-50 backbone to extract morphological features. These features were combined with encoded kinetic parameters and input into an LSTM model. At each time step, a policy network decided whether to skip or retain the current frame’s hidden state for prediction. The selected hidden states (on average around six per video) were then passed to a prediction network to estimate the probability of blastocyst formation. The method was evaluated against the 8-layer CNN approach by [ 26 ] and the STEM network by [ 22 ] on their dataset, outperforming both with a balanced accuracy of 69.43%. A baseline model using only the LSTM without frame selection achieved a lower accuracy of 60.77%. The most recent publication by [ 27 ] was trained on the public dataset introduced by [ 28 ], which contains 704 time-lapse videos. Of these, 522 were labeled as valid blastocysts — an arguably questionable number, as a closer analysis of the dataset reveals that only 490 embryos reached the tB stage, 388 reached tEB , and just 5 were annotated with tHB . Despite this, the authors report a notably high accuracy of 93% using a fine-tuned ResNet-50 on frame-wise morphokinetic annotations combined with a GRU-based sequence model. However, this result is questionable due to the small dataset size and inconsistency with performance metrics reported in other studies. This public available dataset was not selected for this master’s thesis because it does not account for embryo viability. It is intended solely for frame-wise phase prediction, as noted by [ 28 ], and is therefore unsuitable for our classification goal of identifying high-quality blastocysts. Unfortunately, none of the mentioned models aligning with our prediction task have been made publicly available. Only the code for the STEM network proposed by [ 22 ] has been released. As a result, direct comparison is challenging, particularly due to differences in datasets and the lack of complete implementation details. 13 Chapter 3 Dataset Through a formal agreement, the BSC was granted access to a database of embryo time-lapse images provided by the Hospital Clínic de Barcelona, which includes a secure protocol for data transference and storage. This chapter describes the process of assembling the dataset and preparing the data for model training in the subsequent chapters. The statistics of the dataset can be found in Appendix B. 3.1 Metadata Each embryo in the hospital’s database is stored with a unique ID, with records dating back to 2014 and continuing to the present (November 2024). In total, 53,123 embryos are registered in the database. Of these, 48,714 embryos have metadata available, which are considered for further analysis. Due to data privacy regulations, patient information was excluded. Only the unique EmbryoID was retained. As a result, the dataset contains no personal details such as patient age or ethnicity - this information was not available in the original data. Additionally, it is not possible to determine whether multiple embryos originated from the same patient. The metadata includes the following information: • Pronuclear (PN) Number: This value indicates the number of pronuclei in the embryo after insemination. A PN of 2 is considered normal, while other values are classified as abnormal and may suggest poor quality. • Instrument Number: The hospital uses five different time-lapse incubators, of two types: EmbryoScope+ and EmbryoScope-D. These systems differ in image resolution - EmbryoScope+ produces 800 × 800 pixel images, while EmbryoScope-D produces 500 × 500 pixel images. The cropping of the well also differs: in the EmbryoScope+, the well appears larger. Optical bias may also be introduced due to differences in the imaging systems, as well as environmental variation resulting from architectural differences between the incubators. • Embryo Fate: This parameter denotes the final status of the embryo. Possible outcomes include: –Avoid: The embryo was discarded. –Freeze: The embryo was deemed viable and frozen for future use. 15 Chapter 3. Dataset –Unknown: The fate was either not annotated or labeled as ’unknown’. –Transfer: The embryo was considered high quality and was implanted. –Undecided: It was unclear whether the embryo was of sufficient quality. • Morphokinetic Annotations: These annotations mark the time points of key events in embryo development. Below is a summary of all morphokinetic time points, where each of them except tDead must appear in chronological order. –tPB2: Time of the second polar body extrusion. –tPNa: Time when pronuclei appear. –tPNf: Time when pronuclei fade. –tn ( t2, t3, t4, . . . , t9 ): Represents the time points when the embryo undergoes successive cell divisions, where t2 indicates the first division (2-cell stage), t3 indicates the 3-cell stage, and so on up to t9. –tSC : Start of compaction, where individual cells begin to merge into a compact structure. –tM: Morula stage, where the embryo consists of a tightly packed mass of cells. –tSB: Start of blastulation, where the first signs of blastocyst formation appear. –tB: Blastocyst formation, indicating the development of a fluid-filled cavity. –tEB : Expanded blastocyst stage, where the blastocyst enlarges in preparation for implantation. –tHB: Hatching blastocyst, when the blastocyst begins to break out of its zona pellucida. –tDead : Time point when doctors determine that the embryo has ceased development. This event is not necessarily in chronological order, as it can be recorded at any time. It is noted that a standardized embryo quality assessment like the Gardener Score is not present in the metadata. Other studies used this metric as the main factor to label the embryos as either a blastocyst with good quality or bad quality, which makes the accuracy and labels comparable to other studies. Furthermore, two critical pieces of information are missing for the embryos: the patient’s age and the fertilization method. The former is a key determinant of blastocyst formation rates, with the percentage of embryos per patient reaching the blastocyst stage decreasing linearly as maternal age increases [ 29 , 30 ]. In contrast, the fertilization method has no significant impact on the blastocyst formation rate or the quality of the resulting blastocyst [ 31 ]. However, IVF-derived embryos reach the blastocyst stage significantly faster than those fertilized via ICSI [ 32 ]. As discussed in chapter 2, the rate of embryonic development serves as an important indicator of successful blastocyst formation. Omitting this information from the model prevents it from distinguishing whether an embryo’s development is abnormally slow or fast or whether the observed differences are simply due to variations in the fertilization method. 16 Chapter 3. Dataset 3.2 Embryo Sample Selection The selection of embryos for this study was constrained by both technical and biological factors. Each embryo consists of approximately 500 frames, captured across multiple focal planes (either 7 or 11, symmetrically distributed around the center focal plane), with spacing of either 15 µm or 25 µm depending on the EmbryoScope imaging system. Due to security protocols, downloading a single embryo took approximately one minute, making the retrieval of the full dataset infeasible within the available time. Furthermore, after the initial selection described in this chapter, technical restrictions prevented further access to the hospital database, limiting the dataset to the embryos already obtained. To ensure that only relevant embryos were included, all embryos that did not reach the 72-hour mark according to the annotations were excluded, as this was the required prediction time point. Figure 3.1: Most recent time point of annotation in hours for all 48,714 embryos. A black line is added at the prediction time of 72 hours. All embryos left of the line were discarded. Since images were not yet downloaded at this stage, the developmental time of each embryo was estimated based on the timestamp of its most recent annotation. If no further annotations were recorded, it was assumed that the embryo did not progress beyond that point. After applying this filter, 17,140 embryos remained, representing 35.1% of the original 48,714 embryos with valid metadata. These embryos were downloaded from the Hospital and were available for the dataset. This selection process introduces a potential bias in the dataset. The absence of annotations beyond 72 hours could be due to several reasons: • Pre-72h Arrest: Some embryos failed to develop before reaching 72 hours and were considered non-viable early on. Since these cases do not require model-based decisions at day 3, their exclusion was valid. • Transfer/Cryopreservation: Some embryos were transferred to a patient or cryopreserved before reaching the blastocyst stage at day 3. While these embryos were likely of good quality at 72 hours, their final blastocyst developmental outcome remains unknown and they should be excluded. • Discarded at 72h: Some embryos may have been deemed non-viable just at 72 hours, leading to the absence of further annotations. If imaging data at 72 hours existed for these embryos, they could have been labeled as False cases. However, since they were not included, the model was never trained on embryos that were explicitly discarded at this stage. This results in a dataset where the model only learns from embryos that appeared viable at 72 17 Chapter 4. Methods RadImageNet provides a potentially more effective starting point for feature extraction than natural image datasets. All models in this study are evaluated under three initialization settings: with ImageNet-pretrained weights, with RadImageNet-pretrained weights, and with random initialization. This allows for a systematic assessment of how different pretraining strategies influence performance on the given task. Image Assembly The dataset comprises time-lapse sequences of grayscale microscopic images captured at various focal planes. Using a single image as input to the CNN networks discussed in section 4.2.1 is not directly possible, since most convolutional neural networks for image classification are designed for 3-channel RGB inputs. To enable compatibility with such architectures, the single-channel grayscale images must be expanded to three channels. Several strategies are evaluated for this purpose, each offering different forms of spatial or temporal context: 1. Static grayscale: The central focal plane is duplicated across all three channels. This configuration does not add any additional spatial or temporal context. 2. Focal planes: For each frame, the central focal plane is combined with one plane above and one below, forming a three-channel input that captures axial depth information. The interplane spacing is either ± 25 µm or ± 15 µm, depending on the imaging system (EmbryoScope+ or EmbryoScope D, respectively). 3. Temporal sequence: The current frame at prediction time is combined with the two preceding frames from the time-lapse sequence. This configuration introduces short-term temporal information and captures subtle embryonic movements occurring over a time span of approximately 15 minutes to one hour, depending on the image acquisition interval. Grad-CAM Activation Maps To improve interpretability of the CNNs used in this study, Gradient-weighted Class Activation Mapping (Grad-CAM) is employed. Grad-CAM is a widely used visualization technique that highlights the regions of an input image that contribute most to a model’s decision [ 52 ]. It does so by using the gradients of the target class flowing into the final convolutional layer to produce a coarse localization map of important regions. For the trained models, Grad-CAM is applied to a random batch of validation images in order to qualitatively assess which image regions the model focuses on when making a classification decision. This is particularly valuable in biomedical imaging, where model decisions should ideally correlate with biologically meaningful structures. The resulting heatmaps are overlaid on the original input images, enabling visual inspection of whether the network’s focus aligns with expected morphological features. In this study, Grad-CAM is applied using the final convolutional layer of each model, for the EfficientNetB0-Model it is namely the features[7] layer, as this region of the model retains spatial information while still incorporating high-level semantic features. All visualizations are generated post hoc, without modifying the network architecture or affecting model performance. This method provides an intuitive and model-agnostic approach to explainability, complementing evaluation metrics with qualitative insight into the model’s decision-making process. 24 Chapter 4. Methods Augmentation Techniques During the course of this work, several issues within the dataset were identified: •Random obscuration resulting in very low brightness in certain frames •Embryo wells not centered or partially cropped •Out-of-focus frames •JPEG compression artifacts Although such cases are relatively rare, they pose a risk of overfitting, as the model may easily memorize these atypical samples due to their distinctive visual features. Given the large size of the dataset and the limited time available, manual inspection was not feasible. Instead, augmentation techniques are introduced to mimic the inconsistencies found in the outlier images. The goal was to mitigate overfitting on atypical samples and increase the generalization of the model. Figure 4.2: Grad-CAM Overlay of the model prediction on a selection of samples (column-wise). The image on the left includes blur on the right side of the frame and the model is focusing on the blur outside of the well. The two frames to the right show normal Grad-CAMs focusing correctly on the embryo, how it was observed for most examples. On the right, two frames with obscuration are shown, where as well the model does not focus on the embryo even tho the contours are slightly visible. Another issue is the presence of non-uniformity outside of the embryo well, which may contain distinct features and introduces bias. Grad-CAM visualizations revealed that the model occasionally focused on these external regions rather than the embryo itself. To address this issue, an algorithm was developed to black out the area outside the well and center the well within the image. The procedure is described in Appendix C. However, due to the dataset’s size and variability, a number of samples lack a detected well circle and remain unprocessed. Applying cropping to all frames in such cases led to a decrease in performance. Therefore, cropping was applied with a probability of 50%. This probabilistic approach improved performance while providing two key advantages. First, images without a detected well are not particularly standing out, preventing the model from overfitting on these few unique examples. By not cropping every image, the model learns to remain robust to both cropped and uncropped inputs. Second, the inference dataset does not require cropping, as the 25 Chapter 4. Methods model performs well on unprocessed inputs. This flexibility simplifies deployment and allows the model to handle a broader range of input formats. 4.2.2 Embryo Feature Model Figure 4.3: Overview of the Embryo Feature Model, combining the annotations from the medical experts with image input. The feature vector has a length of 24, as described in Section 3.4. The image input consists of one frame sampled 0 to 40 frames before prediction time at 72 HPI and the two preceding frames. Yellow blocks indicate pretrained components. Figure 4.4: Scaled Dot Product attention block from figure 4.3 in detail. Query (Q), Key (K) and Value (V) are linear projections from the inputs A or B. Next, it was evaluated whether the annotation features (see Section 4.1) could enhance model performance when combined with an image input. While the annotations capture the morphokinetics, they do not contain morphological information. Conversely, a single image primarily reflects morphology, although it may implicitly convey developmental speed based on the embryo’s appearance at prediction time. 26 Chapter 4. Methods To integrate both information sources, the annotation feature vector was provided as a secondary input to the model, processed by a dedicated subnetwork alongside a single image input handled by the backbone CNN, as it was described in the previous section. The subnetwork consisted of a single fully-connected (FC) layer with 1024 dimensions. It was pretrained following the same strategy outlined in Section 4.1, but implemented as a neural network with a single output node rather than a traditional ML model. The key challenge lies in effectively fusing the outputs of both networks. Several fusion strategies were evaluated: • Concatenation: The representations from the image and feature modalities are concatenated and passed to a classifier. Additionally, a ModalityDropout variant was explored, where one modality is randomly masked during training. This encourages the model to learn robust representations from both modalities and reduces over-reliance on a single input source. • Mixture of Experts (MoE): A gating network processes the concatenated modality inputs and learns to assign weights to each modality-specific input. The final representation is a weighted sum of the outputs from the two networks, allowing the model to dynamically prioritize information from different modalities depending on the input [53]. • Attention-Based Fusion: The two outputs are passed through an attention block [ 42 ]. Each modality (A and B, with 1024-dimensional inputs) is linearly projected to a query, key, and value space with 512 hidden dimensions. Cross-attention is computed by performing scaled dot-product attention from Query A to Key B and vice versa, each followed by a Softmax operation. These attention maps are then applied to the corresponding value projections (Value B and Value A). A skip connection adds the original modality inputs back to their respective attention outputs. However, the inputs are also projected to 512 dimensions in order to match the dimensions. The resulting representations are layer-normalized and concatenated. The scheme can be seen in figure 4.3 and 4.4. It is additionally explored if a shared FC layer (as seen in figure 4.3) improves the performance of the model. This layer enables joint representation learning by projecting both modalities - image and annotation features - into a common latent space. Such shared layers have been shown to improve multimodal alignment and reduce modality-specific overfitting in related work [54]. 4.2.3 Embryo Video Model Figure 4.5: Overview of the complete Embryo Video Model. The input consists of 20 frames per sample, each processed into a 1280-dimensional feature vector by the EfficientNetB0 backbone. The final output is a single classification prediction. 27 Chapter 4. Methods The temporal evolution of an embryo contains significantly more information than a single image can provide. As shown in Section 5.2.2, combining morphokinetic annotations with image input yielded the best performance so far. However, these annotations must be manually extracted by trained professionals, introducing subjectivity and manual labour. The objective is to develop a model capable of making predictions automatically, without relying on manually curated features. To implicitly capture morphokinetic characteristics, sequential image data from the time-lapse system is leveraged. A LSTM network is employed in conjunction with a backbone feature extractor to process the temporal dynamics of embryo development. LSTMs are a class of recurrent neural networks designed to model sequential data by maintaining an internal memory of prior inputs. Unlike models that treat each input independently, LSTMs can capture temporal dependencies by updating a hidden state ht and a cell state ct at each time step t . This is achieved through gated mechanisms - namely, the input, forget, state candidate and output gates - which regulate the flow of information and enable the network to retain or discard specific features over time. These properties help mitigate challenges such as the vanishing gradient problem, allowing the model to learn long-range temporal patterns effectively [ 55 ]. Consequently, LSTMs are well-suited for modeling embryo development, where the timing and order of morphokinetic events are essential for accurate prediction. Figure 4.6: Detailed structure of the modified LSTM cell for the Video Model. The different gates are indicated by dotted outlines. The symbol σ denotes the sigmoid activation function, while tanh represents the hyperbolic tangent. brefers to bias vectors, Wto input weight matrices, and R to recurrent weight matrices. The vectors h t−1 and c t−1 represent the hidden and cell states from the previous time step, respectively, and x t is the input feature vector at the current time step t. The best-performing EfficientNetB0 model from Section 4.2.1 is used as the feature extractor, producing a 1,280-dimensional feature vector for each frame in the time-lapse sequence. A total of 20 frames are sampled at uniform intervals and passed sequentially to the LSTM network, with the final frame corresponding to the prediction time at 72 HPI. Due to computational and runtime constraints, 28 Chapter 4. Methods incorporating more frames was not feasible. To investigate the effect of temporal resolution, different frame spacing strategies were tested. For classification, only the last hidden state ht of the LSTM - corresponding to the final time step - is used as input to the classification head. In the final architecture, the backbone, LSTM, and classifier head were trained jointly in an end-to-end manner, meaning that the backbone was not simply acting as a frozen feature extractor. This significantly increased the computational load; however, it allowed the backbone to adapt to embryo appearances at previously unseen time points, as the pretrained backbone had originally only been trained with images captured near the prediction time. Initial experiments showed significant overfitting, due to increased model complexity and reduced data augmentation relative to the information density in longer time sequences. After extensive experiments, improved generalization and stable training dynamics were achieved through the following regularization techniques: •Single-layer LSTM with 512 hidden units •Dropout rate of 0.2 applied to the LSTM hidden state ht •Layer normalization applied to both the LSTM cell output ctand hidden state ht •Weight decay of 1×10−5applied to the input-to-hidden weights Wf, Wi, Wo, Wg •A reduced fully connected layer in the classifier head with 256 hidden units •Dropout of 0.2 applied to the fully connected layer (as already applied before) 29 Chapter 5 Results In this chapter, a series of experiments are conducted to compare different models, data augmentation strategies, and hyperparameter configurations. All results will be jointly analyzed in chapter 7 and are summarized in table 7.1. For overall model selection, balanced accuracy is used as the primary evaluation metric. This choice reflects the fact that there is no requirement to favor either the positive or negative class - the objective is to optimize general performance across both. Balanced accuracy is defined as the average of sensitivity (true positive rate) and specificity (true negative rate), and is computed as follows: Balanced Accuracy =1 2T P T P +F N +T N T N +F P (5.1) 5.1 Machine Learning Model Experiments were conducted using Python 3.10.15 on an Apple M1 Mac. The RandomForestClassifier from scikit-learn v1.6.1 [ 56 ] was trained using default settings. As input, the preprocessed feature vector described in Section 3.4 is used. Results are shown in table 5.1. Metric Validation Set Test Set Accuracy 0.7203 0.7601 Sensitivity 0.7695 0.8000 Specificity 0.6747 0.7230 Balanced Accuracy 0.7221 0.7615 Table 5.1: Performance metrics for the machine learning model ( RandomForestClassifier ) on the test and validation sets. 5.2 Deep Learning Models All deep learning models were implemented using PyTorch Lightning v2.4.0 [ 57 ], a high-level wrapper for PyTorch v2.5.0 (pre-release, NVIDIA build nv24.10) [ 58 ]. Training was conducted on the 31 Chapter 5. Results MareNostrum5 supercomputer, equipped with high-performance computing (HPC) clusters with four NVIDIA H100 GPUs (each with 64GB VRAM), using Python 3.10.12 and CUDA compilation tools release 12.6 (V12.6.77). Distributed training was performed using Lightning’s Distributed Data Parallel (DDP) backend, with mixed-precision training enabled via the 16-mixed precision setting. All batch normalization layers were converted to synchronized batch normalization to ensure consistent statistics across devices. Training was conducted for a maximum of 100 epochs, with early stopping applied using a patience of 30 epochs monitoring the balanced validation accuracy. For each run, the model with the highest balanced accuracy on the validation set was selected as the final model. Final validation and test predictions were performed on a single GPU, and all reported metrics in tables represent the mean over multiple training runs with different random seeds. 5.2.1 Single Image Model Results A single frame per embryo is used as input for classification. A random frame is selected from within a window of 40 frames prior to the frame at 72 HPI. This window can span up to 800 minutes, depending on the image acquisition frequency. Initial experiments showed optimal performance using this configuration. The following basic transformation and augmentation pipeline is then applied: •The image is converted to a float32 tensor with three channels. • Color jittering is applied using torchvision’s transforms.v2.ColorJitter with brightness set to 0.2 and contrast set to 0.3. •The image is resized to 512 ×512 pixels. •Random horizontal flipping is applied with a probability of 0.5. •A random fixed rotation of 0°, 90°, 180°, or 270°is applied. • The image is normalized using ImageNet statistics with mean = [0.485, 0.456, 0.406] and standard deviation = [0.229, 0.224, 0.225]. All models are implemented via the torchvision.models package (version 0.20.0a0). For each architecture, the final classification layer is replaced by an identity layer. A custom classification head is appended, consisting of a fully connected layer with a hidden size of 1024, followed by a ReLU activation and a dropout layer with a dropout rate of 0.2. A final linear layer outputs a single logit. The binary classification objective uses torch’s BCEWithLogitsLoss , which provides greater numerical stability than applying a sigmoid activation followed by BCELoss. Additionally, label-smoothing regularization is employed by adding random noise in the range [0.0, 0.1] to the hard labels (0 or 1), following the method described in Section 7 of [ 46 ]. This prevents the model from becoming overly confident and improves generalization. The AdamW optimizer is used, with separate learning rates and weight decay parameters for the backbone and classification head. Weight decay is only applied to the weights of convolutional and linear layers. A warm-up phase of 5 epochs is applied, followed by a cosine annealing learning rate scheduler, which gradually reduces the learning rate to a common minimum by epoch 100. 32 Chapter 5. Results The following hyperparameters were found to produce good results across the evaluated models: •Learning rate (backbone): 1×10−4 •Learning rate (classification head): 1×10−3 •Weight decay (backbone): 1×10−6 •Weight decay (classification head): 1×10−3 •Minimum learning rate: 1×10−5 •Batch size: 32 Backbone selection For initial selection, runs where performed with focal-plane images (table 5.2). However, as seen in table 5.3, the Temporal Sequence performed better overall. Only EfficientNetB0 pretrained on ImageNet will considered as a starting point for further experiments. Model None ImageNet RadImageNet ResNet50 59.09% 68.88% 63.74% EfficientNetB0 59.33% 70.03% - ConvNeXt-Base 53.09% 68.68% - SwinV2-B 53.03% 68.27% - InceptionV3 64.30% 68.30% 61.49% DenseNet121 59.80% 69.10% 61.34% VGG16 55.00% 68.03% - Table 5.2: Comparison of model performance (balanced accuracy on validation set) across different pretrained weights, either random weights, ImageNet weights or RadImageNet weights (where available). EfficientNetB0 performed the best, initialized with ImageNet weights. Metrics represent the mean of five training runs and images are assembled using three focal planes. Input Type Balanced Accuracy Grayscale (Duplicated) 69.31% Focal Planes 70.03% Temporal Sequence 71.41% Table 5.3: Performance of EfficientNetB0 (ImageNet pretrained) using different image input configurations. Temporal Sequence (prediction frame plus the two previous frames) performed the best overall. Metrics are the mean of five runs on the validation set. Augmentation Techniques Due to the irregularities in the dataset described in section 4.2.1, five new augmentation techniques were incorporated into the basic transformation pipeline introduced in Section 5.2.1: •Probabilistic cropping of the embryo well (50% chance), as described in Appendix C 33 Chapter 7. Discussion and Future Work Figure 7.1: Grad-CAM overlays of two frames from an LSTM input sequence. The left image corresponds to the 4-cell stage and the right to the 2-cell stage. The highlighted regions align with individual blastomeres, suggesting the model is implicitly estimating cell count from morphology. 7.2 Using the Annotations as an Additional Learning Task Although the Embryo Video Model performed well, it did not surpass the Machine Learning Model trained solely on manual annotations. This suggests that the Video Model alone does not fully extract all morphokinetic information available in the dataset. That would imply that clinical annotations - while subject to human variability - still carry strong predictive value. Interestingly, despite not being explicitly trained to identify cell stages, the Grad-CAM results indicate that the model partially learns this task (Figure 7.1). Two potential improvements for future work would be thinkable to incorporate the strong informative value of the manual annotations to the dataset, while training a model which still functions fully automatic: • Multi-task Learning: The backbone of the LSTM model could be extended to simultaneously predict the morphokinetic time points (e.g., t2 , t3 , etc.) alongside feature extraction. This could be achieved by adding an auxiliary loss function to classify each frame’s developmental stage, steering the model to focus even more on the morphokinetic stage of the embryo. • Automatic Feature Extraction: An alternative would be to assess the model’s ability to automatically predict the morphokinetic stage and use these predictions to assemble the feature vector from section 4.1 as input to the Machine Learning Model. As discussed in chapter 2, previous work suggests that classification accuracies of up to 88% can be achieved for frame-wise developmental prediction. It would be interesting to compare this automatic pipeline to manual annotations in terms of both performance and consistency. These strategies could help bridge the performance gap between fully automated models and those relying on manual annotations. 7.3 Missing Metadata The current dataset lacks key metadata that could enhance model performance - specifically, fertilization method and patient age. Including whether an embryo was fertilized via ICSI is important, as it influences the timing of early development. Without this information, the model may misinterpret naturally fast development, an indicator of high developmental potential, when it may simply be due to the ICSI procedure. 40 Chapter 7. Discussion and Future Work Similarly, patient age is a well-established factor affecting blastocyst formation rates, with advanced maternal age associated with lower developmental potential. Providing the model with age data would allow it to contextualize embryonic development patterns more accurately, potentially improving both prediction accuracy and clinical applicability. 7.4 Labelling process To improve predictive accuracy, the labelling process should integrate the Gardner scoring system, which explicitly evaluates blastocyst expansion as well as the quality of the ICM and TE. With the current approach (Appendix B), embryo usability is inferred from retrospective clinical outcomes, without direct reference to the Gardner scores - introducing potential label noise due to the variability in clinical decision-making. Using the Gardner score as a labelling target would provide a more standardized and biologically grounded signal for training. However, this requires manual annotation by expert embryologists, which is both time-intensive and resource-demanding. A promising strategy, as demonstrated by the STEM and STEM+ models in [ 22 ], is to decouple the prediction task into two stages: (1) predicting whether an embryo will reach the blastocyst stage, and (2) classifying the quality of the resulting blastocyst. This separation could be implemented via two independent models or a multitask architecture with dual outputs. Reaching the blastocyst stage is typically easier to annotate and less susceptible to subjectivity, as it can be determined visually with relative confidence. In contrast, grading blastocyst quality is inherently more subjective, making it more prone to inter-observer variability. By splitting these tasks, the learning signal becomes more specific and less noisy, which could help the model converge more effectively. Moreover, this separation has practical advantages in clinical settings. It enhances interpretability of the model output, especially important in cases with limited embryos, where even low-quality blastocysts would be used, since they may result in successful pregnancies. 7.5 Enhancing the Dataset Size As discussed in section 3.2, embryos that were discarded by clinicians at 72 hours post-fertilization were excluded from the dataset. Consequently, the model was trained solely on embryos that appeared viable at the 72-hour mark. This introduces a selection bias, potentially limiting the model’s ability to generalize to embryos deemed low-quality at that stage. Investigating the model’s behavior on such cases remains an important direction for future work. However, this is inherently challenging, as discarded embryos were not cultured further and thus lack developmental outcome labels, making accurate labelling difficult. Notably, the largest dataset to date, assembled by Lassen et al. [ 20 ], included discarded embryos in the training process. This approach improved the model’s ability to distinguish between embryos suitable for transfer and those likely to be discarded, helping to mitigate bias toward high-quality embryos with known implantation outcomes. In addition, increasing the dataset size could significantly improve the model’s generalizability. As of November 2024, approximately 10% of the data was unavailable due to a technical issue with the hospital connection. Also, new data continues to be generated, as the clinic remains active in conducting IVF procedures. 41 Chapter 7. Discussion and Future Work Furthermore incorporating data from other clinics, would enable broader evaluation of the model’s generalizability across different clinical environments. Decoupling the prediction of blastocyst formation from quality assessment, as discussed previously, also enables the use of public datasets such as the one from [ 28 ], which includes annotations for blastocyst formation but lacks quality grading. This would enable a standardized assessment across various studies and models. 7.6 Assessing other models The dataset used in this work is comparable in size to that of [ 22 ], which is the largest used for the classification task at hand, providing a solid foundation for exploring more recent and advanced model architectures. Approaches such as Video Transformers or I3D networks offer the potential to capture more complex temporal dynamics from time-lapse embryo development sequences. Moreover, the public availability of the STEM model code offers an opportunity for direct comparison with the models developed in this study. Evaluating the STEM model on this dataset would be an important assessment for future work. 42 Chapter 8 Conclusions In this Master’s Thesis, a new dataset of embryo time-lapse sequences was assembled to train datadriven AI models aimed at assisting clinicians in predicting the likelihood of successful blastocyst development. The best-performing model, based on EfficientNetB0 and enhanced with morphological annotations through a feature fusion mechanism, achieved a balanced accuracy of 77.4%. In parallel, a fully automated video-based model incorporating an LSTM network reached a balanced accuracy of 75.3%, demonstrating competitive performance without relying on clinical input. It was shown that the human-annotated morphokinetics are a very strong predictor of blastocyst development, and fully automated models based on images were not able to match their performance completely. Assembling and labelling the dataset posed significant challenges. Nonetheless, the distribution of labels closely reflects natural blastocyst formation rates, suggesting a realistic and reasonable labelling process. With over 10,000 embryo time-lapse sequences and approximately 5 million images, this dataset is one of the largest to date used for the task of blastocyst formation prediction. Compared to other studies using sufficiently large datasets, the performance metrics achieved are on par with, and in some cases exceed, current state-of-the-art results. However, direct comparison remains difficult due to differences in datasets and methodologies. Still, the results support the validity of both the labelling process and dataset assembly. As the first study using this dataset, the outcomes are encouraging. The Grad-CAM visualizations further support the model’s reliability, showing focused attention on the embryos rather than surrounding artifacts, suggesting limited bias. However, more work and clinical integration has to be done to ensure the model can be used as a clinical support tool in the context of IVF-treatments and to ensure that ethical fairness and generalization are achieved. This work lays a strong foundation for future research on predicting high-quality blastocyst development. In combination with the recommendations discussed in chapter 7, even better results are feasible. The developed models are intended for future deployment in real-world clinical settings and are also planned for public release - potentially becoming the first publicly available models for this task. 43 Bibliography [1] Ashley M. Eskew and Emily S. Jungheim. A history of developments to improve in vitro fertilization. Missouri Medicine, 114(3):156–159, May-Jun 2017. [2] N. Borumandnia, H. Alavi Majd, N. Khadembashi, and H. Alaii. Worldwide trend analysis of primary and secondary infertility rates over past decades: A cross-sectional study. International Journal of Reproductive BioMedicine, 20(1):37–46, Feb 2022. [3] Alexander R.P. Walker, Betty F. Walker, and Fatima Adam. Nutrition, diet, physical activity, smoking, and longevity: From primitive hunter-gatherer to present passive consumer—how far can we go? Nutrition, 19(2):169–173, 2003. [4] Xiaofeng Ye, Xiaoxia Song, Sihang Zhou, Guoqing Chen, and Liping Wang. Association between combined healthy lifestyles and infertility: a cross-sectional study in us reproductive-aged women. BMC Public Health, 25(1):153, 2025. [5] Andrea C. Gore, Valerie A. Chappell, Suzanne E. Fenton, Jodi A. Flaws, Angel Nadal, Gail S. Prins, Jorma Toppari, and R. Thomas Zoeller. EDC-2: The Endocrine Society’s Second Scientific Statement on Endocrine-Disrupting Chemicals. Endocrine Reviews, 36(6):E1–E150, Dec 2015. Epub 2015 Nov 6. [6] Bradley J. Van Voorhis. In vitro fertilization. New England Journal of Medicine, 356(4):379–386, 2007. [7] Patrizia Rubino, Paola Viganò, Alice Luddi, and Paola Piomboni. The icsi procedure from past to future: a systematic review of the more controversial aspects. Human Reproduction Update, 22(2):194–227, 11 2015. [8] Kirstine Kirkegaard, Aishling Ahlström, Hans Jakob Ingerslev, and Thorir Hardarson. Choosing the best embryo by time lapse versus standard morphology. Fertility and Sterility, 103(2):323–332, 2015. [9] J. Conaghan, A.A. Chen, S.P. Willman, K. Ivani, P.E. Chenette, R. Boostanfar, V.L. Baker, G.D. Adamson, M.E. Abusief, M. Gvakharia, K.E. Loewke, and S. Shen. Improving embryo selection using a computer-automated time-lapse image analysis test plus day 3 morphology: results from a prospective multicenter trial. Fertility and Sterility, 100(2):412–419.e5, 08 2013. Epub 2013 May 28. [10] M. Salih, C. Austin, R. R. Warty, et al. Embryo selection through artificial intelligence versus embryologists: a systematic review. Human Reproduction Open, 2023(3):hoad031, 2023. Published 2023 Aug 15. [11] Connie C. Wong, Kevin E. Loewke, Nancy L. Bossert, Barry Behr, Christopher J. De Jonge, 45 Bibliography Thomas M. Baer, and Renee A. Reijo Pera. Non-invasive imaging of human embryos before embryonic genome activation predicts development to the blastocyst stage. Nature Biotechnology, 28(10):1115–1121, 2010. [12] Kuan-Chong Lan, Feng-Jung Huang, Yu-Ching Lin, Fu-Tsai Kung, Chian-Huey Hsieh, Hsin-Wen Huang, Ping-Hui Tan, and Szu-Yuan Chang. The predictive value of using a combined z-score and day 3 embryo morphology score in the assessment of embryo survival on day 5. Human Reproduction, 18(6):1299–1306, Jun 2003. [13] Hui Zou, Jason M Kemper, Elizabeth R Hammond, Fei Xu, Guanghui Liu, Li Xue, Xi Bai, Hong Liao, Shuang Xue, Shuai Zhao, Lei Xia, John Scott, Vanessa Chapple, Mo Afnan, Dean E Morbeck, Ben Willem J Mol, Ying Liu, and Rui Wang. Blastocyst quality and reproductive and perinatal outcomes: a multinational multicentre observational study. Human Reproduction, 38(12):2391–2399, Dec 2023. [14] Nathan H Ng, Julian McAuley, Julian A Gingold, Nina Desai, and Zachary C Lipton. Predicting embryo morphokinetics in videos with late fusion nets & dynamic decoders, 2018. [15] Reza Moradi Rad, Parvaneh Saeedi, Jason Au, and Jon Havelock. Blastomere cell counting and centroid localization in microscopic images of human embryo. 2018 IEEE 20th International Workshop on Multimedia Signal Processing (MMSP), pages 1–6, 2018. [16] Chongwei Wu, Wei Yan, Hongtu Li, Jiaxin Li, Hongkai Wang, Shijie Chang, Tao Yu, Ying Jin, Chao Ma, Yahong Luo, Dongxu Yi, and Xiran Jiang. A classification system of day 3 human embryos using deep learning. Biomedical Signal Processing and Control, 70:102943, 2021. [17] Pegah Khosravi, Ehsan Kazemi, Qiansheng Zhan, Jonas E. Malmsten, Marco Toschi, Pantelis Zisimopoulos, Alexandros Sigaras, Stuart Lavery, Lee A. D. Cooper, Cristina Hickman, Marcos Meseguer, Zev Rosenwaks, Olivier Elemento, Nikica Zaninovic, and Iman Hajirasouliha. Deep learning enables robust assessment and selection of human blastocysts after in vitro fertilization. NPJ Digital Medicine, 2:21, 04 2019. eCollection 2019. [18] M. VerMilyea, J.M.M. Hall, S.M. Diakiw, A. Johnston, T. Nguyen, D. Perugini, A. Miller, A. Picou, A.P. Murphy, and M. Perugini. Development of an artificial intelligence-based assessment model for prediction of embryo viability using static images captured by optical light microscopy during ivf. Human Reproduction, 35(4):770–784, 04 2020. [19] S. M. Diakiw, J. M. M. Hall, M. VerMilyea, A. Y. X. Lim, W. Quangkananurug, S. Chanchamroen, B. Bankowski, R. Stones, A. Storr, A. Miller, G. Adaniya, R. van Tol, R. Hanson, J. Aizpurua, L. Giardini, A. Johnston, T. Van Nguyen, M. A. Dakka, D. Perugini, and M. Perugini. An artificial intelligence model correlated with morphological and genetic features of blastocyst quality improves ranking of viable embryos. Reproductive Biomedicine Online, 45(6):1105–1117, 2022. [20] Jacob Theilgaard Lassen, Mikkel Fly Kragh, Jens Rimestad, Martin Nygård Johansen, and Jørgen Berntsen. Development and validation of deep learning based embryo selection across multiple days of transfer. Scientific Reports, 13(1):4235, 2023. [21] Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset, 2018. [22] Qiuyue Liao, Qi Zhang, Xue Feng, Haibo Huang, Haohao Xu, Baoyuan Tian, Jihao Liu, Qihui Yu, Na Guo, Qun Liu, Bo Huang, Ding Ma, Jihui Ai, Shugong Xu, and Kezhen Li. Development 46 Bibliography of deep learning algorithms for predicting blastocyst formation and quality by time-lapse monitoring. Communications Biology, 4(1):415, 03 2021. [23] G. Coticchio, G. Fiorentino, G. Nicora, R. Sciajno, F. Cavalera, R. Bellazzi, S. Garagna, A. Borini, and M. Zuccotti. Cytoplasmic movements of the early human embryo: imaging and artificial intelligence to predict blastocyst development. Reprod Biomed Online, 42(3):521–528, 03 2021. Epub 2020 Dec 24. [24] Manoj Kumar Kanakasabapathy, Prudhvi Thirumalaraju, Charles L Bormann, Raghav Gupta, Rohan Pooniwala, Hemanth Kandula, Irene Souter, Irene Dimitriadis, and Hadi Shafiee. Deep learning mediated single time-point image-based prediction of embryo developmental outcome at the cleavage stage, 2020. [25] T. Chen et al. Automating blastocyst formation and quality prediction in time-lapse imaging with adaptive key frame selection. In L. Wang, Q. Dou, P.T. Fletcher, S. Speidel, and S. Li, editors, Medical Image Computing and Computer Assisted Intervention - MICCAI 2022, volume 13434 of Lecture Notes in Computer Science, pages 470–480. Springer, Cham, 2022. [26] Astrid Zeman, Anne-Sofie Maerten, Annemie Mengels, Lie Fong Sharon, Carl Spiessens, and Hans Op de Beeck. Deep learning for human embryo classification at the cleavage stage (day 3). In Alberto Del Bimbo, Rita Cucchiara, Stan Sclaroff, Giovanni Maria Farinella, Tao Mei, Marco Bertini, Hugo Jair Escalante, and Roberto Vezzani, editors, Pattern Recognition. ICPR International Workshops and Challenges, pages 278–292, Cham, 2021. Springer International Publishing. [27] Kanak Kalyani and Parag S. Deshpande. A deep learning model for predicting blastocyst formation from cleavage-stage human embryos using time-lapse images. Scientific Reports, 14(1):28019, 2024. [28] Tristan Gomez, Magalie Feyeux, Justine Boulant, Nicolas Normand, Laurent David, Perrine Paul-Gilloteaux, Thomas Fréour, and Harold Mouchère. A time-lapse embryo dataset for morphokinetic parameter prediction. Data in Brief, 42:108258, 2022. [29] Phillip A. Romanski, Ashley Aluko, Pietro Bortoletto, Rony T. Elias, and Zev Rosenwaks. Agespecific blastulation rates in embryo cryopreservation cycles yielding a cryopreserved blastocyst. Fertility and Sterility, 116(1):e11–e12, 2021. [30] R. Sainte-Rose, C. Petit, L. Dijols, C. Frapsauce, and F. Guerif. Extended embryo culture is effective for patients of an advanced maternal age. Scientific Reports, 11(1):13499, 2021. [31] Lisbet Van Landuyt, Anick De Vos, Hubert Joris, Greta Verheyen, Paul Devroey, and André Van Steirteghem. Blastocyst formation in in vitro fertilization versus intracytoplasmic sperm injection cycles: Influence of the fertilization procedure. Fertility and Sterility, 83(5):1397–1403, 2005. [32] Barbara Speyer, Helen O’Neill, Wael Saab, Srividya Seshadri, Suzanne Cawood, Carleen Heath, Matthew Gaunt, and Paul Serhal. In assisted reproduction by ivf or icsi, the rate at which embryos develop to the blastocyst stage is influenced by the fertilization method used: a split ivf/icsi study. Journal of Assisted Reproduction and Genetics, 36(4):647–654, 2019. [33] Mariana Nicolielo, Catherine Jacobs, Andrea Belo, Ana Paula Reis, Renata Erberelli, Fabiana Mendez, Marina Fanelli, Livia Cremonesi, Paulo Cesar Serafini, Eduardo L.A. Motta, Aline R. Lorenzon, and Jose Roberto Alegretti. Embryo culture in time-lapse system provides better 47 Bibliography rates of blastocyst formation, decreases embryo development arrest rate compared to traditional triple-gas culture system. Fertility and Sterility, 112(3):e125–e126, 2019. [34] Alexandra J Kermack, Irina Fesenko, David R Christensen, Kate L Parry, Philippa Lowen, Susan J Wellstead, Scott F Harris, Philip C Calder, Nicholas S Macklon, and Franchesca D Houghton. Incubator type affects human blastocyst formation and embryo metabolism: a randomized controlled trial. Human Reproduction, 37(12):2757–2767, Oct 2022. [35] Gary Bradski. Opencv: Open source computer vision library. https://opencv.org/ , 2000. Accessed: 2025-03-21. [36] Samuel Smith. pytesseract: Python-tesseract. https://github.com/madmaze/pytesseract , 2023. Accessed: 2025-03-21. [37] Leo Breiman. Random forests. Machine Learning, 45(1):5–32, 2001. [38] Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. Nature, 521(7553):436–444, 2015. [39] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition, 2015. [40] Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization, 2019. [41] Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows, 2021. [42] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2023. [43] Mingxing Tan and Quoc V. Le. Efficientnet: Rethinking model scaling for convolutional neural networks, 2020. [44] Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s, 2022. [45] Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks, 2017. [46] Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2818–2826, 2016. [47] Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger. Densely connected convolutional networks, 2018. [48] Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition, 2015. [49] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009. [50] Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. How transferable are features in deep neural networks?, 2014. 48 Bibliography [51] Xi Mei, Zongyu Liu, Paul M. Robson, Benjamin Marinelli, MingDe Huang, Apurva Doshi, Adam Jacobi, Congyu Cao, Kevin E. Link, Tian Yang, Yiqiu Wang, Hayit Greenspan, Tessa Deyer, Zahi A. Fayad, and Yang Yang. Radimagenet: An open radiologic deep learning research dataset for effective transfer learning. Radiology: Artificial Intelligence, 4(5):e210315, 2022. [52] Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradientbased localization. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 618–626, 2017. [53] Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. Empirical evaluation of gated recurrent neural networks on sequence modeling, 2014. [54] Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Ng. Multimodal deep learning. pages 689–696, 01 2011. [55] Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Comput., 9(8):1735–1780, November 1997. [56] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011. [57] William Falcon and The PyTorch Lightning team. Pytorch lightning. https://github.com/ Lightning-AI/lightning , 2019. The lightweight PyTorch wrapper for high-performance AI research. Scale your models, not the boilerplate. [58] Jason Ansel, Edward Yang, Horace He, Natalia Gimelshein, Animesh Jain, Michael Voznesensky, Bin Bao, Peter Bell, David Berard, Evgeni Burovski, Geeta Chauhan, Anjali Chourdia, Will Constable, Alban Desmaison, Zachary DeVito, Elias Ellison, Will Feng, Jiong Gong, Michael Gschwind, Brian Hirsh, Sherlock Huang, Kshiteej Kalambarkar, Laurent Kirsch, Michael Lazos, Mario Lezcano, Yanbo Liang, Jason Liang, Yinghai Lu, CK Luk, Bert Maher, Yunjie Pan, Christian Puhrsch, Matthias Reso, Mark Saroufim, Marcos Yukio Siraichi, Helen Suk, Michael Suo, Phil Tillet, Eikan Wang, Xiaodong Wang, William Wen, Shunting Zhang, Xu Zhao, Keren Zhou, Richard Zou, Ajit Mathews, Gregory Chanan, Peng Wu, and Soumith Chintala. Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (ASPLOS ’24). ACM, April 2024. [59] Lucia Urcelay, Daniel Hinjos, Pablo A. Martin-Torres, Marta Gonzalez, Marta Mendez, Salva Cívico, Sergio Álvarez Napagao, and Dario Garcia-Gasulla. Exploring the role of explainability in ai-assisted embryo selection, 2023. 49 Appendix B Embryo Selection On the following page, the flowchart for embryo labelling is shown. Each EmbryoID with valid annotations undergoes this process to determine whether it should be included in the dataset. The decision is based solely on the annotations associated with each embryo. Due to the dataset’s large size and the long download times, it was not feasible to download all videos first and then decide based on morphology or the presence of frames at specific time points. The possible outcomes of the decision process are: Discard,True Label, or False Label. Each circle in the flowchart shows, in italic, the number of embryos falling into that category. The decisions reflect the clinical procedures followed at the hospital, with each case evaluated by a practicing clinician. It should also be noted that approximately 100 embryos were manually discarded due to specific issues (e.g., empty wells, occluded embryos). However, manual inspection of every embryo was not feasible across the full dataset. 57 All EmbryoIDs on the embryoscope No Yes Most recent annotation time is at least 72 hours Discard No First image of sequence is not later than 5 hours Discard yes no tDead Label? yes no Embryo has tB, tEB or tHB label yes no tB, tEB and tHB not present? Freeze Transfer Avoid Unknown/Undecided 406 Fate? 1. Bad Embryo Blastocyst No 854 abnormal 2 unknown PN number? 9. Avoided abnormal embryos 326 8. Avoided normal embryos 2583 7. Avoided Unknown Embryos Yes 6. Embryo is a cleavage stage transfer, ground truth not known, Discard 837 yes no tEB Label (expanded)? 2 unknown abnormal PN number? 17. Exp. Blast. with normal PN 148 19. Exp. Blast with abnormal PN 14 18. Exp. Blast with unknown PN 16. Avoid. Blast. not reaching exp. stage 432 avoid unknown freeze / transfer undecided Fate? 2 or unknown 3, 4, 5, .. 1 PN number? 13. Good embryo Blastocyst True 4184 14. Good embryo with abnormal PN 108 no yes tEB label (expanded)? 2. incipient blast. that died 1 Embryo Labelling Decision Tree unknown abnormal 2PN number? 10. Completely Unknown Embryo 12. abnorm. non-fate embryo 118 11. Normal non-fate embryo 1296 2 unknown abnormal 3. exp. blast. that died with normal PN 16 4. exp. blast. that died with unknown PN 5. exp. blast. that died with abnormal PN PN number? yes no tEB Label (expanded)? 15a. Expanded nofate Blastocyst 136 15b. Incipient nonfate Blastocycst 270 Abnormal PN number suggest a mistake in outcome or PN number The embryos did not reach a Blastocyst stage, and were avoided. Generally treat as bad examples, even tho this inlcudes some doctors bias, since the reason could be a preliminary doctor's decision. Abnormal PN number suggests bad quality, therefore used as False samples morphologically not good, little cells, malformed, etc. missing data of the fate, quality of the blastocyst can not be validated Discard Embryos False Label True Label PN=1: Embryos can reach blastocyst stage and are valid 13a. Good embryo with PN=1 108 crossed out No samples in dataset Labeling Outcomes Appendix C Embryo Well Cropping The embryo well cropping technique relies on the Hough Circle Transform to accurately detect the circular well structure in each frame. To ensure robust detection, a series of image preprocessing steps is applied. Based on empirical evaluation, the following pipeline yielded the most reliable results. All methods described here were implemented using the python implementation of the OpenCV library [35]. Figure C.1: Step-by-step preprocessing for Hough Circle detection. From left to right: the original image; erosion filter to enhance contrast between the well and the background; application of the Sobel filter to highlight edges and suppress internal textures; further erosion to remove small artifacts such as embryo outlines; Gaussian blurring to reduce noise and smooth the image for more stable circle detection. After preprocessing, the Hough Circle Transform is applied to detect the well. Detection parameters are adjusted based on the machine type (EmbryoScope+ or EmbryoScope-D) and their corresponding image resolution (800×800 or 500×500 pixels). The minimum and maximum circle radii are tuned accordingly to restrict detection to plausible well sizes. A key parameter in the Hough Circle Transform is the circle detection sensitivity param2 . If no circle is initially detected, the sensitivity threshold is gradually lowered until circles are found or a minimum value is reached, allowing more permissive detection in challenging cases. In most instances, multiple candidate circles are identified. To obtain a stable and consistent estimate, the final well position is computed as the average of all detected circle centers and radii. 59 Appendix C. Embryo Well Cropping Figure C.2: Detected circles (left) and the final averaged circle (right), which is used as the estimated well. Finally, the image is cropped to the region corresponding to the estimated well. All areas outside the circular boundary are masked in black. The resulting image is then resized to match the original dimensions, preserving spatial consistency across the dataset. 60