scieee AI-readable full text Open interactive document viewer

Quantitative Evaluation of Repeatability in Medical Microwave Imaging: A Systematic Review

Lopes, Henrique V.; Godinho, Daniela M.; Conceição, Raquel C.

Abstract

Medical microwave imaging (MMWI) prototypes have demonstrated potential in various clinical applications. For successful clinical translation, these systems must be reliable and yield consistent results, that is, they must demonstrate high repeatability. This systematic review critically appraises the existing literature on repeatability in MMWI, with particular attention to the protocols, metrics, and outcomes reported. A systematic search of Scopus, IEEE Xplore, Web of Science, and PubMed was conducted on April 29th 2025, identifying 21 studies that quantitatively evaluated repeatability in MMWI. Two reviewers independently performed full-text screening and data extraction, with disagreements resolved through consultation with a third reviewer. Descriptive analyses were performed using tables and figures to characterise study design, methodologies, and findings. Considerable heterogeneity was observed in sample size, imaging protocols, and the selection of repeatability metrics. Almost all studies were restricted to breast imaging applications. Frequently reported sources of variability included subject repositioning, tissue changes, and hardware-related factors. Although most studies reported acceptable repeatability, the absence of standardised protocols and metrics limits cross-study comparison and the establishment of benchmarks. The review concludes with recommendations to harmonise approaches that contribute to accelerate the clinical translation of MMWI systems.

Full text

IEEE ENGINEERING IN MEDICINE AND BIOLOGY SOCIETY SECTION Received 21 October 2025, accepted 24 November 2025, date of publication 27 November 2025, date of current version 4 December 2025. Digital Object Identifier 10.1109/ACCESS.2025.3638248 Quantitative Evaluation of Repeatability in Medical Microwave Imaging: A Systematic Review HENRIQUE V. LOPES , (Graduate Student Member, IEEE), DANIELA M. GODINHO , (Member, IEEE), AND RAQUEL C. CONCEIÇÃO , (Member, IEEE) Instituto de Biofisica e Engenharia Biomédica, Faculdade de Ciências, Universidade de Lisboa, 1749-016 Lisbon, Portugal Corresponding author: Henrique V. Lopes ([email protected]) This work was funded by the European Union (3BAtwin, GA 101159623). Views and opinions expressed are however those of the authors only and do not necessarily reflect those of the European Union or the European Research Executive Agency. Neither the European Union nor the granting authority can be held responsible for them. This research was funded in whole or in part by the Fundação para a Ciência e a Tecnologia, I.P. (FCT, https://ror.org/00snfqn58) under Grant UID/00645/2025 (https://doi.org/10.54499/UID/00645/2025). For the purpose of Open Access, the author has applied a CC BY public copyright licence to any Author Accepted Manuscript (AAM) version arising from this submission. ABSTRACT Medical microwave imaging (MMWI) prototypes have demonstrated potential in various clinical applications. For successful clinical translation, these systems must be reliable and yield consistent results, that is, they must demonstrate high repeatability. This systematic review critically appraises the existing literature on repeatability in MMWI, with particular attention to the protocols, metrics, and outcomes reported. A systematic search of Scopus, IEEE Xplore, Web of Science, and PubMed was conducted on April 29th 2025, identifying 21 studies that quantitatively evaluated repeatability in MMWI. Two reviewers independently performed full-text screening and data extraction, with disagreements resolved through consultation with a third reviewer. Descriptive analyses were performed using tables and figures to characterise study design, methodologies, and findings. Considerable heterogeneity was observed in sample size, imaging protocols, and the selection of repeatability metrics. Almost all studies were restricted to breast imaging applications. Frequently reported sources of variability included subject repositioning, tissue changes, and hardware-related factors. Although most studies reported acceptable repeatability, the absence of standardised protocols and metrics limits cross-study comparison and the establishment of benchmarks. The review concludes with recommendations to harmonise approaches that contribute to accelerate the clinical translation of MMWI systems. INDEX TERMS Diagnosis, medical microwave imaging, metrics, repeatability, reproducibility, standardization. I. INTRODUCTION Medical microwave imaging (MMWI) has been proposed for a multitude of applications, including breast cancer detection, breast health monitoring, assessment of treatment response in breast cancer, stroke detection and monitoring, bone health monitoring, among others. MMWI uses non-ionising radiation at low power levels, which makes it particularly attractive for repeated use without causing harmful side effects. As a result, interest in its clinical potential has grown substantially. A number of MMWI prototypes have already been developed and tested in clinical trials [1],[2],[3],[4]. With increasing interest in these technologies, the next step is clinical integration — either as a complement The associate editor coordinating the review of this manuscript and approving it for publication was Sejung Yang. to or a replacement for existing imaging modalities. For medical adoption, it is essential that MMWI systems produce consistent and reliable results. Specifically, results must be repeatable under equivalent conditions and not overly sensitive to external or system-related noise, such as temperature fluctuations or hardware drift. Systems with high repeatability are better positioned to detect clinically relevant changes in tissue, whether related to the development of pathological conditions or to responses to treatment. Despite the growing number of experimental MMWI studies and clinical prototypes, there are currently no defined standards or no comprehensive synthesis of how repeatability is evaluated across different systems, protocols, and metrics. This lack of standardisation limits comparability across studies and hinders regulatory approval and clinical 202524 2025 The Authors. This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ VOLUME 13, 2025 H. V. Lopes et al.: Quantitative Evaluation of Repeatability in MMWI: A Systematic Review translation. Preliminary examination of the literature suggests substantial variability in how repeatability is conceptualised and measured in MMWI studies, preventing a clear understanding of how hardware design choices, or reconstruction algorithms influence measurement repeatability, with the term often used interchangeably with reproducibility. Following definitions of the quantitative imaging biomarkers literature [5], this review defines repeatability as the ability of a system to produce consistent results under equivalent conditions, without being overly influenced by external noise. In contrast, reproducibility refers to the ability of a system to yield comparable results when replicated in other settings — for example, in different institutions, by different operators. To the best of the authors’ knowledge, no previous study has systematically examined how the repeatability of MMWI prototypes is evaluated, including the protocols and metrics used, or whether the reported systems demonstrate repeatable results. A systematic review is therefore of utmost importance to assess existing evaluation practices, and identify the main sources of variability that affect repeatability. By synthesising current approaches to assess repeatability, such a review can also provide insight into standardising strategies to improve the robustness and reliability of MMWI systems, contributing to establish the clinical efficacy of MMWI technologies and their progression from research prototypes to clinically deployable technologies. This systematic review addresses these gaps by answering the following research questions: R1: How has repeatability been assessed in the existing MMWI literature? R2: Do the MMWI prototypes reported in the literature demonstrate repeatability? II. METHODS This systematic review followed the Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) guidelines [6]. A. STUDY SELECTION A systematic literature search was conducted to identify studies that evaluated the repeatability of microwave-based systems in biomedical or clinical contexts, including applications, but not limited to, the breast, brain, and stroke. The review focused on studies that employed quantitative metrics to assess repeatability across repeated measurements, and/or multiple sessions. B. DATA SOURCES A systematic search of published studies was conducted in the following databases: •Scopus •IEEE Xplore •Web of Science •PubMed These databases were selected for their comprehensive coverage of the relevant literature in the field of medical microwave imaging, including both peer-reviewed journals and conference proceedings. Conference papers were included given that it is common practice in this field to present peer-reviewed work at conferences. C. SEARCH STRATEGY The defined search strategy included different combinations of the following concepts: microwave, imaging, signal, radar, detection, diagnosis, scanning, screening, monitoring, reproducib*, repeatab*, consisten*, simila*, variability, quatif*, measur*, metri*, evaluat*, successive, asses*, longitudinal, over time, figure of merit, medical, breast, brain, stroke, bone, skin, and lung. The full description of the search strategy is available in the supplementary material. The search was limited to articles published in English, and no date restrictions were applied, with the search being conducted in April 29th, 2025. Duplicates were removed using Rayyan - a web based software for duplicate records removal and screening of papers in systematic reviews [7]. D. ELIGIBILITY CRITERIA The studies included in this systematic review met the following criteria: •Investigated microwave-based systems, such as imaging, radar, or signal processing techniques; •Focused on biomedical imaging or clinical diagnostic contexts, including but not limited to breast, brain, stroke, skin, bone, or lung applications; •Assessed repeatability of results using quantitative metrics, across repeated measurements, multiple sessions, or varying subjects. Only studies published in peer-reviewed journals or conference proceedings were considered for inclusion. E. SOURCE OF EVIDENCE SELECTION Selection and screening process: this step includes title and abstract screening, followed by full text screening. Both stages were performed independently by two reviewers (HL and DG). Title and abstract screening was conducted in Rayyan. The title and abstracts were assessed against the inclusion criteria with each reviewer assessing if each abstract was eligible for full text screening. Disagreements between reviewers were solved based on discussion to arrive to a consensus and when needed a third reviewer (RC) was consulted. Data management: The management of search results was conducted using OneDrive for storage and organisation, and Zotero for reference management. F. MAIN OUTCOMES The main outcomes extracted from the included studies were categorised into seven groups: VOLUME 13, 2025 202525 H. V. Lopes et al.: Quantitative Evaluation of Repeatability in MMWI: A Systematic Review 1) Evidence of repeatability assessment, including definitions of this concept and the specific types of evaluation performed; 2) Quantitative metrics employed to assess repeatability (e.g., Dice coefficient, correlation coefficients); 3) Context of the assessment, encompassing the imaging target, use of physical phantoms, involvement of healthy volunteers or patients, and a description of the microwave system utilised; 4) Study design characteristics, such as the number of repeated measurements, number of subjects, and the interval between measurements; 5) Reported sources of variability, including systemrelated, operator-related, subject-related, and environment-related factors; 6) Reported results of the assessments, including metric values; 7) Prevalence or dominance of specific metrics used to evaluate repeatability. G. DATA EXTRACTION Data extraction was conducted independently by HL and DG using a spreadsheet in Microsoft Excel. Prior to screening the full set of included articles, a pilot test was performed on four reference studies [8],[9],[10],[11] to evaluate the appropriateness and completeness of the data extraction categories. Discrepancies in the extracted data were resolved through discussion with a third reviewer (RC). The following variables were extracted: first author, senior author, year of publication, study objective, characteristics of the microwave imaging system, reconstruction and processing algorithms employed, study design (e.g., use of physical phantoms, healthy volunteers, or patients), sample size, clinical or biomedical application (e.g., breast, brain), definitions of repeatability, the specific metrics used to assess repeatability, type of repeatability assessment, key findings related to repeatability, and reported sources of variability. For the purposes of this review, the type of repeatability assessment was categorised as follows: •Test-retest: repeated measurements acquired on the same subject or phantom within a single session, under identical conditions and without intentional repositioning. •Inter-session: repeated measurements acquired across multiple sessions, generally involving subject or phantom repositioning and/or a temporal interval between acquisitions. H. RISK OF BIASES IN INDIVIDUAL STUDIES Two reviewers (DG and HL) independently performed full-text screening of the studies. The results from each reviewer were discussed until a consensus was reached regarding the combination of the results. Therefore, avoiding the results to be biased towards the judgement of a single reviewer. A third reviewer (RC) was available to be consulted in the event of unresolved disagreement. I. DATA SYNTHESIS After completing the data extraction process, the information recorded in the individual Excel spreadsheets was polished. A descriptive analysis was performed to synthesise information and illustrate trends in the use of quantitative metrics for repeatability assessments, reporting the main outcomes reached, including the ranges of values for the repeatability metrics obtained in each study. The results were presented in tables and figures to facilitate understanding and comparison across studies. Repeatability metrics reported in each study were extracted and subsequently categorised according to their primary function. This categorisation was developed post hoc using an empirical thematic analysis of the mathematical definitions of the metrics and usage contexts, rather than being predetermined. A detailed description of each metric and its classification is provided in the Results section. III. RESULTS A. LITERATURE SEARCH AND STUDY SELECTION The systematic review of the literature was conducted on April 29th 2025, in which 953 records were retrieved from the four databases screened (for more details see Supplemental material). One additional relevant study was identified through manual reference list screening. After removing duplicates, 608 abstracts were screened for eligibility based on the inclusion criteria by two reviewers (DG and HL). A total of 23 papers were selected for full-text screening. During full-text screening phase two studies were excluded due to not reporting quantitative methods to assess repeatability [12],[13]. Ultimately, 21 papers were included for data extraction. This process is described in the PRISMA flow diagram in Figure 1. B. PUBLICATION TREND OVER TIME Of the 21 studies included in this review, the first one was published in 2014. Over the past six years, the rate of publications to quantitatively assess repeatability has increased, averaging 2.3 studies per year compared to 1.2 per year between 2014 and 2019. The peak in publication activity occurred in 2023, with a total of four studies published. The annual distribution of the included studies is presented in Figure 2. As of April 29th 2025, three studies had been published, and it remains uncertain how many additional studies assessing the repeatability of results in medical microwave imaging may be published throughout the remainder of the year. C. DEFINITION OF CONCEPTS RELATED TO REPEATABILITY The 21 studies included in this review quantitatively assessed repeatability; however, the terminology employed to describe this concept varied considerably. Terms such as repeatability, reproducibility, consistency, and reliability were often used interchangeably, with limited clarification of their intended 202526 VOLUME 13, 2025 H. V. Lopes et al.: Quantitative Evaluation of Repeatability in MMWI: A Systematic Review FIGURE 1. PRISMA flow diagram for the systematic literature search and review process. FIGURE 2. Publication trend over time for studies included in the systematic review, i.e. that assess quantitatively the repeatability of their results. meaning. Only one study provided an explicit definition: Wang et al. [11] defined reproducibility as ‘‘the degree to which repeated measurements in healthy study participants provide similar results.’’ The remaining studies did not provide a definition of the concept they were assessing, but rather used the terms interchangeably. D. STUDY CHARACTERISTICS Table 1summarises the medical application, system characteristics, including frequency and number of antennas, and the sample size across the included studies. From the 21 studies found in the literature, the majority (n =20) focused on breast screening, while the remaining study (n =1) investigated internal body temperature. The majority of microwave systems used in the studies varied between two types: (i) ultrawideband (UWB) radar imaging systems, and (ii) transmission-based systems, each reported in nine studies. It is relevant to note that all studies using transmission-based systems originated from the same research group. Three other system types were used in the remaining three studies: (iii) narrowband radar microwave imaging, (iv) MIMO radar with magnetic nanoparticles, and (v) microwave radiometry. While microwave radiometry was used to assess internal body temperature, all other systems targeted breast screening. Results show that of the 21 included studies, 11 assessed repeatability in human participants, 5 in physical breast phantoms only, and 5 in both volunteers and phantoms. Among the 11 studies involving human subjects, 2 were conducted on patients with breast cancer, while the remaining 9 recruited healthy volunteers. For studies employing phantoms, the sample sizes ranged from 1 to 5, with a mean of 3 and a standard deviation of 2.5. Most studies involving healthy volunteers reported sample sizes between 1 and 13 participants (mean: 4.6, SD: 4.3). One study, however, enrolled 35 participants, representing a markedly larger sample than the rest. The two studies that included breast cancer patients used sample sizes of 6 and 15. VOLUME 13, 2025 202527 H. V. Lopes et al.: Quantitative Evaluation of Repeatability in MMWI: A Systematic Review TABLE 1. Summary of medical applications, system characteristics, and sample size across included studies. NA: Not available. The type of repeatability assessment varied across the included studies, as observed in Table 2. In six studies, repeatability was evaluated in both test-retest and intersession basis. In nine studies, repeatability was evaluated using only the inter-session approach, involving repeated acquisitions conducted across different sessions, during which the subject or phantom was repositioned between measurements. Two more studies test repeatability of their system using a inter-session approach, but with the goal of comparing how the results changed in a pathological under treatment scenario, with the goal of verifying that their system is capable of detecting changes related to treatment [19], [24]. One study reported repeatability using the inter-session approach, and also compared the repeatability of two different microwave systems [30]. In two studies, the test-retest approach was used to assess repeatability, but no inter-session repeatability was evaluated [17],[27]. In the remaining study repeatability was reported using the test-retest approach, and also compared two different microwave systems [18]. The characteristics associated with the type of repeatability assessment—such as the number of sessions, measurements per session, and follow-up duration—varied considerably across the included studies. The follow-up period and number of sessions is particularly relevant for the studies evaluating inter-session repeatability. The follow-up time spanned from same-day acquisitions to a period of two years, with nine performing their study for more than one month. The studies that evaluated inter-session repeatability in the same-day performed multiple sessions in the respective day while repositioning the subject or phantom between sessions [17], [18],[25],[26],[27]. However, a longer follow-up period did not necessarily correspond to a greater number of sessions. For example, some studies conducted daily acquisitions over defined periods—27 days in [16] and 28 days in [30]—while others performed up to four sessions distributed across two years [24]. The number of sessions varied from one to 122. For the studies that assessed test-retest repeatability, they commonly had only one session, where the same subjects or phantoms were measured multiple times, with the number of measurements per session ranging from 2 to 12. E. METRICS USED IN INCLUDED STUDIES The majority of the studies (13/21 studies) evaluated repeatability through more than one metric, with a mean of 2 metrics per study (SD: 1.0), across the 21 studies. Eighteen different metrics were identified in the included studies. Figure 3shows each identified metric as well as the distribution of their frequency across the included studies. The most used metric is in fact the group of descriptive statistical measures (in grey in Figure 3), such as mean and standard deviation. This type of measure is normally used to stablish an overall tendency and variability measurement of repeatability after the use of specific repeatability metrics. The following most used metric was the Structural Similarity Index (SSIM), used in five studies. The remaining metrics were used in three or fewer studies. Importantly, these metrics are associated to the type of input data. For example, in 9 studies the repeatability was evaluated by comparing signals. In these cases, the preferred metric was correlation-based metrics, such as 202528 VOLUME 13, 2025 H. V. Lopes et al.: Quantitative Evaluation of Repeatability in MMWI: A Systematic Review TABLE 2. Summary of repeatability assessment protocols across included studies. In cases where the inter-session assessment involved only one session, multiple measurements were conducted within the same day with repositioning to mimic inter-session conditions. NA: Not available. FIGURE 3. Distribution of the number of times each metric was used across the included studies. SSIM: Structural Similarity Index, MHD: Modified Hausdorff Distance, NRMSE: Normalised Root Mean Squared Error, SEM: Standard Error Measurement, ICC: Intraclass Correlation Coefficient, CV: Coefficient of variation. cross-correlation and Pearson correlation coefficient (PCC) (6/9 studies). Image-based metrics were used in 18 studies, with SSIM being the most used (5/18 studies). The metrics employed across the included studies were organised into five thematic groups according to the context that they were used to assess repeatability: spatial consistency, visual and structural similarity, signal pattern similarity, performance summary metrics, and statisticalbased. This categorisation was derived empirically following full-text analysis of each study. The groups are described below, together with the mathematical definitions of the corresponding figures of merit reported in the included studies, where Aand Brepresent two sets of points, with |A| and |B|being the number of points in each set, respectively. 1) SPATIAL CONSISTENCY Spatial consistency metrics are designed to examine the degree to which two datasets remain spatially aligned across repeated scans. Commonly used approaches include distancebased metrics, which compute point-wise dissimilarities between sets, and overlap-based metrics, which quantify the degree of shared area or volume. Among the distance-based approaches, the Euclidean distance, the Modified Hausdorff Distance (MHD) [31], and Fréchet distance [32] were the three figures of merit used. Lavoie et al. [8] computed the Euclidean distance between the locations of maximum intensity values in two images. This metric is denoted as: Euclidean distMaxResponse(A,B)= ∥amax −bmax∥,(1) VOLUME 13, 2025 202529 H. V. Lopes et al.: Quantitative Evaluation of Repeatability in MMWI: A Systematic Review with amax and bmax being the coordinates of the maximum response locations in the images. MHD computes the average spatial distance between two objects, being used to compare image outlines, and is defined as: MHD(A,B)=max 1 |A|X a∈A min b∈B∥a−b∥, 1 |B|X b∈B min a∈A∥b−a∥,(2) where ∥a−b∥is the ℓ2norm (i.e. Euclidean distance) between points aand b. This metric was used in two studies [8],[26]. Lavoie et al. [8] defined two uses of MHD: MHD-outline, which compares the surfaces of reconstructed volumes, and MHD-response, which computes the MHD between regions exceeding a defined intensity threshold. Butterworth et al. [26] used MHD to compute the spatial alignment between point clouds of the breast volume. In addition, Butterworth et al. applied the Fréchet distance between the surfaces of the point clouds, as an alternative to MHD. This metric determines the similarity between two sets while accounting for the order of points along curves or meshes and is typically recommended when comparing continuous shapes such as anatomical contours. The Fréchet distance is defined as: F(A,B)=inf α,β max t∈[0,1] d(A(α(t)),B(β(t))),(3) where αand βare reparameterisations of Aand Bbetween 0 and 1, tis the step between each triangle in the mesh, and dis the distance between the points at parameter t. MHD and Fréchet distance return a value of zero when the shapes are identical and increase with spatial divergence, without imposing an upper limit. In terms of overlap-based metrics, the Dice coefficient [33] was one of the measures to quantify overlap between two sets: Dice coefficient(A,B)=2|A∩B| |A|+|B|.(4) This metric ranges from 0 (no overlap) to 1 (perfect overlap) and was employed in a single study by Butterworth et al. [26]. Another reported overlap-based metric is the volumemismatch metric, which quantifies the proportion of non-overlapping volume between two sets: Volume mismatch(A,B)=|A∪B|−|A∩B| |A∪B|.(5) A value of 0 indicates complete overlap, while 1 corresponds to no overlap. Lavoie et al. [8] used this metric to quantify shape differences of the breast in full 3D reconstructions between repeated acquisitions. 2) VISUAL AND STRUCTURAL SIMILARITY In the context of repeatability assessment in MMWI, visual and structural similarity metrics provide insight into how consistent image features remain across repeated acquisitions. Rather than focusing on spatial displacement or exact geometric correspondence, these approaches evaluate similarity based on local image features or intensity patterns. SSIM [34] was designed to quantify the degree of resemblance between two sets of points or images by analysing the relationships of pixels that are spatially near each other, and was originally motivated to reflect perceived image quality as judged by the human visual system. It compares local statistics of luminance (i.e., brightness), contrast (i.e., dynamic range), and structure (i.e., patterns and textures in the image), which are then aggregated to produce a global similarity score. In the most general formula of SSIM, parameters that define the relative importance of luminance, contrast, and structure can be set. The following equation represents a widely spread simplified version of SSIM in which the importance of each parameter is equal: SSIM(A,B)=(2µAµB+C1)(2σAB +C2) (µ2 A+µ2 B+C1)(σ2 A+σ2 B+C2),(6) where µAand µBare the mean intensities of the two images, σAand σBare the standard deviations of the intensity, σAB is the covariance between the two images, and C1 and C2 are constants that stabilise the division. SSIM is 1 when comparing identical images, but its minimum value depends on specific image content and can even be negative in some cases. SSIM was used in five studies [8],[9],[14],[16],[30] to evaluate repeatability. Another metric used to assess visual similarity was the 2D Pearson correlation coefficient (2PCC), which quantifies the linear relationship between two images. It was used by Smith et al. [10] to evaluate the similarity of the generated images across different sessions by comparing each image to the average image. The 2PCC is defined as: 2PCC(A,B)=Pi,j(aij − ¯a)(bij −¯ b) qPi,j(aij − ¯a)2·qPi,j(bij −¯ b)2 ,(7) where aij and bij are the pixel values at position (i,j) in images Aand B, respectively, and ¯aand ¯ bare the mean pixel values of images Aand B. The 2PCC ranges from -1 to 1, where 1 indicates a perfect positive linear relationship, -1 indicates a perfect negative linear relationship, and 0 indicates no linear relationship. Difference-based metrics, such as Root-Mean-Square Error (RMSE), compute the magnitude of pixel-wise discrepancies between images. RMSE provides a measure of the average magnitude of the differences between two sets of points: RMSE(A,B)=v u u t 1 n n X i=1 (A−B)2,(8) where nis the number of points in sets Aand B. RMSE is always non-negative, with lower values indicating better agreement between the sets. RMSE is commonly normalised 202530 VOLUME 13, 2025 H. V. Lopes et al.: Quantitative Evaluation of Repeatability in MMWI: A Systematic Review to facilitate the comparison of results across different scales or units. Normalised RMSE (NRMSE) is defined in [8] as: NRMSE(A,B)=RMSE(A,B) 1min ,(9) where 1min =min(max(A)−min(A),max(B)−min(B)). The NRMSE ranges from zero to one, where zero indicates no difference and one indicates maximum difference, thus representing the average error per voxel. Lavoie et al. [8] used NRMSE and compared it with SSIM. In three studies [9],[14],[15], the peak pixel intensity was employed as a metric to evaluate repeatability. In [14] and [15], the authors assessed the consistency of the strongest response across sessions by comparing the maximum pixel intensities of two reconstructed images. Specifically, they computed a difference image by subtracting the pixel intensities of the first scan from those of the second. The peak value of this difference image was then analysed relative to the peaks of the original images. The peak difference was compared between scan pairs, with images obtained at different time points providing a proxy for the repeatability of the measurements over time. The study by Porter et al. [9] used a different approach to evaluate the peak pixel intensities across multiple repeated scans. Differential images were formed using the first scan as a calibration baseline, and each image was normalised to the peak intensity of the reference image. 3) SIGNAL PATTERN SIMILARITY Part of the included studies reported metrics that quantify signal pattern similarity. These approaches do not rely on spatial correspondence, but rather on the degree to which signal traces remain consistent across different sessions. One simple hypothesis is to subtract two signals acquired at different time points; a larger difference between them would mean a lower similarity. This approach was used in two studies [18],[27]. Two other commonly used measures in this context are the PCC in 1D and cross-correlation. In the studies reporting these metrics, the signal obtained during each subsequent acquisition for a given antenna was compared against a baseline signal—typically the first scan acquired—to assess repeatability. As it was defined for the 2PCC in 2D, the PCC in 1D is defined as: PCC(X,Y)=Pn i=1(xi− ¯x)(yi− ¯y) qPn i=1(xi− ¯x)2qPn i=1(yi− ¯y)2 ,(10) where nis the number of points in the signals Xand Y,¯xand ¯yare the means of the sets Xand Y, respectively. PCC was used in three studies to evaluate similarity between aligned signals [10],[19],[20]. Cross-correlation generalises the concept of PCC to measure the similarity as a function of a temporal or spatial shift between two signals. This metric is robust to phase shift or time alignment discrepancies that may occur between repeated acquisitions. Cross-correlation was used in three studies to assess the similarity of signal patterns over time [9], [14],[16]. 4) PERFORMANCE SUMMARY METRICS In contrast to studies that employed explicit metrics to evaluate repeatability, three studies instead reported performance metrics commonly used to characterise the quality of reconstructed images in microwave imaging systems. Ley et al. [17] determined the full width at half maximum (FWHM; measures the extent of the detection), and the crosstalk to noise ratio (CNR; ratio between the mean amplitude of the crosstalk signal and the background system noise) for the signals acquired to assess repeatability of their system. Kranold et al. [18], and Correia et al. [25] computed the signal to clutter ratio (SCR; contrast between the target response and background), signal to mean ratio (SMR; target strength relative to average image intensity), and localisation error (LE; deviation between actual and reconstructed target position) of the reconstructed images by their system. Such metrics summarise key aspects of system performance— namely, the capacity to detect and localise targets—which may indirectly reflect repeatability when the same targets are consistently present across repeated acquisitions. 5) STATISTICAL-BASED METRICS While the metrics discussed in previous sections focus primarily on pairwise comparisons between repeated images or signal acquisitions, statistical-based measures offer a complementary perspective by evaluating how such metrics vary across multiple time points or sessions. These approaches quantify the overall consistency of repeated measurements, often across different subjects or conditions. Simple statistical descriptors such as the mean, median, SD, and range were used to summarise the central tendency and variability of measurements [10],[18],[19],[20],[21],[23], [24],[28],[30]. For instance, after evaluating the repeatability between their aquisitions pairwise Porter et al. [30] computed the mean, SD and range of the values to quantify the overall repeatability of their results across all acquisitions. A series of studies from the same research group [10],[19],[20],[23], [24],[28], all employing variations of the same microwave imaging prototype, reported descriptive statistics of the average permittivity values within the reconstructed images. In these systems, each image pixel represents the relative permittivity of the tissue at the corresponding location. By quantifying the variability in average permittivity values across repeated scans using basic summary statistics, these studies aimed to evaluate the stability and repeatability of the imaging output over time. Additionally, Smith et al. [24] conducted a post hoc twoway ANOVA on the average permittivity values of the scans to investigate whether changes observed between baseline and six-week follow-up scans were statistically significant across all volunteers. VOLUME 13, 2025 202531 H. V. Lopes et al.: Quantitative Evaluation of Repeatability in MMWI: A Systematic Review A separate group of statistical metrics was used in one study to assess the global repeatability of microwave imaging outcomes. Wang et al. [11] employed the intraclass correlation coefficient (ICC), the standard error of measurement (SEM), and the coefficient of variation (CV), calculated based on the average relative permittivity of each reconstructed image. ICC was used to assess the proportion of measurement variance related to inter-subject differences. The ICC quantifies the consistency or agreement between repeated measurements by decomposing the total variability into within-subject and between-subject components. In a one-way random-effects model [35], the ICC is given by: ICC =σ2 between σ2 between +σ2 within ,(11) where σ2 between and σ2 within denote the between-subject and within-subject variance, respectively. ICC values range from 0 (no agreement) to 1 (perfect agreement). ICC can accommodate multiple sources of variability across repeated measurements and is therefore well-suited to study designs involving repeated imaging sessions of the same subjects. The SEM quantifies the absolute measurement error and is derived from the within-subject variability: SEM =qσ2 within.(12) Finally, CV was reported to express measurement variability relative to the mean: %CV =σTotal ¯µ×100%,(13) where, ¯µis the mean value. F. REPORTED REPEATABILITY METRICS Tables 3,4,5, and 6summarise the repeatability metrics and their respective values reported across the 21 studies included in this review. These tables focus on signal pattern similarity metrics, spatial consistency metrics, visual and structural similarity metrics, and performance summary or statisticalbased metrics, respectively. Despite variability in the reported metrics, the majority of studies consistently demonstrated high repeatability, with exception of Butterworth et al. [26], with relatively lower repeatability values. Among signal metrics used in multiple studies, cross-correlation and PCC were the most frequent. Cross-correlation ranged from 0.592 to 0.999 across the three studies employing it, while PCC varied between 0.86 and 1.0 in the three studies that reported it. Regarding image-based metrics, the SSIM ranged from 0.41 to 1.0 across four of the five studies that reported it. However, in the study by Lavoie et al. [8], SSIM was varied in the range 0.18 to 0.38. The NRMSE was reported in two studies [8],[29], with values between 0.02 and 0.24. Notably, the Fréchet distance—reported only by Butterworth et al. [26]—showed relatively large values (i.e. lower repeatability), even in phantom-based assessments. Three studies employed both signal and image-based metrics [9],[10],[14]. Two of these studies [9],[10] obtained better metric values when considering signal-based metrics, as aligned with the general trend found, and with a broader variability of values for the imaged-based metrics. In contrast, Porter et al. [14] reported worse values for signal cross-correlation than for SSIM values, with a narrower range of values obtained for the image-based metrics. Among the six studies that assessed both test-retest and inter-session repeatability, five showed higher repeatability for the test-retest condition across all reported metrics [8], [14],[22],[25],[29]. Wang et al. [11] was the only exception, with comparable metric values between the two assessment types. In studies evaluating repeatability with phantoms and human participants under the same repeatability condition, phantoms consistently yielded higher repeatability values [10],[14],[16],[26]. Lavoie et al. [8] was excluded from this comparison, as phantoms and volunteers were used in different assessment types, with phantoms used to assess test-retest repeatability and human subjects inter-session repeatability. G. REPORTED SOURCES OF VARIABILITY In 14 of the included studies, the authors reported factors that may have affected the repeatability of their results. Figure 4illustrates these sources of variability, with each block representing a distinct source and annotated with the studies that identified it. At the bottom of the diagram, the studies that did not report any variability factors are also listed. The most frequently reported sources of variability were subject-related. The most cited factor, reported in seven studies, was volunteer or phantom repositioning between sessions [8],[9],[10],[11],[14],[17],[21]. The second most frequent factor, reported in six studies, was the difficulty in imaging breasts of smaller size [9],[11],[15],[23],[26],[28]. Hardware-related changes, such as variations in electronic noise or scanner performance over time, were identified in three studies [8],[11],[25]. Tissue-related variability, including physiological changes due to the menstrual cycle [9], [16] or pathological differences between healthy and diseased tissue [20], was also reported. In two studies, internal repositioning of breast tissue, such as glandular tissue, was specifically reported as a factor affecting repeatability, even when external positioning remained consistent [10],[26]. Additional individual factors reported by single studies include scan misalignment [8], changes in coupling medium [14], signal interference [21], and operator interference [25]. IV. DISCUSSION This systematic review comprehensively examined how repeatability has been quantitatively evaluated in MMWI systems, including the protocols used, and the repeatability 202532 VOLUME 13, 2025