Roccetti M, Preprint 1 Research Article Methodological Considerations on the External Validity of the Kim HJ et al. COVID19 Vaccination Study (Biomark Res 13:114, 2025): A Quantitative Analysis Marco Roccetti 1, * 1 Department of Computer Science and Engineering, University of Bologna, Italy;
[email protected] * Correspondence:
[email protected] Abstract Background/Objectives: The Kim HJ et al. (2025) cohort study reported a surprising finding: a significantly higher 1-year cancer incidence risk in the COVID-19 vaccinated group. Given the vast public health implications, this result requires immediate and rigorous examination. This analysis provides the first formal external critique, integrating computational reviews to evaluate the methodology and test the external validity of the cohort. Methods: Aggregated data from the Kim HJ et al. matched cohort (n=2,975,035) were used to calculate the overall Crude Incidence Rate (CR) of cancer. This study tested the resulting Crude Incidence Rate (CR) of the cohort against the established o]icial national average CR for South Korea for the reference period (2020–2022). A secondary analysis of the final cohort's size ratio was performed to hypothesize a potential flaw in the 1:4 Propensity Score Matching (PSM) procedure. Results: The cohort's overall CR was found to be 40.78 per 10,000, representing a substantial 22.26% downward deviation from the national average (52.46 per 10,000; SD 2.97). This discrepancy establishes a pronounced epidemiological paradox, strongly suggesting a lack of external validity for the cohort. Based on the exact 4:1 ratio of the final matched groups, the analysis proposes that the PSM procedure was likely inverted or misapplied, with the smaller unvaccinated group being used as the base cohort '1' for matching against the vaccinated group '4', which may have contributed to the suppression of the overall CR. Conclusions: The reliability of the statistical associations reported by Kim HJ et al. is challenged by a possible lack of external validity and the hypothesized methodological ambiguities concerning the PSM. We conclude that independent validation is
Roccetti M, Preprint 2 mandatory and reiterate the call for public access to the underlying Korean National Health Insurance database to resolve these contradictions. Keywords: COVID-19 Vaccination, Crude Incidence Rate, Cancer Incidence, Epidemiological Paradox, External validity , Propensity Score Matching, Inversion or Misapplication of the 1:4 Matching, Cohort Representativeness 1. Introduction The accurate assessment of post-marketing adverse events, particularly those associated with widespread public health interventions like COVID-19 vaccination, is critical for public trust and e]ective health policy [1]. Observational studies drawing from national health insurance databases are essential tools in this process, o]ering large sample sizes and real-world data. However, the reliability of such studies is fundamentally dependent on the methodological rigor applied to cohort selection and statistical adjustment [2]. A core standard of rigor in epidemiology is External Validity. This concept refers to the extent to which the findings of a study can be generalized to other populations, settings, and circumstances outside the study's specific cohort. For a cohort derived from a national registry, high external validity requires that the study’s overall burden of disease (measured by the Crude Incidence Rate, or CR) is statistically consistent with the known national burden of disease for the same period. Failure to meet this standard, often due to sampling or selection issues, means the cohort is not representative of the broader population, rendering its conclusions questionable in a real-world context. The study by Kim HJ et al., published in Biomarkers Research in 2025 [3], is a retrospective, population-based cohort analysis utilizing data from some Korean National Health Insurance database to investigate the 1-year risks of cancers associated with COVID-19 vaccination in South Korea [4]. The study’s finding, suggesting a higher rate of new cancer cases among the vaccinated population compared to the unvaccinated, is a surprising and scientifically challenging result that has yet to be fully addressed and scientifically analyzed with the required depth and urgency, especially considering the global scope of the vaccination programs [5]. The present work serves as a comprehensive critique that integrates two sequential analyses. Initially, a pronounced epidemiological paradox was identified based on raw incidence calculations derived from the study's supplementary data. This paradox established a significant external inconsistency between the study cohort’s aggregate cancer incidence and o]icial national statistics [6]. The follow-up analysis we developed [7] is as an integral part of the overall argument and posits a specific
Roccetti M, Preprint 3 methodological explanation for this paradox: the likely misapplication or inversion of the 1:4 Propensity Score Matching (PSM) procedure used by Kim HJ et al in [3]. In closing, the primary objective of this paper is to quantify and demonstrate the severity of the external inconsistency observed in the Kim HJ et al. cohort, thereby challenging its external validity. We then propose a plausible methodological hypothesis, specifically the inverted Propensity Score Matching (PSM), that would reconcile this numerical discrepancy and ultimately calls into question the reliability of the study’s final association results. It is essential to undertake this critical examination using the known data, as the magnitude of the finding demands the highest level of scientific scrutiny. In the remainder of this paper, Section 2 details the materials and methods for calculating the Crude Incidence Rates and formulating the PSM inversion hypothesis. Section 3 presents the quantitative results, including the derivation of the 22% deviation. Finally, Section 4 discusses the resulting lack of external validity, explores the PSM inversion as a probable cause of the bias, and provides conclusions and recommendations for data transparency. 2. Materials and Methods In this Section, we provide all the necessary details on the data and methods used in this study, allowing readers to easily replicate our findings. 2.1. Sources of Data This analysis is based entirely on publicly available, aggregated data extracted from the Kim HJ et al. study [3] and o]icial South Korean national health statistics as reported in [8-10]. In particular, the raw cohort figures necessary for calculation were obtained from Table S4 ("Cumulative incidences of overall cancers in the matched cohort between vaccinated and unvaccinated individuals") in the Supplementary Material of the Kim HJ et al. manuscript [3]. These figures are as described in the following Table 1.
Roccetti M, Preprint 4 Table 1. Aggregated raw data from the Kim HJ et al. study, extracted from Table S4 of the supplementary materials [3], showing the final matched cohort counts and the reported Propensity Score Matching (PSM) ratio. The O]icial National Cancer Statistical data, including the O]icial Crude Incidence Rate (CR), for all cancers in South Korea were instead sourced from the Korean Central Cancer Registry for the years immediately preceding and during the study period (2020– 2022), as reported in [8-10]. This data provides the robust national baseline against which the study cohort's representativeness is tested. 2.2. Definition and Calculation of Crude Incidence Rate (CR) The Crude Incidence Rate (CR) per 10,000 population is a fundamental epidemiological measure used here specifically to evaluate the external validity of the cohort. Unlike Age-Standardized Rates (ASRs) which adjust for age distribution to allow comparison between populations, the CR reflects the raw burden of disease in a defined population over time [11]. Critically, any cohort derived from a national database should possess an aggregate CR that is statistically consistent with the national average CR for the same time period. A significant deviation signals a foundational problem in the initial sampling or selection process that any given procedure used to construct the cohort (like PSM for example) would fail to correct. The CR is calculated using the established epidemiological formula: CR per 10,000 = (Number of new cases during a given period) / (Average population at risk during the same period) x 10,000 (1) Metric Value Initial Cohort Size 8,407,849 individuals Final Matched Cohort Size 8,407,849 individuals Total Cancer Cases in Matched Cohort 12,133 cancer cases Unvaccinated Group (N) 595,007 individuals Unvaccinated Group (Cases) 1,989 cancer cases Vaccinated Group (N) 2,380,028 individuals Vaccinated Group (Cases) 10,144 cancer cases Propensity Score Matching (PSM) 1:4 Ratio
Roccetti M, Preprint 5 This is followed by a straightforward calculation of o]icial South Korean CR baseline [810], whose values for both sexes per 100,000 population were converted to a per 10,000 basis to establish the national benchmark as described in Table 2 below. Table 2. O]icial National Crude Incidence Rates (CR) for All Cancers in South Korea per 100,000 and the derived CR per 10,000, used to establish the national average baseline for the reference period (2020–2022). Consequently, the o]icial average CR baseline for all cancers for the reference period (2020–2022) can be established as the mean of these values: CR (O]icial Average) = 52.46 per 10,000 (Standard Deviation SD = 2.97). Finally, using the raw figures from Table S4 in the Supplementary material provided by Kim HJ et al. in [3], the following CRs of Table 3 are calculated using Eq. (1) for the cohort of interest. Table 3. Calculated Crude Incidence Rates (CR) for the Kim HJ et al. matched cohort, showing the overall rate for the entire cohort and the rates for the segregated vaccinated and unvaccinated groups. Year CR per 100,000 CR per 10,000 2020 482.9 48.29 2021 540.6 54.06 2022 550.2 55.02 Group Calculation (New Cancer Cases / Population) x 10,000 Crude Incidence Rate (CR) CR (Cohort Overall) (12,133 / 2,975,035) x 10,000 40.78 per 10,000 CR (Vaccinated) (10,144 / 2,380,028) x 10,000 42.63 per 10,000 CR (Unvaccinated) (1,989 / 595,007) x 10,000 33.43 per 10,000
Roccetti M, Preprint 6 2.3. Hypothesis Formulation on Propensity Score Matching (PSM) Inversion A Propensity Score Matching (1:4 PSM) procedure aims to match each individual in the Treatment Group with four comparable individuals from the Control Group. In general, The Propensity Score Matching (PSM) is a quasi-experimental statistical method used to reduce the confounding bias that occurs when estimating the e]ect of a treatment or intervention (like vaccination in our case) in observational studies. The Propensity Score is defined as the conditional probability of an individual receiving the treatment given a set of observed covariates (e.g., age, sex, comorbidities). Hence, the propensity score e(X) is given by e(X)) = Prob (Z = 1 | X), where Z is the treatment assignment and X is the vector of baseline covariates. The PSM calculation procedure involves a multi-step process. First, a logistic regression model is constructed to estimate the propensity score for every individual, based on the set of observed confounders. Once the propensity scores are calculated, the matching phase begins. Di]erent matching algorithms exist (e.g., nearest neighbor, caliper, or kernel matching). In the reported 1:4 PSM, each treated individual (or the base group) is paired with four comparable control individuals whose propensity scores are nearly identical. This process e]ectively creates a synthetic, balanced cohort where the two groups are comparable on all measured confounders, thereby minimizing selection bias. The primary rationale for using PSM is to mimic the randomization process of a randomized controlled trial in non-randomized observational data. By balancing the distribution of baseline covariates between the treated and control groups, PSM aims to isolate the true e]ect of the treatment (e.g., COVID-19 vaccination) from the e]ects of confounding factors that influenced the decision to vaccinate. If the PSM is successfully implemented, any residual di]erence in outcome between the matched groups can be more confidently attributed to the treatment itself. A failure in the PSM process, or a misapplication like the hypothesized inversion, fundamentally undermines this rationale and reintroduces significant bias into the analysis. All this said, given the study’s focus on the COVID-19 vaccine of [3], the standard and appropriate group assignment should have been: Treatment = Vaccinated and Control = Unvaccinated. The hypothesis of Inverted PSM can be hypothesized based on final reported cohort sizes: 595,007 Unvaccinated and 2,380,028 Vaccinated. This distribution is suggesting a reverse assignment: Base Group = Unvaccinated; Matched Group = Vaccinated.
Roccetti M, Preprint 7 3. Results This Results Section presents two types of results: first, the quantification of the Epidemiological Paradox through the comparison of the calculated Crude Incidence Rate (CR) against the national baseline; and second, the numerical evidence supporting the hypothesis of Propensity Score Matching (PSM) inversion. 3.1. Quantification of the Epidemiological Paradox The comparison between the study cohort's aggregated CR and the national average CR revealed a substantial and significant downward deviation, confirming the epidemiological paradox as summarized in the following Table 4. Table 4. Quantification of the Epidemiological Paradox: Comparison of the Kim HJ et al. Cohort's overall Crude Incidence Rate (CR) against the O]icial National Average CR, highlighting the severe downward deviation We have now the paradox summarized as follows: the study’s analysis suggests an elevated cancer risk within the majority group (vaccinated CR is 27.5% higher than unvaccinated CR), which should intuitively push the overall cohort CR higher, yet the overall CR is 22.26% lower than the national baseline. This profound inconsistency represents a strong presumption of the cohort's lack of representativeness, which awaits formal refutation, though such a refutation appears mathematically challenging. To comprehend the full impact of this deviation, one must translate these statistical discrepancies into absolute numbers, which reveal the magnitude of the e]ect. Based on the cohort's overall rate of 40.78 per 10,000 and applying this to South Korea’s population (approx. 51.77 million inhabitants [12]), the cohort rate would translate to approximately 211,273 new annual cancer cases. This is over 60,000 fewer new cases Metric Rate per 10,000 Analysis O]icial National Average CR (2020–2022) 52.46 Baseline for External Validity Kim HJ et al. Cohort Overall CR 40.78 Calculated from Study Data Downward Discrepancy 11.68 (52.46 - 40.78) Percentage Deviation ≈ 22.26% (52.46 - 40.78) / 52.46 x 10%
Roccetti M, Preprint 8 than the 271,957 derived from the o]icial national average rate of 52.46 per 10,000 for the same population. This massive deficit in expected cases underscores the profound lack of representativeness. 3.2. Evidence Supporting the PSM Inversion Hypothesis The hypothesis that the 1:4 PSM was inverted is strongly supported by the final cohort sizes reported in Kim HJ et al. In fact, given the study’s focus on the COVID-19 vaccine, the standard and appropriate group assignment should have been: Treatment = Vaccinated; Control = Unvaccinated. The hypothesis of Inverted PSM is here formulated by analyzing the final reported cohort sizes: 595,007 Unvaccinated and 2,380,028 Vaccinated. This distribution is mathematically consistent with taking the smaller group (approx. 600,000) as the base "1" and matching it to the larger group (approx. 2.4 million) as the "4" component, suggesting the reverse assignment: Base Group = Unvaccinated; Matched Group = Vaccinated. All this is further evidenced by the "numerical signature" where the total matched cohort (2,975,035) is exactly five times the size of the smaller unvaccinated group (595,007), that is the sum of the 1:4 ratio. This result confirms that the smaller unvaccinated group was used as the base '1' for the matching, thus inverting the standard procedure and yielding a final 4:1 ratio which is (erroneously) as follows: Ratio (Vaccinated / Unvaccinated) = 2,380,028 / 595,007 ≈ 4.00004. This calculated ratio confirms the numerical correspondence: the vaccinated group (2,380,028) is precisely four times the size of the unvaccinated group (595,007). This exact numerical construction solidifies the argument for an inversion, where the small, unvaccinated pool defined the base cohort size for the 1:4 matching. 4. Discussion The subsequent Discussion synthesizes the quantitative findings, focusing on several critical areas: 1) summarizing the previous background and the most relevant findings of this paper; 2) establishing the core Lack of External Validity stemming from the 22% downward CR deviation; 3) detailing the PSM Inversion Hypothesis as the primary methodological explanation for this bias; 4) contrasting the cohort against the South Korean Epidemiological Standards and 5) the limitations sorrounding this present study. Finally, Section 5 will conclude this treatment with recommendations for data transparency.
Roccetti M, Preprint 9 4.1. Summary of Prior Background and Key Findings This research has integrated integrates two distinct lines of computational epidemiological analysis to critically evaluate the methodology and findings of the Kim HJ et al. (2025) cohort study concerning the 1-year risks of cancers associated with COVID-19 vaccination in South Korea. The Kim HJ et al. finding, suggesting a higher cancer incidence in the vaccinated group, is a surprising and scientifically challenging result that warrants rigorous, immediate scrutiny using all available data, as undertaken here. Our initial critique established a pronounced epidemiological paradox: while the Kim HJ et al. cohort suggested a higher crude cancer incidence rate (CR) among the vaccinated group compared to the unvaccinated group, the overall cohort CR was found to deviate downwards by over 22% from the o]icial national average CR for South Korea recorded in the immediately preceding years (2020–2022). This fundamental discrepancy has suggested a lack of external validity for the cohort and raised strong concerns regarding its lack of representativeness. Building upon this, the ultimate analysis has proposed a plausible explanation for the paradox: the reported 1:4 Propensity Score Matching (PSM) procedure was potentially inverted or misapplied. This inversion could introduce unidentified confounding factors and artificially bias the overall CR downwards. While we do not intend to reject the association results reported by Kim HJ et al. until direct access to the database is granted, the observed discrepancies in the CR and the apparent inversion of the PSM cannot be overlooked. We have concluded that these methodological ambiguities reiterate the call for public access to the underlying Korean National Health Insurance database for independent validation and resolution of any contradiction. 4.2. The Epidemiological Paradox and Lack of External Validity The core issue addressed by this paper has been the epidemiological paradox which demonstrates a profound lack of external validity for the Kim HJ et al. cohort. This is evidenced by the suppressed overall CR (40.78 per 10,000) compared to the national average (52.46 per 10,000), which indicates that the final sample is strongly suspected to be not representative of the underlying population's cancer incidence. This fundamental inconsistency also suggests that the sampling bias present in the initial cohort selection was not adequately corrected by the Propensity Score Matching, or that the matching procedure itself introduced a new, significant bias. The lack of representativeness of the Kim HJ et al. cohort, evidenced by the suppressed overall CR, calls into question the reliability of the study's conclusions. The discrepancy is so large (11.68 per 10,000) that it cannot be dismissed as a minor statistical (or procedural) fluctuation.