A Biostatistical Reappraisal Unveiling the Mechanism Behind Apparent Cancer Risk Signals in a COVID-19 Vaccinated Cohort Author: Marco Roccetti A/iliation: University of Bologna, Department of Computer Science and Engineering, Bologna, Italy Email:
[email protected] Abstract (Structured) Background: A large-scale South Korean cohort study reported an apparent signal suggesting a higher cancer risk among individuals who received the COVID-19 vaccine [1]. However, a preliminary analysis had already indicated a substantial discrepancy: the cohort's overall cancer incidence was 22.6% lower than the national average (mean 2020-2022). This profound deficit strongly suggests that the finding is the result of a methodological artifact rather than a true biological risk. Using the latest o/icial 2022 national rates for rigorous coherence verification, this deficit is confirmed at 26.3%. Objectives: Our objective was to quantitatively demonstrate that an asymmetry in the underrepresentation of high-risk elderly individuals (>=65 years) is the mechanism that mathematically generates the apparent vaccine risk signal. Methods: To investigate the source of this anomaly, age-stratified cohort data (Table S4 in [1]) were analyzed alongside o/icial South Korean cancer statistics for 2022. The methodology specifically focused on comparing the age-stratified demographic composition and resulting Crude Incidence Rates (CRs) of the cohort against the rigorous South Korean national standards. Results: The total study cohort significantly underrepresented participants aged >=65 by 31.1% compared to the national demographic structure. Critically, the non-vaccinated subgroup aged >=65 showed an even more severe cancer undercount of 45.5% relative to the coherent national rate. This profound, asymmetric underrepresentation within the high-risk non-vaccinated cohort segment fully explains the 26.3% suppression of the overall cohort incidence and, consequently, mathematically produces the spurious excess risk observed among vaccinated individuals. Conclusion: The reported increase in cancer risk associated with COVID-19 vaccination illustrated in [1] is likely to be a statistical artifact stemming from severe, asymmetric selection bias in cohort enrollment. We conclude that methodologically sound and demographically balanced cohorts would almost certainly demonstrate no statistically significant di/erence in cancer incidence between vaccinated and non-vaccinated groups.
Abstract (Unstructured) A large-scale South Korean cohort study reported an apparent signal suggesting a higher cancer risk among COVID-19–vaccinated individuals [1]. However, a preliminary analysis had already indicated that the cohort’s overall cancer incidence was historically 22.6% lower than the national average, a deficit confirmed at 26.3% using the most recently available {2022} national rates. Our primary objective was to quantitatively verify wether an asymmetry in the underrepresentation of high-risk elderly individuals (>=65) was the mechanism that mathematically generateed this apparent risk signal. To investigate this, we performed a rigorous biostatistical comparison, analyzing age-stratified demographic composition and corresponding cancer case counts (Table S4 in [1]) against o/icial Korean national standards (2022) to reconstruct and evaluate Crude Incidence Rates (CRs) The analysis confirmed that the total study cohort markedly underrepresented participants aged >= 65 by 31.1% overall. Importantly, the non-vaccinated subgroup aged >= 65 showed an even more severe cancer undercount of 45.5% relative to the coherent national rate. This profound, asymmetric underrepresentation within the high-risk nonvaccinated cohort segment explains the 26.3% suppression of the overall cohort incidence and, consequently, mathematically produces the spurious excess risk observed among vaccinated individuals. We conclude that the reported increase in cancer risk is most likely a statistical artifact stemming from severe, asymmetric selection bias in cohort enrollment, and that methodologically sound and demographically balanced cohorts would likely demonstrate no statistically significant di/erence in cancer incidence between vaccinated and non-vaccinated groups. 1. Introduction The global rollout of COVID-19 vaccines has been accompanied by intensive observational research aimed at thoroughly assessing post-vaccination health outcomes, including the crucial endpoint of cancer incidence. While a recent large-scale study from South Korea reported an apparently increased cancer risk among their vaccinated cohort [1], a closer examination of their baseline incidence data immediately raised a fundamental concern regarding the external validity of their findings. A prior computational analysis highlighted a striking discrepancy: the reported overall crude cancer incidence rate (40.54 per 10,000, averaged over the period 2020-2022) was dramatically lower, by 22.6%, than the national average rate for the corresponding period [2]. Informed by the strictest criteria of epidemiology and biostatistics [3, 4], we have developed a detailed biostatistical reappraisal employing the latest o/icial South Korean cancer registry data from 2022 [5-7] (where the expected national rate is 55.02 per
10,000). Applying this more current standard, the suppression of the cohort's overall cancer incidence is even more pronounced, confirming a 26.3% deficit. This study aims to move beyond merely noting this statistical paradox. Our objective is to provide a comprehensive quantitative explanation by reconstructing the cohort's demographics using publicly available figures. We will demonstrate explicitly how severe, asymmetric selection bias, particularly within the high-risk elderly group, provides a necessary and su/icient condition to mathematically generate the observed, yet spurious, signal of increased cancer risk in the vaccinated population. 2. Materials and Methods 2.1 Data Sources Raw cohort data, including age-stratified participant and case counts, were carefully extracted from the supplementary materials (Table S4) of [1]. To establish a robust external validity gold standard, national statistics, including o/icial South Korean cancer incidence rates and demographic reports, were sourced from public registries and o/icial data [5-7]. 2.1.1 South Korea: Population Demography and O:icial Cancer Incidence To create a mathematically coherent national benchmark, we rely on a standard national South Korean population distribution approximation of 18.0% individuals aged >=65 and 82.0% individuals aged < 65, as reported in [7] and shown in Table 1 below. Age Group Assumed % in National Population >=65 18.0% < 65 82.0% Table 1: National Demographic Information (South Korea, 2022) Table 2, instead, displays the o/icial national crude cancer incidence rates (CR) per 10,000 individuals for the 2022 Korean population, stratified by age group (Total, >=65, and < 65), which serve as the external validity benchmark for the following analyses [5,6]. Age Group Crude Incidence Rate (CR) per 10,000 Total Population (Target) 55.02 Population >=65 (O/icial) 155.2 Population < 65 (Calculated Coherent) 33.03
Table 2: O/icial Crude Cancer Incidence Rates (South Korea, 2022) As a first comment on Table 2, it should be noticed that while the Total Population and the Population >=65 figures were directly sourced from [5, 6], the Population < 65 measured was calculated coherently with the other figures. Specifically, the Crude Incidence Rate (CR) for the Population < 65 (33.03 per 10,000) was derived using the aggregate weighted incidence principle, where the total national incidence is the weighted sum of the age-specific incidence rates. We solved the equation using the established demographic weights (18.0% for >= 65 and 82.0% for <65) and the o/icial national CR for the >= 65 group (155.2), thereby ensuring the CR for the under 65 figure is entirely consistent with the o/icially published national total. More importantly, we specify now that the analysis that will follow will center only on the aggregate demographic composition of the total cohort (vaccinated and non-vaccinated combined) contrasted against the national population above. We deliberately avoided further stratifying the population by granular vaccination details (e.g., number of doses), as our core thesis, that the profound demographic bias in the high-risk age group is the defining methodological flaw, is fully demonstrable without introducing this unnecessary complexity. At this point, it is time to show the cohort composition and the relative cancer counts as sourced from [1] and summarized in the following Table 3. Group Participants <65 Cases <65 Participants >=65 Cases >=65 Total Participants Total Cases Nonvaccinated 522,722 1,373 72,785 616 595,507 1,989 Vaccinated 2,098,888 6,861 298,140 3,283 2,380,028 10,144 Total 2,621,610 8,234 370,925 3,899 2,992,535 12,133 Table 3: Raw Demographic and Cancer Case Counts for the Cohort of [1] 2.1.2 Methods for Calculation of Cancer Incidence For providing the results that will follow in the next Section, we have extensively utilized the following three formulas which are used to compute the Crude cancer incidence (CR), the relative Demographic Deficit/Surplus and the percentage Di/erence between cancer incidence rates. The three key mathematical formulas employed are as follows [8]: CR (per 10,000) = (Number of Cases / Population at Risk) x 10,000 (1)
Relative Deficit/Surplus % = (Cohort Percentage - Population Percentage) / Population Percentage x 100 (2) Di:erence (% incidence) = (Cohort Incidence Rate - Reference Incidence Rate) / Reference Incidence Rate x 100 (3) 3. Results We first present results quantifying the demographic discordance, that is the di/erence between the actual national South Korean population structure and the composition represented within the cohort of [1], and the resulting over/under-representation of the two age groups against the established 2022 national standards (Table 4). Age Group % in Total Cohort % in Population (Assumed 18.0%/82.0%) Absolute DiXerence (Cohort–Population) Relative Deficit/Surplus >=65 12.4% 18.0% -5.6% -31.1% (Deficit) < 65 87.6% 82.0% +5.6% +6.8% (Surplus) Table 4: Age Representation Comparison for the Cohort of [1] vs. National Population These results clearly indicate a significant 5.6% absolute deficit in the total cohort's representation of individuals aged >=65 compared to the coherent national demographic (18.0%). Moreover, the Relative Deficit/Surplus figure powerfully quantifies the severity of the bias. For the high-risk >=65 group, the relative deficit is -31.1%. Since the CR for the >=65 group (155.2 in Table 2) is approximately 4.7 times higher than the CR for the < 65 group (33.03), this severe relative deficit of -31.1% in the age bracket typically most susceptible to cancer is likely to be the overwhelming primary driver of the overall incidence suppression observed in the cohort. Conversely, the less-susceptible < 65 age group shows a relative surplus of +6.8%. We now provide results relative to the Crude Incidence Rate (CR) measured within the cohort of [1], stratified by age groups (Under 65 in Table 5 and Over 65 in Table 6), where each group's cohort incidence is contrasted with the respective South Korean national age-specific incidence rate established earlier in Table 2.
Group Cohort CR (/10,000) Coherent National CR (/10,000) DiXerence (% incidence) Nonvaccinated 26.3 33.03 -20.4% Vaccinated 32.7 33.03 -1.0% Total 31.4 33.03 -4.9% Table 5: Crude Cancer Incidence Comparison for Participants Under 65 (2022 Standard) Group Cohort CR (/10,000) OXicial National CR (/10,000) DiXerence (% incidence) Nonvaccinated 84.6 155.2 -45.5% Vaccinated 110.1 155.2 -29.1% Total 105.1 155.2 -32.3% Table 6: Crude Cancer Incidence Comparison for Participants Aged 65 and Over (2022 Standard) Importantly, these results show that the critical and dramatically asymmetric di/erence lies particularly in the high-risk group (population aged 65 and over), compared against the o/icial 155.2 rate. In particular, the most pronounced and crucial finding is the asymmetric underrepresentation of cancer cases among non-vaccinated participants aged >= 65: specifically, the -45.5% deficit in the non-vaccinated cohort (Table 5). This severe, unequal undercount in the high-risk, non-vaccinated reference group artificially suppresses the true baseline risk. This fundamental lack of cancer cases in the elderly non-vaccinated population (which would be expected to be substantial given the national incidence rate) most likely explains the spurious result of [1] showing an apparently higher risk for the vaccinated group. This quantitative deficiency is definitely the precise mechanism that has produced the appearance of increased risk in the vaccinated group in [1], providing the comprehensive explanation for the 26.3% overall incidence suppression. 4. Discussion The primary strength of this reappraisal lies in its strictly quantitative and biostatistical focus. We conclusively demonstrate the existence and the exact magnitude of the
demographic bias solely through the careful analysis of publicly available data against established national gold standards. This approach avoids the inherent complexities and controversies of clinical arguments, thereby strengthening the validity of the conclusion. Specifically, our methodology avoids entanglement in endless debates surrounding: • Clinical Causality: The speculative question of whether the vaccine could biologically cause cancer. • Vaccination Status Specificity: Ambiguities regarding the definition of a fully vaccinated or boosted status. By focusing purely on the severe and asymmetric external validity flaw (the 26.3% incidence suppression and the 31.1% elderly deficit), our findings are robust and fully su/icient to explain the published result as a statistical artifact. The quantitative results, confirming the overall 26.3% suppression of the cohort's CR and the dramatic 45.5% cancer undercount in the oldest non-vaccinated subgroup, firmly establish that the apparent increased cancer risk is a systematic biostatistical artifact rooted in compromised external validity. The observed di/erence is not driven by a genuine biological signal related to the vaccine but is rather a direct consequence of the highly unequal and biased selection of non-vaccinated, high-risk elderly individuals. The severe deficit in the elderly non-vaccinated subgroup means the "unvaccinated" reference group used in [1] was demographically unrepresentative, consequently skewing the crucial baseline risk estimate significantly downwards. This purely demographic mechanism, when quantified, fully resolves the statistical paradox described in our prior analysis [2]. The root cause of this demographic distortion almost certainly lies in the study's cohort selection process, highly likely involving a flawed statistical matching procedure, such as an inverted application of Propensity Score Matching (PSM). The resulting cohorted population is demonstrably younger, healthier, and exhibits a lower baseline incidence because the matching strategy severely under-selected high-risk, non-vaccinated individuals aged >= 65. This fundamental flaw led to the non-vaccinated reference group disproportionately representing the younger, lower-risk segment of the general population. This common epidemiological phenomenon, known as the "Healthy User Bias," or "Healthy Vaccinee E/ect," has been already widely documented in observational vaccine studies [9]. By artificially suppressing the expected cancer incidence in the oldest non-vaccinated cohort segment, the study created a baseline that was profoundly unrepresentative. Consequently, the less-suppressed incidence rate in the vaccinated group is made to falsely appear as a spurious excess risk. Ultimately, the non-vaccinated baseline exhibits an artificially low cancer incidence, which then causes the lesssuppressed incidence rate in the vaccinated group to falsely appear as a spurious excess risk.
While this reappraisal provides a robust quantitative explanation, it is obviously also subject to intrinsic limitations that must be transparently acknowledged. Firstly, this entire work constitutes a re-analysis of aggregate published data; we did not perform primary data collection. Lacking access to individual-level, granular data made it impossible to perform advanced, internal bias correction methods (such as detailed propensity score matching or direct rate standardization for small sub-groups) to correct the internal selection distortion. Consequently, our strong conclusions rely on the quantitative analysis of external validity (i.e., comparing the cohort to the national gold standard) rather than correcting the cohort's internal structure. Secondly, our national standard reference is based on published aggregate statistics, requiring a calculation to establish the exact internal coherence of the Under 65 rate. Notwithstanding these limitations, the sheer magnitude of the observed demographic and incidence deficit is so substantial and asymmetric that the demonstrated bias is highly likely to be the singular, primary explanation for the flawed results reported in [1]. 5. Conclusion The apparent signal of excess cancer risk detected in vaccinated individuals in [1] is the likely consequence of the severe, asymmetric underrepresentation of older, high-risk individuals in the study cohort. Our biostatistical analysis confirms that this demographic distortion artificially suppressed the baseline cancer risk. We conclude that a methodologically rigorous and demographically balanced cohort would predictably show no statistically significant di/erence in cancer incidence between vaccinated and non-vaccinated groups, thereby confirming the safety profile of the COVID-19 vaccine with respect to cancer risk. Data Availability Statement All analyzed data are publicly accessible and can be sourced from the study by Kim HJ et al. (Table S4) [1],and from South Korean National Cancer Statistics and Demographic Information (2022) [5-7]. Author Contributions MR conceived and designed the study, carried out all data collection and analysis, interpreted the quantitative results, and was the sole author responsible for writing and revising the manuscript. The author a/irms full responsibility for the integrity of the data and the accuracy of the data analysis presented.
Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. This study was conducted entirely independently by the author using personal and institutional resources. 9. Conflict of Interest The author declares that there is no conflict of interest, financial, personal, or otherwise, that could be construed as influencing the results or the conclusions presented in this paper. 6. References 1. Kim HJ, Kim M-H, Choi MG, Chun EM. (2025) 1-year risks of cancers associated with COVID-19 vaccination: a large population-based cohort study in South Korea. Biomark Res. 13(114). DOI: 10.1186/s40364-025-00831-w 2. Roccetti M. (2025) Methodological considerations on the external validity of the Kim HJ et al. COVID19 vaccination study (Biomark Res 13:114, 2025): A quantitative analysis. Zenodo Preprint. DOI: 10.5281/zenodo.17434738 3. Marathe M, Vullikanti AK. (2013) Computational epidemiology. Comm. ACM 56(7):88-96. DOI: 10.1145/2483852.2483871 4. Stuart EA. (2010) Matching methods for causal inference: a review and a look forward. Stat Sci 25(1):1–21. DOI: 10.1214/09-STS313 5. Park EH, Jung K-W, Park NJ, et al. (2025) Cancer Statistics in Korea: Incidence, Mortality, Survival, and Prevalence in 2022. Cancer Res Treat. 57(2):312-330. DOI: 10.4143/crt.2025.264 6. Statista. South Korea: Cancer crude incidence rate by age, 2022. Statista; 2024. [Accessed 2025 Nov 2]. Available from: https://www.statista.com/statistics/1440818/south-korea-cancer-crudeincidence-rate-by-age/ 7. World Bank. Population ages 65 and above (% of total population) - Korea, Rep. World Population Prospects, United Nations (UN). [Accessed 2025 Nov 2]. Available from: https://data.worldbank.org/indicator/SP.POP.65UP.TO.ZS?locations=KR