Full text
Statistical Analysis of Arithmetic Concordances in the Quranic Corpus Case Study on a Specific Biographical Profile Idriss Gassama∗ [email protected] ORCID: 0009-0007-2722-5380 December 27, 2025 Abstract This study presents a quantitative analysis of arithmetic concordances identified between the biographical metadata of a subject born in 1987 and the occurrences of the name "Idris" in the Quran (Hafs recitation). A High-Performance Monte Carlo simulation performed on 1,000,000,000 (one billion) random profiles tested a 10-criteria model. While the background noise revealed rare partial convergences, no random artifact achieved a perfect match (10/10). Using the exact ClopperPearson method, the statistical significance is calculated at 5.8 Sigma (5.8σ), with p≈3.0×10−9. To address the "Look-Elsewhere Effect," a second global permutation test was conducted across the entire Quranic corpus (N= 77,407 words), where the maximum stochastic noise reached only 3/10. These results, supported by reciprocal bijectivity tests (Inverse Monte Carlo and Lexical Surrogate Analysis), demonstrate a unique structural correlation: only the subject’s exact date points to the target word, and the target word is the only one in the entire vocabulary to satisfy the model’s equations. 1 Introduction The mathematical analysis of ancient texts has often been the subject of academic debate. This study originates from the serendipitous observation of arithmetic concordances between the author’s biographical metadata and the Quranic text. By adopting a strictly objective "black-box" statistical framework, this research does not aim to interpret theological meaning, but to test the null hypothesis (H0) that correspondences between external data and the internal structure of a 7th-century text are a matter of pure chance. We focus on the only two occurrences of the name "Idris" (Sura 19:56, Sura 21:85) and their relationship with the subject’s temporal and onomastic variables. ∗Independent Researcher, Paris (France). 1
2 Methodology 2.1 The Corpus and Standards To minimize degrees of freedom related to text selection, this study restricts itself exclusively to the most statistically and numerically widespread standards: •Reference Text: Quran, Hafs an ’Asim recitation (Standard Cairo Edition). This choice is justified by its global predominance (>95% of printed copies) and its status as the default reference in digital Quranic studies (Corpus Coranicum, Tanzil.net). •Counting System: The counting system of the Noon Center for Quranic Studies is retained as a third-party reference (114 Suras, 6236 Verses, 77407 Words). The use of a pre-existing and independent dataset prevents any ad hoc adjustment of word segmentation by the author to favor results. •Digital Encoding: Abjad System (Standard Arabic Gematria) and Latin Alphabetical Rank. 2.2 Definition of Variables Input parameters are fixed a priori and kept constant: •Temporal Variables (T): D= 19,M= 3,Y= 1987,YXX = 87,H= 8 (Day, Month, Year, Hour). Corresponding Hijri Date: 19/07/1407. •Pivot Constant (K): Defined by the product D×M×H= 456. •Identity Variables (I): Gematria values of names (Idriss: Arabic=275 / Latin=78 ; Gassama: Arabic=103 / Latin=61 ; France: Arabic=391 / Latin=47). •Sociolinguistic Context (C): As the subject is of French nationality, the geographic variable "France" and phonetics in the French language are retained as native constraints of the model (fixing criteria C9 and C25). 2.3 Search Space Constraints To prevent overfitting and combinatorial explosion (Data Dredging), the search space was strictly bounded prior to analysis: 1. Restricted Operator Set (K= 4): Only elementary integer arithmetic is permitted: Addition (+), Subtraction (−), Multiplication (×), and Concatenation (||). Division is explicitly excluded to maintain integer integrity. 2. Input Variable Cap (k≤4): Each equation is limited to a maximum of 4 distinct input variables drawn from the pool of the variables. This constraint drastically reduces combinatorial possibilities by excluding long complex chains. Note that constant textual targets (e.g., Sura numbers) are outputs, not input variables. 2
3. Maximal Hierarchical Depth: The complexity is limited to a nesting depth of 2 (Depth 2). Note: Due to the associative property of addition and multiplication, homogeneous chains (e.g., A×B×C) are considered as a single hierarchical level (Depth 1), whereas mixed operations requiring parentheses (e.g., (A+B)×C) represent a higher complexity level (Depth 2). 4. Non-Repetition Constraint: A "without replacement" rule is applied; a specific temporal constant cannot be used more than once in the same equation equation to force structural coherence. 2.4 Selection of the 10-Criteria Subset While over 30 distinct arithmetic concordances were initially identified, we deliberately restricted the statistical model to a subset of 10 fundamental criteria. This choice is governed by the principle of parsimony (Occam’s Razor). The rationale is twofold: (i) Statistical Saturation: If a model based on 10 elementary relations is sufficient to defy stochastic replication at a high level, further complexity is redundant for the rejection of the null hypothesis. (ii) Robustness against Overfitting: By minimizing the number of parameters, we ensure that the identified signal is not an artifact of excessive data-mining but a result of robust structural correlations. This selection focuses on independent functional domains (positional, temporal, and onomastic) to maintain high statistical power while ensuring simplicity in the mathematical verification process. 2.5 Model Validation Criteria Table 1 details the 10 independent criteria constituent of the model. Table 1: Fundamental Criteria of the Statistical Model Code Type Mathematical Definition C1 Arithmetic Global Rank(Word) = D×M×(YXX )×H C2 Structure Sura/Verse Address: (S=D, V = Π(YXX)) C3 Arithmetic concatenation(M|D)+ Y = Global Verse Rank C4 Fractal 6236 −Gem(Word) =Y×M C5 Calendar Verse Rank + Verse Gem = D×M×Day Rank C7 Pivot Σ(CoordsS+V)+Gem(Word) =Pivot C8 Internal Gem(Name) + Pivot = Internal Rank(Word) C9 Symmetry Iterative Prime Validation (kmin search) C12 Lock Primality Validation (Latin Identity + Verse) C25 Linguistic Phonetic Date (FR) = ΣGematria (AR) 3
3 Results 3.1 Mechanisms (Case Studies) Case A: Symmetry of Prime Numbers (Criterion C9) This criterion utilizes an iterative search for the first mathematical success rather than an arbitrary choice. •The Axis (A): Reversed concatenation of the date (Month|Day): A= 319. •The Gap (K): Standard Pivot: K= 456. •Iterative Search: We seek the smallest integer k≥1such that P1=|A×k−K| and P2=|A×k+K|are simultaneously prime. •Result: The first pair is identified at k= 7. At this iteration, P1= 1777 and P2= 2689. •Verification: 1777 is the 275th prime number (Gem: Idris); 2689 is the 391st prime number (Gem: France). Case B: Linguistic Validation (Criterion C25) The sum of alphabetical ranks for the French phonetic date "dix neuf trois quatre vingt sept" equals 378, which strictly matches the subject’s complete Arabic gematria (Idriss 275 + Gassama 103). 3.2 Statistical Validation (Monte Carlo Simulation) A sample of 109random profiles (Years 1950–2025, using 533 first names and 147 surnames) was generated. Score Distribution (Target: Idris) No false positive reached 10/10. The maximum observed noise for the target "Idris" reached 9/10 with an extremely low frequency. Table 2: Score distribution over N= 109simulations (Corrected) Score Occurrences Frequency (f) 0/10 990,574,565 99.06% 1/10 9,163,383 0.92% 2/10 208,797 2.1×10−4 3/10 16,283 1.6×10−5 4/10 34,709 3.5×10−5 5/10 721 7.2×10−7 6/10 1,472 1.5×10−6 7/10 66 6.6×10−8 8/10 3 3.0×10−9 9/10 1 1.0×10−9 10/10 0 0.00 Total 1,000,000,000 100% 4
Significance Analysis Since the actual observation (10/10) remains unique against a maximum noise of 9/10, we calculate the probability of obtaining a score ≥10 by chance with k= 0 successes out of n= 109trials (Exact Clopper-Pearson method at 95% confidence): Pupper = 1 −(0.05)1/109≈3.0×10−9(1) The conversion to standard deviation (Sigma) via the normal distribution yields: Z= Φ−1(1 −Pupper)≈5.8σ(2) This result validates the hypothesis of an extreme statistical anomaly. Global Robustness: Corpus-Wide Permutation Test To verify the selectivity of the model across the entire text, a "Total Permutation" test was performed by matching 109random profiles with random words from the 77,407-word corpus. Table 3: Global score distribution over 109random Profile-Word pairs Score Occurrences Frequency (f) 0/10 996,311,698 99.63% 1/10 3,682,647 0.36% 2/10 5,652 5.6×10−6 3/10 3 3.0×10−9 4-10/10 0 0.00 The maximum score achieved by pure chance across the entire corpus was strictly 3/10. The 7-point gap between the stochastic noise and the subject’s 10/10 result confirms the significance of 5.8 Sigma. 3.3 Bijectivity Validation (Reciprocal Tests) •Inverse Monte Carlo: Out of 2,000,000 random dates, only 19/03/1987 achieved full convergence for "Idris." •Exhaustive Surrogate Analysis: Out of 21,311 unique words, only "Idris" satisfied the model for the subject’s metadata. 4 Discussion 4.1 Robustness against the Look-Elsewhere Effect A common critique in quantitative text analysis is the "Look-Elsewhere Effect" (or the multiple comparisons problem), where a model might be adjusted post-hoc to fit observed data. However, the constraints defined in Section 2.2 (elementary arithmetic operators, strictly non-repetitive native variables, and low Kolmogorov complexity) effectively restrict the functional search space. We estimate the conservative order of magnitude for these functional combinations to be approximately Nspace ≈105to 106. 5
Even when applying a highly conservative Bonferroni correction based on the upper bound of this search space (106) to the raw Monte Carlo p-value (p≈3.0×10−9), the adjusted p-value remains highly significant: padj =praw ×Nspace ≈0.003 (3) This adjusted value (0.003) comfortably satisfies the standard scientific threshold for significance (α= 0.05), confirming that the observed signal is clearly distinguishable from combinatorial noise. 4.2 Epistemological Robustness: Infinite Conceptual Space and Kolmogorov Complexity A profound critique often addressed to quantitative textual studies concerns the "Infinite Conceptual Space." It argues that in an unbounded search for meaning, any correlation becomes inevitable. We counter this epistemological objection through three distinct layers of evidence: 1. Identity-to-Identity Anchoring: Our search was not exploratory but confirmatory. We did not pick a random word among 77,407; we tested the link between two identical onomastic entities (Subject’s name ↔Quranic name). This drastically restricts the conceptual space from "any possible meaning" to "identity-specific coordinates." 2. Low Kolmogorov Complexity: According to algorithmic information theory, the probability of a correlation being a post-hoc artifact increases with the complexity of the rules required to describe it. Our model relies on a "short program" (elementary arithmetic: +,−,×,||). This low Kolmogorov complexity suggests that the model has captured a robust structural signal rather than a complex noise-based coincidence. 3. Functional Bijectivity: In an infinite conceptual space, one would expect a "oneto-many" relationship (one date pointing to many words). However, our surrogate analysis demonstrates perfect bijectivity. The "combinatorial sealing" evidenced by the 3/10 ceiling in the global permutation test (N= 109) proves that the Quranic text remains inert to random biographical profiles, regardless of the conceptual breadth explored. 4.3 Comparative Analysis and Intellectual Dialogue This research aligns with the field of Computational Theology, where sacred corpora are treated as structured databases. However, it specifically diverges from previous "numerical miracle" studies (e.g., Khalifa, 1974) which often suffered from selection bias and lack of statistical bounds. By introducing probabilistic rigor and functional bijectivity, we distinguish a "signal" from "noise." While previous works showed a date pointing to a word, we demonstrate that only that date points to that word, and that the word is the unique solution in the entire vocabulary. 6
4.4 Limitations and Future Research Directions •Extensional Validation: Applying this model to other historical biographical profiles to further refine the "False Positive" rate. •Cross-Linguistic Comparison: Testing similar alphanumeric functions on the Hebrew Bible or Greek New Testament to determine if these anomalies are corpusspecific. 5 Conclusion This study identifies and quantifies an objective statistical anomaly with a significance of 5.8 Sigma (p≈3.0×10−9). The results from one billion permutations across the entire Quranic corpus, combined with perfect bijectivity, strongly exclude the hypothesis of a generic statistical artifact. These findings reveal a mathematically unique structural correlation between the subject’s biography and the target text, opening new avenues for quantitative textual analysis. Data and Code Availability To ensure reproducibility, the full Python source code for the 109simulations and datasets are available on the Zenodo repository (doi.org/10.5281/zenodo.17676291). 7
References [1] N. Metropolis and S. Ulam, “The Monte Carlo Method,” Journal of the American Statistical Association, vol. 44, no. 247, pp. 335–341, 1949. [2] Efron, B., & Tibshirani, R. J. (1994). An Introduction to the Bootstrap. CRC Press. [3] Good, P. (2005). Permutation, Parametric, and Bootstrap Tests of Hypotheses. Springer Science & Business Media. [4] C. J. Clopper and E. S. Pearson, “The use of confidence or fiducial limits illustrated in the case of the binomial,” Biometrika, vol. 26, no. 4, pp. 404–413, 1934. (Reference for the exact method used for k= 0 successes). [5] H. Jeffreys, Theory of Probability, 3rd ed., Oxford University Press, 1961. [6] T. Sellke, M. Bayarri and J. O. Berger, “Calibration of p-Values for Testing Precise Null Hypotheses,” The American Statistician,55 (1), 62–71 (2001). [7] C. E. Bonferroni, “Teoria statistica delle classi e calcolo delle probabilità,” Pubblicazioni del R Istituto Superiore di Scienze Economiche e Commerciali di Firenze, vol. 8, pp. 3–62, 1936. (Reference for the statistical correction applied to the search space). [8] E. Gross and O. Vitells, “Trial factors for the look-elsewhere effect in high energy physics,” The European Physical Journal C, vol. 70, no. 1, pp. 525–530, 2010. (Theoretical framework for quantifying statistical significance in large search spaces). [9] Al-Qur¯an al-Kar¯ım (Standard Egyptian Edition), Amiri Press, Cairo, 1342 AH [1924 CE]. (Canonical reference text for the Hafs recitation used in this study). [10] Centre Noon for Qur’¯anic Studies, Word, Letter, and Verse Enumeration Tables for the Canonical Hafs .Text (Dataset based on the Medina Codex), Available at: http://www.islamnoon.com/content/887/1 (Accessed 30 June 2025). [11] G. Ifrah, The Universal History of Numbers: From Prehistory to the Invention of the Computer, John Wiley & Sons, 1998. (Reference for the Abjad numeral system and Semitic gematria). [12] B. Jarrar, Irh¯as¯at al-Ij¯az al-Adad¯ı f¯ı al-Qur¯an al-Kar¯ım [Premonitions of Numerical Miracles in the Holy Quran], Noon Center for Qur’anic Studies, Ramallah, 1998. (Seminal work discussing the mathematical balance “Al-Mizan” and the number 456). [13] “Idris (prophet),” Wikipedia, The Free Encyclopedia,https://en.wikipedia.org/ wiki/Idris_(prophet) (Accessed 30 June 2025). 8
A Complete Inventory of the 30 Concordances This table presents all arithmetic and structural anomalies identified during the exploratory study. Table 4: Synthesis of the 30 Numerical Concordances No. Category Description Formula / Proof 01 Arithmetic Global rank of 1st word "Idris" D×M×87 ×H= 39 672 02 Structure Address 19:56 linked to Date S= 19,V= 8 ×7(Short Year) 03 Arithmetic Global Verse Rank (2306) 319(Month|Day) + 1987 = 2306 04 Recursive Difference Total Verses/Name 6236 −275 = 1987 ×3 05 Mixed Rank + Verse Gematria 2306 + 2140 = 19 ×3×78 06 Primes Global Rank via Prime Numbers 193(Day|Month) + P(319(Month|Day)) = 2306 07 Pivot Sum Coordinates + Name Σ(Coords) + 275 = 456 (Pivot) 08 Internal Word Internal Rank (559) 559 + 275 = 456 + 378(275+103) 09 Primes 1st Pair of Primes (Pivot Gap) Axis 319 ×7±456 →P(275) and P(391:G"France") 10 Primes Pair generated by Pivot and Rank Diff(Pa, Pb)=456 →Sura Titles 11 Calendar Solar/Hijri Conversion (Solar−ΣRanks)±Rank = 1433 12 Convergence Convergence P(x)±x P (78) −78 = 319 and P(56) + 56 = 319 13 Metadata Latin Name in Sura Titles Σ(Titles of Name) = 2306 −559 14 Semantic Inverse Rank defines Hijri Year Inv(2306) −2306 = 1407 + G("Hijri") 15 Pivot Verse 21:85 Gematria to Date 1889 −456 = 1433 16 Symmetry Identity (6 letters / 7 letters) 6×7 = 42 (S.22). S.22 has 78 verses. 17 Cluster Quadruple convergence on 22:27 Sum, Product, Context = 378 18 Metadata Verse 22:27 Gematria linked to Titles 3826 −456 = Σ(Titles "Gassama") 19 Fusion Global Rank = Sum of Times 2306 = 873(Year|Month) + 1433 20 Letter Rank of 1st letter Idris 2200 = 275 ×8 21 Theology Balance Idris + Ilyas Σ(Idr+Ily) = 809 = 456+(275+ 78) 22 Primes Sum Rangs Idris+Ilyas 2306 + 3911 = 6217 = P(809) 23 Identity Complete Epithet Gematria "Idriss Gassama Al-Fransi" = 809 24 Calendar Gap Solar/Hijri sums Σ(Solar)−Σ(Hijri) = 378 25 Linguistic Short Date Phonetics (French) Value("dix neuf trois...") = 378 26 Linguistic Hijri Date Phonetics Value("dix neuf sept mille...") = 378 27 Linguistic Full Date (French) Value(Full Date) = 378 + 139 28 Linguistic Latin Name to Hijri Date Latin Name (227)→P(227) = 1433 29 Calendar Temporal Chiasmus 809 Greg. Year 809 = Hijri 193 ; Hijri 809 = Greg 1407 9