scieee AI-readable full text Open interactive document viewer

Parallelism in Neurodegenerative Biomarker Tests: Hidden Errors and the Risk of Misconduct

petzold, axel

Full text

Parallelism in Neurodegenerative Biomarker Tests: Hidden Errors and the Risk of Misconduct Axel Petzold∗and Joachim Pum†and David P. Crabb‡ November 19, 2025 Abstract Biomarkers are critical tools in the diagnosis and monitoring of neurodegenerative diseases. Reliable quantification depends on assay validity, especially the demonstration of parallelism between diluted biological samples and the assay’s standard curve. Inadequate parallelism can lead to biased concentration estimates, jeopardizing both clinical and research applications. Here we systematically review the evidence of analytical parallelism in body fluid (serum, plasma, cerebrospinal fluid) biomarker assays for neurodegeneration and evaluate the extent, reproducibility, and reporting quality of partial parallelism. This systematic review was registered on PROSPERO (CRD42024568766) and conducted in accordance with PRISMA guidelines. We included studies published between December 2010 to July 2024 without language restrictions. Eligible studies included original research assessing biomarker concentrations in body fluids with data suitable for evaluating serial dilution and standard curve parallelism. The data extraction for interrogating parallelism included dilution steps, measured concentrations, and sample types. For each study we generated parallelism plots in a uniform and comparable way. These graphs were used to come to a balanced decision on whether parallelism or partial parallelism were present. The risk of bias was assessed based on sample preparation, buffer consistency, and methodological transparency. Of 44 eligible studies, 19 provided sufficient data for generating 49 partial parallelism plots. Of these plots, only 7 (14%) demonstrated clear partial parallelism. Partial parallelism was typically achieved over a narrow dilution range of about three doubling steps. Most assays deviated from parallelism, risking overor underestimation of biomarker levels if determined at different dilution steps. A high ∗University College London, Queen Square Institute of Neurology, London, WC1N 3BG, United Kingdom. Email: a.p[email protected] †LABanalytics GmbH, Anton-Bruckner-Weg 22, Jena, Germany ‡City St. George’s, University of London, Northampton Square, London, EC1V 0HB, United Kingdom risk of bias was identified in 9 studies using spiked or artificial samples, inconsistent dilution buffers, or incomplete reporting. Several studies assessed sample-to-sample parallelism rather than sampleto-standard, contrary to guidelines by regulatory authorities. In conclusion, partial parallelism was infrequently observed and inconsistently reported in most biomarker assays for neurodegeneration. Narrow dilution ranges and variable methodologies limit generalizability. Transparent reporting of dilution protocols and adherence to established analytical validation guidelines is needed. This systematic review has practical implications for clinical trial design, regulatory approval processes, and the reliability of biomarker-based diagnostics. Keywords: Assay validation; parallelism; matrix effects; hook effect; protein aggregation. 1 Parallelism Page 2 Contents 1 Introduction 2 2 Terminology and concept 4 2.1 Parallelism and standard curves . . . 4 2.2 Determination of parallelism . . . . . 5 3 Clinical context 6 4 Methods 7 4.1 Eligibility Criteria . . . . . . . . . . 7 4.2 Information Sources . . . . . . . . . 7 4.3 Search Strategy . . . . . . . . . . . . 7 4.4 Study Selection . . . . . . . . . . . . 8 4.5 Data Collection Process and Data items.................. 8 4.6 Risk of Bias in Individual Studies . . 8 4.7 Summary Measures . . . . . . . . . . 8 4.8 Synthesis of Results . . . . . . . . . 8 4.9 Risk of Bias Across Studies . . . . . 8 5 Results 9 5.1 Study Selection . . . . . . . . . . . . 9 5.2 Study Characteristics . . . . . . . . 9 5.3 Risk of Bias Within Studies . . . . . 9 5.4 Results of Individual Studies . . . . 9 5.4.1 Amyloid β.......... 9 5.4.2 α-Synuclein . . . . . . . . . . 11 5.4.3 DJ-1.............. 12 5.4.4 Tau Protein . . . . . . . . . . 12 5.4.5 Apolipoprotein E . . . . . . . 14 5.4.6 Neurofilament proteins . . . . 14 5.4.7 Glial fibrillary acidic protein 14 5.4.8 Dipeptide repeat proteins . . 14 5.5 Synthesis of Results . . . . . . . . . 15 5.6 Risk of Bias Across Studies . . . . . 15 6 Discussion 17 7 Conclusion 23 7.1 PRISMA 2020 Checklist . . . . . . . 31 1 Introduction Biomarker assays have become central to advancing research and clinical applications in neurodegeneration. For instance, the astrocytic biomarker glial fibrillary acidic protein (GFAP) has been approved by the U.S. Food and Drug Administration (FDA) as part of a panel to guide decisions on brain imaging after trauma [1]. Similarly, the FDA granted rapid approval for novel treatments in multiple sclerosis (MS) [2] and amyotrophic lateral sclerosis (ALS) [3] on the basis of neurofilament (Nf) biomarkers [4]. The accurate quantification of such biomarkers relies on robust assay performance. To ensure this, the FDA and other regulatory authorities have established comprehensive frameworks for biomarker validation, covering all stages from sample collection and processing to analytical validation, clinical application, quality control, and accreditation. However, the breadth of this framework can also create ambiguity. For example, the terms “sensitivity” and “specificity” are used differently in analytical versus clinical contexts (Table 1). This distinction is illustrated by neurofilament light chain (NfL), which has a reported analytical specificity of 99.3% [5] due to minimal cross-reactivity with other Nf isoforms, yet demonstrates low clinical specificity because blood NfL levels rise across a wide range of neurological diseases [2, 3] and in many conditions that compromise the integrity and function of neurons and their connections [4,6]. The framework for laboratory test evaluation presented in Table 1 requires additional clarification because certain terms are frequently misinterpreted in the literature [8]. A notable example concerns the distinction between “parallelism” and “dilution linearity” Guidance documents such as ISO/IEC 17025 and EURACHEM refer to parallelism when illustrating method validation, but the concept largely originates from applied laboratory practice. In practical use, parallelism and dilution linearity are sometimes conflated because both involve serial dilutions. The key distinction lies in the sample type: dilution linearity experiments use samples spiked with analyte at concentrations designed to minimize matrix effects, whereas spiked samples are not permissible for assessing parallelism [8, 9]. Although this difference may appear subtle, as will be argued in the present critical review, it has important implications for assay validation and interpretation. Beyond these terminological distinctions, core analytical parameters such as precision, accuracy, linearity, and the evaluation of matrix effects remain central to assay validation [7]. Precision and accuracy ensure that results are both reproducible and correct, while linearity confirms that measured values remain proportional to analyte concentrations across the assay’s dynamic range [7]. These parameters can be compromised by matrix effects, which introduce systematic bias [8,10]. This challenge is particularly pronounced when analyzing complex biological samples such as serum and plasma [7, 8, 10], especially in mass spectrometrybased assays [11–13]. To mitigate such issues, harmonization across laboratories and analytical platforms has become a major focus in biomarker validation studies [14–17]. The goal is to ensure comparability of results and reliability of clinical decision-making [8]. As emphasized earlier, one of the hidden sources of error is the lack of parallelism [18, 19]. For this reason, regulatory authorities require demonstration of parallelism between serially diluted samples and the assay caliCritical Reviews in Clinical Laboratory Sciences Parallelism Page 3 Table 1: Framework for laboratory test validation as provided by regulatory authorities. Abbreviations: U.S. Food and Drug Administration = FDA, Conformité Européenne = CE, College of American Pathologists = CAP, Clinical and Laboratory Standards Institute = CLSI, International Organization for Standardization = ISO, analytical measurement range = AMR, limit of the blank = LoB, limit of detection = LoD, limit of quantitation = LoQ. Table adapted from reference [7] with a focus on quantitative laboratory tests. Laboratory test ISO 15189 / 17025 CAP FDA approved / CE certified In-house & modified FDA approved / CE certified FDA approved / CE certified In-house & modified FDA approved / CE certified Precision & Bias Verify Establish Verify Establish CLSI EP15-A3 CLSI EP15-A3 CLSI EP15-A3 CLSI EP15-A3 Method Comparison Experiment Compare with current method Compare with reference method Compare with current method Compare with reference method CLSI EP09-A3 Regression Analysis, Difference Plot CLSI EP09-A3 Regression Analysis, Difference Plot CLSI EP09-A3 Regression Analysis, Difference Plot CLSI EP09-A3 Regression Analysis, Difference Plot Analytical Sensitivity — Establish Verify Establish — CLSI EP17-A2 LoB, LoD, LoQ, Precision Profile Approach, Probit Analysis Documentation from manufacturer or literature, CLSI EP17-A2 Verification of LoB, LoD, LoQ CLSI EP17-A2 LoB, LoD, LoQ, Precision Profile Approach, Probit Analysis Analytical Specificity — Establish Verify Establish — CLSI EP07-A Test for hemolysis, icterus and lipemia, potential crossreactivities Documentation from manufacturer or literature CLSI EP07-A Test for hemolysis, icterus and lipemia, potential crossreactivities Diagnostic Sensitivity — — — — — — — — Diagnostic Specificity — — — — — — — Linearity Verify Establish Verify Establish Evaluation of linearity, Calibration Verification, Verification of AMR Evaluation of linearity, Calibration Verification, Establish AMR Evaluation of linearity, Calibration Verification, Verification of AMR Evaluation of linearity, Calibration Verification, Establish AMR Carryover Verify Establish Verify Establish Standard protocol or Short protocol Standard protocol Standard protocol or Short protocol Standard protocol Measurement Uncertainty Establish Establish Establish Establish ISO 11352 / NORDTEST Method ISO 11352 / NORDTEST Method ISO 11352 / NORDTEST Method ISO 11352 / NORDTEST Method Reference Range / Cut-off Verify Establish Verify Establish Documentation from manufacturer or literature CLSI EP28-A3c Use direct or indirect method for data collection CLSI EP28-A3c (Use direct or indirect method for data collection) OR Documentation from manufacturer or literature CLSI EP28-A3c (Use direct or indirect method for data collection) Critical Reviews in Clinical Laboratory Sciences Parallelism Page 4 bration curve [20, 21]. Demonstrating parallelism ensures that measured concentrations remain consistent and accurate across the relevant dilution range. Yet, the extent to which guidelines [20,21] for testing parallelism are systematically applied in biomarker validation studies remains unknown. Therefore, we systematically reviewed the biomarker literature on neurodegeneration for evidence of parallelism testing, beginning with the first report of its absence in the neurofilament heavy chain (NfH) ELISA, which was attributed to protein aggregate formation [18]. Protein aggregation is a hallmark of many neurodegenerative diseases [22], meaning that the lack of parallelism observed in NfH assays has implications far beyond a single biomarker [18]. Building on this observation [18], we examined subsequent studies to determine whether other biomarkers are similarly affected. Deviations from parallelism [23] can lead to systematic overor underestimation of biomarker concentrations, creating risks for misinterpretation that are particularly consequential in the context of clinical trials [2, 3] and regulatory submissions [1,20,21]. Finally, our critical review underscores the limitations of current reporting practices in biomarker validation studies, which often include inconsistent dilution protocols, incomplete reporting of dilution ranges, and reliance on non-representative spiked or artificial samples. With regard to terminology, we begin by clarifying the concept of parallelism, a term originally rooted in geometry. Although statistical approaches to testing parallelism are available, detailed discussion of these methods is provided in the supplementary materials. In the main text, we instead focus on visual methods, as they are more accessible to readers without a strong mathematical background. This approach is supported by illustrative examples to ensure clarity. We then present the methodology of our systematic review, conducted in strict accordance with the PRISMA 2020 guidelines. The resulting findings are interpreted in the context of clinical laboratory science and extended to address broader methodological and regulatory issues. Finally, we highlight practical recommendations for laboratories, manufacturers, and regulators, with the aim of ensuring applicability beyond the research community. 2 Terminology and concept The formal concept of parallelism can be traced to the Greek mathematician, geometer, and logician Euclid (Εὐκλείδης), who, around 300BCE, authored the seminal treatise Elements. In Book I, Definition 23, Euclid defined parallel lines as “straight lines which, being in the same plane and being produced indefinitely in both directions, do not meet one another in either direction.” However, it is Euclid’s fifth postulate, known as the parallel postulate, that has historically attracted the most scrutiny. Over the centuries, mathematicians sought to prove the fifth postulate using Euclid’s other axioms. These efforts persisted for over two millennia and produced many flawed or incomplete proofs, reflecting a significant historical misconception [24]. It was not until the 19th century that mathematicians like Carl Friedrich Gauss (working on parallelism 1779–1844), Lobachevsky (1829– 30), and Bolyai (1832) demonstrated that entirely consistent non-Euclidean geometries could be constructed by replacing the fifth postulate with alternative versions [24, 25]. In a letter to Bolayi (17DEC-1799) Gauss wrote: “It is true that I have come upon much which by most people would be held to constitute a proof: but in my eyes it proves as good as nothing.” Similarly, in the context of present, critical, systematic review, parallelism, though seemingly straightforward at first glance, demands careful scrutiny. Just as for Euclid’s parallel postulate, analytical parallelism in quantitative assays must not be assumed but must be rigorously tested, substantiated, and importantly lack of parallelism must be understood. 2.1 Parallelism and standard curves The accurate determination of analyte concentrations using a standard curve requires that the dilution series of the test sample exhibits parallelism with the standard curve [26]. This means the two curves must be similar functions, differing only by a scaling factor along the dose axis (in present systematic review this is always the y-axis), so that interpolation yields valid and reproducible results. Without demonstrated parallelism, calculations derived from a standard curve may lead to inaccurate quantification [26]. In its most simple form interpolation is linear. Hence the to verify/establish linearity for a laboratory test in Table 1. Linear interpolation is a method of curve fitting using linear polynomials to construct new data points within the range of a discrete set of known data points [27]. It is absolutely crucial to recognise that for biomarker concentrations from samples calculations are only permitted for data points within the range of the points of the standard curve. Expanding from linear interpolation, quadratic, cubic, four-parameter logistic (4PL) [28,29], 4PL with logarithmic scaling of the dose (y-axis) [30], and five-parameter logistic (5PL) [29,31–33] standard curves have entered contemporary laboratory routine. Extrapolation is not allowed [7,27,30]. It is mandatory to demonstrate parallelism for the selected type of a standard curve and real-world samples [20]. Critical Reviews in Clinical Laboratory Sciences Parallelism Page 5 Despite its importance, the validation of parallelism presents a practical challenge for many laboratory workflows [7]. The concept, as reviewed here and long known in analytical chemistry, has only more recently gained sustained attention in the context of regulatory method validation for biomarker assays [10, 21]. Variability in internal standards (ISV) has been identified as a core factor contributing to measurement error, further underscoring the need for robust validation. Regulatory agencies have begun to address the issue of parallelism. The FDA, through its 2022 M10 Bioanalytical Method Validation (BMV) guidance, and the European Medicines Agency (EMA) both highlight parallelism as a critical validation criterion [20,34]. According to the EMA, parallelism is defined as follows: “Parallelism demonstrates that the serially diluted incurred sample response curve is parallel to the calibration curve” [20]. Spiked samples in parallelism The use of spiked samples in bioassay validation is an interesting strategy to assess analytical performance, particularly in relation to matrix effects and assay specificity [10]. The use of spiked samples is frequently employed in the context of mass spectrometry-based methods, where matrixinduced variability is a major concern to the accuracy of quantification [11, 35]. Current guidelines suggest that variations of less than 20% are generally acceptable when comparing spiked samples with calibrators or quality controls [10,36,37]. However, this approach typically focuses on performance at fixed concentration levels and does not explicitly evaluate assay behaviour across wider dilution ranges or in the presence of other factors that apply to real-world samples [37]. Therefore one of the key limitations of spiked samples is the requirement for high-concentration material to enable serial dilution and minimize matrix effects [37]. Although spiking can be informative under controlled conditions, it introduces several methodological concerns. First, the reference material used for spiking may differ structurally or functionally from the endogenous biomarker. This includes the use of truncated peptide sequences or proteins derived from non-human species [38, 39]. Second, the stability of spiked material can be suboptimal, particularly over long-term storage, where protein aggregation may occur [39]. Aggregation, a known source of error in parallelism assessment [18], cannot be adequately simulated using spiked samples. Moreover, spiked samples do not replicate the complexity of physiological matrices. As noted in prior reviews, their utility is limited to evaluating internal standard variability introduced by physiological differences between control and patient matrices [37]. Even when the spiked analyte behaves additively with the endogenous biomarker, this assumption may not hold under all conditions, and further concerns regarding non-linear interactions persist [40]. For these reasons, real-word, patient-derived samples remain the gold standard for assessing parallelism [8,9]. Their use better reflects the biochemical and biophysical properties relevant to clinical measurement [41]. In the context of this systematic review, the use of spiked samples is discussed as a potential source of bias. 2.2 Determination of parallelism Visual assessment offers a complementary and often more intuitive approach to evaluating parallelism in bioassays [26]. Unlike formal statistical testing (see supplementary materials), graphical methods allow analysts to inspect the behaviour of dilution curves directly, which is particularly valuable in the presence of noisy or incomplete datasets. Statistical expertise required to properly interpret hypothesis testing methods [42–45] is not commonly included in the core training of laboratory scientists, clinicians, or regulatory reviewers. This gap may lead to an overreliance on statistical outcomes, potentially overlooking meaningful deviations in curve behaviour across extended dilution ranges [23]. In real-world applications, complete parallelism is rare, while partial parallelism is often observed within limited dilution ranges [23]. This practical observation has led to the operational definition of partial parallelism, acknowledging that assays may behave acceptably within specific, predefined ranges without meeting idealized criteria across the entire curve [46]. The value of visual methods for detecting lack of parallelism has been recognized in the literature since the landmark paper by Plikaytis [26], and has been expanded by subsequent studies [23, 47–49]. Figure 1 illustrates the most common patterns observed. The interpretation of the partial parallelism plots goes back to Euclid’s geometry. The two lines in the partial parallelism plot need to be parallel. To be more precise they need to be parallel to each other, on the y-axis, for at least part of the graph (Pattern 3 in Figure 1). Hence the term partial parallelism. A vertical offset is permitted, which is also entirely consistent with the requirements for statistical testing [43]. Interpretation of partial parallelism plots is best if data for the standard curves are present. For the purpose of this systematic review partial parallelism plots will be created for each study included. Based on the graphical interpretation, which is presented for each study, a binary decision will be made: presence (see Figure 1 Critical Reviews in Clinical Laboratory Sciences Parallelism Page 6 Figure 1: Visual assessment of parallelism. The figure illustrates characteristic patterns used to compare dilution ranges (x-axis) of the assay calibration standard (blue dotted line) with theoretical sample responses (black lines). Such visual patterns provide an accessible means of determining whether parallelism is maintained or lost and form the basis of our approach for evaluating parallelism across the assays included in this critical and systematic review. Pattern 1 shows an apparent increase in biomarker concentration with greater sample dilution, a phenomenon typically associated with the release of biomarker molecules from protein aggregates [18]. Pattern 2 shows a decrease in measured biomarker concentration with successive dilution steps, commonly observed when concentrations approach the assay’s lower limit of detection (LoD; see Table 1). This underscores why regulatory authorities require validation of analytical sensitivity [7]. Pattern 3 demonstrates that parallelism may be achieved within a limited dilution range but lost again at higher dilutions. This pattern is termed partial parallelism [23]. As highlighted throughout this review, observed parallelism is most often partial. pattern 3) or absence (Figure 1 patterns 1&2, and other patterns of deviation from parallel alignment) of partial parallelism. 3 Clinical context The clinical implications of reporting artificially elevated biomarker concentrations due to lack of parallelism (for example Pattern 1 in Figure 1) can be illustrated using neurofilaments. Both NfL [50,51] and NfH [52,53] have been established as valuable prognostic biomarkers in patients with GuillainBarré syndrome (GBS) [54,55]. GBS is characterized by evolving paralysis, which in severe cases can compromise respiratory function and necessitate intensive care unit (ICU) admission with mechanical ventilation. While most patients experience transient paralysis followed by substantial recovery, a subset develop long-term disabilities, including the permanent loss of ambulation. Early identification of high-risk patients would therefore be of great clinical value, enabling personalized management strategies ranging from ICU admission decisions to treatment escalation. Targeting high-efficacy, but often more costly, treatments to patients at greatest risk offers the dual benefit of optimizing clinical outcomes and improving resource allocation. Enhanced functional recovery in this subgroup would not only improve patient quality of life but also reduce downstream healthcare costs, including those associated with long-term disability, loss of work capacity, increased care needs, and broader economic deadweight losses. The importance of such strategies is reflected in World Health Organization (WHO) guidance developed in collaboration with the Universal Health Coverage Partnership (UHC Partnership). Their report, covering 115 countries and representing more than three billion people, explicitly recognizes the essential role of clinical laboratory services in healthcare delivery [56]. While the report does not specifically address technical parameters such as parallelism, this level of detail is beyond its scope, it implies that such issues may be incorporated under the remit of National Quality Control Laboratories [56]. To illustrate how the lack of demonstrated parallelism can have wide-ranging consequences, we highlight one recent example from the literature of NfL in GBS [57]. In this study, NfL concentrations were measured using a commercially available assay. The manufacturer recommends quantification at a dilution of 1:4. However, for a subset of samples, measurements were performed at a much higher dilution (1:400) without demonstrating parCritical Reviews in Clinical Laboratory Sciences Parallelism Page 7 allelism between the diluted samples and the calibration standard across this extended range. The inferred implications are substantial: •Inflated biomarker concentrations: Reported values ranged from approximately 1 pg/mL to nearly 10,000 pg/mL, which is several orders of magnitude higher than expected under either physiological variation or disease-related pathology. •Potential patient misclassification: Elevated NfL levels were associated with more severe phenotypes defined by electrodiagnostic criteria [54, 58]. The data imply that artificially inflated concentrations may have been interpreted as markers of irreversible axonal degeneration. •Bias in predictive modeling: Predictive models are invaluable for risk stratification, but if their statistical significance arises from artificially elevated biomarker concentrations in a subgroup of severely affected patients, these models are unlikely to replicate in independent cohorts, clinical trials, or clinical practice. •Impact on treatment evaluation: The study reported no treatment effect on NfL levels. One clinical interpretation might be to withhold a potentially effective therapy. However, the extreme variability in reported concentrations likely increased noise and obscured real effects, deviating from prior treatment trials that adhered to validated dilution ranges [2,3]. •Risk of deliberate misuse: A particularly concerning possibility arises in the context of treatment trials. If placebo samples were analyzed at a higher dilution range than treatment samples under Pattern 1 conditions (Figure 1), or at a lower range under Pattern 2 conditions, this would introduce systematic bias favoring a positive drug effect. Such practices would constitute scientific misconduct. This example underscores, on multiple levels, the necessity of ensuring rigorous and reliable biomarker quantification. Inaccuracies not only risk patient misclassification but also undermine predictive modeling, treatment evaluation, and ultimately regulatory confidence. To evaluate how consistently this principle has been upheld in the field, we now proceed with a systematic review of the literature. 4 Methods The protocol for this systematic review was submitted to the PROSPERO registry and published on the PROSPERO website. The study protocol can be searched under the registration number CRD42024568766. The 2020 PRISMA guidelines for reporting systematic reviews are followed [59]. The 27-item PRISMA checklist is uploaded (see section 7.1). 4.1 Eligibility Criteria Inclusion Criteria: All studies involving the analysis of parallelism of biomarkers in body fluids were considered for inclusion. No exclusion criteria were applied based on participant demographics or specific disease conditions. Body fluids included cerebrospinal fluid (CSF), urine, saliva, and other relevant bodily fluids. Various analytical techniques were considered, including immunoassays, mass spectrometry, ELISA, and other relevant methods for quantitative biomarker detection [11–13,21]. Data was permitted to be derived from purely analytical studies, experimental studies or clinical settings. Exclusion Criteria: Studies not involving biomarker sampling from body fluids. Research focused solely on tissue biomarkers or biomarkers derived from non-fluid sources. Articles lacking detailed methodology or insufficient data on biomarker sampling techniques [7]. Case reports, reviews, and opinion articles without original research data. 4.2 Information Sources Two databases were searched, Medline and Google Scholar. 4.3 Search Strategy A search of the MEDLINE database was conducted covering the period after publication of lack of parallelism for NfH [18]. The dates entered to the search strategy were between 9-Dec2010 and 12-July-2024. There were no language restrictions. The Entrez Programming Utilities (E-utilities), provided by the National Center for Biotechnology Information (NCBI), were used in a Python script which can be downloaded from the PROSPERO register. Search Strategy Details: The search terms for the biomarkers (first search term) were: “neurofilament”, “neurofilaments”, “tau protein”, “T-tau”, “Ptau”, “glial fibrillary acidic protein”, “amyloid beta”, “ubiquitin C-terminal hydrolase 1”, “neurogranin”, and “YKL-40”. The search terms for the analytical methods (second search term) were: “method”, “development”, “linearity”, “parallelism”, and “doubling dilution”. The first search term and the second search term were combined individually. These combinations were coded in python. The complete Critical Reviews in Clinical Laboratory Sciences Parallelism Page 8 Figure 2: PRISMA flow diagram for the literature search, based on the PRISM template code [60]. literature search code can be downloaded from the study protocol on the PROSPERO Registry under item number 17. The literature search was performed on 02-SEP-2024. 4.4 Study Selection Neurodegeneration is a prevalent feature in various human diseases, and biomarkers play a crucial role in its indirect assessment. Therefore, no specific disease-related restrictions were imposed. Limitations were focused on the analytical development of biomarker tests [7]. Studies were selected based on the availability of quantitative body fluid samples. Selection Process: All studies identified underwent a full-text review and review of supplementary data where available. Tools Used: A spreadsheet was kept using the PubMed Identifier (PMID) for each study. Studies included and excluded after review were clearly marked as such, including the reason for that decision. 4.5 Data Collection Process and Data items Data Extraction: The data were extracted by hand from the values provided for samples at each dilution step into a spreadsheet. If data were only available in a figure, the corresponding author was contacted by email, containing the PROSPERO study number, asking to share the raw data. Author contact: Two email requests for sharing the data items required were sent to all corresponding authors. The first email was sent on 12-SEP-2024. The second email was sent to the non-responders on 12-OCT-2024. Data Items: The three specific variables for which data were collected included: 1. dilution step, 2. measured biomarker concentration, 3. sample type (standard, body fluid, artificial/spiked). 4.6 Risk of Bias in Individual Studies The risk of bias in individual studies was evaluated by carefully reviewing the methodologies used in sample selection and preparation. Specifically, we assessed whether the samples were native or spiked with protein standards and whether both the samples and standards were diluted using the same buffer, ensuring that it was indeed the correct dilution buffer [7]. Additionally, we recorded the selection of appropriate disease and control samples, noted whether the laboratory analyst was blinded during the experiments, and considered the reproducibility of the experiments through repeat assessments. 4.7 Summary Measures The summary measures used in this review include the deviation from the line of unity, which is represented by the horizontal line at a y-value of one in the partial parallelism plots [23]. The data are categorical, indicating the presence or absence of partial parallelism (yes/no). For studies where parallelism is observed, continuous data were presented, specifically the dilution range within which partial parallelism was demonstrated. These two effect measures will be analysed for each sample individually and, when meaningful (n>5), also as a group mean. 4.8 Synthesis of Results The synthesis of results is primarily visual, utilising partial parallelism plots [23] derived from the included studies. These visual representations are then summarised in table format and complemented by a narrative synthesis. 4.9 Risk of Bias Across Studies The primary focus is on determining whether partial parallelism exists between a sample from an individual with a specific disease and the protein standard used in the test. Practically this requires to provide data for a dilution series of the sample into the assay buffer (the same dilution buffer that is used for the standard curve) [7]. Therefore, a high risk of bias was assigned if no such sample was available and instead, an artificial sample was created by spiking the body fluid collected with the protein. Similarly, if different buffers were used for dilution between the sample and the standard, the bias was also rated as high. A moderate risk of bias Critical Reviews in Clinical Laboratory Sciences Parallelism Page 9 was assigned when an appropriate disease sample was not used, or if the laboratory analyst was not blinded, or if data on reproducibility were missing. In all other cases, the risk of bias was rated as low. 5 Results 5.1 Study Selection The flow diagram in Figure 2 outlines the process of study selection, detailing the number of records identified, included, and excluded [60]. A comprehensive literature search initially yielded 39 articles for further evaluation [16,35,61–97]. Additionally, a search on Google Scholar, using the same terms, identified five more references [5, 98–101]. All 44 articles [5,16,35,61–101], including supplementary analyses where available, were thoroughly reviewed. The corresponding author was contacted twice per email for sharing data if this was necessary for creating partial parallelism plots. For four articles, no presently valid author contact details could be found [69, 70, 73, 75]. For the remaining papers, a careful examination of all data available resulted in the exclusion of 25 articles [16,70,73,75,76,78–97] due to the absence of data necessary for generating partial parallelism plots as per our predefined protocol. Consequently, 19 studies were selected for further evaluation [5,35,61–69,71,72,74,77,98–101]. 5.2 Study Characteristics The characteristics of the included studies are summarised in Table 2. The majority of studies focused on human samples, primarily using plasma or CSF, with only a few investigating serum. Four studies examined brain tissue homogenates, while two used rodent tissue samples. All studies included control samples, and most also incorporated samples from a variety of disease conditions, with Alzheimer’s disease (AD) being the most frequently studied. The biomarkers analyzed across the studies included the Amyloid βfragments, Aβ1−40, Aβ1−42, Aβoligomers, α-synuclein, DJ-1, total tau protein (Tau), phospho tau proteins (pTau181, pTau217, pTau231), Apolipoprotein E (ApoE) isoforms E2, E3, E4, neurofilament light chain (NfL), neurofilament heavy chain (NfH), GFAP, and poly(GP). 5.3 Risk of Bias Within Studies The risk of bias across the included studies is summarised in Table 3. Only one study was assessed as having a low risk of bias [61]. This study provided clear documentation on key factors, including blinding and sample handling. Eight studies were rated as having a moderate risk of bias [5,35,63,64,66,74,98,100], mainly due to incomplete information regarding blinding procedures, particularly in relation to the blinding of the analyst. Another, eight studies were categorised as having a high risk of bias [62, 65, 67, 68, 71, 72, 77, 99]. This rating was due either to the use of spiked samples [62,65,67,68,72,77,99], or to incomplete documentation on whether the samples were spiked and how they were processed [71]. 5.4 Results of Individual Studies Detailed results for each included study are summarised in Table 4. 5.4.1 Amyloid β The amyloid cascade hypothesis, proposed by Hardy and Higgins [102], has positioned amyloid β(Aβ) as a central pathological driver in AD, following proteolytic cleavage of the amyloid precursor protein (APP). Quantification of Aβpeptides, particularly Aβ1−42 and Aβ1−40, and their ratios has since become integral to subsequent revisions of diagnostic criteria for AD [103–105]. However, aggregation-prone properties of Aβ[106,107] complicate immunoassay quantification, as epitope masking and altered conformations impair antibody recognition and disrupt dilutional parallelism [18]. On review of individual partial parallelism plots this emerges as a consistent and persistent analytical problem: •Aβ1−40: None of the assessed immunoassays achieved partial parallelism across three independent studies [65,69,98]. For each study deviations from expected parallelism was shown in the partial parallelism plots (see Figures 12, 14, 16). •Aβ1−42: Likewise, most studies failed to demonstrate partial parallelism [35, 65, 68, 69, 98]. A single study reported successful parallelism [72]; however, this study was flagged as high risk for bias (Table 3), as it involved spiking immunodepleted pooled plasma with synthetic Aβ, a method that may not replicate the conformational complexity of endogenous peptides (see Figures 6, 4, 12, 14, 16, 18). •Other A βspecies: No evidence of partial parallelism was found for truncated Aβpeptides [64], aggregated forms [62], or oligomeric Aβin EDTA plasma samples [67]. However, in the latter study, the use of citrate or heparin as anticoagulants enabled partial parallelism, highlighting the importance of pre-analytical variables for future immunoassay development (see Figure 21). Critical Reviews in Clinical Laboratory Sciences Parallelism Page 16 Figure 8: Partial Parallelism Plot for serial dilution of Amyloidβpeptides in plasma. This plot illustrates the partial parallelism for serial dilutions of various amyloid-βpeptides in plasma. Data were analyzed across complete dilution ranges (five data points) for comparison in the partial parallelism plots. Such data were available for 4 out of 18 (22%) plots presented in Figure 4 of reference [64]. The peptides included Aβ6−40, Aβ5−40, Aβ1−38, Aβ1−40. The plots demonstrate a lack of partial parallelism across the tested dilution ranges. Critical Reviews in Clinical Laboratory Sciences Parallelism Page 17 Figure 9: Partial parallelism plot for NfH quantification across different sample types. Sample A represents buffer spiked with NfH, while Samples B–D are derived from native CSF [66]. The raw data were provided by the corresponding author. The plot shows increasing NfH concentrations with increasing dilution steps, while the standard curve demonstrates partial parallelism only within the 1:1 to 1:16 dilution range. Overall, the results indicate a lack of partial parallelism across the full dilution range tested. tion curves across six laboratories employing one or more of seven assays [76]. Parallelism was reported as mean percentages per laboratory, with results ranging from 74% to 344%, indicating significant variability and inconsistency. Further complicating the assessment, some studies failed to provide key methodological details, such as sample preparation protocols or calibration standards, which may have introduced additional biases to what was reviewed in Table 3. Additionally, variations in experimental design, such as the use of non-standard dilution matrices or unverified spiking procedures, may have further reduced the comparability and reproducibility of parallelism testing. This underscores the need for stricter adherence to standardised guidelines and more transparent reporting of methodological details to minimise bias and improve the reliability of partial parallelism assessments. Our analysis revealed substantial variability in parallelism assessment. The implications of these findings are explored in the discussion section. 6 Discussion This systematic review provides a comprehensive analysis of parallelism testing in neurodegeneration biomarker assays. The principal finding is that all current biomarker tests exhibit only partial [23], rather than full, parallelism. Partial parallelism plots, which can be readily generated from existing data, offer a straightforward visual tool for comparison. However, the range of partial paralFigure 10: Partial Parallelism Plot for Neurofilament Light Chain (NfL) quantification. The plot illustrates the quantification of NfL in serum samples spiked with 500 ng/mL of the peptide standard [5]. The graph to the left displays the individual dilution curves, while the graph to the right shows the averaged data from 10 spiked serum samples. The raw data were provided by the corresponding author. The graph demonstrates that parallelism is demonstrated, on a group level, for a dilution range of 1:2–1:8 for recombinant NfL in serum using the assay buffer. Critical Reviews in Clinical Laboratory Sciences Parallelism Page 18 Figure 11: Partial parallelism plots for serial dilution of CSF measuring pTau217 using PT3xHT43 and PT3xPT82 assays. Data were extracted from Figure 2A in reference [74]. The plots do not demonstrate partial parallelism; however, this may be influenced by data noise. Averaged data from additional samples tested within the dilution range of 1:8 to 1:64 could provide further clarity. Figure 12: Partial Parallelism plot for Aβ1−42 and Aβ1−40 quantified from plasma samples spiked with the respective Amyloidβpeptides. The data were taken from supplementary Table 1 in reference [98]. The plot indicates a lack of partial parallelism across the dilution range (native sample to 1:4) tested. lelism is typically narrow, spanning approximately three doubling dilution steps. This finding has critical implications for interpreting studies that dilute samples beyond this range, as such practices can lead to inaccuracies. Biomarker concentrations may be overestimated when partial parallelism plots deviate upwards (e.g., Figure 12) or underestimated with downward deviations (e.g., Figure 15). The clinical context of such inaccuracies has been discussed for one example where neurofilaments were quantified at a dilution of 1:400 instead of 1:4. As seen in the present review, other biomarkers used in the context of neurodegeneration are affected as well. Notably however, near perfect partial parallelism [7] can sometimes be achieved on a group level (e.g., Figure 10). The distinction between group-level and individual-level partial parallelism (see Figure 10) highlights two key points. First, biomarker assays showing group-level partial parallelism may be suitable for clinical trials, especially as regulatory agencies like the FDA and EMA increasingly accept biomarker-based endpoints for rapid drug approvals in neurodegeneration. Second, the absence of parallelism in individual samples warrants further investigation. Tests measuring proteins involved in aggregation-related pathologies may harbor hidden biases, which could influence results and interpretations [138]. The effect of age will be one of the most obvious demographic factors to be investigated further [21,139]. Third, partial agreement between different assays for quantification of NfL implies presence of partial parallelism between the two tests [140]. However, the same method comparison study also showed overestimation of NfL levels by one assay compared to another at higher concentrations [140], indicating a lack of parallelism outside a narrow dilution range. This is entirely consistent with the clinical example, NfL in GBS, discussed earlier. A 20% failure rate in partial parallelism at the individual sample level indicates that one in five samples may not yield analytically comparable results. This degree of non-parallelism, if attributable to biological phenomena such as protein aggregation inherent to disease pathology [18,141], could introduce significant bias into quantitative interpretations. It is therefore essential that deviations from parallelism are not dismissed as technical artefacts but are instead rigorously interrogated. While matrix effects are frequently cited, they do not account for all scenarios and are predominantly a concern in mass spectrometry-based platforms [11–13]. Additional causes include endogenous protein–protein interactions, biochemical dissimilarity between native analytes and recombinant standards, analyte instability, suboptimal assay sensitivity, and supraphysiological biomarker concentrations exceeding Critical Reviews in Clinical Laboratory Sciences Parallelism Page 19 Figure 13: Partial Parallelism Plot for Aβoligomer quantification. This figure presents the partial parallelism plot for the quantification of Aβin CSF and PBS [62]. Experiments were conducted with increasing iterations: Sample A had zero iterations, Sample B had one iteration, progressing to Sample E with four iterations. The raw data were kindly shared by the corresponding author. The plots demonstrate a lack of partial parallelism across the dilution range tested. The full dilution range is displayed in the plots on the left. The zoomed-in plots on the right highlight that Aβoligomers are overestimated at dilutions beyond 1:10. Critical Reviews in Clinical Laboratory Sciences Parallelism Page 20 Figure 14: Partial Parallelism Plot for amyloid β1−42 and amyloid β1−40 quantified from artificial samples of recombinant human serum albumin which was spiked with full length recombinant amyloid β1−42 and amyloid β1−40 proteins at concentrations about three times above the highest standard. The data were taken from Table 6 in reference [65]. The plots do not demonstrate partial parallelism. Figure 15: Partial Parallelism plot for total Tau spiked into PBS. The data were taken from table 5 in reference [77]. There is lack of partial parallelism. Figure 16: Partial Parallelism plot for spiked CSF Aβ1−40 (78,300 ng/L) diluted in an assay buffer. The data were taken from the top left graph in Figure 2 in reference [69]. The plots do not demonstrate partial parallelism. Figure 17: Partial parallelism plots for total Tau and pTau (181P) assays. Data were derived from “representative results from one sample” in Supplementary Figure 1 of reference [71]. The CSF sample dilution series in buffer was estimated from the xaxis (1, 0.95, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.15, 0.10), and measured concentrations were approximated from the y-axis. Partial parallelism was achieved within a dilution range of 1:1.43 to 1:2.5 for the pTau181 assay but not for the total Tau assay. Figure 18: Partial Parallelism plot for Aβ1−42 from plasma diluted into a sample buffer. Data were taken from the log-scaled (0–100) y-axis from Figure 3A in reference [72]. The plot shows partial parallelism within a dilution range of 1:4–1:16. Critical Reviews in Clinical Laboratory Sciences Parallelism Page 21 Figure 19: Partial Parallelism plots for quantification of poly-glutamin expansions from native plasma samples. The data were taken from Figure 3H in reference [100]. The plot to the left shows the individual dilution curves. The graph to the right shows the average for 6 plasma samples. On a group level, the plot shows partial parallelism within a dilution range of 1:4–1:16. Figure 20: Partial Parallelism plot for quantification of phosphorylated (pTau181, pTau231) tau protein from samples spiked with the peptide standard. The data were taken from Supplementary Table 3 in reference [99]. Partial parallelism is not achieved in this plot. the assay’s dynamic range [11–13,18,141,142]. Therefore a critical finding in this systematic review is the potential publication bias, as study abstracts often overstate the presence of parallelism. Only 14% of the included studies demonstrated clear evidence of partial parallelism [5, 61, 71, 72, 100,101]. Interestingly, several studies reported on parallelism in between different samples instead of in between samples and the standard. This point could be emphasized stronger in future guidelines on harmonization of assay validation and implementation of quality control. The review encountered several limitations stemming from inconsistencies in the reporting of methodological details. Dilution ranges varied widely across studies, with arbitrary steps (e.g., 1:4, 1:5, 1:6 versus 1:1,250, 1:2,500, 1:5,000) [69]. While some dilution ranges could be inferred retrospectively [71], this was often based on a single representative sample, preventing the generation of parallelism plots for other data. Additionally, some reported values exceeded the assay measurement range [71], indicating inappropriate extrapolation beyond the standard curve. For the standard curve the dilution data was frequently not given. Instead, parallelism in between samples, rather than between samples and the standard curve was shown. Table 3 further highlights that not all samples were diluted in the assay buffer, a critical limitation when testing parallelism with a standard curve [7]. Overall the number of samples and standard curves was too small for meaningful statistical modelling [42, 43]. There is need for larger numbers [143] of samples, standards and datapoints (of the dilution ranges) for robust statistical evidence [143,144]. For example, one study that was able to demonstrate partial parallelism, did so for a very narrow dilution range of 1:1.43–1:2.5 based on one single sample that was considered to be representative [71]. The only study with a very large range of partial parallelism 1:100–1:10,000 reports this close to the detection limit (16 fM) of the test [67]. Moreover, there was concerning use of spiked or artificial samples in many studies [5,35,61,62, 65– 69, 72, 74, 77, 99, 100], instead of the recommended use of native samples [8, 9]. For one study it remained unclear if samples were spiked or not [71]. Taken together the use of spiked samples reduces the generalisability of the findings for creating representative parallelism plots. Similar issues with have been reported for other groups of biomarkers [145]. Stronger adherence to established guidelines is recommended [8]. Likewise, generalisability is hampered by the lack of sharing data on the standard curve for the dilution steps presented [5, 35, 61–65, 67–69, 71, 72, 74, Critical Reviews in Clinical Laboratory Sciences Parallelism Page 22 Figure 21: Partial Parallelism Plots for Aβoligomer quantification in different sampling buffers. The figure shows partial parallelism plots for Aβoligomers in EDTA, Citrate, and Heparin buffers. Samples A–E are spiked plasma, and Sample F is spiked PBS [67]. The raw data were kindly shared by the corresponding author. The plots reveal a lack of parallelism for EDTA, while partial parallelism is achieved in Citrate and Heparin buffers between dilutions of 1:100 and 1:10,000 on a group level. Critical Reviews in Clinical Laboratory Sciences Parallelism Page 23 77, 98–100]. On revision of the Figures in present systematic review it is not possible to state with absolute certainty if the standard curve overlays the (blue) line of unity. It would be desirable to have, in future studies, data from a representative number of stand curves available. These data points do not need to be exactly for the same dilution steps as for samples if the visual representation with partial parallelism plots is chosen. This is another advantage of the visual approach compared to statistical techniques. On critical review of the setup for sample dilution it should be mentioned that for example one of the finally excluded studies [16] diluted blood samples from patients with multiple sclerosis into blood samples of health controls. For assessment of parallelism between patient samples and the assay’s standard curve it is mandatory to use the the same buffer solution [7,20,21]. Conflicting parallelism results were found for the pTau181 assay. One study did show partial parallelism [71] (Figure 17B) and one did not [101] (Figures 5A-F). Clearly, on direct visual comparison of these partial parallelism plots it is evident that different dilution ranges were used. The short stretch of partial parallelism (1:1.43–1:1.25) from one study [71] is for a dilution range lower than the one shown for the other (1:4–1:32 and a:5– 1:40) [101]. Because a third study did also not demonstrate partial parallelism for pTau181 [99], a balanced interpretation would suggest absence of parallelism for this biomarker. Another limitation of this systematic review is the restriction to studies published between 2010 and 2024, a subset of the 32-years since the formulation of the amyloid cascade hypothesis (1992–2024) [102]. This time frame was chosen intentionally, as concerns regarding the effect of protein aggregation on dilutional parallelism in immunoassays were first raised in 2010 [18], and only subsequently acknowledged in influential white papers [12, 21, 146] and regulatory guidance documents [20,34,147]. Taken together, this critical and systematic review emphasises the importance of adhering rigorously to guidelines for parallelism testing [12,20,21, 34, 146]. These guidelines are embedded in a well defined framework for laboratory test validation that is endorsed by regulatory authorities (Table 1). If observed, the reason for lack of partial parallelism needs to be explained as it has been done in the context of protein aggregation [18]. Authors should explicitly report the dilution range over which partial parallelism is achieved and highlight potential biases, such as overor underestimation of biomarker concentrations, when dilutions exceed this range. For clinical trials, it is crucial to clearly document the range of partial parallelism, the dilutions performed, and strategies to minimise biases. Regulatory agencies might consider mandating the inclusion of data from random individual samples to demonstrate parallelism at both individual and group levels. Without clear evidence of partial parallelism, it would be inadvisable to quantify biomarkers for neurodegeneration at varying dilution steps or across different time points in clinical trials. In summary, our critical and systematic review highlights both progress and persistent challenges in parallelism testing. These insights inform recommendations for future assay validation and inclusion into the test validation framework of regulatory authorities (e.g. under precision and bias or under linearity in Table 1). 7 Conclusion This systematic review identified partial parallelism in only a small proportion of biomarker tests for neurodegeneration. A likely biological reason is the presence of protein aggregates, a key pathological feature in many neurodegenerative diseases. Where partial parallelism is absent, the data suggest specific dilution ranges where it could potentially be achieved. To advance the field, future research must align with established guidelines [20,147], ensuring transparent reporting that allows independent researchers and regulatory bodies to evaluate partial parallelism through standardised plots. References [1] Abdelhak A, Foschi M, Abu-Rumeileh S, et al. Blood GFAP as an emerging biomarker in brain and spinal cord disorders. Nature Reviews Neurology. 2022 Feb;18(3):158–172. [2] Hauser SL, Bar-Or A, Cohen JA, et al. Ofatumumab versus teriflunomide in multiple sclerosis. New England Journal of Medicine. 2020 Aug;383(6):546–557. [3] Miller TM, Cudkowicz ME, Genge A, et al. Trial of antisense oligonucleotide tofersen for SOD1 ALS. New England Journal of Medicine. 2022 Sep;387(12):1099–1110. [4] Khalil M, Teunissen CE, Lehmann S, et al. Neurofilaments as biomarkers in neurological disorders — towards clinical application. Nature Reviews Neurology. 2024 Apr;20(5):269– 287. [5] Lee S, Plavina T, Singh CM, et al. Development of a highly sensitive neurofilament light chain assay on an automated immunoassay platform. Frontiers in Neurology. 2022 Jul;13. Critical Reviews in Clinical Laboratory Sciences Parallelism Page 24 [6] Petzold A. The 2022 Lady Estelle Wolfson Lectureship on Neurofilaments. Journal of Neurochemistry. 2022 Aug;163(3):179–219. Available from: https://discovery. ucl.ac.uk/id/eprint/10146661/1/ LadyEstelleWolfsonLecture.mp4. [7] Pum J. A practical guide to validation and verification of analytical methods in the clinical laboratory. Advances in Clinical Chemistry, Vol 38. 2019;:215–281. [8] Andreasson U, Perret-Liaudet A, van Waalwijk van Doorn LJC, et al. A practical guide to immunoassay method validation. Frontiers in Neurology. 2015 Aug;6. [9] Burtis C, Ashwood E, editors. Tietz Textbook of Clinical Chemistry. 2nd ed. WB Saunders Company; 1994. [10] Zhang J, Dasgupta A, Ayala R, et al. Internal standard variability: root cause investigation, parallelism for evaluating trackability and practical considerations. Bioanalysis. 2024 May;:1–6. [11] Richards S, Amaravadi L, Pillutla R, et al. 2016 White Paper on recent issues in bioanalysis: focus on biomarker assay validation (BAV): (Part 3–LBA, biomarkers and immunogenicity). Bioanalysis. 2016;8(23):2475– 2496. [12] Piccoli S, Mehta D, Vitaliti A, et al. 2019 White paper on recent issues in bioanalysis: FDA immunogenicity guidance, gene therapy, critical reagents, biomarkers and flow cytometry validation (Part 3 — recommendations on 2019 FDA immunogenicity guidance, gene therapy bioanalytical challenges, strategies for critical reagent management, biomarker assay validation, flow cytometry validation & CLSI H62). Bioanalysis. 2019; 11(24):2207–2244. [13] Fernández-Metzler C, Ackermann B, Garofolo F, et al. Biomarker assay validation by mass spectrometry. The AAPS Journal. 2022; 24(3):66. [14] Petzold A, Altintas A, Andreoni L, et al. Neurofilament ELISA validation. Journal of Immunological Methods. 2010 Jan;352(1-2):23– 31. Available from: https://discovery. ucl.ac.uk/id/eprint/95615/. [15] Oeckl P, Jardel C, Salachas F, et al. Multicenter validation of CSF neurofilaments as diagnostic biomarkers for ALS. Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration. 2016 Apr;17(5-6):404–413. [16] Wilson D, Chan D, Chang L, et al. Development and multi-center validation of a fully automated digital immunoassay for neurofilament light chain: toward a clinical blood test for neuronal injury. Clinical Chemistry and Laboratory Medicine. 2024 Jan;62:322–331. [17] Booth RA, Beriault D, Schneider R, et al. Validation and generation of age-specific reference intervals for a new blood neurofilament light chain assay. Clinica Chimica Acta. 2025 Sep;577:120447. [18] Lu CH, Kalmar B, Malaspina A, et al. A method to solubilise protein aggregates for immunoassay quantification which overcomes the neurofilament hook effect. Journal of Neuroscience Methods. 2011 Feb;195(2):143–150. Available from: https://discovery.ucl. ac.uk/id/eprint/678078. [19] Musso G, Zoccarato M, Gallo N, et al. Hookeffect in maglumi immunoassay for serum anti-gad antibodies in neurological disorders: When “wrong” matrix is the right choice. Clinica Chimica Acta. 2024 May;558:119679. [20] Committee for Medicinal Products for Human Use. ICH guideline M10 on bioanalytical method validation and study sample analysis. European Medical Agency. 2023 Jan;M10. [21] Hersey S, Keller S, Mathews J, et al. 2021 white paper on recent issues in bioanalysis: ISR for biomarkers, liquid biopsies, spectral cytometry, inhalation/oral & multispecific biotherapeutics, accuracy/LLOQ for flow cytometry (part 2 — recommendations on biomarkers/CDx assays development & validation, cytometry validation & innovation, biotherapeutics PK LBA regulated bioanalysis, critical reagents & positive controls generation). Bioanalysis. 2022 May;:627–692. [22] Jucker M, Walker LC. Self-propagation of pathogenic protein aggregates in neurodegenerative diseases. Nature. 2013 Sep; 501(7465):45–51. [23] Petzold A. Partial parallelism plots. APPS Applied Sciences. 2024 Jan; 14(2):602. Available from: https: //discovery.ucl.ac.uk/id/eprint/ 10185313/1/parallel-v9-accepted.pdf. [24] Winger RM. Gauss and non-euclidean geometry. Bulletin of the American Mathematical Society. 1925;31(7):356–358. [25] Jenkovszky L, Lake MJ, Soloviev V. János Bolyai, Carl Friedrich Gauss, Nikolai Critical Reviews in Clinical Laboratory Sciences Parallelism Page 25 Lobachevsky and the New Geometry: Foreword. Symmetry. 2023 Mar;15(3):707. [26] Plikaytis BD, Holder PF, Pais LB, et al. Determination of parallelism and nonparallelism in bioassay dilution curves. Journal of Clinical Microbiology. 1994 Oct;32(10):2441–2447. [27] Meijering E. A chronology of interpolation: from ancient astronomy to modern signal and image processing. Proceedings of the IEEE. 2002 Mar;90(3):319–342. [28] Rodbard D, Hutt D. Statistical analysis of radioimmunoassays and immunoradiometric (labelled antibody) assays. A generalized weighted, iterative, least-squares method for logistic curve fitting. NSA-31-000958. 1974; :165–192OSTI ID:4277004. [29] Rodbard D, Munson P, De Lean A. Improved curve-fitting, parallelism testing, characterization of sensitivity and specificity, validation, and optimization for radioligand assays. Radioimmunoassay and related procedures in medicine; 1978. [30] Yang H. Emerging non-clinical biostatistics in biopharmaceutical development and manufacturing. Chapman and Hall/CRC; 2016. [31] Prentice RL. A generalization of the probit and logit methods for dose response curves. Biometrics. 1976 Dec;32(4):761. [32] Dudley RA, Edwards P, Ekins RP, et al. Guidelines for immunoassay data processing. Clinical Chemistry. 1985 Aug;31(8):1264– 1271. [33] Gottschalk PG, Dunn JR. Determining the error of dose estimates and minimum and maximum acceptable concentrations from assays with nonlinear dose-response curves. Computer Methods and Programs in Biomedicine. 2005 Dec;80(3):204–215. [34] Food and Drug Administration. M10 bioanalytical method validation and study sample analysis. 2022 ; 2023. [35] Cullen VC, Fredenburg RA, Evans C, et al. Development and advanced validation of an optimized method for the quantitation of aβ42 in human cerebrospinal fluid. AAPS Journal. 2012 Sep;14:510–8. [36] Fu Y, Barkley D, Li W, et al. Evaluation, identification and impact assessment of abnormal internal standard response variability in regulated lcms bioanalysis. Bioanalysis. 2020 Apr;12(8):545–559. [37] J Woolf E. Learning how to interpret ‘dangerous’ internal standard behaviors. Bioanalysis. 2019 Sep;11(18):1679–1684. [38] Lawson NS, Haven GT, Williams GW. Analyte stability in clinical chemistry quality control materials. CRC Critical Reviews in Clinical Laboratory Sciences. 1982 Jan;17(1):1–50. [39] Roberts CJ. Protein aggregation and its impact on product quality. Current Opinion in Biotechnology. 2014 Dec;30:211–217. [40] Tu J, Bennett P. Parallelism experiments to evaluate matrix effects, selectivity and sensitivity in ligand-binding assay method development: pros and cons. Bioanalysis. 2017 Jul; 9(14):1107–1122. [41] Loh TP, Lim CY, Sethi SK, et al. Advances in internal quality control. Critical Reviews in Clinical Laboratory Sciences. 2023 May; 60(7):502–517. [42] Faya P, Rauk AP, Griffiths KL, et al. A curve similarity approach to parallelism testing in bioassay. Journal of Biopharmaceutical Statistics. 2020 Mar;30(4):721–733. [43] Gottschalk PG, Dunn JR. Measuring parallelism, linearity, and relative potency in bioassay and immunoassay data. Journal of Biopharmaceutical Statistics. 2005 May; 15(3):437–463. [44] Guy S, Shepherd MF, Bowyer AE, et al. How to assess parallelism in factor assays: coefficient of variation of results with different dilutions or slope ratio? International Journal of Laboratory Hematology. 2023 Dec;45(2):229– 240. [45] Jonkman JN, Sidik K. Equivalence testing for parallelism in the four-parameter logistic model. Journal of Biopharmaceutical Statistics. 2009 Aug;19(5):818–837. [46] Smith WC, Sittampalam GS. Conceptual and statistical issues in the validation of analytic dilution assays for pharmaceutical applications. Journal of Biopharmaceutical Statistics. 1998 Jan;8(4):509–532. [47] Giltinan D. Selected topics in parallel line bioassay: Potency estimation by parallel line biossay techniques. Developments in biologicals. 2002;107:47–56. [48] Klein J, Capen R, Mancinelli R, et al. Validation of assays for use with combination vaccines. Biologicals. 1999 Mar;27(1):35–41. Critical Reviews in Clinical Laboratory Sciences Parallelism Page 32 Supplementary information Acknowledgments The sharing of the raw data needed to construct the partial parallelism plots is acknowledged from the following authors: Oliver Bannach, Dieter Willbold and Marleen J.A. KoelSimmelink. Author contributions A.P. designed and performed all experiments, analysed the data, prepared the figures and wrote the manuscript. D.C. reviewed and co-registered the protocol on PROSPERO and reviewed the manuscript draft. J.P. reviewed the manuscript draft. Disclosure of interest The authors report there are no competing interests to declare. This study was not funded. Ethical approval declarations This is a systematic review on existing data from studies which had ethical permission. Data availability & Data deposition All data have been uploaded to Figshare [150]. The python code for the search strategy is openly available for download from the PRESTO registry. Critical Reviews in Clinical Laboratory Sciences Supplementary Materials to: “Parallelism in Neurodegenerative Biomarker Tests. . . ” Page 33 Statistical methods for determination of parallelism The assessment of parallelism originates in the statistical literature and is typically framed as a binary decision: parallelism is either present or absent [43]. This decision is tested through statistical hypotheses embedded within formal models. Historically, the first hypothesis documented is Euclid’s fifth postulate. It is a hypothesis that can be tested, and as we learned from history, it can take centuries to discover that presumed mathematical solutions eventually turned out to be wrong. The contemporary statistical approaches to testing parallelism all have in common that they build on the probability theory [144]. With that comes the law of large numbers [143]. Only with large numbers there is some guarantee that the averages from random events provide somehow stable long-term results [143,144]. That implies that statistical methods of testing for parallelism on small numbers, are open to criticism. One established approach of statistical testing for parallelism is the extra sum-of-squares analysis of variance (ANOVA) method, which compares residual sum of squares between nested models (RSSEnonpar) [148, 149]. For example this approach forms the basis for using F and χ2statistics to test for parallelism [43]. The authors compare visual and statistical testing, showcasting the strengths and weaknesses of these methods. The key message, for presence of nonparallelism, is to decompose RSSEconst into the component of nonparallel origine or RSSEnonpar and what can be described by random variation RSSEfree. Simplified, RSSEconst =RSSEnonpar+RSSEfree. Here the definition of nonparallelism is the extra error that comes because of lack of similarity between two curves as part of RSSEnonpar. The practical calculation of RSSEnonpar depends, and this is very important to realise, on the assumption of normally distributed data and presence of parallelism. With this the law of large numbers applies because the χ2 test is used, which works on a distributed randome variable: dfconst =Nstd +Nuk −(P+ 1). To statistically test for this one needs to calculate the probabilities for the dose (x) for two models. 1. The free model (SSEfree) is described as: SSEfree(pstd,puk) = Nstd X i=1 wstd iystd i−f(xstd i;pstd)2 + Nuk X i=1 wuk iyuk i−f(xuk i;puk)2 2. The constraint model (SSEconst) is described as: SSEconst(r, p) = Nstd X i=1 wstd iystd i−f(xstd i;p)2 + Nuk X i=1 wuk iyuk i−f(rxuk i;p)2 Very elegantly the authors elaborate, citing comprehensive reviews, that highlight the limitations of various statistical factors in this context [43]. The application of statistical models for parallelism assessment is not without limitations. As noted by Gottschalk et al.,“The existence of similarity between two mathematical functions is not difficult to determine. It is less straightforward to determine the degree of parallelism between two functions that are not exactly similar.” [43]. Such demonstration of lack of similarity has been proposed to be solved by employing Bayesian or frequentist approaches [42]. With Bayesian posterior probability one can test parallel equivalence through: p(γ, xL, xU) = Pr min ρmax x∈[xL,xU]|f(θ1, x)−f(θ2, x +ρ)|< γ  data The auhors give pratical examples, based on simalated data, that deliver a p-value, using this probabilistic approach. The simulated data give biomarker concentration ranging from 0.02– 125,000 arbitrary units (Table 1 in reference [42]). In this example the standard error (SE) equals r1 η2Var hAˆ θi. Therefore, with α= 0.05 and δ= 0.85, the probability caluclates as Pr T2n−p>λ(ˆ θ)−δ SE = 0.062. This is not significant. Consequently, similarity between curves cannot be declared. This implies that there is no evidence for parallelism in bespoke example (for visual comparison see Figure 1 in reference [42]. This is an extreme example to demonstrate lack of parallelism, because the two curves cross over between log2and log4). Critical Reviews in Clinical Laboratory Sciences Supplementary Materials to: “Parallelism in Neurodegenerative Biomarker Tests. . . ” Page 34 As introduced above, the 4PL and 5PL standard curves are now frequently used for fitting doseresponse curves and consequently of relevance for discussed statistical evaluations of parallelism between sample and standard curves [42–45]. One final word of caution is warranted here, when employing highly parameterized non-linear curve models, which can lead to overfitting, there is a risk to introduce bias into the parallelism metrics that provide the data fed inot above described statistical models. As Smith noted, “The condition [of parallelism] and its importance are relatively unknown to bioanalytical chemists and many consulting statisticians” [46]. In that work, the authors illustrated interpretative challenges using simulated data, showing how deviations in only a portion of the curve can complicate decision-making. For this reason, our critical review primarily relied on visual methods to assess the presence or absence of parallelism [23,26,47–49]. Visual inspection offers an intuitive and transparent way to identify deviations, making the concept accessible to a broader scientific audience, including laboratory practitioners who may not have advanced statistical expertise. Importantly, we do not consider visual and statistical methods to be mutually exclusive. Rather, they are complementary: visual assessment provides a practical first-line evaluation, while statistical analyses can serve as confirmatory tools, or address specific questions in selected situations. Critical Reviews in Clinical Laboratory Sciences