Full text
Zenodo doi:10.5281/zenodo.17430468 Personalized prognosis & treatment with Bayesian nonparametric inference An example study on conversion from Mild Cognitive Impairment to Alzheimer’s Disease P.G.L. Porta Mana Western Norway University of Applied Sciences <pgl portamana.org> I. Rye University of Oslo A. Vik Haukeland University Hospital, Bergen M. Kociński University of Bergen A. Lundervold Mohn Medical Imaging and Visualization Centre (MMIV), Department of Radiology, Haukeland University Hospital, Bergen University of Bergen A. J. Lundervold Department of Biological and Medical Psychology, University of Bergen A. S. Lundervold Mohn Medical Imaging and Visualization Centre (MMIV), Department of Radiology, Haukeland University Hospital, Bergen Western Norway University of Applied Sciences 18 November 2022; updated 23 November 2025 [draft] The present work presents a statistically sound, rigorous, and model-free inference method for use in personalized medicine, together with a software implementation. The method, Bayesian nonparametric inference, is designed first to learn from a set of clinical data with relevant predictors and predictands, and then to assist a clinician in the assessment of prognosis & treatment for new patients. It allows the clinician to input, for each new patient, additional patient-dependent clinical information, as well as patient-dependent information about benefits and drawbacks of available treatments. We apply this method and software in a realistic setting for clinical decision-making, incorporating clinical, environmental, imaging, and genetic data, using a data set of subjects suffering from mild cognitive impairment and Alzheimer’s Disease. We show how the method is theoretically optimal, and discuss some of its major advantages for decision-making under risk, resource planning, imputation of missing values, assessing the prognostic importance of predictors, and further uses. Keywords: Clinical decision making, Utility theory, Probability theory, Bayesian nonparametrics, Machine Learning, Artificial Intelligence, Base-rate fallacy 1
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics 1 Introduction: Personalized prognosis, treatment, statistics, and assistive software 1.0 Prologue: Four unique patients Meet Olivia, Ariel, Bianca, Curtis. 1 These four persons don’t know each other, but they have something in common: they all suffer from a mild form of cognitive impairment, and are afraid that their impairment will turn into Alzheimer’s Disease within a couple of years. This is why each of them recently underwent a wide range of clinical examinations and tests, including brain imaging. Today they are receiving the results. Based on their individual results, on available clinical statistical data, and on other relevant information, their clinician will assess their risk of developing Alzheimer’s Disease. Then, together with the patients and their relatives, the clinician will make a decision among four distinct preventive-treatment options, available to each patient. 2 In these tasks, the clinician will be helped by an assistive inference method and software. Besides a shared diagnosis of Mild Cognitive Impairment and associated worries, these patients have other things in common – but also some differences. Let’s take Olivia as reference, and list the similarities and differences between her and the other three patients: • Olivia and Ariel have identical results on the clinical and laboratory measures and age. They would also incur similar benefits and losses from the four available treatment options. Ariel, however, comes from a different geographical region, which presents a higher rate of conversion from Mild Cognitive Impairment to Alzheimer’s Disease. And unlike Olivia, Ariel comes from a family with a heavy history of Alzheimer’s Disease. Because of this geographical and family background and some relevant statistics found in some publications, the clinician judges, before seeing the clinical data, that there’s a 65% probability that Ariel’s cognitive impairment will convert to Alzheimer’s Disease. • Olivia and Bianca have identical clinical results and age; they also come from the same geographical region and have very similar 1 These are purely fictive characters but with clinically realistic conditions; any reference to real persons is purely coincidental. 2 In the present paper we use “prognosis” in a general sense to include also “diagnosis”, and “treatment” quite loosely to mean any course of action a clinician might take, including preventive treatment or even “additional tests”. 2
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics family histories. In fact, we shall see that they have the same probability of developing Alzheimer’s Disease. Bianca, however, suffers from several allergies and additional clinical conditions that render some of the treatment options slightly riskier for her. • Olivia and Curtis have different results on all measures included in the clinical and laboratory examinations; Olivia is also more than 10 years older than Curtis. They otherwise come from the same geographical region, have very similar family histories, and would incur similar benefits or losses from the treatment options. Note that the imaging result for Curtis (hippocampal volume) is missing. Considering the similarities and differences among these patients, which of the four available treatments will be optimal for each of them? The clinician will find that, despite the many factors in common among our four patients – even despite Olivia’s, Ariel’s, and Bianca’s identical clinical results, and Olivia’s and Bianca’s identical probability of conversion to Alzheimer’s Disease – the optimal treatment for each patient is different from those for the other three – how come? 1.1 Assistive inference software: personalized input and output In the example above, we said “in these tasks, the clinician will be helped by an assistive inference method and software”. The need for such computational help is clear from the vast amount of clinical statistical data and the large number of clinical predictors today available to clinicians. But how should such assistive method and software be designed in order to take fully into account patient differences? Although the example above concerns specifically Alzheimer’s Disease, the differences among patients described there apply more generally to most, if not all, clinical problems of prognosis and treatment. These differences can be broadly categorized as “difference in auxiliary or supplementary tests and background information” (Olivia vs Ariel), “difference in benefit and availability of treatments” (Olivia vs Bianca), “difference in clinical predictors” (Olivia vs Curtis), as schematized in the side figure. Each of these difference categories can affect the clinician’s final choice of 3
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics optimal treatment. An assistive method and software should therefore reflect these differences in its input, its output, or both: • In principle, there could be three kinds of input “slots”, where the clinician can input the current patient’s specific values as regards clinical predictors, auxiliary information, and treatment options & benefits. • If input slots are only available for one or two of the categories above, the output should at least be of such a kind as to allow the clinician to integrate the current patient’s specific values of the missing input categories. To appreciate these requirements, one should contrast the input and output of many kinds of machine-learning classification algorithms. These typically only allow the input of a patient’s clinical predictors, with no space for patient-specific auxiliary information or for adjustments of differences in background statistics (think of Olivia vs Ariel). And they typically output only a discrete prognostic label (say, “stable Mild Cognitive Impairment” vs “conversion to Alzheimer’s Disease”), but no measure of the uncertainty about that label. Unfortunately, such output does not allow the clinician to assess treatment benefits and losses for the current patient, for this assessment depends not on the presence (present or future) of a disease, but on the risk of its presence. We shall discuss these points at length in §§ 3.2 and 3.3. The purpose of the present work is to present an assistive method, Bayesian nonparametric inference 3 , together with a software implementation, which meet the requirements above. This method is designed to first learn from a dataset of clinical data with relevant predictors and predictand 4 , and then assist a clinician in the assessment of prognosis & treatment for new patients. It offers these ten features: 1. It can work with clinical predictors comprising any combination of categorical and one-dimensional (continuous, discrete ordinal, unbounded or bounded, uncensored or censored) variates. The 3 More precisely Bayesian nonparametric density inference; see e.g. Rodríguez et al. 2009; Bhattacharya & Dunson 2010; and Walker’s 2010 witty overview. 4 literally “quantity to be predicted” or, more generally, inferred (cf. measurand in metrology, jcgm 2012 2.3). We find this term, used in meteorology and climate science, more precise and less obscure or misleading than “dependent variate”, “response variate”, “outcome variable”, or similar. 4
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics predictand can also be any combination of categorical and onedimensional variates. 2. It treats predictor and predictand variates on equal footing, in the sense that the clinician can at any moment decide to infer some other variate given the rest. 3. It does not require that the current patient be considered in all respects as a member of the population underlying the learning dataset. The patient can be considered a member only conditionally on particular variate values. 4. It accepts three inputs: (i1) the clinical-predictor values for the current patient; (i2) information about which predictand-predictor relationships learned from the dataset can be generalized to the current patient, and a prior prognostic probability representing auxiliary information; (i3) a set of treatment options and their benefits and losses for the current patient. 5. It yields three basic outputs: (o1) any prognostic probabilities or likelihoods about predictors and predictand desired by the clinician, given input (i1); (o2) final prognostic probabilities, given inputs (i1)–(i2); (o3) optimal treatment, given inputs (i1)–(i3); 6. Its input and outputs are modular, in the sense that the clinician can, for instance, give inputs (i1)–(i2) only, get a prognostic probability (o2) as output, and then proceed to treatment assessment by other means or methods. 7. It works even if predictor data are missing, both in the learning dataset and for the current patient. 8. It can quantify the uncertainty of its own outputs, allowing for sensitivity analyses. For example, it can tell how much a prognostic probability could have been different if the learning dataset had been larger, or whether the optimal treatment could be different if a particular missing predictor for the current patient were available. 5
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics 9. It can make various kinds of long-term forecasts, such as frequency of prognoses with given probabilities, frequency of prescribed treatments, and similar – provided that the dataset used for its learning can be considered representative of the full population. 10. It is model-free and extracts the maximal amount of information theoretically contained in the learning dataset, and therefore achieves the maximal prognostic power that the predictors can yield. In other words, it is unbeatable. Let us comment on some of these features. We believe that the capability of working with complex predictands, feature 1., is important for a more realistic and nuanced approach to prognosis. In the case of Alzheimer’s Disease, for instance, a simple dichotomy “has disease” vs. “doesn’t have disease” is possibly an oversimplification 5 . Without feature 3., the capability of auxiliary contextual information, the method would be of no use in the often occurring case of patients having peculiar clinical contexts. The capability of dealing with missing data, feature 7., is important for a concrete implementation in a clinical setting, typically afflicted by imputation problems. Feature 8. is extremely important for a clinician to assess the reliability of final decisions and honestly inform the patient of the possibility of unwanted outcomes. Finally, features 2. and 10., the fact that this method yields the maximal amount of information jointly contained in all variates, makes it valuable in general clinical research. The method can, for example, forecast the maximal accuracy obtainable by any inference algorithm based on the same predictors or a subset of those predictors; and it attains, by construction, that maximal accuracy. Further features of interest in machine learning are discussed in the next section. The software for implementing Bayesian nonparametric inference is at the moment available as a clinician-friendly proto-package 6 in the R programming language7. The methodology underlying Bayesian nonparametrics has been successfully demonstrated for Alzheimer’s Disease with a smaller number of predictors 8 , is used in many applications in astrophysics 9 , and its 5 see e.g. Edmonds et al. 2015;2020, whose methods we find, however, inconclusive. 6https://pglpm.github.io/inferno ; the scripts used in the present work are available at doi: 10.17605/osf.io/zb26t . 7 R Core Team 2023. 8 Antoniano-Villalobos et al. 2014.9Event Horizon Telescope Collaboration 2019;2022; Del Pozzo et al. 2018. 6
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics advantages in neurocritical care and medicine have been emphasized for quite some time10. The next section 2gives an intuitive understanding of the method’s underlying principles and workings. Concrete application of the method is shown in § 3, using the four-patient fictitious scenario of § 1.0 as a concrete example, and subsection 3.4 discusses further applications to general medical research. A summary and discussion is given in § 4. Mathematical details and proofs on which the present work is grounded are given in a companion technical note11. We apologize to readers who may find some discussions or explanations too obvious, or some mathematical details too scarce. We wanted the present work to be accessible to a wide audience, from clinicians and students of medicine to researchers in machine learning and probability theory. 2 Bayesian nonparametric inference This section can be especially of interest to readers from Machine Learning and Artificial Intelligence. It is largely independent of the next one, which describes the method’s application. It can be read after § 3by readers who would like to see the method in action first. 2.1 Underlying theory and characteristics The method to solve clinical decision-making problems such as the one of § 1is none other than Decision Theory: the combination of probability theory and utility theory. It integrates available clinical statistical data with each patient’s unique combination of clinical results, auxiliary information, and treatment benefits, in a mathematical framework, completely determined by basic self-consistency requirements.12 Medicine has the distinction of having been one of the first fields to adopt Decision Theory, with the pioneering work by Ledley – who, 10 Jawa & Maslove 2023; Temp et al. 2021; Antoniano-Villalobos et al. 2014; Sidebotham 2020; Goodman 1999. 11 https://github.com/pglpm/inferno/raw/main/de velopment/manual/optimal_predictor_machine.pdf 12 Jaynes 2003 chs 13–14; von Neumann & Morgenstern 1955; Cox 1946; Savage 1972; Luce & Raiffa 1957; Raiffa & Schlaifer 2000; Raiffa 1970; Lindley 1988; Kreps 1988. 7
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics incidentally, died of Alzheimer’s Disease 13 – and Lusted 14 , who also promoted its algorithmic implementation 15 . Clinical decision-making is today explained and exemplified in brilliant textbooks for medical students and clinicians16. An outline is given in § 3.3. Software for Bayesian nonparametrics provides an algorithmic implementation, as dreamed by Lusted and Ledley 17 , of the main calculations underlying the clinical decision-making process: from the comparison of a patient’s specific predictors with the statistics offered by a clinical database, to the choice of optimal treatment.18 Decision theory is also the normative foundation for the construction of an Artificial Intelligence agent capable of rational inference and decision making 19 . Bayesian nonparametrics can therefore be seen as an ideal machine-learning algorithm. It is “ideal” in the sense of being free from special modelling assumptions (this is why we do not call it a “model”) and from limitations of informational output which affect most common machine-learning algorithms; not “ideal” in the sense of being impracticable. Quite the opposite, the present work shows that this ideal machine-learning algorithm can today be used in a wide range of inference problems at insubstantial computational cost. More concretely, Bayesian nonparametrics is ideal because it computes the probability distribution over all possible long-run frequency distributions from which the learning dataset can originate, these frequency distributions being joint ones for all predictor and predictand variates. This is the maximum possible amount of information that can be extracted from the learning dataset, in a strict information-theoretic sense. From this probability distribution, Bayesian nonparametrics can indeed calculate any quantity outputted by other machine-learning algorithms. For example 20 : • “Discriminative” algorithms: the probability p(𝑌|𝑋) of any set of predictands 𝑌given any set of input predictors 𝑋. • “Generative” algorithms: the probability p(𝑋|𝑌) of any set of input predictors 𝑋given any set of predictand values 𝑌. 13 Shah et al. 2013. 14 Ledley & Lusted 1959a,b;1960; Lusted & Ledley 1960; Lusted 1967. 15 Lusted 1968; Ledley 1959;1960 § 1-5 p. 21. 16 Weinstein & Fineberg 1980; Sox et al. 2013; Hunink et al. 2014. 17 cf. the Appendices in Lusted 1968. 18 In previous drafts we called this software implementation a “Lusted-Jaynes machine” as a homage to Lusted and to Jaynes 2003, who brilliantly explained the inductive logic underlying such a “robot”. 19 Russell & Norvig 2022 part IV; Jaynes 2003 chs 1–2, 13–14. 20 for terminology see e.g. Murphy 2012 § 8.6. 8
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics More generally, Bayesian nonparametrics can compute any joint, marginal, or conditional probabilities 𝑝(𝑍′, 𝑍′′) , 𝑝(𝑍′) , 𝑝(𝑍′|𝑍′′) for any desired subsets of variates 𝑍′, 𝑍′′. • Regression or classification: the expected value E(𝑌|𝑋) of any set of variates 𝑌 , given any other set of variates 𝑋 , including the particular case of 𝑌 predictand, and 𝑋 predictors. The uncertainty or variability around such an average is also automatically computed. • Functional regression: if the predictand 𝑌 or any other variate of interest turns out to be a function 𝑓 of variates 𝑋 , then their conditional probability will be a delta distribution: p(𝑌|𝑋)=δ[𝑌− 𝑓(𝑋)] . Thus Bayesian nonparametrics always recovers a functional relationship if there is one, as well as its noise distribution. Furthermore, this method also quantifies the uncertainty of all outputs above. More precisely, it takes into account how the statistical properties of the learning dataset could be different from those of its original population, owing to sampling fluctuations; and it can compute how much any of the outputs above would probably change if more learning data were collected. In the next section we explain intuitively how Bayesian nonparametrics computes the general probability distribution over long-run frequencies. A couple of special characteristics brought about by such computation can already be summarized here. First, in contrast to machine-learning algorithms such as neural networks, random forests, Gaussian processes, support-vector machines, or generalized linear models, Bayesian nonparametrics does not assume the existence of a function (possibly contaminated by a little noise) from predictors to predictands. This is a very strong assumption, justifiable in the presence of informationally very rich predictors such as images, but otherwise quite unrealistic for many kinds of predictors considered in medicine, especially those that are more readily available and less invasive and, therefore, more desirable. Second, in contrast to algorithms such as neural networks, random forests, support-vector machines, logistic regression, or generalized linear models, Bayesian nonparametrics does not do an optimization during the learning phase, searching for the minimum of some objective function. It does a full hypothesis-space survey. The optimization done by most machine-learning algorithms is an approximate form of this survey, based on the assumption or hope that the most relevant portion 9
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics and GDS variates are integer-valued, hippocampal volume and Age are continuous, and APOE4 and Sex are binary. The values of one or two of these predictors were missing for 30 subjects in the dataset. The Bayesian-nonparametrics software took less than five hours (on a 16-core Intel Core i9-12900K CPU) to calculate the probability distribution for the possible joint population-frequency distributions of the 13 variates. Some results can already be visualized after this inference. Figure 3 shows, on the left, the inferred distributions of RAVLT-del , RAVLT-imm , GDS , and hippocampal volume for the subpopulation of patients that will convert to Alzheimer’s Disease (red) and the subpopulation that will remain with stable Mild Cognitive Impairment (blue). On the right, the inferred frequency of conversion in the full population is plotted (grey), conditional on the same predictors. The thin curves are 100 samples of highly probable population-frequency distributions; the thicker lines are their means, which are also the predictive conditional probabilities. The two subpopulations of patients are clearly distinct in the RAVLTdel , RAVLT-imm , HV variates. These predictors can yield probabilities of conversion as high as 70% or as low as 10%. The two subpopulations are practically indistinguishable in the GDS variate, which, therefore, always gives very uncertain predictions. The learning dataset comprises enough data to greatly reduce our uncertainty about the population distributions, as evident from the very narrow spread of the curves. In fact it leads to identical answers, within numerical-computation error, even if we drastically change the prior illustrated in fig. 1, for example favouring more unimodal distributions or more multimodal distributions. These simple results show the great usefulness of Bayesian nonparametric software for general medical research. 3.1 Patient’s clinical information The 12 predictor values for our four patients are reported in table 2, top. Note that Curtis’s value for the Hippocampal Volume is missing; this is not a problem for Bayesian nonparametrics. Given these predictor values the method can output any probabilities of interest to the clinician. 16
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics 0 2 4 6 8 10 12 14 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 0.45 RAVLT−del population frequency will convert to AD data histogram stable MCI data histogram 0 2 4 6 8 10 12 14 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 RAVLT−del population frequency of conversion to AD 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 RAVLT−imm population frequency will convert to AD data histogram stable MCI data histogram 5 10 15 20 25 30 35 40 45 50 55 60 65 70 75 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 RAVLT−imm population frequency of conversion to AD 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 GDS population frequency will convert to AD data histogram stable MCI data histogram 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 GDS population frequency of conversion to AD 2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 6.5 7.0 0 50 100 150 200 250 300 350 400 450 500 550 600 650 700 HV population frequency density will convert to AD data histogram stable MCI data histogram 2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 6.5 7.0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 HV population frequency of conversion to AD Figure 3 Inferred distributions of some predictor variates, for the subpopulation of patients that will convert to Alzheimer’s Disease (red dashed) and the subpopulation with stable Mild Cognitive Impairment (solid blue). 17
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics Olivia Ariel Bianca Curtis Age 75.4 75.4 75.4 63.8 Sex F F F M HV/10−34.26 4.26 4.26 [missing] APOE4 N N N Y ANART 18 18 18 15 CFT 21 21 21 14 GDS 3 3 3 2 RAVLT-imm 36 36 36 20 RAVLT-del 5 5 5 0 RAVLT-rec 10 10 10 3 TMTA 21 21 21 36 TMTB 114 114 114 126 p(cAD=Y|predictors)0.302 0.302 0.302 0.703 p(predictors |cAD=Y)/10−12 8.97 8.97 8.97 1.14 p(predictors |cAD=N)/10−12 18.6 18.6 18.6 0.343 Table 2 Predictor values for the four patients (see § 3.0), and resulting conditional probabilities. Table 2, bottom, reports three probabilities that are important for the step of the next subsection:35 •p(cAD=Y|predictors) : the probability that the patient will convert to Alzheimer’s Disease, given the patient’s specific predictors and that the patient comes from the same population as the learning dataset. •p(predictors |cAD=Y) : the probability that a patient who will convert to Alzheimer’s Disease would have these specific predictor values. In other words, the likelihood 36 of conversion to Alzheimer’s Disease, given the predictors. •p(predictors |cAD=N) : the probability that a patient who will remain with stable Mild Cognitive Impairment would have these specific predictor values. In other words, the likelihood of stable Mild Cognitive Impairment, given the predictors. 35 All relative uncertainties of the results caused by numerical computation error are below 0.8%, Curtis’s two likelihoods being an exception at 2%. 18
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics 2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 6.5 7.0 0 50 100 150 200 250 300 350 400 450 500 550 600 650 Curtis's HV probability density Curtis Whole population Figure 4 Probability distribution for Curtis’s Hippocampal Volume (green). The full-population distribution (dashed grey) is also plotted for reference. 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 6.5 frequency of conversion to AD, given predictor values probability density Figure 5 Probability distribution for the frequency of conversion to ad in the subpopulation having Olivia’s predictors. The red vertical line is the value of the probability p(cAD=Y|predictors). Bayesian nonparametrics can also answers other questions of interest to the clinician. For instance, what could be the value of Curtis’s Hippocampal Volume? The answer is given in fig. 4, which also shows the full-population distribution as comparison (dashed grey); with 95% probability Curtis’s value is between 2.8 and 5.3, with a median of 3.8. And what is the frequency of conversion to Alzheimer’s Disease among the subpopulation having Olivia’s, Ariel’s, or Bianca’s predictors? The answer is given in the histogram of fig. 5: with 95% probability, the fraction of this subpopulation that eventually converts to Alzheimer’s Disease is between 0.19 and 0.43; this uncertainty range is due to the limited size of the learning dataset. The probability p(cAD=Y|predictors) is equal to the average of such a distribution 37 , provided the patient and dataset can be considered as belonging to the same population. 3.2 Assessment of relevant subpopulation and auxiliary information Rationale As already mentioned, and as will be argued more concretely in the next section, the clinician needs a probability in order to choose a treatment or other course of action for the current patient. This probability is computed by generalizing associations between predictors and predictand 36 p(𝐴|𝐵) is the probability of 𝐴 given 𝐵 , as well as the likelihood of 𝐵 given 𝐴 (Good 1950 § 6.1). 37 e.g. Bernardo & Smith 2000 §§ 4.2–4.3. 19
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics hidden in a dataset of similar patients, as discussed in § 2. The way this generalization is made, however, can differ from patient to patient in two respects: • Only some particular directed associations can be generalized to the current patient, whereas others would be inappropriate to generalize. In some cases, for example when the learning dataset is artificially assembled with balancing or stratification methods, some associations cannot be generalized to any patients at all. • There can be additional information available for the current patient, for instance some clinical predictors not included in the learning dataset, or other “softer” information such as family history or geographic background. There is no sharp separation between these two items. The presence of additional information often automatically implies that some associations cannot be generalized from the learning dataset to the current patient. Let us explain with a familiar example why particular associations cannot be generalized: the base-rate fallacy 38 . Consider a large set of clinical trials, illustrated in the upper table on the side, where each dot represents, say, 10 000 trial patients. In this sample dataset it is found that, among patients having a particular value “+” of some predictors (left column), 5/7=71.4% of them (upper square) eventually developed a disease. The fallacy lies in judging that a new patient from the full population, who also has predictor value “+”, also has a 71.4% probability of developing that disease. In fact, this probability will in general be different. In our example, it is 5/15 =33.3% , as can be seen in the lower table illustrating the full population. This difference would be noticed as soon as the inappropriate probability was used to make prognoses in the full population. A similar situation happens for the other predictor value “−”. The discrepancy in the conditional frequencies of predictand given predictors, between the sample dataset and the full population, arises 38 Bar-Hillel 1980; Jenny et al. 2018; Sprenger & Weinberger 2021; Matthews 1996. 20
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics because the proportion of positive vs negative disease cases in the latter has some value, 16.7%/83.3% in our example, whereas the trial samples (dashed line in the lower table) were hand-chosen so as to have a 50%/50% proportion. This sampling procedure is called “class balancing” in machine learning 39 . More generally this discrepancy can appear whenever a population and a sample do not have the same frequency distribution for the predictand. In this case we cannot rely on the probabilities of “predictand given predictors” obtained from the sample dataset, which we symbolically write as p(predictand |predictors,dataset)(1) A little counting in the side figure reveals, however, that other frequencies may be relied upon. Consider the full population. Among all patients who developed the disease, 5/6=83.3% of them (upper row) had predictor value “+”, while among those who did not develop the disease, 2/6=33.3% (lower row) had predictor value “ − ”. And these frequencies are the same in the sample dataset. These frequencies from the clinical trials can therefore be used to make a prognosis about the full population, using Bayes’s theorem. For brevity, denote the predictors by 𝑋 , the predictand by 𝑌 , the dataset or trials by 𝐷 , and the full-population base rate by 𝑅. Bayes’s theorem yields p(𝑌|𝑋, 𝐷, 𝑅)= p(𝑋|𝑌, 𝐷) · p(𝑌|𝑅) P 𝑌 p(𝑋|𝑌, 𝐷) · p(𝑌|𝑅)(2) In our example we find p(𝑌=Y|𝑋=+, 𝐷, 𝑅) = p(𝑋=+ | 𝑌=Y, 𝐷)·p(𝑌=Y|𝑅) p(𝑋=+ | 𝑌=Y, 𝐷)·p(𝑌=Y|𝑅) + p(𝑋=+ | 𝑌=N, 𝐷)·p(𝑌=N|𝑅) ≈0.833 ·0.167 0.833 ·0.167 +0.333 ·0.833 =0.33 (3) which is indeed the correct full-population frequency. If the samples of the clinical trials had been chosen with the same frequencies as the full population (no “class balancing”), then 39 Provost 2000; Drummond & Holte 2005; Weiss & Provost 2003. 21
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics the probability p(predictand |predictors,dataset) from the dataset would be the appropriate one to use. But the probabilities p(predictors |predictand,dataset) together with Bayes’s theorem as in eq. (2) would also lead to exactly the same probability. We thus see that using the probabilities p(predictors |predictand,dataset) from the dataset is preferable to using p(predictand |predictors,dataset) . The former yield the same results as the latter when use of the latter is appropriate, and allow us to apply corrections when use of the latter is inappropriate. The superiority of using p(predictors|predictand,dataset) probabilities (called “generative” in machine learning, see e.g. 40 ) is illustrated with a toy example in table 3. The use of dataset probabilities different from p(predictand |predictors,dataset) can be necessary even when the dataset has statistics identical with the population it is sampled from. Typical cases are the prognosis of a patient that comes from a peculiar subpopulation or even from a different population 41 . For instance, the first case happens when the clinician has additional information not included among the predictor variates, such as the result of an additional clinical test, or family history; the second case happens when the patient comes from a different geographical region. There is of course no sharp distinction between these two cases. What is important is that, in either case, it can still be possible to use statistical information from the sample dataset to make prognoses. It is sufficient that some conditional statistics may be applicable to the specific patient. For a patient coming from a different region, for example, it may be judged that the conditional probabilities p(predictand |predictors,dataset) still apply. In other words, the patient may still be considered a member of the subpopulation having those specific predictor values. Using more technical language we say that a new patient can be considered exchangeable with the patients constituting the dataset, but only conditional on particular variates. For a clear and logically impeccable presentation not obscured by technical language, 40 Murphy 2012 § 8.6. 41 Lindley & Novick 1981; Quintana et al. 2017; Sox et al. 2013 ch. 4; Hunink et al. 2014 ch. 5. 22
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics Table 3 Superiority of the “predictors | | |predictand” (or “generative”) approach We split our learning dataset into two subsets: • One with 361 subjects and a ratio of 29.9%/70.1% of subjects with cAD=Yvs cAD=N. • One with 343 subjects and a ratio of 63.3%/36.7% of subjects with cAD=Yvs cAD=N. This subset is used as a fictive full population. This partition was made with no systematic sampling of any variates except the predictand cAD. After training on the learning dataset, we make a prognosis for each of the 343 “new” patients, through four separate approaches: (a) using the probabilities p(predictand |predictors,dataset) , as typical of machine-learning algorithms; (b) using p(predictors |predictand,dataset) together with the base rate, as explained above; (c) tossing a coin; (d) always prognosing “ cAD=Y ”, which guarantees 63.3% correct prognoses owing to the base rate of the full population. Finally, the accuracies (number of prognoses giving more than 50% probability to the correct outcome) of these four approaches are calculated. Here are the results from lowest to highest: predictand |predictors 37.3% coin toss 50% always predict conversion 63.3% predictors |predictand & base rate 73.2% The “predictand | predictors” approach (“discriminative” in machinelearning parlance) leads to worse results than a coin toss because of its underlying base-rate fallacy. The “predictors | predictand” approach (“generative” in machine-learning parlance) leads to better results than simply always prognosing the most common base-rate outcome; this shows that the dataset can still provide useful statistical information despite its mismatched base rate. Inference algorithms that only yield “predictand | predictors” outputs, unlike Bayesian nonparametrics, are incapable of extracting this useful information. 23
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics see Lindley 42 . This topic is also tightly related to confounding and to Simpson’s paradox43. This topic is complex and of extreme importance for inference, but its detailed study is not the goal of the present work. Our main point here is that population variability and auxiliary clinical information are important factors that differentiate patients, and a personalized approach ought to take them into account. The method here presented does this naturally, allowing a great flexibility in selecting which statistical features of the sample dataset should be used for each new patient, and the integration of auxiliary clinical information in the form of a prior probability. As discussed in § 3.1, Bayesian nonparametrics allows us to quickly calculate conditional probabilities p(𝑌|𝑋, dataset) for any desired variate subsets 𝑌 and 𝑋 required by the patient’s relevant population. Application to the example study In our example of § 1.0, all statistics of the dataset are considered relevant for Olivia, Bianca, and Curtis. For these patients the clinician can therefore use Bayes’s theorem with the likelihoods of table 2and the dataset conversion rate of 0.463 – or equivalently directly the probabilities p(cAD=Y|predictors,dataset)provided in the same table. For Ariel, however, the clinician judges that a different base rate or prior probability of conversion should be used, equal to 65%, because of her different geographical origin and family history. In her case the clinician uses Bayes’s theorem with the likelihoods of table 2and the prior probability of 0.65. The final probabilities of conversion to Alzheimer’s Disease for our four patients are reported in table 4. Note how the final probability for Ariel is higher than that for Olivia and Bianca, even if the predictor data are the same for these three patients. 42 Lindley 2014 especially around §§ 7.3, 8.6; Lindley & Novick 1981; more technical references are de Finetti 1930;1937; Dawid 2013; Bernardo & Smith 2000 §§ 4.2–4.3, 4.6. 43 Malinas & Bigelow 2016; Sprenger & Weinberger 2021. 24
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics 3.3 Assessments of treatments and benefits; final decision Rationale A crucial point in clinical decision-making is this: the clinician needs to assess, not the presence (present or future) of a disease, but the risk of its presence. Is there a difference between these two problems? and why is the difference important? In clinical practice, we can rarely diagnose or prognose a medical condition with full certainty. Perfect classification is therefore impossible. But also a “most probable” classification, which may be enough in other contexts, is inadequate in clinical ones. The problem is that the clinician has to decide among different courses of action, such as different treatments, more tests, and so on, and the optimal one depends on how probable the medical condition is, not just on whether it is more probable than not. Two examples illustrate this point. Suppose there is a dangerous treatment that extends the patient’s lifetime by 1 year if the disease is on its course, but shortens the patient’s lifetime by 5 years if the disease is not present. Also suppose that some algorithm tells the clinician whether the disease’s presence is “more probable than not”, given some predictor values; in which case the clinician administers the dangerous treatment. It turns out that 60 out of 100 treated patients having these same predictor values eventually develop the disease, so “more probable than not” is correct. However, the final result is that the clinician has added 1×60 =60 years but also subtracted 5×40 =240 years from the combined lifespans of the treated patients! The conclusion is that the treatment cannot be prescribed just because the disease is “more probably present than not”. As an opposite example, suppose that a less dangerous treatment extends the patient’s lifespan by five years if the disease is on its course, but shortens it by one month if the disease is not present. In this case, it may Olivia Ariel Bianca Curtis initial probability p(cAD=Y|aux info)0.463 0.65 0.463 0.463 final probability p(cAD=Y|predictors,dataset,aux info)0.302 0.47 0.302 0.703 Table 4 Final probabilities of conversion computed from dataset and auxiliary information 25
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics (the four corresponding probability histograms, if plotted jointly, would look like distinct vertical lines). It is clear that knowledge of the Hippocampal Volume is extremely unlikely to change Curtis’s optimal treatment from 𝛿 . Considering that the negligible information gained would not outweigh the economic costs (involving an mri-scan) for obtaining this predictor, the clinician decides to proceed without it. Predictor importance The question about Curtis in the previous subsection can be generalized to a whole population. Predictors that are too invasive or too expensive to obtain, but that are uninformative for the prognosis, could be dropped altogether. So how important, in general, is each predictor in prognosing the conversion to Alzheimer’s Disease? As posed, this question is too vague (ill-posed) because it does not exactly specify how a predictor is used, and what “important” means. Let us see why these details matter. X1 X2 The schematic picture on the side illustrates the necessity of specifying a predictor’s context. Individuals in this population can be either blue circles or red triangles △ , and have two predictors 𝑋1 and 𝑋2 . Predictor 𝑋1 , if used by itself, is worthless in distinguishing the two subpopulations, because these have identical marginal distributions (depicted underneath the grey horizontal line). If used in conjunction with 𝑋2 , however, predictor 𝑋1 allows us to identify an individual’s subpopulation with full certainty, as is clear from the two-dimensional view. It is therefore an essential predictor in this case: dropping it would lead to a complete loss of predictive power. An analogous discussion holds for 𝑋2 in the present case. The converse can also happen (not illustrated): a predictor might be “good” if used by itself, and yet it might be discarded without any loss if used in combination with others. In our question about a predictor’s importance, we want to know what happens if the predictor is dropped from the set of all predictors. Regarding the meaning of “importance” or “prognostic power”, we must specify a relevant metric, and predictors could be ranked differently by different metrics. From our discussion so far it is clear that in clinical 32
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics decision-making the canonical metric is the final expected utility – and therefore the choice of optimal treatment – which a predictor’s presence or absence leads to (see § 3.3). This point was illustrated with Curtis’s example in the previous subsection. What if we want to make a similar assessment, not for a single patient, but for the full population? which utility matrix should we use? It can be proved, again from decisiontheoretic principles, that the population average of all utility matrices should be used in this case 49 . This seems a quantity very difficult to assess, but it can also be shown 50 that even a semi-quantitative assessment leads to better results than using some other general-purpose metric. Bayesian nonparametrics allows us to compute the expected value of virtually any prognostic-importance metric, and for any subset of predictors available in the dataset. This computation has moreover two properties of paramount importance: (a) the prognostic power of a set of predictors found with Bayesian nonparametrics is the maximum possible obtainable by any inference algorithm, or in other words it is an intrinsic property of that set of predictors; (b) Bayesian nonparametrics achieves this maximum power. Thus, if Bayesian nonparametrics says that the accuracy obtainable with a given set of predictors is 70%, then we know that no other inference algorithm can reach a higher accuracy than 70%; inference algorithms that reach lower accuracy can in principle be improved upon. Bayesian nonparametrics, by construction, will reach this accuracy. Note that we mean accuracy in the long run, over the full population; an inference algorithm could reach higher accuracies in some test dataset thanks to sampling fluctuations; in fact this is bound to happen from time to time.51 Let us illustrate this kind of “predictor importance” assessment for our dataset. We use (a) two metrics: the accuracy and the mutual information 52 between a set of predictors and the cAD predictand; (b) 27 different sets of predictors: •every predictor, used individually (12 sets); • all cognitive-test predictors used together, jointly with information about depression (GDS) and demographics (Age and Sex). 49 cf. Dyrland et al. 2022a § 4.1. 50 Dyrland et al. 2022a § 4.2. 51 Bayesian nonparametrics can also calculate, with a somewhat expensive computation, the size of such fluctuations, given the size of the test dataset. 52 Shannon 1948; Cover & Thomas 2006. 33
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics •APOE4 and Hippocampal Volume, jointly with demographic information; • all predictors jointly excluding one, each single predictor being excluded in turn (12 sets); •all predictors jointly. Use of the accuracy assumes that the population of patients has only two available treatments having average utility matrix 1 0 0 1 . Mutual information is a model-free measure of the relation between two sets of variates, with diverse operational interpretations 53 and international standards 54 . A set of predictors and a binary variate (such as our conversion to Alzheimer’s Disease) have a mutual information of 1 Sh if and only if there is a non-constant deterministic function from the former to the latter. Our specific questions are the following: “What is the expected value of the accuracy for the next new patient, if we use the given set of predictors?” and “What is the mutual information between the given set of predictors and the predictand, given the presently available data?”. The answers to these questions are reported in fig. 7, ordered from bottom to top according to increasing metric. The ordering of mutual information and accuracy agree within the uncertainty of the numerical computation (Monte Carlo integration). The latter is reported as coverage intervals of ±two standard deviations. The plots reveal several findings, valid within the population selected for the dataset, which can be compared with the analysis in Rye et al.55: • The set of 12 predictors considered in the present work and in Rye et al. 56 can at most yield a prognostic accuracy of around 67.7%±0.7% over the full population, for any inference algorithm. This fact agrees with the (completely independent) findings in Rye et al. 57 , where a maximal accuracy of 68.3% on a test dataset was found using an ensemble model. The present analysis also shows that the ensemble model managed to achieve the maximal accuracy possible with these predictors (but see § 4for limitations of that model). • The mutual information using all 12 predictors is quite low at (0.140 ±0.008)Sh , indicating that we cannot reasonably consider 53 MacKay 2005; Woodward 1964; Minka 2003; Good 1961; Good & Toulmin 1968; Kelly 1956; Kullback 1978. 54 iso 2008. 55 Rye et al. 2022 see especially Fig. 3 and Table 3. 56 Rye et al. 2022.57 Rye et al. 2022. 34
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics 0.50 0.54 0.58 0.62 0.66 0.70 0.74 0.78 accuracy (expected utility) APOE4 Sex GDS Age ANART TMTA CFT HC TMTB APOE4+HC+Age+Sex RAVLT−rec RAVLT−imm RAVLT−del all minus RAVLT−del cognitive+Age+Sex all minus RAVLT−imm all minus HC all minus GDS all minus Age all minus APOE4 all minus TMTA all minus Sex all minus CFT all minus ANART all minus RAVLT−rec all minus TMTB all 0.00 0.02 0.04 0.06 0.08 0.10 0.12 0.14 0.16 0.18 0.20 0.22 mutual information/Sh GDS Sex APOE4 Age ANART TMTA TMTB CFT HC APOE4+HC+Age+Sex RAVLT−rec RAVLT−imm RAVLT−del all minus RAVLT−del all minus RAVLT−imm all minus TMTB all minus RAVLT−rec all minus HC cognitive+Age+Sex all minus TMTA all minus CFT all minus Age all minus GDS all minus ANART all minus Sex all minus APOE4 all Figure 7 Expected accuracy for the next new patient (left), and mutual information (right), of several sets of predictors for the prognosis of conversion to Alzheimer’s Disease. Each graph has been vertically ordered according to increasing values; the two rankings agree within the respective uncertainties. The all predictor set is mathematically guaranteed to be optimal according to both metrics and has therefore been ranked first. Bars show the uncertainty interval (±two standard deviations). the predictand to be an approximate function of the predictors ( 0 Sh corresponds to a coin toss, 1 Sh to a perfect function). Machinelearning algorithms based on functional regression, such as neural networks, are therefore not appropriate for this prognostic problem. •APOE4 , GDS , Age , Sex , and to some degree ANART are poor predictors (within this population) when used alone and when used in combination with all other predictors. The latter point is evident from the fact that the mutual information and accuracy of the combined predictors barely decreases if any one of these four predictors is omitted. • The combined cognitive and demographic variates are better predictors than the joint use of Hippocampal Volume, APOE4 , and demographic variates. 35
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics 2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 6.5 7.0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 Curtis's unknown HV Curtis's probability of conversion to AD Figure 8 Probability of cAD=Y for Curtis, given Curtis’s known predictors and different possible values of his unknown Hippocampal Volume. The thinner curves are 100 probable samples of how this probability would change with a larger learning dataset. Compare this figure with fig. 3,p.17, bottom-right. •RAVLT-imm , RAVLT-del , and to a lesser degree RAVLT-rec are good predictors, both when used alone and when used jointly with all other predictors. Hippocampal Volume is a poorer predictor than any of the RAVLT when used alone, and likely also when used in combination with all others 58 . This last finding is also clear in Curtis’s case: fig. 8shows that his probability of conversion to Alzheimer’s Disease, given his current predictors, would practically be the same for all values of Hippocampal Volume; and it would probably be the same even if the learning dataset contained more points. Bayesian nonparametrics shows that the omission of any one of the 12 predictors, except RAVLT-del and possibly RAVLT-imm , does not lead to an appreciable decrease in accuracy (relative decrease of 0.3% or less) or in mutual information (relative decrease of less than 3%). This puts the prognostic-importance analysis of Rye et al. 59 into perspective. The 58 contrast this with Rye et al. 2022.59 Rye et al. 2022. 36
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics exact quantification of these subtle differences is computationally quite expensive, and we did not carry it out further. 4 Discussion Which requirements does a personalized approach to prognosis and treatment impose on assistive computational technology? This is an important question, because with the increasing amount of statistical clinical data and clinical predictors available for medical care, assistive computational technology is today not merely a useful option, but a necessity in clinical practice. In the present work we started from the perspective of the clinician’s ultimate task, decision-making under risk, and saw that patients’ differences relevant to prognosis and treatment can be approximately divided into three categories: • differences in the values – and availability – of a core set of clinical predictors, for which we have population-wide statistical information; • differences in the availability and values of auxiliary and usually semi-quantitative clinical information, such as geographical or family background; • differences in the availability and values or “utilities” of clinical courses of action, such as preventive treatments or further tests; such values can have a highly variable, patient-dependent subjective component. Luckily there is a theory that takes into account and integrates these differences towards the final goal: Decision Theory, which is the subject of several good textbooks on clinical decision-making 60 after the pioneering work of Lesley & Lusted 1959a,b;1960;1960;1968 (a summary and references were given in § 3.3). Decision-making under risk requires any assistive algorithm to work, explicitly or implicitly, in terms of probabilities, having precise connections with population statistics (§3.3). Without this condition the integration of patient-dependent treatment utilities would be impossible. 60 Weinstein & Fineberg 1980; Sox et al. 2013; Hunink et al. 2014. 37
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics Table 6 Summary of the clinician’s patient-dependent inputs and Bayesian-nonparametrics outputs. Input data and final results that distinguish Ariel, Bianca, Curtis from Olivia are in red. Olivia Ariel Bianca Curtis Clinician’s patient-dependent inputs Predictor values Age 75.4 75.4 75.4 63.8 Sex FFFM HV/10−34.26 4.26 4.26 [missing] APOE4 NNNY ANART 18 18 18 15 CFT 21 21 21 14 GDS 3332 RAVLT-imm 36 36 36 20 RAVLT-del 5550 RAVLT-rec 10 10 10 3 TMTA 21 21 21 36 TMTB 114 114 114 126 Additional information auxiliary info none family history, base rate none none applicable subpopulation all predictor |predictand all all prior probability of conversion 0.463 0.65 0.463 0.463 Available treatments and utilities treatment 𝛼 treatment 𝛽 treatment 𝛾 treatment 𝛿 cAD N Y 10 0 9 3 8 5 0 10 cAD N Y 10 0 9 3 8 5 0 10 cAD N Y 10 0 8 3 7 5 0 10 cAD N Y 10 0 9 3 8 5 0 10 Bayesian-nonparametrics outputs p(cAD=Y|predictors,dataset)0.302 0.302 0.302 0.703 p(predictors|cAD=Y,dataset)/10−12 8.97 8.97 8.97 1.14 p(predictors|cAD=N,dataset)/10−12 18.6 18.6 18.6 0.343 final probability of conversion p(cAD=Y|predictors,dataset,aux info)0.302 0.47 0.302 0.703 exp. utility treatment 𝛼 exp. utility treatment 𝛽 exp. utility treatment 𝛾 exp. utility treatment 𝛿 Optimal treatment 6.98 7.19 7.09 3.02 𝜷 5.27 6.16 6.58 4.73 𝜸 6.98 6.49 6.40 3.02 𝜶 2.97 4.78 5.89 7.03 𝜹 38
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics The handling of these probabilities should moreover be enough flexible to take into account peculiar but common subpopulations of patients having special contexts or auxiliary information (§ 3.2), and the common possibility of missing values for some clinical predictors (§3.1). Most, if not all, popular machine-learning algorithms either do not meet these requirements, or they do so at the cost of unrealistic modelling assumptions. Unfortunately they tend to overly simplify the problem of decision-making under risk, as if it were a simple classification or regression task. We presented an assistive method and software, implementing Bayesian nonparametric inference, that meets all these requirements (§ 2) and carries out the calculations required by decision theory. This method is moreover model-free, not making a-priori assumptions about functional dependencies or particular distributions in the variates. The inference principles on which it is based have recently been recommended for the study of Alzheimer’s Disease 61 , and have been successfully demonstrated in a simpler predictor setting 62 . We showed its application in an example of prognosis and treatment of conversion from Mild Cognitive Impairment to Alzheimer’s Disease for four different patients, where all three categories of differences listed above appeared. The patients were fictitious but the underlying learning database, originating from adni, is real and was explored in a previous work63. Bayesian nonparametrics was also shown to have uses that go beyond individual clinical decision-making but are still of importance to personalized medicine. For instance, it can assess the maximum possible prognostic power of particular sets of predictors, potentially allowing us to discard clinical predictors that are too invasive or expensive and yet prognostically unimportant. In actual deployment, we would recommend the hospital, medical centre, or clinician using Bayesian-nonparametrics software to keep a database of incoming patients, with their predictor values, adding the true values of their predictand later in time, once they become known. The software can then be retrained on such local database when the latter reaches a size comparable to the original one’s, and periodically retrained afterwards. All inferences would thus become increasingly more reliable, 61 Temp et al. 2021; see also asa 2016;2019; Jawa & Maslove 2023; Sidebotham 2020; Goodman 1999.62 Antoniano-Villalobos et al. 2014.63 Rye et al. 2022. 39
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics because the software would base them on updated population statistics that are characteristic to the specific hospital. 4.1 Counters to possible critiques Any inference or decision-making algorithm aspiring to take into account patient differences must perforce have some open “input slots” for such differences. We saw that Bayesian nonparametrics requires inputs about a patient’s specific predictors, relevant statistical relations and auxiliary data, and treatment utilities. The most difficult input to quantify is probably the third: translating benefits and drawbacks of different treatments into numbers. On this complex topic we refer the reader to specially dedicated textbooks on clinical decision making, for example Sox et al. 2013 and Hunink et al. 2014. But some readers may wonder: “can all these additional inputs be avoided?”, fearing that errors could sneak in through them. This question is answered by a mathematical theorem at the very core of decision theory 64 , which is too seldom emphasized: Any decision we make, either (A) comes explicitly or implicitly through some set of utilities and maximization of their expectations, or (B) is logically inconsistent. There is no third alternative. Thus the choice is not between using utilities or not using utilities, but between choosing them explicitly or letting them be chosen in a way we do not know. If we use a decision-making algorithm that does not ask us for utilities, then the algorithm is internally supplying utilities not chosen by us (and probably divorced from our specific problem), or, worse, is committing logical inconsistencies. The first advantage of explicitly operating through utilities, probabilities, decision theory, is that we are, at the very least, sure of not acting in a self-contradictory way. The second advantage is that the utilities used to arrive at a decision appear openly in front of us. We can analyse and change them if we find them inappropriate to a specific problem. If they are hidden, it is more difficult to analyse which are inappropriate and how they should be changed. The fact that an algorithm works according to decision theory is also an assurance of striving towards theoretical optimality. This point has 64 Savage 1972; Luce & Raiffa 1957; Raiffa & Schlaifer 2000; Atkinson et al. 1964; Ferguson 1967; Lindley 1988;1977; Kreps 1988; Bernardo & Smith 2000; Pratt et al. 1996; Lindley 2014; Pettigrew 2019. 40
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics very subtle consequences. Consider a non-optimal algorithm that leads to saving 85 000 patients out of 100 000. Given these numbers it might be deemed a success. But what if a theoretically optimal algorithm leading to 95 000 saved patients is feasible? What shall we say to the families of the 10 000 patients who could have been saved but weren’t? 4.2 Range of application of Bayesian-nonparametrics software First let us emphasize, even if it is obvious, that the quality of the results obtained with Bayesian nonparametric inference depends on the quality of the learning dataset. Any peculiar sampling biases (or numerical errors) in the dataset that are unknown to the clinician will affect the final results. This is of course true for any inference algorithm. But we saw that Bayesian nonparametrics allows the clinician to correct for particular sampling biases present in the dataset, if they are known. The range of application of current Bayesian-nonparametrics software has two kinds of bounds: computational and theoretical. The fact that Bayesian nonparametrics extracts all available information from the dataset makes it computationally expensive (see § 2). At present it cannot be used with high-dimensional predictors: if our dataset had included a predictor such as a 128 ×128 ×128 greyscale mri image, the learning stage would have taken around 100 years. Approximate but much faster algorithms such as neural networks and random forests are thus, at present, still the only options with such predictors. There is, however, the interesting possibility of combining these fast algorithms together with Bayesian-nonparametrics software, as a post-processor of their raw output. The software extracts useful information usually hidden in their output at a low computational cost 65 ; this information can then be used for clinical decision-making as illustrated in the present work. The sole assumption underlying Bayesian nonparametric inference and its practical use with new patients, is that the latter can be assumed to come, at least in some respects, from the same population as the learning dataset (in probability-theory jargon, partial or conditional exchangeability applies; see § 3.2). This precludes using the present Bayesian-nonparametric software to forecast how the statistics of the full population could change in the future. However, the software can 65 Dyrland et al. 2022b. 41
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics Good, I. J. (1961): Weight of evidence, causality and false-alarm probabilities. In: Cherry 1961, ch. 11, 125–136. With Discussion by A. J. Mayne and D. M. MacKay and reply. Good, I. J., Toulmin, G. H. (1968): Coding theorems and weight of evidence. IMA J. Appl. Math. 41, 94–105. doi:10.1093/imamat/4.1.94. Goodman, S. N. (1999): Toward evidence-based medical statistics. 1: The 𝑃 value fallacy. 2: The Bayes factor. Ann. Intern. Med. 130 12 , 995–1013. doi: 10.7326/0003-4819-130-12-1 99906150-00008 ,doi: 10.7326/0003-4819-130-12-199906150-00019 . http://ww w.perfendo.org/docs/bayesprobability/5.3_goodmanannintmed99all.pdf . See also comments and reply in Davidoff 1999; Sulmasy et al. 2000. Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., Altman, D. G., Altman, N. S., et al. (2016): Online supplement and discussion: ASA statement on statistical significance and 𝑝 -values. Am. Stat. 70 2 , 129. doi: 10.1080/00031305.2016.11 54108 supplemental material. See asa 2016 and Wasserstein, Lazar 2016. Hailperin, T. (1996): Sentential Probability Logic: Origins, Development, Current Status, and Technical Applications. (Associated University Presses, London). https://archive.or g/details/hailperin1996-Sentential_probability_logic. Halpern, J. Y. (1999): Cox’s theorem revisited. J. Artif. Intell. Res. 11, 429–435. doi: 10.1613/j air.644. See also Snow 1998. Harper, W. L., Hooker, C. A., eds. (1976): Foundations of Probability Theory, Statistical Inference, and Statistical Theories of Science. Vol. II: Foundations and Philosophy of Statistical Inference. (Reidel, Dordrecht). Hjort, N. L., Holmes, C., Müller, P., Walker, S. G., eds. (2010): Bayesian Nonparametrics. (Cambridge University Press, Cambridge). doi:10.1017/CBO9780511802478. Hunink, M. G. M., Weinstein, M. C., Wittenberg, E., Drummond, M. F., Pliskin, J. S., Wong, J. B., Glasziou, P. P. (2014): Decision Making in Health and Medicine: Integrating Evidence and Values, 2nd ed. (Cambridge University Press, Cambridge). doi: 10.1017/CBO9781 139506779. First publ. 2001. Ishwaran, H., Zarepour, M. (2002): Dirichlet prior sieves in finite normal mixtures. Stat. Sinica 12 3 , 941–963. http://www3.stat.sinica.edu.tw/statistica/J12n3/j12n316/j12 n316.htm. iso (2008): ISO 80000-11:2008: Quantities and units 11: Characteristic numbers. International Organization for Standardization. Jawa, N. A., Maslove, D. M. (2023): Bayes’ theorem in neurocritical care: principles and practice. Neurocrit. Care 2023, 1. doi: 10.1007/s12028-022-01665-2 . Note that some statements in this paper are incorrect or misleading. Table 5 states as advantage of frequentist statistics “not susceptible to prior beliefs”; this is not true. Frequentist statistics does use a prior, but keeps it hidden: see e.g. Pratt 1961 p. 167, Savage et al. 1962 p. 49 item (ii); so in frequentist statistics it becomes difficult to check whether this hidden prior is appropriate to the given problem. On the other hand, if the frequentist statistics did not implicitly use a prior, it would lead to contradictions – and in some cases indeed it does – see references above and Lindley 1977. Table 5 also states as disadvantage of Bayesian statistics “Priors are subjective and may be biased”. This statement is unfair: a prior is no less subjective than choosing “0.05” as a threshold of “statistical significance”: why this particular value? what is its rationale? If one answers that it is a value generally agreed upon by the community, then such general community agreement exists also for priors; thus “subjective” is untrue. But the truth is that there are debates about the appropriate “significance” level, which is therefore as “subjective” as a prior. 48
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics Jaynes, E. T. (1976): Confidence intervals vs Bayesian intervals. In: Harper, Hooker 1976, 175–257. With discussion, comments by M. Maxfield and O. Kempthorne, and reply. Repr. with an introduction in Jaynes 1989 pp. 149–209; https://bayes.wustl.edu/e tj/node1.html. Jaynes, E. T. (1989): Papers on Probability, Statistics and Statistical Physics, repr. (Kluwer, Dordrecht). Edited by R. D. Rosenkrantz. First publ. 1983. Jaynes, E. T. (2003): Probability Theory: The Logic of Science. (Cambridge University Press, Cambridge). Ed. by G. Larry Bretthorst. First publ. 1994. doi: 10.1017/CBO97805117 90423 , https://archive.org/details/XQUHIUXHIQUHIQXUIHX2 , http://www-biba .inrialpes.fr/Jaynes/prob.html. jcgm (2008): JCGM 100:2008: Evaluation of measurement data – Guide to the Expression of Uncertainty in Measurement, corr. version. Joint Committee for Guides in Metrology (jcgm): bipm,iec,ifcc,ilac,iso,iupac,iupap,oiml. http://www.bipm.org/en/public ations/guides/gum.html. Includes various supplements. First publ. 1993. jcgm (2012): JCGM 200:2012: International vocabulary of metrology – Basic and general concepts and associated terms (VIM), 3rd ed. Joint Committee for Guides in Metrology (jcgm): bipm,iec,ifcc,ilac,iso,iupac,iupap,oiml. https://www.bipm.org/en/publication s/guides/vim.html. First publ. 1997. Jenny, M. A., Keller, N., Gigerenzer, G. (2018): Assessing minimal medical statistical literacy using the Quick Risk Test: a prospective observational study in Germany. BMJ Open 8, e020847, e020847corr2. doi: 10.1136/bmjopen-2017-020847 ,doi: 10.1136/bmjopen2017-020847corr2. Johnson, W. E. (1924): Logic. Part III: The Logical Foundations of Science. (Cambridge University Press, Cambridge). https://archive.org/details/logic03john. Johnson, W. E. (1932): Probability: the deductive and inductive problems. Mind 41 164 , 409–423. With some notes and an appendix by R. B. Braithwaite. doi: 10.1093/mind/XLI.164.4 09. Kelly Jr., J. L. (1956): A new interpretation of information rate. Bell Syst. Tech. J. 354, 917–926. http://turtletrader.com/kelly.pdf , https://archive.org/details/bstj35-4 -917. Kreps, D. (1988): Notes On The Theory Of Choice. (Routledge, New York). doi: 10.4324/978 0429498619. Kullback, S. (1978): Information Theory and Statistics. (Dover, New York). Republ. with a new preface and corrections and additions by the author. First publ. 1959. Kyburg Jr., H. E., Smokler, H. E., eds. (1980): Studies in Subjective Probability, 2nd ed. (Robert E. Krieger, Huntington, USA). First publ. 1964. Ledley, R. S. (1959): Digital electronic computers in biomedical science: computers make solutions to complex biomedical problems feasible, but obstacles curb widespread use. Science 130 3384 , 1225–1234. doi:10.1126/science.130.3384.1225. Ledley, R. S. (1960): Digital Computer and Control Engineering. (McGraw-Hill, New York). Written with the assistance of Louis S. Rotolo and James Bruce Wilson. https://arch ive.org/details/bitsavers_columbiaUnuterandControlEngineering1960_4075 2710. Ledley, R. S., Lusted, L. B. (1959a): Reasoning foundations of medical diagnosis: symbolic logic, probability, and value theory aid our understanding of how physicians reason. Science 130 3366 , 9–21. doi:10.1126/science.130.3366.9. Ledley, R. S., Lusted, L. B. (1959b): The use of electronic computers to aid in medical diagnosis. Proc. IRE 4711, 1970–1977. doi:10.1109/JRPROC.1959.287213. 49
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics Ledley, R. S., Lusted, L. B. (1960): Computers in medical data processing. Oper. Res. 8 3 , 299–310. doi:10.1287/opre.8.3.299. Lindley, D. V. (1977): The distinction between inference and decision. Synthese 36 1 , 51–58. doi: 10.1007/BF00485691. Lindley, D. V. (1988): Making Decisions, 2nd ed. (Wiley, London). First publ. 1971. Lindley, D. V. (2014): Understanding Uncertainty, rev. ed. (Wiley, Hoboken, USA). First publ. 2006. Lindley, D. V., Novick, M. R. (1981): The role of exchangeability in inference. Ann. Stat. 9 1 , 45–58. doi:10.1214/aos/1176345331. Liu, C. - C., Kanekiyo, T., Xu, H., Bu, G. (2013): Apolipoprotein E and Alzheimer disease: risk, mechanisms and therapy. Nat. Rev. Neurol. 9 2, 4 , 106–118, 184. doi: 10.1038/nrneurol.2 012.263,doi:10.1038/nrneurol.2013.32. Luce, R. D., Raiffa, H. (1957): Games and Decisions: introduction and critical survey. (Wiley, New York). https://archive.org/details/img-1907_202109. Lusted, L. B. (1967): Logical analysis in medical diagnosis. Berkeley Symp. Math. Stat. Probab. 5/IV, 903–923. https://projecteuclid.org/proceedings/berkeley-symposium-o n-mathematical-statistics-and-probability/Proceedings-of-the-Fifth-Be rkeley-Symposium-on-Mathematical-Statistics-and/Chapter/Logical-analy sis-in-medical-diagnosis/bsmsp/1200513835. Lusted, L. B. (1968): Introduction to Medical Decision Making. (Thomas, Springfield, USA). Lusted, L. B., Ledley, R. S. (1960): Mathematical models in medical diagnosis. J. Med. Educ. 35 3 , 214–222. https://journals.lww.com/academicmedicine/Citation/1960/030 00/Mathematical_Models_in_Medical_Diagnosis.2.aspx. MacKay, D. J. C. (1992a): Bayesian interpolation. Neural Comput. 4 3 , 415–447. https://www .inference.org.uk/mackay/PhD.html,doi:10.1162/neco.1992.4.3.415. MacKay, D. J. C. (1992b): A practical Bayesian framework for backpropagation networks. Neural Comput. 4 3 , 448–472. https://www.inference.org.uk/mackay/PhD.html ,doi: 10.1162/neco.1992.4.3.448. MacKay, D. J. C. (2005): Information Theory, Inference, and Learning Algorithms, version 7.2 (4th pr.) (Cambridge University Press, Cambridge). https://www.inference.org.u k/itila/book.html . First publ. 1995. See also video lectures at https://videolectu res.net/events/course_information_theory_pattern_recognition. Malinas, G., Bigelow, J. (2016): Simpson’s Paradox. In: Zalta 2023. https://plato.stanfor d.edu/archives/fall2016/entries/paradox-simpson. First publ. 2004. Matthews, R. A. J. (1996): Base-rate errors and rain forecasts. Nature 382 6594 , 766. doi: 10.1038/382766a0. McKhann, G., Drachman, D., Folstein, M., Katzman, R., Price, D., Stadlan, E. M. (1984): Clinical diagnosis of Alzheimer’s disease report of the NINCDS-ADRDA Work Group under the auspices of Department of Health and Human Services Task Force on Alzheimer’s disease. Neurology 347, 939–944. doi:10.1212/WNL.34.7.939. Minka, T. P. (2003): Bayesian inference, entropy, and the multinomial distribution. Tech. rep. (MIT media Lab, Cambridge, USA). https://tminka.github.io/papers/multinom ial.html. First publ. 1998. Murphy, K. P. (2012): Machine Learning: A Probabilistic Perspective. (MIT Press, Cambridge, USA). Neal, R. M. (1993): Probabilistic inference using Markov chain Monte Carlo methods. Tech. rep. CRG-TR-93-1. (University of Toronto, Toronto). http://www.cs.utoronto.ca/~radf ord/review.abstract.html,https://omega0.xyz/omega8008/neal.pdf. 50
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics O’Hagan, A., Buck, C. E., Daneshkhah, A., Eiser, J. R., Garthwaite, P. H., Jenkinson, D. J., Oakley, J. E., Rakow, T. (2006): Uncertain Judgements: Eliciting Experts’ Probabilities. (Wiley, Chichester). doi:10.1002/0470033312. Paris, J. B. (2006): The Uncertain Reasoner’s Companion: A Mathematical Perspective, repr. (Cambridge University Press, Cambridge). doi: 10.1017/CBO9780511526596 . See also Snow 1998. Petersen, R. C., Aisen, P. S., Beckett, L. A., Donohue, M. C., Gamst, A. C., Harvey, D. J., Jack, C. R., Jagust, W. J., et al. (2010): Alzheimer’s Disease Neuroimaging Initiative (ADNI): clinical characterization. Neurology 74 3 , 201–209. doi: 10.1212/WNL.0b013e3181cb3e25 . Pettigrew, R. (2019): Epistemic Utility Arguments for Probabilism. In: Zalta 2023. https://p lato.stanford.edu/archives/win2019/entries/epistemic-utility . First publ. 2011. Pólya, G. (1954): Mathematics and Plausible Reasoning: Vol. I: Induction and Analogy in Mathematics. (Princeton University Press, Princeton). https://archive.org/detail s/Induction_And_Analogy_In_Mathematics_1_,doi:10.1515/9780691218304. Pólya, G. (1968): Mathematics and Plausible Reasoning: Vol. II: Patterns of Plausible Inference, 2nd ed. (Princeton University Press, Princeton). First publ. 1954. Porta Mana, P. G. L. (2019): A relation between log-likelihood and cross-validation log-scores. OSF doi: 10.31219/osf.io/k8mj3 , HAL https://hal.science/hal-02267943 , arXiv doi: 10.48550/arXiv.1908.08741. Pratt, J. W. (1961): Book review: Testing Statistical Hypotheses, E. L. Lehmann. J. Am. Stat. Assoc. 56293, 163–167. doi:10.1080/01621459.1961.10482103. Pratt, J. W., Raiffa, H., Schlaifer, R. (1996): Introduction to Statistical Decision Theory, 2nd pr. (MIT Press, Cambridge, USA). First publ. 1995. Provost, F. (2000): Machine learning from imbalanced data sets 101. Tech. rep. WS-00-05-001. (AAAI, Menlo Park, USA). https://aaai.org/Library/Workshops/2000/ws00-05001.php. Quintana, M., Viele, K., Lewis, R. J. (2017): Bayesian analysis: using prior information to interpret the results of clinical trials. J. Am. Med. Assoc. 318 16 , 1605–1606. doi: 10.1001/jama.2017.15574. R Core Team (2023): R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing. https://www.R-project.org. First released 1995. Raiffa, H. (1970): Decision Analysis: Introductory Lectures on Choices under Uncertainty, 2nd pr. (Addison-Wesley, Reading, USA). First publ. 1968. Raiffa, H., Schlaifer, R. (2000): Applied Statistical Decision Theory, repr. (Wiley, New York). First publ. 1961. Rasmussen, C. E. (1999): The infinite Gaussian mixture model. Adv. Neural Inf. Process. Syst. (NIPS) 12, 554–560. https://www.seas.harvard.edu/courses/cs281/papers/rasm ussen-1999a.pdf. Rodríguez, A., Dunson, D. B., Gelfand, A. E. (2009): Bayesian nonparametric functional data analysis through density estimation. Biometrika 96 1 , 149–162. doi: 10.1093/biomet/asn0 54. Rosenkrantz, R. D. (1977): Inference, Method and Decision: Towards a Bayesian Philosophy of Science. (Reidel, Dordrecht). Rossi, P. E. (2014): Bayesian Nonand Semi-parametric Methods and Applications. (Princeton University Press, Princeton). doi:10.1515/9781400850303. 51
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics Russell, S. J., Norvig, P. (2022): Artificial Intelligence: A Modern Approach, 4th global ed. (Pearson, Harlow, UK). http://aima.cs.berkeley.edu/global-index.html . First publ. 1995. Rye, I., Vik, A., Kociński, M., Lundervold, A. S., Lundervold, A. J. (2022): Predicting conversion to Alzheimer’s disease in individuals with Mild Cognitive Impairment using clinically transferable features. Sci. Rep. 12, 15566. doi:10.1038/s41598-022-18805-5. Savage, L. J. (1972): The Foundations of Statistics, 2nd rev. and enl. ed. (Dover, New York). First publ. 1954. Savage, L. J., Bartlett, M. S., Barnard, G. A., Cox, D. R., Pearson, E. S., Smith, C. A. B. (1962): The Foundations of Statistical Inference: A Discussion. (Methuen, London). With a discussion including H. Ruben, I. J. Good, D. V. Lindley, P. Armitage, C. B. Winsten, R. Syski, E. D. Van Rest, G. M. Jenkins. Self, M., Cheeseman, P. C. (1987): Bayesian prediction for artificial intelligence. In: Proceedings of the Third Conference on Uncertainty in Artificial Intelligence (UAI’87), ed. by J. Lemmer, T. Levitt, L. Kanal (AUAI Press, Arlington, USA), 61–69. Repr. in arXiv doi: 10.48550 /arXiv.1304.2717. Shah, S., Beck, J. R., Pauker, S. G. (2013): In memoriam: Robert Steven Ledley, DDS, MS (physics), 1926–2012. Med. Decis. Making 335, 731–733. doi:10.1177/0272989X1348794. Shannon, C. E. (1948): A mathematical theory of communication. Bell Syst. Tech. J. 27 3, 4 , 379–423, 623–656. https://archive.org/details/bstj27-3-379 , https://archiv e.org/details/bstj27-4-623 , http://math.harvard.edu/~ctm/home/text/othe rs/shannon/entropy/entropy.pdf. Sidebotham, D. (2020): Are most randomised trials in anaesthesia and critical care wrong? an analysis using Bayes’ theorem. Anaesthesia 75 10 , 1386–1393. doi: 10.1111/anae.15029 . See also accompanying editorial Charlesworth, Pandit 2020. Smith, J. E., Winkler, R. L. (2006): The optimizer’s curse: skepticism and postdecision surprise in decision analysis. Manag. Sci. 523, 311–322. doi:10.1287/mnsc.1050.0451. Snow, P. (1998): On the correctness and reasonableness of Cox’s theorem for finite domains. Comput. Intell. 143, 452–459. doi:10.1111/0824-7935.00070. Snow, P. (2001): The reasonableness of possibility from the perspective of Cox. Comput. Intell. 171, 178–192. doi:10.1111/0824-7935.00138. Sox, H. C., Higgins, M. C., Owens, D. K. (2013): Medical Decision Making, 2nd ed. (Wiley, New York). doi:10.1002/9781118341544. First publ. 1988. 3rd ed. Sox et al. 2024. Sox, H. C., Higgins, M. C., Owens, D. K., Schmidler, G. S. (2024): Medical Decision Making, 3rd ed. (Wiley, New York). doi:10.1002/9781119627876. First publ. 1988. Sprenger, J., Weinberger, N. (2021): Simpson’s Paradox. In: Zalta 2023. https://plato.sta nford.edu/archives/sum2021/entries/paradox-simpson. Sulmasy, D. P., Morgan, T., Caubet, J. - F., Goodman, S. (2000): Toward evidence-based statistics [comments and response]. Ann. Intern. Med. 132 6 , 507–508. See Goodman 1999; Davidoff 1999. Temp, A. G. M., Lutz, M. W., Trepel, D., Tang, Y., Wagenmakers, E. - J., Khachaturian, A. S., Teipel, S. (2021): How Bayesian statistics may help answer some of the controversial questions in clinical research on Alzheimer’s disease. Alzheimer’s Dement. 17 6 , 917–919. doi:10.1002/alz.12374. Tribus, M. (1969): Rational Descriptions, Decisions and Designs. (Pergamon, New York). doi: 10.1016/C2013-0-01558-7. Van Horn, K. S. (2003): Constructing a logic of plausible inference: a guide to Cox’s theorem. Int. J. Approximate Reasoning 341, 3–24. doi:10.1016/S0888-613X(03)00051-3. 52
Porta Mana et al. Personalized prognosis & treatment with Bayesian nonparametrics von Neumann, J., Morgenstern, O. (1955): Theory of Games and Economic Behavior, 3rd ed., 6th pr. (Princeton University Press, Princeton). https://archive.org/details/in .ernet.dli.2015.215284. First publ. 1944. Walker, S. G. (2010): Bayesian nonparametric methods: motivation and ideas. In: Hjort et al. 2010, ch. 1, 22–34. doi:10.1017/CBO9780511802478.002. Wasserstein, R. L., Lazar, N. A. (2016): The ASA’s statement on 𝑝 -values: context, process, and purpose. Am. Stat. 70 2 , 129–133. doi: 10.1080/00031305.2016.1154108 . See asa 2016 and discussion in Greenland et al. 2016. Weinstein, M. C., Fineberg, H. V. (1980): Clinical Decision Analysis. (Saunders, Philadelphia). Weiss, G. M., Provost, F. (2003): Learning when training data are costly: the effect of class distribution on tree induction. J. Artif. Intell. Res. 19, 315–354. doi:10.1613/jair.1199. Woodward, P. M. (1964): Probability and Information Theory, with Applications to Radar, 2nd ed. (Pergamon, Oxford). doi:10.1016/C2013-0-05390-X. First publ. 1953. Zalta, E. N., ed. (2023): Stanford Encyclopedia of Philosophy, continuously updated. (The Metaphysics Research Lab). https://plato.stanford.edu. First publ. 1995. 53