Bernstein flows for flexible posteriors in variational Bayes
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Dürr, Oliver; Hörtling, Stefan; Dold, Danil; Kovylov, Ivonne; Sick, Beate Article — Published Version Bernstein flows for flexible posteriors in variational Bayes AStA Advances in Statistical Analysis Provided in Cooperation with: Springer Nature Suggested Citation: Dürr, Oliver; Hörtling, Stefan; Dold, Danil; Kovylov, Ivonne; Sick, Beate (2024) : Bernstein flows for flexible posteriors in variational Bayes, AStA Advances in Statistical Analysis, ISSN 1863-818X, Springer, Berlin, Heidelberg, Vol. 108, Iss. 2, pp. 375-394, https://doi.org/10.1007/s10182-024-00497-z This Version is available at: https://hdl.handle.net/10419/315187 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by/4.0/
Vol.:(0123456789) AStA Advances in Statistical Analysis (2024) 108:375–394 https://doi.org/10.1007/s10182-024-00497-z 1 3 ORIGINAL PAPER Bernstein flows forflexible posteriors invariational Bayes OliverDürr1 · StefanHörtling1· DanilDold1· IvonneKovylov1· BeateSick2,3 Received: 18 November 2022 / Accepted: 6 February 2024 / Published online: 3 April 2024 © The Author(s) 2024 Abstract Black-box variational inference (BBVI) is a technique to approximate the posterior of Bayesian models by optimization. Similar to MCMC, the user only needs to specify the model; then, the inference procedure is done automatically. In contrast to MCMC, BBVI scales to many observations, is faster for some applications, and can take advantage of highly optimized deep learning frameworks since it can be formulated as a minimization task. In the case of complex posteriors, however, other state-of-the-art BBVI approaches often yield unsatisfactory posterior approximations. This paper presents Bernstein flow variational inference (BFVI), a robust and easy-to-use method flexible enough to approximate complex multivariate posteriors. BF-VI combines ideas from normalizing flows and Bernstein polynomial-based transformation models. In benchmark experiments, we compare BF-VI solutions with exact posteriors, MCMC solutions, and state-of-the-art BBVI methods, including normalizing flow-based BBVI. We show for low-dimensional models that BF-VI accurately approximates the true posterior; in higher-dimensional models, BF-VI compares favorably against other BBVI methods. Further, using BF-VI, we develop a Bayesian model for the semi-structured melanoma challenge data, combining a CNN model part for image data with an interpretable model part for tabular data, and demonstrate, for the first time, the use of BBVI in semistructured models. Keywords Variational inference· Deep learning· Transformation models· Bayesian neural network 1 Introduction Uncertainty quantification is essential, especially if model predictions are used to support high-stakes decision-making. Quantifying uncertainty in statistical or machine learning models is often achieved by Bayesian approaches, where posterior distributions represent the uncertainty of the estimated model parameters. Oliver Dürr and Beate Sick have contributed equally to this work. Extended author information available on the last page of the article
376 O.Dürr et al. 1 3 Determining the exact posterior distributions is often impossible when the posterior takes a complex shape and the model has many parameters. This is especially true for complex models such as Bayesian neural networks (NNs) or semi-structured models that combine an interpretable model part with deep NNs. Variational inference (VI) is a commonly used approach to approximate complex distributions through optimization (Jordan etal. 1999; Blei etal. 2017). In VI, the complex posterior is approximated by a variational distribution by minimizing a divergence measure between the variational and the true posterior distribution. VI is currently a very active research field tackling different challenges, which can be categorized into the following groups: (1) constructing variational distributions that are flexible enough to match the true posterior distribution, (2) defining optimal variational objective for tuning the variational distribution, which boils down to finding the most suited divergence measure quantifying the difference between a variational distribution and posterior, and (3) developing robust and accurate stochastic optimization frameworks for the variational objective (Dhaka etal. 2020; Blei etal. 2016; Welandawe etal. 2022). Here, we focus on challenge (1) and propose a method to construct a variational distribution that is flexible enough to accurately and robustly approximate complex multidimensional posteriors. To avoid model-specific calculations, we design our method as a blackbox VI (BBVI) approach (Ranganath et al. 2014). In BBVI, the approximative posterior is determined by stochastic gradient descent. The user simply defines the Bayesian model by specifying the likelihood and the prior, after which all subsequent calculations are carried out automatically. Due to its simplicity, BBVI is implemented in many packages for Bayesian modeling, like Stan (Carpenter etal. 2017) and Pyro (Bingham etal. 2019) as an alternative to MCMC. Given BBVI’s scalability to large datasets and its widespread applicability, it has emerged as the preferred technique in the field of machine learning (Welandawe etal. 2022). Our approach uses transformation models (TMs) to construct complex posteriors. Transformation models (TMs) have been introduced for fitting potentially complex outcome distributions for probabilistic regression models (Hothorn et al. 2014). Since then, they have been mainly used to model different outcome types, such as ordinal (Kook et al. 2022; Buri et al. 2020), count (Siegfried and Hothorn 2020), continuous (Lohse et al. 2017), or time-to-event outcomes (Campanella etal. 2022) based on tabular predictors. Moreover, TMs have been used to model multidimensional distributions (Klein etal. 2019). Neural networks can be used to extend TMs to model outcomes for unstructured predictors (e.g., images or text) or a combination of tabular and unstructured predictors (Sick etal. 2021; Baumann etal. 2021; Kook etal. 2022; Rügamer etal. 2021). The basic idea of TMs is to learn a flexible and monotone transformation function that transforms between a simple latent distribution and a potentially complex conditional outcome distribution. In TMs, the transformation function is parameterized as an expansion of basis functions. In the case of continuous target distributions, most often, Bernstein polynomials (Bernšteın 1912) are used because they can easily be constrained to be strictly monotone, and their flexibility can be tuned via the order M. A large order M ensures an accurate approximation of the distribution, which is robust against a further increase of M (Hothorn etal. 2018;
377 1 3 Bernstein flows forflexible posteriors invariational Bayes Ramasinghe etal. 2021); this is also demonstrated in our experiments for the BBVI setting. Independently of TMs, normalizing flows (NFs) have been developed in the deep learning community. NFs and TMs rely on the same idea, but NFs usually construct the transformation by chaining many simple functions, while TMs construct one rather complex transformation function. In NFs, each simple function, such as shifting and scaling, incrementally adds to the complexity of the final transformation. Among the prominent NF implementations are RealNVP (Dinh et al. 2016) and Masked Autoregressive Flow (MAF) (Papamakarios etal. 2017). RealNVP stands out for its efficient, invertible transformations facilitated by a specialized neural network architecture. Its key advantage lies in the efficient computation of the Jacobian matrix’s determinant, essential for direct density estimation in the change of variable function (see6). This efficiency is achieved by iteratively splitting the components of the data into two parts. In each step, the first part of the components is used to train a neural network computing the scale and shift parameters of the transformation. This transformation is then applied to the other components, while the first part remains unaltered. This procedure is repeated multiple times with different partitioning, leading to a triangular Jacobian, thus enabling efficient and invertible transformations. In contrast, MAF adopts a fundamentally different approach to construct transformations (Papamakarios et al. 2017). It utilizes a sequential (autoregressive) framework, facilitated by neural networks. In MAF each output component relies exclusively on its preceding components, a concept often referred to as causality in this context. This design also leads to a triagonal Jacobian matrix and thus a fast computation of the change of variable equation. The MAF ensures that the nth output of NN is solely dependent on the first n−1 inputs, yielding an autoregressive model. However, some NF approaches use a single flexible transformation, such as sum-of-squares polynomials (Jaini etal. 2019) or splines (Durkan etal. 2019). Recently, also Bernstein-based polynomials have been used for modeling unconditional multivariate density distributions (Ramasinghe etal. 2021). NFs were initially introduced for variational inference to approximate potentially complex distributions of latent variables in models such as variational autoencoders (Rezende and Mohamed 2015; Van DenBerg etal. 2018). In the past, often members from simple distribution families have been used to approximate the posterior in BBVI. In the "Bayes by Backprop" method, Blundell et al. used independent Gaussians to approximate the posterior of the weights in a Bayesian Neural Network (BNN). They determined the parameters of these Gaussians using BBVI (Blundell etal. 2015). This approach was made more flexible by using a multivariate Gaussians (Louizos and Welling 2017) as variational distribution. While it is clear that TMs or NFs have the potential to construct flexible variational distributions, the first attempts to use NF-based BBVI were proposed only recently (Agrawal etal. 2020). These NF-based BBVI approaches compare favorably against existing BBVI methods but require a complex training scheme and sometimes exhibit pathological behavior (Dhaka etal. 2021). Here, we introduce Bernstein flow variational inference (BF-VI), which, for the first time, uses TMs in BBVI. We use TMs based on Bernstein polynomials to construct a
378 O.Dürr et al. 1 3 variational distribution that closely approximates a potentially complex posterior in Bayesian models. The proposed method is computationally efficient and applicable to typical statistical models. The proposed method yields superior results in our experiments compared to existing NF approaches (Dhaka etal. 2021). Using BF-VI we further demonstrate, for the first time, how VI can be used to fit Bayesian semi-structured models where interpretable statistical model parts (based on tabular data) and deep NN model parts (based on images) are jointly fitted. We define our method in Sect.2.1 for one-dimensional examples and generalize it to Bayesian models with multivariate posteriors in Sect.2.2. In Sect.3, we benchmark our BF-VI approach against exact Bayesian models, MCMC-Simulations, Gaussian-VI, and NF-based BBVI, showing accurate posterior approximations in low dimensions and superior approximations in higher dimensions when compared to NF-based BBVI and summarize in Sect.4. 2 Bernstein flow variational inference In the following, we describe the Bernstein Flow-VI (BF-VI) approach, which we propose for accurately and robustly approximating potentially complex posteriors in Bayesian models. The main idea is to enable the VI procedure to approximate the joined posterior of the p model parameters by a flexible variational distribution. This is done by modeling the transformation function from a predefined simple latent distribution to a potentially complex variational distribution. The number of parameters p in the Bayesian model determines the dimension of both the latent and the variational distribution. We first explain BF-VI for Bayesian models with a single parameter and hence a one-dimensional posterior and then generalize to models with multivariate posteriors. The code is publicly available on GitHub.1 2.1 One‑dimensional Bernstein flows BF-VI approximates the bijective transformation function g∶Z → 𝜃 between a latent variable Z∈ℝ with predefined distribution FZ∶ℝ → [0, 1] with log-concave and continuous density fZ , and the model parameter 𝜃∈ℝ with a potentially complex distribution F𝜃∶ℝ → [0, 1] so that FZ(z)=F𝜃(g(z)) . Figure1 visualizes this transformation on the scale of the densities, where fZ(z)=f𝜃(g(z)) ∣ 𝜕g(z) 𝜕z∣ according to the change-of-variable formula. Hothorn etal. (2018) give theoretical guarantees for the existence and uniqueness of g =F −1 𝜃 ◦F Z . However, g cannot be computed directly if F𝜃 is not known (in our application F𝜃 is the unknown distribution of the posterior). The core of BF-VI is to approximate g, shown in Fig.1, by fBP Bernstein polynomials (BP)2 as 1 https:// github. com/ tenso rchie fs/ bfvi_ paper. 2 Some authors make a distinction between Bernstein polynomials in which 𝜗i is fixed by the values of the function to be approximated and call expressions like expressions like in (1), where 𝜗i is a fitting parameter, polynomials of Bernstein type.
379 1 3 Bernstein flows forflexible posteriors invariational Bayes with Be i(z)= Be (i+1,M−i+1)(z) being the density of a Beta distribution with parameters i+1 and M−i+1 . To preserve the bijectivity of g, we use in BF-VI w.l.o.g. a strict monotone increasing BP to approximate g. With the approximation of the transformation function g by fBP , it holds that fZ(z) can be approximated by f𝜃(fBP(z)) ∣ 𝜕f BP (z) 𝜕z∣ . Using a BP for approximating g gives the following theoretical guarantees (Farouki 2012) (1) With increasing order M of the BP, the approximation fBP to g gets arbitrarily close (the BP have been introduced for this very purpose in the constructive proof of the Weierstrass theorem by Bernšteın (1912)); (2) the required (1) fBP(z)= M ∑ i=0 Be i(z) 𝜗i M+ 1 Fig. 1 Overview of the transformation model. A Shows the bijective transformation function g∶Z → 𝜃 (or its approximation fBP ) mapping between, B a predefined latent density fZ , and C a potentially complex posterior (or its variational distribution)
380 O.Dürr et al. 1 3 strict monotonicity of the approximation fBP can be easily achieved by constraining the coefficients 𝜗i of the BP to be increasing; (3) BPs are robust versus perturbations of the coefficients 𝜗i ; (4) the approximation error decreases with 1/M (Voronovskaya’s theorem). See Bernšteın (1912); Farouki (2012) for more detailed discussions of the beneficial properties of BPs in general and Hothorn etal. (2018); Ramasinghe etal. (2021) for transformation models. While the output of fBP(z) is unrestricted, a BP requires a input z within [0,1]. We experimented with several approaches to ensure the restriction z∈[0, 1] resulting in slightly different behavior during the training (see AppendixB.2.2). Based on these experiments, we decided to obtain z∈[0, 1] by sampling values from a standard normal distribution, z�� ∼N(0, 1) , then apply the affine transformation l(z��)=𝛼 ⋅ z�� +𝛽 , followed by a sigmoid 𝜎(z�)=1∕(1+e− z �) . Altogether, we approximate the transformation g by f∶Z → 𝜃 via f=fBP ◦ 𝜎 ◦ l , which we call Bernstein flow. To allow the application of unconstrained stochastic gradient descent optimization, which is typically used in the deep learning domain, we enforce the strict monotonicity of the flow f as follows: We optimize unrestricted parameters of f, i.e., 𝜗� 0 ,…𝜗 � M , 𝛼′ ,𝛽′ , and apply the following transformations to determine the parameters of the bijective flow: 𝜗0 =𝜗 � 0 , and 𝜗i =𝜗 i−1 +softplus (𝜗 � i) for i=1, …,M for getting a strictly increasing BP and 𝛼=softplus (𝛼�) , 𝛽=𝛽� for getting an increasing affine transformation. In Appendix A, we show that the resulting variational distribution is a tight approximation to the posterior in the sense that the KL divergence between q𝜆(𝜃) and p(𝜃∣D) decreases with the order of the BP via 1/M. 2.2 Multivariate generalization In the case of a Bayesian model with p parameters, 𝜃1,𝜃2,…,𝜃p , the Bernstein flow bijectively maps a p-dimensional Z′ to a p-dimensional 𝜽 . We realize this flow by choosing p independent standard normal Gaussians as simple latent distribution for the p-dimensional Z′ and apply on each component an affine transformation followed by a sigmoid function to achieve a [0,1] restricted Z . The possible complex dependencies in 𝜽 are modeled in the multivariate generalization fBP of the one-dimensional Bernstein polynomial (see Eq.2 for the definition of the jth component of fBP ). To achieve an efficient computation, we use a triangular map for constructing coefficients 𝜗j i j=2, …p,i=0, ⋯ ,M from Z . This ensures that the jth BP determining 𝜃j only depends on the first j-1 components of Z (see Eq.2). It is known that bijective triangular maps with sufficient flexibility can map a simple p-dimensional distribution into arbitrary complex p-dimensional target distributions (Bogachev etal. (2) 𝜃 j=fBPj(z1∶j)= 1 M+1 M ∑ i=0 𝜗j i(z1,…,zj−1)Be i(zj )
381 1 3 Bernstein flows forflexible posteriors invariational Bayes 2005). We use a masked autoregressive flow (MAF) (Papamakarios etal. 2017) to map Z to the BP coefficients 𝜗j i j=2, …p,i=0, ⋯ ,M from Z . The MAF architecture ensures that that 𝜗j i depend only on those components of the latent variables zj′ with j′≤j (as required in Eq.2). Note that the first coefficients in all BPs 𝝑1 do not depend on z and are therefore not modeled via the MAF. Therefore, the Jacobian ∇fBP w.r.t. z is a triangular matrix, and hence det ∇fBP is given by the product of the diagonal elements of the Jacobian allowing for efficient computation of the resulting p-dimensional variational distribution q𝝀(𝜽) via the multivariate version of the change of variable formula (see Eq.6). The flexibility of such a p-dimensional bijective Bernstein flow is only limited by the order M of the Bernstein polynomial and the complexity of the MAF. In our experiments, we use an MAF with two hidden layers, each with 10 neurons. The weights w of the MAF are part of the variational parameters for fBP . In total, we have 𝝀=(𝝑1, w ,𝜶,𝜷) variational parameters. 2.3 Variational inference procedure In VI the variational parameters 𝝀 are tuned such that the resulting variational distribution q𝝀(𝜽) is as close to the posterior p(𝜽∣D) as possible. Here, we do this by minimizing the KL divergence between the variational distribution and the (unknown) posterior: The KL divergence is commonly used in VI, and a recent study showed that it is easier to train than other divergences and applicable to higher-dimensional distributions (Dhaka etal. 2021). Instead of minimizing (3) usually only the evidence lower bound (ELBO) is maximized (Blundell etal. 2015) which consists of the expected value of the log-likelihood, 𝔼𝜽∼q𝝀(log(p(D∣𝜽))) , minus the KL divergence between the variational distribution q𝝀(𝜽) and the prior p(𝜽) . Note that the ELBO does not explicitly contain the unknown posterior. In practice, we minimize the negative ELBO using stochastic gradient descent facilitated by automatic differentiation. For consistency with Dhaka etal. (2021), we use TensorFlow’s RMSprop in all our experiments, configured with the default settings. We follow the BBVI approach and approximate the expected log-likelihood by averaging over S samples 𝜽s∼q𝝀(𝜽) via (3) KL (q𝝀(𝜽)||p(𝜽∣D)) = ∫q𝝀(𝜽)log ( q𝝀 (𝜽) p(𝜽∣D) ) d𝜽 =log(p(D)) − (𝔼𝜽∼q𝝀(log(p(D∣𝜽))) − KL(q𝝀(𝜽) || p(𝜽)) ) ⏟⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏟⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞ ⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏞⏟ ELBO (𝝀) (4) 𝔼 𝜽∼q𝝀(log(p(Di∣𝜽))) ≈ 1 S ∑ s,i log ( p(Di∣𝜽s) ).
382 O.Dürr et al. 1 3 Hereby, we also assume the usual independence of the i=1, …N training data points Di . To get these samples 𝜽s , we use S samples z′ s from the latent distribution, apply the transformation f=fBP ◦ 𝜎 ◦ l , and then compute the corresponding parameter samples via 𝜽s=f(zs) . We use the same samples 𝜽s∼q𝝀(𝜽) to approximate the Kullback–Leibler divergence between the variational distribution q𝝀(𝜽) and the prior p(𝜽) via: where the probability density q𝝀(𝜽s) can be calculated, from the samples 𝜽s using the change of variable function as: 2.4 Evaluation Evaluating the quality of the fitted variational distributions requires a comparison to the true posterior. In the case of low-dimensional problems, the two distributions can be compared visually. In the case of higher-dimensional problems, this is not possible anymore. While the evidence lower bound (ELBO) is a valuable metric for optimizing the parameters in VI, it is less helpful in comparing different approximations because it depends on the specific parametrization of the model (Yao etal. 2018). Therefore, Yao etal. (2018) introduced k as a more suited approach for comparison, which since then has been used in other studies like (Dhaka etal. 2021) to which we compare. The computation of k is based on the importance ratios which are defined as If the variational distribution q𝝀(𝜽) would be a perfect approximation of the posterior p(𝜽∣D)∝p(D∣𝜽)p(𝜽) , then important ratios rs would be constant. However, because of the asymmetry of the KL divergence used in the optimization objective (see Eq.3), the fitted q𝝀(𝜽) tends to have lighter tails than p(𝜽∣D) , with the effect that the distribution of rs is heavily right-tailed. To quantify the severity of the underestimated tails, a generalized Pareto distribution is fitted to the right tail of the rs . The estimated shape parameter k of the Pareto distribution can be used as a diagnostic tool. A large k indicates a pronounced tail in the rs distribution and, hence, a bad posterior approximation. According to Yao etal. (2018) values of k<0.5 indicate that the variational approximation q𝝀 is good. Values of 0.5 < k<0.7 indicate the variational approximation q𝝀 is not perfect but still useful. (5) KL (q𝝀(𝜽) || p(𝜽)) ≈ 1 S ∑ s log ( q𝝀 (𝜽 s ) p(𝜽 s ) ) (6) q𝝀 (𝜽 s )=p(z � s )⋅∣det ∇z�f BP (𝜎(l(z � s ))) ∣ −1 (7) r s= p(𝜽 s ,D) q 𝝀( 𝜽 s) = p(D∣𝜽 s )p(𝜽 s ) q 𝝀( 𝜽 s)
389 1 3 Bernstein flows forflexible posteriors invariational Bayes NNs and tabular data by interpretable model components. Please note that here, both the conditional distribution of the outcome (y∣B,x) and the unconditional posterior of the parameters are modeled by transformation models. Because of the deep NN model components involved, MCMC is not feasible anymore to determine the posterior. As a dataset, we use the SIIM-ISIC Melanoma Classification Challenge4 data. The data come from 33,126 patients (6626 as test set, 21,200 train, and 5300 validation set) with a confirmed diagnosis of their skin lesions, which is in ≈98 % benign ( y=0 ) and in ≈2 % malignant ( y=1 ). The provided data D=(B,x) are semi-structured since it comprises (unstructured) image data B from the patient’s lesion along with (structured) tabular data x, i.e., the patient’s age. We fit the conditional outcome distribution ( Y∣D)∼ Ber ( 𝜋 D) by modeling the probability for a lesion to be malignant 𝜋D= p ( y =1∣ D )=𝜎( h ) applying the sigmoid function 𝜎( ⋅ ) to a fitted transformation function h∶Y → Z . We study three models for h depending on x alone, B alone, and in combination B and x: M1 (DL-Model) h=𝜇(B) : As a baseline, we use a deep convolutional neural network (CNN) based on the melanoma image data (see Fig. 6c) with a total of 419,489 weights to take advantage of the predictive power of DL on complex image data. For this DL model, we use deep ensembling (Lakshminarayanan etal. 2017) by fitting three CNN models with different random initializations and averaging the Fig. 6 The architecture of the used NN models or model parts. a Dense NN with one hidden layer to model nonlinear dependencies from the input (used in the NN-based nonlinear regression example). b Dense NN without hidden layer to model linear dependencies from tabular input data (used in M1 and M3 of the melanoma experiment). c CNN to model nonlinear dependencies from the image input (used in M1 and M3) 4 https:// chall enge2 020. isicarchi ve. com
390 O.Dürr et al. 1 3 predicted probabilities. The achieved test predictive performance and its comparison to other models are discussed in the last paragraph of this section. M2 (Logistic Regression) h=𝜇0+𝛽1 ⋅ x : When using only tabular features x, interpretable models can be built. We consider a Bayesian logistic regression with age as the only explanatory variable x and use a BNN without a hidden layer to set up the model (see Fig.6b with only one input feature x). In logistic regression, a latent variable is modeled by a linear predictor h=𝜇0+𝛽1 ⋅ x , which determines the probability for a lesion to be malignant via 𝜋x =𝜎 ( 𝜇 0 +𝛽 1 ⋅x ) allowing to inter et e𝛽1 as the odds ratio, i.e., the factor by which the odds for lesions to be malignant changes when increasing the predictor x by one unit. In Fig.7, we compare the exact MCMC posterior of 𝛽1 with the BF-VI approximation, demonstrating that BF-VI accurately approximates the posterior. M3 (semi-structured) h=𝜇(B)+𝛽1 ⋅ x : This model integrates image and tabular data and combines the predictive power of M1 with the interpretability of M2. We use a (non-Bayesian) CNN that determines 𝜇(B) and BF-VI for the NN without a hidden layer that determines 𝛽1 (see Fig.6b and c). Both NNs are jointly trained by optimizing the ELBO. The resulting posterior for 𝛽1 differs from the simple logistic regression (see Fig.7), indicating a diminished effect of age after including the image. Again, e𝛽1 can be interpreted as the factor by which the odds for a lesion to be malignant change when increasing the predictor age by one unit and holding the image constant. While the main interest of our study is on the posteriors, we also determine the predictive performance on the test set. To quantify and compare the test prediction performances, we look at the achieved log scores (M1: − 0.076, M2: − 0.085, M3: − 0.076) and the AUCs with 95 % CI (M1: 0.83(0.79,0.86), M2: 0.66(0.61, 0.71), M3: 0.82(0.79,0.85)). For both measures, higher is better. Interestingly, the imagebased models (M1, M3) have higher predictive power than M2, which only uses tabular data. The semi-structured model M3, including tabular and image information, has a similar predictive power compared to M1, which only uses images. The benefit of the semi-structured model here is that it provides interpretable parameters for the tabular data along with uncertainty quantification without losing predictive performance. Fig. 7 Posteriors for the ageeffect parameter 𝛽1 in the melanoma models M2 and M3 0 2 4 6 0.0 0. 40 .8 β 1 density M2: BF−VI M2: MCMC M3: CNN+BF−VI
391 1 3 Bernstein flows forflexible posteriors invariational Bayes 4 Summary andoutlook The proposed BF-VI is flexible enough to approximate any posterior in principle without being restricted to variational distributions with known parametric distribution families like Gaussians. In benchmark experiments, BF-VI accurately fits non-trivial posteriors in low-dimensional problems in a BBVI setting. For higher-dimensional models, BF-VI outperforms published results from other NF-VI methods (Dhaka etal. 2021) on the studied benchmark datasets. Still, we observe that the posterior cannot be fitted perfectly in high dimensions by BF-VI, especially since the tails of the approximation are too short. We attribute this limitation to known difficulties in the optimization process and the asymmetry of the KL divergence. These challenges of VI were not in the focus of our study, and we leave it to further research. To the best of our knowledge, we are the first to demonstrate how BBVI can be used in semi-structured models. We used BF-VI on the public melanoma challenge dataset, integrating image data and tabular data by combining a deep CNN and an interpretable model part. We see a valuable application of BF-VI in models with interpretable parameters, i.e., statistical or semi-structured models where we can model complex posterior distributions of the interpretable parameters. Especially in semi-structured models with deep NN components that cannot be fitted with MCMC, BF-VI allows determining the variational distribution for the interpretable model parts. Moreover, efficient SGD optimizers can be used in BF-VI to fit all model parts jointly. We plan to extend our research on BF-VI for semi-structured models in the future and investigate the quality of the posterior approximations. Supplementary Information The online version contains supplementary material available at https:// doi. org/ 10. 1007/ s1018202400497-z. Acknowledgements The research of BS was supported by the Novartis Research Foundation (FreeNovation2019). The research of the OD and DD was partly supported by the Federal Ministry of Education and Research of Germany (BMBF) in the project DeepDoubt (grant no. 01IS19083A). We, further, would like to thank Nadja Klein, Lucas Kook, and Rebekka Axthelm for fruitful discussions. Funding Open Access funding enabled and organized by Projekt DEAL. This study was funded by Bundesministerium für Bildung, Wissenschaft, Forschung und Technologie (01IS19083A), Novartis Stiftung für Medizinisch-Biologische Forschung. Declarations Conflict of interest None. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line
392 O.Dürr et al. 1 3 to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/ licenses/by/4.0/. References Agrawal, A., Sheldon, D.R., Domke, J.: Advances in black-box VI: normalizing flows, importance weighting, and optimization. Adv. Neural Inf. Process. Syst. 33, 17358–17369 (2020) Baumann, P.F.M., Hothorn, T., Rügamer, D.: Deep conditional transformation models. In: Machine Learning and Knowledge Discovery in Databases. Research Track, pp. 3–18. Springer, Cham (2021) Bernšteın, S.: Démonstration du théoreme de weierstrass fondée sur le calcul des probabilities. Commun. Soc. Math. Kharkov 13, 1–2 (1912) Bingham, E., Chen, J.P., Jankowiak, M., Obermeyer, F., Pradhan, N., Karaletsos, T., Singh, R., Szerlip, P.A., Horsfall, P., Goodman, N.D.: Pyro: deep universal probabilistic programming. J. Mach. Learn. Res. 20, 28–1286 (2019) Blei, D., Ranganath, R., Mohamed, S.: Variational inference: foundations and modern methods. In: NIPS tutorial (2016) Blei, D.M., Kucukelbir, A., McAuliffe, J.D.: Variational inference: a review for statisticians. J. Am. Stat. Assoc. 112(518), 859–877 (2017) Blundell, C., Cornebise, J., Kavukcuoglu, K., Wierstra, D.: Weight uncertainty in neural network. In: International Conference on Machine Learning, pp. 1613–1622. PMLR (2015) Blundell, C., Cornebise, J., Kavukcuoglu, K., Wierstra, D.: Weight uncertainty in neural network. In: International Conference on Machine Learning, pp. 1613–1622. PMLR (2015) Bogachev, V.I., Kolesnikov, A.V., Medvedev, K.V.: Triangular transformations of measures. Sbornik Math. 196(3), 309 (2005) Buri, M., Curt, A., Steeves, J., Hothorn, T.: Baseline-adjusted proportional odds models for the quantification of treatment effects in trials with ordinal sum score outcomes. BMC Med. Res. Methodol. 20, 1–14 (2020) Campanella, G., Kook, L., Häggström, I., Hothorn, T., Fuchs, T.J.: Deep conditional transformation models for survival analysis (2022). arXiv preprint arXiv: 2210. 11366 Carpenter, B., Gelman, A., Hoffman, M.D., Lee, D., Goodrich, B., Betancourt, M., Brubaker, M., Guo, J., Li, P., Riddell, A.: Stan: a probabilistic programming language. J. Stat. Softw. (2017). https:// doi. org/ 10. 18637/ jss. v076. i01 Dhaka, A.K., Catalina, A., Andersen, M.R., Magnusson, M., Huggins, J., Vehtari, A.: Robust, accurate stochastic optimization for variational inference. Adv. Neural Inf. Process. Syst. 33, 10961–10973 (2020) Dhaka, A.K., Catalina, A., Welandawe, M., Andersen, M.R., Huggins, J., Vehtari, A.: Challenges and opportunities in high dimensional variational inference. Adv. Neural Inf. Process. Syst. 34, 7787– 7798 (2021) Dinh, L., Sohl-Dickstein, J., Bengio, S.: Density estimation using real NVP. In: International Conference on Learning Representations (2016) Durkan, C., Bekasov, A., Murray, I., Papamakarios, G.: Neural spline flows. Adv. Neural Inf. Process. Syst. 32, 7511–7522 (2019) Farouki, R.T.: The Bernstein polynomial basis: a centennial retrospective. Comput. Aided Geom. Des. 29(6), 379–419 (2012) Hothorn, T., Kneib, T., Bühlmann, P.: Conditional transformation models. J. R. Stat. Soc. Ser. B Stat. Methodol. 76(1), 3–27 (2014) Hothorn, T., Moest, L., Buehlmann, P.: Most likely transformations. Scand. J. Stat. 45(1), 110–134 (2018) Huggins, J., Kasprzak, M., Campbell, T., Broderick, T.: Validated variational inference via practical posterior error bounds. In: International Conference on Artificial Intelligence and Statistics, pp. 1792– 1802. PMLR (2020)
393 1 3 Bernstein flows forflexible posteriors invariational Bayes Jaini, P., Selby, K.A., Yu, Y.: Sum-of-squares polynomial flow. In: International Conference on Machine Learning, pp. 3009–3018. PMLR (2019) Jordan, M.I., Ghahramani, Z., Jaakkola, T.S., Saul, L.K.: An introduction to variational methods for graphical models. Mach. Learn. 37(2), 183–233 (1999) Klein, N., Hothorn, T., Barbanti, L., Kneib, T.: Multivariate conditional transformation models. Scand. J. Stat. 49(1), 116–142 (2019) Kook, L., Herzog, L., Hothorn, T., Dürr, O., Sick, B.: Deep and interpretable regression models for ordinal outcomes. Pattern Recognit. 122, 108263 (2022) Lakshminarayanan, B., Pritzel, A., Blundell, C.: Simple and scalable predictive ncertainty estimation using deep ensembles. In: Advances in Neural Information Processing Systems, vol. 30 (2017) Lohse, T., Rohrmann, S., Faeh, D., Hothorn, T.: Continuous outcome logistic regression for analyzing body mass index distributions. F1000Research 6, 1933 (2017) Louizos, C., Welling, M.: Multiplicative normalizing flows for variational Bayesian neural networks. In: International Conference on Machine Learning, pp. 2218–2227. PMLR (2017) Magnusson, M., Bürkner, P., Vehtari, A.: PosteriorDB: a set of posteriors for Bayesian inference and probabilistic programming Papamakarios, G., Pavlakou, T., Murray, I.: Masked autoregressive flow for density estimation. In: Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 2335– 2344 (2017) Ramasinghe, S., Fernando, K., Khan, S., Barnes, N.: Robust normalizing flows using Bernstein-type polynomials (2021). arXiv preprint arXiv: 2102. 03509 Ranganath, R., Gerrish, S., Blei, D.: Black box variational inference. In: Artificial Intelligence and Statistics, pp. 814–822. PMLR, Cambridge (2014) Rezende, D., Mohamed, S.: Variational inference with normalizing flows. In: International Conference on Machine Learning, pp. 1530–1538. PMLR (2015) Rügamer, D., Baumann, P.F.M., Kneib, T., Hothorn, T.: Probabilistic time series forecasts with autoregressive transformation models. arXiv preprint (2021) Sick, B., Hathorn, T., Dürr, O.: Deep transformation models: tackling complex regression problems with neural network based transformation models. In: 2020 25th International Conference on Pattern Recognition (ICPR), pp. 2476–2481. IEEE (2021) Siegfried, S., Hothorn, T.: Count transformation models. Methods Ecol. Evol. 11(7), 818–827 (2020). https:// doi. org/ 10. 1111/ 2041210X. 13383 Van DenBerg, R., Hasenclever, L., Tomczak, J.M., Welling, M.: Sylvester normalizing flows for variational inference. In: 34th Conference on Uncertainty in Artificial Intelligence 2018, UAI 2018, pp. 393–402. Association For Uncertainty in Artificial Intelligence (AUAI) (2018) Welandawe, M., Andersen, M.R., Vehtari, A., Huggins, J.H.: Robust, automated, and accurate black-box variational inference (2022). arXiv preprint arXiv: 2203. 15945 Yao, Y., Vehtari, A., Simpson, D., Gelman, A.: Yes, but did it work?: Evaluating variational inference. In: International Conference on Machine Learning, pp. 5581–5590. PMLR (2018) Yao, Y., Vehtari, A., Gelman, A.: Stacking for non-mixing Bayesian computations: the curse and blessing of multimodal posteriors. J. Mach. Learn. Res. 23, 1–45 (2022) Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
394 O.Dürr et al. 1 3 Authors and Affiliations OliverDürr1 · StefanHörtling1· DanilDold1· IvonneKovylov1· BeateSick2,3 * Oliver Dürr [email protected] * Beate Sick beate.sic[email protected]; [email protected] Stefan Hörtling Stefanhoer[email protected] Danil Dold [email protected] Ivonne Kovylov iv[email protected] 1 IOS, Konstanz University ofApplied Sciences, Alfred Wachtel Straße 8, 78462Konstanz, Germany 2 IDP, Zurich University ofApplied Sciences, Technikumstrasse 81, 8401Winterthur, Switzerland 3 EBPI, University ofZurich, Hirschengraben 84, 8001Zurich, Switzerland