Testing covariance separability for continuous functional data
Abstract
EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.
Full text
Dette, Holger; Dierickx, Gauthier; Kutta, Tim Article — Published Version Testing covariance separability for continuous functional data Journal of Time Series Analysis Provided in Cooperation with: John Wiley & Sons Suggested Citation: Dette, Holger; Dierickx, Gauthier; Kutta, Tim (2024) : Testing covariance separability for continuous functional data, Journal of Time Series Analysis, ISSN 1467-9892, John Wiley & Sons, Ltd, Oxford, UK, Vol. 46, Iss. 3, pp. 402-420, https://doi.org/10.1111/jtsa.12764 This Version is available at: https://hdl.handle.net/10419/319337 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. http://creativecommons.org/licenses/by-nc-nd/4.0/
JOURNAL OF TIME SERIES ANALYSIS J. Time Ser. Anal. 46: 402–420 (2025) Published online 11 August 2024 in Wiley Online Library (wileyonlinelibrary.com) DOI: 10.1111/jtsa.12764 ORIGINAL ARTICLE TESTING COVARIANCE SEPARABILITY FOR CONTINUOUS FUNCTIONAL DATA HOLGER DETTE GAUTHIER DIERICKX AND TIM KUTTA Department of Mathematics, Ruhr-University Bochum, Bochum, Germany Analyzing the covariance structure of data is a fundamental task of statistics. While this task is simple for low-dimensional observations, it becomes challenging for more intricate objects, such as multi-variate functions. Here, the covariance can be so complex that just saving a non-parametric estimate is impractical and structural assumptions are necessary to tame the model. One popular assumption for space-time data is separability of the covariance into purely spatial and temporal factors. In this article, we present a new test for separability in the context of dependent functional time series. While most of the related work studies functional data in a Hilbert space of square integrable functions, we model the observations as objects in the space of continuous functions equipped with the supremum norm. We argue that this (mathematically challenging) setup enhances interpretability for users and is more in line with practical preprocessing. Our test statistic measures the maximal deviation between the estimated covariance kernel and a separable approximation. Critical values are obtained by a non-standard multiplier bootstrap for dependent data. We prove the statistical validity of our approach and demonstrate its practicability in a simulation study and a data example. Received 11 January 2023; Accepted 26 June 2024 Keywords: Dependent multiplier bootstrap; space-time data; separability; Banach space; functional time series. MSC subject classification: 62G10; 62R10. 1. INTRODUCTION Over the last decades, the analysis of high-dimensional space-time data has become a cornerstone of geostatistics. New technologies allow the collection of high-frequency and high-resolution measurements for variables such as temperature, magnetic fields or pollutant concentrations (see, for example, Gromenko et al. (2012); Aue et al. (2018); King et al. (2018)). One way to analyze such data is to smooth it over space and time, which yields time series of spatio-temporal processes. This approach of reconstructing and analyzing random functions follows the paradigm of functional data analysis (FDA) and has recently gained attraction in the geostatistical community (for overviews see, e.g., Martínez-Hernández and Genton (2020) and the monograph of Mateu and Giraldo (2022)). While the analysis of such processes promises deep scientific insights, their high complexity can push statistical methods to their limit. For example, commonly used tools such as PCA and Kriging hinge on an approximation of the processes’ covariance operator – an object that can be too massive to be stored or to be inverted for the purpose of prediction. For a concise overview of these computational challenges, we refer to Table 1 in Masak et al. (2023). To reduce complexity, many works impose structural assumptions on the covariance, one of them being separability. Roughly speaking, separability states that the covariance of a space-time process can be decomposed into a purely temporal and a purely spatial component. The implied elimination of space-time interactions cuts the number of model parameters and makes the covariance tractable again (see Genton (2007)). Besides, separability ∗Correspondence to: Holger Dette, Department of Mathematics, Ruhr-University Bochum, Bochum, Germany. Email: [email protected] © 2024 The Author(s). Journal of Time Series Analysis published by John Wiley & Sons Ltd. This is an open access article under the terms of the Creative Commons Attribution-NonCommercial-NoDerivs License, which permits use and distribution in any medium, provided the original work is properly cited, the use is non-commercial and no modifications or adaptations are made.
TESTING COVARIANCE SEPARABILITY 403 entails a product structure for the principal components, facilitating the construction of estimators and inference methods (see Gromenko et al.,2012,2016). To rigorously define separability, consider a stochastic process {X(s,t)|s∈K1,t∈K2}, with one argument in a spatial domain K1and one in a temporal domain K2. Then, under suitable conditions (see, e.g., Janson and Kaijser, 2015), its covariance operator can be defined point-wise as C(s,t,s′,t′)∶=E[(X(s,t)−E[X(s,t)])(X(s′,t′)−E[X(s′,t′)])]. We call Cseparable, if there exist two functions C1(spatial), C2(temporal), such that C(s,t,s′,t′)=C1(s,s′)×C2(t,t′)∀s,s′∈K1,∀t,t′∈K2. As pointed out before, the product structure of separability prunes model parameters, making the model statistically and computationally more tractable. However, separability is not a free lunch for space-time data. Indeed, erroneously assuming a separable model can lead to inconsistent estimates and biased inference results. This point is crucial, as separability is rarely self-evident, as noticed by many authors (see Scaccia and Martin, 2005; Genton, 2007; Aston et al.,2017, among many others). To address this issue, statistical tests have been proposed to examine separability, such as for finite dimensional data by Matsuda and Yajima (2004), Scaccia and Martin (2005), Fuentes (2006), and Crujeiras et al. (2010). More recently, non-parametric methods tailored to functional observations have been devised by Aston et al. (2017), Constantinou et al. (2017,2018), Bagchi and Dette (2020), and Dette et al. (2022). While these latter works differ in terms of their inference strategies, they share the mathematical setup of modelling observations in a space of square integrable functions. This approach is standard in FDA (see the monographs of Horváth and Kokoszka, 2012; Hsing and Eubank, 2015) and provides the most immediate extension of finite dimensional methodology to function spaces. Nevertheless, the choice of L2-spaces is usually more informed by mathematical convenience, than by practicability. Indeed, as many undergraduate textbooks point out, the L2-distance is hard to geometrically interpret and often stuns the novice by its unintuitive notion of convergence. More specific to FDA, basing a theory on L2-spaces usually ignores structural features of the functions, such as continuity. Notice that the bulk of FDA relies on continuous, non-parametric curve estimation as preprocessing. In such cases, insisting on an L2-framework can feel disjoint from a user’s intuition and defy heuristic interpretations of results. Following recent work of Degras (2011); Cao et al. (2012); Degras (2017) and Dette et al. (2020), we opt instead to conduct FDA on the in our opinion more natural space of continuous functions equipped with the supremum norm. On this more intricate space, we study functional space-time processes and advance a new test for separability of the covariance, which is based on an estimate of the maximum deviation between the covariance operator and an approximation by a covariance operator from a separable process. We combine the profound theory of weak convergence of stochastic processes (see, e.g., van der Vaart and Wellner, 1996 or Giné and Nickl, 2016 among many others) with new differentiability results for these separability measures to study the asymptotic properties of the corresponding estimates. To improve the finite sample properties and to avoid the estimation of complicated nuisance parameters we propose a multiplier bootstrap for dependent data as a more practicable alternative. In particular, our results hold under weaker dependence and stationarity assumptions compared to previous works. The rest of this article is organized as follows. Section 2provides a brief introduction to random variables on the space of continuous functions and the notion of separability. Subsequently, we discuss three methods from the related literature to approximate a covariance kernel Cby a separable version (we develop theory for all three approximations at a later point). In Section 3, we present the test statistic for the hypothesis of separability and prove its weak convergence. To approximate the limiting distribution, we discuss a non-standard multiplier bootstrap for dependent data. Section 4is dedicated to the finite sample properties of our test, which we study by virtue of simulations as well as a data example. Finally, all proofs and technical details are deferred to the Appendix S1. J. Time Ser. Anal. 46: 402–420 (2025) © 2024 The Author(s). wileyonlinelibrary.com/journal/jtsa DOI: 10.1111/jtsa.12764 Journal of Time Series Analysis published by John Wiley & Sons Ltd.
404 H. DETTE, G. DIERICKX and T. SKUTTA 2. MATHEMATICAL CONCEPTS Here, we lay the mathematical foundations of FDA in the space of continuous functions. Section 2.1 begins with a review of random, continuous functions, as well as basic concepts, such as functional expectations and covariance kernels. Subsequently, we define the model assumption of separability, which is the focus of our below statistical analysis. In Section 2.2, we discuss three methods to approximate a kernel Aby a separable version Axand show that each approximation map (A→ Ax) is well defined and differentiable (see Theorem 2.2). 2.1. Mathematical preliminaries Let (K)∶={f∶K→R|fcontinuous}, denote the space of continuous, real valued functions defined on a non-empty, compact set K⊂Rd, which equipped with the common supremum norm (or ‘sup-norm’ for short) ||f||∶= sup t∈K|f(t)|,(2.1) is a Banach space. If there is no danger of confusion, we sometimes refer to a function f∈(K)by its evaluation f(t). Moreover, we also use the notation ||⋅||to refer to the max-norm (maximum absolute entry) for matrices or vectors (corresponding to a finite set Kin (2.1)). Letting (Ω,,P)denote a complete probability space, we call a map X∶(Ω,,P)→(K)a random (K)-valued function, if Xis Borel-measurable w.r.t. the sup-norm. We point out that Borel-measurability of X(as a Banach space valued function) is equivalent to measurability of the real valued marginals X(t)∶(Ω,,P)→R for all t∈K. Supposing that the first absolute moment of Xexists, in the sense that E||X||<∞, the expectation of Xis well defined (in the Bochner sense), with EX∈(K)and equal to the point-wise expectation E[X(t)] for any t∈K. If the stronger moment condition E||X||2<∞holds, we can define the expectation of the product function {(X(s)−EX(s))(X(t)−EX(t))}s,t∈Kon the space (K2)∶=(K×K). We call this the covariance kernel and define it point-wise as C(s,t)∶=E[(X(s)−EX(s))(X(t)−EX(t))]. For further details on defining moment of Banach space valued random variables we refer to Janson and Kaijser (2015). As common in the study of functional data, we sometimes consider for a continuous kernel function A∈(K2)the corresponding integral operator {f(t)}t∈K→ {∫KA(s,t)f(s)ds}t∈K. For ease of notation, we usually identify the integral operator with the associated kernel (notice that the relation is one-one) and define the expression A[f]∶={ ∫KA(s,t)f(s)ds}t∈K. In particular, we call the integral operator associated with Cthe covariance operator. Finally, we observe that (K)can be understood as a subspace of L2(K), the space of square integrable functions, equipped with the canonical semi-norm ||f||L2∶= {∫Kf(s)2ds}1∕2. Recall that the L2-topology is weaker than the uniform topology, since we consider functions with compact support. In the next step, we consider the mathematical property of separability.LetK1⊂Rp,K2⊂Rqbe compact and non-empty sets. Then we say that a kernel A∈((K1×K2)2)is separable, if there exist kernels Ai∈(K2 i)for i=1,2, s.t. A(s,t,s′,t′)=A1(s,s′)A2(t,t′)∀s,s′∈K1,∀t,t′∈K2. Notice that here we take the point-wise product of the functions A1,A2(this means that the integral operator corresponding to Ais the tensor product of the integral operators associated with A1,A2; see e.g. Dette et al.,2022). Later, we will study the case of a separable covariance kernel Cof a random function X. To assess separability, wileyonlinelibrary.com/journal/jtsa © 2024 The Author(s). J. Time Ser. Anal. 46: 402–420 (2025) Journal of Time Series Analysis published by John Wiley & Sons Ltd. DOI: 10.1111/jtsa.12764
TESTING COVARIANCE SEPARABILITY 405 we will employ an estimator CNand compare it to a separable approximation. The construction of such separable approximations is the subject of the next section. 2.2. Separable approximations Here, we discuss the approximation of a kernel function A∈((K1×K2)2)by the (separable) product of two functions B=A1×A2with Ai∈(K2 i)and i=1,2. This approximation is key to our subsequent analysis, as it characterizes separability by the vanishing goodness-of-fit measure ||A−B||=0. How then should we construct a separable approximation? The most natural way might be to optimize over all separable maps w.r.t. the norm ||⋅||,thatis,tofind Bopt ∈argmin{||A−B||∶Bseparable}. Unfortunately, this type of approximation is computationally infeasible. To illustrate this point, let us consider the analogue problem of finding an optimal separable approximation for a d×d-matrix w.r.t. to the max-norm (also denoted by ||⋅||). This problem reflects the search for an optimal, separable approximation for a discretized version of A(as it would be saved on a computer in a real application). In Remark 2.1 “⊗” denotes the well-known Kronecker product for matrices. Remark 2.1. Let M∈Rd×dbe a real-valued matrix with d=pq and p,q∈N. Then the problem of finding a solution to the following minimization problem “minimize ||M−M1⊗M2|| with M1∈Rp×p,M2∈Rq×q”, is NP-complete. Proof. The problem of finding optimal rank-1-approximations with regard to the max-norm is an NP-complete problem, according to Gillis and Shitov (2017). According to Section 2 of Genton (2007) the problem of finding separable approximations is equivalent to finding rank-1-approximations. ◾ The notion of NP-completeness in the above Remark 2.1 refers to a set of problems which are (by today’s knowledge) not efficiently solvable (for details, see Garey and Johnson (1979)). Evidently, if it is not feasible to find max-norm approximations for matrices, it has a fortiori to be true for continuous kernels (which then cannot be optimally approximated by a separable kernel even on a grid with reasonable effort). This insight suggests to use different, suboptimal, but computationally feasible approximation methods. For this reason, we will discuss three distinct procedures of separable covariance approximation in the next section that are commonly used in the L2-literature. While these covariance estimators have been used before, the resulting separability measures are different, since they are based on the sup-norm rather than an integrated norm. 2.2.1. Partial trace approximations Let A∈((K1×K2)2)be a kernel function. We can then define the marginal kernels Atr 1∈(K2 1)and Atr 2∈(K2 2) point-wise as Atr 1(s,s′)∶=∫K2 A(s,w,s′,w)dwand Atr 2(t,t′)=∫K1 A(u,t,u,t′)du,(2.2) and therewith the separable approximation Atr of Aas Atr(s,t,s′,t′)∶= Atr 1(s,s′)Atr 2(t,t′) ∫K1∫K2A(u,w,u,w)dudw.(2.3) J. Time Ser. Anal. 46: 402–420 (2025) © 2024 The Author(s). wileyonlinelibrary.com/journal/jtsa DOI: 10.1111/jtsa.12764 Journal of Time Series Analysis published by John Wiley & Sons Ltd.
406 H. DETTE, G. DIERICKX and T. SKUTTA Notice that Atr is well-defined, if the denominator in (2.3) is non-zero. For a covariance kernel (symmetric and positive semi-definite) this is automatically satisfied, unless the kernel is trivial. Partial trace approximations are well-known from quantum physics. Quantum states can be represented in different ways: by their wave functions (i.e. a probability distribution function), or by some operator. Pure states are the basic building blocks from which mixed states can be assembled. These mixed states are then represented as a random wave function or as a tensor product of operators. In quantum-mechanical information theory one is interested in relating the entropy of mixed states to the pure states they are made of. Here partial traces come into play since they allow in some sense to decompose the mixed states. We refer the interested reader to Bhatia (2003) for a linear algebraic introduction to partial trace operators and to the review of Wehrl (1978) (and references therein) for applications to quantum information theory. As far as we know partial traces have recently been applied in statistics for separability tests of space-time processes in L2-spaces (see Constantinou et al.,2018, Aston et al.,2017). In a recent work of Masak et al. (2023) generalizations of partial traces have been investigated in the context of ‘almost separable matrices’. 2.2.2. Partial product approximations Closely related to the notion of partial traces are partial products, that was introduced in Bagchi and Dette (2020). The partial product approximation depends on a user determined function 𝜓∈(K2 2), where typical choices are discussed in Bagchi and Dette (2020) (one being the constant 1). We then define for A∈((K1×K2)2)point-wise the marginal kernels Apr 1(s,s′)∶=∫K2 2 A(s,w,s′,w′)𝜓(w,w′)dwdw′and Apr 2(t,t′)∶=∫K2 1 A(u,t,u′,t′)Apr 1(u,u′)dudu′(2.4) and therewith the separable approximation Apr as Apr(s,t,s′,t′)∶= Apr 1(s,s′)Apr 2(t,t′) ∫K2 1(Apr 1(u,u′))2dudu′.(2.5) As for partial traces, the approximation is well defined for a non-vanishing numerator, that is, for Apr 1≠0.(2.6) This condition is fulfilled if A≠0 for an appropriate choice of 𝜓. Recently, partial products have been used in Dette et al. (2022) for the quantification of separability in L2-spaces and more recently for the efficient derivation of optimal separable approximations w.r.t. the L2-norm in Masak et al. (2023). These latter approximations are discussed next. 2.2.3. SPCA approximations SPCA (separable principal component analysis) provides the last approximation method that we want to investigate. With respect to the L2-norm, SPCA-approximations are optimal, whereas for the sup-norm they do not occupy this special position (see Remark 2.1). SPCA-approximations have been known in finite dimensions for several decades (Van Loan and Pitsianis, 1993; Genton, 2007) and have recently been generalized to infinite-dimensional Hilbert spaces by Dette et al. (2022). The name ‘SPCA’ has been introduced by Masak et al. (2023), who proposed an efficient algorithm for the calculation of these approximations by virtue of the partial product (see Section 2.2.2 in this reference). Consider the kernels ÃPCA 1(s,s′,s,s′)∶=∫K2 2 A(s,w,s′,w′)A(s,w,s′,w′)dwdw′, wileyonlinelibrary.com/journal/jtsa © 2024 The Author(s). J. Time Ser. Anal. 46: 402–420 (2025) Journal of Time Series Analysis published by John Wiley & Sons Ltd. DOI: 10.1111/jtsa.12764
TESTING COVARIANCE SEPARABILITY 407 and ÃPCA 2(t,t′,t,t′)∶=∫K2 1 A(u,t,u′,t′)A(u,t,u′,t′)dudu′, both of which are continuous, symmetric (w.r.t. to the first and second pair of components) and positive definite. According to Mercer’s theorem (theorem 4.49 in Steinwart and Christmann (2008)) we can write them down as ÃPCA 1(s,s′,s,s′)=∑ i≥1 𝜆ivi(s,s′)vi(s,s′),ÃPCA 2(t,t′,t,t′)=∑ i≥1 𝜆iui(t,t′)ui(t,t′). Here {𝜆i,vi}i∈Nand {𝜆i,ui}i∈Nare the eigensystems of the respective integral operators and the eigenvalues are supposed to be in descending order. It is not difficult to show that if the strict inequality 𝜆1>𝜆 2,(2.7) holds, the eigenfunction v1,u1are well defined (up to sign) and continuous (see for both results the Appendix S1). Supposing that this is true, we define the marginals APCA 1(s,s′)=√𝜆1v1(s,s′),APCA 2(t,t′)=√𝜆1u1(t,t′), and therewith the SPCA approximation APCA(s,t,s′,t′)=APCA 1(s,s′)APCA 2(t,t′).(2.8) We conclude this section with a general result regarding the approximation maps A→ Ax,foreveryx∈ {tr,pr,PCA}. We demonstrate that each of these maps is well-defined (in the sense that the resulting kernels are indeed continuous and positive definite). Moreover, we conclude that the approximation maps are Fréchet differentiable, which is critical for the subsequent application of the functional Delta-method (see, for instance, section 3.9 in van der Vaart and Wellner, 1996). Theorem 2.2. Let K1⊂Rp,K2⊂Rqbe compact, non-empty sets and let Ã∈((K1×K2)2)beacovariance kernel with Ã≠0. Moreover, suppose that equation (2.6)and(2.7) are satisfied. Then the maps {Fx i∶((K1×K2)2)→(K2 i)∶A→ Ax i Fx∶((K1×K2)2)→((K1×K2)2)∶A→ Ax are for i=1,2andx∈{tr,pr,PCA}well defined in a sufficiently small, open neighborhood of Ãand Fréchet differentiable in Ã. Moreover, Fx i[ A]and Fx[ A]are again covariance kernels. The proof of this theorem can be found in the Appendix S1, where we also state the explicit form of the derivatives. In the next section, we will use this result in the context of statistical inference for spatiotemporal data. 3. TESTING SEPARABILITY FOR A CONTINUOUS COVARIANCE KERNEL Here, we develop a statistical test for the hypothesis of a separable covariance kernel in the space of continuous functions. First, we specify the statistical framework in Section 3.1 and, second, investigate test statistics for the hypothesis of separability in Section 3.2. Lemma 3.5 entails weak convergence of these test statistics to suprema of Gaussian processes (under the null hypothesis) and hence provides the theoretical tools for separability tests. To approximate the asymptotic quantiles of the test statistics, we present in Section 3.3 a multiplier bootstrap for dependent data. J. Time Ser. Anal. 46: 402–420 (2025) © 2024 The Author(s). wileyonlinelibrary.com/journal/jtsa DOI: 10.1111/jtsa.12764 Journal of Time Series Analysis published by John Wiley & Sons Ltd.
408 H. DETTE, G. DIERICKX and T. SKUTTA 3.1. Notations and assumptions Let K1⊂Rpand K2⊂Rqbe compact, non-empty sets and (Xn)n∈Zbe a time series of random functions in the Banach space (K1×K2)(for a definition, see Section 2.1). In the following, we will assume that (Xn)n∈Z satisfies fourth-order stationarity, in the sense that for any indices (n1,…,n4)∈Z4and k∈Zthe vectors (Xn1,…,Xn4)and (Xn1+k,…,Xn4+k)have the same distribution. In particular, supposing that E||X1||2<∞, both the mean function EXn(s,t)and the covariance C(s,t,s′,t′)∶=E[Xn(s,t)Xn(s′,t′)]−E[Xn(s,t)]E[Xn(s′,t′)], do not depend on nand are well defined on the spaces (K1×K2)and ((K1×K2)2)respectively. In the following, we want to construct a test for the hypothesis of a separable covariance operator, that is, H0∶Cis separable vs. H1∶Cis not separable,(3.1) where separability is defined in Section 2.1. For this purpose, suppose that we observe a sample of Nrandom functions X1,…,XN(from the time series (Xn)n∈Z). For estimating Cwe use the (standard) empirical covariance estimator defined as: CN(s,t,s′,t′)∶= 1 N N ∑ n=1(Xn(s,t)−XN(s,t))(Xn(s′,t′)−XN(s′,t′)),(3.2) where XN(s,t)∶=1 N∑N n=1Xn(s,t),(s,t)∈K1×K2, denotes the sample mean estimator function. In the next section, we will construct a test statistic for the hypothesis H0, by comparing CNwith a separable approximation. The use of this statistic is motivated by the fact that, under suitable assumptions on the dependence structure, CNis a consistent estimator of Cby the law of large numbers (on Banach spaces). To quantify dependence, we introduce the popular concept of 𝛼-mixing sequences. Definition 3.1. Let (Xn)n∈Zbe a sequence of random variables on some Banach space. For index sets I,J⊂Zwe define the set distance dist(I,J)=min{|i−j|∶i∈I,j∈J}, and use the notation J∶= 𝜎({Xi∶i∈I})for the 𝜎-algebra generated by the family of random variables {Xi∶ i∈I}.Forr∈N0the rth 𝛼-mixing coefficient is then defined as 𝛼(r)=sup{|P(A∩B)−P(A)P(B)|∶A∈I,B∈J,dist(I,J)≥r,I,J⊂Z}. The sequence (Xn)n∈Zis called 𝛼-mixing if 𝛼(r)→0, as r→∞. We can now state the theoretical assumptions for the separability test, developed in this section. Assumption 3.2. (i) The sequence (Xn)n∈Zconsists of centered, random functions in (K1×K2)and is fourth-order stationary. (ii) There exist non-negative random variables M1,M2,…, parameters 𝛽∈(0,1]and J>0 with the constraint J𝛽>2(p+q)+1 such that sup nE(||Xn||JMJ n)<∞,and |Xn(s,t)−Xn(s′,t′)|≤Mnmax{||s−s′||𝛽,||t−t′||𝛽},(3.3) holds for all (s,t),(s′,t′)∈K1×K2and all n∈Z. wileyonlinelibrary.com/journal/jtsa © 2024 The Author(s). J. Time Ser. Anal. 46: 402–420 (2025) Journal of Time Series Analysis published by John Wiley & Sons Ltd. DOI: 10.1111/jtsa.12764
TESTING COVARIANCE SEPARABILITY 409 (iii) For some 𝛾>max{J,8}it holds that E||X1||𝛾<∞. iv) The sequence (Xn)n∈Zis 𝛼-mixing in the sense of Definition (3.1), with 𝛼(r)≤𝜅∕(1+r)a,where𝜅>0is some constant and a>2𝛾∕(𝛾−8). We briefly discuss each of these Assumptions. Remark 3.3. (i) We require second-order stationarity s.t. the empirical mean XNand the empirical covariance operator CNare consistent estimators. The stronger assumption of fourth-order stationarity guarantees existence of the long-run variance operator of √N( CN−C). An examination of our proofs shows that these assumptions can be further relaxed to conditions on the moments of Xn. However, for parsimony of presentation, we do not discuss these mathematically weaker (but harder to understand) adaptations. Our stationarity assumption is weaker than those in the related literature (both in L2-and-spaces), where either independence or strict stationarity is considered (see Constantinou et al.,2017; Aston et al.,2017; Bagchi and Dette, 2020; Dette and Kokot, 2022). (ii) To devise asymptotic tests for separability, we derive a CLT on the Banach space of continuous functions. Such results require the validation of tightness conditions, which depend on the geometry of the underlying space (reflected by the entropy rate). In the case of i.i.d. observations, sufficient conditions for a CLT can be found in Jain and Marcus (1975), and for the dependent case (under 𝛼-and𝜙-mixing) in Dmitrovskii et al. (1984). More specifically, theorem 1 of Jain and Marcus (1975) requires a random Lipschitz condition, which together with an entropy condition entails a functional CLT on (K).Inthis article, we further relax this assumption by requiring only a random 𝛽-Hölder condition. This condition specifically includes sample paths of the Brownian motion, that are almost surely 𝛽-Hölder continuous, for 𝛽<1∕2. However, additional smoothness helps to reduce moment conditions imposed on the data (see the next assumption). (iii)-(iv) The existence of sufficiently many moments is the key to proving our CLT (and its bootstrap variant). As might be expected, stronger moment conditions can be traded off against weaker smoothness assumptions for the data functions, as well as milder conditions on temporal dependence. We quantify dependence by the decay rate of strong mixing coefficients, where a slower decay expresses stronger time-dependence. In comparison to the related literature, our mixing conditions are rather weak and include large classes of dependent time series. (v) The assumption of 𝛼-mixing used in this article is less restrictive than the alternative concepts of 𝛽-or 𝜙-mixing typically used in Banach spaces (Dehling, 1983). However, 𝛼-mixing, and by implication most other mixing assumptions, have some known weaknesses. For instance, Andrews (1984) demonstrated that a (specific) AR-(1)process with Bernoulli innovations fails to be strongly mixing for an arbitrarily small, positive AR-parameter. Such counterintuitive findings have fostered interest in alternative notions of dependence, such as the popular concept of m-approximability (see Hörmann and Kokoszka, 2010). In FDA m-approximability has become an important competitor to mixing assumptions as some existing models, such as AR-processes, can be easily fit into this framework. However, up to this point, m-approximability has exclusively been developed for functional data on Hilbert spaces. It will be an interesting subject of research to extend the scope of m-approximability to general Banach spaces. We conclude by pointing out that at least under some additional assumptions (such as existence of densities) popular processes of AR and ARMA type are strongly mixing (see Mokkadem, 1988 and Dedecker et al.,2007). For additional details on mixing processes we refer to Bradley (2007,2005). 3.2. Central limit theorems in ((K1×K2)2) We begin our theoretical derivations by proving weak convergence of the standardized empirical covariance operator √N( CN−C)to a Gaussian process G. This result, together with the corresponding bootstrap convergence, J. Time Ser. Anal. 46: 402–420 (2025) © 2024 The Author(s). wileyonlinelibrary.com/journal/jtsa DOI: 10.1111/jtsa.12764 Journal of Time Series Analysis published by John Wiley & Sons Ltd.
416 H. DETTE, G. DIERICKX and T. SKUTTA Table II. Empirical rejection probabilities of the bootstrap test (3.9) under the hypothesis (c=0) and the alternative (c=1), for data Xnfollowing model (4.3) (AR(1)-process, with parameter 𝜌=0.5) Rejection probability under H0Rejection probability under H1 SN=50 N=100 N=150 N=200 N=50 N=100 N=150 N=200 4 7.0 5.2 5.5 4.3 50.4 83.0 97.1 99.9 6 6.4 4.5 4.7 5.8 49.6 81.2 96.9 99.8 8 7.3 5.4 3.8 5.8 48.6 80.0 96.7 99.9 10 7.1 4.6 4.5 6.1 45.1 80.9 96.7 99.6 12 5.4 6.0 4.9 5.8 43.7 81.6 96.5 99.6 14 5.6 5.4 5.0 6.4 43.4 80.7 96.5 99.7 20 5.6 5.8 5.3 6.2 45.2 79.3 96.2 100.0 30 6.0 5.7 4.9 6.4 44.0 81.1 95.2 99.6 In a second simulation setup, we consider an AR(1) process to simulate data. More precisely, with centered, i.i.d. Gaussian random functions (en)n∈Zcharacterized by the covariance kernel defined in (4.1), we investigate the process Xn(s,t)∶=en(s,t)+𝜌⋅Xn−1(s,t),(4.3) where 0 ≤𝜌<1 is the AR-parameter. In our simulations we set 𝜌=0.5, consider the same sample sizes, discretizations, nominal level and bootstrap generations as before, and choose the bootstrap bandwidths as lN=2,3,4,4. Our results are collected in Table II with empirical rejection probabilities under the hypothesis of separability (c=0) and the alternative (c=1). Our results in Tables Iand II attest a satisfactory performance of the bootstrap test. The nominal level is reasonably approximated in most cases, even though for N=50, the smallest sample size considered, it is slightly less precise. The power of the bootstrap test is high in most scenarios. In the case of the MA(1)-process (I)we can make comparisons to the L2-benchmarks. While for small values of S, the benchmark test fares slightly better, the bootstrap’s performance does not deteriorate for larger S, where it clearly outperforms the benchmark. Even raising Sto 20 or 30 does not impinge on performance in any systematic way (exactly what a theory of continuous processes would suggest). Computationally, the bootstrap is particularly user-friendly: It allows a straightforward parallelization in the generation of bootstrap samples and is hence easy to implement. We have run our simulations on a standard desktop computer (3.2 GHz Apple M1 Pro, Octa-core, 16 GB RAM) and any test evaluation needed less than a minute (for S≤10 less than 2 seconds), which underpins the practicability of this approach. Remark 4.1. We conclude this section with a small sensitivity analysis. In our simulations, we have considered different sample sizes Nand different discretizations across space S, but have kept the number of time points T=50 fixed. We briefly show that the effect of varying Tis relatively small. For this purpose, we consider the second scenario in our simulation study of the MA(1) process, with parameter 𝜌=0.5. We fix the sample size N=100 and consider different temporal discretizations T=10,20,30,50 and spatial discretizations S=4,6,8,10,12,14,20,30 under the hypothesis and the alternative. The simulated rejection probabilities of the bootstrap test (3.9) are presented in Table III. We observe that the results are rather stable with respect to different choices of Sand T. 4.2. A data example We apply our test to the acoustic phonetic dataset of acoustic (log-)spectograms of Aston et al. (2017). A brief discussion with additional references can be found in section 4.2 of their paper. Both the raw and preprocessed wileyonlinelibrary.com/journal/jtsa © 2024 The Author(s). J. Time Ser. Anal. 46: 402–420 (2025) Journal of Time Series Analysis published by John Wiley & Sons Ltd. DOI: 10.1111/jtsa.12764
TESTING COVARIANCE SEPARABILITY 417 Table III. Empirical rejection probabilities of the bootstrap test (3.9) under the hypothesis (c=0) and the alternative (c=1), for data Xnfollowing model (4.3) (AR(1)-process, with parameter 𝜌=0.5) and sample size N=100, where we consider different spatial and temporal discretizations Rejection probability under H0Rejection probability under H1 ST=10 T=20 T=30 T=50 T=10 T=20 T=30 T=50 4 6.5 5.0 5.5 4.9 84.1 84.1 82.7 83.0 6 5.9 4.4 4.7 5.9 83.6 85.4 84.4 81.2 8 7.0 5.6 3.8 4.3 85.2 84.5 83.2 80.0 10 5.4 6.2 4.5 6.0 82.1 82.7 83.2 80.9 12 4.5 5.8 4.9 4.1 83.2 81.6 81.8 81.6 14 5.3 7.0 5.0 5.1 81.5 82.5 82.1 80.7 20 6.8 6.2 5.3 5.2 82.4 81.9 82.4 79.3 30 6.0 5.8 4.9 5.2 81.7 81.3 81.5 81.1 Table IV. The relative measure of separability w.r.t. the trace approximation of the five Romance languages French Italian Portuguese American Spanish Iberian Spanish Relative measure 0.631 0.895 0.868 0.947 0.882 data together with detailed descriptions are contained in the file Acoustic_Data_And_Code.zip available from Series C datasets of Volume 67 (https://rss.onlinelibrary.wiley.com/hub/journal/14679876/series-c-datasets /67_5). The dimension of the data is 80 in space and 100 in time. The data consist of recordings of spoken words (in this case the numbers one to ten) in five different Romance languages. For statistical use they have been transformed in the form of acoustic log-spectograms. For our purposes only the preprocessed data are used (namely the files WarpedPSD.RData and SVRF_WarpedPSD_SuppMat.RData). There are in total 219 data ‘functions’ in a frequency-time domain as described in sections 2–4 of Pigoli et al. (2018). As already indicated by the analysis in Pigoli et al. (2018) the null hypothesis (3.1) of separability seems to be violated in this instance. Before applying our test statistic we look at relative measures of separability. More concretely, for the separable trace approximation, the relative measure is given by || CN−Ftr( CN)|| || CN|| . Indeed, as a preliminary analysis, we found that the relative measure of separability for the language covariance operators are rather high, see Table IV. For each of the five languages we ran our bootstrap statistic of 1000 repetitions on the residual (i.e. centered) surface data for each language separately. As a result the hypothesis of separability for each of these languages is rejected with a highly significant p-value of less than 0.1%. See the first row of Table Vfor the p-values of our test. For the sake of completeness the second and third row contain the p-values obtained in Aston et al. (2017)fortwo frequency and three time dimensions and eight frequency and 10 dimensions respectively. Our method provides several advantages: It is less dependent on tuning parameters compared to the Studentized version of the empirical bootstrap of Aston et al. (2017), who have to choose additional parameters of dimensions of eigendirections (note that their test can only detect deviations from separability along those eigendirections). Moreover, when too few eigendirections are chosen (see section 4.2 of Aston et al.,2017) the test can lack power. Second, to calculate the bootstrap we only need to save the data, not the whole covariance operator, hence avoiding storage problems. To investigate the stability of our findings, w.r.t. to different discretizations in space and time, we display in rows 4–6 of Table Vp-values of our procedure, when only a fraction of the data (10%,20%,50%) in both space and J. Time Ser. Anal. 46: 402–420 (2025) © 2024 The Author(s). wileyonlinelibrary.com/journal/jtsa DOI: 10.1111/jtsa.12764 Journal of Time Series Analysis published by John Wiley & Sons Ltd.
418 H. DETTE, G. DIERICKX and T. SKUTTA Table V. p-Values of three different bootstrap tests of five Romance languages French Italian Portuguese American Spanish Iberian Spanish Test (3.9)<0.001 <0.001 0.001 <0.001 <0.001 Emp. stud. test (2,3)0.078 0.197 0.022 0.360 0.013 Emp. stud. test (8,10)0.001 0.002 0.001 0.001 <0.001 Test (3.9)|S|=8,|T|=10 0.01 <0.001 0.142 0.021 0.004 Test (3.9)|S|=16,|T|=20 0.008 <0.001 0.129 0.007 0.027 Test (3.9)|S|=40,|T|=50 0.005 0.002 0.139 0.005 0.026 Note: First row: the test (3.9) proposed in this article. Second and third row: the studentized version of the empirical bootstrap test proposed by Aston et al. (2017) with frequency and time dimensions (2,3)and (8,10)respectively. Fourth till sixth row: p-values of our test with 10,20 and 50%of the data points in Space and Time respectively. time is used. We observe that most p-values remain consistently small, with some exception for the Portuguese language data, where the sample size is smallest (25). This finding is consistent with our expectation that some alternatives require finer grids to be detectable than others. ACKNOWLEDGEMENTS This article was partially supported by the DFG Research unit 5381 Mathematical Statistics in the Information Age, project number 460867398, DFG project number 457238972. The authors would like to thank the associate editor and two referees for their constructive comments on earlier versions of this article. Open Access funding enabled and organized by Projekt DEAL. DATA AVAILABILITY STATEMENT Both the raw and preprocessed data used in this article are contained in the file Acoustic_Data_And_Code.zip available from Series C datasets of Volume 67 (https://rss.onlinelibrary.wiley .com/hub/journal/14679876/series-c-datasets/67_5). The data that support the findings of this study are available from the corresponding author on reasonable request. SUPPORTING INFORMATION Additional Supporting Information may be found online in the supporting information tab for this article. REFERENCES Andrews DWK. 1984. Non-strong mixing autoregressive processes. Journal of Applied Probability 21:930–934. Aston JAD, Pigoli D, Tavakoli S. 2017. Tests for separability in nonparametric covariance operators of random surfaces. Annals of Statistics 45(4):1431–1461. Aue A, Rice G, Sönmez O. 2018. Detecting and dating structural breaks in functional data without dimension reduction. Journal of the Royal Statistical Society, Series B (Statistical Methodology) 80(3):509–529. Bagchi P, Dette H. 2020. A test for separability in covariance operators of random surfaces. Annals of Statistics 48(4):2303–2322. Bhatia R. 2003. Partial traces and entropy inequalities. Linear Algebra and its Applications 370(1):125–132. Bradley RC. 2007. Introduction To Strong Mixing Conditions. Volumes 1,2 And 3 Heber City: Kendrick Press. Bradley RC. 2005. Basic properties of strong mixing conditions. a survey and some open questions. Probability Surveys 2:107–144. Bücher A, Kojadinovic I. 2013. A dependent multiplier bootstrap for the sequential empirical copula process under strong mixing. Bernoulli 22(2):927–968. Bücher A, Kojadinovic I. 2019. A note on conditional versus joint unconditional weak convergence in bootstrap consistency results. Journal of Theoretical Probability 32:1145–1165. wileyonlinelibrary.com/journal/jtsa © 2024 The Author(s). J. Time Ser. Anal. 46: 402–420 (2025) Journal of Time Series Analysis published by John Wiley & Sons Ltd. DOI: 10.1111/jtsa.12764
TESTING COVARIANCE SEPARABILITY 419 Cao G, Yang L, Todem D. 2012. Simultaneous inference for the mean function based on dense functional data. Journal of Nonparametric Statistics 24(2):359–377. Constantinou P, Kokoszka P, Reimherr M. 2017. Testing separability of space–time functional processes. Biometrika 104(2):425–437. Constantinou P, Kokoszka P, Reimherr M. 2018. Testing separability of functional time series. Journal of Time Series Analysis 39(5):731–747. Crujeiras RM, Fernández-Casal R, González-Manteiga W. 2010. Nonparametric test for separability of spatio-temporal processes. Environmetrics 21:382–399. Dedecker J, Doukhan P, Lang G, León JR, Louhichi S, Prieur C. 2007. Weak Dependence: With Examples and Applications. Lecture Notes in Statistics, Vol. 190 New York: Springer. Degras D. 2011. Simultaneous confidence bands for nonparametric regression with functional data. Statistica Sinica 21:1735–1765. Degras D. 2017. Simultaneous confidence bands for the mean of functional data. WIREs Computational Statistics 9(3):e1397. Dehling H. 1983. Limit theorems for sums of weakly dependent Banach space valued random variables. Zeitschrift für Wahrscheinlichkeitstheorie und Verwandte Gebiete 63:393–432. Dette H, Kokot K. 2022. Detecting relevant differences in the covariance operators of functional time series: a sup-norm approach. Annals of the Institute of Statistical Mathematics 74(2):195–231. Dette H, Kokot K, Aue A. 2020. Functional data analysis in the banach space of continuous functions. Annals of Statistics 48(2):1168–1192. Dette H, Dierickx G, Kutta T. 2022. Quantifying deviations from separability in space-time functional processes. Bernoulli 28(4):2909–2940. Dmitrovskii VA, Ermakov SV, Ostrovskii EI. 1984. The central limit theorem for weakly dependent Banach-valued variables. Theory of Probability and Its Applications 28:89–104. Fuentes M. 2006. Testing for separability of spatial–temporal covariance functions. Journal of Statistical Planning and Inference 136(2):447–466. Garey MR, Johnson DS. 1979. Computers and Intractability: A Guide to the Theory of NP-Completeness. Series of Books in the Mathematical Sciences, 1st ed. USA: W. H. Freeman. Genton MG. 2007. Separable approximations of space-time covariance matrices. Environmetrics 18:681–695. Gillis N, Shitov Y. 2017. Low-rank matrix approximation in the infinity norm. Linear Algebra and its Applications 581:367–382. Giné E, Nickl R. 2016. Mathematical Foundations of Infinite-Dimensional Statistical Models. Cambridge Series in Statistical and Probablistic Mathematics Cambridge University Press, New York. Gromenko O, Kokoszka P, Zhu L, Sojka J. 2012. Estimation and testing for spatially indexed curves with application to ionospheric and magnetic field trends. The Annals of Applied Statistics 6(2):669–696. Gromenko O, Kokoszka P, Reimherr M. 2016. Detection of change in the spatiotemporal mean function. Journal of the Royal Statistical Society, Series B (Statistical Methodology) 79(1):29–50. Hall P. 1992. The Bootstrap and Edgeworth Expansion Springer, New York. Hörmann S, Kokoszka P. 2010. Weakly dependent functional data. Ann. Statist. 38:1845–1884. Horváth L, Kokoszka P. 2012. Inference for Functional Data with Applications. Springer Series in Statistics Springer, New York. Hsing T, Eubank R. 2015. Theoretical Foundations of Functional Data Analysis, with an Introduction to Linear Operators Wiley, New York. Jain NC, Marcus MB. 1975. The central limit theorem for C(S)-valued random variables. Journal of Functional Analysis 19:216–231. Janson S, Kaijser S. 2015. Higher Moments of Banach Space Valued Random Variables. Memoirs of the American Mathematical Society, Vol. 238. Providence, Rhode Island: American Mathematical Society. King MC, Staicu A-M, Davis JM, Reich BJ, Eder B. 2018. A functional data analysis of spatiotemporal trends and variation in fine particulate matter. Atmospheric Environment 184:233–243. Martínez-Hernández I, Genton M. 2020. Recent developments in complex and spatially correlated functional data. Brazilian Journal Of Probability And Statistics 34:204–229. Masak T, Sarkar S, Panaretos VM. 2023. Separable expansions for covariance estimation via the partial inner product. Biometrika 110(1):225–247. Mateu J, Giraldo R. 2022. Geostatistical Functional Data Analysis. Hoboken: Wiley. Matsuda Y, Yajima Y. 2004. On testing for separable correlations of multivariate time series. Journal of Time Series Analysis 24(4):501–528. Mokkadem A. 1988. Mixing properties of ARMA processes. Stochastic Processes and Their Applications 29:309–315. Paparoditis E, Politis DN. 2001. Tapered block bootstrap. Biometrika 88(4):1105–1119. J. Time Ser. Anal. 46: 402–420 (2025) © 2024 The Author(s). wileyonlinelibrary.com/journal/jtsa DOI: 10.1111/jtsa.12764 Journal of Time Series Analysis published by John Wiley & Sons Ltd.
420 H. DETTE, G. DIERICKX and T. SKUTTA Paparoditis E, Sapatinas T. 2016. Bootstrap-based testing of equality of mean functions or equality of covariance operators for functional data. Biometrika 103:727–733. Pigoli D, Hadjipantelis PZ, Coleman JS, Aston JAD. 2018. The statistical analysis of acoustic phonetic data: exploring differences between spoken romance languages. Journal of the Royal Statistical Society, Series C (Applied Statistics) 67:1103–1145. Rice G, Shang HL. 2017. A plug-in bandwidth selection procedure for long-run covariance estimation with stationary functional time series. Journal of Time Series Analysis 38(4):591–609. Scaccia L, Martin RJ. 2005. Testing axial symmetry and separability of lattice processes. Journal of Statistical Planning and Inference 131(1):19–39. Steinwart I, Christmann A. 2008. Support Vector Machines Springer Science and Business Media, New York. van der Vaart AW, Wellner JA. 1996. Weak convergence and empirical processes. With applications to statistics Springer Series in Statistics, New York. Van Loan CF, Pitsianis N. 1993. Approximation with kronecker products. Linear Algebra for Large Scale and Real-Time Applications 232:293–314. Wehrl A. 1978. General properties of entropy. Reviews of Modern Physics 50:221–260. wileyonlinelibrary.com/journal/jtsa © 2024 The Author(s). J. Time Ser. Anal. 46: 402–420 (2025) Journal of Time Series Analysis published by John Wiley & Sons Ltd. DOI: 10.1111/jtsa.12764