scieee AI-readable full text Open interactive document viewer

Asymptotic study of canonical correlation analysis: from matrix and analytic approach to operator and tensor approach

Fine, Jeanne

Abstract

Asymptotic study of canonical correlation analysis gives the opportunity to present the different steps of an asymptotic study and to show the interest of an operator and tensor approach of multidimensional asymptotic statistics rather than the classical, matrix and analytic approach. Using the last approach, Anderson (1999) assumes the random vectors to have a normal distribution and the non zero canonical correlation coefficients to be distinct. The new approach we use, Fine (2000), is coordinate-free,distribution-free and permits to have no restriction on the canonical correlation coefficients multiplicity order. Of course, when vectors have a normal distribution and when the non zero canonical correlation coefficients are distinct, it is possible to find again Anderson’s results but we diverge on two of them. In this methodological presentation, we insist on the analysis frame (Dauxois and Pousse, 1976), the sampling model (Dauxois, Fine and Pousse, 1979) and the different mathematical tools (Fine, 1987, Dauxois, Romain and Viguier, 1994) which permit to solve problems encountered in this type of study, and even to obtain asymptotic behavior of the analyses random elements such as principal components and canonical variables.)

Full text

Statistics & Operations Research Transactions SORT 27 (2) July-December 2003, 165-174 Statistics & Operations Research Transactions Asymptotic study of canonical correlation analysis: from matrix and analytic approach to operator and tensor approach Jeanne Fine∗ Universit´e Paul Sabatier, France Abstract Asymptotic study of canonical correlation analysis gives the opportunity to present the different steps of an asymptotic study and to show the interest of an operator and tensor approach of multidimensional asymptotic statistics rather than the classical, matrix and analytic approach. Using the last approach, Anderson (1999) assumes the random vectors to have a normal distribution and the non zero canonical correlation coefficients to be distinct. The new approach we use, Fine (2000), is coordinate-free, distribution-free and permits to have no restriction on the canonical correlation coefficients multiplicity order. Of course, when vectors have a normal distribution and when the non zero canonical correlation coefficients are distinct, it is possible to find again Anderson’s results but we diverge on two of them. In this methodological presentation, we insist on the analysis frame (Dauxois and Pousse, 1976), the sampling model (Dauxois, Fine and Pousse, 1979) and the different mathematical tools (Fine, 1987, Dauxois, Romain and Viguier, 1994) which permit to solve problems encountered in this type of study, and even to obtain asymptotic behavior of the analyses random elements such as principal components and canonical variables.) MSC: 62E20, 62H20, 62H25, 47N30 Keywords: multivariate analysis, canonical correlation analysis, asymptotic study, operator, coordinatefree, distribution-free 1 Classical approach 1.1 Population canonical correlation analysis Let Xand Ybe two random vectors, pand qdimensional respectively (p≤q) defined on a same probability space (Ω,A,P), centered and admitting order 4 moments. We ∗Address for correspondence: Laboratoire de Statistique et Probabilit´ es, Universit´ e Paul Sabatier, 118, route de Narbonne, 31062 Toulouse cedex, France. e-mail: [email protected]. Received: October 2003 Accepted: December 2003 166 Asymptotic study of canonical correlation analysis: from matrix and analytic... assume the matrix covariance VXof Xto be non-singular and we denote by HXthe vector space of real-valued random variables (r.r.v.) linear combinations of Xcomponents. We introduce similarly VYand HYand we denote by VXY the cross covariance the components of Xwith the ones of Y. The aim of canonical correlation analysis (CCA) of (X,Y) is to measure the relationships between Xand Y. CCA may be defined as the search for f1and g1, r.r.v. of HXand HYwith unit variance and maximal correlation ρ1then, iteratively, for j=2, ..., r, (r≤p), as the search of fjand gj, r.r.v. of HXand HYwith unit variance, uncorrelated with the ( fk)k<jand (gk)k<jand with maximal correlation ρj. The r.r.v. fj and gjare called jth canonical variables and the real ρjof [0,1] is called jth canonical correlation coefficient. Let RX=V−1 2 XVXY V−1 YVYXV−1 2 Xand the same for RYpermuting Xand Yroles. It is easy to verify that RXand RYhave the same non-zero eigenvalues denoted by (ρ2 j)j=1,...,r when written in a decreasing order. We set : λj=ρ2 j,for j=1, ..., r,and, except in the particular case r=p=q, we set λj=ρj=0 for j>r. For j>rwe define fjand gjas r.r.v. of HXand HYrespectively with unit variance and uncorrelated with the ( fk)k<jand (gk)k<j. CCA of (X,Y) is then : ((ρj)j=1,...,r+1,(fj)j=1,...,p,(gj)j=1,...,q).(1) CCA of (X,Y) depends only on HXand HY, which are also generated by components of X0:=V−1 2 XXand Y0:=V−1 2 YYrespectively. We show that, if (uj)j=1,...,pand (vj)j=1,...,q denote unit eigenvectors bases of RXand RYassociated with (λj)j=1,...,pand (λj)j=1,...,q respectively, we can obtain canonical variables fjand gjas linear combinations of X0 and Y0components, using ujand vjcomponents as coefficients, that is, by setting: fj=huj,X0ipet gj=hvj,Y0iq, where h., .ipand h., .iqdenote Rpand Rqusual scalar products. Decomposition (1) is not unique because each canonical variable associated with a simple eigenvalue may be replaced by its opposite and the set of canonical variables associated with a multiple eigenvalue may be replaced by an other set according to the choice of RXand RYeigenvectors associated with this eigenvalue. 1.2 Sample canonical correlation analysis Let (Xl,Yl)l=1,...,nbe a n-sample i.i.d. as (X,Y). We index by nthe elements defined previously and calculated on the sample : µn X, µn Y,Vn X,Vn Y,Vn XY ,Rn X,Rn Y. Let (λn j)j=1,...,pbe the decreasing sequence of the peigenvalues of Rn X(and of the p largest eigenvalues of Rn Y, the other ones, if q>p, being null), (un j,vn j)j=1,...,pa sequence of associated unit eigenvectors of Rn Xand of Rn Yand ( fn j,gn j)j=1,...,pthe canonical variables sequence, vectors of Rn, obtained by : Jeanne Fine 167 ∀l∈ {1, ..., n}(fn j)l=hun j,(Vn X)−1 2(Xl−µn X)ipand (gn j)l=hvn j,(Vn Y)−1 2(Yl−µn Y)iq. If the case arises (q>p), we let λn p+1=0, we complete (vn j)j=1,...,pwith Rn Y eigenvectors in order to obtain an orthonormal basis (vn j)j=1,...,qof Rqand we define the canonical variables associated. At last, for all jin {1, ..., p+1}let ρn j=qλn j. Sample CCA of (X,Y) is then: ((ρn j)j=1,...,p+1,(fn j)j=1,...,p,(gn j)j=1,...,q).(2) 1.3 Asymptotic study Asymptotic study of CCA consists in establishing a.s. convergence of the canonical elements sequences of the sample CCA (2) to the corresponding canonical elements of the population CCA (1) and in establishing convergence in distribution of the standardized canonical elements sequences. Difficulties are numerous : canonical variables are estimated (“predicted”) by Rn vectors, the space dimension increasing with sample size. Then the use is to restrict asymptotic study to Rn Xand Rn Yeigenvectors : (un j)j=1,...,pand (vn j)j=1,...,qrespectively, called canonical vectors and also to the Rpand Rqvectors defined by: xn j=(Vn X)−1 2un j and yn j=(Vn Y)−1 2vn jrespectively, called canonical factors; these vectors permit to obtain directly canonical variables by: ∀l∈ {1,...,n}(fn j)l=hxn j,Xl−µn Xipet (gn j)l=hyn j,Yl−µn Yiq. Multiple eigenvalues case is difficult to process because the eigenvectors associated with are not uniquely defined. Then the use is to restrict asymptotic study to the case where all eigenvalues are simple. Uniqueness is then verified by choosing systematically the unit vector (between the two ones) which has the first non null coordinate in respect with the canonical basis positive. As for all multidimensional analyses, covariance matrices of sample random matrices Vn X,Vn XY ,Rn X, ... are “super-matrices” (that is, matrices of matrices). Tools such as the “vec” operator which transforms matrix into vector, have been introduced in order to handle theses super-matrices; the difficulty comes from the necessity of fixing the order of lines and columns elements. In other respects, we know that the sequence ( √n(Vn X−VX)) converges in distribution to a centered normal variable, the covariance super-matrix of which is known in some special cases, when Xhas a normal or elliptical distribution for example. In the CCA frame work, we need to study convergence in distribution of the sequence (√n(Vn Z−VZ)) with Z=(X,Y), then to study convergence in distribution of the sequence (√n(Rn Z−RZ)) with RZ=(RX,RY) and Rn Z=(Rn X,Rn Y),before studying convergence 168 Asymptotic study of canonical correlation analysis: from matrix and analytic... of the Rn Zeigenelements sequences and convergence of the sample canonical elements sequences. This asymptotic study is much more complex than the one of Principal Component Analysis (PCA) because principal values and principal vectors of the X PCA are eigenelements of the Xcovariance matrix. It is only in 1999 that Anderson publishes a CCA asymptotic study when (X,Y) has a normal distribution and when all non zero eigenvalues are simple. Canonical factors components and canonical correlation coefficients of the population CCA are differentiable functions of VZ. Results are then obtained from Taylor expansions. So, this classical approach may be qualified as matricial and analytic. In order to simplify calculations, we propose to change variables from (X,Y) to (X0,Y0), which is equivalent to changing the basis in Rpand Rq. We then have: VX0=Ip, and RX=RX0=VX0Y0VY0X0, and similarly for VY0and RY. 2 Operator and tensor approach 2.1 Introduction Difficulties previously described are ensued only from the fact that matricial tool is not convenient. Working directly on linear operators in Euclidean spaces avoid indices problems and can be easily extended to an Hilbertian frame. Moreover, instead of studying eigenvectors associated with simple eigenvalues, it is possible to study eigenprojectors associated with multiple eigenvalues. Eaton (1983) also advices a “vector space approach” of the multidimensional statistics. Dauxois and Pousse (1976) enlarge the PCA definition of a Rprandom vector to a Hilbert random variable and even to a Hilbert random function, that is, a Hilbert random variable depending on a parameter in order to process temporal or spatial data. They redefine each factorial analysis (PCA, CCA, Correspondence Analysis, Discriminant Analysis, ...) in an operatorial and stochastic frame, that leads them to define, between others, nonlinear analyses. The first asymptotic study in this frame has been realized by Romain (1979) for a Hilbert random function PCA (see also Dauxois, Pousse and Romain, 1982), study completed by Arconte (1980) who also started on CCA asymptotic study but all tools were not available to continue the study. Dauxois, Romain and Viguier (1994) propose to use some tensor products and establish a dictionary between matricial and operatorial formula. This work permits to compare common results obtained in both frames, but also to obtain more easily complex results; writings in respect with eigenvectors basis are established after concise formulations with operators. These new tools permit to realize in Fine (2000) the CCA asymptotic study without restriction, that is, without assumption on the (X,Y) distribution, in the general case where eigenvalues may be multiple and without excluding canonical variables Jeanne Fine 169 asymptotic study (CCA random elements). Therefore, our approach may be qualified as an operator and tensor approach. We give below the different steps of the CCA asymptotic analysis, some tools used and some examples of results. 2.2 Different steps of the CCA asymptotic study, tools, results 1) Population CCA First, the matter is to define CCA of a pair of Euclidean random variables (population CCA). Again, we use classical approach notations substituting (Rp,h., .ip) and (Rq,h., .iq for pand qdimensional Euclidean spaces (X,h., .iX) and (Y,h., .iY) respectively. Obviously, we work without reference to any basis (free coordinate). Let L2(P) the Hilbert space of r.r.v. defined on (Ω,A,P) and admitting order 2 moments, scalar product of which associates E(f g) to ( f,g). The operator ΦXfrom Xto L2(P) which associates hx,XiXto xplays an essential role in the operator approach of multidimensional statistics. In particular, Xis a normal Euclidean random variable if, and only if, ∀x∈ X,hx,XiXis a normal r.r.v.. The expected value of Xis the unique element of X(Riesz theorem), denoted by E(X), verifying: ∀x∈ X,hx,E(X)iX=E(hx,XiX). For all (x,y)∈ X×Y, we denote by x⊗ythe operator from Xto Ywhich associates hx0,xiXyto x0; it is an element of the Hilbert space σ2(X,Y) of operators from Xto Ywith the scalar product : hA,Bi2=tr(AB∗).Due to the Riesz theorem, we may then define covariance operators VXof X,VYof Y, and crossed covariance operators VXY and VYX of Xand Y:VX=E((X−µX)⊗(X−µX)),... As in the CCA classical approach (§1.1), Xand Yare assumed to be centered. The adjoint operator Φ∗ Xof ΦXis the operator from L2(P) to Xwhich associates E(f X) to f and then we have : Φ∗ X◦ΦX=VX,Φ∗ X◦ΦY=VXY , ... It is convenient to represent operators relationships in the following commutative diagram, also called a duality scheme; here, each space is identified with its dual space. The HXand HYspaces are image spaces of ΦXand ΦYrespectively and the orthogonal projectors of L2(P) on these subspaces are: ΠX= ΦX◦V−1 X◦Φ∗ Xand ΠY= ΦY◦V−1 Y◦Φ∗ Y. X Φ∗ X ←− L2(P) Φ∗ Y −→ Y V−1 X↓↑ VX↑I VY↑↓ V−1 Y X−→ ΦXL2(P)←− ΦYY Operators RXand RY, and also CCA of (X,Y), are defined as previously (symbols ◦ are deleted in order to reduce notation). 170 Asymptotic study of canonical correlation analysis: from matrix and analytic... As in the classical approach, in order to facilitate calculations, we change the scalar product on Xso that the covariance operator of Xis the identity of X, and similarly for Y. We then have: RX=VXY VYX and RY=VYXVXY . 2) Sample model and sample CCA We use a sample model (Dauxois, Fine and Pousse, 1979) establishing a link between the sample used in Data Analysis and the i.i.d. sample of Statistics. A sample (Xl,Yl)l∈N∗ i.i.d. as (X,Y) is built from an element ωof ΩN∗setting, for all lof N∗(πldenoting the lth projection of ΩN∗onto Ω) : Xl=X◦πland Yl=Y◦πlthat is Xl(ω)=X(ωl) and Yl(ω)=Y(ωl). We then provide L2(P) with the scalar product (random scalar product as it depends on ω): ∀(f,g)∈L2(P)×L2(P),En(f g)=1 n n X l=1 f(ωl)g(ωl). Then we have: En(X)=µn X,Φn X=h., X−µn XiX,Vn X=1 n n X l=1 (Xl−µn X)⊗(Xl−µn X), . . . This sample model is the clue to distinguish randomness implied by the model (L2(P) elements) from randomness implied by sampling. It permits to obtain the canonical variables asymptotic distribution. The duality scheme of sample CCA is the same as the population one after substituting L2(P) for (L2(P),En) and indexing operators by n. Sample operators Rn Xand Rn Yand sample CCA are defined as previously. 3) Convergence of sample operators sequence Limit theorems in Euclidean or Hilbert spaces permit to obtain a.s. convergence and convergence in distribution of the covariance operators sequence without assumption on the distribution of (X,Y) except the existence of order 4 moments. For CCA, we obtain (remind we let Z=(X,Y)): Wn Z:=√n(Vn Z−VZ)D −→ WZ∼N(0 ; KZ), where KZis the covariance operator of Z⊗Z. In what concerns the sample operators Rn Z=(Rn X,Rn Y), elements of σ2(Z) (with Z=X×Y), a.s. convergence derives from the fact that it is possible to write Rn Xand Rn Y as continuous function of Vn Z. Let Un Z=√n(Rn Z−RZ) (:=(Un X,Un Y)). We write Un X= Ψn X(Wn Z) where (Ψn X) is a sequence of random operators from σ2(Z) to σ2(X) a.s. converging to ΨX. We then deduce the convergence in distribution Jeanne Fine 171 of (Un X) to UX= ΨX(WZ), centered normal variable, covariance operator of which being LX= ΨX◦KZ◦Ψ∗ X, and the same result for (Un Y) permuting Xand Yroles. The proposition used here, is easy to prove from classical results in metric spaces (Billingsley, 1968). We obtain for example: UX=−1 2(WXRX+RXWX)+WXY VYX +VXY WYX −VXY WYVYX ∼N(0 ; LX) 4) Convergence of eigenelements and CCA elements sequences Whatever may be the “factorial” method, which is an analysis or a model obtained from a spectral (or singular-value) decomposition, all results concerning eigenelements (eigenvalues, eigenprojectors, eigenvectors associated with simple eigenvalues, ...) are easily obtained thanks to perturbation theory of linear operators (Kato, 1980). In Fine (1987), this theory has been adapted to bounded perturbations that permits to use it, due to the iterated logarithm law, in the asymptotic study frame. So we obtain a.s. expansions of eigenelements of a symmetric positive operators sequence. We may also consult Dossou-Gbete and Pousse (1991) for limit results but, for the convergence in distribution of some CCA elements, limit results are not sufficient when perturbation expansions permit to conclude. For example, for the canonical factors associated to a simple eigenvalue λi, we have: xi=uibecause VX=IXand xn i=(Vn X)−1 2un iso: √n(xn i−xi)=−(Vn X)−1 2((Vn X)1 2+IX)−1[√n(Vn X−IX)]un i+[√n(un i−ui)]. We know that ( √n(Vn X−IX)) converges in distribution to WXand ( √n(un i−ui)) to SXiUXxi(with SXi=(RX−λiIX)−) but, thanks to perturbation expansions, it is possible to establish: √n(xn i−xi)D −→ 1 2WXxi+SXiUXxi∼N(0 ; LXi) 5) Asymptotic covariance operators in the elliptical case We have already seen that the asymptotic covariance operator of ( √n(Rn X−RX)) is LX= ΨX◦KZ◦Ψ∗ Xwhere KZis the asymptotic covariance operator of ( √n(Vn Z−VZ)) and where the operator ΨXfrom σ2(Z) to σ2(X) can be written explicitly. All the distribution limits of eigenelements or CCA elements sequences are centered normal variables (or function of centered normal variables), covariance operator of which being written as function of KZin the same way. Now, we may write explicitly these asymptotic covariance operators in the case where Zhas an elliptical distribution with mean µZ, covariance operator VZand kurtosis κ(real parameter, which, when it is null, leads to a N(µZ,VZ) distribution). We then know that KZis the operator from σ2(Z) to itself which associates to T: KZ(T)=(1 +κ)VZ(T+T∗)VZ+κhVZ,Ti2VZ. 172 Asymptotic study of canonical correlation analysis: from matrix and analytic... At this step, we need more algebraic tools. The tensor product in spaces of type σ2 is denoted by ˜ ⊗.For example: ∀(A,B)∈σ2(Z)×σ2(Z),∀T∈σ2(Z),A˜ ⊗B(T)=hT,Ai2B. We define also product ` ⊗in spaces of type σ2.For example: ∀(A,B)∈σ2(Z)×σ2(Z),∀T∈σ2(Z),A` ⊗B(T)=BT A∗. We define the commutation operator Cwhich associates to an operator Tits adjoint T∗. At last, we substitute the space product Z=X×Yfor the Hilbertian sum Z=X⊕Y; this permits to plunge all the operators into σ2(Z) in order to simplify notation. Projector PXfrom σ2(Z) onto σ2(X) becomes in this frame a symmetric operator of σ2(Z). The operator KZfrom σ2(Z) to itself may be written as: KZ=(1 +κ)VZ ` ⊗VZ(I+C)+κVZ˜ ⊗VZ. Let (xi)i=1,...,pbe an orthonormal basis formed by canonical factors of X, then (xi⊗xj)i,j=1,...,pis an orthonormal basis of σ2(X) and ((xi⊗xj)˜ ⊗((xk⊗xl))i,j,k,l=1,...,p is an orthonormal basis of σ2(σ2(X)). After calculations obtained in a concise way, it is easy to decompose operators in respect with this type of basis. For example, we obtain for the asymptotic covariance operator of ( √n(Rn X−RX)): LX=(1 +κ)(I+C)[−3 4R2 X ` ⊗IX+R2 X ` ⊗RX+RX ` ⊗IX−5 4RX ` ⊗RX](I+C) and, in respect with the basis of canonical factors (remember that (λj)j=1,...,pis the decreasing sequence of eigenvalues of RX): LX=1 2(1 +κ) p X j=1 p X k=1Ã−3 4λ2 j−3 4λ2 k+λ2 jλk+λjλ2 k+λj+λk−5 2λjλk! (xj⊗xk+xk⊗xj)˜ ⊗(xj⊗xk+xk⊗xj) When (X,Y) has a normal distribution and when all eigenvalues are simple, it is possible to rediscover Anderson’s results but we diverge on two of them. 6) Convergence of CCA random elements sequences As previously announced (§2.1.2) the sample model permits to obtain a.s. convergence and convergence in distribution of canonical variables sequences. We have for example, for the canonical variable associated with a simple eigenvalue λi: √n(fn i−fi)) D −→ h1 2WXxi+SXiUXxi,XiX∼N(0 ; MXi) Jeanne Fine 173 with, in the particular case where (X,Y) has an elliptical distribution: MXi =1 4(2 +3κ)fi⊗fi+(1 +κ)X j,i (1 −λi)(λi+λj−2λiλj)(λi−λj)−2fj⊗fj 7) Inferential applications and conclusion These results on CCA asymptotic study permit to tackle easily inferential applications (confidence interval estimation, statistical tests, ...) which imply CCA elements, particularly the proximity measures built on canonical correlation coefficients. See Anderson (1999) and Dauxois and Nkiet (2002). Further aspects and results may be consulted in Fine (2000). This methodological presentation shows that the operator approach performs quite well in solving asymptotic problems in multivariate statistics. 3 References Anderson, T. W. (1999). Asymptotic Theory for Canonical Correlation Analysis. Journal of Multivariate Analysis, 70, 1-29. Arconte, A. (1980). ´ Etude asymptotique de l’analyse en composantes principales et de l’analyse canonique. Th` ese de 3` eme cycle, Universit´ e de Pau et des Pays de l’Adour. Billingsley, P. (1968). Convergence of probability measures. Wiley, New York. Dauxois, J., Fine, J. and Pousse A. (1979). ´ Echantillonnage en segmentation, ´ etude de la convergence. Statistique et Analyse des Donn´ees, 3, 45-53. Dauxois, J. and Nkiet, G. M. (2002). Measures of Association for Hilbertian subspaces and some applications. Journal of Multivariate Analysis, 82, 263-298. Dauxois, J. and Pousse, A. (1976). Les analyses factorielles en calcul des probabilit´es et en statistique: essai d’´etude synth´etique. Th` ese de Doctorat d’ ´ Etat, Universit´ e Paul Sabatier, Toulouse. Dauxois, J., Pousse, A. and Romain, Y. (1982). Asymptotic theory for the principal component analysis of a vector random function; some applications to statistical inference. Journal of Multivariate Analysis, 12, 136-154. Dauxois, J., Romain, Y. and Viguier, S. (1994). Tensor products and statistics. Linear Algebra and its Applications, 210, 59-88. Dossou-Gbete, S. and Pousse, A. (1991). Asymptotic study of eigenelements of a sequence of random self adjoint operators. Statistics, 22, 479-491. Eaton, M. L. (1983). Multivariate statistics. A vector space approach. Wiley, New York. Fine, J. (1987). On the validity of the perturbation method in asymptotic theory. Statistics, 18, 401-414. —(2000). ´ Etude Asymptotique de l’Analyse Canonique. Pub. Inst. Stat. Univ. Paris, 44, 2-3, 21-72. Kato, T. (1980). Perturbation theory for linear operators. Springer-Verlag, New York. Romain, Y. (1979). ´ Etude Asymptotique des approximations par ´echantillonnage de l’analyse en composantes principales d’une fonction al´eatoire. Quelques applications. Th` ese 3` eme cycle. Universit´ e Paul Sabatier. Toulouse.