scieee AI-readable full text Open interactive document viewer

What does intrinsic mean in statistical estimation?

García, Glòria; Oller i Sala, Josep Maria

Abstract

In this paper we review different meanings of the word intrinsic in statistical estimation, focusing our attention on the use of this word in the analysis of the properties of an estimator.We review the intrinsic versions of the bias and the mean square error and results analogous to the Cram'er-Rao inequality and Rao-Blackwell theorem. Different results related to the Bernoulli and normal distributions are also considered.

Full text

Statistics & Operations Research Transactions SORT 30 (2) July-December 2006, 125-170 Statistics & Operations Research Transactions What does intrinsic mean in statistical estimation?∗ c Institut d’Estad´ ıstica de Catalunya [email protected] ISSN: 1696-2281 www.idescat.net/sort Gloria Garc´ ıa1and Josep M. Oller2 1Pompeu Fabra University, Barcelona, Spain 2University of Barcelona, Barcelona, Spain. Abstract In this paper we review different meanings of the word intrinsic in statistical estimation, focusing our attention on the use of this word in the analysis of the properties of an estimator. We review the intrinsic versions of the bias and the mean square error and results analogous to the Cram´ er-Rao inequality and Rao-Blackwell theorem. Different results related to the Bernoulli and normal distributions are also considered. MSC: 62F10, 62B10, 62A99. Keywords: Intrinsic bias, mean square Rao distance, information metric. 1 Introduction Statistical estimation is concerned with specifying, in a certain framework, a plausible probabilistic mechanism which explains observed data. The inherent nature of this problem is inductive, although the process of estimation itself is derived through mathematical deductive reasoning. In parametric statistical estimation the probability is assumed to belong to a class indexed by some parameter. Thus the inductive inferences are usually in the form of point or region estimates of the probabilistic mechanism which has generated some specific data. As these estimates are provided through the estimation of the parameter, a label of the probability, different estimators may lead to different methods of induction. *This research is partially sponsored by CGYCIT, PB96-1004-C02-01 and 1997SGR-00183 (Generalitat de Catalunya), Spain. Address for correspondence: J. M. Oller, Departament d’Estad´ ıstica, Universitat de Barcelona, Diagonal 645, 08028-Barcelona, Spain. e-mail: [email protected] Received: April 2006 126 What does intrinsic mean in statistical estimation? Under this approach an estimator should not depend on the specified parametrization of the model: this property is known as the functional invariance of an estimator. At this point, the notion of intrinsic estimation is raised for the first time: an estimator is intrinsic if it satisfies this functional invariance property, and in this way is a real probability measure estimator. On the other hand, the bias and the mean square error (MSE) are the most commonly accepted measures of the performance of an estimator. Nevertheless these concepts are clearly dependent on the model parametrization and thus unbiasedness and uniformly minimum variance estimation are non-intrinsic. It is also convenient to examine the goodness of an estimator through intrinsic conceptual tools: this is the object of the intrinsic analysis of statistical estimation introduced by Oller & Corcuera (1995) (see also Oller (1993b) and Oller (1993a)). These papers consider an intrinsic measure for the bias and the square error taking into account that a parametric statistical model with suitable regularity conditions has a natural Riemannian structure given by the information metric. In this setting, the square error loss is replaced by the square of the corresponding Riemannian distance, known as the information distance or the Rao distance, and the bias is redefined through a convenient vector field based on the geometrical properties of the model. It must be pointed out that there exist other possible intrinsic losses but the square of the Rao distance is the most natural intrinsic version of the square error. In a recent paper of Bernardo & Ju´ arez (2003), the author introduces the concept of intrinsic estimation by considering the estimator which minimizes the Bayesian risk, takingasa loss function a symmetrized versionofKullback-Leiblerdivergence (Bernardo & Rueda (2002)) and considering a reference prior based on an information-theoretic approach (Bernardo (1979) and Berger & Bernardo (1992)) which is independent of the model parametrization and in some cases coincides with the Jeffreys uniform prior distribution. In the latter case the prior, usually improper, is proportional to the Riemannian volume corresponding to the information metric (Jeffreys (1946)). This estimator is intrinsic as it does not depend on the parametrization of the model. Moreover, observe that both the loss function and the reference prior are derived just from the model and this gives rise to another notion of intrinsic: an estimation procedure is said to be intrinsic if it is formalized only in terms of the model. Observe that in the framework of information geometry, a concept is intrinsic as far as it has a well-defined geometrical meaning. In the present paper we review the basic results of the above-mentioned intrinsic analysis of the statistical estimation. We also examine, for some concrete examples, the intrinsic estimator obtained by minimizing the Bayesian risk using as an intrinsic loss the square of the Rao distance and as a reference prior the Jeffrey’s uniform prior. In each case the corresponding estimator is compared with the one obtained by Bernardo & Ju´ arez (2003). Gloria Garc´ ıa and Josep M. Oller 127 2 The intrinsic analysis As we pointed out before, the bias and mean square error are not intrinsic concepts. The aim of the intrinsic analysis of the statistical estimation, is to provide intrinsic tools for the analysis of intrinsic estimators, developing in this way a theory analogous to the classical one, based on some natural geometrical structures of the statistical models. In particular, intrinsic versions of the Cram´ er–Rao lower bound and the Rao–Blackwell theorem have been established. We first introduce some notation. Let (χ,a, µ) be a measure space and Θbe a connected open set of Rn. Consider a map f:χ×Θ−→ Rsuch that f(x, θ)≥0 and f(x, θ)µ(dx) defines a probability measure on (χ,a) to be denoted as Pθ. In the present paper a parametric statistical model is defined as the triple {(χ,a, µ); Θ;f}. We will refer to µas the reference measure of the model and to Θas the parameter space. In a general framework Θcan be any manifold modelled in a convenient space as Rn, Cn, or any Banach or Hilbert space. So even though the following results can be written with more generality, for the sake of simplicity we consider the above-mentioned form for the parameter space Θ. In that case, it is customary to use the same symbol (θ) to denote points and coordinates. Assume that the parametric statistical model is identifiable, i.e. there exists a one-toone map between parameters θand probabilities Pθ; assume also that fsatisfies the regularity conditions to guarantee that the Fisher information matrix exists and is a strictly positive definite matrix. In that case Θhas a natural Riemannian manifold structure induced by its information metric and the parametric statistical model is said to be regular. For further details, see Atkinson & Mitchel (1981), Burbea (1986), Burbea & Rao (1982) and Rao (1945), among others. As we are assuming that the model is identifiable, an estimator Uof the true probability measure based on a k-size random sample, k∈N, may be defined as a measurable map from χkto the manifold Θ, which induces a probability measure on Θknown as the image measure and denoted as νk. Observe that we are viewing Θas a manifold, not as an open set of Rn. To define the bias in an intrinsic way, we need the notion of mean or expected value for a random object valued on the manifold Θ. One way to achieve this purpose is through an affine connection on the manifold. Note that Θis equipped with Levi–Civita connection, corresponding to the Riemannian structure supplied by the information metric. Next we review the exponential map definition. Fix θin Θand let TθΘbe the tangent space at θ. Given ξ∈TθΘ, consider a geodesic curve γξ: [0,1] →Θ, starting at θand satisfying dγξ dt t=0=ξ. Such a curve exists as far as ξbelongs to an open star-shaped neighbourhood of 0 ∈TθΘ. In that case, the exponential map is defined as expθ(ξ)= γξ(1). Hereafter, we restrict our attention to the Riemannian case, denoting by k·kθthe 128 What does intrinsic mean in statistical estimation? norm at TθΘand by ρthe Riemannian distance. We define Sθ={ξ∈TθΘ:kξkθ=1} ⊂ TθΘ and for each ξ∈Sθwe define cθ(ξ)=sup{t>0 : ρ(θ, γξ(t)) =t}. If we set Dθ={tξ∈TθΘ: 0 ≤t<cθ(ξ) ; ξ∈Sθ}and Dθ=expθ(Dθ), it is well known that expθmaps Dθdiffeomorphically onto Dθ. Moreover, if the manifold is complete the boundary of Dθis mapped by the exponential map onto the boundary of Dθ, called the cut locus of θin Θ. For further details see Chavel (1993). TQ θ v 0 ξ  Θ • θ • • γ(1) v ξ γ(1) exp(·) θ ª ªξ= γlength ξ ª ªV= γlength v v-ξ ª ªV-ξ ≠η length η geodesic triangle D θ radialdistancesarepreserved Tangentspace Manifold Figure 1:The exponential map For the sake of simplicity, we shall assume that νk(Θ\Dθ)=0, whatever true probability measure in the statistical model is considered. In this case, the inverse of the exponential map, exp−1 θ, is defined νk–almost everywhere. For additional details see Chavel (1993), Hicks (1965) or Spivak (1979). For a fixed sample size k, we define the estimator vector field A as Aθ(x)=exp−1 θ(U(x)), θ ∈Θ. which is a C∞random vector field (first order contravariant tensor field) induced on the manifold through the inverse of the exponential map. For a point θ∈Θwe denote by Eθthe expectation computed with respect to the probability distribution corresponding to θ. We say that θis a mean value of Uif and Gloria Garc´ ıa and Josep M. Oller 129 . .  TQ -1 A = exp(U(x)) θ θ θ Θ θ geodesic U(x) Figure 2:Estimator vector field only if Eθ(Aθ)=0. It must be pointed out that if a Riemannian centre of mass exists, it satisfies the above condition (see Karcher (1977) and Oller & Corcuera (1995)). We say that an estimator Uis intrinsically unbiased if and only if its mean value is the true parameter. A tensorial measure of the bias is the bias vector field B, defined as Bθ=Eθ(Aθ), θ ∈Θ. An invariant bias measure is given by the scalar field kBk2defined as kBθk2 θ, θ ∈Θ. Notice that if kBk2=0, the estimator is intrinsically unbiased. The estimator vector field Aalso induces an intrinsic measure analogous to the mean square error. The Riemannian risk of U, is the scalar field defined as EθkAθk2 θ=Eθρ2(U, θ), θ ∈Θ. since kA(x)k2 θ=ρ2(U(x), θ). Notice that in the Euclidean setting the Riemannian risk coincides with the mean square error using an appropriate coordinate system. Finally note that if a mean value exists and is unique, it is natural to regard the expected value of the square of the Riemannian distance, also known as the Rao distance, between the estimated points and their mean value as an intrinsic version of the variance of the estimator. To finish this section, it is convenient to note the importance of the selection of a loss function in a statistical problem. Let us consider the estimation of the probability of success θ∈(0,1) in a binary experiment where we perform independent trials until the first success. The corresponding density of the number of is given by 130 What does intrinsic mean in statistical estimation? . .  TQ -1 B=E( )exp(U(x)) θ θθ θ Θ θ biasvector geodesic P θ U meanvalueof theimagemeasure P exp °U –1 θ θ meanvalueoftheimagemeasure Figure 3:Bias vector field f(k;θ)=(1 −θ)kθ;k=0,1,... If we restrict our attention to the class of unbiased estimators, a (classical) unbiased estimator Uof θ, must satisfy ∞ X k=0 U(k)(1 −θ)kθ=θ, ∀θ∈(0,1), where it follows that P∞ k=0U(k)(1 −θ)kis constant for all θ∈(0,1). So U(0) =1 and U(k)=0 for k≥1. In other words: when the first trial is a success, Uassigns θequal to 1; otherwise θis taken to be 0. Observe that, strictly speaking, there is no (classical) unbiased estimator for θsince Utakes values in the boundary of the parameter space (0,1). But we can still use the estimator Uin a wider setting, extending both the sample space and the parameter space. We can then compare Uwith the maximum likelihood estimator, V(k)=1/(k+1) for k≥0, in terms of the mean square error. After some straightforward calculations, we obtain Eθ(U−θ)2)=θ−θ2 Eθ(V−θ)2)=θ2+θLi2(1 −θ)+2θ2ln(θ)/(1 −θ) where Li2is the dilogarithm function. Further details on this function can be found in Abramovitz (1970), page 1004. The next figure represents both mean square error of U and V. Gloria Garc´ ıa and Josep M. Oller 131 0.0 0.05 0.10 0.15 0.20 0.25 0.0 0.2 0.4 0.6 0.8 1.0 θ Figure 4:MSE of U (dashed line) and V (solid line). It follows that there exist points in the parameter space for which the estimator U is preferable to Vsince Uscores less risk; precisely for θ∈(0,0.1606) where the upper extreme has been evaluated numerically. This admissibility contradicts the common sense that refuses U: this estimator assigns θto be 0 even when the success occurs in a finite number of trials. This points out the fact that the MSE criterion is not enough to distinguish properly between estimators. Instead of using the MSE we may compute the Riemannian risk for Uand V. In the geometric model, the Rao distance ρis given by ρ(θ1, θ2)=2argtanhp1−θ1−argtanhp1−θ2, θ1, θ2∈(0,1) which tends to +∞when θ1or θ2tend to 0. So Eθρ2(U, θ))= +∞meanwhile Eθρ2(V, θ))<+∞. The comparison in terms of Riemannian risk discards the estimator Uin favour of the maximum likelihood estimator V, as is reasonable to expect. Furthermore we can observe that the estimator U, which is classically unbiased, has infinite norm of the bias vector. So Uis not even intrinsically unbiased, in contrast to V which has finite bias vector norm. 3 Intrinsic version of classical results In this section we outline a relationship between the unbiasedness and the Riemannian risk obtaining an intrinsic version of the Cram´ er–Rao lower bound. These results are obtained through the comparison theorems of Riemannian geometry, see Chavel (1993) 132 What does intrinsic mean in statistical estimation? and Oller & Corcuera (1995). Other authors have also worked in this direction, such as Hendricks (1991), where random objects on an arbitrary manifold are considered, obtaining a version for the Cram´ er–Rao inequality in the case of unbiased estimators. Recent developments on this subject can be found in Smith (2005). Hereafter we consider the framework described in the previous section. Let Ube an estimator corresponding to the regular model {(χ,a, µ); Θ;f}, where the parameter space Θis a n–dimensional real manifold and assume that for all θ∈Θ,νk(Θ\Dθ)=0. Theorem 3.1. [Intrinsic Cram´er–Rao lower bound] Let us assume that E ρ2(U, θ) exists and the covariant derivative of E(A)exists and can be obtained by differentiating under the integral sign. Then, 1. We have Eρ2(U, θ)≥div(B)−Ediv(A)2 kn +kBk2, where div(·)stands for the divergence operator. 2. If all the sectional Riemannian curvatures K are bounded from above by a nonpositive constant Kand div(B)≥ −n, then Eρ2(U, θ)≥div(B)+1+(n−1)√−K kBkcoth√−K kBk2 kn +kBk2. 3. If all sectional Riemannian curvatures K are bounded from above by a positive constant Kand d(Θ)< π/2√K, where d(Θ)is the diameter of the manifold, and div(B)≥ −1, then Eρ2(U, θ)≥div(B)+1+(n−1)√Kd(Θ)cot√Kd(Θ)2 kn +kBk2. In particular, for intrinsically unbiased estimators, we have: 4. If all sectional Riemannian curvatures are non-positive, then Eρ2(U, θ)≥n k 5. If all sectional curvatures are less or equal than a positive constant Kand d(Θ)< π/2√K, then Eρ2(U, θ)≥1 kn The last result shows up the effect of the Riemannian sectional curvature on the precision which can be attained by an estimator. Observe also that any one–dimensional manifold corresponding to one–parameter family of probability distributions is always Euclidean and div(B)=−1; thus part 2 of Gloria Garc´ ıa and Josep M. Oller 133 Theorem (3.1) applies. There are also some well known families of probability distributions which satisfy the assumptions of this last theorem, such as the multinomial, see Atkinson & Mitchel (1981), the negative multinomial distribution, see Oller & Cuadras (1985), or the extreme value distributions, see Oller (1987), among many others. It is easy to check that in the n-variate normal case with known covariance matrix Σ, where the Rao distance is the Mahalanobis distance, the sample mean based on a sample of size kis an estimator that attains the intrinsic Cram´ er–Rao lower bound, since Eρ2(X, µ)=E(X−µ)TΣ−1(X−µ)= =Etr(Σ−1(X−µ)(X−µ)T)= =trΣ−1E(X−µ)(X−µ)T=tr(1 kI)=n k where vTis the transpose of a vector v. Next we consider a tensorial version of the Cram´ er-Rao inequality. First we define the dispersion tensor corresponding to an estimator Uas: Sθ=Eθ(Aθ⊗Aθ)∀θ∈Θ Theorem 3.2. The dispersion tensor S satisfies S≥1 kTr 2,4[G 2,2[(∇B−E(∇A)) ⊗(∇B−E(∇A))]] +B⊗B where Tr i,jand Gi,jare, respectively, the contraction and raising operators on index i, j and ∇is the covariant derivative. Here the inequality denotes that the difference between the right and the left hand side is non-negative definite. Now we study how we can decrease the mean square Rao distance of a given estimator. Classically this is achieved by taking the conditional mean value with respect to a sufficient statistic; we shall follow a similar procedure here. But now our random objects are valued on a manifold: we need to define the conditional mean value concept in this case and then obtain an intrinsic version of the Rao–Blackwell theorem. Let (χ,a,P) be a probability space. Let Mbe a n–dimensional, complete and connected Riemannian manifold. Then Mis a complete separable metric space (a Polish space) and we will have a regular version of the conditional probability of any M–valued random object fwith respect to any σ-algebra D ⊂ aon χ. In the case where the mean square of the Riemannian distance ρof fexists, we can define E(ρ2(f,m)|D)(x)=ZMρ2(t,m)Pf|D(x,dt), where x∈χ,Bis a Borelian set in Mand Pf|D(x,B) is a regular conditional probability of fgiven D. 140 What does intrinsic mean in statistical estimation? Bb ξ=Eξξb−ξ= k X t=0 (ξb(t)−ξ) k t!sin2tξ 2cos2(k−t)ξ 2 B∗ ξ=Eξ(ξ∗)−ξ= k X t=0     2 arcsin     rt n     −ξ      k t!sin2tξ 2cos2(k−t)ξ 2 The squared norm of the bias vector Bband of B∗are represented in Figures and respectively. 0.0 0.2 0.4 0.6 0.8 0.0 0.5 1.0 1.5 2.0 2.5 3.0 ξ Figure 8:kBbk2for k =1(solid line), k =2(long dashed line), k =10 (short dashed line) and k =30 (dotted line). Now, when the sample size is fixed, the intrinsic bias corresponding to ξbis greater than the intrinsic bias corresponding to ξ∗in a wide range of values of the model parameter, that is the opposite behaviour showed up by the Riemannian risk. 4.2 Normal with mean value known Let X1,...,Xkbe a random sample of size kfrom a normal distribution with known mean value µ0and standard deviation σ. Now the parameter space is Θ = (0,+∞) and the metric tensor for the N(µ0, σ) model is given by g(σ)=2 σ2 Gloria Garc´ ıa and Josep M. Oller 141 0.0 0.05 0.10 0.0 0.5 1.0 1.5 2.0 2.5 3.0 ξ Figure 9:kB∗k2k=1,2(solid line) (the same curve), k =10 (dashed line) and k =30 (dotted line). We shall assume again the Jeffreys prior distribution for σ. Thus the joint density for σand (X1,...,Xk) is proportional to 1 σk+1exp       −1 2σ2 k X i=1 (Xi−µ0)2        depending on the sample through the sufficient statistic S2=1 kPk i=1(Xi−µ0)2. When (X1,...,Xk)=(x1,...,xk) put S2=s2. As Z∞ 0 1 σk+1exp −k 2σ2s2!dσ=2k 2−1 ks2k 2 Γ k 2! the corresponding posterior distribution π(· | s2) based on the Jeffreys prior satisfies π(σ|s2)=ks2k 2 2k 2−1Γk 2 1 σk+1exp −k 2σ2s2! Denote by ρthe Rao distance for the N(µ0, σ). As we did in the previous example, instead of directly determining σb(s)=arg min σe∈(0+∞)Z+∞ 0ρ2(σe, σ)π(σ|s2)dσ we perform a change of coordinates to obtain a Cartesian coordinate system. Then we compute the Bayes estimator for the new parameter’s coordinate θ; as the estimator obtained in this way is intrinsic, we finish the argument recovering σfrom θ. Formally, the 142 What does intrinsic mean in statistical estimation? change of coordinates for which the metric tensor is constant and equal to 1 is obtained by solving the following differential equation: 1= dσ dθ!2 σ2 with the initial conditions θ(1) =0. We obtain θ=√2 ln(σ) and θ=−√2 ln(σ); we only consider the first of these two solutions. We then obtain ρ(σ1, σ2)=√2 ln σ1 σ2! =|θ1−θ2|(3) for θ1=√2lnσ1and θ2=√2logσ2. In the Cartesian setting, the Bayes estimator θb(s2) for θis the expected value of θwith respect to the corresponding posterior distribution, to be denoted as π(· | s2). The integral can be solved, after performing the change of coordinates θ=√2 ln(σ) in terms of the gamma function Γand the digamma function Ψ, that is the logarithmic derivative of Γ. Formally, θb(s2)=ks2k 2 2n−3 2Γk 2Z+∞ 0 ln(σ) σk+1exp −k 2σ2s2!dσ =√2 2 ln k 2s2!−Ψ k 2!! The Bayes estimator for σis then σb(s2)=rk 2exp −1 2Ψ k 2!! √s2 Observe that this estimator, σbis a multiple of the maximum likelihood estimator σ∗=√s2. We can evaluate the proportionality factor, for some values of n, obtaining the following table. nqk 2exp−1 2Ψk 2 nqk 2exp−1 2Ψk 2 1 1.88736 10 1.05302 2 1.33457 20 1.02574 3 1.20260 30 1.01699 4 1.14474 40 1.01268 5 1.11245 50 1.01012 6 1.09189 100 1.00503 7 1.07767 250 1.00200 8 1.06725 500 1.00100 9 1.05929 1000 1.00050 Gloria Garc´ ıa and Josep M. Oller 143 The Riemannian risk of σbis given by Eσρ2(σb, σ)=Eθ(θb−θ)2=1 2Ψ′(k 2) where Ψ′is the derivative of the digamma function. Observe that the Riemannian risk is constant on the parameter space. Additionally we can compute the square of the norm of the bias vector corresponding to σb. In the Cartesian coordinate system θand taking into account that k S2 σ2is distributed as χ2 k, we have Bb θ=Eθ(θb)−θ=Eσ√2 lnqk 2S2−Ψk 2−√2lnσ =√2Eσ      ln      rk S2 2σ2            −Ψ k 2!=0 That is, the estimator σbis intrinsically unbiased. The bias vector corresponding to σ∗ is given by B∗ θ=Eθ(θ∗)−θ=Eσ√2 ln√S2−√2lnσ =√2Eσ      ln      rS2 σ2             =1 √2 Ψ k 2!!+1 √2ln k 2! which indicates that the estimator σ∗has a non-null intrinsic bias. Furthermore, the Bayes estimator σbalso satisfies the following interesting property related to the unbiasedness: it is the equivariant estimator under the action of the multiplicative group R+that uniformly minimizes the Riemannian risk. We can summarize the current statistical problem to the model corresponding to the sufficient statistic S2which follows a gamma distribution with parameters n 2σ2and k 2. This family is invariant under the action of the multiplicative group of R+and it is straightforward to obtain that the equivariant estimators of σwhich are function of S=√S2are of the form Tλ(S)=λS, λ ∈(0,+∞) a family of estimators which contains σb=qk 2exp−1 2Ψk 2S, the Bayes estimator, and the maximum likelihood estimator σ∗=S. In order to obtain the equivariant estimator that minimizes the Riemannian risk, observe that the Rao distance (3) is an invariant loss function with respect to the induced group in the parameter space. This, together with the fact that the induced group acts transitively on the parameter space, makes the risk of any equivariant estimator to be constant on all the parameter space, among them the risk for σband for S, as it was shown before. Therefore it is enough to minimize the risk at any point of the parameter space, for instance at σ=1. We want to 144 What does intrinsic mean in statistical estimation? determine λ∗such that λ∗=arg min λ∈(0,+∞)E1ρ2(λS,1) It is easy to obtain E1ρ2(λS,1)=E1√2 ln(λS)2=1 2Ψ′k 2+1 2 Ψk 2+ln2λ2 k!2 which attains an absolute minimum at λ∗=rk 2exp −1 2Ψk 2! so that σbis the minimum Riemannian risk equivariant estimator. Finally, observethat the results in Lehmann (1951) guarantee the unbiasedness of σb, as we obtained before, since the multiplicative group R+is commutative, the induced group is transitive and σbis the equivariant estimator that uniformly minimizes the Riemannian risk. 4.3 Multivariate normal, Σ known Let us consider now the case when the sample X1,...,Xkcomes from a n-variate normal distribution with mean value µand known variance–covariance matrix Σ0, positive definite. The joint density function can be expressed as f(x1,...,xk;µ)=(2π)−nk 2|Σ0|−k 2etr −k 2Σ−1 0s2+(¯x−µ)(¯x−µ)T! where |A|denote the determinant of a matrix A, etr(A) is equal to the exponential mappingevaluated at thetrace of the matrix A, ¯xdenotes 1 kPk i=1xiand s2standsfor 1 kPk i=1(xi− ¯x)(xi−¯x)T. In this case, the metric tensor Gcoincides with Σ−1 0. Assuming the Jeffreys prior distribution for µ, the joint density for µand (X1,...,Xk)=(x1,...,xk) is proportional to |Σ0|−1 2exp −k 2(¯x−µ)TΣ−1 0(¯x−µ)! Next we compute the corresponding posterior distribution. Since ZRnexp −k 2(¯x−µ)TΣ−1 0(¯x−µ)!|Σ0|−1 2dµ= 2π k!n 2 the posterior distribution π(· | ¯x) based on the Jeffreys prior is given by π(µ|¯x)=(2π)−n 2|1 kΣ0|−1 2exp −1 2(µ−¯x)T 1 kΣ−1 0!(µ−¯x)! Gloria Garc´ ıa and Josep M. Oller 145 Observe here that the parameter’s coordinate µis already Cartesian, the Riemannian distance expressed via this coordinate system coincides with the Euclidean distance between the coordinates. Therefore, the Bayes estimator for µis precisely the sample mean ¯x. µb(¯x)=¯x which coincides with the maximum likelihood estimator. Arguments of invariance that are analogous to those in the previous example apply here, where µb=¯ Xis the minimum Riemannian risk equivariant estimator under the action of the translation group Rn. The induced group is again transitive so the risk is constant at any point of the parameter space; for simplicity we may consider µ=0. A direct computation shows that E0ρ2(¯ X,0)=E0¯ XTΣ−1 0¯ X=n k Following Lehmann, Lehmann (1951), and observing that the translation group is commutative, ¯ Xis also unbiased, as can easily be verified. References Abramovitz, M. (1970). Handbook of Mathematical Functions. New York: Dover Publications Inc. Atkinson, C. and Mitchel, A. (1981). Rao’s distance measure. Sankhy`a, 43, A, 345-365. Berger, J. O. and Bernardo, J. M. (1992). Bayesian Statistics 4. (Bernardo, Bayarri, Berger, Dawid, Hackerman, Smith & West Eds.), On the development of reference priors, (pp. 35-60 (with discussion)). Oxford: Oxford University Press. Bernardo, J. and Ju´ arez, M. (2003). Bayesian Statistics 7. (Bernardo, Bayarri, Berger, Dawid, Hackerman, Smith & West Eds.), Intrinsic Estimation, (pp. 465-476). Berlin: Oxford University Press. Bernardo, J. M. (1979). Reference posterior distributions for bayesian inference. Journal of the Royal Statistical Society B, 41, 113-147 (with discussion) Reprinted in Bayesian Inference 1 (G. C. Tiao and N. G. Polson, eds). Oxford: Edward Elgar, 229-263. Bernardo, J. M. and Rueda, R. (2002). Bayesian hypothesis testing: A reference approach. International Statistical Review, 70, 351-372. Burbea, J. (1986). Informative geometry of probability spaces. Expositiones Mathematicae, 4. Burbea, J. and Rao, C. (1982). Entropy differential metric, distance and divergence measures in probability spaces: a unified approach. Journal of Multivariate Analysis, 12, 575-596. Chavel, I. (1993). Riemannian Geometry. A Modern Introduction. Cambridge Tracts in Mathematics. Cambridge University Press. Erd´ elyi, A., Magnus, W., Oberhettinger, F., and Tricomi, F. G. (1953-1955). Higher Transcendental Functions, volume 1,2,3. New York: McGraw–Hill. Hendricks, H. (1991). A Cram´ er-Rao type lower bound for estimators with values in a manifold. Journal of Multivariate Analysis, 38, 245-261. Hicks, N. (1965). Notes on Differential Geometry. New York: Van Nostrand Reinhold. Jeffreys, H. (1946). An invariant form for the prior probability in estimation problems. Proceedings of the Royal Society, 186 A, 453-461. Karcher, H. (1977). Riemannian center of mass and mollifier smoothing. Communications on Pure and Applied Mathematics, 30, 509-541. 146 What does intrinsic mean in statistical estimation? Kendall, W. (1990). Probability, convexity and harmonic maps with small image I: uniqueness and fine existence. Proceedings of the London Mathematical Society, 61, 371-406. Lehmann, E. (1951). A general concept of unbiassedness. Annals of Mathematical Statistics, 22, 587-592. Oller, J. (1987). Information metric for extreme value and logistic probability distributions. Sankhy`a, 49 A, 17-23. Oller, J. (1993a). Multivariate Analysis: Future Directions 2. (Cuadras and Rao Eds.), On an Intrinsic analysis of statistical estimation (pp. 421-437). Amsterdam: Elsevier science publishers B. V., North Holland. Oller, J. (1993b). Stability Problems for Stochastic Models. (Kalasnikov and Zolotarev Eds.), On an Intrinsic Bias Measure (pp. 134-158). Berlin: Lect. Notes Math. 1546, Springer Verlag. Oller, J. and Corcuera, J. (1995). Intrinsic analysis of statistical estimation. Annals of Statistics, 23, 15621581. Oller, J. and Cuadras, C. (1985). Rao’s distance for multinomial negative distributions. Sankhy`a, 47 A, 7583. Rao, C. (1945). Information and accuracy attainable in the estimation of statistical parameters. Bulletin of the Calcutta Mathematical Society, 37, 81-91. Smith, S. (2005). Covariance, Subspace, and Intrinsic Cram´ er-Rao Bounds. IEEE Transactions on Signal Processing, 53, 1610-1630. Spivak, M. (1979). A Comprehensive Introduction to Differential Geometry. Berkeley: Publish or Perish. Discussion of “What does intrinsic mean in statistical estimation?” by Gloria Garc´ ıaand JosepM.Oller Hola Jacob Burbea University of Pittsburgh [email protected] In this review paper, Garc´ ıa and Oller discuss and study the concept of intrinsicality in statistical estimation, where the attention is focused on the invariance properties of an estimator. Statistical estimation is concerned with assigning a plausible probabilistic formalism that is supposed to explain the observed data. While the inference involved in such an estimation is inductive, the formulation and the derivation of the process of estimation are based on a mathematical deductive reasoning. In the analysis of this estimation, one likes to single out those estimates in which the associated estimator possesses a certain invariance property, known as the functional invariance of an estimator.Such an estimator essentially represents a well-defined probability measure, and as such it is termed as an intrinsic estimator. The present paper revolves around this concept within the framework of a parametric statistical estimation. In this context, the probability is assumed to belong to a family that is indexed by some parameter θwhich ranges in a set Θ, known as the parameter space of the statistical model. In particular, the resulting inductive inferences are usually formulated in the form of a point or a region estimates which ensue from the estimation of the parameter θ. In general, however, such an estimation depend on the particular parametrization of the model, and thus different estimators may lead to different methods of induction. In contrast, and by definition, intrinsic estimators do not depend on the specific parametrization, a feature that is significant as well as desirable. In order to develop a suitable analysis, called intrinsic analysis, of such a statistical estimation, it is required to assess the performance of the intrinsic estimators in terms of intrinsic measures or tools. At this stage, it is worthwhile to point out that, for example, the mean square error and the bias, which are commonly accepted measures of a performance of an estimator, are clearly dependent on the model parametrization. In particular, minimum variance estimation and unbiasedness are non intrinsic concepts. To avoid such situations, the intrinsic analysis exploits the natural geometrical structures that the statistical models, under some regularity conditions, possess to construct quantities which have a well-defined geometrical meaning, and hence also intrinsic. 156 Kendall, W. S. (1992). Convex geometry and nonconfluent Γ-martingales. II. Wellposedness and Gmartingale convergence. Stochastics Stochastics Reports, 38, 135-147. Kume, A. and H. Le (2003). On Fr´ echet means in simplex shape spaces. Advances in Applied Probability, 35, 885-897. Le, H. (2004). Estimation of Riemannian barycentres. LMS Journal of Computation and Mathematics,7, 193-200 (electronic). Srivastava, A. and E. Klassen (2004). Bayesian and geometric subspace tracking. Advances in Applied Probability, 36, 43-56. Steven Thomas Smith1 MIT Lincoln Laboratory, Lexington, MA 02420 [email protected] The intrinsic nature of estimation theory is fundamentally important, a fact that was realized very early in the field, as evidenced by Rao’s first and seminal paper on the subject in 1945 (Rao, 1945). Indeed, intrinsic hypothesis testing, in the hands of none other than R. A. Fisher played a central role establishing one of the greatest scientific theories of the twentieth century: Wegener’s theory of continental drift (Fisher, 1953). (Diaconis recounts the somewhat delightful way in which Fisher was introduced to this problem (Diaconis, 1988)). Yet in spite of its fundamental importance, intrinsic analysis in statistics, specifically in estimation theory, has in fact received relatively little attention, notwithstanding important contributions from Bradley Efron (Efron, 1975), Shun-ichi Amari (Amari, 1985), Josep Oller and colleagues (Oller, 1991), Harrie Hendriks (Hendriks, 1991), and several others. I attribute this limited overall familiarity with intrinsic estimation to three factors: (1) linear estimation theory, though in itself is implicitly intrinsic, is directly applicable to the vast majority of linear, or linearizable, problems encountered in statistics, physics, and engineering, obviating any direct appeal to the underlying coordinate invariance; (2) consequently, the number of problems demanding an intrinsic approach is limited, though in some fields, such as signal processing, nonlinear spaces abound (spheres, orthogonal and unitary matrices, Grassmann manifolds, Stiefel manifolds, and positive-definite matrices); (3) intrinsic estimation theory is really nonlinear estimation theory, which is hard, necessitating as it does facility with differential and Riemannian geometry, Lie groups, and homogeneous spaces—even Efron acknowledges this, admitting being “frustrated by the intricacies of the higher order differential geometry” [Efron, 1975, p. 1241]. Garc´ ıa and Oller’s review of intrinsic estimation is a commendable contribution to addressing this vital pedagogical matter, as well as providing many important insights and results on the application of intrinsic estimation. Their explanation of the significance of statistical invariance provides an excellent introduction to this hard subject. 1. This work was sponsored by DARPA under Air Force contract FA8721-05-C-0002. Opinions, interpretations, conclusions, and recommendations are those of the author and are not necessarily endorsed by the United States Government. 158 Intrinsic estimation concerns itself with multidimensional quantities that are invariant to the arbitrary choice of coordinates used to describe these dimensions— when estimating points on the globe, what should it matter if we choose “Winkel’s Tripel” projection, or the old Mercator projection? Just like the arbitrary choice of using the units of feet or meters, the numerical answers we obtain necessarily depend upon the choice of coordinates, but the answers are invariant to them because they transform according to a change of variables formula involving the Jacobian matrix. Yet the arbitrary choice of metric—feet versus meters—is crucial in that it specifies the performance measure in which all results will be expressed numerically. There are two roads one may take in the intrinsic analysis of estimation problems: the purely intrinsic approach, in which the arbitrary metric itself is chosen based upon an invariance criterion, or the general case in which the statistician, physicist, or engineer chooses the metric they wish to use, invariant or not, and demands that answers be given in units expressed using this specific metric. Garc´ ıa and Oller take the purely intrinsic path. They say, “the square of the Rao distance is the most natural intrinsic version of the square error,” and proceed to compute answers using this Rao (or Fisher information) metric throughout their analysis. One may well point out that this Fisher information metric is itself based upon a statistical model chosen using various assumptions, approximations, or, in the best instances, the physical properties of the estimation problem itself. Thus the Fisher information metric is natural insofar as the measurements adhere to the statistical model used for them, indicating a degree of arbitrariness or uncertainty even in this “most natural” choice for the metric. In addition to the choice of metric, the choice of score function with which to evaluate the estimation performance is also important. This Fisher score yields intrinsic Cram´ er-Rao bounds, and other choices of score functions yield intrinsic versions of the Weiss-Weinstein, Bhattacharyya, Barankin and Bobrovsky-Zakai bounds (Smith, Scharf and McWhorter, 2006). Moreover, there may be legitimately competing notions of natural invariance. The invariance that arises from the Fisher metric is one kind, as recognized. But in the cases where the parameter space is a Lie group Gor a (reductive) homogeneous space G/H(H⊂Ga Lie subgroup), such as found in the examples of unitary matrices, spheres, etc. cited above, invariance to transformations by the Lie group Gis typically of principal physical importance to the problem. In such cases, we may wish to analyze the square error using the unique invariant Riemannian metric on this space, e.g., the square error on the sphere would be great circle distances, not distances measured using the Fisher information metric. In most cases, this natural, intrinsic metric (w.r.t. Ggroup invariance) is quite different from the natural, intrinsic (w.r.t. the statistical model) Fisher metric. I am aware of only one nontrivial example where these coincide: the natural Gl(n,C)-invariant Riemannian metric for the space of positive-definite matrices Gl(n,C)/U(n) is the very same as the Fisher information metric for Gaussian covariance 159 matrix estimation [Smith, 2005, p. 1620]: gR(A,A)=tr(AR−1)2(1) (ignoring an arbitrary scale factor). In this example, the square error between two covariance matrices is given by, using Matlab notation and decibel units, d(R1,R2)=norm(10 ∗log 10(eig(R1,R2))) (2) (i.e., the logarithm of the 2-norm of a vector of generalized eigenvalues between the covariance matrices), precisely akin to the distance between variances in Equation (3) from Garc´ ıa and Oller’s review. Furthermore, and perhaps most importantly, the statistician or engineer may well respond “So what?” to this metric’s invariance, whether it be G-invariance or invariance arising from the statistical model. For example, Hendriks (Hendriks, 1991) considers the embedded metric for a parameter manifold embedded in a higher dimensional Euclidean space; the extrinsic Euclidean metric possesses very nice invariance properties, but these are typically lost within arbitrary constrained submanifolds. The choice of metric is, in fact, arbitrary, and there may be many good practical reasons to express one’s answers using some other metric instead of the most natural, invariant one. Or not— the appropriateness of any proposed metric must be assessed in the context of the specific problem at hand. For these reasons, an analysis of intrinsic estimation that allows for arbitrary distance metrics is of some interest (Smith, 2005), (Smith, Scharf and McWhorter, 2006), as well the special and important special case of the Fisher information metric. Another critical factor affecting the results obtained is the weapon one chooses from the differential geometric arsenal. Garc´ ıa and Oller present results obtained from the powerful viewpoint of comparison theory (Cheeger and Ebin, 1975), a global analysis that uses bounds on a manifold’s sectional curvature to compare its global structure to various model spaces. The estimation bounds derived using these methods possess two noteworthy properties. First, Oller and Corcuera’s expressions (Oller and Corcuera, 1995) are remarkably simple! It is worthwhile paraphrasing these bounds here. Let θ θ θbe an unknown n-dimensional parameter, ˆ θ θ θ(z) an estimator that depends upon the data z,A(z|θ θ θ)=exp−1 θ θ θˆ θ θ θthe (random) vector field representing the difference between the estimator ˆ θ θ θand the truth θ θ θ,b(θ θ θ)=E[A] the bias vector field, “expθ θ θ”the Riemannian exponential map w.r.t. the Fisher metric, ¯ Kdef =maxθ θ θ,HKθ θ θ(H) the maximum sectional curvature of the parameter manifold over all two-dimensional subspaces H, and D=supθ θ θ1,θ θ θ2d(θ θ θ1,θ θ θ2) be the manifold’s diameter. Then, ignoring the sample size kwhich will always appear in the denominator for independent samples, Theorem 3.1 implies that the mean square error (MSE) about the bias—which is the random part of 160 the bound (add b2for the MSE about θ θ θ)—is EA−b2≥ ⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ n−1(div E[A]−E[div A])2, (general case); n−1div b+1+(n−1)|¯ K|1 2bcoth|¯ K|1 2b2, ¯ K≤0, div b≥−n; n−1div b+1+(n−1) ¯ K1 2Dcot( ¯ K1 2D)2, ¯ K≥0, D<π/(2 ¯ K1 2), div b≥−1; n,¯ K≤0; n−1,¯ K≥0, D<π/(2 ¯ K1 2), (3) where div b(θ θ θ)=1 |g|1 2 i ∂|g|1 2bi(θ θ θ) ∂θi(4) is the divergence of the vector field b(θ θ θ) w.r.t. the Fisher metric G(θ θ θ) and a particular choice of coordinates (θ1,θ 2,...,θ n), and |g|1 2=|det G(θ θ θ)|1 2is the natural Riemannian volume form, also w.r.t. the Fisher metric. Compare these relatively simple expressions to the ones I obtain for an arbitrary metric [Smith, 2005, Theorem 2]. In these bounds, the covariance Cof A−babout the bias is given by C≥MbG−1MT b−1 3Rm(MbG−1MT b)G−1MT b+MbG−1Rm(MbG−1MT b)T(5) (ignoring negligible higher order terms), where Mb=I−1 3b2K(b)+∇b,(6) (G)ij =g(∂/∂θi,∂/∂θj) is the Fisher information matrix, Iis the identity matrix, (∇b)ij=(∂bi/∂θj)+kΓi jk bkis the covariant differential of b(θ θ θ), Γi jk are the Christoffel symbols, and the matrices K(b)andRm(C) representing sectional and Riemannian curvature terms are defined by K(b)ij =⎧ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩ sin2αi·K(b∧Ei)+O(b3),if i=j; sin2α ij·Kb∧(Ei+Ej) −sin2α ij·Kb∧(Ei−Ej) +O(b3), if ij,(7) 161 αi,α ij,andα ij are the angles between the tangent vector band the orthonormal tangent basis vectors Ei=∂/∂θi,Ei+Ej,andEi−Ej, respectively, and Rm(C) is the mean Riemannian curvature defined by the equality Rm(C)Ω,Ω=ER(X−b,Ω)Ω,X−b,(8) where R(X,Y)Zis the Riemannian curvature tensor. Ignoring curvature, which is reasonable for small errors and biases, as well as the intrinsically local nature of the Cram´ er-Rao bound itself, the intrinsic Cram´ er-Rao bound simplifies to the expression C≥(I+∇b)G−1(I+∇b)T.(9) The trace of Equations (5) and (9) (plus the square length b2) provides the intrinsic Cram´ er-Rao bound on the mean square error between the estimator ˆ θ θ θand the true parameter θ θ θ. Let’s take a breath and step back to compare how these bounds relate to one another, beginning with the simplest, one-dimensional Euclidean case as the basis for our comparison. The biased Cram´ er-Rao bound for this case is provided in an exercise from Van Trees’ excellent reference [Van Trees, 1968, p. 146f]: the variance of any estimator ˆ θof θwith bias b(θ)=E[ˆ θ]−θis bounded from below by Var( ˆ θ−θ−b)≥(1 +∂b/∂θ)2 E(∂log f(z|θ)/∂θ)2.(10) The denominator is, of course, the Fisher information. The numerator, (1 +∂b/∂θ)2, is seen in all Equations (3)–(9) above. The term div E[A]−E[div A]appearinginthe numerator of Equation (3) is, in this context, precisely div E[A]−E[div A]=∂b/∂θ −E[(∂/∂θ)(ˆ θ−θ)] (11) =1+∂b/∂θ, (12) i.e., the one-dimensional biased CRB, as promised. Likewise, the expression I+∇bfrom Equation (9) reduces to this simplest biased CRB form as well. In the general Euclidean case with a Gaussian statistical model log f(z|μ μ μ)=−(z−μ μ μ)TR−1(z−μ μ μ)+constants, the mean square error of any unbiased estimator of the mean E[z]=μ μ μ(with known covariance) is bounded by E(ˆ μ μ μ−μ μ μ)T(ˆ μ μ μ−μ μ μ)≥tr R.(13) Note that this is the mean square error measured using the canonical Euclidean metric μ μ μ2 2=iμ2 i, i.e., the 2-norm. As discussed above, Garc´ ıa and Oller use the intrinsic 162 Fisher metric to measure the mean square error, which results in the simpler expression E(ˆ μ μ μ−μ μ μ)TR−1(ˆ μ μ μ−μ μ μ)≥tr I=n,(14) which is seen in part 4 of Equation (3), as well as in Garc´ ıa and Oller’s paper. It may be argued that the relatively simple expressions for the bounds of Equation (3) actually conceal the true complexity found in these equations because these bounds depend on the sectional curvature and diameter of the underlying Riemannian manifold, which, especially for the case of the Fisher information metric, is typically quite involved, even for the simplest examples of Gaussian densities. The second noteworthy property of the bounds of Equation (3) is their global, versus local, properties. It must be acknowledged that parts of these bounds, based as they are upon the global analysis of comparison theory (specifically, Bishop’s comparison theorems I and II (Oller and Corcuera, 1995)), are not truly local bounds. As the square error becomes small (i.e., as either the sample size or signal-to-noise-ratio becomes large) relative to the inverse of the local curvature, the manifold appears to be locally flat (as the earth appears flat to us), and the curvature terms for a truly local Cram´ er-Rao bound should become increasingly negligible and approach the standard Euclidean case. This local effect is not captured by the global comparison theory, and explains why the curvature terms remain present for errors of all sizes in Equation (3). It also explains why a small change of curvature, i.e., a small change in the assumed statistical model, will result in very different numerical bounds, depending upon which of the five cases from Equation (3) applies, i.e., these bounds are not “tight”—they fall strictly below the asymptotic error of the (asymptotically efficient) maximum likelihood estimator. To see this, consider the 2-dimensional unit disk. It is flat; therefore, the fourth part of the bound applies, i.e., the lower bound on the square error is 2/k. Now deform the unit disk a little so that it has a small amount of positive curvature, i.e., so that it is a small subset of a much larger 2-sphere. Now the fifth part of Equation (3) applies, and no matter how tiny the positive curvature, the lower bound on the square error is 1/(2k). This is a lower bound on the error which is discontinuous as a function of curvature. A tighter estimation bound for a warped unit disk, indeed for the case of general Riemannian manifolds, is possible. Nevertheless, the relative simplicity of these bounds lends themselves to rapid analysis, as well as insight and understanding, of estimation problems whose performance metric is Fisher information. The tight bounds of Equations (5)–(9) can be achieved using Riemann’s original local analysis of curvature (Spivak, 1999), which results in a Taylor series expansion for the Cram´ er-Rao bound with Riemannian curvature terms appearing in the second-order part. These are truly local Cram´ er-Rao bounds, in that as the square error becomes small relative to the inverse of the local curvature, the curvature terms become negligible, and the bounds approach the classical, Euclidean Cram´ er-Rao bound. Typical Euclidean 163 Cram´ er-Rao bounds take the form (Van Trees, 1968) Var( ˆ θ−θ−b)≥beamwidth2 SNR ,(15) where the “beamwidth” is a constant that depends upon the physical parameters of the measurement system, such as aperture and wavelength, and “SNR” denotes the signalto-noise-ratio of average signal power divided by average noise power, typically the power in the deterministic part of the measurement divided by the power in its random part. The intrinsic Cram´ er-Rao bounds of (5)–(9) take the form (loosely speaking via dimensional analysis and an approximation of A=exp−1 θ θ θˆ θ θ θ≈ˆ θ θ θ−θ θ θusing local coordinates) Cov(ˆ θ θ θ−θ θ θ−b)beamwidth2 SNR 1−beamwidth2·curvature SNR +O(SNR−3).(16) Note that the curvature term in Equation (16), the second term in a Laurent expansion about infinite SNR, approaches zero faster than the bound itself; therefore, this expression approaches the classical result as the SNR grows large—the manifold becomes flatter and flatter and its curvature becomes negligent. The same is true of the sample size. Also note that, as the curvature is a local phenomenon, these assertions all depend precisely where on the parameter manifold the estimation is being performed, i.e., the results are local ones. In addition, the curvature term decreases the Cram´ erRao bound by some amount where the curvature is positive, as should be expected because geodesics tend to coalesce in these locations, thereby decreasing the square error, and this term increases the bound where there is negative curvature, which is also as expected because geodesics tend to diverge in these areas, thereby increasing the square error. Finally, even though this intrinsic Cram´ er-Rao bound is relatively involved compared to Equation (3) because it includes local sectional and Riemannian curvature terms, the formulae for these curvatures is relatively simple in the case of the natural invariant metric on reductive homogeneous spaces (Cheeger and Ebin, 1975), arguably a desirable metric for many applications, and may also reduce to simple bounds [Smith, 2005, pp. 1623–24]. These comments have focused on but a portion of the wide range of interesting subjects covered well in Garc´ ıa and Oller’s review. Another important feature of their article that warrants attention is the discussion at the end of section 4.2 about obtaining estimators that minimize Riemannian risk. It is worthwhile comparing these results to the problem of estimating an unknown covariance matrix, analyzed using intrinsic methods on the space of positive definite matrices Gl(n,C)/U(n) (Smith, 2005). As noted above, the Fisher information metric and the natural invariant Riemannian metric on this space coincide, hence Garc´ ıa and Oller’s risk minimizing estimator analysis may be applied directly to the covariance matrix estimation problem as well. Furthermore, the 164 development of Riemannian risk minimization for arbitrary metrics, not just the Fisher one, appears promising. This body of work points to an important, yet unanswered question in this field: Aside from intrinsic estimation bounds using various distance metrics, does the intrinsic approach yield practical and useful results useful to the community at large? Proven utility will drive greater advances in this exciting field. References Amari, S. (1985). Differential-Geometrical Methods in Statistics. Lecture Notes in Statistics, 28. Berlin: Springer-Verlag. Amari, S. (1993). Methods of Information Geometry, Translations of Mathematical Monographs, 191 (transl. D. Harada). Providence, RI: American Mathematical Society, 2000. Originally published in Japanese as “Joho kika no hoho” by Iwanami Shoten, Publishers, Tokyo. Cheeger, J. and Ebin, D. G. (1975). Comparison Theorems in Riemannian Geometry. Amsterdam: NorthHolland Publishing Company. Diaconis, P. (1988). Group Representations in Probability and Statistics. IMS Lecture Notes-Monograph Series, 11, ed. S. S. Gupta. Hayward, CA: Institute of Mathematical Statistics. Efron, B. (1975). Defining the curvature of a statistical problem (with applications to second order efficiency) (with discussion). Annals of Statistics, 3, 1189-1242. Fisher, R. A. (1953). Dispersion on a sphere. Proceedings of the Royal Society London A, 217, 295-305. Hendriks, H. (1991). A Cram´ er-Rao type lower bound for estimators with values on a manifold. Journal of Multivariate Analysis, 38, 245-261. Oller, J. M. (1991). On an intrinsic bias measure. In Stability Problems for Stochastic Models, Lecture Notes in Mathematics 1546, eds. V. V. Kalashnikov and V. M. Zolotarev, Suzdal, Russia, 134–158. Berlin: Springer-Verlag. Oller, J. M. and Corcuera, J. M. (1995). Intrinsic analysis of statistical estimation. Annals of Statistics, 23, 1562–1581. Rao, C. R. (1945). Information and the accuracy attainable in the estimation of statistical parameters. Bulletin of the Calcutta Mathematical Society, 37, 81–89, 1945. See also Selected Papers of C. R. Rao, ed. S. Das Gupta, New York: John Wiley, Inc. 1994. Smith, S. T. (2005). Covariance, subspace, and intrinsic Cram´ er-Rao bounds. IEEE Transactions Signal Processing, 53, 1610–1630. Smith, S. T., Scharf, L. and McWhorter, L. T. (2006). Intrinsic quadratic performance bounds on manifolds. In Proceedings of the 2006 IEEE Conference on Acoustics, Speech, and Signal Processing (Toulouse, France). Spivak, M. (1999). A Comprehensive Introduction to Differential Geometry, 3d ed., vol. 2. Houston, TX: Publish or Perish, Inc. Van Trees, H. L. (1968). Detection, Estimation, and Modulation Theory,Part1.NewYork:JohnWiley, Inc. Rejoinder It is not frequent to have the opportunity to discuss some fundamental aspects of statistics with mathematicians and statisticians such as Jacob Burbea, Joan del Castillo, Wilfrid S. Kendall and Steven T. Smith, in particular those aspects considered in the present paper. We are really very fortunate in this respect and thus we want to take this opportunity to thank the former Editor of SORT, Carles Cuadras, for making it possible. Some questions have arisen on the naturalness of the information metric and thus on the adequacy of the intrinsic bias and Riemannian risk in measuring the behaviour of an estimator, see the comments of Professors S.T. Smith and J. del Castillo. Therefore it might be useful to briefly reexamine some reasons to select such metric structure for the parameter space. The distance concept is widely used in data analysis and applied research to study the dissimilarity between physical objects: it is considered a measure on the information available about their differences. The distance is to be used subsequently to set up the relationship among the objects studied, whether by standard statistical inference or descriptive methods. In order to define a distance properly, from a methodological point of view, it seems reasonable to pay attention to the formal properties of the underlying observation process rather than taking into account the physical nature of the objects studied. The distance is a property neither of physical objects nor of an observer: it is a property of the observation process. From these considerations, a necessary first step is to associate to every physical object an adequate mathematical object. Usually suitable mathematical objects for this purpose are probability measures that quantify the propensity to happen of the different events corresponding to the underlying observation process. In a more general context we may assume that we have some additional information concerning the observation process which allows us to restrict the set of probability measures that represents our possible physical objects. These ideas lead us to consider parametric statistical models to describe the knowledge of all our possible universe of study. The question is now: which is the most convenient distance between the probability measures corresponding to a statistical model? Observe that this question pressuposes that the right answer depends on the statistical model considered: if we can assume that all the possible probability measures corresponding to a statistical analysis are of a given type (at least approximately) this information should be relevant in order to quantify, in a right scale, the differences between two probability measures of the model. In other