scieee AI-readable full text Open interactive document viewer

The h-index as an almost-exact function of some basic statistics

Bertoli-Barsotti, Lucio

Abstract

As is known, the h-index, h, is an exact function of the citation pattern. At the same time, and more generally, it is recognized that h is "loosely" related to the values of some basic statistics, such as the number of publications and the number of citations. In the present study we introduce a formula that expresses the h-index as an almost-exact function of some (four) basic statistics. On the basis of an empirical study-in which we consider citation data obtained from two different lists of journals from two quite different scientific fields-we provide evidence that our ready-to-use formula is able to predict the h-index very accurately (at least for practical purposes). For comparative reasons, alternative estimators of the h-index have been considered and their performance evaluated by drawing on the same dataset. We conclude that, in addition to its own interest, as an effective proxy representation of the h-index, the formula introduced may provide new insights into "factors" determining the value of the h-index, and how they interact with each other.

Full text

The h-index as an almost-exact function of some basic statistics Lucio Bertoli-Barsotti 1 •Tommaso Lando 2 Received: 17 July 2017 / Published online: 9 September 2017 ÓThe Author(s) 2017. This article is an open access publication Abstract As is known, the h-index, h, is an exact function of the citation pattern. At the same time, and more generally, it is recognized that his ‘‘loosely’’ related to the values of some basic statistics, such as the number of publications and the number of citations. In the present study we introduce a formula that expresses the h-index as an almost-exact function of some (four) basic statistics. On the basis of an empirical study—in which we consider citation data obtained from two different lists of journals from two quite different scientific fields—we provide evidence that our ready-to-use formula is able to predict the h-index very accurately (at least for practical purposes). For comparative reasons, alternative estimators of the h-index have been considered and their performance evaluated by drawing on the same dataset. We conclude that, in addition to its own interest, as an effective proxy representation of the h-index, the formula introduced may provide new insights into ‘‘factors’’ determining the value of the h-index, and how they interact with each other. Keywords h-Index Journal ranking Weibull distribution Lambert Wfunction Mathematical Subject Classification 62P99 JEL Classification C46 &Lucio Bertoli-Barsotti [email protected] 1 Department of Management, Economics and Quantitative Methods, University of Bergamo, Via dei Caniana 2, 24127 Bergamo, Italy 2 Department of Finance, VS ˇB -TU Ostrava, Sokolska `33, 70121 Ostrava, Czech Republic 123 Scientometrics (2017) 113:1209–1228 DOI 10.1007/s11192-017-2508-6 Introduction The purpose of this paper is to present a formula with which to determine (estimate) the hindex, h, under incomplete information conditions (IIC). By IIC we mean the situation in which, for different kinds of reasons, we do not know the whole set of citation data, the entire citation profile that would allow us to obtain the actual exact value of the h-index. This is the case, for example, when only few ‘‘basic’’ citation statistics (other than the hindex) are published, or known to us. To be concrete, we will refer to simple citation indicators—to use the words of Hirsch (2005), ‘‘single-number criteria commonly used to evaluate scientific output’’—as: 1. total number of citations C; 2. total number of citations for the t(t21;2;3;... fg ) most-cited publications, Ct; thus, Ct¼Pt i¼1ciðÞ, where ciðÞrepresents the number of citations to publication i, and where publications are ranked in decreasing order of the number of citations: c1ðÞc2ð ÞcTðÞ. 3. total number of publications T; 4. total number of ‘‘significant’’ publications, that is, those with at least a predetermined number of citations keach (k21;2;3;... fg ), Tk. In this paper we focus on these indicators in their simplest versions, that is: C,C1,Tand T1. The purpose of the analysis is twofold: to estimate the h-index (when it cannot be determined directly from the data) and hence at the same time to identify the main factors which influence the level of the h-index. A crucial question is therefore the extent to which the h-index can be satisfactorily predicted from knowledge of only the above basic statistics—i.e. under IIC. More formally, we are searching for a formula ^ h¼^ hS 1;...;Sr ðÞ;ð1Þ 1r4, Sj2S,1jr, where S¼ C;C1;T;T1 fg . To be noted is that the formula ^ hcan be interpreted as a genuine estimator of the h-index, h, i.e. ^ hffih, because it does not depend on values of unknown parameters. Possible estimators under IIC of the h-index can be found in the literature: – A very simple proxy for the h-index is given by hH¼ffiffiffiffiffiffiffiffiffi C=a p. This model, which can be traced back to Hirsch (2005), is not a genuine estimator of the h-index because hHis still a function of an unknown parameter, a, and it is not specified (by the formula itself) how to estimate this parameter in terms of the above basic statistics. Nevertheless, an estimator for the h-index can be obtained by substituting the unknown parameter awith a fixed constant (Hirsch found ‘‘empirically’’ that alay between 3 and 5). Redner (2010) found that ‘‘ ffiffiffiffi C pis essentially equivalent to the hindex, up to an overall factor that is close to 2’’ (put otherwise, he found that the distribution ratio ffiffiffiffi C p=2hhas an empirical distribution ‘‘sharply peaked about 1’’). This suggests the approximating formula ^ h¼hR¼ffiffiffiffi C p=2ð2Þ with r¼1, S¼ C fg , which we could then call the Redner formula—probably the simplest estimator of the h-index, under IIC. 1210 Scientometrics (2017) 113:1209–1228 123 – While hRis a model-free proxy for the h-index, more elaborate solutions has been attempted in the literature by assuming specific probabilistic distributions for the citation rate. For example, a formula that follows model (1), with r¼4, has been recently introduced by Bertoli-Barsotti and Lando (2017), ^ h¼~ h1ðÞ W¼1 log 1 ~ m1 1  WT1 1~ m1 1log 1 ~ m1 1   ;ð3Þ where ~ m1¼CC1 ðÞ=T11ðÞis nothing but a ‘‘trimmed’’ version of the simple sample mean C=T1, and where WðÞrepresents the so-called Lambert-Wfunction (Corless and Jeffrey 2015). The Lambert-Wfunction is the function WzðÞsatisfying z¼WzðÞeWzðÞ , and can be currently computed using mathematical software, for example the Mathematica Ò software package (Wolfram Research, Inc. 2014), or the R statistical computing environment (R Development Core Team 2012). The use of a ‘‘trimmed’’ version of the sample mean is a simple technique with which to make the sample mean more robust with respect to a single outlier—a single highly-cited paper that could substantially inflate the mean, as is well known. Formula ~ h1ðÞ Wðr¼4;S¼ C;C1;T;T1 fgÞis based on the assumption that the citation rate of papers (cited at least once) follows a shifted-geometric distribution (SGD) with parameter QðQ[1Þwith probability function pyðÞ¼QyQ1ðÞ y1,y¼1;2;...;pyðÞ represents the probability of observing the number of citations yof a paper (cited at least once), while Qrepresents the expectation of the SGD. Then, ^ nyðÞ¼Tp yðÞexpresses the ‘‘expected’’/estimated number of articles with ycitations. – As an alternative approach, an important class of models is the one defined by the formula ^ h¼c0C2=3T1=3ð4Þ where c0is a fixed and known positive constant (Schubert and Gla ¨nzel 2007). From model (4), specific ready-to-use formulas are obtained by taking, in particular: (a) c0¼41=3(Iglesias and Pecharroman 2007; see also Ionescu and Chopard 2013; Panaretos and Malesios 2009; Vinkler 2009,2013), (b) c0¼0:75 (Schubert and Gla ¨nzel 2007), (c) c0¼1 Prathap (2010a,b). Following the notation of Bertoli-Barsotti and Lando (2017), let hSG c0 ðÞ¼c0C2=3T1=3. Note that these formulas are functions of the data only through two out of the four basic statistics (r¼2, S¼ C;T fg ), and they are based on the assumption of a continuous-type distribution. The formula hSG 1ðÞis also known as the ‘‘p-index’’ (Prathap 2010a,b). – Another approach which deserves mention for completeness, even if it does not yield a ready-to-use formula, is that proposed by Iglesias and Pecharroman (2007). Adopting a different perspective, i.e. the rank-size formulation, and starting from the assumption that the number ckðÞof citations of the paper of rank k, is approximately distributed following a stretched exponential type PDF fk;g;bðÞ¼Cg1=bC1þb1  1exp gkb  ;k[0;ð5Þ (not to be confused with a Weibull PDF, see below), Iglesias and Pecharroman suggest deriving a formula for the h-index as the solution of the equation Scientometrics (2017) 113:1209–1228 1211 123 fx;g;bðÞ¼x:ð6Þ Interestingly, the solution may be derived in closed form (even if authors did not realize this) by means of the Lambert-Wfunction. Unfortunately, this solution still depends on the value of an unknown free parameter, specifically b[see their Eqs. (16) and (17)]. Hence, their formula could become a genuine estimator of the h-index—of the form ^ h¼^ hC;T;T1 ðÞ,r¼3—only by constraining the unknown parameter bto assume a fixed (but arbitrary) value b0. A new formula for the h-index under the Weibull assumption Let Ny ðÞbe the empirical citation distribution function, i.e. the function giving the number of papers which have been cited ytimes at most. Then, in particular, nyðÞ¼NyðÞNy1ðÞ, for y¼1;2;...,n0ðÞ¼N0ðÞ, is the number of papers that have been cited exactly ytimes. We assume that the citation rate of a paper is a random variable Xthat is distributed as a two-parameter Weibull distribution, with CDF Fx;a;bðÞ¼1exp axb  ,x[0, and 0 otherwise, where a[0 and b[0. The probability density function is then fx;a;bðÞ¼abxb1exp axb  ;ð7Þ for x[0, and 0 otherwise. The Weibull distribution is a rather flexible model: the PDF is reverse J-shaped for b1 and bell-shaped otherwise. Since our assumption involves a continuous distribution, a suitable discretization rule is needed. In particular, for every y,y¼0;1;2;..., let Texp ayb  express the ‘‘expected’’ number of articles with at least ycitations. Hence, ^ ny ðÞ¼ TRyþ1 yfx;a;bðÞdx¼TFyþ1;a;bðÞFy;a;bðÞðÞrepresents the expected number of articles with ycitations exactly, and ^ Ny ðÞ¼TF y þ1;a;b ðÞ the expected number of papers which have been cited ytimes at most. As a special case, F1;a;b ðÞ F0;a;b ðÞ ¼1eað8Þ can be interpreted as a model for the so-called uncitedness factor,TT1 T¼n0ðÞ T(Hsu and Huang 2012; see also Egghe 2013; Burrell 2013). A Weibull model for the h-index is then yielded by the solution of the equation Texp axb  ¼x;x2<.ð9Þ Replacing axbwith tin the equation, we have tebt¼aTb:ð10Þ Thus, replacing btwith s, we obtain the equivalent equation ses¼abTb:ð11Þ Hence, by definition of the above mentioned Lambert-Wfunction, we find the solution s¼WabTb  and, since x¼s ab  1=b , we finally arrive at the formula 1212 Scientometrics (2017) 113:1209–1228 123 x¼WabTb  ab  1=b :ð12Þ An empirical counterpart of the above theoretical model for the h index may now be obtained by substituting the parameters aand bwith estimates, aand b, based on suitable functions of the citation data only through the basic statistics C;C1;Tand T1. This can be done firstly by using the uncitedness factor to derive the equation 1 ea¼TT1 T, that can be solved (under the assumption 0\T1\T) for the variable aas a¼log T T1  ;ð13Þ as an estimate of parameter a, and secondly, by using the trimmed sample citation rate, m¼CC1 T1þ0:5;ð14Þ as an estimate of the expectation of X, that is EXðÞ¼ga;bðÞ¼a1=bC1þ1 b  [0. Note that, by construction, our approximation slightly overestimates the true average number of citations, so that a correction for continuity by one-half is needed. We then find bas the solution (method of moments) of the equation m¼ga ;bðÞ;ð15Þ that can be solved numerically. It should be noted that the existence and uniqueness of the solution of Eq. (15) are not always warranted a priori. Indeed, it can be proved that the necessary and sufficient condition for existence and uniqueness of the solution is m[1 (see ‘‘Appendix’’). We should then consider ‘‘out of range’’ the cases where m1, and exclude them from the analysis. With aand breplaced by a¼aT;T1 ðÞand b¼bC;C1;TðÞin formula (12) one finally obtains (r¼4, S¼ C;C1;T;T1 fg ) ^ h¼hWW ¼Wa bTb  ab ! 1=b ;ð16Þ where the suffix WW is motivated by the fact that the formula is based on a Weibull distribution and on the Lambert-Wfunction. Analysis Two datasets This section empirically investigates the effectiveness of formula hWW as an estimate of the actual value of the h-index, h. We will compare estimates derived from hWW with the real values of the h-index. In order to facilitate possible comparisons with other formulas (see below), we choose to use the same two datasets as in Bertoli-Barsotti and Lando (2017), where the authors present an empirical study based on citation data obtained from two different sets of journals belonging to two different scientific fields: (1) the S&MM list and (2) the EE&F list. Scientometrics (2017) 113:1209–1228 1213 123 1. S&MM list The former dataset includes the 231 journals as selected from a former list of 568 journals identified as important (in the opinion of a group of experts) in the area ‘‘Statistics and Mathematical Methods’’ (S&MM). Overall, the S&MM dataset included 485,628 citations of 99,409 publications from these journals (for details see Bertoli-Barsotti and Lando 2017). For each journal, the actual value hof the h-index was computed—on the basis of citations retrieved from the Scopus database in last week of December 2015—as the largest number of papers published in the journal between 2010 and 2014 and which obtained at least hcitations each, from the time of publication until December 2015. Thus, citation data referred to a 6-year citation window, 2010–2015, and a 5-year publication window, 2010–2014. The four basic statistics C,C1,Tand T1were derived as well. The list of the 231 journals in the S&MM dataset is reported in Table 1. 2. EE&F list The second dataset included the 100 journals (with a minimum number of 50 publications) top ranked according to the Scopus Impact per Publication (IPP; the IPP is defined as the ratio of citations in a year to papers published in the three previous years divided by the number of papers published in those same years) in 2014, within the Scopus subject area of ‘‘Economics, Econometrics and Finance’’ (EE&F). The citation data of all 100 journals in the EE&F list were retrieved during the last week of April 2016. The dataset obtained included 19,889 publications receiving a total of 74,096 citations. In this case, differently from the above dataset, in order to obtain citation and publication windows as similar as possible to those employed for the computation of the IPP 2014 by Scopus, the citations used were those received during 2014 of papers published within the previous 3 years 2011–2013 (for further details see Bertoli-Barsotti and Lando 2017). For each journal the actual value hof the h-index was then computed as the largest number of papers published in the journal between 2011 and 2013 and which obtained at least hcitations each in the year 2014. The list of the journals in the EE&F dataset is reported in Table 2. Estimation of the h-index with the formula hWW Table 1for the S&MM list and Table 2for the EE&F list report, for each journal, identified by its ISSN code, the four basic statistics, C,C1,Tand T1, the h-index, h,as computed using the above procedure, and the value provided by the formula hWW in its rounded-off version hWW hi, that is, in symbols, hWW hi ¼bhWW þ0:5c;ð17Þ where  bcis the floor function (recall that the floor function of xgives the greatest integer less than or equal to x). Note that, from an operational point of view, all estimating formulas (1) generate real numbers. However, for estimation purposes, these numbers should be rounded-off to the nearest integer, not only in order to produce numbers in the same range of values as the h-index but also to avoid ‘‘false precision’’. (Hicks et al. 2015). To give an example illustrating the calculation of this estimate, let us consider the case of the Journal of the American Statistical Association (ISSN 0162-1459, from the S&MM list). We have C¼5231;C1¼156;T¼663 and T1¼519. Hence a¼log T T1  ¼log 663ðÞlog 519ðÞ¼0:2449 ð18Þ 1214 Scientometrics (2017) 113:1209–1228 123 Table 1 Basic statistics for the S&MM list of journals and the approximation of the Hirsch h-index calculated by means of the hWW formula (rounded values). The value hWW is not uniquely defined (N/D) for the first journal on the list (because of a too small average number of citations per paper). (Data retrieved in December 2015) # ISSN code CC 1 TT 1 hh WW hi 1 1405-7425 42 6 152 24 3 N/D 2 1012-9367 276 14 360 111 6 8 3 0017-095X 158 13 166 71 5 6 4 0315-3681 557 44 427 177 9 10 5 1081-1826 201 12 140 77 6 6 6 0957-3720 323 15 228 122 7 7 7 0002-9890 589 87 351 171 9 9 8 0361-0926 2033 28 1555 754 11 12 9 0117-1968 163 20 120 61 5 6 10 1210-0552 405 31 205 119 9 9 11 1056-2176 290 22 222 101 7 8 12 0165-4896 583 16 320 198 10 9 13 0315-5986 166 24 83 48 6 6 14 0736-2994 577 19 283 176 9 9 15 0399-0559 153 32 86 47 5 6 16 1303-5010 658 56 334 154 11 12 17 0927-7099 463 16 296 162 8 8 18 1351-1610 313 23 150 92 8 8 19 1292-8100 191 22 78 52 7 7 20 0361-0918 1036 45 635 369 9 10 21 0269-9648 263 16 172 84 7 8 22 1532-6349 308 15 141 93 7 8 23 0217-5959 522 33 261 155 9 9 24 1018-5895 424 25 189 115 9 9 25 0266-4763 2164 323 901 518 13 14 26 1471-678X 336 23 138 92 8 8 27 0304-4068 737 25 433 265 9 9 28 0020-7276 480 13 265 158 8 9 29 0023-5954 813 36 337 208 11 11 30 1220-1766 526 31 193 137 10 9 31 1226-3192 457 20 271 137 10 9 32 1618-2510 305 31 172 90 8 8 33 1083-589X 739 20 353 209 10 11 34 1048-5252 643 17 283 189 10 10 35 1004-3756 443 27 140 96 9 10 36 1009-6124 979 56 466 240 12 13 37 1120-9763 434 18 492 165 8 9 38 1369-1473 282 24 140 76 8 8 39 1230-1612 346 32 128 84 8 9 40 0026-1335 544 24 283 171 10 9 41 0218-348X 476 30 167 129 9 9 Scientometrics (2017) 113:1209–1228 1215 123 Table 1 continued # ISSN code CC 1 TT 1 hh WW hi 42 0167-7152 3169 40 1546 945 16 14 43 0032-4663 154 13 103 58 6 6 44 0282-423X 405 20 196 116 9 9 45 1748-670X 1933 36 822 543 14 13 46 0094-9655 1649 55 695 425 14 14 47 0039-0402 365 34 129 86 9 9 48 0894-9840 615 29 331 184 9 10 49 0398-7620 679 66 303 170 10 11 50 0219-0257 336 31 159 102 7 8 51 0319-5724 511 36 206 129 10 10 52 0020-3157 772 60 285 189 11 11 53 0898-2112 597 26 228 149 11 10 54 1524-1904 669 42 301 155 12 12 55 0963-5483 719 24 272 179 11 11 56 1547-5816 770 37 290 201 11 11 57 0001-8678 821 37 269 201 11 11 58 0021-9002 1168 35 477 321 13 12 59 0257-0130 719 18 260 179 11 11 60 1026-0226 2306 34 1036 610 15 15 61 0378-3758 3899 71 1334 907 18 18 62 0377-7332 1353 38 597 348 15 13 63 1560-3547 735 25 249 182 11 11 64 0893-4983 793 36 297 200 12 11 65 1387-5841 645 26 305 178 10 10 66 0167-6377 1702 33 582 399 14 14 67 1747-7778 837 294 135 93 10 12 68 1054-3406 1098 40 429 277 13 12 69 1619-4500 493 38 125 89 12 11 70 0143-9782 761 31 258 179 12 11 71 1432-2994 512 29 207 146 9 9 72 0219-4937 304 21 178 102 7 7 73 0033-5177 1734 42 878 522 14 13 74 1748-006X 779 31 238 184 11 11 75 1381-298X 364 23 113 82 9 9 76 0277-6693 825 61 217 160 14 12 77 1435-246X 735 43 263 175 11 11 78 1572-5286 587 25 158 114 12 12 79 1134-5764 458 59 246 128 8 9 80 0932-5026 829 26 396 210 11 12 81 0926-2601 769 78 286 196 10 10 82 0890-8575 333 47 119 74 8 9 83 0219-5259 803 32 254 179 12 12 84 0515-0361 447 37 150 89 11 10 1216 Scientometrics (2017) 113:1209–1228 123 Table 1 continued # ISSN code CC 1 TT 1 hh WW hi 85 0095-4616 626 46 192 135 11 11 86 0233-1934 1191 24 490 304 13 13 87 0167-5923 663 38 216 152 12 11 88 1469-7688 2100 77 653 404 17 18 89 1083-6489 1321 32 488 330 13 13 90 1392-5113 747 52 202 138 13 13 91 1863-8171 404 34 118 77 10 10 92 1380-7870 379 39 170 103 9 8 93 1862-4472 1866 32 652 438 15 15 94 0219-8762 905 65 300 185 15 13 95 0218-1274 5537 136 1370 1013 26 22 96 0747-4938 649 54 149 113 12 12 97 0020-7985 1280 28 417 268 16 15 98 0047-259X 3329 89 915 650 21 19 99 0303-6898 868 31 256 188 12 12 100 1471-082X 405 35 134 88 9 10 101 0924-6703 413 38 117 79 9 10 102 0346-1238 337 28 128 79 9 9 103 0748-8017 2076 31 534 380 19 18 104 1389-4420 793 124 184 124 15 13 105 0146-6216 737 30 215 155 12 12 106 0160-5682 3870 90 853 663 21 20 107 0960-0779 2712 118 570 443 20 19 108 0246-0203 1019 33 266 206 14 13 109 0306-7734 563 101 147 83 12 12 110 1350-7265 1499 40 375 294 15 15 111 0021-9320 910 22 274 207 12 12 112 0218-4885 1036 81 297 202 13 13 113 1945-497X 885 57 162 130 15 14 114 1352-8505 564 64 192 130 10 10 115 0003-1305 670 43 241 133 13 12 116 1076-2787 900 49 224 163 14 13 117 1862-5347 524 63 125 79 11 12 118 0022-4715 5302 91 1246 966 24 21 119 1133-0686 617 54 246 127 12 12 120 1539-1604 1075 183 286 194 13 13 121 1434-6028 7722 72 1849 1420 27 23 122 0304-4149 2652 44 791 577 15 16 123 0143-2087 1089 152 228 155 15 15 124 0323-3847 1221 129 327 230 15 14 125 0266-4666 1295 33 303 208 17 17 126 0925-5001 3452 61 849 611 22 20 127 1085-7117 682 49 183 129 13 12 Scientometrics (2017) 113:1209–1228 1217 123 and Redner (2010), for formula hR]. To measure the magnitude of the observed accuracy, for each of the six estimation formulas respectively numbered as: (1) hWW , (2) ~ h1ðÞ W, (3) hSG 0:63ðÞ, (4) hSG 0:75ðÞ, (5) hSG 1ðÞ, (6) hR, (a) we calculated the absolute relative error (ARE) of the estimator ^ hjiðÞ  of the actual h-index, hj, for each journal j,j¼1;...;J, AREjiðÞ¼ ^ hjiðÞ  hj  hj ;ð23Þ where ^ hjiðÞ  ¼b^ hjiðÞþ0:5cis the rounded-off version of formula i,i¼1;2;...;6, then, (b) as a criterion with which to assess the overall quality of the formula, we computed the mean absolute relative error (MARE), 22,520,017,515,012,510,07,55,0 22,5 20,0 17,5 15,0 12,5 10,0 7,5 5,0 estimatedh-index h-index Fig. 2 Scatterplot of the empirical value of the h-index hversus its predicted value by hWW , for the EE&F list of journals. The dashed line is identity, so ideally all the points should overlie this line Table 3 Relative accuracy, computed in terms of MARE, of different estimators of the h-index; rrepresents the number of basic metrics on which the estimation formula is based for each dataset, the smallest error is indicated by a boldface number \hWW [~ h1 ðÞ W DE hSG 0:63ðÞ hi hSG 0:75ðÞ hi hSGð1Þ hi hR hi r442221 S&MM list (230 cases) 0.060 0.076 0.271 0.141 0.162 0.224 EE&F list (100 cases) 0.056 0.050 0.217 0.081 0.251 0.192 1224 Scientometrics (2017) 113:1209–1228 123 MARE ^ hiðÞ  ¼X J j¼1 AREjiðÞ=J:ð24Þ The results are summarized in Table 3. Conclusion This paper has addressed the need to gain better understanding of how simple citation metrics are related to the h-index, or rather, to a ‘‘good’’ proxy representation of the h index. This also responds to the more basic requirement of ‘‘building bridges’’ between different types of known and available measures of impact/impact indicators—under IIC. Differently from other studies (that consider the problem of defining a ‘‘model’’ of the h-index), our concern has not been to estimate the parameters (sometimes even considered at the unit level, i.e. single journal, or single scientist; see e.g. Petersen et al. 2011)ofa parametric model for the h-index under the assumption of knowing the entire citation pattern; rather, we addressed the quite different and more practical problem of finding a proxy representation of hthrough a universal formula that only depends on few summary statistics of the data. The formula hWW is ‘‘universal’’ in the sense that it gives a proxy representation of hthat holds for any given journal and any dataset. The issue of determining an indicator under IIC is closely related to the search for a solution of the problem of recovering and comparing impact indicators from different databases. As a simple but significant example of this issue, we may cite the specific problem of determining/estimating the IF for journals using the Google Scholar-based hindex as a predictor (Bertocchi et al. 2015). As confirmed in our case study analysis, the h-index can be viewed as an almost-exact function of C;C1;Tand T1, through hWW , i.e. that the basic statistics C;C1;Tand T1 provide salient information for the evaluation of the h-index with high precision. In practice, while computation of the h-index hrequires knowledge of the entire citation profile (or at least large part of it, e.g. the so-called h-core), formula hWW requires knowledge of only a few elementary summary statistics, but reproduces the actual value of hquite well. In truth, in our computations we found that the estimates yielded by hWW were slightly biased downwards for quite high values of the h-index but, as can be seen from Table 3, overall the formula hWW yields very accurate approximations to the empirical value of the h-index, with values of the MARE ranging around 5–6%, not too dissimilar from those obtained by formula ~ h1ðÞ W(Bertoli-Barsotti and Lando 2017). Both formulas ~ h1ðÞ W and hWW exhibit comparable levels of accuracy (the advantages of the formula ~ h1ðÞ W,as compared to formula hWW , may be that: (i) it yields an explicit expression of the basic indicators C;C1;Tand T1, while the latter not, and (ii) it is based on a simpler probabilistic model). Even though the Pearson correlation, q,isnot an adequate measure of the accuracy of the estimation and should not be used to compare the effectiveness of the different estimators considered (and this is the reason why this concept has been banished from this study), for the sake of completeness we point out that: (1) for the S&MM dataset (230 journals), we found qh;hWW  ¼0:99, qh;~ h1ðÞ W  ¼0:98, qh;hSG ðÞ¼0:98 and qh;hR ðÞ ¼0:96; (2) for the EE&F dataset we found qh;hWW  ¼0:97, qh;~ h1ðÞ W  ¼0:98, qh;hSG ðÞ ¼0:97 and qh;hR ðÞ ¼0:90. Ultimately, despite the differences between the Scientometrics (2017) 113:1209–1228 1225 123 datasets considered—in terms of scientific areas, time windows for publication and citation, types of ‘‘citable’’ documents considered, mean level of the basic indicators C;C1;T and T1(with values of respectively 2111, 95, 432 and 312 for the S&MM dataset and 741, 33, 199 and 159 for the EE&F dataset)—we may conclude that, on the whole, hWW provides fairly accurate approximations to the real value of the h-index, at least for not too large values of T(e.g. T\2000), m(e.g. m\20) and h(e.g. h\40), such as those considered in this study. Acknowledgements Funding was provided by Czech Science Foundation (Grant No. 17-23411Y). Open Access This article is distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license, and indicate if changes were made. Appendix Conditions for existence and uniqueness of a solution of Eq. (15) For every fixed a¼a[0, ga ;bðÞ!þ1as b!0 and ga ;bðÞ!1asb!þ1. Moreover, since o obga ;bðÞ¼ ga ;bðÞ b2log aw1þ1 b  ;ð25Þ where wis the digamma function, i.e. the function defined by wzðÞ¼d dzlog CzðÞ¼ C0zðÞ=CzðÞ(see Johnson et al. 2005, pp. 8–9), we find that the inequality o obga ;bðÞ\0ð26Þ holds if and only if it holds log aw1þ1 b  \0:ð27Þ Now, the function w1þ1 b  is (convex and) strictly decreasing from þ1 at 0 to lim b!1w1þ1 b  ¼w1ðÞ¼C01ðÞ¼c;ð28Þ where cis the Euler–Mascheroni constant (c¼C01ðÞffi0:5772), at þ1. Hence w1þ1 b  [c[log afor every b[0 if and only if 0\aexp cðÞffi0:561. Thus the following two cases are possible. 1226 Scientometrics (2017) 113:1209–1228 123 (a) If 0\aexp c ðÞ , the inequality (26) holds. In this case the function ga ;b ðÞ is strictly decreasing from þ1 at 0 to 1 at þ1, with a limit approached from above. We conclude that, in this case, Eq. (15) has a unique solution if and only if m[1; otherwise, if m1, Eq. (15) has no solution. (b) On the other hand, if a[exp cðÞ, the derivative function o obga ;bðÞchanges its sign from negative to positive at b¼b0, for some b0[0; hence ga ;bðÞis strictly decreasing for every 0\b\b0, and strictly increasing for every b[b0, and the point b0is a global minimum for ga ;b ðÞ . Moreover since, as seen before, lim b!1ga ;bðÞ¼1, then 0\ga ;b0 ðÞ\1, and the limit at infinity is approached from below. We conclude that, in this case too, Eq. (15) has a unique solution if and only if m[1; conversely, if m1 Eq. (15) may have two solutions, or no solution at all. In both cases (a) and (b), Eq. (15) has one and only one solution if and only if m[1. References Bertocchi, G., Gambardella, A., Jappelli, T., Nappi, C. A., & Peracchi, F. (2015). Bibliometric evaluation vs. informed peer review: Evidence from Italy. Research Policy, 44, 451–466. Bertoli-Barsotti, L., & Lando, T. (2017). A theoretical model of the relationship between the h-index and other simple citation indicators. Scientometrics, 111(3), 1415–1448. Burrell, Q. L. (2013). A stochastic approach to the relation between the impact factor and the uncitedness factor. Journal of Informetrics,7, 676–682. Corless, R. M., & Jeffrey, D. J. (2015). The Lambert W Function. In N. J. Higham, M. Dennis, P. Glendinning, P. Martin, F. Santosa, & J. Tanner (Eds.), The Princeton companion to applied mathematics (pp. 151–155). Princeton: Princeton University Press. Egghe, L. (2013). The functional relation between the impact factor and the uncitedness factor revisited. Journal of Informetrics,7, 183–189. Gla ¨nzel, W. (2006). On the h-index—A mathematical approach to a new measure of publication activity and citation impact. Scientometrics, 67, 315–321. Hicks, D., Wouters, P., Waltman, L., De Rijcke, S., & Rafols, I. (2015). The Leiden Manifesto for research metrics. Nature, 520(7548), 429. Hirsch, J. E. (2005). An index to quantify an individual’s scientific research output. Proceedings of the National Academy of Sciences,102, 16569–16572. Hsu, J. W., & Huang, D. W. (2012). A scaling between impact factor and uncitedness. Physica A,391, 2129–2134. Iglesias, J., & Pecharroman, C. (2007). Scaling the h-index for different scientific ISI fields. Scientometrics, 73, 303–320. Ionescu, G., & Chopard, B. (2013). An agent-based model for the bibliometric h-index. The European Physical Journal B, 86, 426. Johnson, N. L., Kemp, A. W., & Kotz, S. (2005). Univariate discrete distributions. New York: Wiley. Malesios, C. (2015). Some variations on the standard theoretical models for the h-index: A comparative analysis. Journal of the Association for Information Science and Technology, 66, 2384–2388. Panaretos, J., & Malesios, C. (2009). Assessing scientific research performance and impact with single indices. Scientometrics, 81, 635–670. Petersen, A. M., Stanley, H. E., & Succi, S. (2011). Statistical regularities in the rank-citation profile of scientists. Scientific Reports, 1, 181. Prathap, G. (2010a). Is there a place for a mock h-index? Scientometrics, 84, 153–165. Prathap, G. (2010b). The 100 most prolific economists using the p-index. Scientometrics, 84, 167–172. R Development Core Team. (2012). R: A language and environment for statistical computing. Vienna: R Foundation for Statistical Computing. http://www.R-project.org. Redner, S. (2010). On the meaning of the h-index. Journal of Statistical Mechanics: Theory and Experiment, 2010(03), L03005. Scientometrics (2017) 113:1209–1228 1227 123 Schreiber, M., Malesios, C. C., & Psarakis, S. (2012). Exploratory factor analysis for the Hirsch index, 17 h-type variants, and some traditional bibliometric indicators. Journal of Informetrics, 6, 347–358. Schubert, A., & Gla ¨nzel, W. (2007). A systematic analysis of hirsch-type indices for journals. Journal of Informetrics, 1, 179–184. Vinkler, P. (2009). The p-index: A new indicator for assessing scientific impact. Journal of Information Science, 35, 602–612. Vinkler, P. (2013). Quantity and impact through a single indicator. Journal of the American Society for Information Science and Technology, 64, 1084–1085. Wolfram R. (2014). Mathematica 10.0. Champaign, IL: Wolfram Research Inc. 1228 Scientometrics (2017) 113:1209–1228 123