scieee AI-readable full text Open interactive document viewer

Estimating population coefficient of variation using a single auxiliary variable in simple random sampling

Singh, Rajesh,Mishra, Madhulika

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Singh, Rajesh; Mishra, Madhulika Article Estimating population coefficient of variation using a single auxiliary variable in simple random sampling Statistics in Transition New Series Provided in Cooperation with: Polish Statistical Association Suggested Citation: Singh, Rajesh; Mishra, Madhulika (2019) : Estimating population coefficient of variation using a single auxiliary variable in simple random sampling, Statistics in Transition New Series, ISSN 2450-0291, Exeley, New York, NY, Vol. 20, Iss. 4, pp. 89-111, https://doi.org/10.21307/stattrans-2019-036 This Version is available at: https://hdl.handle.net/10419/210692 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/4.0/ STATISTICS IN TRANSITION new series, December 2019 89 STATISTICS IN TRANSITION new series, December 2019 Vol. 20, No. 4, pp. 89–111, DOI 10.21307/stattrans-2019-036 Submitted – 28.08.2018; Paper ready for publication – 01.10.2019 ESTIMATING POPULATION COEFFICIENT OF VARIATION USING A SINGLE AUXILIARY VARIABLE IN SIMPLE RANDOM SAMPLING Rajesh Singh 1 , Madhulika Mishra 2 ABSTRACT This paper proposes an improved estimation method for the population coefficient of variation, which uses information on a single auxiliary variable. The authors derived the expressions for the mean squared error of the proposed estimators up to the first order of approximation. It was demonstrated that the estimators proposed by the authors are more efficient than the existing ones. The results of the study were validated by both empirical and simulation studies. Key words: coefficient of variation, simple random sampling, auxiliary variable, mean square error. 1. Introduction It is a prominent fact in the theory of sample surveys that suitable use of auxiliary information increases the efficiency of the estimators used for estimating the unknown population parameters. Some important works illustrating use of auxiliary information at estimation stage are Singh et al. (2005), Singh et al. (2007), Khoshnevisan et al. (2007), Singh et al. (2009), Singh and Kumar (2011), Malik and Singh (2013) and Singh et al. (2018). Over a vast period of time a substantial amount of work has been done by several authors for the estimation of population mean, population variance but little attention has been given to the estimation of the population coefficient of variation. Das and Tripathi (1992–93) first proposed the estimator for the coefficient of variation when samples were selected using simple random sampling without replacement (SRSWOR) scheme. Other works include Patel and Shah (2009) and Ahmed, S.E. (2002). Breunig (2001) suggested an almost unbiased estimator of the coefficient of variation. Sisodia and Dwivedi (1981) suggested a modified ratio estimator using the coefficient of variation of auxiliary variable. 1 Department of Statistics, Banaras Hindu University, Varanasi-221005, India. ORCID ID: http://orcid.org/0000-0002-9274-8141. 2 Corresponding Author: Department of Statistics, Banaras Hindu University, Varanasi-221005, India. E-mail: [email protected]. ORCID ID: http://orcid.org/0000-0002-6408-0746. 90 R. Singh, M. Mishra: Estimating population coefficient… Rajyaguru and Gupta (2005) also worked on the problem of estimation of the coefficient of variation under simple random sampling and stratified random sampling. The coefficient of variation is extensively used in biology, agriculture and environmental sciences. A brief summary of the paper is as follows. Section 1 is introductory in nature, comprises the works that have been already done in the sampling literature. In Section 2 we considered five estimators for comparison purposes and their properties. In Section 3, we proposed two log type estimators for the coefficient of variation, one general type estimator and one wider type. In Section 4, an empirical study was carried out in support of our results. In Section 5, we carried out a simulation study to validate our theoretical results and have presented them with the help of bar graphs. In Section 6 we finally concluded our results. Let us consider a finite population P = (P1, P2……… PN) of size ‘N’ consisting of distinct and identifiable units. Let the study and auxiliary variables be denoted by Y and X, and let Yi and Xi be their values corresponding to ith unit in the population (i = 1, 2………. N). We define:   N ii Y N Y 1 1 as the population mean for the study variable   N ii X N X 1 1 as the population mean for the auxiliary variable       N iiy YY N S 1 2 2 1 1 as the population mean square for the study variable       N iix XX N S 1 2 2 1 1 as the population mean square for the auxiliary variable   XXYY N Si N iixy    1 1 1 as the population covariance between the study and auxiliary variable, X and Y. Let us suppose that a sample of size ‘n’ has been drawn from this population of size ‘N’ units using SRSWOR technique. For this sample let yi and xi denote values of the ith sample unit corresponding to study variable Y and auxiliary variable X respectively. For the sample observations, we define:   n ii y n y 1 1 as the sample mean for the study variable Y STATISTICS IN TRANSITION new series, December 2019 91   N ii x n x 1 1 are the sample mean for the auxiliary variable X       n iiy yy n s 1 2 2 1 1 as the sample mean square for the study variable        n i ix xx n s 1 2 21 1 as the sample mean square for the auxiliary variable   xxyy n si n iixy    1 1 1 as the sample covariance term. Now, let us define 1 0 Y y , 1 1 X x , 1 2 2 2 y y S s and 1 2 2 3 x x S s such that         0 3210    22 0n f-1 y C        ,   22 1n f-1 x C        ,    1 n f-1 40 2 2         ,     1 n f-1 04 2 3           xyCC         n f-1 10 ,   3020 n f-1  y C          1230 n f-1  y C        ,   2121 n f-1  x C        ,   0331 n f-1  x C        ,    1 n f-1 2232          Here, N n f : Sampling fraction, Y S y  y C and X S x  x C are the population coefficient of variation for the study variable Y and auxiliary variable X, respectively. Also xy  denotes the correlation coefficient between X and Y. 92 R. Singh, M. Mishra: Estimating population coefficient… In general,          N i sr irs XxYy i 1 1-N 1  and 2 s 02 2 r 20 rs     rs respectively. 2. Existing estimators  The usual unbiased estimator to estimate the population coefficient of variation using information on a single auxiliary variable is defined below:    0 2 1 2yy 0 1 1S y s ˆ    Y Ct y y C             822 1- 2 2 20 2 2 00 (2.1) Its mean squared error (MSE) is given by:                 30 40 22 y0 4 1 1 C   yy CC n f tMSE (2.2)  Solanki et al. (2015) introduced a difference type estimator for the population coefficient of variation y C as:   xxd CCC ˆ C ˆ 2y   (2.3) MSE of d C is given by:                                                03 04 2 22 21 12 2 30 40 22 y 4 1 4 1 22 4 1 1 C         xx x y xy yyyd CC C C CC CCC n f CMSE (2.4) STATISTICS IN TRANSITION new series, December 2019 93  Solanki et al. (2015) defined another class of estimator for the population coefficient of variation y C as:   xxyd CCCC  ˆˆ 21 *  (2.5) MSE of * d C is given by:   ECCDCCCCCBCACCMSE xyyxyyxyd 2 2 121 222 2 22 1 *222   (2.6) Here,   30 223 n f-1 1  yy CCA                       03 04 2 4 1 n f-1   xx CCB                4 1 2228 1 n f-1 C 22 21 12 0304 2      x y xy x x C C CC C C (2.7)                28 1 n f-1 1 30 40 2   y y C CD . 2 C 8 1 C n f-1 E 03x04 2 x                On differentiating equation (2.6) with respect to 1  and 2  , we obtain their optimum values as: 2 1 CAB CEBD opt     (2.8)         2 x y 2C C CAB CDAE opt  (2.9) On substituting these optimum values of 1  and , 2  in equation (2.6), we obtain the Minimum MSE for the estimator * d C as:   ECCDCCCCCBCACCMSE xyoptyoptxyoptoptyxoptyoptd 2 2 121 222 2 22 1 *222 min   (2.10) 94 R. Singh, M. Mishra: Estimating population coefficient…  Adichwal et al. (2016) proposed a two-parameter ratio-product-ratio estimator for the population coefficient of variation as:           yyr C Xx Xx C Xx Xx tˆ 1 1 1 ˆ 1 -1 1                        (2.11)          y xx xx y xx xx rC Ss Ss C Ss Ss tˆ 1 1 1 ˆ 1 -1 22 22 22 22 2                        (2.12) MSE of the estimators 1r t and 2r t are respectively given by:       2 2 21y1 2 1 4 1 CMSE yyxyrCC n f tMSE          (2.13)           2 04 2 1222 y2 1 21 1 4 1 CMSE y y rC C n f tMSE            (2.14) 3. Proposed estimators We have proposed some estimators for the coefficient of variation based on information on a single auxiliary variable. Motivated by Mishra and Singh (2017), we propose improved log type estimators for estimating the population coefficient of variation given by: estimators t1 and t2 as: a)          x x C C tˆ logC ˆ y1  (3.1) b)            x x C C wwt ˆ log1C ˆ 21y2 (3.2) Expressing the estimator 1 t and in terms of s' and then taking expectations up to the first order of approximation, we get MSE of the estimator as:                                     03 04 2 x 2 30 40 2 y 2 y1 4 1 C 1 4 1 C 1 C     xy C n f C n f tMSE                     224 1 1 221 12 22     x y xyy C C CC n f C (3.3) 1 t STATISTICS IN TRANSITION new series, December 2019 95   32 2 1 2 y1 2C ACAAtMSE y   (3.4) Here,   ,C 4 1 C n f-1 A 30y 40 2 y1                  ,C 4 1 C n f-1 A 03x 04 2 x2                (3.5)   . 2 C 2 C 4 1 CC n f-1 A 12y 21x22 xy3                   To obtain the optimum value of  , we partially differentiate the expression (3.4) with respect to  and we obtain the optimum value as: 2 3 y C- A A opt   (3.6) Putting this optimum value of  in equation (3.4), we get the minimum value for   1 tMSE as:           2 2 3 1 2 y1 C min A A AtMSE (3.7) Expressing the estimators 2 t in terms of s' and then taking expectations up to the first order of approximation we get MSE of the estimator 2 t as:                                   30 22 1 2 30 40 22 y2 1 2 1 31 4 1 1 C )(   yyyyy C n f C n f wCCC n f tMSE                               30 40 2 1 2 03 04 22 22 3 8 1 2 1 2 4 1 1     yyyxx CC n f wCCC n f w                  224 1 4 1 2 1 221 12 04 22 2 21      x y xy x y C C CC C n f wwC 96 R. Singh, M. Mishra: Estimating population coefficient…               224 1 1 221 12 22 2     x y xyy C C CC n f wC (3.8)   6252141 2 3 2 22 2 1 2 1 2 y2 222C BwCBwwCBwCBwBwCBtMSE yyyy  (3.9) Here,               30 40 2 14 1 1   yy CC n f B 30 2 2 1 2 1 31  yy C n f C n f B                            03 04 2 34 1 1   xx CC n f B               30 40 2 42 3 8 1 2 1   yy CC n f B (3.10)                 224 1 4 1 2 1 21 12 04 22 2 5      x y xy xC C CC C n f B               224 1 1 21 12 22 6     x y xy C C CC n f B To obtain the optimum value of 1 w and ,w2 we differentiate the expression (2.21) with respect to 1 w and 2 w and obtain the optimum values as:           2 532 4365 1 BBB BBBB wopt (3.11)            32 2 5 5426 2 BBB BBBB Cw yopt (3.12) STATISTICS IN TRANSITION new series, December 2019 103 Table 1. MSE and PRE of the estimators (cont.( ESTIMATOR POPULATION-1 POPULATION-2 MSE PRE MSE PRE 1 t 0.00123 651.7356 0.0299 127.4598 2 t 0.001038 771.9898 0.0283 134.6127 3 t 0.001203 666.4304 0.0297 128.345 4 t 0.001203 666.4304 0.0297 128.345 We can summarize the results from Table 1 as: All the proposed estimators 1 t , 2 t , 3 t and 4 t are more efficient than the usual unbiased estimator 0 t .The estimator 1 t turns out to be nearly as efficient as the difference type estimator d C while all the remaining estimators, 2 t , 3 t and 4 t are more efficient than the estimators d C , * d C , 1r t and 2r t . Among all the estimators, 2 t is the most efficient because of the smallest value of MSE and highest value of PRE. 5. Simulation studies This section describes the procedure that we adopted for the simulation study. We have used R programming for calculating MSE of the existing and proposed estimators. We followed the procedure adopted by Reddy et al. (2010) and have generated bivariate population with a specified correlation coefficient between the study and auxiliary variable. The algorithm is as follows: 1. Generate two independent random variables X from ),( 2  N and Z from ),( 2 11  N using Box-Muller method (Jhonson, 1987). 2. Set ZXY 2 1   where 195.0,85.0,75.00   . 3. Consider the population with the parameters 5.2  , 2 2  , 5 1  3 2 1  and repeat the steps 1-2 2000 times. 4. From the population of size N=2000, draw 1500 simple random samples ),.....,2,1(),( nixy ii  without replacement of size 70,50,30n . 104 R. Singh, M. Mishra: Estimating population coefficient… 5. For each of the sample, compute MSE of the estimators o t , Cd , * Cd , 1 t , 2 t , 1r t and 2r t . 6. Compute the average MSE of the estimator by the following formula:        1500 1 1500 1 j jimseiMSE where 2121 *,,,,, rro tandtttCdCdti . Table 2. Table showing MSE and PRE of the existing and proposed estimators for different values of  and n  n Estimator MSE PRE o t 0.006053626 100.0000 Cd 0.004924006 122.9410 * Cd 0.004617663 131.0970 1r t 0.005080027 119.1651 2r t 0.005532107 109.4270 1 t 0.004668748 129.6626 2 t 0.004552470 132.9744 50 o t 0.003450581 100.0000 Cd 0.002835622 121.6869 * Cd 0.002671694 129.1533 1r t 0.002688403 117.5860 2r t 0.002650186 109.5933 1 t 0.002934516 128.3506 2 t 0.003148534 130.2015 70 o t 0.002412659 100.0000 Cd 0.001990824 121.1889 * Cd 0.001879170 128.3896 1r t 0.002062564 116.9737 2r t 0.002200267 109.6530 1 t 0.001887289 127.8373 2 t 0.001868644 129.1128 STATISTICS IN TRANSITION new series, December 2019 105 0.85 30 o t 0.006358341 100.0000 Cd 0.004912595 129.4294 * Cd 0.003809327 166.9151 1r t 0.004876890 130.3769 2r t 0.005133219 123.8665 1 t 0.003831045 165.9688 2 t 0.003739557 170.0293 50 o t 0.003621428 100.0000 Cd 0.002828737 128.0228 * Cd 0.002203557 164.3447 1r t 0.002825058 128.1895 2r t 0.002910006 124.4474 1 t 0.002210627 163.8190 2 t 0.002180249 166.1016 70 o t 0.002527309 100.0000 Cd 0.001982556 127.4561 * Cd 0.001547597 163.3054 1r t 0.001984483 127.3535 2r t 0.002027634 124.6433 1 t 0.001551035 162.9434 2 t 0.001536162 164.5210 106 R. Singh, M. Mishra: Estimating population coefficient… 95 30 o t 0.008647395 100.0000 Cd 0.005851489 147.7811 * Cd 0.002426053 356.4389 1r t 0.005461214 158.3420 2r t 0.005113439 169.1094 1 t 0.002430095 355.8459 2 t 0.002364604 365.7015 50 o t 0.004896276 100.0000 Cd 0.003355658 145.9111 * Cd 0.001397620 350.3295 1r t 0.003172583 154.3309 2r t 0.002841595 172.3073 1 t 0.001398209 350.1820 2 t 0.001376286 355.7601 70 o t 0.0034018746 100.0000 Cd 0.0023435450 145.1593 * Cd 0.009782472 347.7520 1r t 0.0022234723 152.9983 2r t 0.0019631488 173.2866 1 t 0.0009784789 347.6697 2 t 0.0009677328 351.5304 From the table, we can observe that for a particular value of  the value of MSE of the estimators decreases as the sample size increases. Also, we can see that in each of the cases among the proposed estimators 1 t and 2 t , 2 t is more efficient amongst all the existing estimators o t , Cd , * Cd , 1r t , 2r t and the proposed estimator 1 t while the estimator 1 t turns out to be more efficient than the existing estimators o t , Cd , 1r t , 2r t and nearly as efficient as the estimator * Cd . Hence, it turns out that the proposed estimator performs better than the existing estimators, therefore it is desirable to use the estimator in practice. STATISTICS IN TRANSITION new series, December 2019 107 We have also shown the results through a bar diagram as below: Bar graph showing MSEs of the existing and proposed estimators for 75.0  and (n1, n2, n3)= (30, 50, 70) Explanation: It can be seen from the bar graph that for 75.0  , MSE of all the estimators decreases as the value of the sample size (n) increases. And for a particular value of n, estimator 2 t has the least MSE among all the other estimators. 108 R. Singh, M. Mishra: Estimating population coefficient… Bar graph showing MSEs of the existing and proposed estimators 85.0  and (n1, n2, n3) = (30, 50, 70) Explanation: It can be seen from the bar graph that for 85.0  , MSE of all the estimators decreases as the value of the sample size (n) increases. And for a particular value of n, estimator 2 t has the least MSE among all the other estimators. STATISTICS IN TRANSITION new series, December 2019 109 Bar graph showing MSE of the existing and proposed estimators 85.0  and (n1, n2, n3)= (30, 50, 70) Explanation: It can be seen from the bar graph that for 95.0  , MSE of all the estimators decreases as the value of the sample size (n) increases. And for a particular value of n, estimator 2 t has the least MSE among all the other estimators. Combined Explanation: From the above three bar graphs it can be summarized that for every value of   95.0,85.0,75.0  , the increase in the sample size causes a decrease in the mean square error of all the estimators. It is also evident that for a particular value of n, 2 t has the minimum MSE as compared to the other estimators. 6. Conclusion In this paper we have proposed estimators for the population coefficient of variation and compared them with some existing estimators and saw from the empirical and simulation studies that the proposed estimator 2 t performs better 110 R. Singh, M. Mishra: Estimating population coefficient… than all the existing estimators o t , Cd , * d C , 1r t , 2r t and the proposed estimator 1 t . As regards 1 t , it performs better than the estimators o t , Cd , 1r t , 2r t but is no more better than the estimator * d C . For a better understanding of our results we have also considered a graphical approach and considered bar graphs to depict our results. Acknowledgement The authors are grateful and obliged to the Editor-in-Chief Prof. Włodzimierz Okrasa and the anonymous referees who devoted a part of their valuable time to provide us with their fruitful recommendations, which in turn helped us to improve the manuscript. REFERENCES AHMED, S. E., (2002). Simultaneous estimation of Co-efficient of Variation. Journal of Statistical Planning and Inference, 104, pp. 31–51. ARCHANA, V., RAO, A., (2014). Some Improved Estimators of Co-efficient of Variation from Bi-variate normal distribution: A Monte Carlo Comparison. Pakistan Journal of Statistics and Operation Research, 10(1). BREUNIG, R., (2001). An almost unbiased estimator of the co-efficient of variation. Economics Letters, 70(1), pp. 15–19. DAS, A. K., TRIPATHI, T. P., (1981). A class of estimators for co-efficient of variation using knowledge on coefficient of variation of an auxiliary character, In annual conference of Ind. Soc. Agricultural Statistics, Held at New Delhi, India. DAS, A. K., TRIPATHI, T. P., (1992). Use of auxiliary information in estimating the coefficient of variation, Alig. J. of. Statist, 12, pp. 51–58. FINITE POPULATION-II, Model Assisted Statistics and application, 1(1), pp. 57– 66. KHOSHNEVISAN, M., SINGH, R., CHAUHAN, P., SAWAN, N., SMARANDACHE, F. (2007). A general family of estimators for estimating population means using known value of some population parameter(s), Far East Journal of Statistics, 22(2), pp. 181–191. MALIK, S., SINGH, R., (2013). An improved estimator using two auxiliary attributes, Appli. Math. Compt., 219, pp. 10983–10986. MISHRA, P., SINGH, R., (2017). A new log-product type estimator using auxiliary information, Jour. Sci. Res., 61(1&2), pp. 179–183. STATISTICS IN TRANSITION new series, December 2019 111 MURTHY, M. N., (1967). Sampling theory and methods, Sampling theory and methods. PATEL, P. A., RINA, S., (2009). A Monte Carlo comparison of some suggested estimators of co-efficient of variation in finite population, Journal of Statistics sciences, 1(2), pp. 137–147. RAJYAGURU, A., GUPTA, P., (2005). On the estimation of the co-efficient of variation from SINGH, H. P., TAILOR, R., (2005). Estimation of finite population mean with known coefficient of variation of an auxiliary character, Statistica, 65(3), pp. 301–313. SINGH, H. P., SINGH, R., (2002). A class of chain ratiotype estimators for the coefficient of variation of finite population in two phase sampling, Aligarh Journal of Statistics, Vol. 22, pp. 1–9. SINGH, S., (2003). Advanced Sampling Theory With Applications: How Michael Selected Amy (Vol. 2), Springer Science and Business Media. SINGH, H. P., SINGH, R., ESPEJO, M. R., PINEDA, M. D., NADRAJAH, S., (2005). On the efficiency of a dual to ratio-cum-product estimator in sample surveys. Mathematical proceedings of the Royal Irish Academy, 105A (2), pp. 51–56. SINGH, R., CHAUHAN, P., SAWAN, N., (2007). A family of estimators for estimating population means using known correlation coefficient in two-phase sampling. Statistics in Transition, 8(1), pp. 89–96. SINGH, R., KUMAR, M., CHAUDHARY, M. K., KADILAR, C., (2009). Improved Exponential Estimator in Stratified Random Sampling, Pak. J. Stat. Oper. Res., 5(2), pp 67–82. SINGH, R., KUMAR, M., (2011). A note on transformations on auxiliary variable in survey sampling. Mod. Assis. Stat. Appl., 6:1, pp. 17–19. SINGH, R., MISHRA, P., BOUZA, C. N., (2018). Estimation of population mean using information on auxiliary attribute: A review, RG, DOI: 10.13140/RG.2.2.20477.87524. SISODIA, B. V. S., DWIVEDI, V. K., (1981). Modified ratio estimator using coefficient of variation of auxiliary variable, Journal-Indian Society of Agricultural Statistics.