Full text
RESEARCH ARTICLE Copyright Bereznyatskiy AN, Badina SV, Babkin RA. This isan open access article distributed under the terms ofthe Creative Commons Attribution License (CC-BY 4.0), which permits unrestricted use, distribution, and reproduction inany medium, provided the original author and source are credited The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data Alexander N. Bereznyatskiy 1, Svetlana V. Badina 2, Roman A. Babkin 3 1 Central Economics and Mathematics Institute of the Russian Academy of Sciences, Moscow, 117418, Russia 2 Lomonosov Moscow State University, Moscow, 119991, Russia; Bernardo O’Higgins University, Santiago, 8320000, Chile 3 Plekhanov Russian University of Economics, Moscow, 117997, Russia Received 23 March 2024 ♦ Accepted 24 February 2025 ♦ Published 3October 2025 Citation: Bereznyatskiy AN, Badina SV, Babkin RA (2025) The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data. Population and Economics 9(3):80-93. https://doi.org/10.3897/ popecon.9.e123730 Abstract This study investigates the spatial and temporal characteristics ofpopulation distribution inthe city ofMoscow, with afocus onthe constancy ofthese dynamics over time, using high-frequency data from mobile network operators. The primary unit ofanalysis isa 500× 500meter grid cell covering the city’s territory. The analysis spans the period from 2018to2020. The stability ofpopulation distribution isassessed along two dimensions: the spatial distribution ofthe permanent population and the interannual variation ofdaily population gradients across grid cells. The hypothesis ofdistributional constancy istested through the construction ofregression models, followed bya residual analysis using Geographic Information Systems (GIS). Spatial heterogeneities identified inthe residuals are further examined. The findings indicate that, for anumber ofMoscow districts, the spatial population distribution closely follows apattern consistent with the Zipf– Mandelbrot law, while other districts deviate significantly from this model. This pattern was further validated using Rosstat data for Moscow districts covering the years 2013–2022, confirming both the functional form and the classification ofdistricts into two distinct groups. The first group primarily includes areas within Moscow’s 2012administrative boundaries and high-density zones inNew Moscow, whereas the second group comprises the remaining, less densely populated areas. The proposed methodology demonstrates that the distribution ofthe permanent population across Moscow remained generally stable inboth spatial and temporal terms over the study period. Conversely, the hypothesis oftemporal constancy indaily population gradients isrejected, likely due tothe influence ofcomplex factors such asthe evolving urban transportation infrastructure and cumulative measurement errors. Keywords big data, mobile operator data, urban population density, regression analysis, Zipf– Mandelbrot law JEL codes: C01, C21, R14, R23 Population and Economics 9(3): 80–93 DOI 10.3897/popecon.9.e123730
Population and Economics 9(3): 80–93 81 Introduction Understanding the dynamics ofpopulation distribution within urban areas isa critical component ofmodern city modeling. Traditionally, population studies have focused ontotal population changes ordemographic structures, often overlooking the spatial and structural dimensions ofurban settlement patterns (Argunov etal. 2017). However, for avariety ofapplications– including population risk assessment (Tenerelli etal. 2015), the evaluation ofurban policy impacts, and the regulation oftraffic flows (Saghapour etal. 2016)– it isessential toconduct adetailed statistical analysis ofsettlement structures and their temporal dynamics. Such analyses enable the identification ofstable territorial formations and anomalous spatial behaviors, facilitating further modeling ofinfluencing factors (e.g., urban development projects, housing renovation programs; see (Badina etal. 2023)) and forecasting changes insettlement patterns over time. Until recently, the primary limitation inconstructing such detailed models has been the lack ofappropriate data: • with high spatial granularity; • updated athigh temporal frequencies (e.g., minuteorhour-level observations). Recent advances in data availability from cellular network operators and providers ofpublic Wi-Fi infrastructure offer unprecedented opportunities. These data sources enable the tracking ofsubscriber locations atfine spatial and temporal resolutions, effectively merging the spatial precision ofsatellite imagery with the quantitative depth oftraditional demographic statistics. Several studies have already taken advantage ofsuch data inanalyzing settlement dynamics, including notable work focused onthe Moscow region (Makhrova etal. 2016; Makhrova and Babkin 2018; Mayakov 2022). Internationally, comparative analyses ofvarious data sources (e.g. (Lenormand etal. 2014)) have confirmed the high representativeness ofmobile phone data for capturing urban dynamics. Numerous studies have also addressed the broader topic ofurban structure (Babkin etal. 2022b). One ofthe earliest and most prominent initiatives using mobile data was the Austro-American Real-Time Graz project, launched in2005(Ratti 2005). Leveraging data from A1, Austria’s largest mobile network operator, the project successfully mapped labor migration patterns and shifts inpopulation distribution inthe city ofGraz. Similarly, Czech researchers used correlations between human mobility patterns and the spatial distribution ofresidential and commercial real estate toclassify daily rhythm types across Prague (Nemeškal etal. 2020). They demonstrated that mobile phone data– particularly indicators such asduration and timing ofuser presence– can reliably differentiate between residential, work, transportation, and service areas within the urban landscape. Comparable studies have since been conducted ina wide range ofglobal cities, including New York and Los Angeles (Isaacman etal. 2011), Harbin (Yuan etal. 2012), Rome (Calabrese etal. 2013), Singapore (Zhong etal. 2016; Jiang etal. 2017), London (Zhong etal. 2016), Beijing (Zhong etal. 2016), and Jakarta (Ruslani etal. 2019). There are numerous examples ofproject implementations where mobile operator data serve not merely asa supplementary source, but asthe primary and often the only reliable source ofrelevant information. This transition allows researchers and planners tobypass traditional data collection methods and instead leverage real-time information derived from mobile devices and satellite imagery (Giugale 2012). Mobile phone data has proven especially valuable incontexts where official population and mobility statistics are absent, outdated, orunreliable (Lu etal. 2012; Ruslani etal. 2019; Deville etal. 2014).
A.N. Bereznyatskiy et al.: The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data 82 Atthe same time, the widespread adoption ofGeographic Information Systems (GIS) inurban studies has significantly enhanced the capacity for spatial analysis ofpopulation distribution (Dolgacheva etal. 2022). When combined with granular big data, GIS technologies offer powerful tools for exploring complex patterns ofurban settlement. Equally important isthe rapid development ofopen-source statistical modeling tools that integrate GIS capabilities. Software environments such asR (notably through packages described in(Bivand etal. 2013)), released under the GNU license, now offer robust functionalities that greatly simplify spatial statistical analyses for researchers. This study presents one such approach for analyzing the “sustainability”– used interchangeably here with the term “constancy”– of the territorial distribution ofthe population inthe city ofMoscow. The stability ofthe underlying processes isinvestigated from two complementary perspectives: • the stability ofthe spatial distribution ofthe permanent population; • the stability ofthe spatial distribution ofpopulation gradients (or pulsations) across the city. Toaddress these objectives, the study employs methods ofapplied statistical analysis and cartographic visualization. Methodology This study isbased ona statistical dataset derived from primary information provided bycellular operators, detailing the positions ofnetwork subscribers within 500× 500meter grid cells across the city ofMoscow, with atemporal resolution of30minutes. The anonymized data were supplied bythe Department ofInformation Technology ofthe City ofMoscow. The full dataset includes approximately 10.000spatial cells, each tagged with unique identifiers that enable geographic referencing tospecific city areas. The data span the period from 2018to2020(for more details, see (Badina etal. 2022)). Prior toanalysis, the dataset underwent preprocessing toeliminate potential double counting– such asinstances where asingle user possesses multiple SIM-cards. Given the focus onterritorial distribution, the accuracy ofsubscriber positioning isa relevant concern. Areview ofthe literature oncellular-based positioning systems (Resch and Romirer-Maierhofer 2005; Babkin etal. 2022a) indicates that the achievable accuracy ranges from afew meters toseveral hundred meters, depending largely onthe density ofbase stations. Maximum precision istypically observed inlarge urban areas. The choice ofa 500× 500meter grid was therefore made toalign with the optimal spatial resolution achievable using the available data. To assess the stability ofthe permanent population distribution, asubsample ofthe dataset was created byextracting records corresponding to2:00a.m.– a time assumed torepresent the typical residential location ofsubscribers. For each month, the median population values were computed for each cell (a similar approach was applied inanalyzing daily population gradients). This assumption rests onthe premise that most individuals are athome during this hour, thus minimizing distortions from intra-urban movement– a finding supported byprevious research (Badina and Babkin 2021; Babkin etal. 2022b). The population distribution and its temporal consistency are illustrated inFig. 1. The analysis ofdaily population fluctuations across the city was based ona derived metric: the daily population gradient, defined asthe ratio between the maximum and minimum population observed ina given cell over a24-hour period. These pulsations are also visualized inFig. 1.
Population and Economics 9(3): 80–93 83 Figure 1. The change inthe population ofindividual Moscow cells during the day, 2019, the number at0-00hours istaken as100%. Source: compiled bythe authors based ondata from mobile operators. № 7113 (Akademichesky District) 212 196 180 164 148 132 116 100 84 % 000 100 200 300 400 500 600 700 800 900 1000 1100 1200 1300 1400 1500 1600 1700 1800 1900 2000 2100 2200 2300 Hour № 83998 (Donskoy District) № 23526 (Alexeyevsky District) Moscow, aggregate № 86675 (Akademichesky District) № 34546 (Alexeyevsky District) № 59734 (Bibirevo District) The research methodology follows the sequence outlined below. At each time point t (where tcan represent hours, days, months, years), the population within each spatial cell isrecorded asxi, where i= 1…N, N isthe total number ofcells covering the city ofMoscow. Alternatively, the analysis may beconducted ata more aggregated level, such asadministrative districts. Thus, weobtain apopulation distribution vector (data set): X x x x t t t t N = 1 2 . . . , which, for each period t, shows the distribution of the population across the territory ofMoscow. Since all cells have astandard size of500 by500meters, the concepts ofpopulation density structure and population distribution across cells can beconsidered mathematically equivalent. Further, ifthe population density structure attimes t1and t2does not undergo major changes during the transition from one period toanother, this means that vector X t 1 and vector X t 2 are consistent with each other, and the stability ofthe settlement structure over
A.N. Bereznyatskiy et al.: The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data 84 time isobserved. Otherwise, weare dealing with variability inthe density structure over the period t1, t2. The question ishow toevaluate the consistency ofthe two data sets. In this paper, itis proposed toassess data consistency using aregression model inwhich one data set ismapped onto the other. Subsequently, the hypothesis ofthe invariance ofthe population distribution across the territory ofMoscow istested using this approach, both inthe case ofthe permanent population and inthe analysis ofthe stability ofpulsations over time. Considering both the random component ofthe data (measurement errors, etc.) and the randomness ofpossible observed changes, the hypothesis istested using aregression test model ofthe following type: XX t m t n =+ + 01 αα ε, (1) where α0, α1 – unknown regression parameters (estimated for two test vectors using the least squares method); X tm – the vector ofpopulation distribution across the city ata given time tm; X tn – the vector ofpopulation distribution across the city ata given time tn; ε– arandom residual reflecting the discrepancy between the two data sets. The less consistent the vectors are, the greater the residual component ofthe regression will be when estimating the parameters. For example, inthe case ofself-regression (i.e. tm=tn), the residual component will bezero, which isexpected– asthere are nochanges inthe structure. The hypothesis about the constancy ofthe population distribution over time (for the selected interval) will beconfirmed ifa regression model that isstatistically adequate isobtained. Weare primarily interested inthe coefficient ofdetermination R2, which ranges from 0to1. Avalue of0indicates complete inconsistency between the two data sets, while avalue of1indicates full consistency. Special attention ispaid tothe residual component ε. The values ofthe residuals are recorded inGIS and displayed onmaps for visual analysis, providing more complete information about the phenomenon inquestion. Results and their discussion The methodology described inthe previous section was applied toa sample ofMoscow data for the period 2018–2020. February 2018was taken ast1, and January 2020ast2. The overall appearance ofthe data cloud and the regression line isshown inFig. 2. Figure 2displays apronounced elliptical shape with asmall number ofoutliers. Itis assumed that the constructed regression model for the two vectors will have good statistical properties and, overall, will not reject the hypothesis ofthe constancy ofthe population structure inMoscow for the period from February 2018toJanuary 2020. Table 1presents the results ofthe statistical evaluation ofthe regression model parameters. The test results show that, ingeneral, the hypothesis ofthe stability ofthe distribution ofthe permanent population across Moscow’s cells isnot rejected over atwo-year interval (the coefficient ofdetermination R2) for the regression model was 0.92, indicating good consistency between the two data sets). However, analysis ofthe model residuals and their visualization (Fig. 5) reveals the presence ofoutliers. Although itwould bepossible todisre-
Population and Economics 9(3): 80–93 85 gard individual deviations against the background ofthe overall residual pattern, the clear spatial clustering ofthese outliers onthe map suggests that they are not random. Toexplore this result ingreater detail, anadditional analysis was conducted onthe rank patterns ofcells (ranking bythe number ofpermanent residents per cell) and the corresponding population values. This approach– Zipf frequency analysis, modified byMandelbrot– is well known inlinguistics and has been adapted for use inurban demography (PavFigure 2. Visualization ofcomponents and regression lines for vectors ofdistribution ofpermanent population inMoscow cells, people (on the horizontal axis– the number ofcells for February 2018, onthe vertical axis– for January 2020). Source: calculated bythe authors based ondata from mobile operators. 12000 10000 8000 6000 4000 2000 January 2020 0 0 1000 2000 3000 4000 5000 6000 7000 8000 9000 10000 February 2018 Table 1. Regression model testing the stability ofthe permanent population distribution across Moscow’s cells for the period February 2018– January 2020 Estimation ofthecoefficients ofthe model Standard error t-statistics Significance level α094.3 5.3 17.8 *** α11.04 0.003 339 *** R2= 0.92 F (1.10087) = 114679** Number ofobservations = 10089 Source: authors’ calculations based ondata from mobile operators. Note: *10%, **5%, ***1% level ofsignificance
A.N. Bereznyatskiy et al.: The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data 86 lov 2020; Jingang etal. 2019; Urzúa 2000). Inthis study, the technique isapplied tospecific territories within the city (Li etal. 2009). The frequency diagram ofthe rank-size distribution ofcells bypopulation ispresented inFig. 3and clearly demonstrates two distinct modes ofpopulation distribution across Moscow’s cells. Experiments with model fitting toapproximate the distribution confirm the results ofthe visual analysis. Toanalyze the result over alonger time horizon, the data were aggregated from individual cells tothe level ofMoscow districts. Population distribution models were constructed for the period from 2013to2022(Fig. 4). Figure 3. The rank-frequency chart for Moscow cells, January 2020. Source: calculated bythe authors based ondata from mobile operators. 13 12.5 12 11 10 9 logarithm of the population of cells 8 01234 56 logarithm of cell rank 11.5 10.5 9.5 8.5 Figure 4. Chart rank-frequency for Moscow districts and model values, left graph– 2022, right graph– models for 2013, 2022incomparison. Source: calculated bythe authors according toRosstat data. 200000 100000 0 observed data 250000 150000 50000 1 12 23 34 45 56 67 78 89 100 111 122 133 144 modelmodel 2013 model 2022 200000 100000 0 250000 150000 50000 1 8 15 22 29 36 43 50 57 64 71 85 92 99 78 106 113 120 127 134 141
Population and Economics 9(3): 80–93 87 In terms ofthe population distribution model, all Moscow districts can begrouped into two classes. InFig. 4, this isshown asthe merging ofa logarithmic and alinear model atan inflection point. Figure 6presents the two identified clusters and the number ofcells onthe map. Ascan beseen, the distribution ofpopulation densities and the two clusters are closely interrelated. Itwas found that asubset ofareas with high population density iswell described bya logarithmic population distribution model, while the remaining areas correspond toa linear model. Importantly, the classification bytype ofpopulation distribution model remains stable across all periods from 2013to2022(calculations based onRosstat data). The parameters ofthe models, calibrated for each period separately, are also stable, asclearly shown bythe right-hand graph inFig. 4. It isworth referring once again tothe results oftesting for the constancy ofthe population distribution across Moscow’s cells and comparing Fig. 5with Fig. 6. Inthis context, the origin ofthe outliers inthe regression model becomes evident. Two different mechanisms within the settlement system give rise topatterns inthe residuals. This mechanism remains stable, atleast throughout the period from 2013to2022. 5.0 Administrative Okrugs: 0.5 0.25 0 -0.25 -0.5 -2.0 Regression model residuals CAO – Central NWAO – North-Western NAO – Northern NEAO – North-Eastern EAO – Eastern SEAO – South-Eastern SAO – Southern SWAO – South-Western WAO – Western NMAO – Novomoskovsky TAO – Troitsky ZelAO – Zelenogradsky ZelAO NEAO NAO NWAO CAO EAO WAO SWAO SEAO SAO NMAO TAO Figure 5. Visualization ofthe residual component ofthe regression model– the stability test ofthe distribution ofthe permanent population ofMoscow for the period February 2018– January 2020. Source: calculated bythe authors based ondata from mobile operators.
A.N. Bereznyatskiy et al.: The Stability of Population Distribution in Moscow: Analysis Based on Mobile Operators’ Data 88 During the 2012 administrative reform, various suburban territories were annexed toMoscow, forming New Moscow. The modeling conducted atthe level ofsmall territorial cells shows that these still constitute two distinct areas, despite their formal unification. This isunsurprising. The territories inquestion have undergone long-term development (in terms ofdemographics, economics, etc.), and itwould benaïve toexpect that their formal merger would immediately alter the demographic and economic systems formed over centuries. This type ofphenomenon isknown asa temporary decline inurban population density over time (Gang etal. 2019), when acity’s territory expands ata pace far exceeding the population growth ofthe newly incorporated areas. Let usnow consider the second part ofthe task– testing the stability ofthe territorial distribution ofpopulation gradients inMoscow. The testing mechanism issimilar tothe one previously discussed, with the only difference being the composition ofthe vector Xt. Inthis case, instead ofthe permanent population, the values ofthe daily gradients calculated for the periods ofFebruary 2018and January 2020are used for each cell. Figure 7shows the data cloud and the regression line for the corresponding values inFebruary 2018and January 2020. In contrast toFig. 2, there isno observable alignment ofpoints along astraight line; asexpected, the hypothesis ofgradient constancy islikely tobe rejected. The results ofthe regression parameter estimation are presented inTable 2. According tothe results ofthe regression assessment, the hypothesis regarding the constancy of the distribution ofdaily pulsations in Moscow’s cells for the period February 2018– January 2020isindeed rejected. The regression model has alow coefficient ofdetermination, R2= 0.3, whereas inthe case ofthe permanent population distribution, itwas 0.92. Apparently, daily population changes inindividual cells are significantly more complex than inthe case ofthe permanent population, which isreflected inthe marked discrepancy inthe distribution ofpulsations byregion during the interval from February 2018toJanuary 2020. During this period, changes may have occurred inthe transport scheme (such asthe development ofmetro lines ormodifications tohighway interchanges), which likely influenced the test results. Figure 6. Number ofpermanent population ofcells (left map) and clustering according tothe law ofdistribution ofcell densities ofthe city ofMoscow. Source: calculated bythe authors based ondata from mobile operators. 10000 + 5000 2500 1000 500 250 0Administrative Okrugs: CAO – Central NWAO – North-Western NAO – Northern NEAO – North-Eastern EAO – Eastern SEAO – South-Eastern SAO – Southern SWAO – South-Western WAO – Western NMAO – Novomoskovsky TAO – Troitsky ZelAO – Zelenogradsky Zipf model Linear model Administrative Okrugs: CAO – Central NWAO – North-Western NAO – Northern NEAO – North-Eastern EAO – Eastern SEAO – South-Eastern SAO – Southern SWAO – South-Western WAO – Western NMAO – Novomoskovsky TAO – Troitsky ZelAO – Zelenogradsky