scieee AI-readable full text Open interactive document viewer

Assessing the impact of interregional mobility on COVID19 spread in Spain using transfer entropy

Ponce de León, Miguel; valencia, alfonso; Alex Arenas; Pontes, Camila

Abstract

Human mobility played a key role in shaping the spatiotemporal dynamics of COVID19 transmission. This study employs Transfer Entropy (TE), an information-theoretic approach, to investigate the directional relationship between interregional mobility and COVID19 spread in Spain. Specifically, we use the mobility-associated risk time series, derived from phone-based origin–destination data and local infection prevalence, to estimate the flow of potentially infected individuals between regions. TE is then applied to measure the information flow from mobility-associated risk to regional case counts, enabling us to uncover spatio-temporal patterns of mobility-driven transmission. Using real-world data, we identified provinces that acted as outbreak drivers during the COVID19 pandemic in Spain and detected temporal shifts in the strength and direction of mobility’s influence. Our findings align with key epidemiological events, such as the 2020 summer outbreak in Lleida linked to seasonal workers, and highlight the effects of non-pharmaceutical interventions, including bar closures in Catalunya, on transmission dynamics. Finally, we validated our approach using simulations from a metapopulation SIR model with known transmission pathways, showing that TE can recover mobility-induced transmission structure while reducing indirect or spurious associations. Altogether, our work provides a novel approach to study the effect of interregional mobility on epidemic spread and to uncover spatio-temporal patterns of mobility-driven transmission, offering valuable insights to inform the timing and regional targeting of non-pharmaceutical interventions.

Full text

Assessing the impact of interregional mobility on COVID19 spread in Spain using transfer entropy Miguel Ponce-de-Leon1, Camila Pontes1, Alex Arenas2,3 & Alfonso Valencia1,4 Human mobility played a key role in shaping the spatiotemporal dynamics of COVID19 transmission. This study employs Transfer Entropy (TE), an information-theoretic approach, to investigate the directional relationship between interregional mobility and COVID19 spread in Spain. Specifically, we use the mobility-associated risk time series, derived from phone-based origin–destination data and local infection prevalence, to estimate the flow of potentially infected individuals between regions. TE is then applied to measure the information flow from mobility-associated risk to regional case counts, enabling us to uncover spatio-temporal patterns of mobility-driven transmission. Using real-world data, we identified provinces that acted as outbreak drivers during the COVID19 pandemic in Spain and detected temporal shifts in the strength and direction of mobility’s influence. Our findings align with key epidemiological events, such as the 2020 summer outbreak in Lleida linked to seasonal workers, and highlight the effects of non-pharmaceutical interventions, including bar closures in Catalunya, on transmission dynamics. Finally, we validated our approach using simulations from a metapopulation SIR model with known transmission pathways, showing that TE can recover mobility-induced transmission structure while reducing indirect or spurious associations. Altogether, our work provides a novel approach to study the effect of interregional mobility on epidemic spread and to uncover spatiotemporal patterns of mobility-driven transmission, offering valuable insights to inform the timing and regional targeting of non-pharmaceutical interventions. Keywords Epidemic, Human mobility, Causal inference, Transfer entropy The COVID19 pandemic emerged as an unprecedented global health crisis, challenging healthcare systems, economies, and societal norms worldwide. As nations endeavour to fortify their preparedness against future epidemics, integrating innovative data-driven strategies becomes imperative1. One promising approach is the use of anonymized phone-based mobility data (APMD) and reports of new infections to guide the design of public health responses such as non-pharmaceutical interventions (NPIs)2–5. Over the past decade, APMD applications have emerged as versatile tools, serving to model human mobility6,7, analyze urban spatial structures7,8, infer friendship networks9, map dynamic populations10, predict poverty and wealth11, and aid disaster management12. In the realm of epidemiology, this data has been seamlessly integrated into various classes of epidemiological models to account for population mobility between regions13–18. While population mobility undoubtedly plays a role in the spread of infectious diseases, establishing a causal relationship or coupling between mobility patterns and regional disease outbreaks remains a challenging task. Several studies highlight that the relationship between mobility and the spread of COVID19 is highly contextdependent, shaped by factors such as pathogen-specific dynamics, geographic and demographic variability, and periods of behavioural or policy-driven changes. For instance, Perofsky et al. demonstrated that mobility positively predicted the transmission of endemic respiratory viruses but exhibited a negative or lagging correlation with SARS-CoV-2 during periods of stringent behavioural restrictions. The authors conclude that mobility is most predictive of respiratory virus transmission during significant behavioural shifts and the early stages of epidemic waves19. Similarly, Grantz et al. found that mobility reductions in highly populated counties were linked to declining infection rates during the early stages of the COVID19 pandemic, although this relationship was inconsistent across other regions and timeframes20. Another study integrating multiple mobility 1Barcelona Supercomputing Center, Barcelona, Spain. 2Universitat Rovira i Virgili, Tarragona, Spain. 3Pacific Northwest National Laboratory, Richland, WA 99354, USA. 4ICREA, Passeig Lluís Companys 23 , 08010 Barcelona, Spain. email: [email protected] OPEN Scientific Reports | (2025) 15:31504 1 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports data sources over one year revealed that mobility does not consistently provide significant predictive power for COVID19 spread, underscoring the limitations of mobility data as a universal predictor21. These studies suggest that mobility has a stronger influence on transmission during periods of dramatic behavioural change or at the onset of epidemics. Nonetheless, finding a way to detect whether an increase in the number of cases in a region is likely to ‘cause’ an outbreak in a different region is critical to help in the development of effective NPIs3,22. We found that the big volumes of data in the form of time series generated during the COVID19 pandemic23–27 constitute an excellent opportunity to develop, refine and validate methodologies aimed at elucidating potential causal relationships among various factors, including population mobility, emergence of novel variants, regional outbreaks, etc. Causality inference is a complex and longstanding issue, drawing from various disciplines such as philosophy, economics, statistics, and biology28. Various definitions exist to describe the causal relationship between two variables X and Y29. A widely accepted definition posits that X causes Y if an intervention on X leads to a change in Y28. Leveraging this definition, causal relationships can be deduced from time series data using statistical and data-driven methodologies. These approaches, devoid of explicit mechanistic models of data generation, are commonly termed as “model-free” methods (for an extensive review, see29,30]). Model-free methods can be broadly categorised into two classes31. The first class comprises methods that rely on specific assumptions about the underlying processes generating the observed data. One prominent example is Granger Causality (GC), a statistical test assessing whether variable X enhances the predictive capability of variable Y32,33. While GC has found success across various domains, its utility is constrained by several restrictive assumptions, including linearity and stationarity, among others (see31] for a comprehensive review). Additionally, GC presupposes that certain properties of X can change independently of Y and that information flow is predominantly unidirectional. However, in complex dynamical systems, interactions between components can exhibit bidirectional tendencies, resulting in cyclic information flows34,35. Moreover, dependencies among system components can evolve over time and often manifest nonlinearly, rendering many of GC’s assumptions invalid. In the context of epidemic processes, nonlinear dependencies and transient couplings between regions are anticipated due to the intricate interplay between daily population mobility and social interactions, which drive the spread of infection. Consequently, traditional causal inference methods such as GC encounter challenges in deciphering cause and effect within complex systems where causal relationships can become entangled36. The second class encompasses information-based methods that quantify the additional insights into the dynamics of Y gained from knowledge of X29,37. One notable example is Transfer Entropy (TE), a technique for assessing the directed (time-asymmetric) transfer of information between two stochastic processes38. TE serves as a generalisation of Granger Causality (GC), with both methods being entirely equivalent for Gaussian variables39. While TE is often described as a model-free or nonparametric method, this refers to the fact that it does not assume a predefined model structure for the underlying processes40. However, its application does involve methodological requirements, such as discretization of continuous data, approximate stationarity, and assumptions about the temporal dependencies e.g., the Markov order of the processes. Consequently, TE has found application in studying causal relationships within nonlinear systems across diverse domains such as neuroscience37, complex systems41,42, and economics and social science43,44, among others. TE has been used in the field of epidemic modelling to examine age-structured transmission dynamics during the autumn wave of the 2009 influenza pandemic in the USA45. More recently, TE was applied to quantify the relationship between mobility metrics and COVID19 cases and deaths in some European regions21. Specifically, it was assessed how each region’s internal mobility influenced the number of COVID19 cases in that same region in a static manner, considering the overall correlation during the full period of the pandemic. Here, we present a different analysis that takes into account aspects such as the influence one region has over the others and the temporal dynamics of the pandemic. In this study, we propose a data-driven framework to assess the role of interregional mobility in shaping epidemic dynamics using Transfer Entropy. Our approach focus on quantifying the directional information flow from the time series of inter regional mobility-associated risk (a proxy for the number of potentially infected individuals traveling between regions), to the time series of newly reported COVID19 cases. This enables us to examine whether interregional flows of individuals, weighted by infection prevalence at the source, are associated with subsequent daily reported cases at the destination region. We apply this framework to the COVID19 pandemic in Spain, combining daily confirmed case counts with origin–destination mobility matrices derived from anonymized phone-based mobility data to compute MAR between provinces27. At the national scale, we characterize both global and local patterns of information flow across provinces and identify periods in which specific regions acted as major drivers of information transfer. We further demonstrate the applicability of our method to policy evaluation by analyzing the effect of a nonpharmaceutical intervention (NPI) implemented in Catalunya4, on the coupling between mobility and disease spread. We show that the observed decrease in mobility-associated risk during the intervention period preceded a subsequent reduction in new cases, consistent with the expected impact of the NPI. Finally, we validate our approach in controlled settings using epidemic simulations based on a metapopulation SIR model. These simulations confirm that TE can recover mobility-driven information flow under simplified and controlled scenarios and highlight the importance of accounting for the temporal structure of mobility networks when assessing causality in spatial epidemic spread. Altogether, our work provides a novel approach to study the effect of interregional mobility on epidemic spread and to uncover spatio-temporal patterns of mobility-driven transmission, offering valuable insights to inform the timing and regional targeting of nonpharmaceutical interventions. Scientific Reports | (2025) 15:31504 2 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports/ Material and methods Datasets The COVID19 Flow-Maps are a cross-referenced geographical information system that includes time-series accounting for population mobility and daily reports of COVID19 cases in Spain at different scales of time and spatial resolution27. In this work, we have used different data sets from COVID19 Flow-Maps available at Zenodo. The datasets include daily reports on COVID19 cases46, census population47, and population mobility48 at different levels of geographic resolution49. Since the different data records span different periods, we selected the maximum possible time interval that includes reported entries in all the datasets, including case reports, population, and mobility. We trimmed all the data records from 01-03-2020 to 01-05-2021, the maximum time range to which mobility and population data sets are available. Table 1 summarizes the different datasets used in this work, which are explained in the further subsections. COVID19 case reports We retrieved daily cases for all of Spain reported at the level of provinces and the level of Basic Health Areas (BHA) for Catalunya and Madrid46. Each record corresponds to a geo-referenced time series where each record has an associated date, the corresponding identifier of the layer, a code of the zone (i.e. a province or BHA) as well as the number of cases reported on that date (daily incidence). Then, for a given region i∈L and a date t∈T , we denote the daily reported new cases as Ci(t) . To minimize the impact of artefacts in case reports (e.g., reporting delays or weekly patterns), we applied a 7-day rolling average to smooth the time series Ci(t) . In this context, by ’smoothing’ we refer to the process of averaging the number of new daily COVID19 cases over a 7-day window to reduce short-term fluctuations and irregularities in the data, such as reporting delays or weekly cycles, by producing a more stable and consistent representation of the underlying trends in case counts. The cases dataset for Madrid is reported by day until the end of June and after by week. Therefore, to avoid artefacts due to the differences in the reporting frequency for the case of Madrid, we applied a 14-day moving average. Throughout this work, Ci(t) denotes the smoothed time series of reported cases unless otherwise specified. We also calculate the 10-day cumulative incidence, denoted by C10 i(t) , and use this quantity as a proxy of the number of infected individuals Ii(t) at time t27,50. We also normalized Ci(t) by the total population Ni and rescaled this value to express the incidences as new cases by 100,000 inhabitants and denote it as ˆ Ci(t)=Ci(t)/Ni·100,000 . Population and daily mobility data Mobility and population data records come from a study conducted by the Spanish Ministry of Transport, Mobility and Urban Agenda (MITMA, Ministerio de Transportes, Movilidad y Agenda Urbana) which analyses the mobility and distribution of the population in Spain from February 14th 2020 to May 9th 2021 ( h t t p s : / / w w w . m i t m a . g o b . e s / m i n i s t e r i o / C O V I D 1 9 / e v o l u c i o n - m o v i l i d a d - b i g - d a t a). The data records are based on a sample of more than 13 million anonymized mobile phone lines provided by a single mobile operator whose subscribers are evenly distributed and include two different mobility indicators. Daily origin-destination matrices (ODM) account for the number of trips between 2850 mobility areas that cover almost the entire territory of Spain, reconstructed by combining cell phone antenna coverage areas with districts and municipalities. In this work, we used ODMs48 and population data47, reported at the province and Basic Health areas levels. The raw data are obtained from a single mobile phone operator covering approximately 30% of the Spanish population and are processed using sample expansion methods to estimate mobility flows for the entire population, as detailed in the original dataset publication27). For mobility indicators and population, we will use the following notations: Mij (t) is the daily number of trips from zone i to j with both i, j ∈L and reported at date t, and will use Ni to refer to the total population for zone i∈L . Mobility-associated risk To assess the effect of population mobility on the spread and outbreaks of COVID19, we have previously introduced the Mobility Associated Risk (MAR)27. The MAR Rij (t) is an estimator of the potential number of infected individuals moving from zone i to j on the date t and is given by the following expression: Region Dataset Layer N ◦ patches Spain COVID19 cases Provinces 50 Spain Population Provinces 50 Spain Mobility Provinces 50 † Catalunya COVID19 cases BHA 374 Catalunya Population BHA 374 Catalunya Mobility BHA 374 Madrid COVID19 cases BHA 286 Madrid Population BHA 286 Madrid Mobility BHA 286 Table 1. Summary of the different datasets used in this work. † The cities of Ceuta and Melilla were excluded from the study because of the low coverage found in the phone-based mobility data. Scientific Reports | (2025) 15:31504 3 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports/ R ij (t)= I i (t) Ni Mij (t ) (1) , where the first term on the right side represents the density of infected individuals in the source zone i per the total number of inhabitants Ni and Mij ( t ) , is the number of trips from region i to j at time t. Rij ( t ) is the estimated number of infected individuals commuting between regions i and j at time t. For a given date t, the values Rij ( t ) can be arranged in a | L |×| L | matrix R(t) which represents a temporal directed weighted network where the nodes correspond to the different zones i, j ∈L and the flow corresponds to the estimated number of infected individuals commuting between regions i and j at time t. When working with simulated data, Ii is directly calculated through numerical integration of the model equations (see Eq. 8). Therefore, to obtain the mobility-associated risk Rij ( t ) between i and j for the simulated data Ii is divided by the total population Ni at region or patch i and then, multiply this value by the number of commuters Mij ( t ) . When dealing with real-world data, COVID19 incidence time series report the number of new daily cases reported Ci( t ) for a region i. The calculation of Rij ( t ) , as defined in Eq. (1) requires the number of infected individuals Ii( t ) in the region i at time t, a quantity unknown in real-world scenarios. Hence, we used the accumulated incidence over the ten days C10 i( t ) preceding the report as a proxy of the number of infected individuals Ii( t ) at time t. Then, we scale C10 i( t ) by the total population Ni and use this quantity together with the trips Mij ( t ) and calculate Rij ( t ) from i to j27. Transfer entropy Basic definitions Transfer entropy (TE) is an information-theoretic measure that quantifies the directional flow of information between two time-dependent discrete random processes, X(t) and Y(t), defined over a discrete alphabet or set of symbols A38. Specifically, TE measures the reduction in uncertainty about the future value yt+1 ∈A when incorporating knowledge of both the past of Y(t) and the past of X(t), denoted y(k) t and x(l) t , respectively. This approach provides insights into potential directed (non-symmetric) dependencies between variables and has been widely used to infer directional influences in time series. To compute TE, joint and conditional probabilities are estimated from the empirical frequencies of symbol sequences in the observed time series. The transfer entropy from X to Y is defined as: TE X→Y(k,l)= ∑ x,y p(yt+1,y(k) t,x (l) t) log2 p ( yt+1|y (k) t,x (l) t ) p(yt+1 | y(k) t) (2) where y(k) t=(yt,...,y t−k+1) and x(l) t=(xt,...,x t−l+1) represent the past states of Y and X up to orders k and l, respectively. Following the original formulation by Schreiber38, we adopt the commonly used setting of k=l=1 in our analyses. Assumptions of transfer entropy analysis Although Transfer Entropy (TE) is often described as a model-free method, its application to empirical time series involves several methodological assumptions. First, TE is typically defined for discrete symbolic sequences; thus, continuous-valued data must be transformed accordingly (see Discretisation below). Second, TE implicitly assumes that the underlying stochastic processes exhibit Markovian properties of a given embedding order. In our case, this assumption is justified by modelling epidemic dynamics at the metapopulation level, where transitions between states have been shown to follow Markovian approximations51. While we acknowledge that alternative non-Markovian approaches are gaining traction in the literature52, we consider the Markov assumption both reasonable and empirically grounded for our chosen modelling framework. Finally, TE analysis requires approximately stationary time series. When this condition is not met, a common approach is to compute first-order differences, which often results in a stationary series. In our analysis of TE between mobility and new case reports, we apply this differencing procedure to ensure stationarity in all the time series. Calculation of transfer entropy between mobility-associated risk and incidence time series In this work, we aim to assess the impact of mobility on the dynamics of the epidemic spread of COVID19. We hypothesise that if an outbreak in a target region j is triggered by infected individuals arriving from a source region i, there might be a detectable flow of information from the time series of mobility-associated risk Rij (t) (see Eq. 1) to the new infections in region j (scaled by 100.000 inhabitants) denoted by ˆ Ci(t) . Both Rij (t) and ˆ Ci(t) are continuous-valued time series: the former represents the estimated flux of infected individuals between regions (combining trip volume and the density of infected individuals), while the latter corresponds to the daily new cases per 100,000 inhabitants. To quantify the directional relationship between them, we measure the transfer entropy (TE) from Rij (t) to ˆ Cj(t) , as defined in Eq.2. Time series stationarity. Since the TE framework requires the input time series to be stationary, we first tested the stationarity of both Rij (t) and ˆ Ci(t) using the Augmented Dickey-Fuller (ADF) test. In most cases, neither time series could be assumed stationary (see Supplementary Table 1). To address this, we applied a firstorder differencing transformation to both series to achieve stationarity: ∆R ij (t)=R ij (t)−R ij (t−1) ∆ ˆ Cj ( t )= ˆ Cj ( t )− ˆ Cj ( t −1) (3) Scientific Reports | (2025) 15:31504 4 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports/ We then applied the ADF test to the transformed time series ∆Rij (t) and ∆ˆ Cj(t) and found that all of them were stationary, with p-values below the 0.05 significance level, except for one case: the mobility-associated risk Rij (t) between the provinces of Zaragoza and Badajoz (see Supplementary Table 1). Discretization of continuous time series. In addition to the stationary condition, the framework of Transfer Entropy also requires the input data to be discrete. A common method for transforming a continuous-valued series X(t) into a symbolically encoded series S(t) involves partitioning the values into a limited number of bins43,53. This method enables each value in the time series X(t) to be assigned an integer label (−1,0,1) S (t)= { − 1for X ( t ) ≤q1 0for q1<X(t)<q 2 1for X(t) ≥ q2 (4) The definitions of q1 and q2 can be done by either defining a set of uniform set of bins or by calculating quantiles from the Kernel Density Estimation of the distribution of observed values. In our case, the time series ∆Rij (t) and ∆ˆ Cj(t) units in the number of individuals, and therefore, a natural choice for partitioning the observations into discretized values is to set q1=−q and q2=q with q=1 . While we test the sensitivity of the results to the selection of q, the results presented in the Results section are for q=1 . Finally, we will use Xij (t)∈ {−1,0,1} and Yj(t)∈ {−1,0,1} , with i, j ∈L and t∈T to refer to the transformed time-series of ∆Rij (t) and ∆ˆ Cj(t) , respectively. Then, to estimate the effect of the population mobility from i to j in the number of newly reported cases we will measure the transfer entropy between Xij (t) and Yj(t) that for simplicity will be denoted as TEi→j . Considering incubation periods and reporting delays. To account for the incubation period and delays in case reports, we introduce a time delay δ (days) between the mobility-associated time series Xij (t) and the new case series Yj(t+δ) . Previous studies have reported incubation periods ranging from 5 to 6 days, with most symptomatic cases developing within 11 days54,55. Global and local transfer entropy. For each pair of regions i and j, we compute the global transfer entropy TEi→j using the full-length time series Xij (t) and Yj(t+δ) . To capture temporal variations in information flow, we also compute a local (dynamic) transfer entropy TEi→j(t) using a sliding window approach. Specifically, we define a window of size ω days and slide it across the time series in daily increments. For each time point t, we compute the TE using the mobility-associated risk time series Xij (s) over the interval s∈[t, t +ω] , and the delayed incidence time series Yj(s) over s∈[t+δ, t +ω+δ] . The resulting TE value is assigned to the right edge of the window, i.e., to time t+ω+δ . This assignment convention ensures consistent alignment with the most recent time point included in the estimation window, which is particularly relevant for retrospective or real-time monitoring applications. We selected ω= 28 days to smooth out short-term fluctuations due to weekly reporting artifacts while preserving sufficient temporal granularity. The lag δ=7 days accounts for typical delays between mobility-induced exposure and case reporting, and both parameters are used consistently across the synthetic simulations and the real COVID19 data analysis. Testing the effect of mobility. To investigate the role of mobility, we compute both global and local transfer entropy using the differenced time series of infections Yi(t) and Yj(t) , thereby excluding the influence of mobility. This allows us to quantify the information transfer between the epidemic dynamics of regions i and j independently of mobility data. To avoid confusion, we explicitly denote this case-only transfer entropy as TEcases i→j . Effective transfer entropy Some authors have claimed that transfer entropy estimates are likely to be biased due to small sample effects and thus propose a modified version of the transfer entropy, called the Effective Transfer Entropy56. Given a source and target time series X(t) and Y(t), respectively, the effective transfer entropy ETEX→Y is computed by subtracting from the TEX→Y (Eq. 2) the term TE X shuffled→ Y as follows: ETE X → Y (k,l)=TE X → Y −TE X shuffled→ Y (5) where TE X shuffled→ Y is the transfer entropy calculated using a shuffled version of the time series of the variable X(t). This definition is an attempt to reduce both the statistical dependencies between the time series X(t) and Y(t) and the time series dependencies of X(t). This procedure’s base relies on generating a new time series by a random realignment of the values from the time series X(t) to obtain the shuffled time series Z(t) and measuring the TE between the shuffled variable Z(t) and Y(t). This procedure is repeated n times and averaged to obtain TE X shuffled→ Y56. Directionality index An essential property of Transfer Entropy is its asymmetry, i.e. TEX→Y=TEY→X , this property allows quantifying the directional coupling between components of complex systems. Therefore, the difference between TEX→Y and TEY→X can provide information regarding the dominant direction of the information flow53. The directionality index, DIX→Y , measures the asymmetry in the transfer of information between two processes and is given by: DI X → Y =TE X → Y −TE Y → X (6) Scientific Reports | (2025) 15:31504 5 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports/ A positive value of DIX→Y means process X(t) is transferring a net flow of information to process Y(t), whereas a negative value of DIX→Y means that the net flow of information goes from Y(t) to X(t). Therefore, we use the net directional flow of information, DI+ X→Y given as follows: DI + X → Y= { DI + X→Y, if DI + X→Y> 0 0,otherwise (7) to measure the dominant direction of the information flow between two components X(t) and Y(t). Epidemic simulations We conducted epidemic spread simulations using a Susceptible-Infected-Recovered (SIR) compartmental model within a metapopulation framework. In this model, a population of N individuals resides across a set L of regions connected by a network of daily commuters, who may travel to a different region for work each day. Here, Nij denotes the fraction of individuals living in the region i who commute daily to j for work, and N i =∑j Nij represents the total population residing in i (Fig. 1). Moreover, we use Sij (t) and Iij (t) to denote the number of susceptible and infected individuals living in region i and working in region j at time t, respectively. If the first infected individuals are introduced in a region i, the epidemic will spread to other areas via infected daily commuters or through susceptible individuals who commute to region i and become infected in their work region and subsequently carry the infection back to their home region. To simulate infection during working time in the region i, we draw the number of new infections from a binomial distribution based on the total pool of susceptible individuals S i =∑j S ji , which includes individuals living and working in i ( Sii ) and all other individuals working in i. The force of the infection is defined as λ work i=( β 2 Iii +∑ jIji Nii+ ∑j Nji ) where β ( days−1 ) is the disease transmission rate and the terms Iii and Iji represent the infected individuals living and working in i those commuting from j to i, respectively. A similar approach is used to simulate infection events at home locations for the remainder of the day. Here, the number of new infections is drawn from a binomial distribution based on the total pool of susceptible individuals S i =∑j Sij with a force λ home i=β 2(Iii +∑ jIij Nii+ ∑j Nij ) . Below, there is a set of differential equations that govern the dynamics of the systems: dS ij dt =−Sij (λwork i+λhome i) dIij dt =−µIij +Sij (λwork i+λhome i ) dZij dt =µIij (8) To avoid notational confusion with the mobility-associated risk metric Rij (t) , we denote the number of recovered individuals living in region i and working in region j as Zij (t) , instead of the usual Rij (t) used in standard SIR-type models. Further details on this simulation approach can be found in previous works13,14. For Fig. 1. SIR metapopulation model. Panel A represents a metapopulation composed of n zones or regions connected through a network of daily commuters. For two areas i and j, Nij denotes the total population that lives in zone i and works in region j. Panel B shows a schematic representation of the SIR model in a metapopulation. S, I and R denote the three different stages of disease and Sij (t) , Iij (t) , Zij (t) denote the number of susceptible, infected and removed individual living in zone i and working in zone j at time t, respectively. At any given time t, the state of the model and any given time can be arranged as a tensor of dimension n×n×3 . Scientific Reports | (2025) 15:31504 6 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports/ our study, we conducted numerical simulations using EpiCommute14, a Python implementation of the approach developed by Tizzoni et al.13, available at https://github.com/franksh/EpiCommute/. We evaluated three metapopulation structures of increasing complexity: a simple chain, a ring, and an extended star network (Supplementary Fig. 21). In each case, we assumed populations of equal size ( N= 100,000 individuals) and a fixed number of commuters ( Nij =5,000 ) if regions i and j are connected. We also considered a realistic scenario, modelling the metapopulation as the provinces of Spain, connected by recurrent mobility patterns inferred from anonymized phone-based mobility data (see the subsection on Datasets for further details). Since the mobility in the SIR models is static, we simulated two scenarios, considering two different recurrent mobility matrices obtained from the average mobility over one week before the lockdown (from 14-03-2020 to 21-03-2020) and during the lockdown (from 14-04-2020 to 21-04-2020). The corresponding mobility matrices were computed by averaging daily flows over one week before the lockdown (March 14–21, 2020) and one week after (April 14–21, 2020). For each metapopulation model, we simulated the epidemic dynamics using a basic reproductive number of R0=2.5 , which represents a moderate level of transmissibility consistent with COVID19 estimates, typically ranging between 2.43 and 3.1057. The average infectious period was set to µ−1=5 days, in line with prior studies14. Due to the stochastic nature of the simulations, we performed 1,000 independent runs per scenario and averaged the results across all replicates. Each simulation was initialized with three infected individuals introduced into a single subpopulation; for the Spain model, this seed population was located in Madrid. To analyze the simulated outbreaks, we extracted time series representing both mobility-associated infection risk and daily incidence. Specifically, we used Iij (t) to denote the number of infected individuals residing in region i and traveling to region j at time t. This quantity reflects the actual number of infectious travelers and was used instead of a proxy for mobility-associated risk. Daily new infections in each region j, denoted Cj(t) , were computed by taking the difference in the number of susceptible individuals between consecutive time steps: C j(t)= ∑ k Sjk(t)− ∑ k Sjk(t− 1) Ones we obtained the time series we compute the first order difference and discretize the data following the approach used for the real datasets (see Calculation of Transfer Entropy between mobility-associated risk and incidence time series) and subsequently computed both global and dynamic effective transfer entropy (ETE) and directional influence (DI) using ω= 28 (days) and δ=7 (days). Results and discussion The results of this study are structured to systematically explore the interplay between population mobility and the spread of infectious diseases, using Transfer Entropy (see Fig. 2). The first section analyses the relationship between mobility patterns and COVID19 case data across Spanish provinces, identifying regions that significantly influenced the epidemic’s progression. The second section focuses on the dynamic nature of these interactions, using local or dynamic TE to reveal temporal changes in information flow. In the subsequent section, the impact of non-pharmaceutical interventions is assessed, with the closure of bars and restaurants in Catalunya serving as a case study to illustrate how mobility reductions significantly correlate with the infection trends. Finally, simulated epidemic scenarios validate the framework, demonstrating the critical role of mobility data in avoiding spurious correlations and accurately capturing the drivers of disease spread. Measuring information flow between mobility patterns and epidemic spread We used the proposed approach to investigate the role of population mobility in the spread of COVID19 in Spain from 1 March 2020 to 1 May 2021. Phone-based anonymised mobility data was used to compute the mobilityassociated risk Rij (t) between regions and the time series of daily reported cases by 100,000 inhabitants ˆ Ci(t) for the different regions i (see Fig. 2A-B and Datasets for more details). We first focus the analysis at the national level, considering as regions the fifty provinces of Spain and excluding the autonomous cities of Ceuta and Melilla (see Supplementary Fig. S5 and Material and Methods Sections). We measured the global TEi→j between each pair of provinces in Spain (see Fig. 2C). To discretise the time series of mobility-associated risk and daily reported COVID19 cases, we set q=1 (see Calculation of Transfer Entropy between mobility-associated risk and incidence time series for further details on the calculation). The results show that Madrid is the province that exhibits the highest TEi→j , which is expected due to fact that Madrid’s high population density, its central location and Madrid city is Spain’s Capital (see Fig. 3A and C and Supplementary Table S1). Furthermore, we analysed the directional information difference DI+ i→j , and total directional information DIi , among provinces (Fig. 3B). The resulting patterns differ markedly from those seen in the analysis of total transfer entropy TEi→j (Fig. 3A). For instance, Madrid, which shows the highest value of total TEi , shows a slightly negative value for the DIi , indicating that the information received by Madrid from other provinces exceeds the amount it transmits to them (Fig. 3B). Comparing province rankings based on TEi and DIi reveals substantial changes, highlighting potential asymmetries in information transfer between neighbouring regions. To further assess this asymmetry, we compared TEi→j and TEj→i , finding a significant correlation ( R=0.62 and p-value = 10−20 ) (see Supplementary Fig. S6A), suggesting that information flow between provinces is relatively balanced. However, notable exceptions exist, such as between A Coruña and Lugo, or Asturias and Lugo, where one province exhibits a substantially stronger influence over the other. These cases illustrate instances of pronounced directional influence between regions. For each source province, we estimate the distribution of the maximum values maxjTEij for each province i (see Supplementary Fig. s7). We calculate the 10% percentile of these maximum values and use it as a threshold Scientific Reports | (2025) 15:31504 7 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports/ to filter out small interactions. After applying this filter, we represent the TEi→j as a directed network on the map, finding that the flow of information primarily occurs between neighbouring provinces (Fig. 3A). Interestingly, the results show almost no Transfer Entropy for certain provinces, including the Canary Islands, the Balearic Islands, and Soria. To investigate the contribution of mobility in the observed patterns of information flow, we computed the TEcases i→j , i.e., the information flow between the time series of cases without considering the mobility (see Calculation of Transfer Entropy between mobility-associated risk and incidence time series). We found significant transfer entropy between almost every pair of provinces (Fig. 3D), suggesting that incorporating mobility helps filter out interactions that may otherwise be confounded or indirect when using case data alone (see Fig. 3C). It is important to note that the presence of significant TE between case time series, even when mobility data is not explicitly considered, does not necessarily indicate misleading or invalid results. Instead, it may reflect dependencies arising from unobserved or indirect mobility pathways, regional synchrony, or shared external influences. Nevertheless, by incorporating direct mobility measurements in the form of the mobilityassociated risk, we aim to better resolve the specific pathways through which inter-regional influence occurs, thereby enhancing the interpretability of the observed information flow and grounding it in known mechanisms of disease transmission. We also compared the asymmetry between TEcases i→j and TEcases j→i and observed a remarkable reduction in the correlation coefficient ( R=0.19 ), emphasizing the critical role of mobility in shaping these dynamics (see Supplementary Fig. S6B). This decrease in symmetry suggests that case-only time series reflect more indirect or confounded dependencies, whereas mobility-informed TE captures more Fig. 2. Measuring transfer entropy between time series of phone-based mobility data and incidence to explore patterns of disease spread. Panel A displays the COVID19 incidence ˆ Cj(t) for region j, representing new daily cases per 100,000 inhabitants. This panel also includes the first-order difference ∆ˆ Cj(t) and the resulting time series Yj(t) , derived by discretising ∆ˆ Cj(t) . Panel B shows the mobility-associated risk between source region i and target region j ( Rij (t) ), together with the first-order difference ∆Rij (t) and the resulting time series Xij (t) , derived by discretizing ∆Rij (t) . Panel C illustrates the calculation of the global Transfer Entropy TEi→j , which quantifies the information flow between the complete time series Xij (t) and Yj(t) to examine the influence of source region i on the case dynamics in target region j. Panel D illustrates the calculation of the local or dynamic Transfer Entropy TEi→j(t) using sliding windows of size ω days to account for the local changes in the flow of information, considering the values of Xij (t) in the time interval [t, t +ω] and the values of Yj(t) in the time interval [t+δ, t +δ+ω] . The base map was retrieved from CNIG h t t p s : / / c e n t r o d e d e s c a r g a s . c n i g . e s / C e n t r o D e s c a r g a s / under CC-BY 4.0 Licence h t t p s : / / w w w . i g n . e s / r e s o u r c e s / l i c e n c i a / C o n d i c i o n e s _ l i c e n c i a U s o _ I G N . p d f. Scientific Reports | (2025) 15:31504 8 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports/ balanced, bidirectional influence patterns, as expected from the generally reciprocal nature of human mobility between regions We further examined the relationship between inter-provincial distances and TEi→j , finding a negative correlation R=−0.56 where TEi→j decays with distance (see Supplementary Fig. S8A). This result is expected since the average number of trips Mij also shows a negative relation with distance (see Supplementary Fig. S8C). If the analysis is performed between distance and TEcases i→j , we found that the correlation coefficient drops to R=−0.24 (see Supplementary Fig. S8B). To address potential biases in Transfer Entropy calculations arising from the potentially limited sample size, we conducted a complementary analysis using Effective Transfer Entropy ETEi→j . The results from ETEi→j were consistent with those from TEi→j (see Fig. 3C and Supplementary Fig. S9), as expected given the strong correlation between TEi→j and ETEi→j (see Supplementary Fig. S10). We also computed the ETEcases i→j and found that even with the correction, there is still information flow between almost every pair of provinces (see Supplementary Fig. S11). Finally, we performed a sensitivity analysis and confirmed that the observed TE and ETE patterns are robust to changes in the delay parameter δ days and the discretisation threshold q (see Supplementary Fig. S1, S2, S3 and S4), with optimal and stable results obtained around δ=7 days and q≈1 (see Supplementary section Sensitivity Analysis of TE and ETE to Parameters δ and q). Overall, these results suggest that incorporating mobility data into TE analysis not only improves the interpretability of inter-regional influence but also helps eliminate misleading correlations present in case-only analyses. Fig. 3. Transfer entropy TEi→j between provinces. Panel A shows the TEi→j between provinces represented by arrows, and the colour of each patch represents the total TEi transferred by each province i. Panel B shows the DI+ i→j between provinces represented by arrows, and the colour of each patch represents the total DIi transferred by each province i. In panels A and B, values were filtered using the 10% percentile from the distribution of the maxjTEij and max jDI + i→j , respectively, as explained in the main text. Panels C and D show the Transfer Entropy TEi→j between provinces of Spain, considering and not considering mobility, respectively. TE calculations were performed setting δ=7 (days) and q=1 . The Canary Islands are excluded from Panels A and B due to the absence of measured values for TEj→i and DI+ i→j . The base map was retrieved from CNIG https://centrodedescargas.cnig.es/CentroDescargas/ under CC-BY 4.0 Licence h t t p s : / / w w w . i g n . e s / r e s o u r c e s / l i c e n c i a / C o n d i c i o n e s _ l i c e n c i a U s o _ I G N . p d f. Scientific Reports | (2025) 15:31504 9 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports/ information flow. These changes are mirrored in the patterns of dynamic DI over time (Supplementary Figs. S28 and S29), highlighting how mobility restrictions reshaped the epidemic influence network. Altogether, these findings reinforce the relevance of mobility data in uncovering potential causal pathways of epidemic spread and emphasize the value of dynamic information-theoretic measures in characterizing evolving transmission dynamics. We also examined the relationship between inter-provincial distances and TEi→j in both the pre-lockdown and lockdown scenarios. A significant negative correlation was observed in both cases– R=−0.19 ( p= 10−7 ) pre-lockdown and R=−0.18 ( p= 10−5 ) during lockdown–indicating that transfer entropy tends to decrease with increasing distance between provinces (Supplementary Fig. S30A and B). These findings are consistent with those obtained from the real-world data (Supplementary Fig. S8). To test whether this spatial dependency arises from mobility structure or is an artifact of the simulation, we conducted additional experiments under a homogeneous mixing assumption. In these simulations, all outgoing trips from each province were redistributed uniformly across other regions, thereby removing the natural correlation between mobility and geographic distance. Under this setup, the correlation between inter-provincial distances and TEi→j became statistically non-significant (Supplementary Fig. S30C), supporting the idea that the spatial pattern of information flow is driven by structured mobility patterns rather than random mixing. Finally, we compared the global ETE values obtained from real-world data with those derived from simulations. The overall Pearson correlation was R=0.255 with a highly significant p-value of 10−38 (Supplementary Fig. S31). While modest, this correlation reflects the fundamental differences between simulated and real epidemic dynamics. In the simulations, mobility is static, and each province experiences only a single epidemic wave. In contrast, the real-world epidemic involves dynamic changes in mobility and multiple waves of transmission. Despite these simplifications, the simulations capture essential features of the observed TE patterns, providing valuable insights into the mechanisms underlying epidemic spread. Taken together, these results confirm that the observed TE signals are not merely statistical artifacts of noisy real-world data, but can mechanistically arise from structured, mobility-driven transmission processes. The use of synthetic outbreaks enabled us to isolate and scrutinize the causal pathways of epidemic spread, demonstrating how TE can capture influence patterns aligned with known transmission routes. By comparing networks with and without incorporating mobility data, we identified the conditions under which spurious or indirect interactions may appear. This reinforces the validity of our methodological framework and highlights the value of mobility-informed TE analysis in detecting meaningful, time-resolved influence patterns in epidemic dynamics. Conclusions In this study, we introduced a mobility-informed Transfer Entropy (TE) framework to explore the directional relationships between population mobility and infectious disease spread. By applying this methodology to both synthetic epidemic simulations and real-world COVID-19 data, we demonstrated that incorporating explicit mobility patterns helps reduce spurious associations and enhances the detection of potentially meaningful information flows within complex epidemiological systems. A key challenge in inferring causal relationships in such systems is disentangling the impact of mobility from confounding factors such as non-pharmaceutical interventions, behavioural adaptations, and reporting delays. Our findings showed that conventional TE analyses based solely on incidence time series are prone to false positives, capturing apparent dependencies that do not reflect true transmission pathways. By integrating mobility data coupled with the epidemic state of the source region, our approach reduces the likelihood of such artifacts and offers a more nuanced understanding of disease dynamics–particularly in identifying regions that may act as transmission hubs. Nevertheless, we emphasize that TE, while useful for detecting directional dependencies, does not by itself establish causality. Rather, it offers a probabilistic measure of information transfer that must be interpreted in light of epidemiological context and other sources of evidence. Our framework addresses some limitations of conventional TE applications, but it remains subject to challenges such as unobserved confounders and temporal aggregation effects. The interpretation of TE signals–whether they reflect propagation, suppression, or stochastic coincidence–requires careful domain-specific analysis. Despite these caveats, the temporal and spatial patterns we identified were consistent with known features of the COVID-19 epidemic in Spain, lending support to the utility of our method. By integrating mobility data into information-theoretic analysis, our framework contributes to the broader development of data-driven tools for real-time epidemic assessment. Future work should aim to combine such approaches with complementary causal inference methods to strengthen the identification of transmission drivers and inform more targeted public health interventions. Data availability The datasets generated and/or analysed during the current study are available in the Zenodo repository, h t t p s : / / d o i . o r g / 1 0 . 5 2 8 1 / z e n o d o . 1 2 2 0 7 2 1 2 , https://zenodo.org/communities/flow-maps/. Received: 18 February 2025; Accepted: 21 August 2025 References 1. Pascal P. K., et al. Enhancing global preparedness during an ongoing pandemic from partial and noisy data. PNAS Nexus, 2(6), pgad192, (2023). 2. Oliver, N. et al. Mobile phone data for informing public health actions across the COVID-19 pandemic life cycle. Sci. Adv. 6(23), eabc0764 (2020). Scientific Reports | (2025) 15:31504 16 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports/ 3. Perra, N. Non-pharmaceutical interventions during the COVID-19 pandemic: A review. Phys. Rep. 913, 1–52 (2021). 4. Smith, M, Ponce-de-Leon, M & Valencia, Al. Evaluating the policy of closing bars and restaurants in Cataluña and its effects on mobility and COVID19 incidence. Sci. Rep. 12(1), 9132 (2022). 5. Mitjà, O. et al. Experts’ request to the Spanish Government: Move Spain towards complete lockdown. The Lancet 395(10231), 1193–1194 (2020). 6. González, M. C., Hidalgo, C. A. & Barabási, A.-L. Understanding individual human mobility patterns. Nature 453(7196), 779–782 (2008). 7. Noulas, A., Scellato, S., Lambiotte, R., Pontil, M. & Mascolo, C. A Tale of many cities: Universal patterns in human urban mobility. PLoS ONE 7(5), e37027 (2012). 8. Louail, T. et al. From mobile phone data to the spatial structure of cities. Sci. Rep. 4(1), 5276 (2014). 9. Eagle, N., Pentland, A. S. & Lazer, D. Inferring friendship network structure by using mobile phone data. Proc. Natl. Acad. Sci. 106(36), 15274–15278 (2009). 10. Deville, P. et al. Dynamic population mapping using mobile phone data. Proc. Natl. Acad. Sci. 111(45), 15888–15893 (2014). 11. Blumenstock, J., Cadamuro, G. & On, R. Predicting poverty and wealth from mobile phone metadata. Science 350(6264), 1073– 1076 (2015). 12. Yabe, T. et al. Mobile phone location data for disasters: A review from natural hazards and epidemics. Comput. Environ. Urban Syst. 94, 101777 (2022). 13. Tizzoni, M. et al. On the use of human mobility proxies for modeling epidemics. PLoS Comput. Biol. 10(7), e1003716 (2014). 14. Schlosser, F. et al. Covid-19 lockdown induces disease-mitigating structural changes in mobility networks. Proc. Natl. Acad. Sci. 117(52), 32883–32890 (2020). 15. Arenas, A. et al. Modeling the Spatiotemporal Epidemic Spreading of COVID-19 and the Impact of Mobility and Social Distancing Interventions. Phys. Rev. X 10(4), 041055 (2020). 16. Ferrari, A. et al. Simulating SARS-CoV-2 epidemics by region-specific variables and modeling contact tracing app containment. npj Digit. Med. 4(1), 1–8 (2021). 17. Balcan, D. et al. Modeling the spatial spread of infectious diseases: The GLobal Epidemic and Mobility computational model. J. Comput. Sci. 1(3), 132–145 (2010). 18. Mazzoli, M. et al. Interplay between mobility, multi-seeding and lockdowns shapes COVID-19 local impact. PLoS Comput. Biol. 17(10), e1009326 (2021). 19. Perofsky, A. C. et al. Impacts of human mobility on the citywide transmission dynamics of 18 respiratory viruses in preand postCOVID-19 pandemic years. Nat. Commun. 15(1), 4164 (2024). 20. Grantz, K.H. et al. The use of mobile phone data to inform analysis of COVID-19 pandemic epidemiology. Nat. Commun. 11(1), 4961 (2020). 21. Delussu, F., Tizzoni, M. & Gauvin, L. The limits of human mobility traces to predict the spread of covid-19: A transfer entropy approach. PNAS nexus 2(10), 302 (2023). 22. Brauner, J. M. et al. Inferring the effectiveness of government interventions against COVID-19. Science 371(6531), eabd9338 (2021). 23. Alamo, T., Reina, D. G., Mammarella, M. & Abella, A. Covid-19: Open-Data Resources for Monitoring, Modeling, and Forecasting the Epidemic. Electronics 9(5), 827 (2020). 24. Harrison, P.W. et al. The COVID-19 Data Portal: Accelerating SARS-CoV-2 and COVID-19 research through rapid open access data sharing. Nucleic Acids Res. 49(W1), W619–W623 (2021). 25. Latif, S. et al. Leveraging data science to combat COVID-19: A comprehensive review. IEEE Trans. Artif. Intell. 1(1), 85–103 (2020). 26. Shuja, J., Alanazi, E., Alasmary, W. & Alashaikh, A. COVID-19 open source data sets: A comprehensive survey. Appl. Intell. 51(3), 1296–1325 (2021). 27. Leon, Miguel P.-de et al. Covid-19 flow-maps an open geographic information system on covid-19 and human mobility for spain. Sci. data 8(1), 1–16 (2021). 28. Judea, P. Causality. Cambridge University Press, Cambridge, 2 edition, 2009. 29. Alex,E. Y. and Wenying, S.. Data-driven causal analysis of observational biological time series. eLife 11, e72518, (2022). 30. Ashley,R. C., Sarah K. H., Elaine, L., Daniel, M., and Joshua,S. W. A Primer for Microbiome Time-Series Analysis. Frontiers in Genetics, 11, (2020). 31. Shojaie, A. & Fox, E. B. Granger Causality: A Review and Recent Advances. Annu. Rev. Stat. Appl. 9(1), 289–319 (2022). 32. Clive,W. J. G. Investigating causal relations by econometric models and cross-spectral methods. Econometrica: journal of the Econometric Society, pages 424–438, (1969). 33. Granger, C. W. J. Testing for causality: A personal viewpoint. J. Econ. Dyn. Control 2, 329–352 (1980). 34. Galea, S., Riddle, M. & Kaplan, G.A. Causal thinking and complex system approaches in epidemiology. Int. J. Epidemiol. 39(1), 97–106 (2010). 35. Sugihara, G. et al. etecting causality in complex ecosystems. Science 338(6106), 496–500 (2012). 36. Daniel, H., Erik, L., Maik, S., & Klaus,R. P. Topological Causality in Dynamical Systems. Phys. Rev. Lett. 119(9), 098301, (2017). 37. Vicente, R., Wibral, M., Lindner, M. & Pipa, G. Transfer entropy—a model-free measure of effective connectivity for the neurosciences. J. Comput. Neurosci. 30(1), 45–67 (2011). 38. Schreiber, T. Measuring information transfer. Phys. Rev. Lett. 85(2), 461 (2000). 39. Barnett, L., Barrett, A. B. & Seth, A. K. Granger Causality and Transfer Entropy Are Equivalent for Gaussian Variables. Phys. Rev. Lett. 103(23), 238701 (2009). 40. Lungarella, M., Ishiguro, K., Kuniyoshi, Y. & Otsu, N. Methods for quantifying the causal structure of bivariate time series. Int. J. Bifurcation Chaos 17(03), 903–921 (2007). 41. Fatimah Abdul Razak and Henrik Jeldtoft Jensen. Quantifying ‘Causality’ in Complex Systems: Understanding Transfer Entropy. PLOS ONE, 9(6):e99462, June 2014. 42. Porfiri, M. & Ruiz Marín, M. Inference of time-varying networks through transfer entropy, the case of a Boolean network model. Chaos: An Interdisciplinary Journal of Nonlinear Science 28(10), 103123 (2018). 43. Thomas, D. & Franziska,J. P. Using transfer entropy to measure information flows between financial markets. Stud. Nonlinear Dyn. Econom. 17(1), 85–102 (2013). 44. Borge-Holthoefer, J. et al. The dynamics of information-driven coordination phenomena: A transfer entropy analysis. Sci. Adv. 2(4), e1501158 (2016). 45. Kissler, S. M., Viboud, C., Grenfell, B. T. & Gog, J. R. Symbolic transfer entropy reveals the age structure of pandemic influenza transmission from high-volume influenza-like illness data. J. R. Soc. Interface 17(164), 20190628 (2020). 46. [dataset], Miguel, P.-de L. et al. COVID19 Flow-Maps Daily Cases Dataset. Zenodo https://doi.org/10.5281/zenodo.4634868, (2021). 47. [dataset], Miguel Ponce-de, L et al. COVID19 Flow-Maps Population Dataset (Maestra 2). Zenodo h t t p s : / / d o i . o r g / 1 0 . 5 2 8 1 / z e n o d o . 4 6 3 5 2 5 8 (2021). 48. [dataset], Miguel Ponce-de, L et al. COVID19 Flow-Maps Daily-Mobility Dataset (Maestra 1). Zenodo h t t p s : / / d o i . o r g / 1 0 . 5 2 8 1 / z e n o d o . 4 6 3 4 8 9 5 (2021). 49. [dataset], Miguel Ponce-de, L. et al. COVID19 Flow-Maps GeoLayers Dataset. Zenodo https://doi.org/10.5281/zenodo.4634662, (2021). Scientific Reports | (2025) 15:31504 17 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports/ 50. Hamada, S. B., Hongru, D., Maximilian, M., Ensheng D., Marietta,M. S. & Lauren,M. G. Association between mobility patterns and covid-19 transmission in the usa: a mathematical modelling study. Lancet Infect. Dis. 20(11), 1247–1254, (2020). 51. Granell, C., Gómez, S., Gómez-Gardeñes, J. & Arenas, A. Probabilistic discrete-time models for spreading processes in complex networks: A review. Ann. Phys. 536(10), 2400078 (2024). 52. Feng, M., Cai, S.-M., Tang, M. & Lai, Y.-C. Equivalence and its invalidation between non-markovian and markovian spreading dynamics on complex networks. Nat. Commun. 10(1), 3748 (2019). 53. Nicoló,A. C. & Paolo, P. Effective transfer entropy to measure information flows in credit markets. Stat. Methods Appl. 31(4), 729–757, (2022). 54. Lauer, S. A. et al. The incubation period of coronavirus disease 2019 (COVID-19) from publicly reported confirmed cases: Estimation and application. Ann. Intern. Med. 172(9), 577–582 (2020). 55. Cheng, Ch. et al. The incubation period of COVID-19: A global meta-analysis of 53 studies and a Chinese observation study of 11 545 patients. Infect. Dis. Poverty 10(05), 1–13 (2021). 56. Marschinski, R. & Kantz, H. Analysing the information flow between financial time series. Eur. Phys. J. B - Condensed Matter and Complex Systems 30(2), 275–281 (2002). 57. D’Arienzo, M. & Coniglio, A. Assessment of the sars-cov-2 basic reproduction number, r0, based on the early phase of covid-19 outbreak in italy. Biosaf. Health 2(2), 57–59 (2020). 58. Oriol, G. Álava ya sufre una mayor incidencia del virus que Lombardía al decretarse la cuarentena. h t t p s : / / e l p a i s . c o m / s o c i e d a d / 2 0 2 0 - 0 3 - 0 9 / a l a v a - y a - s u f r e - u n a - m a y o r - i n c i d e n c i a - d e l - v i r u s - q u e - l o m b a r d i a - a l - d e c r e t a r s e - l a - c u a r e n t e n a . h t m l,(2020). 59. Oriol, G. Un funeral en Vitoria causó más de 60 infectados. h t t p s : / / e l p a i s . c o m / s o c i e d a d / 2 0 2 0 - 0 3 - 0 6 / m a s - d e - 6 0 - p e r s o n a s - s e - c o n t a g i a r o n - a - l a - v e z - e n - u n - f u n e r a l - e n - v i t o r i a . h t m l, (2020). 60. Jessica,M. E. & Virginia,L. El brote entre temporeros de Aragón deja una cuarta comarca en la fase 2. h t t p s : / / e l p a i s . c o m / s o c i e d a d / 2 0 2 0 - 0 6 - 2 3 / e l - b r o t e - e n t r e - t e m p o r e r o s - d e - a r a g o n - d e j a - u n a - c u a r t a - c o m a r c a - e n - l a - f a s e - 2 . h t m l (2020) 61. Coronavirus: Aíslan cincuenta temporeros en un pueblo de Lleida por un positivo. h t t p s : / / w w w . l a v a n g u a r d i a . c o m / l o c a l / l l e i d a / 2 0 2 0 0 6 0 3 / 4 8 1 5 9 0 7 1 7 1 6 5 / a i s l a n - c i n c u e n t a - t e m p o r e r o s - s e r o s - l l e i d a - p o s i t i v o - c o v i d - 1 9 - c o r o n a v i r u s . h t m l, ( 2020). 62. Gómez-Gardeñes, J., Soriano-Paños, D. & Arenas, A. Critical regimes driven by recurrent mobility patterns of reaction-diffusion processes in networks. Nat. Phys. 14(4), 391–395 (2018). Acknowledgements This work has received funding from the Horizon 2020 project CREXDATA (ID: 101092749) and the MePreCiSa project of the UNICO I+D Cloud program, co-financed by the Ministry for Digital Transformation and of Civil Service and the EU-Next Generation EU as financing entities, within the framework of the PRTR and the MRR. CP was supported by the fellowship “Juan de la Cierva - Formación” from the Ministry of Education and Science of Spain (ID: FJC2021-046655-I). AA acknowledges support from the Spanish Ministerio de Ciencia e Innovación (PID2021-128005NB-C21), and the Joint Appointment Program at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated for the U.S. Department of Energy (DOE) by Battelle Memorial Institute under Contract No. DE-AC05-76RL01830. Author contributions M.P., C.P., A.A. and A.V. contributed to the conception of the work. M.P. and C.P. contributed to the acquisition, analysis, or interpretation of data. M.P. implemented the code used in the work. M.P. and C.P. have drafted the work or substantively revised it. All authors have approved the submitted version and are personally accountable for their own contributions. Declarations Competing interests The authors declare no competing interests. Additional information Supplementary Information The online version contains supplementary material available at h t t p s : / / d o i . o r g / 1 0 . 1 0 3 8 / s 4 1 5 9 8 - 0 2 5 - 1 7 2 1 8 - 4 . Correspondence and requests for materials should be addressed to M.P.-d.-L. Reprints and permissions information is available at www.nature.com/reprints. Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit h t t p : / / c r e a t i v e c o m m o n s . o r g / l i c e n s e s / b y - n c - n d / 4 . 0 / . © The Author(s) 2025 Scientific Reports | (2025) 15:31504 18 | https://doi.org/10.1038/s41598-025-17218-4 www.nature.com/scientificreports/