scieee AI-readable full text Open interactive document viewer

An Agent-Based Approach of the Portuguese Population Projection and the Social Security Sustainability

Renato da Silva Fernandes

Full text

An Agent-Based approach of the Portuguese population projection and the Social Security sustainability Renato da Silva Fernandes Dissertação de Mestrado apresentada à Faculdade de Ciências da Universidade do Porto em Matemática 2015 An Agent-Based approach of the Portuguese population projection and the Social Security sustainability Renato da Silva Fernandes MSc FCUP 2015 2.º CICLO An Agent-Based approach of the Portuguese population projection and the Social Security sustainability Renato da Silva Fernandes Mestrado em Engenharia Matemática Departamento de Matemática 2015 Orientador Prof. Doutor Pedro Campos, FEP Coorientador Prof. Doutora Ana Rita Gaio, FCUP Todas as correções determinadas pelo júri, e só essas, foram efetuadas. O Presidente do Júri, Porto, ______/______/_________ Resumo O impacto da estrutura por idades da populac¸ ˜ ao portuguesa na seguranc¸a social ´ e estudado nesta tese, focando nos reformados, pois o pagamento de pens˜ oes de reforma ´ e de longe a func¸ ˜ ao mais importante e mais dispendiosa da seguranc¸a social [Sta15c]. Nos ´ ultimos anos, a idade m´ ınima de reforma aumentou, a percentagem das contribuic¸ ˜ oes para a seguranc¸a social aumentou, as f´ ormulas para o c´ alculo das pens˜ oes foram alteradas e a func¸ ˜ ao do estado na assistˆ encia social diminuiu. A incerteza relativa ` a sustentabilidade do sistema ´ e uma justificac¸ ˜ ao poss´ ıvel para estas alterac¸ ˜ oes. Neste trabalho foram criados alguns modelos baseados em agentes com a especificac¸ ˜ ao da idade e g´ enero dos agentes para uma populac¸ ˜ ao humana e para as suas atividades econmicas, atrav´ es de mecanismos de simulac¸ ˜ ao, regress˜ ao e decis˜ ao. Estes modelos s˜ ao ent˜ ao aplicados ` a populac¸ ˜ ao portuguesa e ao seu sistema de seguranc¸a social numa tentativa de mostrar o crescimento da populac¸ ˜ ao Portuguesa e para tentar descobrir se a seguranc¸a social Portuguesa ´ e sustent´ avel. Um bom conhecimento da evoluc¸ ˜ ao da populac¸ ˜ ao por faixas et´ arias permite a realizac¸ ˜ ao de estudos mais aprofundados. Assim ´ e poss´ ıvel compreender a interac¸ ˜ ao da populac¸ ˜ ao com o mercado de trabalho. Adicionalmente, s˜ ao encontrados padr˜ oes migrat´ orios e ´ e fornecida uma descric¸ ˜ ao populacional mais detalhada. Tamb´ em ´ e poss´ ıvel obter melhores estimativas para o crescimento da populac¸ ˜ ao como um todo. Este trabalho mostra que o tamanho popuac¸ ˜ ao Portuguesa ir´ a decrescer at´ e o ano de 2041, este decr´ escimo ´ e em grande parte referente a pessoas com menos de 65 anos de idade. Isto ´ e principalmente justificado pelo aumento das emigrac¸ ˜ oes, que por sua vez leva a uma reduc¸ ˜ ao nos nascimentos, uma vez que as mulheres emigram em maior volume na sua idade f´ ertil. Adicionalmente, mostra-se que o sistema de Seguranc¸a Social Portuguesa precisaria de ter pelo menos 250 mil milh˜ oes de euros de capital e estar a receber juros nesse mesmo valor desde o ano de 2011, para que o atual sistema de pens˜ oes seja sustent´ avel. Abstract The impact of the age structure on social security is studied in this thesis. The focus of the thesis is on the Portuguese population and the elderly pensioners, because it is by far considered to be the most important and is the most costly function of social security [Sta15c]. In the past years, the retirement age has increased, the contribution percentage to the Social Security system has increased as well, the pension formulas have changed and social assistance has been reduced. This is maybe because of the risk that the system was not sustainable. In this work, some age and gender specific Agent-based models for a Human population and its economical activity are created, through a combination of simulation, regression and decision making mechanisms. These models are then applied to the Portuguese population and its social security system in an attempt to shed some light on how the Portuguese population will grow, and on the sustainability of the Portuguese Social Security. A good knowledge on how population age groups evolve enables further studies to be performed. It is then possible to understand how population and labor market interact. Furthermore, patterns on migration are found and a more detailed population characterization is described. It is also possible to obtain better estimates for the entire population growth. This work shows that the Portuguese population size will decrease until the year of 2041, mainly in the number of persons of age below 65 years-old. This decrease is mostly justified by an increase on the emigration which also leads to a decrease on births because women usually emigrate in their fertile age. Also, it is shown that the Portuguese Social Security system would need to have a base capital of at least 250 thousand million euros and receiving interest on it since 2011, so that the current pension system remain sustainable. Acknowledgements I would like to thank my supervisor Prof. Pedro Campos for suggesting me this thesis theme even though I am from a different faculty and for helping me making the bridge between all the subthemes that take part on this thesis, as well as my co-supervisor Prof. Ana Rita Gaio for helping me maintaining the formalism and scientific rigor demanded for the Master degree in Mathematical Engineering. Also I thank all my friends and colleagues with a special emphasis to Ricardo Cruz for keeping up my moral during this degree and helping me finishing this degree and this thesis. Finally, I send my love to my mother Rosa, father Jos´ e, grandmother Teresa and uncle Manuel. FCUP 1 An Agent-Based approach of the Portuguese population projection Chapter 1 Introduction Today there are over 7 thousand million people alive [CIA13], composed of children, adults and elderly people. The proportion of population in each age group plays an important role both in the whole population growth and in social-economic activities. Demography has been studied for hundreds of years. While many developments have been made so far, in the last few years the development has been less than significant when considering full population growth projections (Coale and Trussell [CT96]). This is a dire problem which can show that either the current developments are not intended for full population growth or that they are not effective. With the evolution of modeling techniques, more specific studies are being made, nowadays. Some models focus on specific demographic phenomena, like nuptiality or mortality, and some models focus on phenomena for different population characteristics, such as age and gender. We have found that trying to model several closely dependent entities, namely fertility, mortality and migrations, while being specific on the age and gender of the population was a true challenge for the author of this work. This is why we had got motivation to study this theme and wanted to move forward as further as possible. In addition, the Portuguese Social Security system sustainability is a very mediated event in Portugal and as the author is also on the verge to enter the system, this system’s sustainability is a great matter of personal concern. One of the possible key aspects for this system’ sustainability is the population proportion that in the future will be contributing for the system and the proportion that will be supported by it. So, having in mind the very close link between the Portuguese population growth and the Portuguese Social Security system, it brought us additional motivation to study the economic events around the problem. However, the aspect that influenced our decision most on whether we would pursue the research was the modeling technique to be used: the agent-based modeling. Agent-based modeling is a recent simulation technique which instead of modeling directly the event, models the intervenor that influences the event. In other words, this is a bottom-up approach. For the study at hand the chosen intervenor to be modeled had to be a human being and with a sufficiently large set of intervenors it is possible to recreate the Portuguese population. With this modeling technique we saw many innovation possibilities, which quickly built our interest over this project. Therefore, a new age and gender specific agent-based model was created to study the demographic and social-economic evolution of the Portuguese population. This model has got some interesting particularities. A base model is created to support multiple fertility, mortality, migration, 2 FCUP An Agent-Based approach of the Portuguese population projection employment and social security system models and this model can easily be expanded to support other entities or events. Therefore, if there is the will to try different growth models for each considered entity or event, there is no need to make a new connection function. New models can be easily added to the base model that is presented here. Another particularity is on the present models, in which some use an additional modeling technique called Mic-Mac modeling, where the actions of an intervenor can affect the whole country and where events in the whole country independent to the intervenor actions can also affect each intervenor. In addition, to the presentation of the models, their results and projections are also disclosed. Offering several possible scenarios for the evolution of the Portuguese population size and for the sustainability of the Portuguese Social Security system, always presented in a summarized and concise way. The following chapters are organized in the following manner: Chapter 2: a historical tour through the more remarkable aspects of modeling and the development of demographic and social security models by time and perspective will be displayed. Chapter 3: the main engine of the developed agent-based model will be shown. This will showcase the main population structure and some applicable models for fertility, mortality and migration. Chapter 4: the population model presented in Chapter 3 is extended to support some economic behaviors. The main mechanism to model the economic status of an agent is detailed. In addition, some models which can be used to model employment and social security systems are presented. Chapter 5: the starting data is presented alongside with its sources. The initialization mechanism of both population and social security is also shown. Finally, the main results are exposed. Chapter 6: the results and respective models are discussed, focusing on the models’ strengths and weaknesses. Chapter 7: the main conclusions of the work are described and some insights for Portugal are drawn. A summarized overview on the created model is observable in Figure 1.1 on the following page, showing all the developed models and the variables for which several values will be tested. The diagram sequence goes from the top to the bottom. FCUP 3 An Agent-Based approach of the Portuguese population projection Population Model Fertility MacMic Mic-Mac Migration Mortality MacMic Mic-Mac Heterogeneity Employment Mic Mic-Mac Variables Retirement Age 68 years66 years 70 years Economic Scenario Stability Recession Prosperity Social Security System Contributionss-Based Earnings-Based Portuguese Initialization Figure 1.1: Model Overview Diagram. FCUP 5 An Agent-Based approach of the Portuguese population projection Chapter 2 Literature Review A model is an attempt to translate the real world into a mathematical formulation. As it is widely known, the real world is a very complex system and to translate it into one single model is impossible. Therefore, many simpler models are needed to translate the world, yet simplifying them implies a loss of accuracy. A model operates over some observable data. These observable data are the model variables. This chapter presents basic ideas from demographic modeling. It starts with a general modeling section, where some modeling methods are shown and then progresses to a section of demographic modeling approaches, showing the demographic perspectives that have been studied. In addition, the origin of social security systems are presented along side some basic social security models for the elderly pension. 2.1 Modeling Approaches A given phenomenon can be modeled by a multitude of methods. Though, there are special characteristics in common among them. These characteristics can be used to group them in big modeling approaches. Here four groups are presented which are not necessarily incompatible: The first two are the base models, the macro and the micro models. The first attends to model aggregated phenomena while the second tries to model disaggregated phenomena, hence the names macro for big entities and micro for small entities. The remaining two groups have their own distinctions, but are combinations of the previous two. One is mic-mac group of models (micro-macro) where both micro and macro approaches are combined to establish connections and complementariness between them. Finally, agent-based models will be presented, which can be any combination of the previous models and the main distinction from them is that instead of modeling the phenomena directly, they model the intervenor. 2.1.1 Macro Models Macro models attempt to govern the environment and usually resort to differential equations, regression methods, time series and others. A general rule is that macro models consist of solving a set of equations in order to obtain the general behavior of the considered phenomena. 6 FCUP An Agent-Based approach of the Portuguese population projection The Lotka-Volterra equations are widely used to model the growth of populations (Lotka [Lot10]). They often resulted in exponential growths and some variants, such as the Malthusian model (Malthus [Mal98]) or the Verhulst model (Verhulst [Ver38]). Lotka-Voltera equations are defined as: dx dt =αx−βxy (2.1.1.1) dy dt =δxy −σy(2.1.1.2) where xis the number of preys, yis the number of predators, tis the time and α,β,δ,σare positive real parameters describing the interaction between the two species. When large amounts of past data are available, a phenomenon may be modeled by regression methods, attempting to explain the phenomenon through a given set of explanatory variables; or may be modeled using time series, finding patterns on the past (Legendre [Leg05]). If the set of used variables is sufficiently small, macro models may achieve results, analytically or numerically, in short time scales. 2.1.2 Micro Models Micro models, in comparison to macro models, attempt to model disaggregated entities or events of the entire environment. Additionally, for models to be more realistic, when possible, researchers introduce stochasticity in the parameters. However, this addition has a severe drawback, the mathematical analysis becomes intractable. Instead, independently from its complexity, a model can always be simulated. The only limitations for the complexity of the model are the time that is acceptable to wait for the results or the computer power, such as ram, processor speed or graphics cards (Coale and Trussell [CT96]). Micro models can be used to validate or evaluate the performance of many theoretical models for which testing data is lacking (Coale and Trussell [CT96]). They can also be used to study rare events and may lead to the discovery of unusual events not expected from the mere study of the theoretical model. Usually micro models are dealt with using the Monte Carlo method, from Neumann, Ulam, and Richtmmyer [NUR47], in juxtaposition to the theoretical analysis. 2.1.3 Mic-Mac Models Mic-Mac models attempt to model a given phenomenon combining the two previously presented approaches. The mic (micro) part models smaller entities, such as small groups of individuals, individuals or particles. This part is usually modeled allowing some stochasticity and is solved through simulation. The mac (macro) part models large scale events, such as the weather, economic growth or net migration. This part is usually deterministic (Gaag, Beer, and Willekens [GBW05]). In these models, the mic part is expected to influence the mac part and vice-versa. Also, this study can reveal the effect of each point of view on the other. The great challenge within mic-mac models relies on the joint modeling of the mic and mac approaches. As one influences the other, the order by which they are executed is also an important choice. FCUP 7 An Agent-Based approach of the Portuguese population projection 2.1.4 Agent-Based Models Agent-based modeling constitutes a recent approach to model phenomena. These have been used, for example, to model the life of the inhabitants of a town (Heiland [Hei03]) and the predator-prey dynamics (Wilensky [Wil99]). Each participant, which will be referred as agent in the future, in the given phenomena is modeled according to the following structure: •definition of each participant permissions and constraints, like allowing individuals to move freely or forbidding two individuals to stand in the same place at the same time; •definition of their rules and interaction patterns, like a buy-sell pattern or a simple rule where every individual must eat to remain alive; •definition of environmental rules, such as climate influences on a person actions or the economic recession. These models allow for a very large number of variables and/or constraints and admit stochasticity. To model and test a phenomena is not always needed past data; with simple and acceptable assumptions these models can be initialized. All this is achieved by the way these models are run. They do not try to solve a system of equations (possibly differential and usually stochastic), which with the allowed number of variables would most likely be impossible to solve. They also do not try to supply a model obtained through regression methods, nor do they try to discover a time series that governs the data, as they usually do not have enough data to achieve such results. These models run through simulation so all they require, besides a decent model and starting variables, is a sufficiently powerful computer to deal with all the interactions between the individuals and the environment. To detect the most usual outcomes, the simulation requires to be run several times (Billari, Ongaro, and Prskawetz [BOP03]). The main aim of a model is to obtain the most probable outcome, however agent-based models and simulation models can also provide rarer events (Chattoe [Cha03]). Rare events are useful for the prediction of worst case scenarios and provide a broader horizon on the situation being modeled. Agent-based models may also explain how a given phenomenon occurs, instead of only offering the final outcome. The interaction history can be stored and a theoretical analysis can be made, building the paths for the possible outcomes (Chattoe [Cha03]). As in most modeling approaches, sensitivity analysis to a variable can be performed, as well as a study of the impact of adding or removing a variable. Due to the large flexibility of this modeling method, hybrid models can be created whereby other modeling methods are incorporated within this method. Incorporating models is possible when data is previously provided or is generated by the simulation itself (Wooldridge [Woo02]). 2.2 Demographic Models In the Greek language, ”demo” means ”the people” and ”graphy” means ”writing, description”. And so, demography is the scientific and statistical study of populations, including of human beings. Demography encompasses the study of the size, structure, and distribution of these populations, and spatial and/or temporal changes in them in response to time, birth, migration, ageing, and death (Preston, Heuveline, and Guillot [PHG01]). There are essentially two different approaches to the study of demography (Preston, Heuveline, and Guillot [PHG01], Caswell [Cas01]): 8 FCUP An Agent-Based approach of the Portuguese population projection 1. Cross-sectional, studying members of a population of different ages in a specified time interval; 2. Cohort, studying the evolution of a population with the same event (e.g. birthing year). Events or entities, such as the birth rates or the number of deaths, are modeled using simple models with reduced number of variables. These simple models may, then, be used to construct more complex and complete models for the whole population structure. If the small models are complex then the whole population structure model would, most likely, become too complex to be dealt with. Demographic models have often one of three objectives (Preston, Heuveline, and Guillot [PHG01], Caswell [Cas01]): 1. To model time series, which attempt to capture empirical regularities and are mostly used as theoretical frameworks, or to determine the quality of demographic data; 2. To attempt to estimate levels and trends in mortality, fertility, and other events or entities; 3. To smooth recorded age-specific time series for fertility, mortality and others. Although most demographic models are macro models which try to model large events or the entire population; there are a few micro-models which try to model the individual (Coale and Trussell [CT96]). Also, a case was found where the modeling was done through the mic-mac approach (Gaag, Beer, and Willekens [GBW05]). These models can be either deterministic, where the outcome is certain, or stochastic, where outcome depends on chance. Usually, most phenomena can be treated both in a deterministic and in a stochastic way (Coale and Trussell [CT96]). 2.2.1 Conventional Life Table Under a closed population model (no migrations), the size of a cohort for any given year can be computed. Let dibe the number of persons in the cohort that die at age iand l0be the number of persons that were born in that cohort. Then lx=l0− x−1 X i=0 di(2.2.1.1) is the number of persons aged xyears-old. In particular, lx l0is the proportion of individuals in the cohort surviving until age x. Defining Txby Tx=Z∞ x lsds ∼ ∞ X i=x i li, (2.2.1.2) which corresponds to the number of person-years to be lived by individuals surviving until age x, and e0=T0 l0is the average number of years lived by the members of the cohort, i.e., the life expectancy at birth (Preston, Heuveline, and Guillot [PHG01]). The study of a cohort corresponds to the study of a population subset that has the same year of birth finding all characteristics of this cohort implies the wait of a reasonable number of years, as every member must die. It is thus impossible to study a currently living cohort, prior to all members dying. In order to solve this problem, the assumption that the death rates within a cohort do not differ from the current death rates in the population with age ifor the individuals of the same age is made. FCUP 9 An Agent-Based approach of the Portuguese population projection Therefore lx= x Y i=o l0qi(2.2.1.3) where qxis the death rate of the individuals in the population that are aged x−1 and that die before they reach xyears old (Caswell [Cas01]). 2.2.2 Mathematical Models of Conception and Birth Keyfitz [Key77], Bongaarts and Potter [BP83], and among others, created the first mathematical models for conception and birth, which have been vastly disseminated until now. Of these, birth interval models have been found to provide broader explanations for the birth cycles (Coale and Trussell [CT96]). Menarche Menopause First Birth Second Birth Third Birth Birth Return of Ovulation Conception Birth Conception Return of Ovulation Intrauterine Death Conception Postpartum Infecundable interval Conception wait Gestation Gestation Conception wait Infecundable interval Reproductive Life Span Birth Interval (without intrauterine death) Time added by intrauterine death Figure 2.1: Birth Intervals. From Figure 2.1, it is possible to observe that the fertile time of a woman can be split into multiple intervals. These intervals range from the date of menarche, which is the first menstruation, until the menopause, in which menstruation stops from occurring. Between these two moments a woman can have several births. The interval starting and ending in consecutive births is also divided into smaller intervals, Bint =Pint +Cint +Gint (2.2.2.1) where Bint is the between births interval; Pint is the postpartum interval; Cint is the conception interval; Gint is the gestation interval. If an intrauterine death occurs, then interval between the intrauterine death and the next birth is shorter than the usual interval between two births because the infecundable interval after an in- 10 FCUP An Agent-Based approach of the Portuguese population projection trauterine death is shorter than a postpartum infecundable interval (Coale and Trussell [CT96]). These models were essentially used as explanatory models for the effectiveness of anti-contraceptive methods, sterility, the effects of heterogeneity and sex-selection. This was done studying the difference between birth intervals when these variables were and were not present (Coale and Trussell [CT96]). 2.2.3 Age Schedules The description of an event’s value over the age of a individual is called an age schedule. This offers a very extensive report for the event at study and allows the creation of high quality life tables and it makes possible to test if already constructed life tables are coherent. For an age schedule to be created it is needed all the accounts of a given event for age cohort, which, as already explained, proves to be impossible in some situations. So it is very important to use mechanisms that can work with missing data. Age Schedules of Mortality As in section 2.2.1 on page 8, the number of deaths at age x(dx) is very important for the construction of a Life Table. So, it is very important to have very good models to predict age schedules of mortality. In 1662, Graunt [Gra62] stated that l6 l0≈64% and l76 l0≈1%, in London, where l0,l6,l76 were computed as in Equation 2.2.1.1 on page 8. He also found that the remainder values were almost proportional to those, which translate to an almost linear growth in mortality. From this, he made one of the first Life Tables in history. Later, in 1825, Gompertz [Gom25] changed the linear growth rate in mortality to an exponential growth, which was further improved in 1860 by Makeham [Mak60], by adding a constant term. Many more models appeared since then, but they will not be discussed here. It is worth noticing, however, that currently developed countries publish several variations of the models they use, varying their parameters to offer the biggest range of possible national outcomes. Often, linear regressions on age schedules of mortality were used to correct life tables, on the years on which the information was incorrect or missing, so that future projections became more accurate, as stated by Coale and Trussell [CT96]. Age Schedules of Nuptiality Nuptiality is usually modeled using empirical data, such as the the distribution of women’s age at marriage, or by making simple and intuitive social assumptions. The first model for nuptiality was developed by Coale and McNeil [CM72] who have found that the average year of the first marriage varied according to the geographic location. In most Asian and African countries, the women’s mean age for the first marriage is bellow 15 years-old, while in Western Europe, it is over 25 years-old. A second model was created by Hernes [Her72] and it is built over two assumptions; 1. The social pressure of a member of the cohort to marry is an increasing function of the percentage of the cohort members already married; 2. The odds to get married decreases as the age increases. As stated by Hernes [Her72], this model appears to be in a high accordance with the United States marriage schedules. FCUP 17 An Agent-Based approach of the Portuguese population projection •To fix ptand vary ct, keeping the yearly pensions constant and to recompute the contributions based on the vowed pensions; •To vary both ctand pt, in order to keep PPAYG(t) = CPAYG(t) in a fair manner for both tax payers and retired population. FCUP 19 An Agent-Based approach of the Portuguese population projection Chapter 3 Population Model An agent-based model for a Human population will be presented here. The basic functionalities are offered, i. e., an agent may give birth to another agent, may die or may leave/enter the country. An agent may get several other functionalities as well. In the case to be presented, the model was applied to the Portuguese population to test the sustainability of the Portuguese Social Security System. Therefore, this model requires the agents to have an activity, employment and retirement status. Additionally, they need to have an education and working proficiency. Finally, they must have a remuneration and pay their social security contributions, if they are working, or they must receive a retirement pension if they are retired. In this chapter, the main population model will be presented, followed by a set of possible functions for the management of the births, deaths and migrations. 3.1 Closed Population Model A closed population model is a population model in which it is assumed that there are no migrations, or that the net migration is 0. Therefore, only the fertility and mortality are considered. The variables used in the closed model that will be presented can be computed using multiple methods. The simplest way to compute and understand them will be presented; however, during programming, other methods were used in order to enhance the performance speed and to reduce the RAM usage. Firstly, the notation and the main components of the model are established. Variables starting with A, G and X correspond to agent variables, global variables and random variables, respectively. The indices a, s, k and y are used to denote age, sex, agent identification and year, respectively. Any variable indexed by a, s, k, y represents the realization of the variable in the agent k, aged a years-old and from sex s, in the year y. A similar interpretation applies to any subset of these indices. The following variables are then defined: y0,yf: starting and finishing simulation years; AAlive a,s,k,y : vital status, associated with a binary code where AAlive a,s,k,y = 1 means that the agent is alive, while AAlive a,s,k,y = 0 means that the agent is dead. Clearly if an agent is dead in a given year, then that agent will remain death in the following year, AAlive a,s,k,y = 0 ⇒AAlive a+1,s,k,y+1 = 0; 20 FCUP An Agent-Based approach of the Portuguese population projection GAlive a,s,y : number of living agents and is given by GAlive a,s,y =X k AAlive a,s,k,y; (3.1.0.1) GMaleFreq y: relative frequency of male agents, which corresponds to total alive male agents divided by the total alive population, GMaleFreq y=P a GAlive a,M,y P a,s GAlive a,s,y ; (3.1.0.2) GBirths a,s,y : number of births of sex s, obtained from female agents aged a years-old; GDeaths a,s,y : number of deaths. If an agent’s age is greater than 0, then he/she must have been alive in the previous year and so, the value is the count of agents which were alive in the year y−1 and are death in the year y. If the age is equal to 0, then the number of deaths is the difference between the total number of agents born in the year y and the total number of agents of age 0 alive at the end of year y, GDeaths a,s,y = #{AAlive a,s,k,y = 0 ∧AAlive a−1,s,k,y−1= 1},a6= 0, (3.1.0.3) GDeaths 0,s,y =X a GBirths a,s,y −GAlive 0,s,y. (3.1.0.4) Different update methods were used for the fertility and mortality in order to find the ones whose results closest to the reality. For the fertility case, there is a mic model in section 3.2.1, a mac model in section 3.2.2, a mic-mac model in section 3.2.3 on the following page and an agent heterogeneity model in section 3.2.4 on the following page. For the mortality case there is a mic model in section 3.3.1 on page 22, a mac model in section 3.3.2 on page 22, a mic-mac Model in section 3.3.3 on page 22 and an agent meterogeneity Model in section 3.3.4 on page 23. 3.2 Fertility 3.2.1 Mic Model A simple micro model is defined here, where the next year fertility rate is computed using the data generated from the current simulated year. Let GFertR a,y+1 be the fertility rate indexed by the age a of the mother and the birth year y plus one; it is given by GFertR a,y+1 =P s GBirths a,s,y P sGDeaths a,s,y +GAlive a,s,y. (3.2.1.1) 3.2.2 Mac Model A simple macro model is defined here, where the next year fertility rate is computed using the current fertility rate and the expected fertility rate growth. Let FCUP 21 An Agent-Based approach of the Portuguese population projection GFertEvo be a vector of size equal to the number of years to be simulated. This vector is composed by the expected mean fertility rate growth for those years and GFertEvo yare GFertEvo elements; GFertR a,y+1 be the fertility rate indexed by the age a of the mother and the birth year plus one with GFertR a,y+1 =GFertR a,y ×GFertEvo and GFertR a,y0=P s GBirths a,s,y0−1 P sGDeaths a,s,y0−1+GAlive a,s,y0−1. (3.2.2.1) 3.2.3 Mic-Mac Model A mic-mac model is defined here, where the next year fertility rate is computed using the current year generated data. The expected growth of the fertility rate is also used as a controlling factor, forcing the overall fertility rate to grow according to controlling factor. Let GFertEvo be a vector with size equal to the number of years to be simulated. This vector is composed of the expected mean fertility rate growth for those years and GFertEvo yare GFertEvo elements; GFertR a,y+1 be the fertility rate indexed by the age a of the mother and the birth year plus one, given by GFertR a,y+1 =P s GBirths a,s,y P sGDeaths a,s,y +GAlive a,s,y×GBirthEvo y. (3.2.3.1) 3.2.4 Agent Heterogeneity One of the greatest advantages of agent-based models is the possibility of having heterogeneous individuals. That functionality is described and explored here. The following variable is created in order to achieve heterogeneity among the agents: XFertR a,y+1 a random variable with σxFertR a,y+1 =min{0.02, GFertR a,y+1 3,1−GFertR a,y+1 3}such that XFertR a,y+1 ∼N(GFertR a,y+1,σxFertR a,y+1 ); (3.2.4.1) This random variable must be between 0 and 1, so the standard deviation σmust be bounded. A factor of 3 is used in each fraction because a normal distributed sample has about 99.75% of its values between a range of three standard deviations from the mean. With this constraint, it is almost guaranteed that the variable will remain between 0 and 1, as desired. With this distribution, each agent is assigned the parameter: AFertR a+1,k,y+1 which is the probability for the female agent k, aged a+1 years-old, to give birth in the year y+1; it is given by AFertR a+1,k,y+1 =xFertR a+1,y+1,xFertR a+1,y+1 ∈XFertR a+1,y+1. (3.2.4.2) 22 FCUP An Agent-Based approach of the Portuguese population projection 3.3 Mortality 3.3.1 Mic Model A simple micro model is defined here, where the next year mortality rate is computed using the data generated on the current simulated year. Let GMortR a,s,y+1 be the mortality rate at year y+1 for agents aged a years-old and with sex s; it is given by GMortR a,s,y+1 =GDeaths a,s,y GDeaths a,s,y +GAlive a,s,y . (3.3.1.1) 3.3.2 Mac Model A simple macro model is defined here, where next year mortality rate is computed using the current mortality rate and the expected mortality rate growth. Let GMortEvo be a vector of size equal to the number of years to be simulated. It is composed of the expected mean mortality rate growth for those years and GMortEvo yare GMortEvo elements; GMortR a,s,y+1 be the mortality rate at year y+1 for agents aged a years-old and with sex s; it is given by GMortR a,s,y+1 =GMortR a,s,y ×GMortEvo yand (3.3.2.1) GMortR a,s,y0=GDeaths a,s,y0−1 GDeaths a,s,y0−1+GAlive a,s,y0−1 . (3.3.2.2) 3.3.3 Mic-Mac Model A mic-mac model is defined here, where the next year mortality rate is computed using the current year generated data. The expected mortality rate growth is also used as a controlling factor, forcing the overall mortality rate to grow according to controlling factor. Let GMortEvo be a vector of size equal to the number of years to be simulated. It is composed of the expected mean mortality rate growth for those years and GMortEvo yare GMortEvo elements; GMortR a,s,y+1 be the mortality rate at year y+1 for agents of age a and sex s with GMortR a,s,y+1 =GDeaths a,s,y GDeaths a,s,y +GAlive a,s,y ×GMortEvo y(3.3.3.1) If, for some characteristics, the population size is very small, then this formula is replaced by the correspondent mac model formula. FCUP 23 An Agent-Based approach of the Portuguese population projection 3.3.4 Agent Heterogeneity Mortality agent heterogeneity is made in a completely analogous method to the fertility counterpart ( 3.2.4 on page 21). Then AMortR a+1,s,k,y+1 which is the probability for the agent k of aged a+1 years-old and with sex s to die in the year y+1; it is given by AMortR a+1,s,k,y+1 =xMortR a+1,s,y+1,xMortR a+1,s,y+1 ∈XMortR a+1,s,y+1; (3.3.4.1) 3.4 Algorithm The previous variables have to be updated at each time step. In real life any of the previously stated occurrences may happen in any given time. However, to maintain the control over the evolution, an update order needs to be forced. The evolution process is done according to the following steps: Step 1. Increase the simulation year by one; Step 2. Age every living agents by one; Step 3. Give birth to new agents according to the birth rates of the agents, i.e., two values u1and u2are randomly sampled from U(0, 1) and if u1<AFertR a,s,k,y then a new agent is born. If u2<GMaleFreq ythen the agent is a set as male, otherwise it is set as female. Step 4. Randomly ”kill” agents, i.e., u3is randomly sampled from U(0, 1) and if u3<AMortR a,s,k,y then set AAlive a,s,k,y = 0. Step 5. Compute the next year birth and death parameters and male proportion rates according to the chosen update model. Step 6. Define each agent’s birth and death parameters for the following year. 3.5 Migration 3.5.1 Emigration Model Since, in agent-based models, it is possible to map each intervenor individually, the modeling of the emigration is based on the will of each agent. Each agent determines his/her gain when moving to another country and the gain when staying in Portugal. This gain is computed using several variables as suggested by Bal´ aˇ z, Williams, and Fifekov´ a [BWF14], namely health and safety indicators as well as wages and living costs. In addition the number of Portuguese people that were born in Portugal and that are living in each considered country, the distance of each country to Portugal and the spoken language in the foreign country is also used to compute the gains, as already studied by Anjos and Campos [AC10]. Moreover, observation of Portuguese data reveals that emigration is dependent on age and common sense dictates that work success is also an important factor in the emigration decision. So these two parameters are also considered, but they will not be used to directly model the gain in emigrating but to model the will to emigrate. The considered countries were the ones with more than 50 emigrants in 2011. 24 FCUP An Agent-Based approach of the Portuguese population projection The following variables are considered. Henceforth, cis an index denoting a given country while c0denotes Portugal. GHealth cis a health indicator with AHealthW kas its corresponding weight; GSafety cis a safety indicator with ASafetyW kas its corresponding weight; GWage c,l is a wage indicator with AWageW kas its corresponding weight; GPop cis an indicator for the Portuguese population size, with APopW kas its corresponding weight; GDist cis an indicator for the distance to Portugal, with ADistW kas its corresponding weight; GLang cis an indicator for the Portuguese language, with ALangW kas its corresponding weight; GLimit cis the Portuguese emigration limit, defined by the destination country; GECounter cis a counter for the number of emigrants. The first five indicators range between 0 and 1. The used indicator must be the same for all countries and it is preferable that the data source is the same, because the same indicator may vary in different sources. GLang cequals 1 if Portuguese is the native language and 0 otherwise. The wage indicator changes every year according to the country expected mean wage growth. For the parameter Glimit cthere is no weight because this parameter cannot be influenced by any agent. For each weight a variable is created in order to achieve heterogeneity: XHealthW kfor the health weight such that XHealthW k∼N(0, 0.75 ×AHealthW k); XSafetyW kfor the safety weight such that XSafetyW k∼N(0, 0.75 ×ASafetyW k); XWageW k,s for the wage weight such that XWageW k,e ∼N(0, 0.75 ×AWageW k,e ); XPopW kfor the Portuguese population size weight such that XPopW k∼N(0, 0.75 ×APopW k); XLangW kfor the spoken language weight such that XLangW k∼N(0, 0.75 ×ALangW k); and then the weights are recomputed as AHealthW k=AHealthW k+xHealthW k,xHealthW k∼XHealthW k; (3.5.1.1) ASafetyW k=ASafetyW k+xSafetyW k,xSafetyW k∼XSafetyW k; (3.5.1.2) AWageW k,e =AWageW k,e +xWageW k,e ,xWageW k,e ∼XWageW k,e ; (3.5.1.3) APopW k=APopW k+xPopW k,xPopW k∼XPopW k; (3.5.1.4) ALangW k=ALangW k+xLangW k,xLangW k∼XLangW k. (3.5.1.5) One more parameter is used, in order to take into account age and employment history . Also let AWill a,k,y be the will to emigrate from Portugal of agent k with age a in the year y. This parameter is allowed to change every year, depending on the agent’s current age and work success in the previous year. From Portuguese data concerning the age distribution of emigrants, a resized Weibull probability density function was found to be the function that best describes the emigration behavior depending on age as in Section 5.2 on page 39. The update of the will of each agent needed a calculation of the derivative of the Weibull probability density function. Let FCUP 25 An Agent-Based approach of the Portuguese population projection W(x;λ,k) = k λx λk−1 e−(x λ)k(3.5.1.6) be the probability density function of the Weibull distribution, with parameters λand k. The derivative of the probability density function is then given by w(x;λ,k) = dW(x;λ,k) dx (3.5.1.7) =k λ"(k−1)xk−2 λk−1e−(x λ)k−x λk−1ke−(x λ)k(x λ)k x#(3.5.1.8) =k λe−(x λ)k(k−1)xk−2 λk−1−xk−2 λk−1kx λk(3.5.1.9) =k λe−(x λ)kxk−2 λk−1(k−1) −kx λk(3.5.1.10) =kxk−2 λke−(x λ)kk1−x λk−1(3.5.1.11) (3.5.1.12) The shape (λ) and scale (k) parameters for the Weibull distribution function are estimated by statistical fitting, as shown in Section 5.2 on page 39. Values are also jittered, and three variables are created: two variables to perform dilatation in the x and y-axis of the function and a third variable to define a base level for the will to emigrate. Defining: AShape kas the shape parameter; AScale kas the scale parameter; Ax-axis kas the x-axis dilatation parameter; Ay-axis kas the y-axis dilatation parameter; ABase kas the base level parameter; as well as the following random variables: XShape kfor the shape parameter such that XShape k∼N(0, 0.1 ×AShape k); XScale kfor the scale parameter such that XScale k∼N(0, 0.1 ×AScale k); Xx-axis kfor the x-axis dilatation parameter such that Xx-axis k∼N(0, 0.1 ×Ax-axis k); Xy-axis kfor the y-axis dilatation parameter such that Xy-axis k∼N(0, 0.1 ×Ay-axis k); XBase kfor the base level parameter such that XBase k∼N(0, 0.1 ×ABase k). Then the main parameters are recomputed as AShape k=AShape k+xShape k,xShape k∼XShape k; (3.5.1.13) AScale k=AScale k+xScale k,xScale k∼XScale k; (3.5.1.14) Ax-axis k=Ax-axis k+xShape k,xx-axis k∼Xx-axis k; (3.5.1.15) Ay-axis k=Ay-axis k+xScale k,xy-axis k∼Xy-axis k; (3.5.1.16) ABase k=ABase k+xBase k,xBase k∼XBase k; (3.5.1.17) (3.5.1.18) 26 FCUP An Agent-Based approach of the Portuguese population projection For these parameters, a standard deviation of 10% of the base value is used because they are very sensitive to small changes in comparison with the weight parameters. Also defining ASuccess k,l as a work success parameter, with ASuccess k,y =         −1 , if AEmp a,s,k,y−1= 1 0 , if AAct a,s,k,y−1= 0 1 , if AEmp a,s,k,y−1= 0 (3.5.1.19) ASuccessW kas the weight given to work success ASuccess k,y . Like the previous weights, this will suffer a 75% random variation for each agent so, let XSuccessW kbe a random variable for the work success weight such that XSuccessW k∼N(0, 0.75 ×ASuccessW k). (3.5.1.20) Then the work success weight is recomputed as ASuccessW k=ASuccessW k+xSuccessW k,xSuccessW k∼XSuccessW k. (3.5.1.21) Finally the will to emigrate can be computed as AWill a+1,k,y+1 =Ay-axis k×w(Ax-axis k×i;AShape k,AScale k) + ASuccessW k×ASuccess k,y (3.5.1.22) and AWill a,k,y0=Ay-axis k×W(Ax-axis k×i;AShape k,AScale k) + ABase k. (3.5.1.23) Then, for each country, a gain is computed for either emigrating or staying in the home country. The emigration gain is given by AGain c,a,k,y =AHealthW k×GHealth c+ASafetyW k×GSafety c+AWageW k,e ×GWage c,e + (3.5.1.24) +APopW k×GPop c+ADistW k×GDist c+ALangW k×GLang c×AWill a,k,y (3.5.1.25) while the gain to stay in Portugal is AGain c,a,k,y =AHealthW k×GHealth c+ASafetyW k×GSafety c+AWageW k,e ×GWage c,e + (3.5.1.26) +APopW k×GPop c+ALangW k×GLang c×1−AWill a,k,y(3.5.1.27) The GWage c,y,e parameter is updated assuming that the net migration on the foreign countries remains constant as well as the number of jobs. The following variables are defined for each destination country of the Portuguese emigration: GForPop c,y as the foreign population size; GForEmpR c,y as the foreign population employment rate; GForEmpR c,y,e as the foreign population employment rate indexed by education; GForVarR c,y as the foreign population employment rate variation; FCUP 33 An Agent-Based approach of the Portuguese population projection GDir a,s,y = GMaxAge−a X t=0 GStatProp a+t,j,y GStatProp a,s,y (1 + GGDP)−1(4.2.2.2) GInd a,s,y = GMaxAge−a X t=0 GStatProp a+t,s,y GStatProp a,s,y 1−GStatProp a+t+1,s,y GStatProp a+t,s,y !(1 + GGDP)−(t+1)GWidFrac(1 + GWidInc)(t+1), (4.2.2.3) where GMaxAge is the maximum allowed age in the model; GStatProp a,j,y is the stationary population proportion as explained in Section 5.3.3 on page 52; GWidFrac is the percentage of the pension which is payed to the widow(er); GWidInc is the expected increase in the percentage of the pension which is payed to the widow(er) each year. Then APension k,y is computed by APension k,y = ARetAge k X a=0 ARem k,a,e,j(1 + GGDP)(ARetAge k−a)GδAge ARetAge k,y. (4.2.2.4) 4.2.3 Portuguese Model In addition to the previous theoretical social security models, the model used by the Portuguese Social Security will also be considered. This model is a mixture of the two previous models. In the early years, the Portuguese model followed an earnings-based model, which was then reformed into a contributions-based model, due to sustainability issues. So the current model has functionalities from both, in order not to penalize the Portuguese population which started working on the previous regime. The following variables are defined: ALastRem k: the vector of the last 40 work remunerations, corrected to present values.This vector might be smaller than 40 if the agent k has worked for less than 40 years; ALastRem k,t:ALastRem kelements; ASizeRem k: the size of ALastRem k; ARefRem k: the reference remuneration used by the Portuguese Social Security system to compute the pensions value, ARefRem k=P t ALastRem k,t 14ASizeRem k (4.2.3.1) GIAS : the reference value for the computation of aids and other expenses, and revenues of the Portuguese State administration [06]; AInvLastRem k: the sub-vector of ALastRem kcontaining its last 15 elements at most. AInvLastRem kis also ordered from the greatest to the smallest value; AInvLastRem k,t:AInvLastRem kelements. If ASizeRem k≤10 then AP1 k=ARefRem k. (4.2.3.2) 34 FCUP An Agent-Based approach of the Portuguese population projection else AP1 k= 10 P t=1 AInvLastRem k,t 14 ∗10 . (4.2.3.3) If ASizeRem k≤20 then AP2 k=ARefRem k×0.02 ×ASizeRem k(4.2.3.4) else AP2 k=                                              0.023ARefRem kASizeRem k, if ARefRem k≤1.1GIAS [0.023 ×1.1GIAS +0.0225(ARefRem k−1.1GIAS)]ASizeRem k, if 1.1GIAS <ARefRem k≤2GIAS [(0.023 ×1.1 + 0.0225 ×2)GIAS +0.022(ARefRem k−2GIAS)]ASizeRem k, if 2GIAS <ARefRem k≤4GIAS [(0.023 ×1.1 + 0.0225 ×2 + 0.022 ×4)GIAS +0.021(ARefRem k−4GIAS)]ASizeRem k, if 4GIAS <ARefRem k≤8GIAS [(0.023 ×1.1 + 0.0225 ×2 + 0.022 ×4 + 0.021 ×8)GIAS +0.02(ARefRem k−8GIAS)]ASizeRem k, if 8GIAS <ARefRem k (4.2.3.5) The model also defines AC1 kas the working years of agent k until year 2006; AC2 kas the working years of agent k from year 2007 until present; AC3 kas the working years of agent k until year 2001; AC4 kas the working years of agent k from year 2002 until present. Finally, the pension value is set as APension k,y =         AP2 k, if AFirstContr >2002 AP1 kAC1 k+AP2 kAC2 k AWYears k , if AFirstContr ≤2002 and ARetAge k≤2016 AP1 kAC3 k+AP2 kAC4 k AWYears k , if AFirstContr ≤2002 and ARetAge k>2016 (4.2.3.6) 4.3 Algorithm Most parameters related to the social security are updated every year, according to their values in the previously simulated year and the 2011 Portuguese data from Statistics Portugal and GEE databases [Sta15c],[GEE12]. Step 1. Update the activity of the agents in the current year yc. If GActR a,j,yc−GActP a,s >Gthen u is randomly sampled from U(0, 1) for each agent aged a years-old, with sex s and with AAct a,s,k,y = 1; if u<G, then, the agent k is set as inactive AAct a,s,k,y = 0 in order to reduce the excess of active population. If GActR a,s,yc−GActP a,s <−Gthen uis randomly sampled from U(0, 1) for each agent aged a years-old, with sex s and with AAct a,s,k,y = 0; if u<G, then then agent k is set as active AAct a,s,k,y = 1 in order to fix the lack of active population; FCUP 35 An Agent-Based approach of the Portuguese population projection Step 2. Update the employment status of the agents; Step 3. Update the retirement status of the agents with ARet a,s,k,y = 0. If GMinRet ≤a≤GMinRet + 10 and GAct a,s,k,y = 0 then uis randomly sample from U(0, 1) and if u<GRetR a,s,y , then the agent k is set as retired ARet a,s,k,y, = 1. When an agent gets retired, its retirement pension is defined and remains the same until the agent’s death. Step 4. Update the agents education and job qualification. When an agent gets employed for the first time these parameters are defined. They are decided randomly applying the Inversion Method (Devroye [Dev86]) to randomly sampled u1and u2from U(0, 1) and using the rates of the education and job qualification for the age of the agent. On the following years the parameters are updated using the same mechanism but with the constraint that they have to be increasing with time. Step 5. Update the wages. The wage of an employed agent is defined as AWage a,s,k,y,e,j =GWage s,e,j × 1 + GGDPy-y0; Step 6. Impose that all agents with AEmp a,s,k,y = 1 pay their social security contributions; Step 7. Pay the retirement pensions to all agents satisfying AAlive a,s,k,y = 1 and ARet a,s,k,y = 1; Step 8. Update the social security’s capital GSS. FCUP 37 An Agent-Based approach of the Portuguese population projection Chapter 5 Model Initialization, Data and Results The model presented sometimes requires information from before the start of the simulation. Here, the mechanism that was used to setup the simulation will also be shown. Furthermore, a broad spectrum of variables is needed for the initialization of the previous model. The data used was obtained from several sources, national and international. Often, the data that was found was not exactly the one that was needed, so some arrangements were made in order to achieve the desired information. Preferably, the data source for a determined variable should always be the same, but unfortunately that was not always the case. The source data and the needed modifications will be presented in this chapter. At the end, some results of the program will be presented, focusing in a set of sub-models that portrait the outcome of the other sub-models. These results showcase the Portuguese population structure and its social security system. Figure 5.1 on the next page systematizes and summarizes the previous model application and study, showcasing the most important aspects. The diagram sequence goes from the top to the bottom and starts from the end of the diagram in Chapter 1 on page 1. 38 FCUP An Agent-Based approach of the Portuguese population projection Initialization Data Simulation Clustering Determination Hierarchical Clustering K-Means Identification Supervised Decision Tree Association Rules Results Population Growth Stable Population Validation Overview Social Security Sustainability Sensitivity Analysis (initial capital in thousand million) 5.62.8 8.40 11.2 42 28 56 14 70 Initialization Figure 5.1: Diagram of how results will be presented. 5.1 Model Initialization The model that is closed to migration does not require information prior to the starting year, but all other models do. Still, to initialize the migration model, the information needed from before the simulation starting year is only used to fit a few distributions, and so the supplied data concern only the distribution parameters. However, the model for the social security system requires previous data to be inserted into the model. The computation of the retirement pension is highly dependent on the remuneration and work situation of an agent through all agent’s working career. The data from the past Portuguese activity rates PActR y, employment rates PEmpR y, mean wages Pwage yand GDP growth PGDP ysince 1950 were obtained from Statistics Portugal and Banco de Portugal [Sta15c; Ban15]. At the start of the simulation, the following procedure was performed for each living agent k: Step 1. Find the year ybfor which the age of agent k was 15 years-old (minimum age for employment) Step 2. For each year y, with yb<y<y0, generate random numbers u1,u2from U(0, 1). If FCUP 39 An Agent-Based approach of the Portuguese population projection u1<PActR y∧u2<PEmpR y, then AEmp a,s,k,y = 1 and AWage a,s,k,y,e,j =Pwage y. In section 4.2 on page 31, the social security capital was assumed to have an annual increase that is directly proportional to the increase of the GDP. The same will be done for previous contributions but, in the initialization, no pension will be paid. In addition, and for simplicity, each agent’s contribution AContr k,y will be already increased in accordance to the GDP increase: AContr k,y =ARem k,y,e,j ×GContr (5.1.0.1) and GTContr y=X k AContr k,y × y0 Y j=y PGDP j , y <y0. (5.1.0.2) Therefore, in the beginning of the simulation GSS y0=X y<y0 GTContr y. (5.1.0.3) There is one more important information relating to the model initialization. This one is not related to the necessary past data, but it is related to the population size. For the computational simulation to work it was necessary to use a 2% sample of the Portuguese population. A higher sample would require many computer resources that were not available (processor speed and RAM) and a lower sample would make some age groups to be too small or inexistent. 5.2 Data GAlive a,s,y02011 Portuguese Census [Sta15c] GBirths a,s,y0Statistics Portugal [Sta15c] GDeaths a,s,y0Statistics Portugal [Sta15c] GFertEvo Resident Population Projections 2012-2060 [Sta15c] GMortEvo Resident Population Projections 2012-2060 [Sta15c] Table 5.1: Closed Population Model data source. GFertEvo and GMortEvo are growth rates, but the existing data is about the expected fertility and mortality. Therefore, in order to obtain the growth rate for each year, the value of the following year from the data was divided by the value of the current year. The chosen health GHealth cand safety GSafety cindicators were the percentage of the public expenses spent on health and safety for each country. For GHealth cthe needed data that was available, unlike GSafety cwhich was available only as the nominal expenditure in safety and total GDP. The population indicator GPop cwas set equal to the percentage of Portuguese born residents in country c over the total resident population of the same country. The distance indicator GDist cthat was used was the distance between each country’s capital to the Portuguese capital (Lisbon). The available data were the coordinates of each city according to its latitude and longitude. So defining lat1,lon1as the coordinates in degrees of the first capital, lat2,lon2 as the coordinates in degrees of the second capital and setting R= 6371 as the radius of the planet 40 FCUP An Agent-Based approach of the Portuguese population projection GHealth cThe World Factbook [CIA13] GSafety c The World Factbook [CIA13] United Nations Statistics Division [UN15] Rea [Rea09] (New Zealand source) Statistics Austria [Sta15a] (Austria source) Comptroller General of the Union [Com15] (Brazil source) Statistics Canada [Sta15b] (Canada source) National Institute of Statistics, Geography and Informatics [NI15] (Mexico source) GPop cUnited Nations Statistics Division [UN15] GDist cCIA [CIA13] GLimit c OECD Statistics [OEC15] Emigration Observatory [Emi15] (through apparent limits on data) W(x;λ,k) Eurostat [Eur15] GForPop c,y0 OECD Statistics [OEC15] United Nations Statistics Division [UN15] Japan Institute for Labour Policy and Training [Jap15] (Japan source) France’s National Institute for Statistics and Economic Studies [Fra15] (France source) GForEmpR c,y0,s OECD Statistics [OEC15] Japan Institute for Labour Policy and Training [Jap15] (Japan source) GNetMig cThe World Factbook [CIA13] GMWage cOECD Statistics [OEC15] GPPP cWorld Bank [Wor15] GImmi c,y OECD Statistics [OEC15] XImmiAge Immigration Annual Estimations [Sta15c] GImmiProp OECD Statistics [OEC15] Table 5.2: Migration Model data source. Earth in kilometers, then dist =1−cos ((lat2−lat1)×π 180 ) 2(5.2.0.4) + cos(lat1×π 180)×cos(lat2×π 180)×(1 −cos (lon2−lon1)×π 180  2(5.2.0.5) dist =R×2×arcsin(√dist) (5.2.0.6) where dist is the distance in kilometers. The parameters GHealth c,GSafety c,GPop c,Gwage c,y,e ,GDist cwere forced to be between 0 and 1, dividing them by their corresponding maximum, ranging between all values of c. FCUP 41 An Agent-Based approach of the Portuguese population projection The parameter W(x;λ,k) is an age distribution of Portuguese emigrants. The data containing the age of the Portuguese emigrants was fitted to a Weibull and Gamma distributions as they were the distributions that had the closest shape to the data density function. Figure 5.2 on the next page shows the data density function and the fitted distributions. The Weibull distribution was chosen as it presented the best fit. GNetMig c,y0,e values are on a 1 : 1000 ratio. Therefore, the data was scaled using the population size of each country. XImmiAge is an age distribution of immigrants in Portugal. The data with the age of the Portuguese immigrants was fitted to a Weibull, Gamma and Log-Normal distributions as they were the distributions that had the closest shape to the data density function. Figure 5.3 on the next page shows the data density function and the fitted distributions. The Weibull distribution was again chosen as it exhibited the closest fit, even though it is indistinguishable from the gamma distribution in the figure. GAct a,s Active Population (2011 Series) [Sta15c] GEmpProp a,s Active Population (2011 Series) [Sta15c] GRetProp a,s Statistics Portugal [Sta15c] GEdu a,s Active Population (2011 Series) [Sta15c] GJQual a,s 2011 Staffing Tables [GEE12] GWage s,e,j 2011 Staffing Tables [GEE12] GGDP Trading Economics [Tra15] (projections) GWTenure a,s OECD Statistics [OEC15] GIAS European Social Fund [IGF15] Table 5.3: Social Security Model data source. The data found for GWTenure a,s was not as desired. The data that was needed was the mean work tenure by age of the worker at the beginning of the work contract, but the data that was found was for the mean work tenure by age of the worker at the end of the work contract. This problem was solved by shifting the data backwards considering the mean work tenure. For example, if the data showed that there were 200 people aged 60 years-old and their mean work tenure was 15 years, then the 200 value would be moved to people aged 60 −15 = 45 years-old having a mean work tenure of 15 years. If the work tenure was expressed with more precision than years, then the value was rounded to years. This was applied for all ages. Inspired by Addison and Portugal [AP87], the work tenure was modeled using a statistical distribution. Due to unique characteristics of the data, namely a heavy right tail with a bump, the only distribution that was found to be capable of reproducing such behavior was the Log-Cauchy distribution. The graphic in Figure 5.4 on page 43, shows the closeness of the Cauchy distribution to the logarithm of the data, which is equivalent to show that the Log-Cauchy distribution is close to the data distribution itself. 42 FCUP An Agent-Based approach of the Portuguese population projection 0.00000 0.00025 0.00050 0.00075 0.00100 0 24 49 74 99 Age Probability Density Gamma Weibull Emigration Fitting Figure 5.2: Portuguese emigrants fitted to Weibull and Gamma distributions. 0.000 0.005 0.010 0.015 0 24 49 74 99 Age Probability Density Gamma Log−Normal Weibull Immigration Fitting Figure 5.3: Immigrants to Portugal fitted to Weibull, Gamma and Log-Normal distributions. FCUP 49 An Agent-Based approach of the Portuguese population projection be tested for the population open to migration. Nevertheless, model validation techniques can be applied to the population closed to migration with a sub-model consisting only of the models for the mic fertility and mic mortality. A sliding window verification is performed using 100 simulations for the mic fertility and mic mortality sub-model, initiated with the data from the years of 1981, 1991 and 2001, until the years of 1991, 2001 and 2011. For each of these years there was a population census in Portugal and the data are therefore available. Comparing the output from these simulations with the correct data from 1991, 2001 and 2011 census data, the error of the simulations can be computed as (simulated −real). These errors are stored in tables such as Table 5.6. Start-off instants 1981 1991 2001 Gaps 10 years Error in 1991 Error in 2001 Error in 2011 20 years Error in 2001 Error in 2011 30 years Error in 2011 Table 5.6: Model Validation. Error tables are made for the Total Population, Total Population by age, Total Population by gender and Total Population by gender and by age. With these, an overview of the simulation error can be presented. Figures 5.10 on the next page, 5.11 on page 51 and 5.12 on page 51 show the error from the three starting years 1981, 1991 and 2001 respectively, alongside with their comparison with the years of 1991, 2001 and 2011 when possible. In all figures, the most dramatic errors are at the lowest and highest ages. In Figure 5.12 on page 51 additional errors are found at every ages, showing that the simulation under-predicted what would happen in the next 10 years. It should be noticed, however, that according to the database from the Statistics Portugal [Sta15c], the immigration in that time frame was very high, as well as the net migration. Start-off instants 1981 1991 2001 Gaps 10 years 4734 −3877 −16491 20 years 9462 −8225 30 years 17979 Table 5.7: Total Population Error. 50 FCUP An Agent-Based approach of the Portuguese population projection −400 0 400 800 1200 1991 Female population Male population −400 0 400 800 1200 2001 −400 0 400 800 1200 0 25 50 75 100 age 2011 0 25 50 75 100 age Figure 5.10: Error means bar plot starting at 1981. FCUP 51 An Agent-Based approach of the Portuguese population projection −200 −100 0 100 200 300 2001 Female population Male population −200 −100 0 100 200 300 0 25 50 75 100 age 2011 0 25 50 75 100 age Figure 5.11: Error means bar plot starting at 1991. −300 −200 −100 0 100 200 0 25 50 75 100 age 2011 Female population 0 25 50 75 100 age Male population Figure 5.12: Error means bar plot starting at 2001. 52 FCUP An Agent-Based approach of the Portuguese population projection Start-off instants 1981 1991 2001 Gaps 10 years 2996 −849 −8589 20 years 6201 −3177 30 years 10448 Table 5.8: Total Female Population Error. Start-off instants 1981 1991 2001 Gaps 10 years 1738 −2029 −7902 20 years 3261 −5048 30 years 7531 Table 5.9: Total Male Population Error. Tables 5.7 on page 49, 5.8 and 5.9 show total population, total female population and total male population errors organized as in table 5.6 on page 49. Not surprisingly, the error increases alongside the considered gap size, i. e., the longer the simulation lasts, the higher the errors will be produced. From year to year, the error is not independent; the error from the previous year influences the present year error. Therefore, there is an error propagation, so it is not interesting to increase the simulation time indefinitely. Also, except for the starting year of 1981, the total female population error is lower than the male counterpart. 5.3.3 Stationary and Stable Population The stable population model, as presented in Section 2.2.5 on page 12, is used to study populations with constant age-specific fertility, mortality and net migration. Portugal, as a developed country, should tend towards a structure similar to those of other developed countries, because the population distribution of most developed countries is very close to their stable population equivalent (Bras [Bra08]). 1. Mortality-wise, there are variations between population distributions, but they are small; 2. Fertility-wise, studies say global fertility is expected to remain constant, but that is not necessarily the case for each age-specific fertility; 3. Migration-wise, the age-specific net migration is not yet null. But even so, the stable population equivalent is still valid, because an approximation for the stable population equivalent might translate into a tendency towards the satisfaction of the three starting conditions for the stable population model. In the following pages, the algorithm for the stable population will be presented. Even though stable population can be achieved immediately by complex analytic means, it is simpler if the stationary population is first computed and only then the stable population. The stationary population is a specific case of the stable population model, for whose intrinsic growth rate (r) equals to 0. FCUP 53 An Agent-Based approach of the Portuguese population projection Usually this computation is not made for each age, but it is made for 5 year age groups (Preston, Heuveline, and Guillot [PHG01]). The data required is only the female population, female deaths, and births for each age group. The data on male population and deaths will only be used to show the outcome for the male population. Defining: lfive i,sas the number of persons in the age group i∈ {1, ... , 20}with sex s; dfive i,sas the number of deaths in the age group i∈ {1, ... , 20}with sex s; srfive i,sas the survival rate in the age group i∈ {1, ... , 20}with sex s, given by: sfive i,s= 1 −dfive i,s dfive i,s+lfive i,s ; (5.3.3.1) bfive ias the number of births in the age group i∈ {1, ... , 20}. Defining pstati i,sand rstati i,sas the proportions of the stationary population and real stationary population respectively, then: pstati i,s= 5 × i Y x=1 srfive i,s(5.3.3.2) rstati i,s=pstati i,s×X i lfive i,s(5.3.3.3) For the stable population, the following computations are also required: r0=Dpstati i,F,bfive iE; (5.3.3.4) r1=Dpstati i,F×(2 + 5(i−1)), bfive iE, where (2 + 5(i−1)) is the mean age of each group; (5.3.3.5) r2=Dpstati i,F×(2 + 5(i−1))2,bfive iE; (5.3.3.6) and m1=r1 r0 ; (5.3.3.7) m2=r2 r0 ; (5.3.3.8) k2=m2−m2 1; (5.3.3.9) r= m1−qm2 1−2×k2×ln(r0) k2 (5.3.3.10) The intrinsic growth rate ris already defined, so it will now be defined the proportions of the stable population and real stable population, pstable i,sand rstable i,srespectively, as pstable i,s=e(2+5(i−1))×r×pstatic i,s; (5.3.3.11) rstable i,s=e(2+5(i−1))×r×rstatic i,s. (5.3.3.12) To determine if the Portuguese population is getting closer to its stable population equivalent, the difference between the simulated population and the stable population equivalent in age groups for 54 FCUP An Agent-Based approach of the Portuguese population projection each year is computed, and checked if the gap is getting bigger or smaller. −400000 −380000 −360000 −340000 −320000 2010 2020 2030 2040 Sub−Model 1 −400000 −350000 −300000 −250000 −200000 2010 2020 2030 2040 Sub−Model 2 −400000 −350000 −300000 −250000 2010 2020 2030 2040 Sub−Model 3 −400000 −360000 −320000 −280000 2010 2020 2030 2040 Sub−Model 4 −400000 −375000 −350000 −325000 −300000 2010 2020 2030 2040 Sub−Model 5 −400000 −375000 −350000 −325000 −300000 2010 2020 2030 2040 Sub−Model 6 −400000 −375000 −350000 −325000 −300000 2010 2020 2030 2040 Sub−Model 7 −400000 −360000 −320000 −280000 2010 2020 2030 2040 Sub−Model 8 −400000 −375000 −350000 −325000 −300000 −275000 2010 2020 2030 2040 Sub−Model 9 −400000 −360000 −320000 2010 2020 2030 2040 Sub−Model 10 −400000 −375000 −350000 −325000 −300000 2010 2020 2030 2040 Sub−Model 11 Female Population Male Population Figure 5.13: Distance between the 11 sub-models and their Stable Population Equivalent. The evolution of the simulated population for the 11 sub-models, as previously considered, in relation to their stable population equivalent can be observed in Figure 5.13. The difference between the simulated population and its stable population equivalent is given by (simulated −stable). As can be observed, the distance is shortening every year, which may lead to conclude that we are approximating the stable population equivalent. However, without guarantees that, in the future, the Stable Population model’s starting conditions will be met, there is no assurance that we will converge to it. 5.3.4 Population Evolution Overview Although it is very important to know how the population will be in 2041, the knowledge over the intermediary years may be also very informative. In this section, a yearly representation of the Portuguese population from 2011 until 2041 will be presented. The most important reasons for the 11 possible outcomes will be shown and explained. FCUP 55 An Agent-Based approach of the Portuguese population projection 50000 100000 150000 200000 2010 2020 2030 2040 Sub−Model 1 50000 100000 150000 200000 2010 2020 2030 2040 Sub−Model 2 50000 100000 150000 200000 2010 2020 2030 2040 Sub−Model 3 50000 100000 150000 200000 2010 2020 2030 2040 Sub−Model 4 50000 100000 150000 200000 2010 2020 2030 2040 Sub−Model 5 50000 100000 150000 200000 2010 2020 2030 2040 Sub−Model 6 50000 100000 150000 200000 2010 2020 2030 2040 Sub−Model 7 50000 100000 150000 200000 2010 2020 2030 2040 Sub−Model 8 50000 100000 150000 200000 2010 2020 2030 2040 Sub−Model 9 50000 100000 150000 200000 2010 2020 2030 2040 Sub−Model 10 50000 100000 150000 200000 2010 2020 2030 2040 Sub−Model 11 Female Population Male Population Total Population Figure 5.14: Total Population projection for the 11 sub-models. 56 FCUP An Agent-Based approach of the Portuguese population projection 0 10000 20000 30000 40000 50000 2010 2020 2030 2040 Sub−Model 1 0 10000 20000 30000 40000 50000 2010 2020 2030 2040 Sub−Model 2 0 10000 20000 30000 40000 50000 2010 2020 2030 2040 Sub−Model 3 0 10000 20000 30000 40000 50000 2010 2020 2030 2040 Sub−Model 4 0 10000 20000 30000 40000 50000 2010 2020 2030 2040 Sub−Model 5 0 10000 20000 30000 40000 50000 2010 2020 2030 2040 Sub−Model 6 0 10000 20000 30000 40000 50000 2010 2020 2030 2040 Sub−Model 7 0 10000 20000 30000 40000 50000 2010 2020 2030 2040 Sub−Model 8 0 10000 20000 30000 40000 50000 2010 2020 2030 2040 Sub−Model 9 0 10000 20000 30000 40000 50000 2010 2020 2030 2040 Sub−Model 10 0 10000 20000 30000 40000 50000 2010 2020 2030 2040 Sub−Model 11 Female Population under 15 years Male Population under 15 years Total Population under 15 years Figure 5.15: Total Population under 15 years projection for the 11 sub-models. FCUP 57 An Agent-Based approach of the Portuguese population projection 0 50000 100000 150000 2010 2020 2030 2040 Sub−Model 1 0 50000 100000 150000 2010 2020 2030 2040 Sub−Model 2 0 50000 100000 150000 2010 2020 2030 2040 Sub−Model 3 0 50000 100000 150000 2010 2020 2030 2040 Sub−Model 4 0 50000 100000 150000 2010 2020 2030 2040 Sub−Model 5 0 50000 100000 150000 2010 2020 2030 2040 Sub−Model 6 0 50000 100000 150000 2010 2020 2030 2040 Sub−Model 7 0 50000 100000 150000 2010 2020 2030 2040 Sub−Model 8 0 50000 100000 150000 2010 2020 2030 2040 Sub−Model 9 0 50000 100000 150000 2010 2020 2030 2040 Sub−Model 10 0 50000 100000 150000 2010 2020 2030 2040 Sub−Model 11 Female Population between 15 and 65 years Male Population between 15 and 65 years Total Population between 15 and 65 years Figure 5.16: Total Population between 15 and 65 years projection for the 11 sub-models. 58 FCUP An Agent-Based approach of the Portuguese population projection 0 20000 40000 2010 2020 2030 2040 Sub−Model 1 0 20000 40000 2010 2020 2030 2040 Sub−Model 2 0 20000 40000 2010 2020 2030 2040 Sub−Model 3 0 20000 40000 2010 2020 2030 2040 Sub−Model 4 0 20000 40000 2010 2020 2030 2040 Sub−Model 5 0 20000 40000 2010 2020 2030 2040 Sub−Model 6 0 20000 40000 2010 2020 2030 2040 Sub−Model 7 0 20000 40000 2010 2020 2030 2040 Sub−Model 8 0 20000 40000 2010 2020 2030 2040 Sub−Model 9 0 20000 40000 2010 2020 2030 2040 Sub−Model 10 0 20000 40000 2010 2020 2030 2040 Sub−Model 11 Female Population over 65 years Male Population over 65 years Total Population over 65 years Figure 5.17: Total Population over 65 years projection for the 11 sub-models. Figures from 5.14 on page 55 to 5.17 show an inevitable population reduction in all sub-models and in all age sub-groups except for the population over 65 years, in which case, after a small decrease, the tendency trend changes and starts to increase. This reveals that the Portuguese population is tending to an aged population structure. For the 8 sub-models with population closed to migration, the population decrease is almost linear, while, for the population open to migration, it shows a quadratic or cubic descent, stabilizing in the last years. From Figure 5.15 on page 56, Sub-Model 2 shows a very drastic reduction of the population under 15 years of age, closing to 0 in 2041, though this is not due to a reduction in fertility, but it is due to a reduction of the number of females in the fertile period, as revealed in Figure 5.16 on the previous page. Recalling Subsection 5.3.1 on page 44, Sub-Model 2 refers to a sub-model with very high emigration, probably because of the simulated economy being in recession. This sub-model is one in which the active-inactive dependency ratio is very elevated. This is the most pessimist scenario present in this work. In opposition, Sub-Model 1 and 3 show very positive developments. Although there is still a reduction in the population with less than 15 years-old and between 15 and 65 years-old, there are FCUP 65 An Agent-Based approach of the Portuguese population projection remaining for the social security, because this variable causes a big variability in the results. In addition, the lower and upper bounds vary, depending on the model — this reveals a correlation between the sustainability of the social security and the inherent population age and employment structure, where the bounds are as low as the active-inactive dependency ratio. Through all submodels, a initial capital of 14 thousand million euros or higher is likely to guarantee the sustainability of the social security, while with a initial capital of 5.6 thousand million euros or lower it is likely to occur a bankruptcy of the same. So, while it is not possible to state if the Portuguese Social Security will be sustainable or not, it can be affirmed that if its capital is lower than 5.6 thousand million euros, it is unlikely to be sustainable. Through all graphics it can also be noticed that using the earnings-based model or the contributionsbase model, the Portuguese Social Security would be in a much safer position. However, it is not feasible to make a drastic change on the social security model. After all, besides having to maintain it sustainable, Portugal’s government has to maintain its population happy and because of that, they try not to do severe changes on the pension system. This is clear from the Portuguese formulas; it can be deduced from the formulas that in earlier years, the Portuguese Social Security was using pension formulas similar to the ones of the earnings-based model, and, in later years, it is making use of formulae based on the contributions-based model. FCUP 67 An Agent-Based approach of the Portuguese population projection Chapter 6 Discussion In this chapter, a discussion over the main results will be presented, focusing on the detection of possible weaknesses of the proposed models and in suggestions on how to improve them. Also, some insights on the proposed models will be offered, showcasing the already known defects, characteristics and advantages. 6.1 Results Population dynamics are a very complex system with high dependability between parameters. Any error committed predicting a year will be propagated to the following one. Finding better modeling approaches that minimize each year error is desired, and, for this, demographers and mathematicians must come together to discuss and improve each others approaches. From the lack of data, all models presented here, except one, cannot be verified correctly for accuracy. Only insights and philosophical explanations can be shown to validate the models. Hopefully, within some years more data will be available, and the validation of a future model will be possible to be made. One of the best methods to verify if a model is correct is by initializing it in past years and check how far from the reality they are when predicting the present. In Section 5.3.2 on page 48 a mechanism for this evaluation is shown for which the data available is enough. As observed already from Table 5.7 on page 49, with the exception of the starting year of 2001, the error for the first 10 years is below 5 thousand persons and for the next 10 years the error is below 10 thousand persons. As for the starting year of 2001, the error for the first 10 years is near 15 thousand, which in comparison with the remainder ones is a very serious error. As previously stated, the period of 2001 until 2011 is marked with a positive net migration and that could produce a big influence on the results. On the remaining years, the greatest error is present on the ages below 10, 20 or 30, depending on the considered gap. This comes from the mic fertility model, which does not consider the fertility tendencies and, as visible on Figure 5.10 on page 50, the error is lower the closer the age is from the gap size. From these results, undoubtedly, a model incorporating the mic-mac fertility model and the migration model would make big changes to the predictions and possibly reduce the error. With regard to all possible population’s evolution that were evaluated, the general overview, which is supported by all the considered sub-models, is that the Portuguese Population will decrease in the years after 2011. This is also stated by Statistics Portugal [Sta15c] in a study called “Resident Population Projections” for the years 2012-2060. Figure 6.1 on page 69 is extracted from the aforementioned study, and shows that the Portuguese population projections will decrease until 2060. 68 FCUP An Agent-Based approach of the Portuguese population projection However, this reduction is lighter than the ones presented in Subsection 5.3.4 on page 54 for the models with migration. This change is due to the migration predictions by Statistics Portugal that can be observed in Figure 6.2 on the following page, because Statistics Portugal predicts that the Portuguese net migration will increase starting on 2012 for either pessimistic and optimistic scenarios. Figure 5.19 on page 60 shows that, from the three sub-models with migration, only the sub-model with prosperous economy has a similar projection to Statistics Portugal’s optimistic scenario. All the other migration projections showcased in this thesis are much more pessimistic than Statistics Portugal. Additionally, evaluating the tendency of the Portuguese net migration before 2011 in Figure 6.2 on the following page, the scenarios produced in this work seem to be the ones that better replicate that tendency. The higher emigration explains why projections from this work are more pessimistic than the ones from Statistics Portugal. As already explained in Subsection 5.3.4 on page 54, the highest emigration volume is in fertile ages, so, along with an higher emigration, there is also a lower count of births in each year. Also, an high emigration in the first years deepens the effects on future years because newborns in the first years become mothers and fathers in the later ones. Furthermore, Statistics Portugal’s projections are in general optimistic on the evolution of fertility (Figure 6.3 on the following page) in which, considering the past fertility evolution trend, the pessimistic scenario is actually a very positive one. Although this predictions can be correct in the long term, closer to 2011 they seem very peculiar, where at least a slight decrease past 2011 should be expected before a possible recovery in subsequent years. In spite of this, the pessimistic hypothesis was used for the mac and mic-mac models for the fertility, as it was the most convincing scenario that was found in published data. Statistics Portugal’s optimistic scenario is the result of 2013’s Fecundity Survey conducted also by Statistics Portugal, where in average, women reported to be expecting 1.8 more children until 2060 besides their current children, which seems highly subjective and does not offer much confidence. In order to improve the sub-models without using previously predicted values for fertility, the following suggestion is presented. Many studies were already conducted for Matching models; in this case, the matching would be between two persons (Kohler [Koh00], Billari and Kohler [BK04], Zinn [Zin12], White and Potter [WP13], Dribe, Hacker, and Scalone [DHS14]). Using social-economic values and past accounts on births from couples of Male-Female, Female-Female and Male-Male parents, a birth decision model could be designed. At the beginning of this project, this was one of the models that was to be created, but later it was found that adding this model was too complicated for the available time, mainly because a decent Matching model would require several months to be developed. As for the Portuguese Social Security sustainability, the presented results show how much capital was needed in 2011 for the Portuguese Social Security to remain sustainable for the next 30 years for each of the considered sub-models. It is important to reinforce that the present values of Portuguese Social Security’s capital are publicly unknown and, therefore, all that can be said is that, if the Portuguese Social Security’s capital is lower than the lower bounds, then the odds are high that the system bankrupts. It is also important to notice that the Portuguese Social Security expenses are not just allocated to the retirement pension, since there are many other expenses for the system, such as unemployment subsidies or study scholarships. In the revenue’s side, the Portuguese Social Security’s capital is not only funded by the working contributions but, if needed, the State may move capital from other funds at its disposal. FCUP 69 An Agent-Based approach of the Portuguese population projection Figure 6.1: Statistics Portugal’s Resident Population Projections from 2012 until 2060. Figure 6.2: Statistics Portugal’s Net Migrations Projections from 2012 until 2060. Figure 6.3: Statistics Portugal’s Children per Woman Projections from 2012 until 2060. 70 FCUP An Agent-Based approach of the Portuguese population projection 6.2 Models Many models for various situations were proposed. As already stated a model is an attempt to translate the reality and, obviously, a model is highly likely to not be 100% accurate. Some of their shortcomings and advantages were presented in the previous section. Now, the characteristics already expected to happen during their design and implementation will be presented. Fertility (Section 3.2 on page 20) and mortality models (Section 3.3 on page 22) are very similar. Therefore, the discussion will not be focused on any of them but in the pair simultaneously. The mic model is a model with high variability for each year, since each year is very dependent on random effects — it is possible that two consecutive years present a very low and a very high fertility/mortality rate for any given age. But for a sufficiently big population per age, on average the sates will remain constant. This statement is due to the following line of thoughts. First, the process of having a child or dying is the same as a Bernoulli Process and for a sufficiently large population, this can also be seen as a binomial distribution. Second, for a sufficiently large number of trials n(or population in this case) and probability p, the binomial distribution B(n,p) average occurrences equals to np and averaging this value through the population gives the yearly rate of p. So, for a sufficiently large population, the rates will remain constant. The mac model, however, has no variability and each simulation run produces the same global rate for each age. The mic-mac model maintains the same variability as the mic model, but on average tends to a previously predicted direction (either increases, stabilizes or decreases). Due to the already explained higher variability on the global rates for ages with small population density, in such cases, the mic-mac model should be forced to work exclusively with the mac model, as it is the case for the mortality model. It is also important to discuss the closed population algorithm (Section 3.1 on page 19). In this algorithm, the most important detail is the order of the events. It was chosen to age the agents first, then to give birth to new agents and only after to do the random “killing”. The choice for the order was mainly to keep the algorithm simple while explaining the observable data. The existing data shows the number of deaths of persons of age 0 (newborns), which in the demographic field is a very important value, since it is used to evaluate the development indicators of a country; it also shows the number of births in the same year and persons alive at the end of the year. So, for simplicity, it was necessary to make the random “killing” to happen only after the births. And the aging comes first so that in the end of the simulation there are agents with age 0 (newborns that survived) and there are no agents of age 100 (which have all died at the “killing” step). The 99 age boundary was defined to control the population growth; if no boundary was defined, there would be a low (but positive) chance of existing agents of age 130 in the end of the simulation. The emigration model attempts to translate the choice of a person to emigrate to a determined country through a gain function dependent on its own preferences (Subsection 3.5.1 on page 23). The preferences of a person are very diverse and it is hard to model them. One of the difficulties lays on the amount and type of the country’s aspects that influences a person. The chosen aspects were objective ones, which could be ordered by amount or distance, but a person decision is not only based on objective characteristics. For example, the climate was not taken into consideration for this model, but it is an important characteristic when choosing a destination country. The climate could not be used because it is a subjective characteristic and the country ordering through climate would be different for each person. Also, many objective characteristics were not used to maintain some simplicity. One other problem is the modeling of the person’s preferences. Even if all possible FCUP 71 An Agent-Based approach of the Portuguese population projection countries’ characteristics were present, a person does not make its choice using all available possibilities. A characteristic could be unimportant to a person while for another it could be one of the most important characteristics. A good way to deal with that would be through a large scale survey to the Portuguese population, to obtain the personal interests of a significant number of persons. For the immigration model (Subsection 3.5.2 on page 27), however, there is little to be said. It is not possible to make a decision model for agents that do not exist in the model, therefore the model has to be completely exogenous. Other modeling ways could be considered, like using Markov Chains or incomplete/censored data regressions; also more variables could be used, like Portuguese economic scenario and emigration volume. The model used here is not perfect but it is a very simple model and it is easy to compute. For the employment model, there is both the mic model (Subsection 4.1.1 on page 30) and the mic-mac model (Subsection 4.1.2 on page 30). The mic model will not be discussed further, because the behavior is the same of the mic models for fertility and mortality. For the mic-mac model, some considerations must be made. This model makes the age-specific employment rate change; this is due to the fact that this model, instead of using the employment rates, uses the work tenure distributions. This makes the model much more dynamic than simpler models. This model can also be used together with legislative rules for work contract periods instead of just using work tenure. However, this model is highly dependent on the chosen distributions and a proper study of the data is very important before using it. A clear disadvantage of using this model is the necessary data. Ideally, the data should be age-specific, but that is hard to find; furthermore, the data must also be ordered on the date of contract. Therefore, this model is more delicate and its usage should be done with some care, so the produced results are the ones intended. Another thing worth discussing, besides the models themselves, is the programming language that was used. Initially this model was programmed in NetLogo (Wilensky [Wil99]). As it is a prototyping programming language, the model development was quite fast and easy, but as the model became more and more complex, problems started to appear. The first problem was the memory usage — with the mic models alone, along with the Earnings-based model, the model, in NetLogo, required 2GB of available memory to function. This was a problem because NetLogo standard memory usage limit was 1GB, but with a simple manipulation on the batch files of NetLogo it was possible to increase this limit. With the same basic sub-model, after solving the memory issue, another problem arose. This time was the computing time — a single simulation run required about 1 hour and 30 minutes to be completed. This problem was then solved with the help of the Faculty of Engineering of the University of Porto’s Grid, which allowed about 100 simulations to be run at the same time. This was however the most problematic issue. At the time, more sub-models were already planned to be modeled and running several simulations with them would become a major problem. To solve this problem, the whole model was reprogrammed in Java [AG98]. In this new programming language, the model’s memory requirement fell bellow the 1GB limit and a simulation run time reduce from the 1 hour and 30 minutes to a stunning 30 seconds; and with further improvements to the code, the computing time was reduced even more, reaching the minimum time of 5 seconds per simulation run. Now, with the more complex sub-models, the memory usage remains a problem, requiring 4GB of memory, but the run time is no longer a problem since each simulation has been taking at most 1 minute. FCUP 73 An Agent-Based approach of the Portuguese population projection Chapter 7 Conclusion Population growth has been studied for centuries; many models and many approaches have been tried ever since. In the beginning, linear models were used but they were unable to capture the entire reality. Then, models became more and more complex, using more variables, which have brought a large improvement on population growth predictions. Nowadays, besides modeling population growth, the growth for each age and gender is also modeled, as is fertility and mortality rates and also migrations. Without a doubt, models are now able to capture more parts of the reality. However, as stated by Coale and Trussell [CT96], although more parts are explained and models became more sophisticated, Human population growth models have not improved significantly so far. However, Human population growth is a very uncertain phenomenon, and is very different from general population models. Human population is dependent on their members and the actions of those members are almost random to an outside observer. Of course people’s choices are not made by chance, however because local knowledge is distributed, the reasoning behind each person’s decisions is not fully known to the statistician. Here, this work presented a model for the modeling of a person’s actions. Far from being perfect, it is, nevertheless, a starting point. A core algorithm was created for people’s actions, accepting multiple models for each life components. For now, the fertility part is not based on the decisions of a person but is more of a random effect with some control. It has three different models to choose from, ranging from a random model to a fully controlled model. The mortality part is also not modeled as a personal decision, but neither should it fully be, as the time of death of a person is partly decision-based (how healthy a person lives, for example) but is also random (a healthy person may have a car accident and die, for example). This component also has three different models to choose from, ranging from a random model to a fully controlled model. The migration model, however, is based on a person’s decision capacity. There are many approaches on how to model the decision of a person; here, a gain maximization algorithm was used. This algorithm uses the person’s preferences (for health, safety, ...), its life history (employment status and remuneration) and exogenous factors (age, economic scenario, ...). Independently of the combination of the chosen models, the predictions suggest that there will be a population decrease on the next 30 years, but the slope varies. A sample size of 2% of the population was simulated. The predicted population reduction ranged from 10% of the total population to 50%. The different outcomes were largely influenced by the usage of the migration model and different economic scenarios. If the migration model is not present, the reduction appears to be linear, but adding the migration model, the reduction appears to have a quadratic or cubic behavior. In addition, when the migration model is used, the population growth seems to stabilize towards the 74 FCUP An Agent-Based approach of the Portuguese population projection end of the simulation, which sheds light into the future of the Portuguese population. If nothing more, these results show the importance of using decision agent-based models in detriment to deterministic or purely stochastic models. The Portuguese Social Security sustainability is an important issue for most Portuguese. It is a hot discussion topic in Portugal, largely observed in news reports. Studies from people of many different areas, like mathematics, economy, politics and so on..., were presented to the general public. However, there is a big problem; there are many studies which state that the Portuguese Social Security will bankrupt in the following years, while others assert it will maintain its sustainability. Joining in this hot topic, this work also sheds light on this issue. There are plenty uncertainties related to the Portuguese Social Security sustainability and many facts are not made available to the general public. Therefore, in this work, there is no attempt at making a definitive statement on the topic. All that is supplied are possible scenarios; and with those scenarios some insights are provided. The big question to be answered is if the Portuguese Social Security is currently sustainable. The given answer here is: it depends. This work reveals two great variables which will affect the Portuguese Social Security outcome. One is the age-specific growth of the Portuguese population and its respective work structure; the other is the remaining capital on the Portuguese Social Security funds nowadays. Alongside these two great variables, there are two obvious conclusions to return. First, the higher the emigration, the worse the Portuguese Social Security sustainability will endure, as emigration is mainly focused on the active population. Second, the least the capital remaining nowadays, the higher the chances of bankruptcy. Additionally, two important and more precise insights can be drawn from this work. An upper and lower bound were found for the social security’s capital in 2011. The upper bound is 14 thousand million euros and if the value at 2011 was higher than that, then the sustainability of the social security is a probable event. The lower bound is 5 thousand million euros and if the value at 2011 was lower than that, then it is most likely that the Portuguese Social Security is unsustainable. It is also important to remember that the simulations were run for a sample of 2% of the population and so, the results provided are also for a 2% sample. Therefore, extrapolating the results to the total Portuguese population, the upper and lower bounds would be 700 thousand million euros and 250 thousand million euros respectively.