scieee AI-readable full text Open interactive document viewer

Can teaching be taught? Improving teachers' pedagogical skills at scale in rural Peru

Castro, Juan Francisco,Glewwe, Paul,Heredia-Mayo, Alexandra,Majerowicz, Stephanie,Montero, Ricardo

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Castro, Juan Francisco; Glewwe, Paul; Heredia-Mayo, Alexandra; Majerowicz, Stephanie; Montero, Ricardo Article Can teaching be taught? Improving teachers' pedagogical skills at scale in rural Peru Quantitative Economics Provided in Cooperation with: The Econometric Society Suggested Citation: Castro, Juan Francisco; Glewwe, Paul; Heredia-Mayo, Alexandra; Majerowicz, Stephanie; Montero, Ricardo (2025) : Can teaching be taught? Improving teachers' pedagogical skills at scale in rural Peru, Quantitative Economics, ISSN 1759-7331, The Econometric Society, New Haven, CT, Vol. 16, Iss. 1, pp. 185-233, https://doi.org/10.3982/QE2079 This Version is available at: https://hdl.handle.net/10419/320331 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc/4.0/ Quantitative Economics 16 (2025), 185–233 1759-7331/20250185 Can teaching be taught? Improving teachers’ pedagogical skills at scale in rural Peru Juan F. C astro Department of Economics, Universidad del Pacifico Paul Glewwe Department of Applied Economics, University of Minnesota Alexandra Heredia-Mayo Department of Economics, Universidad del Pacifico Stephanie Majerowicz School of Government, Universidad de los Andes Ricardo Montero Department of Applied Economics, University of Minnesota We evaluate the impact of a large-scale teacher coaching program in Peru, a context with high teacher turnover, on teachers’ pedagogical skills and student learning. Previous studies find that small-scale coaching programs can improve teaching of reading and science in developing countries. However, scaling up can reduce programs’ effectiveness, and teacher turnover can erode compliance and cause spillovers onto non-program schools. We develop a framework that defines different treatment effects when teacher turnover is present, and explains which effects can be estimated. We evaluate this teacher coaching program, exploiting random assignment of that program’s expansion to 3797 rural schools in 2016. After two years, teachers assigned to the program increased their aggregate pedagogical skills by 0.20 standard deviations. The program also increased student Juan F. Castro: [email protected] Paul Glewwe: [email protected] Alexandra Heredia-Mayo: [email protected] Stephanie Majerowicz: [email protected] Ricardo Montero: [email protected] We would like to thank seminar participants at the Department of Applied Economics of the University of Minnesota, the Department of Economics of Universidad del Rosario, the LACEA 2019 Annual Meeting, the Department of Agricultural and Consumer Economics at the University of Illinois, and the Department of Agricultural Economics and Rural Development at Seoul National University for their valuable comments. We also thank several anonymous referees, whose comments were very helpful for improving our paper. We are also grateful to Hugo Fernández for excellent research assistance, and to Diana Horvath for preparing the replication materials. Any remaining errors are ours alone. The randomized evaluation was planned by Peru’s Ministry of Education. The student assessments and teacher observation instrument used in this study were designed by the Ministry of Education for general internal use. We used anonymized data provided by the Ministry of Education. The replication package for this paper is at: https://doi.org/10.5281/zenodo.13738582. ©2025 The Authors. Licensed under the Creative Commons Attribution-NonCommercial License 4.0. Available at http://qeconomics.org.https://doi.org/10.3982/QE2079 186 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) learning; after 1 year, Grade 2 students’ mathematics and reading scores increased by 0.106 and 0.075 standard deviations (of the distributions of those test scores), respectively. After three years, the cumulative effect increases slightly, to 0.114 and 0.100, respectively. One reason why these impacts are low is that some uncoached teachers moved into treated schools in years 2 and 3. Following our framework, we estimate that the impacts on students of having a “fully” coached teacher for all three years are 0.18 and 0.16 standard deviations for mathematics and reading comprehension, respectively. Keywords. Education, teacher coaching, pedagogical skill, student learning, teacher turnover. JEL classification. I21, O15. 1. Introduction Teacher quality is an essential determinant of student learning (Das, Dercon, Habyarimana, and Krishnan (2007), Clotfelter, Ladd, and Vigdor (2010), Chetty, Friedman, and Rockoff (2014)). Yet many teachers lack mastery in the subjects they teach or lack the pedagogical skills to teach them effectively. This is especially true for teachers in developing countries (World Bank (2018)). Can these teachers’ skills be improved? Every year, developing countries spend over $1 billion on teacher training (Loyalka, Popova, Li, Liu, and Shi (2019)). Popova, Evans, and Arancibia (2016) find that about two-thirds of the World Bank educational projects between 2000 and 2012 included inservice teacher training. Such training is attractive because it can be centrally designed and coordinated by the Ministry of Education and is usually supported by teachers’ unions (Evans and Popova (2016)). In this study, we evaluate the impact of a large-scale teacher coaching program, operating in a context of high teacher turnover, on teachers’ pedagogical skills as well as on student learning outcomes. Evidence on the impacts of in-service training in developing countries is mixed, and programs vary widely in form and content. A survey by Evans and Popova (2016) found that programs with face-to-face training, follow-up visits, engagement of teachers to obtain their ideas, and adaptation to local context, tend to have larger effects on student learning. Coaching programs often have these features as they involve school visits, classroom observations, and personalized feedback for teachers by trained peers or coaches. Thus, coaching programs are a promising alternative to traditional in-service training that offers intensive sessions to large numbers of teachers at a centralized venue. When programs are offered at the school level but are intended to operate through teachers, and teachers can move between schools, estimates of the average treatment effect (ATE) of the program based on a randomized control trial may be biased. In particular, movement of teachers across schools may lead to spillovers that will introduce biases when comparing treated and control schools, even when all schools comply with their random assignment and there are no biases due to the selection or attrition of students. Education interventions that operate through teachers often have all teachers in a school share treatment status (i.e., all teachers are either treated or untreated). Most Quantitative Economics 16 (2025) Can teaching be taught? 187 studies of the effectiveness of these types of interventions focus on student outcomes and compare treatment schools with control schools, and some of them evaluate results after enough time has passed for teachers to switch schools (Lucas, McEwan, Ngware, and Oketch (2014), Jukes et al. (2017), Cilliers, Fleisch, Prinsloo, and Taylor (Cilliers et al. (2020))). These studies usually address potential biases due to student attrition, yet they rarely mention the possibility of teacher turnover or the potential bias it may induce. This risk of bias may occur not only for education interventions but also for any estimation of treatment effects in cluster randomized control trials (RCTs) with movement of service providers or program beneficiaries across clusters. Indeed, high turnover is reported for many non-education contexts. For example, Kovner, Brewer, Fatehi, and Jun (2014) report that 17.5% of new nurses in the U.S. leave their jobs within 1 year of starting, and Banerjee, Chattopadhyay, Duflo, Keniston, and Singh (2021) find, in their control sample, that one-third of police officers in India changed stations over an 18-month period. Despite its frequency, turnover is usually ignored in program evaluations. For example, Georgiadis and Pitelis (2016) compare treated and control enterprises (clusters) in a job training program but do not discuss the possibility of workers moving across firms. We make a methodological contribution by developing a framework that clarifies the assumptions and data needed to obtain unbiased estimates of average treatment effects (ATE), intent to treat effects (ITT), and average causal response (ACR, an extension of local average treatment effects (LATE)) in a clustered RCT with movement of service providers across clusters. In our context, this framework explains how treatment effects differ, depending on whether one focuses on a particular set of teachers, following them if they move to other schools (in which case the outcome variables are those teachers’ skills), or on the teachers and students in particular schools (in which case the outcome variables are the skills of these schools’ teachers and the learning progress of these schools’ students). Both sets of treatment effects are highly relevant from a policy perspective. The first set is relevant for policies that focuses on improving the skills of a particular group of teachers, such as teachers whose pedagogical skills are thought to be deficient. The second set is relevant for policies aimed at improving the teaching skills and learning progress, respectively, of the teachers and students in a particular group of schools, such as schools where students’ academic performance is particularly low. We show how the latter set of effects depends not only on the direct effect of the program on participating teachers’ skills but also on the indirect effect of the program on teacher composition: which teachers stay in these schools, which teachers leave these schools, and which teachers move into these schools. Previous research based on cluster RCTs where service providers move across clusters has ignored these composition effects. We show that, in general, it is not possible to estimate average treatment effects (ATEs) for teacher skill and student learning, although under certain conditions lower bounds for ATEs can be estimated. We also show that comparisons of teachers in treated and control schools after turnover has occurred will, in general, lead to biased estimates of intent to treat (ITT) effects for teachers in the program schools when the program started. However, it is possible to estimate these ITT effects if one has a sample of teachers that follows them when they change schools, or using the data of teachers in treated 188 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) and control schools after turnover has occurred if that turnover is unrelated to the program. This last result is important because following teachers who change schools and, more generally, following service providers who change clusters, can be difficult, which raises the risk of attrition bias in ITT estimates. We estimate the effects on teachers’ pedagogical skills and on student learning of a teacher coaching program implemented in rural multigrade schools in Peru. Trained coaches visit classrooms and give specific advice to teachers on their pedagogical practices, providing customized strategies to improve them. Identification exploits random assignment of 6218 schools (3797 treated schools, 2421 control schools) when the program expanded in 2016. Teacher skills were measured in late 2017 (after nearly 2 years of treatment) by observing teacher-student interactions and a broad range of instructional practices in a randomly selected subsample of 166 treated and 174 control schools. Student skills were tested in grades 2 (late 2016) and 4 (late 2018) for all public schools with five or more students in those grades, which provides student test score data for 2567 of the 6218 randomly assigned schools. As in many developing countries, Peru’s rural schools have very high rates of teacher turnover;1of the teachers in the subsample of 340 schools with teacher skills data, about 43% had moved between 2016 and the start of 2017. Importantly, classroom observation data were collected not only in these 340 schools, but also in many (but not all) of the schools that received the teachers who moved from these schools to other schools between 2016 and 2017. Our main findings are as follows. For the teachers who, after turnover occurred (i.e., in 2017), were teaching in the schools assigned to the program, we find that the ITT effect of 2 years of coaching on their pedagogical skills is 0.20 standard deviations (s.d.) of the distribution of those skills. This is also our preferred estimate of the ITT effect on the skills of the teachers in the program schools when the program began, many of whom left those schools in the next year. We also show that this ITT estimate is, under plausible assumptions, a lower bound of the ATEs for both sets of teachers. Turning to specific skills, the largest ITT effects are for lesson planning and, to a lesser extent, encouraging students’ critical thinking. We also estimated treatment effects of the program on student learning after 1 and 3 years (we have no data for the second year). After 1 year, the program increased learning among the Grade 2 students who took the 2016 National Student Evaluation by 0.106 s.d. in mathematics and 0.075 s.d. in reading comprehension (of the distributions of those test scores in the control schools). These are both ITT and ATE effects, since all teachers followed their random assignment in the first year. After 3 years of exposure, the ITT effect increases only slightly, to 0.114 s.d. for mathematics and 0.100 s.d. for reading comprehension; these estimates, which are lower bounds for ATE (which cannot be estimated in year 3), reflect the fact that many teachers in program schools in year 3 did not have 3 full years of coaching, and some teachers who had moved to control schools 1High teacher turnover is common in developing countries: Zeitlin (2021) reports turnover of about 20% per year in Rwanda, and Schaffner, Glewwe, and Sharma (2024) report 18–21% turnover per year for teachers in Nepal. Quantitative Economics 16 (2025) Can teaching be taught? 189 by year 3 had been coached in previous years. The average causal response (ACR) estimates after 3 years, which adjust the ITT estimates to estimate the impact of 3 years of exposure to teachers who were coached in all 3 years, are 0.180 s.d. for mathematics and 0.162 s.d. for reading comprehension. Our estimates for the effect of coaching on pedagogical skills are smaller than those found in developed countries (0.49 s.d. on instructional practices, see Kraft, Blazar, and Hogan (2018)). This may reflect the scale of the program, and Peru’s high rate of teacher turnover. Yet we address two unresolved questions on coaching’s impact on teachers’ pedagogical skills in developing countries. We show that: (i) A program implemented at scale, even with high teacher turnover, can still exhibit positive impacts; and (ii) General pedagogical skills can be increased. Furthermore, while our estimated effects on student learning may seem small, they are similar, and in one sense larger, than those typically found in developing countries. Evans and Yuan (2022) surveyed 224 education studies and found that the median effect on learning outcomes is 0.10 s.d., and these effect sizes decrease with the size of the study. For large studies, those with over 5000 students, the median effect is only 0.05 s.d. To our knowledge, no prior study has evaluated the effects on pedagogy and student learning of a large-scale teacher coaching program in a developing country. Most in-service training programs evaluated in those countries are small-scale pilots or efficacy trials run by researchers or NGOs (Evans and Popova (2016)). For example, Cilliers et al. (2020) estimated the impact of coaching and centralized teacher training on student reading skills implemented in 180 public schools in South Africa, and Albornoz, Anauati, Furman, Luzuriaga, Podesta, and Taylor (Albornoz et al. (2020)) estimated the impact of teacher coaching to improve student learning of science implemented in 70 public schools in Argentina. In contrast, we evaluate a program implemented in 3797 rural schools in Peru. The issue of scale is relevant for coaching programs’ effectiveness because of two features of this type of in-service training. First, the program’s success depends on the supply of qualified coaches. If these skills are scarce, expanding the program likely will reduce its quality, and thus its effectiveness. Second, classroom observation and personalized feedback requires coaches to travel to several schools. This can be costly and can complicate program delivery if scaling-up implies serving schools in very remote areas. This is very likely for rural schools in developing countries, whose teachers often require additional training. Teacher turnover not only complicates identification of program effects, as discussed above, but may also make coaching less effective by reducing compliance. Teachers who leave a school before the program ends may not receive the full “dose” of coaching, and program schools that receive new teachers may have staff who are only partially coached. We know of only one other study that considered teacher turnover when evaluating a teacher training program. Matsumura, Garnier, Correnti, Junker, and Bickel (2010) estimated the effect of a literacy coaching program in 32 elementary schools in Texas. Stressing how such turnover can thwart schools’ efforts to improve instruction through teacher training, the authors estimated the program’s effect on the reading skills of the 190 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) students of teachers recruited to replace those who left their school in the first year of the program. They found a positive association between teachers’ program participation and their students’ reading skills. However, the nonrandom composition of their sample (recruited teachers in program and nonprogram schools may not be comparable) casts doubt on the causal interpretation of their results. Finally, the literature thus far does not provide a clear indication as to whether coaching can improve general pedagogical skills. Most evaluations of coaching programs focus on pedagogy for a specific topic or course. For example, Albornoz et al. (2020) focus on improving teaching of science, and Cilliers et al. (2020) focus on reading. Kraft, Blazar, and Hogan (2018) highlight a lack of causal evidence on the effect of coaching for subjects other than reading or literacy. Some papers measure the effect of training on teacher time allocation (Bruns, Costa, and Cunha (2018)) or on using specific types of teaching (Kotze, Fleisch, and Taylor (2019)), but not on their teaching skills. The pedagogical skills of public-school teachers in developing countries are generally low, and a key policy question is whether coaching can improve a broad set of teaching skills. The rest of the paper is organized as follows. Section 2describes the program and explains the evaluation design. Section 3presents our analytical framework, defines several treatment effects, and explains which can be estimated. Sections 4and 5present estimates of the program’s impact on teachers’ pedagogical skills and on student learning, respectively. Section 6provides concluding remarks, policy implications, and advice for future research. The Supplemental Appendix (Castro, Glewwe, Heredia-Mayo, Majerowicz, and Montero (Castro et al. (2024a))) contains additional tables and derivations. 2. The coaching program and its evaluation design 2.1 Teacher hiring and movement in Peru There are two types of teachers in the Peruvian school system: Tenured (civil servant) teachers (nombrados), who have a permanent position in a particular school, and contract teachers (contratados) on temporary 1-year contracts who are filling in for tenured teachers who are temporarily absent or for unfilled vacancies in particular schools. In the schools we consider—multigrade and monolingual—most (70–75%) of teachers are tenured. Teachers become tenured through a selection process with two stages. The first stage consists of a nationally administered exam that covers reading comprehension, logical reasoning, and knowledge of pedagogical practices. Teachers with the minimum passing grade on the exam proceed to a second stage that is carried out by regional education offices and includes an interview and in-classroom observation of teaching practices. Teachers who do not reach a minimum passing grade in an exam in the first stage of the selection process, or who receive a passing grade but are unsuccessful at the second stage, can fill temporary teaching positions as contract teachers (and can continue trying to obtain tenure). Contract teachers have annual contracts: at the end of each school year, they must apply for either a renewed contract at their current school or for a contract position at another school. When applying to new schools, contract teachers Quantitative Economics 16 (2025) Can teaching be taught? 191 can apply to as many schools as they want within one region. They are then ranked according to their scores on the latest exam, and teachers with the best scores get their top priority of schools. Teachers can maximize their probability of getting placed by ranking as many schools as they are willing to go to, and by selecting less popular schools (e.g., schools located in remote rural areas). Tenured teachers tend to move less frequently given their permanent position in their schools, but they can request a transfer to another tenured position. In order to do this, they must meet three requirements: have been in a tenured position for at least 3 years, have been in the current tenured position for at least 2 years, and cannot move to another school within the same school district (Peru has about 250 school districts (UGELs)). 2.2 The coaching program In 2010, the Peruvian government initiated coaching programs to improve public primary school teachers’ pedagogical practices. As per Ministry of Education guidelines, the school district authority (UGEL) hires coaches for teachers in the schools targeted by the program, who are selected from top-performing teachers. Coaches must have a pedagogical college or university degree, 5 or more years of primary school teaching experience, and at least 1 year of experience training or providing support to teachers. Administrative data show that coaches rank much higher than other teachers in the Ministry’s teacher evaluations. Coaches were paid the equivalent of US$ 1200 per month, about double the average teacher’s wage. The Ministry of Education sets the standards for hiring coaches, and for the general program design, but the UGELs select and hire the coaches. Each coach works with eight teachers, and UGELs decide how to match coaches to teachers. Coaches are hired annually. About 20% continue for another year, but only 5% stay in the same school the next year. The coaching program is a substantial investment by Peru’s government, costing over US$ 130 million per year.2By 2016, teachers in over 14,000 public schools with more than 900,000 students were being coached under several coaching programs. Over 90% of these schools are primary schools. There are three versions of the program for primary schools: (i) bilingual coaching (for schools where most students speak a Peruvian indigenous language); (ii) monolingual multigrade coaching (for schools where most students speak Spanish and there are fewer teachers than grades taught); and (iii) monolingual full-teacher coaching (for schools large enough to have one teacher per grade and where most students speak Spanish). This paper evaluates the second type of coaching program,3which operates primarily in rural areas.4Over 90% of Peru’s rural public primary schools are multigrade, which 2It was not implemented in 2021 and 2022 due to Covid-19, after which it was restarted, but on a smaller scale. 3Although the three types of coaching programs have some differences (such as the teacher-to-coach ratio or the bilingual certification of coaches), what happens during the coaching sessions is very similar in all three types. 4About 95% of the 6218 schools in our study are located in rural areas. 192 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) typically have two teachers and about 30 students. Rural multigrade schools are the majority of schools with coaching programs. The monolingual multigrade program is particularly expensive because the target schools tend to be very far apart, so the program requires a large number of coaches and significant travel expenses. This version of the program alone, called Acompañamiento Pedagógico Multigrado (APM) in Spanish, cost the government about US$ 40 million in 2016 and served 174,000 students. This implies an annual cost of US$ 228 per student, which is over 20% of the total expenditure per student in Peru’s primary schools (in 2015, average spending per primary school student was 2800 soles, or about US$ 940). A coach’s work consists of several tasks. First, the coach meets the school principal and gathers information about the educational context. Then the coach attends all teachers’ class sessions (one teacher per day) to observe their classroom performance and make an initial diagnostic assessment. The coach uses this assessment to identify the competencies that the teachers must improve and develops an improvement plan with each teacher. During the school year, the coach observes eight more of each teacher’s class sessions at regular intervals. The program is usually implemented for 3 consecutive years. After each classroom observation, the coach and the teacher meet to discuss the progress made in terms of the improvement plan. The coach sends monthly and quarterly reports to the UGEL, and to the school principal, on each teacher’s progress, and on areas for improvement. At the end of the year, the coach provides a final feedback session for each teacher, collecting his or her impressions of the process, and then writes a final report for each teacher on the achievements, actions, and areas requiring further effort, referencing the initial improvement plan. In addition to the classroom observations, each coach organizes eight workshops per year for his or her teachers to discuss pedagogical practices and encourage the exchange of ideas. In the workshops, all the teachers for a given coach gather with the coach to discuss a particular pedagogical topic of interest. The coach encourages and guides the exchange of ideas and successful practices among teachers and provides theoretical support on the chosen subject. At the end of each workshop, the group chooses a new topic for the next gathering. Instead of content knowledge of the material, the program focuses on strengthening pedagogical skills and on developing the ability of teachers to periodically reflect on their own strengths and weaknesses and adjust their behavior accordingly: “The pedagogical coaching promotes the development and strengthening of skills related to understanding the student in her context, curricular planning, guiding learning, ensuring a safe school environment, and evaluating student learning. In addition, it promotes the development of critical thinking skills like self-reflection and analysis, through exercises that seek reflection and critical analysis of the teacher’s own performance.” (APM Manual) APM uses a cascade system. Each coach is trained, supported, and monitored by a pedagogical specialist. Each specialist is required to monitor each coach at least twice per year during the coach’s classroom visits. The specialist also provides two workshops per year directly to teachers. Coaches and specialists follow the “Framework for Good Quantitative Economics 16 (2025) Can teaching be taught? 199 In year 2, there may be interactions (denoted by γk 1,2) of the coaching in years 1 and 2: y2 j=y1 j+λj+δkT2 j+γk 1,2T1 jT2 j =θ2 j+δkT1 j+T2 j+γk 1,2T1 jT2 j,fork=R, L, D, M. (3) The second line substitutes out y1 jusing (2), and θ2 jdenotes θ1 j+λj=y0 j+2λj. If, for example, the second year’s impact of coaching is less than that of the first year, then the interaction term γk 1,2 is <0. Also, γk 1,2 can include depreciation of teacher skills produced by the program. For year 3, further interaction effects are needed. The equation for y3 jis y3 j=y2 j+λj+δkT3 j+γk 1,2T1 jT2 j+T1 jT3 j+T2 jT3 j+γk 1,2,3T1 jT2 jT3 j =θ3 j+δkT1 j+T2 j+T3 j+γk 1,2T1 jT2 j+T1 jT3 j+T2 jT3 j+γk 1,2,3T1 jT2 jT3 j, for k =R, L, D, M, (4) where the second line uses (3) to substitute out y2 j,andθ3 j=θ2 j+λj=y0 j+3λj.Notethat the interaction effect for any combination of 2 years of coaching is assumed to be the same, regardless of which 2 years they are; allowing for different interaction effects for each possible pair of years would do little beyond complicating the notation. The triple interaction γk 1,2,3 can include depreciation of the skills of teachers who are coached for all 3 years. For the APM program, three standard treatment effects can be defined for teacher skills. The first is the average treatment effect (ATE), APM’s impact on the average teacher (when all teachers are treated, i.e., receive coaching). The counterfactual is that no teachers are treated, or equivalently that the program does not exist. ATE for year t is defined as ATEtchr(t)≡Eyt 1−yt 0=Eyt 1−Eyt|No program exists,(5) where the “tchr” subscript indicates that the treatment effect refers to teachers’ skills. For y, the superscript is still years since the program started, but subscripts indicate potential outcomes (1 =treated, 0 =not treated). Implicit in this definition is that the two potential outcomes in year t (yt 1and yt 0) maintain the same potential outcome status (treated or not treated) since year 1, so a teacher who is treated in year 1 is treated for all years between 1 and t, and a teacher who is not treated in year 1 is not treated for check this assumption with our data, yet there are three reasons why it is unlikely that a coached teacher will have sizeable impacts on the skills of other teachers in the same school. First, about 20% of the schools in the test score data, and 49% in the teacher skill data, have only one teacher; for these schools peer effects are not possible. Second, almost all schools that have more than one teacher have only two or three teachers, and they all teach different grades. For example, one teacher teaches grades 1–3 and another teaches grades 4–6. Third, coaching is generally teacher-specific, addressing the pedagogical weaknesses of a specific teacher and the needs of that teacher’s students; other teachers are likely to have different pedagogical weaknesses and students with different needs; this further reduces opportunities for peer effects. If peer effects do occur, such a SUTVA violation would lead to underestimation of ITT effects, so our ITT estimates would be lower bounds for the true ITT parameters. 200 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) all years between 1 and t. The population of teachers for which this treatment effect is defined is all teachers who were teaching in multigrade monolingual schools in Peru in year 1. A more specific example of equation (5)isforyear2(t=2), which is the only year for which teacher skill data are available. This can be expressed as ATEtchr(2)≡Ey2 1−y2 0=2δ+γ1,2, where δ=δRpR+δLpL+δDpD+δMpM,γ1,2 =γR 1,2pR+γL 1,2pL+γD 1,2pD+γM 1,2pM,andp k is the proportion of type k teachers. Appendix B in Castro et al. (2024a) gives expressions for ATEtchr(1)and ATEtchr(3). Next, consider the intention to treat (ITT) effect. This is the program’s impact on skills in year t of teachers randomly assigned to APM schools in year 1, regardless of the school they were in (APM or non-APM) in later years. The counterfactual is random assignment to a non-APM school in year 1, regardless of where they taught in later years. It is defined as ITTtchr(t)≡Eyt|Rtchr,year1 =1−Eyt|Rtchr,year 1 =0,(6) Rtchr,year1 refers to the teacher’s school in year 1, which can differ from his or her school in year t. An example of equation (6) is for year 2, the year with teacher skill data:14 ITTtchr(2)≡Ey2|Rtchr,year1 =1−Ey2|Rtchr,year 1 =0 =δ+pRδR+γR 1,2+pLγL 1,2 +pMτγM 1,2, where τis the proportion of teacher positions in APM schools among the population of all monolingual multigrade schools. The intuition is that δis the effect of the first year, when all teachers follow their random assignment, and the other terms are the effects on the teachers treated in the second year (remainers, likers, and the movers who randomly end up in APM schools in year 2). The counterfactual for remainers is being in a non-APM schools for both years, while the counterfactual for likers, and for movers who randomly (with probability τ) end up in an APM school in year 2, is being in a nonAPM school in year 1 and an APM school in year 2. A final important point is that, unlike ATEtchr(2),ITT tchr(2)depends on τ.Ina “small-scale” RCT, τwould be almost zero and so could be ignored, but in an “at-scale” RCT τwill be larger and will affect ITTtchr(2). The intuition is that a proportion τof movers in APM schools in year 1 will also be in APM schools in year 2, which “turns on” the interaction effect from 2 years of coaching; if the proportion of APM schools had been very small, very few movers who moved into APM schools in year 2 would have been treated in year 1. In addition, there is a more subtle impact of τon ITTtchr(2): it determines the level of competition among “potential likers” to move into APM schools, and similarly the extent of competition among “potential dislikers” to move into non-APM schools. This will 14Note a slight abuse of notation: “R” is used in two different ways. If it is “normal” size (not a superscript) it indicates a school’s random assignment, but if it is a superscript, it denotes remainer teachers. Quantitative Economics 16 (2025) Can teaching be taught? 201 ultimately determine the proportions of teachers who are actual likers and dislikers, and thus the proportions of teachers who are remainers and movers. However, if there are no likers or dislikers, then the value of τwould not affect the proportions of remainers and movers. Another treatment effect that is often estimated for randomized control trials is a local average treatment effect (LATE).15 It is defined only for a binary treatment variable, but the APM treatment variable can have more than two values since teachers can switch schools: the treatment can be 0, 1, 2, or 3 years. Angrist and Imbens (1995)extended LATE to nonbinary treatments, which they call an average causal response (ACR). The general definition is ACRtchr(t)≡ t  s=1 Eyt s−yt s−1|Tt 1≥s>Tt 0ProbTt 1≥s>Tt 0 t  r=1 ProbTt 1≥r>Tt 0 ,(7) where Tt 0is the (potential) number of years of coaching up through year t for teachers assigned to non-APM schools in year 1, and Tt 1is the (potential) years of coaching up through year t for a teacher assigned to an APM school in year 1.16 The subscripts on y indicate the value of y given a (potential) number of years of treatment (which varies from 0 to 3), not the value of y given a binary “treated or not treated” variable, in contrast to the definition of ATEtchr(t). Consider equation (7) for year 2, the only year with teacher skill data: ACRtchr(2)≡Ey2 1−y2 0|T2 1≥1>T2 0ProbT2 1≥1>T2 0 ProbT2 1≥1>T2 0+ProbT2 1=2>T2 0 +Ey2 2−y2 1|T2 1=2>T2 0ProbT2 1=2>T2 0 ProbT2 1≥1>T2 0+ProbT2 1=2>T2 0 =δ+pRδR+γR 1,2+pLγL 1,2 +pMτγM 1,2/1+pR=ITTtchr(2)/1+pR. The intuition behind this equation is the following. The term E[y2 1−y2 0|T2 1≥1>T2 0]is the impact on teacher skills of receiving 1 year of treatment, relative to having zero years of treatment, as indicated by the subscripts on the y terms, for teachers who would have had at least 1 year of treatment by year 2 if assigned to an APM school in year 1 (T2 1≥1), but would not have been treated by year 2 if assigned to a non-APM school in year 1 (T2 0<1). Of the four teacher types, this includes all remainers and dislikers, and movers who randomly switched to a non-APM school in year 2 (for whom T2 0=0andT 2 1=1). 15For the APM context, there is no ATT (average treatment effect on the treated) because ATT requires that some teachers assigned to the treatment (Rtchr,year 1 =1) are never treated. Such teachers do not exist in the APM context because all the teachers who were randomly assigned to the APM schools were treated in year 1. 16For the general case, possible values for both Tt 0and Tt 1are integers from 0 to t. Yet, for the APM program, all teachers followed their random assignment in year 1, so possible values for Tt 0are 0 to t −1, and for Tt 1are 1 to t. 202 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) The term E[y2 2−y2 1|T2 1=2>T2 0]is the impact on teacher skills of receiving a second year of the treatment, relative to having 1 year of treatment, as indicated by the subscripts on the y terms, for teachers who would have had 2 years of treatment in year 2 if assigned to an APM school in year 1 but only 0 or 1 year of treatment in year 2 if assigned to a nonAPM school in year 1. This includes all remainers, all likers, and movers who randomly switched to APM schools in year 2 (for whom T2 0=1andT 2 1=2). Turning to the sums of the probabilities in the denominators, Prob[T2 1≥1>T2 0]is the probability that a teacher is a remainer, a disliker, or a mover who randomly switches to a non-APM school in year 2, and Prob[T2 1=2>T2 0]is the probability that a teacher is a remainer, a liker, or a mover who randomly switches to an APM school in year 2. Their sum is greater than 1; remainers are “counted twice” since they are included in both probabilities. Likers, dislikers, and movers are “counted” only once. In effect, ACRtchr(2)is an average of: (a) the (average) impact on teacher skills of going from no treatment to 1 year of treatment for remainers, dislikers, and those movers who randomly move to a non-APM school in year 2; and (b) the (average) impact on those skills of going from 1 to 2 years of treatment for remainers, likers, and the movers who randomly move to APM schools in year 2. Thus, ACRtchr(2)is the average of the impact on teacher skills for each additional year of treatment due to random assignment in year 1 to an APM school, with remainers getting “double weight” since that assignment raises their years of treatment by 2 years, but for all others that assignment raises years of treatment by only 1 year. Importantly, note that, for any t, ACRtchr(t)is a per year (not a cumulative) impact, averaging over years of treatment induced by schools’ random assignment to APM in year 1. The cumulative effect is ACRtchr(t)multiplied by the years of coaching induced by a school’s random assignment to APM (the denominator in (7)): this equals ITTtchr(t). A final aspect of ACRtchr(2)to note is that, like ITTtchr(2),itisa function of τ, since its numerator is ITTtchr(2). The three treatment effects discussed so far focus on particular teachers, and so they follow teachers who move to other schools. But many teacher training or coaching programs focus on particular schools, so it is useful to define treatment effects for the teachers currently in the schools that implemented APM. There are two possibilities for treatment effects that focus on schools.17 The first is an average treatment effect (ATE) on teacher skills for those schools, where the counterfactual is no program at all, which we denote as ATEsch.Thisisdefinedasfollowsfor year t: ATEsch(t)≡Eyt|R=1−Eyt|Program does not exist.(8) As above, consider again the specific case of year 2, the only year with teacher skill data: ATEsch(2)=2δR+γR 1,2pR+δL(1+τ)+γL 1,2τpL/τ+δM(1+τ)+γM 1,2τpM(μ/τ) +θ2,LpL(1−τ)/τ−θ2,DpD+θ2,MpM(μ/τ)−1, 17ACRsch(t)is not well-defined since teachers who move into the 6218 schools have no instrumental variable. Quantitative Economics 16 (2025) Can teaching be taught? 203 where μis the proportion of all movers who move to an APM school in year 2 or year 3, and the θ2,k terms are averages of θ2 jfor year 2 for type k teachers.18 The first line of ATEsch(2)is the “direct” treatment effect and the second is a “composition” effect, which accounts for differences in average θbetween likers, who move into APM schools in year 2, and dislikers, who move out of APM schools in year 2 (and also accounts for changes in the distribution of movers across the two types of schools, who compete with likers to get into APM schools and with dislikers to get into non-APM schools). Note that ATEsch(2), and more generally ATEsch(t)with t≥2, also depends on τ.Intuitively,τdetermines the proportions of likers and movers in APM schools (and of dislikers and movers in nonAPM schools), yet this is no longer the case if there are no likers or dislikers, as explained below. The second treatment effect for teacher skills that focuses on schools is ITTsch;it is similar to ATEsch except that the counterfactual is the skills of teachers in non-APM schools: ITTsch(t)≡Eyt|R=1−Eyt|R=0.(9) For year 2, this is ITTsch(2)=2δR+γR 1,2pR+δL(1+τ)+γL 1,2τpL/τ+δM(1+τ)+γM 1,2τpM(μ/τ) −δDpD(τ/(1−τ))+δMτpM(1−μ)/(1−τ) +θ2,LpL/τ+θ2,MpM(μ/τ) −θ2,DpD/(1−τ)+θ2,MpM(1−μ)/(1−τ). The first two lines are the (net) treatment effect; the last two are the composition effect. As with ATEsch(t),ITT sch(t)depends on the proportion of schools that are treated (τ) when t ≥2. 3.3 Treatment effects for student learning Next, consider treatment effects on student skills. Assume that the skill (measured by a test score) of student i at the end of year t, denoted by st i, is determined by his or her skill at the end of the previous year (st−1 i) and the skills of his or her teacher in year t (yt j), where j is the teacher that student i had in year t, and πis the impact of teacher skill on student skills: st i=σst−1 i+πyt j. (10) 18To see where the μ/τ term comes from, note that the number of teaching positions in a school rarely changes. If the number of those positions is fixed in all schools, this definition of μ(where μis determined by the application process that also determines the proportions of teachers who are likers, dislikers, movers, and remainers; see Section 2.1), implies that, among all teachers in APM and non-APM schools, the proportion who are movers in APM schools in year 2 or 3 is μpM. Focusing on APM schools, this proportion must be divided by τ, yielding (μ/τ)pM. Similar derivations show the proportion of movers in non-APM schools in year 2 or 3 is [(1−μ)/(1−τ)]pM. 204 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) Each school is randomly assigned to be either an APM (R =1) or non-APM (R =0) school, an assignment that is fixed over time. Analysis of student skills is simplified by the fact that few students change schools (see Section 4.2), and each school follows its random assignment. We define three treatment effects for student skills. The first two, ATEstud and ITTstud, are analogous to the two treatment effects defined for their schools (ATEsch and ITTsch). All three treatment effects for years 2 and 3 are complex due to several possible “histories” for students’ teachers in those years. For example, in year 2 a student’s teacher in an APM school could be a liker who was in an APM school in years 1 and 2, or a liker who was in a non-APM school in year 1 but in an APM school in year 2. Another example is a student in an APM school in year 3; if he or she was taught by a treated teacher in year 1 (this is certain as the student was in an APM school in year 1), and by a teacher in year 2 who had APM in year 2 but not year 1, and by a teacher in year 3 who had APM in years 2 and 3 but not year 1, he or she was exposed to 4 years of teacher coaching, and the cumulative learning gain from this exposure is averaged over the 4 years. The general definition of ATEstud for year t is ATEstud(t)≡Est|R=1−Est|Program does not exist. (11) Applying this definition to year 1 yields ATEstud(1)=πδ. Applying it to year 3 (recall that test score data exist only for 2016 and 2018) yields (see Appendix B in Castro et al. (2024a) for the derivations): ATEstud(3)=σATEstud(2)+πATEsch(3)=σσATEstud(1)+πATEsch(2)+πATEsch(3) =σ2πδ +σπ2δR+γR 1,2pR+δL(1+τ)+γL 1,2τpL/τ +δM(1+τ)+γM 1,2τpM(μ/τ) +π3δR+3γR 1,2 +γR 1,2,3pR+δL(2+τ)+γL 1,2(2τ+1)+γL 1,2,3pL/τ +δM(1+2τ)+γM 1,2τ(2+τ)+τ2γM 1,2,3pM(μ/τ) +σπθ2,LpL(1−τ)/τ−θ2,DpD+θ2,MpM(μ/τ)−1 +πθ3,LpL(1−τ)/τ−θ3,DpD+θ3,MpM((μ/τ)−1). For ATEstud(3), the first four lines are the treatment effect, and the last two lines are the composition effect. Again, for t =2ort=3, ATEstud(t)depends on τ. Turn next to ITT. The general definition for year t is ITTstud(t)≡Est|R=1−Est|R=0. (12) For year 1, ITTstud(1)=ATEstud(1)=πδ, as all teachers follow their schools’ random assignment in year 1. For year 3, applying the general definition yields (Appendix B in Castro et al. (2024a) gives details): ITTstud(3)=σITTstud(2)+πITTsch(3)=σσITTstud(1)+πITTsch(2)+πITTsch(3) =σ2πδ +σπ2δR+γR 1,2pR+δL(1+τ)+γL 1,2τpL/τ Quantitative Economics 16 (2025) Can teaching be taught? 205 +δM(1+τ)+γM 1,2τpM(μ/τ) −δDpD(τ/(1−τ))+δMτpM(1−μ)/(1−τ) +π3δR+3γR 1,2 +γR 1,2,3pR+δL(2+τ)+γL 1,2(2τ+1)+τγL 1,2,3pL/τ +δM(1+2τ)+γM 1,2τ(2+τ)+τ2γM 1,2,3pM(μ/τ) −πδDpD(τ/(1−τ))+δM2τ+τ2γM 1,2pM(1−μ)/(1−τ) +σπθ2,LpL/τ+θ2,MpM(μ/τ) −θ2,DpD/(1−τ)+θ2,MpM(1−μ)/(1−τ) +πθ3,LpL/τ+θ3,MpM(μ/τ) −θ3,DpD/(1−τ)+θ3,MpM(1−μ)/(1−τ). The first six lines are the (net) treatment effect, and the last four are the composition effect. Note again that, for t =2or3,thatITT stud(t)depends on τ. The third treatment effect for students is the (average) impact of an additional year of teacher coaching on student learning, averaged over all additional years of that coaching that a student experiences. In effect, this is a transfer of the ACRtchr treatment effects on teacher skill onto student learning, which is complicated by the many different “histories” a student can have of treated teachers in years 2 and 3. We call these treatment effects ACRstud, though they differ from ACRtchr (and so differ from the Angrist and Imbens ACR effects) since students are not directly treated but instead are indirectly treated by exposure to treated teachers. The general definition of ACRstud in year t (1, 2, or 3) is ACRstud(t)≡Est|R=1−Est|R=0 Ehtchr(t)|R=1−Ehtchr(t)|R=0, (13) where htchr(t)is the cumulative “history” from year 1 to year t of a student’s exposure to teachers with APM coaching. For example, a student in a treated school in year 2 had a coached teacher in year 1, but in year 2 the teacher could have 1 or 2 years of coaching (e.g., 1 for a teacher in a non-APM school in year 1), so the student’s htchr(2)could be 2 or 3. The expected value of htchr(t)averages over the types of teachers in the school from year 1 to year t. For year 1, ACRstud(1)=ATTstud(t)=ITTstud(t)since all teachers follow their random assignment in year 1, so ACRstud(1)=πδ. For year 3, applying the definition in (13) yields ACRstud(3)=Es3|R=1−Es3|R=0 Ehtchr(3)|R=1−Ehtchr(3)|R=0 =ITTstud(3) 1+5pR+(3+2τ)pL/τ +(2+3τ)pM(μ/τ)−2τpD/(1−τ)+3τpM(1−μ)/(1−τ). 206 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) To understand this derivation, note that the numerator is ITTstud(3). The first expression in brackets in the denominator, 1 +5pR+(3+2τ)pL/τ +(2+3τ)pM(μ/τ),is E[htchr(3)|R=1], the average cumulative exposure to years of teacher coaching of a student in an APM school in year 3. The “1 +” term is exposure to a coached teacher in year 1. In years 2 and 3, the probability of getting a remainer teacher is pR, and the probabilities of getting a liker or mover teacher are pL/τ and pM(μ/τ), respectively. If a student gets a remainer teacher in year 2, he or she is exposed to 2 more years of accumulated coaching since that teacher has had 2 years of coaching by year 2, and if the student gets a remainer teacher in year three he or she will get 3 more years of accumulated coaching, for a total of five additional years (beyond year 1). If the student gets a liker teacher in year 2, the average liker teacher will have had (1+τ) years of coaching (one in year 2 and one more for a proportion τof those teachers in year 1), and if the student gets a liker teacher in year 3, that will expose him or her to an additional 2 +τyears of accumulated coaching, so overall exposure to liker teachers will provide 3 +2τyears of accumulated coaching. Finally, exposure to a mover teacher in year 2 leads to 1 +τadditional years of accumulated coaching, and exposure to a mover in year 3 adds another 1 +2τ(since movers move randomly every year). Similar calculations for students who randomly end up in non-APM schools in year 1 (and the next 2 years) lead to accumulated coaching from dislikers and movers of 2τpD/(1−τ)+3τpM((1−μ)/(1−τ)). As with ACRtchr(2),ACR stud(3)is a per year effect (averaged over the relevant years of exposure to coached teachers). To obtain a cumulative effect for exposure to fully coached teachers in all 3 years, which is a weighted average over the four teacher types (where the weights are probabilities of teacher types being in treated schools), multiply ACRstud(3)by 6. 3.4 Treatment effects if no likers or dislikers The treatment effects for years 2 and 3 in Sections 3.2 and 3.3 are much simpler if there are no likers or dislikers, leaving only remainers and movers. The equations simplify to19 ATEtchr(2)=2δ+γ1,2,whereδ=δRpR+δMpMand γ1,2 =γR 1,2pR+γM 1,2pM, ITTtchr(2)=δ+pRδR+γR 1,2+pMτγM 1,2, ACRtchr(2)=δ+pRδR+γR 1,2+pMτγM 1,2/2pR+pM=ITTtchr(2)/1+pR, ATEsch(2)=2δR+γR 1,2pR+δM(1+τ)+γM 1,2τpM, ITTsch(2)=δ+δR+γR 1,2pR+pMτγM 1,2. Note that ITTsch(2)=ITTtchr(2),butATE sch(2)= ATEtchr(2). ATEstud(3)=σ2πδ +σπ2δR+γR 1,2pR+δM(1+τ)+γM 1,2τpM +π3δR+3γR 1,2 +γR 1,2,3pR+δM(1+2τ)+γM 1,2τ(2+τ)+τ2γM 1,2,3pM, 19They follow from the results in Sections 3.2 and 3.3:p L=pD=0 and μ=τif there are no likers or dislikers. Quantitative Economics 16 (2025) Can teaching be taught? 207 ITTstud(3)=σ2πδ +σπ2δR+γR 1,2pR+δM+γM 1,2τpM +π3δR+3γR 1,2 +γR 1,2,3pR+δM+γM 1,2τ2+τ2γM 1,2,3pM, ACRstud(3)=πσ2δ+(3+2σ)δR+(3+σ)γR 1,2 +γR 1,2,3pR+δM(σ+1)+γM 1,2τ(σ+2)+γM 1,2,3τ2pM 1+5pR+2pM =ITTstud(3) 1+5pR+2pM. Note that there are no composition effects for ATEsch(2),ITT sch(2),ATE stud(3),and ITTstud(3). Also, the absence of likers and dislikers (pL=pD=0) implies that there are only remainers and movers, and that μ=τ(movers are equally distributed over APM and non-APM schools since they do not compete with likers or dislikers to move into an APM or non-APM school). 3.5 What do OLS and IV regressions estimate? Most, but not all, of these treatment effects can be estimated by OLS or IV regression. We have two samples of teachers, one (imperfectly) follows the teachers who were in APM and non-APM schools in year 1 (Sample 1), and the other focuses on the teachers in the APM and non-APM schools in any given year (Sample 2). OLS regression of Sample 1 teachers’ skills in year t on a constant term and a dummy variable for assignment to an APM school in year 1 yields an unbiased estimate of the ITTtchr(t)treatment effect.20 For example, consider year 2: ˆ βy 1OLS,t=2=Ey2|Rtchr,year 1 =1−Ey2|Rtchr,year1 =0 =δ+pRδR+γR 1,2+pLγL 1,2 +τpMγM 1,2 =ITTtchr(2). The “1” subscript indicates Sample 1 teachers. Appendix B in Castro et al. (2024a) presents this derivation, as well as those for years 1 and 3. It also presents the derivations for the other OLS and IV estimators in this subsection, for all 3 years, and shows that OLS estimation applied to Sample 2 teachers estimates ITTtchr(t)(recall that ITTtchr(t)= ITTsch(t)if there are no likers or dislikers). Next, consider IV estimation using Sample 1 teachers. Let TTot, t denote the number of years that a teacher has participated in the program up through year t. IV regression uses random assignment as an instrument for TTot, t to estimate the (average) impact of a year of exposure to the program on teacher skills. This yields unbiased estimates of ACRtchr(t). For year 2: ˆ βy 1IV,t=2=Ey2|Rtchr,year 1 =1−Ey2|Rtchr,year 1 =0 ETTot, 2 |Rtchr,year1 =1−ETTot, 2 |Rtchr,year1 =0 =δ+pRδR+γR 1,2+pLγL 1,2 +pMτγM 1,2/1+pR=ACRtchr(2). 20Almost all of the regressions in this paper have other explanatory variables, but since random assignment is by definition uncorrelated with these other variables, the first line in the ˆ βy 1OLS,t=2equation still holds by the Frisch–Waugh theorem. Regressions without these explanatory variables (e.g., Table 6)yields very similar results. 208 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) One can also apply OLS to Sample 2 teachers, the teachers who, in any given year, teach in the schools that were randomly assigned in year 1 to be APM or non-APM schools. An OLS regression of Sample 2 teachers’ skills in year t on a constant and a dummy for teaching in an APM school in year t yields an unbiased estimate of ITTsch(t). So, for year 2: ˆ βy 2OLS,t=2=Ey2|R=1−Ey2|R=0 =2δR+γR 1,2pR+δL(1+τ)+γL 1,2τpL/τ −[δDpDτ/(1−τ)+pMδMμ−τ2/τ−τ2+μγM 1,2 +θ2,LpL/τ+θ2,MpM(μ/τ) −θ2,DpD/(1−τ)+θ2,MpM(1−μ)/(1−τ)=ITTsch(2). In general, IV estimation cannot be used for Sample 2 teachers in year 2 since some of those teachers moved into both APM schools and non-APM schools that were not part of the initial random assignment, such as teachers working in monolingual multigrade schools in year 1 that had ECE scores above the threshold that determined eligibility for the randomized expansion (see Section 2.3). These Sample 2 teachers have no instrument, so IV estimation cannot be done for Sample 2 teachers. Next, consider OLS regression for student test scores, more specifically regressing those scores on a constant and a dummy indicating being in an APM school. OLS regression of students’ test scores in year t on a constant and a dummy for being enrolled in an APM school in year t yields an unbiased estimate of ITTstud(t). For years 1 and 3, this implies that ˆ βs OLS,t=1=Es1|R=1−Es1|R=0=πδ =ATEstud(1)=ITTstud(1)=ACRstud(1), ˆ βs OLS,t=3=Es3|R=1−Es3|R=0 =σ2πδ +σπ2δR+γR 1,2pR+δL(1+τ)+γL 1,2τpL/τ +δM(1+τ)+γM 1,2τpM(μ/τ)−δDpDτ/(1−τ)+δMτpM(1−μ)/(1−τ) +π3δR+3γR 1,2 +γR 1,2,3pR+δL(2+τ)+γL 1,2(2τ+1) +τγL 1,2,3pL/τ+δM(1+2τ)+γM 1,2τ(2+τ)+τ2γM 1,2,3pM(μ/τ) −πδDpDτ/(1−τ)+δM2τ+τ2γM 1,2pM(1−μ)/(1−τ) +σπθ2,LpL/τ+θ2,MpM(μ/τ) −θ2,DpD/(1−τ)+θ2,MpM(1−μ)/(1−τ) +πθ3,LpL/τ+θ3,MpM(μ/τ) −θ3,DpD/(1−τ)+θ3,MpM(1−μ)/(1−τ) =ITTstud(3). Quantitative Economics 16 (2025) Can teaching be taught? 215 Figure 2. Balance in teacher characteristics for the original and observed in year 2 teachers who worked in an evaluation sample school in 2016 (Sample 1). Note: All regressions include UGEL fixed effects. Standard errors clustered at the school level. Estimates indicate differences in the standardized characteristics of control and treatment groups. Thick and thin lines indicate 90% and 95% confidence intervals, respectively. We do not present the differences in teacher experience and pedagogical degree for the original sample because we do not have information on those variables for the teachers that were not observed at the end of year 2. Figure 3. Balance in school characteristics in the original and observed evaluation sample schools. Note: All regressions include UGEL fixed effects. Estimates indicate differences in the standardized characteristics of control and treatment groups. Thick and thin lines indicate 90% and 95% confidence intervals, respectively. 216 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) We do not compare Sample 2 teachers’ baseline characteristics in 2017 (year 2) between APM and non-APM schools to check for balance at baseline because random assignment of schools in 2016 (year 1) does not ensure such balance across these two groups of schools in year 2. In particular, if certain types of teachers self-select into APM or non-APM schools in year 2, Sample 2 teachers’ baseline characteristics may be correlated with the treatment status of the schools where they worked in year 2. Third, we used data from exams given to teachers in 2014 and 2015 that were used as part of the process by which contract teachers could become permanent civil service teachers and civil service teachers apply for promotion. We found that teachers who scored higher on those exams were less likely to move from other schools in Peru to either an APM school or a non-APM school in our 6218 randomized expansion schools, and also that teachers who scored higher were less likely to move out of the 6218 randomized expansion schools to other schools in Peru (see Table A10 in Castro et al. (2024a)). Most importantly, there is no relationship between these test scores and whether the teachers moved to an APM or a non-APM school, which shows that there is no systematic movement of better (or worse) teachers into APM or non-APM schools. For the student-level data, there is little attrition. Using administrative data on enrollment, we found almost all the students who started in our sample in 2016 (year 1). Student turnover, unlike teacher turnover, is relatively rare in rural primary schools, especially among those targeted by APM since almost 95% are in rural areas where there are very few schools to choose from. Excluding students in their final year of primary school, and averaging over the years 2013 to 2016, only 6.9% of the students in our 6218 primary schools in a given year were not in the same school in the next year. 4.3 Teacher turnover and the proportions of the four types of teacher We use administrative data on the location of teachers as well as the framework established in Section 3to examine teacher turnover and the proportions of the four types of teachers in the sample.22 Table 4shows the 2016–2017 turnover behavior of Sample 1 teachers (i.e., the 12,18923 teachers in the 6218 randomized schools in 2016). By comparing the proportions of teachers in APM and non-APM schools in year 1 who moved to an APM school in year 2 (the difference between equations (A4) and (A1) in Table A5 of Castro et al. (2024a)), we estimate that σpL=−0.024, where σis the proportion of likers in an APM school in a given year (e.g., year 1) who remain in the same school in the next year (e.g., year 2), rather than moving to a different APM school.24 22Table A5 in Castro et al. (2024a) shows where teachers assigned to APM and non-APM schools in the randomization year end up in each type of school 1 year later according to their type and initial sorting. 23Table 4excludes 951 teachers (7.8% of the 12,189 teachers) in the 2016 randomization sample who were not found in the administrative data in 2017; they most likely left the public education system. 24To see how this was calculated, this definition of σimplies that the proportion of likers who move to another APM school is 1-σ.Recallthatμis the proportion of movers in any school who (randomly) move to an APM school in the following year. Thus, of all teachers in an APM school in year 1, pL(1−σ)+pMμis the proportion who move to other APM schools in year 2, and our data show that this proportion is 0.121 (see Table A5 in Castro et al. (2024a)). Similarly, the proportion of teachers in non-APM schools in year 1 who move to an APM school in year 2 is pL+pMμ, and this proportion equals 0.097 in our data. The difference Quantitative Economics 16 (2025) Can teaching be taught? 217 Table 4. Distribution of year 1 teachers by their destination school in year 2. Treatment Arm in 2016 2016–2017 Turnover Teachers Percent APM school Stayed in the same school 4222 63.2 Moved to an APM school 806 12.1 Moved to a non-APM school 1649 24.7 Total 6677 100.0 Non-APM school Stayed in the same school 2847 62.4 Moved to an APM school 440 9.7 Moved to a non-APM school 1274 27.9 Total 4561 100.0 Similarly, by comparing the proportions of teachers in APM and non-APM schools who moved to a non-APM school from year 1 to year 2 (the difference between equations A5 and A2 in Table A5 of Castro et al. (2024a)), we estimate that νpDequals −0.032, where νis the proportion of dislikers in a non-APM school in a given year (e.g., year 1) who remain in the same school in the next year (e.g., year 2), rather than moving to a different non-APM school. Both σpLand νpDare very close to 0. For σpLto equal 0, either σor pL(or both) must equal 0. If σ=0, then all likers change from one APM school to another APM school in the following year. Similarly, ν=0 implies that all dislikers already in a non-APM school in a given year move to another non-APM school the next year. Such turnover seems very unlikely since most teachers (63%) remained in the same school even before the randomized expansion of the APM program (see Table 5). By definition, likers and dislikers have strong incentives to move between schools if, in year 1, they find themselves in a school that is the opposite of their preference (likers starting in a non-APM school or dislikers starting in an APM school), but when they are placed in the school of their preferred type, we would expect turnover to be similar to what was observed in the sample before the program started, 36.6%, not 100%. Therefore, both σ=0andν=0 seem very unlikely. The other option, which we consider the most realistic, is that pLand pD are equal to 0: there are no likers or dislikers. The conclusion that there are no likers or dislikers is a strong claim, so we offer two additional pieces of supporting evidence. First, we analyze how teacher turnover changed over time. If there are likers and dislikers, we would expect an increased movement of teachers in the first year after the randomized expansion of APM as likers and dislikers move to the schools of their preferred type. Since schools stick to their random assignment in later years, we would expect that most of this extra turnover would occur in year 2 (2017), although some could occur in later years if some “potential” likers and dislikers are unable to move to their preferred schools in year 2. Therefore, if there are likers or dislikers, there should be a large spike in the number of teachers moving across treatment arms between 2016 and 2017, followed by a gradual return to regular levels between these two proportions equals σpL,whichis−0.024 in our data. Note that this difference includes the estimates for the mentioned parameters as well as random differences in proportions that arise due to sampling. Thus, small negative estimates are possible if a parameter equals 0. 218 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) Table 5. Teacher turnover between APM and non-APM schools. APM schools 2015 to 2016 2016 to 2017 2017 to 2018 2018 to 2019 Stayed in same school 63% 65% 62% 65% Moved to an APM school 8% 10% 10% 9% Moved to a Non-APM school 14% 14% 15% 12% Moved out of target schools 15% 12% 14% 13% Non-APM schools 2015 to 2016 2016 to 2017 2017 to 2018 2018 to 2019 Stayed in same school 63% 66% 62% 65% Moved to an APM school 11% 12% 12% 11% Moved to a Non-APM school 11% 10% 11% 10% Moved out of target schools 15% 12% 15% 14% All schools 2015 to 2016 2016 to 2017 2017 to 2018 2018 to 2019 Stayed in same school 63% 65% 62% 65% Moved to an APM school 10% 11% 11% 10% Moved to a Non-APM school 12% 12% 13% 11% Moved out of target schools 15% 12% 14% 14% Note: This table shows the year-to-year turnover status of teachers who started each 2-year period in a school within one of the 6218 randomized expansion schools. of movement (from movers randomly moving between APM and non-APM schools, and likers and dislikers moving to another school of their preferred type). Table 5shows the evolution of teacher movement across treatment arms from 2015 to 2019. There is no spike in the movement from APM to non-APM schools from 2016 to 2017; it remains at 14%, the same rate as from 2015 to 2016, and slightly less than from 2017 to 2018. A similar pattern holds for movement from non-APM to APM schools, which from 2016 to 2017 increased slightly to 12% (from 11% from 2015 to 2016) and remained at 12% from 2017 to 2018. These trends are consistent with the claim of no likers or dislikers. A second piece of additional evidence for the claim of no likers or dislikers is comparisons of the characteristics of teachers who worked in the randomized pedagogical skill sample in 2017 (Sample 2). If there were likers or dislikers, one would expect the characteristics of teachers to differ between APM and non-APM schools after turnover, as likers would be only in APM schools and dislikers would be only in non-APM schools. Figure 4shows estimates of treatment effects of APM on a wide set of teacher characteristics in the randomized expansion sample in 2017. We find no effect for any of the characteristics, suggesting that there was no systematic selection of teachers into either APM or non-APM schools, further supporting the claim of no likers or dislikers. 5. The treatment effects of APM 5.1 Teacher skills Overall teacher skills This subsection presents estimates of E[y2|Rtchr,year1 =1]− E[y2|Rtchr,year1 =0], that is, estimates of ITTtchr(2)in equation (6), and E[y2|R=1]− Quantitative Economics 16 (2025) Can teaching be taught? 219 Figure 4. Treatment effects on the composition of teacher characteristics among the teachers in randomized pedagogical skill sample schools in 2017 (Sample 2). Note: All regressions include UGEL fixed effects. Estimates indicate differences in the standardized characteristics of control and treatment groups. Thick and thin lines indicate 90% and 95% confidence intervals, respectively. E[y2|R=0], estimates of ITTsch(2)in equation (9), using OLS regressions for the 455 Sample 1 teachers and the 640 Sample 2 teachers (see Table 3), respectively. We also present the estimates obtained by regressing y2on the predicted years of treatment, instrumented by random assignment in year 1, using Sample 1 teachers. As explained in Section 3.5, this IV approach provides a consistent estimate of the ACRtchr(2)treatment effect. For all estimates, the dependent variable, y2, is an index of pedagogical skills that averages the standardized scores of the eight indicators obtained from classroom observations (see Section 2.3). We present estimates with and without teacher characteristics as covariates when using Sample 1.25 Table 6presents these results. Before discussing the results, recall the claim (Section 4.3) that our population of teachers has no likers or dislikers. Recall also (Section 3.4) that, if there are no likers or dislikers, both ˆ βy 1OLS,t=2and ˆ βy 2OLS,t=2estimate ITTtchr(2), which equals ITTsch(2). Thus, all OLS estimates in Table 6consistently estimate the same parameter. The first and second columns of Table 6present estimates of ITTtchr(2).Theestimate in column (1), which does not control for teacher characteristics, indicates that offering APM for 2 years increases teachers’ pedagogical skills by 0.28 standard deviations (s.d.). The estimate in column (2), when teacher characteristics are added as covariates, is very similar: 0.30 s.d. The estimate for ITTsch(2)in column (3), 0.20 s.d., is 25The use of teacher characteristics as covariates is appropriate only for Sample 1 because characteristics of Sample 2 teachers can be affected by the treatment. In Table A6 of Castro et al. (2024a), we test for interactions between the treatment status and the characteristics of Sample 1 teachers. We find no evidence of heterogeneity by teacher experience, type of contract, position in the teacher career, or sex. These results support the linearity assumption for the teacher skills production function in equation (1). 220 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) Table 6. Aggregate skill: ordinary least squares (OLS) estimates and IV estimates. Ordinary Least Squares Estimates IV Estimates Sample 1 Sample 2 Sample 1 (1) (2) (3) (4) (5) Treatment 0.275 0.300 0.197 0.152 0.166 (0.103) (0.097) (0.098) (0.052) (0.048) [0.008] [0.002] [0.046] [0.003] [0.001] Experience – 0.000 – – −0.000 (0.009) (0.008) Contract teacher – 0.139 – – 0.133 (0.155) (0.138) Teacher career level – 0.109 – – 0.107 (0.044) (0.039) Sex (men =1) – −0.300 – – −0.301 (0.095) (0.085) Age – −0.028 – – −0.027 (0.009) (0.008) R20.29 0.37 0.23 0.29 0.37 Sample size 455 455 640 455 455 Note: All regressions include UGEL fixed effects. Standard errors clustered at the school level are presented in parentheses, and p-values shown in brackets. somewhat lower, even though ITTsch(2)should equal ITTtchr(2). Recall that Sample 1 teachers had high rates of attrition due to difficulties finding teachers who moved; this implies that remainers are very likely overrepresented in Sample 1. In contrast, the proportions of remainers and movers in Sample 2 should correspond to their proportions in the population of teachers in the 6218 randomized expansion schools. Thus, the column (3) estimate is our preferred estimate of ITTtchr(2), which also equals ITTsch(2);the effect after 2 years on teachers’ aggregate pedagogical skill of assigning them to an APM school in year 1 is a 0.20 s.d. increase in those skills Our estimate that ITTtchr(2)=ITTsch(2)=0.20 sheds some light on other parameters of interest. Recall that, in general, ATEtchr(2)≥ITTtchr(2), and if there are no likers and dislikers then ATEsch(2)≥ITTsch(2). Thus, the effect of 2 years of APM coaching on the aggregate pedagogical practice of the average teacher, ATEtchr(2), and the effect of APM on the aggregate pedagogical practice of the teachers in APM schools in year 2, ATEsch(2), are at least as large as, and likely larger than, 0.2 s.d. Columns (4) and (5) in Table 6present our IV estimates of ACRtchr(2)using Sample 1 teachers. They show that, averaging over all years of coaching received, an additional year of coaching increases by 0.15 to 0.17 s.d. the average pedagogical skill of all teachers, but this average gives remainers a “double weight” because random assignment to an APM school induces them to obtain 2 years of coaching. Consistent with the fact that ACRtchr(2)equals ITTtchr(2)/(1+pR), this IV estimate, which is a per year estimate, is somewhat larger than (half of) the Sample 1 estimate of ITTtchr(2),anestimateofcumulative impact over 2 years, in column (2). Quantitative Economics 16 (2025) Can teaching be taught? 221 Specific pedagogical skills The discussion thus far has focused on the aggregate index of pedagogical skills, but one can also estimate ITTtchr(2)for each of the eight more specific pedagogical skills shown in Table 1.Table7shows these results. To minimize spurious statistical significance due to multiple hypothesis testing, Table 7also presents adjusted p-values, using the Romano and Wolf (2016)stepdownmethodtoaccountfor multiple hypothesis testing; these are in brackets below the standard errors. The estimates in Table 7indicate that the biggest impact of assigning teachers to the APM program, in terms of both the size and the statistical significance of the estimated parameters, is on teachers’ lesson planning; the point estimates are 0.34 s.d. for Sample 1 and 0.39 s.d. for Sample 2. There is also evidence that APM raises teachers’ pedagogical skills in developing their students’ critical thinking, although the statistical significance is at best only marginal after controlling for multiple hypothesis testing. 5.2 Student learning This subsection explores the impact of the APM coaching program on student learning, as measured by the National Student Evaluation (ECE) taken 1 and 3 years after the program began (i.e., 2016 and 2018). We compare student test scores in the APM and non-APM schools in the much larger student test score sample. This sample is not restricted to the 340 schools with pedagogical practices data, but it is restricted to those schools that participated in the 2016 ECE and the 2018 ECE. As explained earlier, only schools with five or more students in the relevant grade take the ECE, so we have test scores for only 2567 of the 6218 randomized expansion schools. Table 8presents estimates of the APM coaching program’s treatment effects on average ECE scores for the sample of 2567 schools in 2016 and 2018, after 1 and 3 years of coaching. The ECE is taken at the end of the school year (which is also the end of the calendar year), so the 2016 ECE yields estimates of the APM program’s impact after 1 year for students in grade 2. All teachers complied with their random assignment in 2016, so this is an estimate of ATEstud(1), the average treatment effect of 1 year of APM on student learning. In 2018, the ECE was conducted again, but this time it was done in grade 4, which in general contains the same students who were tested in 2016 in grade 2, except that it excludes students who repeated grade 2 or 3 (about 7–8% of students repeat each year). The 2018 ECE allows us to test for the impact of the program after 3 full years of implementation. Students almost always comply with treatment assignments, yet many teachers switched schools between 2016 and 2018, so we cannot estimate the average treatment effect, ATEstud(3)for 3 years. Rather, we estimate ITTstud(3),whichis alowerboundofATE stud(3)if there are no likers or dislikers. 5.2.1 Results after 1year Table 8presents estimates of the program’s treatment effects on standardized test scores for mathematics and reading comprehension.26 Columns 26Recall that ECE scores exist only for schools with five or more students in a given grade; this greatly reduces the number of schools in the student test score sample. Table A1 in Castro et al. (2024a)showsthat almost all characteristics of the schools with test scores are very similar to those for the 6218 randomized expansion schools. The baseline balance in Table 2is for this smaller subsample of schools, which is the relevant sample for analysis. 222 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) Table 7. Disaggregated skills: Ordinary least squares estimates. (1) (2) (3) (4) (5) (6) (7) (8) Lesson Planning Time Management Critical Thinking Student Participation Class Feedback Written Feedback Classroom Relationships Behavior Management Panel A. Sample 1 Treatment 0.338 0.083 0.248 0.162 0.199 0.136 0.069 0.114 (0.106) (0.108) (0.092) (0.103) (0.107) (0.096) (0.112) (0.098) [0.018] [0.692] [0.064] [0.474] [0.339] [0.499] [0.692] [0.561] N 448 450 450 450 450 448 450 450 R-squared 0.307 0.221 0.281 0.364 0.371 0.332 0.263 0.277 Panel B. Sample 2 Treatment 0.387 −0.073 0.186 0.062 0.094 0.173 0.022 0.019 (0.091) (0.099) (0.090) (0.089) (0.103) (0.097) (0.091) (0.088) [0.002] [0.917] [0.274] [0.917] [0.898] [0.428] [0.966] [0.966] N 633 633 633 632 633 631 633 632 R-squared 0.245 0.171 0.200 0.260 0.277 0.236 0.209 0.238 Note: Effects are measured in standard deviations. Regressions of panel A include the following control variables: experience, contract teacher, teacher career level, sex, and age. All regressions include UGEL fixed effects. Standard errors clustered at the school level are reported in parentheses and adjusted p-values for multiple hypotheses testing are reported in brackets. We calculate the adjusted p-values using the stepdown method of Romano and Wolf (2016). Quantitative Economics 16 (2025) Can teaching be taught? 223 Table 8. Results on student learning after 1 and 3 years of coaching. Mathematics Reading 1 Year 3 Years Combined 1 Year 3 Years Combined OLS OLS IV IV OLS OLS IV IV (1) (2) (3) (4) (5) (6) (7) (8) Treatment 0.106 0.114 0.075 0.100 (0.034) (0.033) (0.032) (0.031) [0.002] [0.001] [0.019] [0.001] Cumulative years treated 0. 030 0.107 0.027 0.076 (0.009) (0.034) (0.008) (0.032) [0.001] [0.002] [0.001] [0.017] Cumulative year treated × year 3 dummy variable −0.075 −0.049 (0.032) (0.029) [0.018] [0.095] Sum of above two rows 0.032 0.027 (0.009) (0.008) [0.000] [0.001] Coefficient on random assignment in first-stage regression 3.739 3.739 (0.048) (0.048) F-statistic (for cumulative years treated) 6123 5315 6127 5316 F-statistic (for cumulative) years treated ×year 3) 3854 3856 Control Mean 0.003 0.004 0.004 0.004 0.004 0.003 0.003 0.004 Observations 22,198 18,261 18,261 40,459 22,199 18,275 18,275 40,474 Schools 2547 2053 2053 2547 2547 2053 2053 2547 R20.142 0.182 0.184 0.143 0.162 0.168 0.169 0.153 Note: This table shows treatment effects of the coaching program on standardized student test scores. Columns 1 and 5 show the ITT effects after 1 year of treatment in 2016, while columns 2 and 6 show the ITT effects after 3 years of treatment in 2018. Columns 3 and 7 present 2SLS estimates of ACR using the random treatment assignment as an instrument for the total coaching years to which students were exposed through their teachers over the course of 3 years. Finally, columns 4 and 8 combined the IV regressions for years 1 and 3 (because of almost perfect compliance in year 1, IV and OLS estimates are almost identical); see the text for how to interpret the coefficients for these regressions. All specifications include school district (UGEL) fixed effects and control for school size (number of teachers and students), which is not balanced at baseline (See Table A3 for additional specifications). All results use standardized exam scores and can be interpreted as standard deviations. Regressions are run at the student level, with robust standard errors, clustered by school, presented in parentheses, and p-values shown in brackets. (1) and (5) show estimates of ATEstud(1)after 1 year of implementation, columns (2), (3), (6), and (7) show ITT and ACR estimates after 3 years of the program, and columns (4) and (8) present combined ACR results for years 1 and 3. While the program was designed by the Ministry of Education, it was implemented by each local school district (UGEL),27 so our preferred specification, shown in this table, includes school district fixed effects, which also control for any differences in actual program implementation 27Peru’s 225 school districts (UGELs) are managed by school boards, which implement education policies in their districts. Each UGEL is overseen by its Regional Education Board (Dirección de Educación Regional). 224 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) within each region. All Table 8regressions also control for school size (number of teachers and students), which was slightly unbalanced at baseline. We cluster standard errors at the school level in all regressions, following Abadie, Athey, Imbens, and Wooldridge (2023), since the treatment is assigned at the school level. The APM coaching program has significantly positive impacts on student learning. After 1 year, average test scores increase by 0.106 and 0.075 standard deviations (s.d.) in math and reading comprehension, respectively. These are average treatment effects, ATEstud(1), and they suggest that coaching that provides regular, individualized support to teachers can be an effective policy to increase student learning. For perspective, note that the effect after 1 year is similar in magnitude to the median effect on learning outcomes of 234 education studies in low and middle income countries reviewed by Evans and Yuan (2022). And when compared to the median for large studies (those with over 5000 students), the effect of the APM program after only 1 year is almost double that median effect (0.05 s.d.). Table A3 in Castro et al. (2024a) shows how estimates change when using regional, rather than school district, fixed effects, and when excluding controls. Both of those changes reduce the size of the coefficient slightly, but the results are generally robust to these changes.28 Table A3 includes another specification, column (4), that adds to the analysis the panel data available from 2010 to 2018 and adds school-level fixed effects and state-specific time trends, without any controls; its results are very close to those of main OLS specification in Table 8. 5.2.2 2 results after 3years Columns (2) and (6) of Table 8show the effects of the APM program in 2018, after 3 years. Recall that in 2018 the standardized test is for grade 4, so that, except for repeaters, we follow the same students observed in 2016 in grade 2 after 2 more years of exposure to APM. The estimated program effects, which are now ITT effects (ITTstud(3)) and so are lower bounds for ATE (ATEstud(3)), remain positive after 3 years of the program and are slightly higher (than the estimates after 1 year shown in columns (1) and (5)): 0.114 s.d. for math, 0.100 s.d. for reading comprehension. These ITT results show the average effect on students learning after 3 years for schools that were randomly assigned to the APM program in 2016. Yet the exposure of students to treated teachers and, therefore, the effective treatment dose, differs widely among APM schools as a result of teacher turnover. To estimate the impact on students of being exposed to 1 more year of teacher coaching, we use random assignment in 2016 to instrument students’ exposure to coached teachers in each school. We have data on teachers’ school assignments, so we constructed a variable that captures the intensity of coaching for the teachers present in each year (since the program started) in a given school. This incorporates the coaching history of all teachers that the students had over thecourseof3years. 29 28The exception is reading comprehension scores after 1 year of APM; they are significant only if controls are included. Yet the treatment effects after three years are robust even when excluding controls for both subjects. 29Strictly speaking, we construct and average “history” over all teachers in a given school in a given year, since we cannot match students to individual teachers. Note, however, that 20% of the schools in our student test score sample had only one teacher, so for these schools we are matching students to their specific teacher. Quantitative Economics 16 (2025) Can teaching be taught? 231 Castro, Juan F., Paul Glewwe, Alexandra Heredia-Mayo, Stephanie Majerowicz, and Ricardo Montero (2024a), “Supplement to ‘Can teaching be taught? Improving teachers’ pedagogical skills at scale in rural Peru’.” Quantitative Economics Supplemental Material, 15, https://doi.org/10.3982/QE2079.[0190,0193,0195,0196,0200,0204,0207,0209, 0213,0214,0216,0217,0219,0221,0224] Castro, Juan F., Paul Glewwe, Alexandra Heredia-Mayo, Stephanie Majerowicz, and Ricardo Montero (2024b), “Replication package for: Can teaching be taught? Improving teachers’ pedagogical skills at scale in rural Peru.” Zenodo. https://doi.org/10.5281/ zenodo.13738582. [0210] Chetty, Raj, John Friedman, and Jonah Rockoff (2014), “Measuring the impacts of teachers I: Evaluating bias in teacher value-added estimates.” American Economic Review, 104 (9), 2593–2632. [0186] Cilliers, Jacobus, Brahm Fleisch, Cas Prinsloo, and Stephen Taylor (2020), “How to improve teaching practice? An experimental comparison of centralized training and inclassroom coaching.” Journal of Human Resources, 55 (3), 926–962. [0187,0189,0190] Clotfelter, Charles, Helen Ladd, and Jacob Vigdor (2010), “Teacher credentials and student achievement in high school: A cross subject analysis with fixed effects.” Journal of Human Resources, 45 (3), 655–681. [0186] Das, Jishnu, Stefan Dercon, James Habyarimana, and Pramila Krishnan (2007), “Teacher shocks and student learning. Evidence from Zambia.” Journal of Human Resources,42 (4), 820–862. [0186] de Ree, Joppe, Karthik Muralidharan, Menno Pradhan, and Halsey Rogers (2018), “Double for nothing? Experimental evidence on an unconditional teacher salary increase in Indonesia.” Quarterly Journal of Economics, 133 (2), 993–1039. [0229] Evans, David and Anna Popova (2016), “What really works to improve learning in developing countries? An analysis of divergent findings in systematic reviews.” World Bank Research Observer, 31 (2), 242–270. [0186,0189] Evans, David and Fei Yuan (2022), “How big are effect sizes in international education studies?” Educational Evaluation and Policy Analysis, 44 (3), 532–540. [0189,0224] Fauth, Benjamin, Jasmin Decristan, Anna-Theresia Decker, Gerhard Büttner, Ilonca Hardy, Eckhard Klieme, and Mareike Kunter (2019), “The effects of teacher competence on student outcomes in elementary science education: The mediating role of teaching quality.” Teaching and Teacher Education, 86, 102882. [0196] Gage, Nicholas, Terrance Scott, Regina Hirn, and Ashley MacSuga-Gag (2018), “The relationship between teachers’ implementation of classroom management practices and student behavior in elementary school.” Behavioral Disorders, 43 (2), 302–315. https: //doi.org/10.1177/0198742917714809. [0196] Georgiadis, Andreas and Christos Pitelis (2016), “The impact of employees’ and managers’ training on the performance of small-and medium-sized enterprises: Evidence 232 Castro, Glewwe, Heredia-Mayo, Majerowicz, and MonteroQuantitative Economics 16 (2025) from a randomized natural experiment in the UK service sector.” British Journal of Industrial Relations, 54 (2), 409–421. [0187] Jukes, Matthew, Elizabeth Turner, Margaret Dubeck, Katherine Halliday, Hellen Inyega, Sharon Wolf, Stephanie Simmons Zuilkowski, and Simon Brooker (2017), “Improving literacy instruction in Kenya through teacher professional development and text messages support: A cluster randomized trial.” Journal of Research on Educational Effectiveness,10 (3), 449–481. [0187] Kotze, Janeli, Brahm Fleisch, and Stephen Taylor (2019), “Alternative forms of early grade instructional coaching: Emerging evidence from field experiments in South Africa.” International Journal of Educational Development, 66, 203–213. [0190] Kovner, Christine, Carol Brewer, Farida Fatehi, and Jin Jun (2014), “What does nurse turnover rate mean and what is the rate?” Policy, Politics, & Nursing Practice, 15 (3–4), 64–71. [0187] Kraft, Matthew, David Blazar, and Dylan Hogan (2018), “The effect of teacher coaching on instruction and achievement: A meta-analysis of the causal evidence.” Review of Educational Research, 88 (4), 547–588. [0189,0190] Loyalka, Prashant, Anna Popova, Guirong Li, Chengfang Liu, and Henry Shi (2019), “Does teacher training actually work? Evidence from a large-scale randomized evaluation of a national teacher training program.” American Economic Journal: Applied Economics, 11 (3), 128–154. [0186] Lucas, Adrienne, Patrick McEwan, Moses Ngware, and Moses Oketch (2014), “Improving early-grade literacy in East Africa: Experimental evidence from Kenya and Uganda.” Journal of Policy Analysis and Management, 33 (4), 950–976. [0187] Matsumura, Lindsay, Hellen Garnier, Richard Correnti, Brian Junker, and Donna DiPrima Bickel (2010), “Investigating the effectiveness of a comprehensive literacy coaching program in schools with high teacher mobility.” The Elementary School Journal, 111 (1), 35–62. [0189] Ministry of Education (2019a), Administrative Files on School Characteristics, Teacher Characteristics and Student Outcomes [database]. Ministry of Education of Peru, Lima, last accessed 2019-08-01. [0210] Ministry of Education (2019b), Monitoring of School Practices [database]. Office of Strategic Monitoring and Evaluation. Ministry of Education of Peru, Lima, last accessed 2019-10-17. [0210] Popova, Anna, David Evans, and Violeta Arancibia (2016), “Training teachers on the job: What works and how to measure it.” Policy Research Working Paper 7834. The World Bank: Washington, DC. [0186] Romano, Joseph and Michael Wolf (2016), “Efficient computation of adjusted p-values for resampling-based stepdown multiple testing.” Statistics & Probability Letters, 113, 38–40. [0221,0222] Quantitative Economics 16 (2025) Can teaching be taught? 233 Schaffner, Julie, Paul Glewwe, and Uttam Sharma (2024), “Why programs fail: Lessons for improving public service quality from a mixed methods evaluation of an unsuccessful teacher training program.” Forthcoming, World Bank Economic Review.[0188] Stronge, James, Thomas Ward, and Leslie Gran (2011), “What makes good teachers good? A cross-case analysis of the connection between teacher effectiveness and student achievement.” Journal of Teacher Education, 62 (4), 339–355. https://doi.org/10. 1177/0022487111404241. [0196] Wisniewski, Benedikt, Klaus Zierer, and John Hattie (2020), “The power of feedback revisited: A meta-analysis of educational feedback research.” Frontiers in Psychology, 10, 3087. [0196] World Bank (2018), World Development Report: Learning to Realize Education’s Promise. The World Bank, Washington, DC. [0186] Zeitlin, Andrew (2021), “Teacher turnover in Rwanda.” Journal of African Economies,30 (1), 81–102. [0188] Co-editor Garance Genicot handled this manuscript. Manuscript received 27 January, 2022; final version accepted 4 November, 2024; available online 14 November, 2024. The replication package for this paper is available at https://doi.org/10.5281/zenodo.13738582. The Journal checked the data and codes included in the package for their ability to reproduce the results in the paper and approved online appendices.