scieee AI-readable full text Open interactive document viewer

Surrogate outcomes and transportability

Tikka, Santtu,Karvanen, Juha

Full text

This is a self-archived version of an original article. This version may differ from the original in pagination and typographic details. Author(s): Title: Year: Version: Copyright: Rights: Rights url: Please cite the original version: CC BY-NC-ND 4.0 https://creativecommons.org/licenses/by-nc-nd/4.0/ Surrogate outcomes and transportability © 2019 Elsevier Inc. Accepted version (Final draft) Tikka, Santtu; Karvanen, Juha Tikka, S., & Karvanen, J. (2019). Surrogate outcomes and transportability. International Journal of Approximate Reasoning, 108. https://doi.org/10.1016/j.ijar.2019.02.007 2019 Surrogate Outcomes and Transportability S. Tikka∗, J. Karvanen Department of Mathematics and Statistics, University of Jyvaskyla, P.O. Box 35 (MaD) FI-40014, Finland Abstract Identification of causal effects is one of the most fundamental tasks of causal inference. We consider an identifiability problem where some experimental and observational data are available but neither data alone is sufficient for the identification of the causal effect of interest. Instead of the outcome of interest, surrogate outcomes are measured in the experiments. This problem is a generalization of identifiability using surrogate experiments [1] and we label it as surrogate outcome identifiability. We show that the concept of transportability [2] provides a sufficient criteria for determining surrogate outcome identifiability for a large class of queries. Keywords: Causality, do-calculus, Experiment, Graph, Identifiability, Mediator. 1. Introduction In the formal framework of causal inference it is sometimes possible to make experimental claims using observational data alone. First, we construct a causal model by encoding our knowledge into a graph and specify a probability distribution over the observed variables. An experiment can now be carried out symbolically in the model through an intervention, which is an action that forces variables to take specific values irrespective of the mechanism that would determine their values otherwise. The question is whether the observed probability distribution alone is enough to determine the effect of this intervention. This problem, known as the identifiability problem, has been studied extensively in literature and solutions in the form of graphical criteria [3, 4] as well as algorithms have been proposed [5, 6, 7]. Various extensions to the identifiability problem have emerged in recent years. These include concepts such as transportability, where identifiability is considered in a target population, but information for the task is available from multiple source populations [8, 9]. ∗Corresponding author. Email addresses: [email protected] (S. Tikka), [email protected] (J. Karvanen) The presence of unobserved confounders often renders causal effects of interest non-identifiable from observational data alone. This leads us to ask whether experimental data can be of use in the identification task. The concept of surrogate experiments or z-identifiability considers this problem in a setting where in addition to the observed probability distribution, experimentation is allowed on a set of variables that is disjoint from the interventions of the target causal effect [1] and the experimental distribution of these surrogate experiments is available over all variables. By experimental distribution we mean a distribution of a set of outcomes variables when some variables have been intervened on. We consider a more general problem than z-identifiability: instead of assuming that a experimental distribution over all variables is available, we assume that a collection of experimental distributions is available where every variable has not necessarily been observed. This kind of setting can occur for example in mediation analysis, where we have previously performed an experiment where the mediator was the outcome variable. Another example is a setting where we are interested in two outcome variables but have only measured one of them in a previous experiment. In a practical study we usually have access to information about population characteristics when performing an experiment. Sometimes not all of these characteristics are be measured in conjunction with the experiment itself which leads to incomplete knowledge regarding the experimental distribution. Suppose that we are interested in the experimental distribution of another variable, one that was not measured during the experiment. The question is whether this distribution can be obtained from the observational data and the outcome of the previous experiment, which we refer to as the surrogate outcome. We label this generalization of identifiability as surrogate outcome identifiability. Remarkably, a connection can be drawn between surrogate outcome identifiability and transportability. Transportability is concerned with identifiability across conceptual domains where both observational and experimental data are available from each domain. In practical terms, a domain can be for example a city, and data from multiple domains in this case could be for example the age distributions of the populations of these cities. Naturally, discrepancies between causal mechanisms can arise between domains, which has to be taken into account in the causal modeling framework. Typically, we are interested in the effect of an intervention in a single domain, known as the target domain, and the domains providing additional information for the task are known as source domains. However, existing methods for determining transportability only allow a single experiment to take place within a single domain, whereas our surrogate identifiability is concerned with multiple distributions from differing experiments in a single domain. We incorporate the framework of transportability by depicting each available experiment of the surrogate outcome problem as a source domain of a transportability problem with the same experiments. An introductory example illustrates the difference between surrogate outcome identifiability and z-identifiability. We are interested in the causal effect of X1and X2on Y1and Y2in the graph of Fig. 1, which is easily determined to be non-identifiable from the joint distribution P(v) alone for example via the ap2 plication of the ID algorithm [6, 12]. Suppose now that two surrogate outcomes were measured in previous experiments providing us with two experimental distributions, P(y2|do(x2), x1, z, w) and P(y1|do(x1), z, w). The availability of these two distributions cannot be represented as z-identifiability problem, since they are conditional causal effects and they have common interventions with the target causal effect. We cannot directly regard this problem as a transportability problem either, since we are concerned with only a single domain. The causal effect can now be identified with the help of the two experimental distributions, which we will show later in Section 3. X1 X2 Z Y1Y2 W Figure 1: A graph where the causal effect P(y1, y2|do(x1, x2)) is not identifiable from P(v) alone. In this paper we propose a way to transform a surrogate outcome problem into a transportability problem. We show that the identifiability of the transformed problem is a sufficient condition for identifiability of the surrogate outcome problem. We derive an identifiability algorithm for surrogate outcome problems and implement it as a part of the R package causaleffect [11, 12]. 2. Notation and definitions We assume that the reader is familiar with graph theoretic concepts fundamental to causal inference and refer them to works such as [13]. We use capital letters to denote vertices and the respective variables and small letters to denote their values. We sometimes write singleton sets {X}as Xfor clarity. A directed graph with a vertex set Vand an edge set Eis denoted by (V, E). For a graph G= (V, E) and a set of vertices W⊆Vthe sets Pa(W)G,Ch(W)G,An(W)G and De(W)Gdenote a set that contains Win addition to its parents, children, ancestors and descendants in G, respectively. A subgraph of a graph G= (V, E) induced by a set of vertices W⊂Vis denoted by G[W]. This subgraph retains all edges Vi→Vjof Gsuch that Vi, Vj∈W. The graph obtained from G by removing all incoming edges of Xand all outgoing edges of Zis written as G[X, Z]. A back-door path from Xto Yis a path with an edge incoming to Xand Y. A topological ordering ϕof Gis an ordering of its vertices in which 3 every node is smaller than its descendants in G. The set of vertices smaller than a vertex Viin ϕis denoted by V(i) ϕ. To facilitate analysis of identifiability and the generalization to surrogate outcomes, we must first define the probabilistic causal model [4]. Definition 2.1 (Probabilistic causal model).Aprobabilistic causal model is a quadruple M= (U, V, F, P(u)), where Uis a set of unobserved (exogenous) variables that are determined by factors outside the model, Vis a set {V1, . . . , Vn}of observed (endogenous) variables that are determined by variables in U∪V.Fis a set of functions {fV1, . . . , fVn}such that each fViis a mapping from (the respective domains of) U∪(V\ {Vi})to Vi, and such that the entire set Fforms a mapping from Uto V, and P(u)is a joint probability distribution of the variables in the set U. Each causal model induces a graph through the following construction: A vertex is added for each variable in U∪Vand a directed edge from Vi∈U∪V into Vj∈Vwhenever fVjis defined in terms of Vi. Conventionally, causal inference focuses on a sub-class of models with additional assumptions: each Ui∈U appears in at most two functions of F, the variables in Uare mutually independent and the induced graph of the model is acyclic. Models that satisfy these additional assumptions are called semi-Markovian causal models. The induced graph of a semi-Markovian model is called a semi-Markovian graph. In semiMarkovian graphs every Ui∈Uhas at most two children. In semi-Markovian models it is common not to depict background variables in the induced graph explicitly. Unobserved variables Ui∈Uwith exactly two children are not denoted as Vj←Ui→Vkbut as a bidirected edge Vj↔Vkinstead. Furthermore, unobserved variables with only one or no children are omitted entirely. We also adopt these abbreviations. For semi-Markovian graphs the sets Pa(·)G,Ch(·)G,An(·)G and De(·)Gcontain only observed vertices. Additionally, a subgraph G[W] of a semi-Markovian graph Gretains any bidirected edges between vertices in W. A graph induced by a probabilistic causal model also encodes conditional independences among the variables in the model through a concept known as d-separation. We use the definition in [14] which explicitly accounts for the presence of bidirected edges making it suitable for semi-Markovian graphs. Definition 2.2 (d-separation).A path Pin a semi-Markovian graph Gis said to be d-separated by a set Zif and only if either Pcontains one of the following three patterns of edges: I→M→J,I↔M→Jor I←M→J, such that M∈Z, or Pcontains one of the following three patterns of edges: I→M←J, I↔M←J,I↔M↔J, such that De(M)G∩Z=∅. Disjoint sets Xand Y are said to be d-separated by Zin Gif every path from Xto Yis d-separated by Zin G. Since we are dealing entirely with semi-Markovian graphs, we will henceforth refer to them simply as graphs. If no conditional independence statements other 4 than those already encoded in the graph are implied by the distribution of the variables in the model, we say that the distribution is faithful [10]. A causal model allows us to manipulate the functional relationships encoded in the set F. An intervention do(x) on a model Mforces Xto take the specified value x. The intervention also creates a new sub-model, denoted by Mx, where the functions in Fthat determine the value of Xhave been replaced with constant functions. The interventional distribution of a set of variables Yin the model Mxis denoted by P(y|do(x)). This distribution is also known as the causal effect of Xon Y. Three inference rules known as do-calculus [3] provide the means for manipulating interventional distributions. 1. Insertion and deletion of observations: P(y|do(x), z, w) = P(y|do(x), w),if (Y⊥⊥ Z|X, W)G[X]. 2. Exchange of actions and observations: P(y|do(x, z), w) = P(y|do(x), z, w),if (Y⊥⊥ Z|X, W)G[X,Z]. 3. Insertion and deletion of actions P(y|do(x, z), w) = P(y|do(x), w),if (Y⊥⊥ Z|X, W)G[X,Z(W)], where Z(W) = Z\An(W)G[X]. Regarding the identifiability problem, the goal is to transform P(y|do(x)) into an expression that does not contain the do-operator using do-calculus. A causal effect that admits this transformation is called identifiable, which is formally defined in e.g. [6]. Do-calculus has been shown to be complete with respect to the identifiability problem [6, 5] as well as the transportability and z-identifiability problems [2, 1]. Special graphs known as c-components (confounded components) are crucial for causal effect identification [6]. Definition 2.3 (c-component).Let G= (V, E)be a graph. A c-component C= (VC, EC)(of G) is a subgraph of Gsuch that every pair of vertices in Cis connected via a bidirected path (a path consisting entirely of bidirected edges). A c-component Cis maximal if there are no vertices in VCthat are connected to V\VCin Gvia bidirected paths and Cis an induced subgraph of G. The joint distribution of a causal model admits the so-called c-component factorization with respect to the set of maximal c-components of the induced graph Gof the model, denoted by C(G). Henceforth we will use the term c-component to refer to maximal c-components for brevity. If in addition to the joint observed probability distribution P(v) experimentation is allowed on a set Z, the identifiability problem is known as zidentifiability [1]. The set Zis known as the set of surrogate experiments. 5 Definition 2.4 (z-identifiability).Let G= (V, E)be a graph and let X,Yand Zbe disjoint sets of variables such that X, Y, Z ⊂V. The causal effect of Xon Yis said to be z-identifiable from Pin Gif P(y|do(x)) is uniquely computable from P(v)together with the interventional distributions P(v\z0|do(z0)), for all Z0⊆Z, in any model that induces G. As an example of a z-identifiable causal effect, we consider the identification of P(y|do(x)) from P(x, y, z, w) and P(x, y, w |do(z)) in the graph of Fig. 2. This effect is not identifiable without the experimental distribution, which can be verified for example by using the ID algorithm of [6]. W Z XY Figure 2: A graph where the causal effect of Xon Yis z-identifiable using experiments on Z. We derive the effect using do-calculus: P(y|do(x)) = X w P(y|do(x), w)P(w|do(x)) =X w P(y|do(z, x), w)P(w|do(x)) =X w P(y|do(z, x), w)P(w) =X w P(y|do(z), x, w)P(w) where the second equality follows from the third rule of do-calculus, since (Y⊥⊥ Z|X)G[X,Z]. The third equality follows from the third rule of do-calculus, since (W⊥⊥ X)G[X]and the fourth equality follows from the second rule of docalculus, since (Y⊥⊥ X|W)G[Z,X]. The term P(y|do(z), x) is identifiable from P(x, y, w |do(z)) via marginalization and conditioning and P(w) is identifiable from P(x, y, z, w) via marginalization. The available information in a z-identifiability problem consists of a single observational distribution and experimental distributions resulting from interventions on subsets of Z. Our goal is to extend this problem to a setting where experimentation is allowed on the subsets of multiple surrogate experiments. Furthermore, we do not require that the distribution of the entire set Vis known under these experiments or that the experiments have to be disjoint from X, the intervention in the target causal effect. We formalize these notions in the following definition. Definition 2.5 (Surrogate outcome query).Asurrogate outcome query is a quadruple (X, Y, G, S), where G= (V, E)is a graph, X, Y ⊂Vare disjoint sets 6 of variables. The set of surrogate outcomes S={(Z1, W1),...,(Zn, Wn)}is a collection of intervention–outcome pairs (Zi, Wi)such that for all i= 1, . . . , n it holds that Wi⊂De(Zi)G\Zi,Zi⊂An(Wi)G\Wi, De(Wi)G∩Zi=∅and An(Wij )G\Wi=An(Wi)G\Wifor each Wij ∈Wi. While requirements for the sets Ziand Wimay appear complicated, they are only a formal statement of the fact that we require all variables subject to experimentation to precede all of the outcome variables in the causal order. We also assume that outcomes in a single intervention–outcome pair have the same ancestors. This assumption is made for technical reasons and outcomes with different ancestry can still be represented through separate intervention–outcome pairs. The intuition behind these assumptions is that each intervention–outcome pair should correspond to a single experiment where every manipulated variable has a potential effect on the outcomes. For example, in the graph of Fig. 3, we would not consider ({Z1, Z2},{W1, W2}) to be a valid intervention–outcome pair, since manipulating Z2cannot affect W1. Z1W1Z2W2 Figure 3: An example graph on the proper form of intervention–outcome pairs. Identifiability of a causal effect defined by a surrogate outcome query is characterized by the following definition. Definition 2.6 (Surrogate outcome identifiability).Let (X, Y, G, S)be a surrogate outcome query. Let Ii=∪Z0⊆ZiP(wi|do(z0),An(wi)G[Z0]\(w∪z0)), where (Zi, Wi)∈ S, and let I=∪n i=1Ii∪P(v). Then the causal effect of Xon Yis said to be surrogate outcome identifiable from Iin Gif P(y|do(x)) is uniquely computable from Iin any model that induces G. The precise formulation of the sets Iiand the experimental distributions is needed to closely connect surrogate outcome identifiability to transportability as we will show later in Section 3. While the assumption that the interventional distributions are always available for every subset Z0of every Ziis technical, it can have a real-world interpretation as well. For example, it is realistic to assume that when the joint effect of two medical treatments is studied, either the effect of each individual treatment is already known or they can be estimated from the same experiment. In many cases, it may be unethical to test for the joint effect if it is not known that the individual treatments are safe and efficient. As an example on surrogate outcome identifiability, we consider the graphs of Fig. 4 and attempt to identify the causal effect of Xon Yfrom P(x, y, z) and P(z|do(x)). This corresponds to setting S={(X, Z)}in Definition 2.6. It should be noted that this problem cannot be expressed as a z-identifiability problem, since the experimental distribution that is available contains an intervention on Xand it is not a full experimental distribution over the variables X, Y and Z, but is instead restricted to Zonly. 7 X Z Y (a) XZY (b) Figure 4: Graphs where the causal effect P(y|do(x)) is not identifiable from P(x, y, z) alone, but is identifiable via surrogate outcomes using P(z|do(x)). We can derive the effect as follows in both Fig. 4(a) and 4(b): P(y|do(x)) = X z P(y|do(x), z)P(z|do(x)) =X z P(y|x, z)P(z|do(x)). Both terms in this expression are computable from I: the term P(y|x, z) can be obtained via conditioning from P(x, y, z) and the term P(z|do(x)) is already included in I. Here the second equality follows from the second rule of docalculus, since (Y⊥⊥ X|Z)G[X]. In this trivial example we can easily determine the correct sequence of applications of do-calculus to reach the desired expression. In general, it is difficult to find such a sequence or determine whether such a sequence even exists. For tasks such as identifiability, the solution was to construct an algorithm that either derives the expression for the effect, or returns a graph structure that can be used to construct two models where the distributions over the observed variables agree, but the interventional distributions differ. Instead of developing a similar algorithm for surrogate outcome identifiability, we will describe this problem as a transportability problem, for which a complete solution already exists in the form of an algorithm [15]. 3. Identifying surrogate outcome queries using transportability In order to describe the connection between surrogate outcomes and transportability we first provide the definition of a transportability diagram. Definition 3.1 (Transportability diagram).Let (M, M∗)be a pair of probabilistic causal models relative to domains (π, π∗), sharing a graph G. The pair (M, M∗)is said to induce a transportability diagram Dif Dis constructed as follows: every edge in Gis also an edge in D,Dcontains an extra edge Ti→Vi whenever there might exist a discrepancy fVi6=f∗ Vior P(ui)6=P∗(ui)between Mand M∗. In the above definition, a domain is simply a formalization of the intuitive notion of different contexts of the same phenomena. The domains serve as indices to differentiate between the different causal models that are depicted by the same graph Gand to associate the available observational and experimental 8 Since An(Z, W )G={Z, W}, line 2 and then line 1 are triggered for the last term, which is simply P∗(z, w). The recursive calls for the first two terms both trigger line 2 due to Y1not being an ancestor of Y2and {X2, Y2}not being ancestors of Y1. After these calls we have P∗(y1, y2|do(x1, x2)) =X w,z P∗(y2|do(x1, x2, z, w))P∗(y1|do(x1, z, w))P∗(z, w). Line 10 is triggered next for both of the first two terms because (Y1⊥⊥ {T(1) X1, T(1) Y2} | {X1, Z, W, X2})D(1)[X1,Z,W,X2]. and (Y2⊥⊥ T(2) X2| {X1, Z, W })D(2)[X1,Z,W ]. This means that intervention on X1is activated for the first term and intervention on X2is activated for the second terms. Finally, line 7 is triggered for both remaining terms and we obtain P∗(y1, y2|do(x1, x2)) = X w,z P(1)(y2|do(x2), x1, z, w))P(2)(y1|do(x1), z, w)P∗(z, w), as the final expression. We obtain a solution for the original surrogate outcome problem by simply omitting the domain indicators from this expression P(y1, y2|do(x1, x2)) = X w,z P(y2|do(x2), x1, z, w))P(y1|do(x1), z, w)P∗(z, w). We can also derive the effect using the causaleffect R package with the following commands: library(causaleffect) library(igraph) > fig1 <- graph.formula(x_1 -+ y_2, x_1 -+ y_1, w -+ y_1, w -+ y_2, + z -+ y_1, x_2 -+ y_2, z -+ y_2, z -+ x_2, w -+ z, z -+ w, + z -+ x_2, x_2 -+ z, y_1 -+ x_1, x_1 -+ y_1, simplify = FALSE) > fig1 <- set.edge.attribute(fig1, "description", 9:14, "U") > s1 <- list( + list(Z = c("x_2"), W = c("y_2")), + list(Z = c("x_1"), W = c("y_1")) > ) > cat(surrogate.outcome(y = c("y_1", "y_2"), x = c("x_1","x_2"), + S = s1, G = fig1)) \sum_{w,z}P_{x_2}(y_2|x_1,w,z)P(w,z)P_{x_1}(y_1|w,z) The package uses the notation Px1(y1|w, z) to denote P(y1|do(x1), w, z). In the next section we prove the correctness of TRSO and show that the omission of the domain indicators from its output always produces a valid expression for the original surrogate outcome identifiable causal effect. 15 4. Correctness of the modified transportability algorithm First, we recall that do-calculus is complete with respect to transportability and prove some useful lemmas. Theorem 4.1 (do-calculus characterization).The rules of do-calculus together with standard probability manipulations are complete for establishing transportability of causal effects. Proof. See [15]. Theorem 4.1 shows that a sequence of valid operations necessarily exists for a transportable causal effect. We define this sequence explicitly. Definition 4.1 (do-calculus sequence).Let Gbe a graph or a transportability diagram, let pbe an identifiable or transportable causal effect and let Ibe a set of available information. A do-calculus sequence for pin Gis a pair δp= (R,P), where Ris an n−tuple (R1, . . . , Rn)such that each Riis either a member of the index set {m, c, r}or a quintuple (Yi, Zi, Xi, Wi, ri)such that (Yi⊥⊥ Zi|Xi, Wi)G0and G0=G[Xi]if ri= 1,G0=G[Xi, Zi]if ri= 2 and G0=G[X, Zi(Wi)] if ri= 3 and Zi(Wi) = Zi\An(Wi)G[X].P= (p1, . . . , pn)is a sequence of probability distributions such that if Riis of the first type described above, piis obtained from pi−1via marginalization for Ri=m, conditioning for Ri=cand the chain-rule if Ri=r. If Riis of the second type, then pi is obtained from pi−1using rule number riof do-calculus licensed by the sets Xi, Yi, Ziand Wi. Furthermore, when p0=pis transformed as dictated by the sequence δp, an expression pnis obtained such that each term that appears in pnis a member of Ior computable from Iwithout do-calculus. The idea is to use a do-calculus sequence of a transportable causal effect to construct a do-calculus sequence for its query transformation counterpart. However, as do-calculus statements stem from d-separation in the underlying graph, we must first establish that d-separation is invariant to the presence of transportability nodes. Lemma 4.1. Let Dbe a transportability diagram and let Gbe its subgraph obtained by removing all transportability nodes Tfrom D. Let X, Y, Z be disjoint sets of vertices of Dsuch that they do not contain transportability nodes. Then Xand Yare d-separated by Z∪T0for every T0⊆Tin Dif and only if Xand Yare d-separated by Zin G. Proof. (i) Suppose that Xand Yare d-separated by Z∪T0in D. By assumption X, Y and Zdo not contain any transportability nodes. This means that no path from Xto Ycontains transportability nodes, since a path containing such a node would necessarily have it as one of the path’s endpoints by Definition 3.1. Furthermore, a transportability node cannot be a descendant of a collider by 16 definition. Thus all paths from Xto Yremain d-separated if we remove all transportability nodes from D. (ii) Suppose that Xand Yare d-separated by Zin G. Adding transportability nodes to Gcannot create any new paths between Xand Ysince a transportability node is only connected to other vertices of the graph through a single vertex. As before, a transportability node cannot be a descendant of a collider by definition. Thus all paths between Xand Yare d-separated by Z∪T0in D for any subset T0⊆T. Corollary 4.1. Let QSbe a surrogate outcome query with a graph Gand let QTbe its query transformation with a collection of transportability diagrams D. Then any conditional independence statement (Y⊥⊥ X|Z)Dithat holds in some transportability diagram Diof Dalso holds in Gif the sets Yand Xdo not contain transportability nodes. Conversely, every conditional independence statement of Gholds in every diagram of D. Proof. The proof immediately follows from Theorem 4.1 by noting that Gcan be obtained from each element of Dby removing all transportability nodes. We show that there always exists do-calculus sequence such that every operation that manipulates transportability nodes does not manipulate any other vertices at the same time. Lemma 4.2. Let (X, Y, D, D∗,Π, π∗,Z, Z∗)be a transportability query and let Tbe the set of all transportability nodes over the domains of Dand the target diagram D∗. If p=P(y|do(x)) is a transportable causal effect with transportability information Iof Definition 3.3, then there exists a do-calculus sequence dp= (R,P)such that whenever Riis of the form (Yi, Zi, Xi, Wi, ri)then either Zi∩T=∅or Zi⊂T. Proof. Let d0 p= (R0,P0) be any do-calculus sequence for p. It suffices to consider R0 i∈ R0of the form (Yi, Zi, Xi, Wi, ri). If ri∈ {2,3}there is nothing to prove, since the second and third rules of do-calculus manipulate interventions which are not allowed on transportability nodes. Let ri= 1 and suppose that Zi\T6= ∅. Then from Definition 4.1 we know that (Yi⊥⊥ Zi|Xi, Wi)G[X]which implies that (Yi⊥⊥ Zi\T|Xi, Wi)G[X]and (Yi⊥⊥ Zi∩T|Xi, Wi)G[X]. Now, let R contain every member of R0except that each R0 iwith ri= 1 and Zi\T6=∅ is replaced by R1 i= (Yi, Zi∩T, Xi, Wi, ri) and R2 i= (Yi, Zi\T, Xi, Wi, ri). Similarly, let Pcontain every member of P0except that each p0 i, where the corresponding R0 ihas the aforementioned property, is replaced by p1 iand p2 i where p1 iis obtained from p0 i−1by applying the first rule of do-calculus with the sets Yi, Zi∩T, Xiand Wi, and p2 iis obtained from p1 iby applying the first rule of do-calculus with the sets Yi, Zi\T, Xiand Wi. By construction, dp= (R,P) is a do-calculus sequence for pwith the desired property. Theorem 4.2. TRSO is sound. 17 Proof. Lines 1 through 9 are identical to the original formulation of the transportability algorithm and their soundness was established in [15] with the exception that the order of lines 6 and 10 is reversed. Line 10 is different from the original. On this line we first find the c-component D[C0] of Dsuch that C⊂C0. This c-component necessarily exists since the c-components of D[V\X] are always subsets of the c-components of D. The c-component D[C0] is unique because the vertex sets of c-components of any graph are disjoint. Next we check if there is an active intervention. If there is no such intervention (I=∅), we remove the ability for experimentation entirely by setting Z0=∅. If there is an active experiment (I6=∅), we check whether it falls into the category of allowed experiments by evaluating if Pa(C0)D∩TS=∅. If it does not, the recursive call fails. If there were no active experiments (I=∅) or active experimentation is permissible (Pa(C0)D∩TS=∅), we simply continue the recursion in the c-component D[C0]. The checks for allowing experimentation do not affect the soundness of the recursive function call that follows them on line 10. This recursive call was shown to be sound in [15]. The next result characterizes an important feature of the transport formulas produced by a successful application of TRSO. Lemma 4.3. Let (X, Y, D, G, Π, π∗,Z,∅)be the query transformation of a surrogate outcome query (X, Y, G, S). If TRSO succeeds in transporting the causal effect p=P∗(y|do(x)), then for every term that appears in the expression for pof the form P(i)(c|do(z0), d0)one of the following holds: either P(i)(c|do(z0), d0) = P(i)(w∗|do(z0), x0)P∗(x0),(1) or P(i)(c|do(z0), d0) = P∗(c|An(c)G[Z0]\c),(2) or P(i)(c|do(z0), d0) = P(i)(w∗|do(z0), x0),(3) where x0=An(w∗)G[Z0]\(w∗∪z0),W∗⊂Wiand there exists a set Zisuch that Z0⊆Ziand (Zi, Wi)∈ S. Furthermore, the right-hand sides of (1), (2) and (3) are identifiable from the information set Iof Definition 2.6 when the domain indicators are omitted. We defer the proof of Lemma 4.3 to Appendix A. We are ready to prove Theorem 3.1 Proof of Theorem 3.1. Assume that p=P∗(y|do(x)) is transportable from Π to π∗in Dwith information I2by applying TRSO(y, x, P ∗(v),∅,0, G, D, ϕ). Let δd= (R,P) be a do-calculus sequence for pof the form given by Lemma 4.2. This sequence is valid since the algorithm is sound by Theorem 4.2. We can 18 categorize the do-calculus steps into two distinct types: those that do not modify the transportability nodes present in the expression, and those that do. In other words, if Tis the set of all transportability nodes over the domains D and D∗, the first category contains steps Ri= (Yi, Zi, Xi, Wi, ri) such that Zi∩T=∅. By Corollary 4.1, the conditional independence statements in a transportability diagram involving sets Ziof this type are also valid in a corresponding graph where transportability nodes have been removed. This means that if (Yi⊥⊥ Zi|X, W)Djthen (Yi⊥⊥ Zi|X, W \T)Dj[V\T]. This allows us to construct a new do-calculus sequence as follows: For any Ri∈ R of the form (Yi, Zi, Xi, Wi, ri) with Zi∩T=∅we let R0 i= (Xi, Yi, Zi, Wi\S, ri). If Zi∩T⊆T, we let R0 i=∅. For any Riin the index set m, c, r we let R0 i=Ri. Let the sequence R0now consists of those R0 ithat are non-empty. The sequence P0of distributions is constructed from q=P(y|do(x)) through the sequence of manipulations described by R0. We apply Lemma 4.3 for each term of the form P(c|do(z), d0) in the resulting formula for qsuch that there exists no pair (Zi, Wi)∈ S with Wi=C. This means that additional manipulation steps R0 iand distributions p0 iare added that correspond to the transformation of the distribution on the left-hand side to the distribution on the right-hand side in one of the conditions of Lemma 4.3. It remains to show that every term in this resulting formula for qis included in the information set I1or can be computed from it without do-calculus. Then δq= (R0,P0) gives a do-calculus sequence for q. Let qmbe the last element of the sequence P0. Any term in the transport formula pnthat involves the target domain is unaffected by do-operators since no variable is available for experimentation in the target domain by Definition 3.4. Therefore, the corresponding term in qmcan be obtained from I1since this information set includes the full observed probability distribution P(v). Any term in pnthat involves a source domain is necessarily affected by a do-operator, since the term would otherwise be identified from the target domain directly. Since Lemma 4.3 has already been applied, all such terms take the form P(i)(c|do(z0), d0). Lemma 4.3 also guarantees, that the corresponding term P(c|do(z0), d0) in qmis computable from the information set I1. The proof of Theorem 3.1 provides a construction of a do-calculus sequence for a surrogate outcome identifiable causal effect through the query transformation. In practice, we do not have to retrace the entire derivation to obtain the identifying expression. It is enough to apply Lemma 4.3 to each relevant term in the resulting expression and them replace the terms in the transport formula with their respective counterparts from the information set of the surrogate outcome query. Appendix B contains examples on this process. The following corollary describes the process of obtaining an expression for a surrogate outcome identifiable causal effect using the query transformation. Corollary 4.2. Let (X, Y, D, G, Π, π∗,Z,∅)be the query transformation of a surrogate outcome query (X, Y, G, S)and suppose that there exists a transport formula ptfor P∗(y|do(x)) given by TRSO(y, x, P ∗(v),∅,0, G, D, ϕ). Then 19 P∗(y|do(x)) is surrogate outcome identifiable and its expression is obtained from ptby manipulating every term in ptof the form P(i)(c|do(zi), d0)in accordance to Lemma 4.3 and by omitting the domain indicators. Proof. The result is a direct consequence of the construction for δqin the proof for Theorem 3.1. We illustrate the application of TRSO and Corollary 4.2 through an example. We use surrogate outcomes to identify P(y|do(x)) in graph Gof Fig. 9(a) from P(v) and P(w|do(x), a1, a2, b1, b2). By Definition 3.4, transportability nodes are added for De(X)G\ {W}={X, B1, B2, Y }and for the vertices in the same c-component as Wthat are not ancestors of Win G[X]. The vertex A1is in the same c-component as W, but since it is still an ancestor of Wwhen edges incoming to Xare removed, no transportability node is added for it. The transportability diagram of the corresponding query transformation is depicted in Fig. 9(b). X W Y A1A2 B1B2 (a) XWY A1A2 B1B2 TB1TB2 TY TX (b) Figure 9: Graphs for illustrating the application of TRSO and Corollary 4.2. (a) Graph G corresponding to the surrogate outcome identifiability problem and the target domain of its query transformation. (b) Transportability diagram obtained from a query transformation corresponding to the domain where intervention on Xis available. The application of TRSO succeeds in transporting the causal effect and produces the following expression for P∗(y|do(x)) X a1,a2,b1,b2,w P∗(y|a1, x, a2, b1, b2, w)P(1)(w|do(x), a1, a2, b1, b2)P∗(a2|a1, x) ×P∗(b1|a1, x)P∗(a1)X a0 1,x0 P∗(b2|a0 1, x0, b1)P∗(x0|a0 1)P∗(a0 1). In this case we obtain the expression for the surrogate outcome identifiable causal effect P(y|do(x)) by simply omitting all domain indicators from the expression for P∗(y|do(x)) as licensed by Corollary 4.2 X a1,a2,b1,b2,w P(y|a1, x, a2, b1, b2, w)P(w|do(x), a1, a2, b1, b2)P(a2|a1, x) ×P(b1|a1, x)P(a1)X a0 1,x0 P(b2|a0 1, x0, b1)P(x0|a0 1)P(a0 1). 20 5. Discussion We take advantage of transportability by depicting experimental data as distinct domains and by using transportability nodes to prevent certain variables from being observed under an intervention. The positioning of the transportability nodes results in the need for Lemma 4.3 to parse the output of TRSO. It may be possible to consider other variations of the query transformation, where transportability nodes are omitted from additional vertices based on dseparation in the graph or by some other criteria. In the extreme case we could operate without any connection to transportability by omitting transportability nodes and relying on do-calculus entirely, but this approach can quickly become intractable for larger graphs. Our formulation avoids this, and the output can be directly transformed into a valid formula for a surrogate outcome identifiable causal effect. The query transformation has practical importance because an implementation of TRSO is readily available in the R package causaleffect. Transportability via the query transformation of Definition 3.4 does not provide a complete characterization of surrogate outcome identifiability. As an example, we consider the graph of Fig. 10 and identifiability of P(y1, y2|do(x1, x2)) from P(v), P(y2, z |do(x2), x1, y1) and P(y1|do(x2)). X1 X2 Y1Y2 Z Figure 10: An example where the query transformation does not produce a transportable causal effect, but the causal effect of interest is surrogate outcome identifiable. We derive the effect using do-calculus: P(y1, y2|do(x1, x2)) = P(y1|do(x1, x2))P(y2|do(x1, x2), y1) =P(y1|do(x2))P(y2|do(x2), x1, y1) =P(y1|do(x2)) X z P(y2, z |do(x2), x1, y1), where the second equality follows from rules three and two by (Y1⊥⊥ X1)G[X1,X2] and (Y2⊥⊥ X1|Y1)G[X2,X1]. It is easy to verify that the query transformation of this problem is not transportable using TRSO or the original transportability algorithm in [15]. 21 We performed a simple simulation study to assess the strength of TRSO. We generated 10000 instances of random graphs with 6 vertices and random sets of surrogate outcomes. For every instance, the causal effect p(y1, y2|do(x1, x2)) was verified to be non-identifiable from P(v) alone using the ID algorithm. We used a simple exhaustive breadth-first forwards search that implements the rules of do-calculus and standard probability manipulations to confirm surrogate outcome identifiability or non-identifiability for each instance. Out of the 10000 instances 1514 were found to be surrogate outcome identifiable by the search and 1332 by TRSO which corresponds to 88 % coverage. Based on this result, TRSO seems to be able to identify most of the surrogate outcome identifiable instances. Conflict of interest statement The authors declare that they have no conflict of interest. Acknowledgments This work belongs to the thematic research area “Decision analytics utilizing causal models and multiobjective optimization” (DEMO) supported by Academy of Finland (grant number 311877). We thank the anonymous reviewers for their comments which helped to substantially improve this paper. Appendix A Proof of Lemma 4.3. Let Gdenote the graph of the original recursive call and assume without loss of generality that G=G[An(Y)G] for clarity. Let Ddenote the graph of the current recursion stage. Consider a term of the form P(i)(c|do(z0), d0) that appears in the output formula. Only line 6 of TRSO introduces permanent interventions into the expression by using the available experiments (I=Zi∩X), so it must have been triggered and it cannot be triggered again in the same recursive branch since we check that I=∅on this line. Before triggering line 6, only a combination lines 2, 3 and 4 can be triggered, corresponding to removal of non-ancestors of Y, introducing additional interventions via the third rule of do-calculus, and performing the c-component factorization, respectively. It follows that after these steps, the local distribution Pof one the recursive calls after triggering lines 2 and 4 in sequence is of the form P(An(b)G) where Bis the vertex set of some c-component of G[V\X] Line 6 is triggered next, activating an available experiment which means that there are no transportability nodes incoming to Bin G, since (B⊥⊥ Ti|X)(i) D. After this call and application of line 2, the local distribution is now of the form P(i)(An(b)G[Z0]\z0|do(z0)), where Z0=X∩Ziis the now active intervention. Since the sets Xand Ylocal to this recursive call partition Vand non-ancestors have been removed, it is 22 only possible to trigger line 1, 9 or 10 next, since we know that this call does not fail. We proceed to prove each case. Case of line 1. Then the intervention set Xis empty and the effect is identified as X V\B P(i)(An(b)G[Z0]\z0|do(z0)) = P(i)(b|do(z0)). However, since Xand Ylocal to this call still partition Vand V= An(B)G[Z]\ Z0and Xis an empty set, we have that P(i)(b|do(z0)) = P(i)(An(b)G[Z0]\z0|do(z0)), meaning that Bcontains its own ancestors in G[Z0]. If there exists a intervention– outcome pair (Zi, Wi)∈ S such that Wi=B, then the set An(wi)G[Z0]\(wi∪z0) is empty and we have that P(i)(b|do(z0)) = P(i)(wi|do(z0),An(wi)G[Z0]\(wi∪z0)), which corresponds to (3). When the domain indicator is omitted from this term, it is clearly identifiable from I. If instead there exists a intervention–outcome pair (Zi, Wi)∈ S such that for a subset W∗⊂Wiit holds that W∗⊂B, then P(i)(b|do(z0)) = P(i)(w∗|do(z0), b \w∗)P(i)(b\w∗|do(z0)). Since line 6 was triggered previously, Bcannot have any incoming transportability nodes, which means that B\W∗must be an ancestor of W∗in G[Zi] but not a descendant of Ziaccording to the construction of the transportability diagrams of a query transformation in Definition 3.4. This means that B\W∗ is an ancestor of W∗also in G[Z0] and we have that P(i)(w∗|do(z0), b \w∗)P(i)(b\w∗|do(z0)) = P(i)(w∗|do(z0), b \w∗)P(i)(b\w∗), which follows from the third rule of do-calculus since we have established that B\W∗ imust be a non-descendant of Z0. Furthermore, since B\W∗can only contain ancestors of W∗in G[Z] it follows that P(i)(w∗|do(z0), b \w∗) = P(i)(w∗|do(z0),An(w∗)G[Z0]\(w∗∪z0)).(4) Now, we obtain from Definition 2.5 that An(W∗)G[Z0]\W∗= (An(Wi)G[Z0]\Wi)∪(An(W∗)G[Z0]∩(Wi\W∗)) = (An(Wi)G[Z0]\Wi)∪(Wi\W∗). This means that the right-hand side of (4) can be obtained via conditioning by writing P(i)(wi|do(z0),An(wi)G[Z0]\(wi∪z0)) PWi\W∗P(i)(wi|do(z0),An(wi)G[Z0]\(wi∪z0)), 23 which is identifiable from Iafter omitting domain indicators, which means that (4) is identifiable as well. Therefore this case with W∗⊂Wicorresponds to (1) If instead there is no such W∗ iwe have P(i)(b|do(z0)) = P(i)(b), since now it must be the case that every member of Bis a non-descendant of Z0by Definition 3.4. This corresponds to option (2), and since P(v) is always available, the term is identifiable from Iafter the omission of domain indicators. Case of line 9. In this case, the effect is identified as a conditional distribution X B\YY Vj∈B PV\V(j) ϕP(i)(An(b)G[Z0]\z0|do(z0)) PV\V(j−1) ϕP(i)(An(b)G[Z0]\z0|do(z0)) =X B\YY Vj∈B P(i)(vj|do(z0),An(vj)G[Z0]\(vj∪z0)). As in the case of line 1, if there exists a W∗ i⊂Bsuch that W∗ i⊂Wiand (Zi, Wi)∈ S, then the product inside the sum takes the form Y Vj∈B P(i)(vj|do(z0),An(vj)G[Z0]\(vj∪z0)) =Y Vj∈W∗ i P(i)(vj|do(z0),An(vj)G[Z0]\(vj∪z0)) ×Y Vj∈(B\W∗ i) P(i)(vj|An(vj)G[Z0]\(vj∪z0)).(5) Here, individual terms of the form P(i)(vj|do(z0),An(vj)G[Z0]\(vj∪z0)) are obtained via conditioning exactly as the right-hand side of (4) by applying the same logic to Vjinstead of W∗itself, which is valid for vertices Vj∈W∗. The terms in the first product of (5) correspond to (3) and the terms in the second product correspond to (2). The equality again follows from the third rule of do-calculus that renders Vj∈B\W∗ iunaffected by the intervention on Z0. If no suitable W∗ iexists the product is simply Y Vj∈B P(i)(vj|do(z0),An(vj)D[Z0]\(vj∪z0)) = Y Vj∈B P(i)(vj|An(vj)D[Z0]\(vj∪z0)), where the terms in the product correspond to (3) and the third rule of docalculus is used again. These product terms are directly identifiable from P(v) after omitting domain indicators. Case of line 10. If line 10 was triggered with I=∅we are done, since the set of available experiments was set to ∅. If it was triggered with I6=∅, 24