<?xml version="1.0" encoding="UTF-8"?><?xml-model type="application/xml-dtd" href="http://jats.nlm.nih.gov/publishing/1.1d3/JATS-journalpublishing1.dtd"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1d3 20150301//EN" "http://jats.nlm.nih.gov/publishing/1.1d3/JATS-journalpublishing1.dtd">
<article xmlns:ali="http://www.niso.org/schemas/ali/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.1d3"  specific-use="IPAR UPVEHU" article-type="research-article" xml:lang="es">




<front>
<journal-meta>
<journal-id journal-id-type="UPVEHU-id">3701</journal-id>
<journal-title-group>
<journal-title specific-use="original" xml:lang="en">THEORIA. Revista de Teoría, Historia y Fundamentos de la Ciencia</journal-title>
</journal-title-group>
<issn pub-type="ppub">0495-4548</issn>
<issn pub-type="epub">2171-679X</issn>
<publisher>
<publisher-name>Universidad del País Vasco/Euskal Herriko Unibertsitatea</publisher-name>
<publisher-loc>
<country>España</country>
<email>theoria@ehu.es</email>
</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="art-access-id" specific-use="UPVEHU-id">https://orcid.org/0000-0002-0738-9958</article-id>
<article-id pub-id-type="doi">https://doi.org/10.1387/theoria.22904</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject> MONOGRAPHIC SECTION </subject>
</subj-group>
</article-categories>
      <title-group id="title-group-1">
        <article-title id="article-title-1">
          <bold id="bold-1">Computational causal discovery: Advantages and assumptions</bold>
          <italic id="italic-1">Descubrimiento causal computacional: ventajas y asunciones</italic>
        </article-title>
      </title-group>
      <contrib-group id="contrib-group" content-type="author">

        <contrib id="contrib-author" contrib-type="person">
          <name>
            <surname>Zhang</surname>
            <given-names>Kun</given-names>
          </name>
          <aff>Department of Philosophy, Carnegie Mellon University </aff>
        </contrib>
      </contrib-group>
<pub-date pub-type="epub-ppub">
<season>January</season>
<year>2022</year>
</pub-date>
<volume>37</volume>
<issue>1</issue>
<fpage>75</fpage>
<lpage>86</lpage>
<history>
<date date-type="received" publication-format="10 06 2021">
<day>10</day>
<month>06</month>
<year>2021</year>
</date>
<date date-type="accepted" publication-format="31 01 2022">
<day>31</day>
<month>01</month>
<year>2022</year>
</date>
</history>
<permissions>
<ali:free_to_read/>
</permissions>
<abstract xml:lang="es">
<title>Resumen</title>
<p>Quiero dar la enhorabuena a James Woodward por Flagpoles anyone? (Woodward, 2022), una contribución que supone un nuevo hito tras la publicación de Making Things Happen: A Theory of Causal Explanation (Woodward, 2003). Making Things Happen ofrece una elegante teoría intervencionista para entender la explicación y la causación. Esta nueva contribución (Woodward, 2022) se apoya en esa teoría y da grandes pasos hacia la inferencia empírica de relaciones causales a partir de datos no experimentales. En este artículo, me centro en algunos métodos computacionales emergentes para encontrar relaciones causales a partir de evidencia no experimental y trato de complementar la contribución de Woodward discutiendo: 1) cómo estos métodos se conectan con la teoría intervencionista de la causalidad; 2) cómo de informativos son los resultados de estos métodos, incluyendo si producen gráficos causales dirigidos y cómo tratan los confusores (causas no medidas comunes a dos variables  medidas); y 3) las asunciones subyacentes a la corrección asintótica de los resultados de estos métodos de descubrimiento causal. Diferentes métodos pueden basarse en aspectos diferentes de a distribución conjunta de los datos. Esta discusión pretende dar una explicación técnica de tales asunciones.</p>
</abstract>

<trans-abstract xml:lang="en">
<title>Abstract</title>
<p>I would like to congratulate James Woodward for another landmark accomplishment, after publishing his Making Things Happen: A Theory of Causal Explanation (Woodward, 2003). Makes Things Happens gives an elegant interventionist theory for understanding explanation and causation. The new contribution (­Woodward, 2022) relies on that theory and further makes a big step towards empirical inference of causal relations from non-experimental data. In this paper, I will focus on some of the emerging computational methods for finding causal relations from non-experimental data and attempt to complement Woodward’s contribution with discussions on 1) how these methods are connected to the interventionist theory of causality, 2) how informative the output of the methods is, including whether they output directed causal graphs and how they deal with confounders (unmeasured common causes of two measured variables), and 3) the assumptions underlying the asymptotic correctness of the output of the methods about causal relations. Different causal discovery methods may rely on different aspects of the joint distribution of the data, and this discussion aims to provide a technical account of the assumptions.</p>
</trans-abstract>

<kwd-group xml:lang="es">
<title>Palabras clave</title>
<kwd>dirección causal</kwd>
<kwd>teoría intervencionista</kwd>
<kwd>modelo causal lineal, no-Gaussiano</kwd>
<kwd>confusores</kwd>
<kwd>fidelidad</kwd>  

</kwd-group>

<kwd-group xml:lang="en">
<title>Keywords</title>
<kwd>causal direction</kwd>
<kwd>interventionist theory</kwd>
<kwd>linear, non-Gaussian causal model</kwd>
<kwd>confounders</kwd>
<kwd>faithfulness</kwd>


</kwd-group>

<counts>
<fig-count count="2"/>
<table-count count="0"/>
<equation-count count="7"/>
<ref-count count="25"/>
</counts>
</article-meta>
</front>
      <article-title>Computational causal discovery: Advantages and assumptions </article-title>
      <title>Descubrimiento causal computacional: ventajas y asunciones </title>
      <p>Kun Zhang* </p>
      <p>Department of Philosophy, Carnegie Mellon University </p>
      <abstract>Abstract: I would like to congratulate James Woodward for another landmark accomplishment, after publishing his Making Things Happen: A Theory of Causal Explanation (Woodward, 2003). Makes Things Happens gives an elegant interventionist theory for understanding explanation and causation. The new contribution (­Woodward, 2022) relies on that theory and further makes a big step towards empirical inference of causal relations from non-experimental data. In this paper, I will focus on some of the emerging computational methods for finding causal relations from non-experimental data and attempt to complement Woodward’s contribution with discussions on 1) how these methods are connected to the interventionist theory of causality, 2) how informative the output of the methods is, including whether they output directed causal graphs and how they deal with confounders (unmeasured common causes of two measured variables), and 3) the assumptions underlying the asymptotic correctness of the output of the methods about causal relations. Different causal discovery methods may rely on different aspects of the joint distribution of the data, and this discussion aims to provide a technical account of the assumptions. </abstract>
      <subject>Keywords: causal direction; interventionist theory; linear, non-Gaussian causal model; confounders; faithfulness. </subject>
      <abstract>Resumen: Quiero dar la enhorabuena a James Woodward por Flagpoles anyone? (Woodward, 2022), una contribución que supone un nuevo hito tras la publicación de Making Things Happen: A Theory of Causal Explanation (Woodward, 2003). Making Things Happen ofrece una elegante teoría intervencionista para entender la explicación y la causación. Esta nueva contribución (Woodward, 2022) se apoya en esa teoría y da grandes pasos hacia la inferencia empírica de relaciones causales a partir de datos no experimentales. En este artículo, me centro en algunos métodos computacionales emergentes para encontrar relaciones causales a partir de evidencia no experimental y trato de complementar la contribución de Woodward discutiendo: 1) cómo estos métodos se conectan con la teoría intervencionista de la causalidad; 2) cómo de informativos son los resultados de estos métodos, incluyendo si producen gráficos causales dirigidos y cómo tratan los confusores (causas no medidas comunes a dos variables  medidas); y 3) las asunciones subyacentes a la corrección asintótica de los resultados de estos métodos de descubrimiento causal. Diferentes métodos pueden basarse en aspectos diferentes de a distribución conjunta de los datos. Esta discusión pretende dar una explicación técnica de tales asunciones. </abstract>
      <subject>Palabras clave: dirección causal; teoría intervencionista; modelo causal lineal, no-Gaussiano; confusores; fidelidad. </subject>
      <p>* Correspondence to: Kun Zhang. Department of Philosophy, Carnegie Mellon University, Pittsburgh, PA 15213, USA – kunz1@cmu.edu – https://orcid.org/0000-0002-0738-9958 </p>
      <p>How to cite: Zhang, Kun (2022). «Computational causal discovery: Advantages and assumptions»; Theoria. An International Journal for Theory, History and Foundations of Science, 37(1), 75-86. (https://doi.org/10.1387/theoria.22904).  </p>
      <p>Received: 2021-06-10; Final version: 2022-01-31. </p>
      <p>ISSN 0495-4548 - eISSN 2171-679X / © 2022 UPV/EHU </p>
      <p> This work is licensed under a  Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License </p>
      <title-1>1. Introduction: Interventionist Theory of Causality and Discovering Causality </title-1>
      <p>Given two variables X and Y, we are concerned with the causal direction between them or the direction of explanation. Causal relations are potentially exploitable for the purpose of applying proper manipulations to achieve a certain goal, and it is naturally desirable to provide an interventionist account of causality (Woodward, 2003). Woodward (2022, p. 8) provides a simple version of the interventionist theory: </p>
      <disp-quote>(M) X causes Z if and only if (i) it is possible to intervene to change the value of X and (ii) under some such intervention on X, the value of Z would change. </disp-quote>
      <p>An intervention on X is an unconfounded manipulation of X that changes any other variable, if at all, only through the change in X. That is, the intervention on X directly changes X, but does not directly change any other variable in the system. There are “hard” and “soft” interventions. A hard intervention breaks the connection from direct causes of X (except the intervention) to X, and a soft intervention does not break the connection from those direct causes to X but provides X with an exogenous source of variation that is independent from other causes of X (Eberhardt and Scheines, 2007). </p>
      <p>As noted by Woodward (2003, 2022), the notion of intervention is itself a causal notion and as such has a notion of causal direction built into it. However, it provides a way to verify causal claims, if one is able to actually apply changes to the system that are confirmed to be interventions. For instance, that might be possible if time order information is available such that one can check whether the change to X would directly change any other variable. Even without time order information, sometimes we can make sure that the applied changes are valid interventions, thanks to the partial knowledge of the process; for instance, gene knockout is an intervention on a gene that can be exploited for inferring regulatory networks (Pinna et al., 2010). This also suggests that applying proper interventions is usually too expensive, too time-consuming, or even impossible. </p>
      <p>On the other hand, if one thinks of the observed data (which clearly have multiple values) as produced under unknown or natural “interventions”, then causal direction may be revealed by finding whether certain types of “interventions” actually exist in the data. This idea makes it possible to find causal direction by analyzing observed data, as argued by Woodward (2022) and demonstrated by the algorithms that were recently proposed for distinguishing cause from effects (Shimizu et al., 2006; Zhang and Chan, 2006; Hoyer et al., 2009; Zhang and Hyvärinen, 2009; Janzing et al., 2012; Huang et al.). </p>
      <p>In fact, the past decades witnessed much progress in discovering causal information from non-experimental data, known as causal discovery, and its successful applications. Since the 1990s, conditional independence relationships in the data have been exploited to recover the underlying causal structure to a certain extent (Spirtes et al., 1993; C­hickering, 2002). Recently it has been shown that algorithms based on properly defined functional causal models (FCMs) are able to find causal direction between two variables and hence estimate the underlying causal Directed Acyclic Graph (DAG) uniquely. Can we trust the “causal” information produced by such computational causal discovery methods? How are the methods related to the interventionist theory of causality? Under what assumptions can those method produce causal information? Below we will focus on those questions. </p>
      <title-1>2. Relating Interventionist Theory to Causal Direction Determination between Two Variables </title-1>
      <p>Woodward (2022) suggests that “when one infers causal direction on the basis of non-experimental information what one is in effect doing is inferring what would happen if various interventions were to be performed without actually doing the interventions, relying instead on other features present in such situations—the independence/invariance features.” Indeed, we can think of the goal of causal discovery from non-experimental data as finding the footprint of unknown interventions that were applied (by nature, for instance) on the data. </p>
      <p>Woodward formulates the Causal to Statistical Independence (CSI) assumption, which says that variables that are causally independent are statistically independent. This assumption can be seen as a weaker version of the Causal Markov Condition (CMC) (Kiive­ri et al., 1984; Glymour et al., 1987), and it is generally plausible. CSI motivated the following principle for inferring causal direction (Woodward, 2022, p. 26): </p>
      <disp-quote>(P) Suppose there are 3 variables, H, A and S such that either (i) H and A cause S or (ii) A and S cause H. (This assumption, in some way, implies that there are no omitted common causes etc.) Suppose the patterns of dependence among these three variables are as follows: H ╨ A, H ╨/ S, A ╨/ S, where ╨ means statistical independence and ╨/ means statistical dependence. Then (i) is the correct causal order. </disp-quote>
      <p>Woodward (2022, p. 28) then connects principle P to the interventionist theory M in the following way—this connection is desirable since one aims to use principle P to find causality which can be understood in terms of interventions: </p>
      <disp-quote>According to the interventionist framework, the claim that H causes S and S does not cause H corresponds to the condition that there are interventions on H that will change S but no interventions on S that will change H. Assuming that these are the only two possibilities (i) and (ii) and that there is no common causes, as stated in principle P, the pattern of (in)dependencies H ╨ A, H ╨/ S, and A ╨/ S suggests that A functions as a soft intervention variable on S, since it is exogenous and independent of the only other possible cause of S, namely H. [Let us denote this statement by J1.] Observation shows that changes in this intervention variable A for S are not associated with changes in H, suggesting that S does not cause H. Moreover, if we assume that S causes H, then, under this assumption, there will not be, among the variables in the system, any intervention variable for H that is independent of S, since the only remaining variable, A, is correlated with S. [Denote by J2 this statement.] Hence, the dependence pattern suggests there is a route to changing S that is independent of H (which is what we expect if H causes S)—namely the route involving A—but no route to changing H that is independent of S, which is what we expect if S causes H. </disp-quote>
      <p>Woodward (2022) [Section 9] then goes one step further, to consider the problem of inferring causal direction when only two variables, X and Y, are giving and justify a set of methods to solve this problem. Assume the causal influence from the cause variable and the unmeasured factors (noise) to the effect variable follows some constrained Functional Causal Model (FCM) class, such as the Linear, Non-Gaussian, Acyclic Model (LiNGAM) (Shimizu et al., 2006), Post-NonLinear (PNL) causal model (Zhang and Chan, 2006; Zhang and Hyvärinen, 2009), and the Additive Noise Model (ANM) (Hoyer et al., 2009). First of all, Woodward noted that given only two measured variables X and Y, the error term is unobserved and must be inferred, in order to apply principle P. When we find an error U which is independent of X but no error U l which is independent of Y, we infer that U and X are causes of Y. Such asymmetry between two variables X and Y that are assumed to be directly causally related (the estimated error term is independent from the hypothetical cause in only one direction), under the assumption of a properly constrained FCM class, has inspired several approaches to causal discovery that are able to recover the underlying causal DAG uniquely. </p>
      <p>Why does the asymmetry in the (in)dependence between the error term and the hypothetical cause imply causal direction between two variables, denoted by H and S? This can be examined from an interventionist perspective, as an application of the arguments in the above quote to justify principle P for finding causal directions among 3 variables from the interventionist perspective. We notice that principle P makes use of three variables H, A, and S, while in the two-variable case, U is not observed but constructed from variables X and Y, in light of the constrained FCM class. In order to apply this principle to find causal direction when only two variables, H and S, are given, or more specifically, in order for statements J1 and J2 to hold true when only variables H and S are measured, the following conditions are expected to hold: </p>
      <bullet>C1)	Variable A is a variable that actually exists in the system under consideration. (As a consequence, J1 is true.) </bullet>
      <bullet>C2)	There does not exist another variable in the system, A’, such that S ╨ A’ and A’ ╨/ H (otherwise H and S will be symmetric). (As a consequence, J2 is true.) </bullet>
      <p>As stated in the condition of principle P, it is assumed that H and A are not confounded in the first place. Can we make an alternative set of assumptions that are technically testable or appear weak in terms of the underlying causal model, while guaranteeing that the estimated causal direction from the two given variables is asymptotically correct? We will discuss the required assumptions in Section 4. Specifically, we will see some technical assumptions to imply condition C2 and assumptions on hidden variables in the system under which one can find causal direction even without the unconfounding assumption. Before that, let us review the assumptions that are required for conditional independence-based methods for causal discovery. </p>
      <title-1>3. Assumptions for Conditional Independence-Based Approaches </title-1>
      <p>Woodward (2022) focuses on finding the causal direction between two variables which are believed to be directly causally related. The issue of discovering causal information based on conditional independence relations among the measured variables seems somehow irrelevant to it. However, here for completeness of the discussion and for comparative purposes, let us briefly review such methods and discuss their assumptions. </p>
      <p>Widely-used conditional independence-based approaches to causal discovery include the PC algorithm and FCI (Spirtes et al., 2001). The PC algorithm returns a Markov Equivalence Class (MEC) of DAGs, and all DAGs in the MEC share the same adjacency and conditional independence relations. FCI allows confounders in the system and returns a Partial Ancestral Graph (PAG). Although the Greedy Equivalence Search (GES) (Chickering, 2002) is a score-based method for causal discovery that assumes no confounding, it usually assumes a linear-Gaussian model or multinomial data and as a consequence, it also makes use of conditional independence relations among the measured variables, together with a penalty determined by the complexity of the whole causal model, to find the MEC. </p>
      <p>Generally speaking, conditional independence relations among the measured variables, which reflect only part of the information implied by the joint distribution, do not contain sufficient information for producing a complete picture of the causal relations. First, even under the assumption of no confounding, the methods, such as PC and GES, return a MEC, which may contain multiple DAGs. For instance, if applied to two variables, they cannot determine their causal direction. Second, although there exist algorithms that can produce asymptotically correct results in the presence of confounders, such results are usually not strong enough to determine whether confounders exist. FCI is a remarkable algorithm whose result is asymptotically correct even with confounders. However, in the result by FCI, one usually cannot distinguish between unconfounded pairs of variables with direct causal relations between them and confounded pairs without direct causal relations in between—whenever it is possible to have confounders of the pair of variables, the algorithm will indicate it and, as a consequence, the output usually contains very few variable pairs that are directed causally related without cofounding. </p>
      <p>Standard assumptions for the asymptotic correctness of the output of the above algorithms are the CMC and Faithfulness assumption. The Faithfulness assumption (Spirtes et al., 2001; Zhang, 2013; Zhang and Spirtes, 2008) states that there is no ‘accidental’ conditional independence relation between the variables according to the distribution. More precisely, it says that all conditional independence relations among the variables are implications of the CMC applied to the DAG representing the true causal relations among the variables. Combining the CMC and Faithfulness assumptions, one is then able to recover some information of the underlying DAG from the measured data, given that we have enough data. Faithfulness can be viewed as one version of the “simplicity” assumption of the underlying DAG—if two variables are conditionally independent given any subset of the remaining variables, then they are not adjacent in the causal graph. </p>
      <title-1>4. Independent Noise-Based Approaches: Assumptions on Causal Mechanisms </title-1>
      <p>As Woodward (2022) noted, suitable assumptions on the causal mechanism (which are not made in conditional independence-based methods), such as a LiNGAM model to describe the causal effect, help find causal direction between two variables and hence recover the whole causal DAG, as supported by much empirical evidence. Below we start with an illustration of how and why LiNGAM, which is taken as an example of those properly defined constrained FCMs, helps find causal direction, and then discuss the required assumptions. </p>
      <title-3>4.1. Asymmetry between Two Variables: Illustration </title-3>
      <p>Consider the causal process X → Y with causal model Y = dX + U, where d is the linear coefficient and U is the noise (unmeasured factor) that is independent from X. As a concrete example, one can think of them as atmospheric pressure and the reading of a barometer, respectively. Suppose two linear regressions are done, one predicting X from Y and the other predicting Y from X. The residual of the regression of X on Y is the difference of the measured variable and its predicted value from the regression. In the anti-causal direction the residual is a random variable, U’ = X − αZ, that is a function of the predictor variable and the predicted variable, or a function of the underlying noise terms. The coefficient α can be estimated by minimizing the total squares of the residual, which implies that the residual and predictor Z are uncorrelated. Bearing in mind that for simplicity of the presentation, all variables are assumed to be standardized (i.e., they have a zero mean and unit variance), one can see that </p>
      
      <disp-formula-group>
          <graphic id="graphic-1" mime-subtype="jpg" mimetype="image" xlink:href="form01-Zhang-Theoria371.jpg" />
      </disp-formula-group>

      <p>where ov(·) and ar(·) denote covariance and variance, respectively. The residual of regressing Y on X (in the causal direction) is U, and the residual for regression in the anti-causal direction is </p>

      <disp-formula-group>U’ = X − αY = X − d(dX + U) = (1 − d2)X − dU. </disp-formula-group>

      <fig id="fig-1">

         <label><regular>Figure 1</regular></label>   
         <caption> <title>Illustration of the asymmetry between cause and effect in the linear, non-Gaussian case, where X causes Y with Y = dX + U. From left to right: the scatter plots for X and Y, for X and the residual of regressing Y on X, for Y and X, and for Y and the residual of regressing X on Y, respectively.  Top row: cause X and noise U are both Gaussian; bottom row: they follow a uniform distribution (a particular type of non-Gaussian distributions). </title></caption>

         <alt-text>Figure 1. Illustration of the asymmetry between cause and effect in the linear, non-Gaussian case, where X causes Y with Y = dX + U. From left to right: the scatter plots for X and Y, for X and the residual of regressing Y on X, for Y and X, and for Y and the residual of regressing X on Y, respectively.  Top row: cause X and noise U are both Gaussian; bottom row: they follow a uniform distribution (a particular type of non-Gaussian distributions). </alt-text>
 
         <graphic id="graphic-2" mime-subtype="jpg" mimetype="image" xlink:href="https://ojs.ehu.eus/index.php/THEORIA/article/download/22904/20958/93645" />
      </fig>

      <p>Can we see the asymmetry between X and Y by looking at the properties of the residuals U and U’? </p>
      <p>U is assumed to be independent from X. However, since Y = dX + U, U’ and Y both involve independent variables X and U, and they CANNOT be independent if at least one of X and U is non-Gaussian, as implied by the Darmois-Skitovich theorem (Kagan et al., 1973) (this is related to condition C2): two linear combinations of a set of independent components cannot be independent from each other if they share any non-Gaussian independent component. So by testing for the independence between residuals and predictors, the direction of the X − Y causal link can be identified. For illustrative purposes, Figure 1 provides scatter plots of variables X and Y and scatter plots of the predictor and regression residual in the causal (left part) and anti-causal (right part) directions, in a joint Gaussian case (top row) and in the case with cause X and the noise term following uniform distributions (bottom row). In the uniform case (as a particular non-Gaussian case), one can see that in the anti-causal direction, residual U’ and predictor Y are clearly statistically dependent (because the conditional distribution of one of them given the other taking some specific value is not identical to its marginal distribution). </p>
      <title-3>4.2. What Assumptions on Causal Mechanisms Are Required to Guarantee Asymmetry? </title-3>
      <p>The linear, non-Gaussian model relies on the linearity of the causal mechanism, a particular type of parametric assumption, as well as the non-Gaussianity assumption of the noise terms, making it possible to estimate causal models from non-experimental data. Linear models are thought to be simple in the sense that they involve few parameters (e.g., as linear coefficients).1 In fact, if there is no proper constraint on the FCM, for any two given random variables, one can always write one of them as a function of the other and some independent error term, as shown by Hyvärinen and Pajunen (1999) and Zhang et al. (2015), where the function is related to the conditional distribution of the variables and may be very complex. This is also the case in the example given in Figure 1—although U’, the residual of regressing X on Y is not independent from Y in the uniform case, X can still be written as a rather complex, clearly non-linear function of Y and some error term that is independent from Y. Assuming linear, non-Gaussian causal models, one would infer the causal direction X → Y in this example. The linearity assumption can be checked by inspection on the scatter plots—given the measured data points, one can plot one variable against another variable, and the pattern should be approximately linear. In fact, statistical test of independence between the residual and the predictor (hypothetical cause) can also serve as a test of linearity—if the causal model is linear (resp., non-linear), then the residual produced by linear regression in some direction will be independent from (resp., dependent on) the predictor. If needed, one can also resort to specific tests of linearity, such as the Theil test (Theil, 1950), for this purpose. Test of non-Gaussianity can be performed implicitly or explicitly. If the residual is independent from the predictor only in one direction (i.e., only one DAG gives rise to independent residuals), then the predictor and residual cannot be jointly Gaussian. Alternatively, one may directly exploit statistical tests, such as the Shapiro-Wilk test, on the estimated residual or the predictor to check whether it follows the Gaussian distribution. </p>
      <p>In the causal discovery community, this type of asymmetry between X and Y is essential for distinguishing cause from effect: under the assumed (parametric or non-parametric) FCM class, only in the causal direction one can find an independent noise term. This is asymmetry directly implied by the data distribution and the FCM class. In order to give causal claims, e.g., X → Y, Woodward (2022) explicitly assumes that there is no confounder for X and Y. Below we give an alternative formulation of the assumption in the spirit of Faithfulness, to guarantee the connection between asymmetry implied by data and causal direction in the interventionist framework. </p>
      <title-3>4.3. Assumptions on Hidden Variables to Connect Asymmetry to Causal Direction </title-3>
      <p>Under the assumption that there was no confounder for X and Y, the above analysis does not require the traditional Faithfulness assumption for inferring causal directions asymptotically correctly (Shimizu et al., 2006). Note that in reality we usually are not sure whether there are confounders and if yes, how measured variables are confounded in the true processes—can we still trust the result produced by the above LiNGAM analysis (performing linear regression and testing for independence between the residual and hypothetical cause)? Do we need any Faithfulness-like assumptions at all to guarantee the correctness of the statement? </p>
      <p>Let us look at an illustrative example. </p>
      <p>Example. Let us consider four variables—lifestyle, mortality risk, food consumption, and physical activity—denoted by X, Y, W, and UX, respectively. Suppose that we can directly measure X and Y but not W and UX. Measured variables X and Y are generated by the two hidden, independent, non-Gaussian variables, W and UX, according to the following specific linear model, as shown in Figure 5(a), in which variables in the shaded area are not observed: </p>

      <disp-formula-group>X = UX + f W, Y = f X − f UX + (1 − f 2)W, </disp-formula-group>
      
      <p>where f is a positive number smaller than 1. </p>
      
      <p>Because of the specification of the coefficients, we can rewrite the above model for Y as </p>
      
      <disp-formula-group>Y = f X − f UX + (1 − f 2)W = f (UX + f W) − f UX + (1 − f 2)W = W.	(1) </disp-formula-group>

      <p>That is, mortality risk is determined by food consumption, although physical activity is also its cause. As a consequence, the statistical relationship between X and Y is </p>

      <disp-formula-group>X = UX + f W = f Y + UX. </disp-formula-group>

      <p>Although there is no deterministic relation between measured variables, in this case (X causes Y, together with specific confounders W and UX), we cannot identify the correct causal direction with the LiNGAM analysis—in fact, the distribution of X and Y in this case can also be represented by the causal model in Figure 5(b), in which Y causes X without a confounder. What led to the wrong conclusion produced by the LiNGAM analysis? How can we avoid it? We notice here </p>
      <bullet>1.	that the confounder UX (physical activity) actually has a zero effect on Y (mortality risk), although Y is its descendant in the causal graph, and </bullet>
      <bullet>2.	that although X (lifestyle) and Y (mortality risk) are non-deterministically related, they are completely determined by the confounders. </bullet>
      <p>Now suppose neither of the above two properties holds. That is, we make the following assumptions: </p>
      <bullet>AP1.	Any hidden variable (which may be a confounder or unobserved noise variable) has a non-zero total causal effect on any of its descendants. </bullet>
      <bullet>AP2.	Each measured variable has non-zero noise relative to its parents (among measured variables and unmeasured confounders).2 </bullet>
      <p>Under assumptions AP1 and AP2, one can find causal direction in both the unconfounded and confounded cases. In the considered example, let us denote by g11, g12, g21, and g22 the direct causal effects (linear coefficients) of UX on X, UX on Y, W on X, and W on Y, respectively, and still let f be the causal effect of X on Y. Then X and Y are generated from non-Gaussian independent variables, EX (unobserved noise in X), EY (unobserved noise in Y), UX, and W, in the following way: </p>

      <disp-formula-group>X = g11 UX + g21 W + EX, Y = g12 UX + f X + g22W + EY = (g12 + f g11)UX + (g22 + f g21)W + f EX + EY, </disp-formula-group>

      <p>or in matrix form: </p>

      <disp-formula-group>
      <graphic id="graphic-3" mime-subtype="jpg" mimetype="image" xlink:href="form02-Zhang-Theoria371.jpg" />
      </disp-formula-group>
      
      <p>The above model, or specifically, the coefficients in matrix A1, is identifiable from X and Y with overcomplete ICA (Eriksson and Koivunen, 2004; Hoyer et al., 2008). We can then see that under assumptions AP1 and AP2, we can find the causal relation between X and Y from the estimated matrix A1. From the second column of A1, we know that Y does not cause X—otherwise, given that EY influences Y, as indicated by the non-zero entry of the second entry of this column, EY must influence X as well, according to Assumption AP1; this means that the first entry of this column cannot be zero, which is not the case. Similarly, we know that it is impossible that X does not influence Y—because if X did not influence Y, its noise term, EX would influence only X, but not Y, and accordingly, there would be some column of A1 in which the first entry is non-zero while the second is zero, but there is no such column in A1. So the causal relation is X → Y. </p>
 


       <fig id="fig-2">      
       
         <Tabla xmlns:aid="http://ns.adobe.com/AdobeInDesign/4.0/" aid:table="table" aid:trows="1" aid:tcols="2">
       
            <Celda aid:table="cell" aid:crows="1" aid:ccols="1" aid:ccolwidth="133.2283464566929">
               <graphic id="graphic-4" mime-subtype="jpg" mimetype="image" xlink:href="fig02-Zhang-Theoria371.jpg" />   
               <p>(a) A causal model with specific confounders, in which variables in the shaded area are not observable.</p>
            </Celda>

            <Celda aid:table="cell" aid:crows="1" aid:ccols="1" aid:ccolwidth="104.8818897637795">
               <graphic id="graphic-5" mime-subtype="jpg" mimetype="image" xlink:href="fig03-Zhang-Theoria371.jpg" />   
               <p>(b) A causal model with no confounders but a reverse direction between X and Y.</p>
            </Celda>
       
         </Tabla>
          
 
         <label><regular>Figure 2</regular></label>   
         <caption> <title>Two causal models used in the Example.  While Causal model (a) is the true one, it and model  (b) produce the same joint distribution of X and Y. </title></caption>

         <alt-text>Figure 2. Two causal models used in the Example.  While Causal model (a) is the true one, it and model  (b) produce the same joint distribution of X and Y. </alt-text>
 
         
      </fig>


      <p>Both Assumptions AP1 and AP2 are needed to find the correct causal direction between X and Y. If Assumption AP1 does not hold while AP2 holds, or more specifically, if EX and EY are zero, then one can only recover the last two columns of A1, which does not imply an acyclic relation between X and Y. If Assumption AP2 does not hold while AP1 holds, e.g., the second entry of the third column of A1 is zero, i.e., g12 + f g11 = 0, then an alternative causal model consistent with Eq. (2) may be that X and Y do not have a direct causal influence between them (as seen from the second and third columns of A1) and that there are two confounders, EX and W (each influences both X and Y, as seen from the first and fourth columns). If both AP1 and AP2 are violated, linear, non-Gaussian methods may wrongly infer that Y causes X, as seen at the beginning of the E­xample. </p>
      <title-1>5. Conclusion </title-1>
      <p>Woodward (2022) discusses principles for finding causal direction between two random variables. As philosophical reflections on and justifications of the methods for distinguishing cause from effect recently proposed in machine learning, he connects the computational principles to his interventionist account of causality. In this paper we focus on independent noise-based methods, and attempt to provide alternative formulations of the assumptions that guarantee the correctness of the discover direction. Conditional independence-based methods for causal discovery and their assumptions are briefly mentioned for completeness of the comparison. </p>
      <p>In order to relate the information that is discovered from empirical data to the underlying causal structure, proper assumptions have to be made. Conditional independence-based methods assume some type of faithfulness: conditional independence in the data is not an accidental statistical property, but a reflection of the underlying causal graph. Independent noise-based methods assume “simplicity” of the causal mechanism, as encoded by the functional class of the functional causal models, to guarantee that cause and effect are asymmetric—only in the correct causal direction, the estimated noise term is independent from the hypothetical cause. Moreover, if one does not directly assume out confounders, some faithfulness-type assumption is needed to guarantee that the asymmetry that is discovered from empirical data actually implies causal direction. </p>
      <p>Some of the assumptions, such as Woodward (2022)’s Causal to Statistical Independence (CSI) assumption, are widely accepted in the machine learning and philosophy communities. Some of the assumptions, such as linearity of the relations and non-Gaussianity of the noise, are generally testable. Some of them, including the faithfulness-type assumption AP1, are not generally testable. </p>
      <p>It is worth noting that causal discovery is typically different from traditional machine learning problems such as regression, classification, and clustering, although both of them learn from data: causal discovery aims to find the underlying truth, while machine learning usually aims at good predictions. Therefore, first, it is essential to connect the principles underlying the computational methods to the interventionist theory of causality, to make sure that the causal discovery result actually has a causal interpretation. Second, researchers and practitioners in causal discovery have to pay close attention to the assumptions to guarantee the correctness (relative to the ground truth) of the result produced by computational methods for causal discovery, as pointed out by Woodward (2022). </p>
      <title-1>Acknowledgements </title-1>
      <p>I would like to acknowledge the support by the United States Air Force under Contract No. FA8650-17-C-7715, by National Institutes of Health under Contract No. R01HL159805, and by a grant from Apple. The United States Air Force or National Institutes of Health is not responsible for the views reported in this article. I am grateful to Clark Glymour for stimulating discussions and all his support. </p>
      <ref>References </ref>
      <ref>Chickering, D. M. (2002). Optimal structure identification with greedy search. Journal of machine learning research, 3(Nov):507-554. </ref>
      <ref>Eberhardt, F. &amp; Scheines, R. (2007). Interventions and causal inference. Philosophy of Science, 74: 981-995. </ref>
      <ref>Eriksson, J. &amp; Koivunen, V. (2004). Identifiability, separability, and uniqueness of linear ICA models. IEEE Signal Processing Letters, 11(7):601-604. </ref>
      <ref>Glymour, C., Scheines, R., Spirtes, P., &amp; Kelly, K. (1987). Discovering Causal Strucure. Academic Press. </ref>
      <ref>Hoyer, P. O., Shimizu, S., Kerminen, A. J., &amp; Palviainen, M. (2008). Estimation of causal effects using linear non-gaussian causal models with hidden variables. International Journal of Approximate Reasoning, 49(2):362-378. </ref>
      <ref>Hoyer, P. O., Janzing, D., Mooji, J., Peters, J., &amp; Schölkopf, B. (2009). Nonlinear causal discovery with additive noise models. In Advances in Neural Information Processing Systems 21, Vancouver, B.C., Canada. </ref>
      <ref>Huang, B., Zhang, K., Zhang, J., Ramsey, J., Sanchez-Romero, R., Glymour, C., &amp; Schölkopf, B. (2020). Causal discovery from heterogeneous/nonstationary data. Journal of Machine Learning Research, 21:1-53. </ref>
      <ref>Hyvärinen, A. &amp; Pajunen, P. (1999). Nonlinear independent component analysis: Existence and uniqueness results. Neural Networks, 12(3):429-439. </ref>
      <ref>Janzing, D., Mooij, J., Zhang, K., Lemeire, J., Zscheischler, J., Daniušis, P., Steudel, B., &amp; Schölkopf, B. (2012). Information-geometric approach to inferring causal directions. Artificial Intelligence, 182:1-31. </ref>
      <ref>Kagan, A. M., Linnik, Y. V., &amp; Rao, C. R. (1973). Characterization Problems in Mathematical Statistics. Wiley, New York. </ref>
      <ref>Kiiveri, H., Speed, T., &amp; Karlin, J. B. (1984). Recursive causal models. Journal of theAustralian Mathematical Society (Series A), 36:30-52. </ref>
      <ref>Mandt, S., Wenzel, F., Nakajima, S., Cunningham, J., Lippert, C., &amp; Kloft, M. (2017). Sparse probit linear mixed model. Machine Learning, 106:1621-1642. </ref>
      <ref>Pinna, A., Soranzo, N., &amp; de la Fuente, A. (2010). From knockouts to networks: Establishing direct cause-effect relationships through graph analysis. PLoS ONE, 5. </ref>
      <ref>Schwarz, G. (1978). Estimating the dimension of a model. The Annals of Statistics, 6:461-464. </ref>
      <ref>Shimizu, S., Hoyer, P. O., Hyvärinen, A., &amp; Kerminen, A. J. (2006). A linear non-Gaussian acyclic model for causal discovery. Journal of Machine Learning Research, 7:2003-2030. </ref>
      <ref>Spirtes, P., Glymour, C., &amp; Scheines, R. (1993). Causation, Prediction, and Search. Spring-Verlag Lectures in Statistics. </ref>
      <ref>Spirtes, P., Glymour, C., &amp; Scheines, R. (2001). Causation, Prediction, and Search. MIT Press, Cambridge, MA, 2nd edition. </ref>
      <ref>Theil, H. (1950). A rank-invariant method of linear and polynomial regression analysis i. Proceedings of the Royal Netherlands Academy of Sciences, 53:386-392. </ref>
      <ref>Woodward, J. (2003). Making things happen: A theory of causal explanation. Oxford University Press, New York. </ref>
      <ref>Woodward, J. (2022). Flagpoles anyone? Causal and explanatory asymmetries. THEORIA. An International Journal for Theory, History and Foundations of Science, 37(1), 7-52 (https://doi.org/10.1387/theoria.21921). </ref>
      <ref>Zhang, J. (2013). A comparison of three occam’s razors for markovian causal models. British Journal of Philosophy of Science, 64:423-448. </ref>
      <ref>Zhang, J. &amp; Spirtes, P. (2008). Detection of unfaithfulness and robust causal inference. Minds and Machines, 18:239-271. </ref>
      <ref>Zhang, K. &amp; Chan, L. (2006). Extensions of ICA for causality discovery in the Hong Kong stock market. In Proc. 13th International Conference on Neural Information Processing (ICONIP 2006). </ref>
      <ref>Zhang, K. &amp; Hyvärinen, A. (2009). On the identifiability of the post-nonlinear causal model. In Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence, Montreal, Canada. </ref>
      <ref>Zhang, K., Wang, Z., Zhang, J., &amp; Schölkopf, B. (2015). On estimation of functional causal models: General results and application to post-nonlinear causal model. ACM Transactions on Intelligent Systems and Technologies. </ref>
      <ref>KUN ZHANG is an associate professor of philosophy and an affiliate faculty in the machine learning department at Carnegie Mellon University. He has been actively developing methods for automated causal discovery from various kinds of data, investigating machine learning problems including transfer learning and representation learning from a causal perspective, and studying philosophical foundations of causation and various machine learning tasks. </ref>
      <p>ADDRESS: Department of Philosophy, Carnegie Mellon University, Pittsburgh, PA 15213, USA.  E-mail: kunz1@cmu.edu  ORCID: 0000-0002-0738-9958 </p>
      <p>Notes </p>
      <footnote>1Here is a brief explanation of why we prefer linear models from a model selection perspective. If multiple models explain the given data equally well, i.e., with the same likelihood, then suitable model selection approaches, such as Bayesian Information Criterion (BIC) (Schwarz, 1978), would prefer the model with the fewest number of free parameters. </footnote>
      <footnote>2In some work, confounding is modeled by correlated noise; see, e.g. (Mandt et al., 2017). Assumption AP2 is stronger than it. For instance, in Example 1 X and Y are non-deterministically correlated, but Assumption AP2 is violated.</footnote>
   </article>
