5. Why The Shape Of The Graph Matters

Stefan Kober

We have already seen that the effect of conditioning depends on where a variable sits in the causal structure.

A common cause and an intermediate variable behaved differently even though the same probability operation could be applied to both.

With three variables, there are only a few basic local shapes to consider. Looking at them side by side makes the differences easier to see.

The Fork

Return to the hospital example:

General health affects both treatment and recovery.

This shape is called a fork.

More generally:

$Z$ is a common cause of $X$ and $Y$.

That is the role general health played in our hospital model.

Observing treatment could tell us something about general health, and general health could tell us something about recovery.

So treatment and recovery could be associated partly because both were connected to $G$.

When we compared treated and untreated patients with the same general-health status, that route could no longer explain the difference between them.

For the causal question we were asking, conditioning on $G$ helped separate the effect of treatment from the common-cause path.

That is why adjustment worked in the previous chapter.

The Chain

Suppose instead like in the second example of chapter two:

Here $Z$ is not a common cause.

It lies on a causal route from $X$ to $Y$.

Suppose again treatment reduces blood pressure, and reduced blood pressure improves recovery:

Part of the treatment's effect on recovery now operates through blood pressure.

Imagine comparing treated and untreated patients only among people with the same blood pressure level.

Within that comparison, differences in blood pressure have been removed.

But blood pressure is one of the ways treatment changes recovery in this model.

So holding $B$ fixed also removes that part of the treatment's effect from the comparison.

The probability operation is familiar. We condition on a third variable. But its causal role is now very different.

In the fork, conditioning helped remove a route that was not part of the treatment's effect.

In the chain, conditioning removes a route through which the treatment itself can work. That is not what you want to determine the treatment effect.

The arrows make that difference visible.

The Collider

Suppose:

Both $X$ and $Y$ can affect $Z$.

The arrows meet at $Z$, so this shape is called a collider.

Consider a school that admits applicants for either strong academic results or strong athletic results.

Let:

$$ A=\text{academic strength}, $$

$$ S=\text{athletic strength}, $$

and:

$$ D=\text{admitted}. $$

We can represent the small causal structure as:

For simplicity, suppose academic and athletic strength are unrelated in the whole applicant population.

Now look only at admitted applicants.

Among them, learning that an applicant was not especially strong academically gives us a reason to expect some other route to admission.

In this small model, athletic strength is one such route.

Likewise, if an admitted applicant had exceptionally strong academic results, less athletic strength may have been needed for admission.

Once we restrict attention to admitted applicants, academic and athletic strength can become statistically related even though they were unrelated in the full applicant population.

Conditioning on the common effect has created an association.

That is the opposite of what happened in the fork.

There, conditioning on the common cause helped remove an association carried through that cause.

So there is no general causal rule saying that conditioning on more variables improves the analysis.

The same probabilistic operation can help remove a misleading association, remove part of the effect we are trying to measure, or create a new association.

Which of those happens depends on the causal structure. Graphs make those differences available for inspection. The $do$-notation then lets us mark the intervention whose consequences we want to follow through that structure.

What The Shapes Show

The fork, chain, and collider used the same basic graphical vocabulary.

The same probability operation, conditioning, was also available in every case.

Yet its causal significance changed with the arrangement of the arrows.

In the fork, conditioning could remove an association carried through a common cause.

In the chain, it could remove part of the effect we wanted to examine.

At a collider, it could create an association that was not present before.

The arrangement does more than display several causal claims conveniently. Once its interpretation is accepted, the structure constrains what the same probabilistic operation can tell us about an intervention.

This also makes disagreement easier to locate.

Two people may agree about the probabilities and about how conditioning works while disagreeing about whether a variable is a common cause, an intermediate variable, or a common effect.

The disagreement has moved from the calculation to the causal representation.

Making that dependency explicit does not resolve the disagreement.

It makes more of what carries the causal conclusion available for examination.