6. When Does The Answer Follow?

Stefan Kober

In the hospital example, our records came from a process in which general health helped determine treatment.

We never observed the situation represented by:

$$ do(T=1). $$

Yet, once we supplied the causal graph, the observational records were enough to determine:

$$ P(R\mid do(T=1))=0.65. $$

Why were they enough?

The question is whether the observations and causal assumptions already fix the intervention answer.

In this example, they did. The graph told us why general health had to be adjusted for, and the records supplied the recovery rates within each health group and the distribution of general health. Once those pieces were held in place, the result was fixed:

$$ 0.5\cdot0.9+0.5\cdot0.4=0.65. $$

We could not keep those inputs and freely choose another value for:

$$ P(R\mid do(T=1)). $$

When an interventional quantity can be determined from the observational information together with the causal assumptions, it is said to be identified.

When The Answer Is Not Fixed

Having observational information and a causal model does not guarantee that an intervention quantity will be identified.

We can see why by changing the hospital example slightly.

Our adjustment worked because general health was represented in the causal model and recorded in the data.

We knew:

$$ P(R\mid T,G), $$

$$ P(R\mid T,\neg G), $$

and how common the two health conditions were in the population.

That allowed us to reconstruct the intervention result.

Now suppose the causal structure remains:

But general health was not recorded. We observe treatment and recovery, but not $G$.

The aggregate records can still tell us:

$$ P(R\mid T) $$

and:

$$ P(R\mid\neg T). $$

But our graph says that these comparisons mix the effect of treatment with the influence of general health on both treatment and recovery.

The adjustment we used before is no longer available, because the observational information needed to condition on $G$ is missing.

And the remaining observations need not determine a unique intervention answer.

Different causal processes compatible with the causal assumptions and the same observed distribution of treatment and recovery can imply different values for:

$$ P(R\mid do(T=1)). $$

In that case, the effect is not identified from the observational information available to us under the model.

Clearly there is a fact about what would happen under intervention. But the observations and causal assumptions supplied so far do not fix that fact, and more calculation using the same information cannot remove the remaining openness.

If general health was not measured, and no other observed variable or further causal assumption makes the effect recoverable in another way, we need different data or an actual intervention.

Identification marks that limit: it tells us whether the causal quantity is already fixed by the observational information under the causal assumptions we have supplied.