Appendix B: Examples Of Frequentist Handling Of Evidence
Probability And Statistical Inference
Earlier in the essay, probability was introduced through cases where the setup was already given and we calculated what followed from it.
A fair die has a represented possibility space and assigned weights. From those, we can calculate how likely different outcomes are.
Bayesian updating already moved us into statistical inference.
There the direction was partly reversed.
We had hypotheses about what might be producing the observations, some initial weights assigned to those hypotheses, and new evidence. The evidence then constrained how those weights changed.
So statistical inference is broadly when we use observed data to learn something about the process, population, or hypothesis that may have produced it.
In the Bayesian case, the probabilities belonged to the hypotheses and were updated with the evidence. Frequentist methods place the probability differently. In frequentist methods, the probabilities usually belong to possible data or statistics under a specified model, and the observed data is compared with that distribution.
Imagine an urn whose composition we do not know. We draw some balls and record what we see.
A Bayesian treatment might represent several possible compositions of the urn, assign probabilities to them, and update those probabilities in light of the sample.
A frequentist treatment does not need to assign probabilities to the possible compositions themselves. Instead, it asks how samples would behave if a particular composition or sampling model were in place, and then compares the observed sample with those possible results.
That difference will become clearer in the examples.
We will look at two cases: randomization in an experiment, and sampling from a population.
Randomization
One important route into frequentist statistics comes from the work of Ronald Aylmer Fisher.
Fisher developed methods in which the randomization of an experiment gives probability a precise role in assessing the observed result.
An anecdotal example used often by the statistician Phillip I. Good makes the mechanism visible¹.
Good discusses an experiment involving cell cultures grown either with a conventional nutrient solution or with vitamin E added.
After some losses, six usable observations remained, three in each condition:
| vitamin E | control |
|---|---|
| 121 | 34 |
| 118 | 22 |
| 110 | 12 |
The sum of the vitamin E measurements is:
$$ 121+118+110=349. $$
The sum of the control measurements is:
$$ 34+22+12=68. $$
Their difference is:
$$ 349-68=281. $$
But is this difference large or small? Compared to what other differences?
The experiment gives us a way to answer that.
The treatment labels were assigned randomly. Now suppose that vitamin E had no effect on any of these measured outcomes.
Under that hypothesis, it would not matter which three specimens had received the vitamin E label. The six measurements would remain:
$121,\ 118,\ 110,\ 34,\ 22,\ 12$.
We can thus keep the measurements fixed and redistribute three vitamin E labels and three control labels over them in every way allowed by the experiment.
There are:
$$ \binom{6}{3}=20 $$
such assignments.
Under this randomization design, each has probability:
$$ \frac1{20}=0.05. $$
These 20 assignments give us a reference distribution against which we can compare the observed difference.
For each possible assignment we calculate the same statistic: the sum of the vitamin E measurements minus the sum of the control measurements.
| vitamin E | control | difference |
|---|---|---|
| 121, 118, 110 | 34, 22, 12 | 281 |
| 121, 118, 34 | 110, 22, 12 | 129 |
| 121, 110, 34 | 118, 22, 12 | 113 |
| 118, 110, 34 | 121, 22, 12 | 107 |
| 121, 118, 22 | 110, 34, 12 | 105 |
| 121, 110, 22 | 118, 34, 12 | 89 |
| 121, 118, 12 | 110, 34, 22 | 85 |
| 118, 110, 22 | 121, 34, 12 | 83 |
| 121, 110, 12 | 118, 34, 22 | 69 |
| 118, 110, 12 | 121, 34, 22 | 63 |
| 121, 34, 22 | 118, 110, 12 | -63 |
| 118, 34, 22 | 121, 110, 12 | -69 |
| 121, 34, 12 | 118, 110, 22 | -83 |
| 110, 34, 22 | 121, 118, 12 | -85 |
| 118, 34, 12 | 121, 110, 22 | -89 |
| 110, 34, 12 | 121, 118, 22 | -105 |
| 121, 22, 12 | 118, 110, 34 | -107 |
| 118, 22, 12 | 121, 110, 34 | -113 |
| 110, 22, 12 | 121, 118, 34 | -129 |
| 34, 22, 12 | 121, 118, 110 | -281 |
The entire probabilistic structure is visible.
The observed difference $281$ is the largest positive difference among the 20 possible assignments.
Only one assignment produces a result at least this favorable to vitamin E.
So, under the no-effect hypothesis and this randomization procedure $p=\frac1{20}=0.05.$
If vitamin E really made no difference to these measured outcomes, a result at least this favorable to vitamin E would occur in only one of the 20 possible assignments.
That makes the observed result hard to reconcile with the no-effect setup. N.B.: It does not mean that the probability of the hypothesis is $5\%$. If we repeatedly used a valid test with a $5\%$ rejection threshold in situations where the no-effect hypothesis was actually true, the procedure would falsely reject that hypothesis only about $5\%$ of the time or less.
That is the frequentist error guarantee. It concerns the long-run behavior of the procedure, not the probability that this particular hypothesis is true.
It is a common convention to reject the no-effect hypothesis when the p-value is at most $5\%$.
This differs from the Bayesian case in an important way. There we assigned probability to competing hypotheses and changed those probabilities when evidence arrived. Here we have not assigned a probability to the no-effect hypothesis. We assume it for the purpose of the test and ask how often results like the one we observed would arise under it.
Both approaches use a model of what the evidence would look like under a hypothesis. Bayesian updating uses that relation to redistribute probability among hypotheses. The randomization test uses it to estimate how likely the observed result and more extreme results are compared to the other outcomes expected under the hypothesis being tested.
Sampling From A Population
Randomization is only one route within frequentist statistics.
Another large family of methods is based on sampling.
Instead of asking how treatment assignments could have varied under an experimental design, we ask what observations drawn from a population can tell us about that population.
The familiar urn picture makes the basic relation visible.
Imagine an urn containing a very large number of red and blue balls.
We do not know the proportion of each color.
Let $p$ be the proportion of red balls in the urn.
If we knew that $p=0.60$, probability theory could tell us how samples from the urn would tend to vary.
A sample of ten balls could easily contain somewhat more or fewer red balls than the population proportion would suggest.
If we repeatedly drew random samples of the same size from the same population, the resulting sample proportions would form a probability distribution.
Now reverse the problem.
Suppose we do not know $p$.
We draw a random sample and observe $60$ red balls among $100$.
The sample proportion is:
$$ \hat p=\frac{60}{100}=0.60. $$
Another sample from the same urn might give $0.57$, and another $0.64$.
The statistical problem is to understand how much variation the sampling process itself can produce.
Without going into the calculation, samples like these would be much easier to obtain from a population with $p$ near $0.60$ than from one with $p$ near $0.90$ or $0.10$.
Probability theory lets us work out how compatible different population values are with the observed sample under the assumed sampling process.
In a small finite case, we could make this completely explicit, as we did throughout the main text.
Suppose the urn contained only ten balls, six red and four blue, and we drew three without replacement. Then we could lay out every possible sample of three balls and count how often each number of red balls occurs.
A sample with three red balls would appear in some of those possible samples. A sample with no red balls would appear in others. The probability distribution over the possible samples could be reconstructed directly from the finite arrangement.
If we now started from an observed sample instead, we could compare it with the distributions produced by different possible urn compositions. A composition under which samples like ours occur often fits the observation better than one under which they are rare.
For larger populations and samples we usually do not enumerate every possibility by hand. The mathematics compresses the same kind of relation. Once a population model and a sampling process are supplied, they constrain which samples are more or less likely.
The observed sample can then be compared with those possibilities.
But the application does not establish its own prerequisites. The sample may not have been produced in the assumed way, or the population represented in the model may not be the one relevant to the question.
These prerequisites still have to be sufficiently convincing for the procedure to apply.
Frequentist And Bayesian Uses Of Probability
In both frequentist examples, probability describes possible data under a setup that has been supplied, and the observation is assessed in relation to those possibilities.
In the randomization example, the six measured outcomes were already there. We held them fixed and asked what differences could have appeared if the treatment labels had been assigned differently. The randomization procedure supplied those possible assignments and their probabilities under the hypothesis of no effect.
In sampling the population is not fully known. We observe only a sample from it and ask how samples would tend to vary under different possible population values.
Randomization and sampling require different things to be in place before probability can do any work.
The randomization example needed experimental units, measured outcomes, treatment labels, and a procedure by which those labels were assigned.
The sampling example needed a population, together with a process by which observations were sampled from it. The urn picture makes that concrete: the balls stand for the population, and drawing balls stands for sampling.
Bayesian updating used probability differently. There, probability was assigned to competing hypotheses and redistributed when evidence arrived.
The frequentist and Bayesian approaches represent different kinds of uncertainty.
Randomization works with possible rearrangements of an experimental setup, while sampling starts from a population or generating process and a relation between it and the observations.
In both cases, the probabilities concern possible outcomes of a process taking place in the world: other assignments that could have occurred, or other samples that could have been drawn.
Bayesian updating needs a space of hypotheses and a way of relating each hypothesis to the evidence. Here the process being modeled is how inquiry changes the relative weight of competing hypotheses in light of evidence. Probability is used to represent something like the relative plausibility of those hypotheses. The possibilities being weighted are not alternative assignments or samples, but alternative accounts of what may be the case.
That difference changes how the observation is used.
In the randomization test we asked: if the treatment had no effect, how often would an assignment produce a difference at least this large?
In sampling we asked: if the population had a particular composition, how would samples from it tend to vary?
In Bayesian updating we asked how the evidence should redistribute weight among the represented hypotheses.
Bayesian updating has the central update rule we have already seen: probabilities are assigned to hypotheses and Bayes' theorem updates them from the data.
In frequentist inference, the data can count for or against a hypothesis or population value without that hypothesis or value itself being assigned a probability.
There is no single frequentist update rule corresponding to Bayes' theorem. Different procedures build probability distributions from different parts of the setup: random assignments, repeated samples, test statistics, estimators, and other constructions.
In the frequentist examples above, the constraint came from probabilities that could be traced back to represented possibilities generated by an experimental or sampling arrangement.
In the randomization case we could lay those possibilities out almost completely. In sampling, the same relation is usually compressed into mathematical models of how repeated samples would behave.
Bayesian updating puts probability on the hypotheses themselves.
Across the examples, the constraint came from representing relevant possibilities and assigning or updating their weights in ways that could be inspected.
Is any of these ways of using probability in inference better than the others?
It can be, depending on what we are asking of it. One procedure may be better than another for a specific problem by a criterion such as calibration, error control, predictive performance, transparency, computational cost, or how well its assumptions fit the problem.
But then the criterion itself can become part of the inquiry. Why should that criterion matter here? Does it conflict with another one? Does it still seem convincing once its consequences and assumptions are made explicit?
And that is exactly the kind of problem Part II has been following all along.
- cf. Phillip I. Good, Resampling Methods, 3rd Ed., Birkhäuser Boston, 2006, ch. 3.