5. What Probability Adds
We can now return to the question that motivated the investigation.
Where did the probabilistic constraint come from?
The answer changed as we moved from dice to hypotheses.
Possibilities
The die began with a represented possibility space:
$$ \Omega={1,2,3,4,5,6}. $$
That representation was constructed.
We chose what the experiment was and what would count as its outcome.
But once those choices remained in place, the possibility space was easy to inspect.
The outcomes excluded one another. Their exhaustiveness for the specified experiment could be supported by examining the object itself.
The machine example exposed the same dependence under less favorable conditions. A hypothesis absent from the model cannot acquire posterior probability until the model changes.
Possibility generation remains outside the calculation even when the calculation itself is impeccable.
Weight
The loaded die exposed another source.
Counting six possibilities did not establish that each had probability:
$$ \frac16. $$
Equal weighting required another conviction. In the fair-die case, symmetry and the setup made that weighting easy to accept. With hypotheses, it became harder to see where the weights came from. A prior might come from frequencies, another model, or expert judgment.
The probability notation can represent all of these. It does not preserve the path by which their convincing force arose.
This is one thing probability adds to the earlier investigation of logic.
A possibility need not simply remain or disappear. Several possibilities can stay live while carrying different weights. The space of possibilities can remain open while becoming constrained internally.
Counting And Proportion
In the simplest cases, probability inherited considerable force from counting.
Three equally weighted even results among six gave:
$$ \frac36. $$
One path among 36 gave:
$$ \frac1{36}. $$
These calculations were convincing because the relevant arrangements could themselves be reconstructed.
But the counting depended on what came before it. The possibility space had to be adequate for the question.
And where simple counting was converted into probability by equal weighting, that weighting required its own support.
The loaded die made the division visible.
Conditional Structure
Repeated dice gave us a stable next stage.
Whatever happened first, the second fair throw retained the same distribution.
We later compressed this into:
$$ P(B\mid A)=P(B). $$
Cards without replacement broke that stability. The probability of the next ace changed depending on what had already happened.
That forced a distinction between:
$$ P(B) $$
and:
$$ P(B\mid A). $$
The more general relation:
$$ P(A\land B)=P(A)P(B\mid A) $$
then showed the earlier independent rule as a special case.
Discovering the limit did not make the simpler relation less convincing. It made clearer what had to remain unchanged for the simple multiplication rule to apply.
Compression
Probability also repeated a pattern already encountered in logic.
With two dice, the whole structure remained easy to inspect. We could draw the 36 possibilities and count the relevant paths.
With more throws, exhaustive inspection quickly became inconvenient.
Ten throws already produced $6^{10}$ paths. The local relation could nevertheless remain stable.
So the tree was compressed into:
$$ \left(\frac16\right)^{10}. $$
Conditional probability let us preserve the structure when later probabilities depended on what had happened before.
Bayes' theorem then rearranged the same joint relation into a reusable formula.
A trained user does not need to reconstruct the whole tree each time. If the compressed rule becomes doubtful or needs explanation, however, the earlier construction is still available for examination.
Bayesian Redistribution
The move to hypotheses then added something else.
Suppose several explanations remain live. We represent them, assign priors, represent an observation, and supply likelihoods. An event can then redistribute the probability assigned to those hypotheses.
Given the model, the posterior is constrained. The user does not separately choose how much probability each hypothesis will receive after the update.
That makes Bayesian updating interesting for Conviction Formation Theory.
Several possibilities can remain live, while the procedure constrains their relative weights in a way the agent cannot simply choose.
The agent participates by constructing the representation, selecting or accepting inputs, and initiating the calculation. Once enough has been fixed, however, the formal consequence is no longer independently free.
Where The Constraint Stops
The same investigation also made the limits easier to locate.
The update did not establish that our three machine hypotheses were exhaustive, that the prior or likelihoods were well supported, or that the observation and dependence relations had been represented appropriately. Producing a posterior also did not establish that the model suited the problem.
Those things had to become convincing elsewhere, through observation, experience, measurement, other models, judgment, training, or some combination of them.
They can later become objects of inquiry themselves. They can be recursively examined.
Making these dependencies explicit does something else we already saw in logic.
It makes questions available.
Why these possibilities?
Why these weights?
Why this likelihood?
Is an important hypothesis missing?
Has a dependence been represented correctly?
We can ask such questions deliberately. Whether they then become genuinely puzzling is another matter. Trying to answer one may expose that a number we had been willing to use has little support, or that a possibility space we had treated as adequate is harder to defend than we expected.
Probability participates in conviction formation not only through the constraint of its calculations. Its formalization can also expose places from which further inquiry begins.
Once the model is in place, its probabilistic consequences can be extremely constraining.
But strength inside the representation does not travel backward and establish every condition from which the representation began.
Two Kinds Of Uncertainty
This gives us a useful distinction.
A probability distribution can represent uncertainty among possibilities inside a model.
Suppose:
$$ P(H_1)=0.50, \quad P(H_2)=0.30, \quad P(H_3)=0.20. $$
That distribution represents uncertainty among:
$$ H_1,H_2,H_3. $$
It does not automatically represent our uncertainty about whether those were the right hypotheses to include.
Maybe $H_4$ is missing.
The existing distribution cannot give $H_4$ probability until the model is changed.
Uncertainty inside the model is not the same thing as uncertainty about the model.
Bayesian updating can be exact about the first while the second remains unresolved.
We can represent some uncertainty about the model inside a larger model. But then the same question returns at the next level: how convincing is that larger model?
Constraint Without Final Authority
The result resembles what we found in logic.
Logic did not supply its own premises, interpretations, or every feature of its representation.
But once enough remained fixed, some possibilities disappeared.
Probability likewise receives possibilities, weights, likelihoods, representations, and dependence assumptions.
Once enough remains fixed, the permissible redistribution of weight narrows sharply.
In both cases, the supplied conditions can remain open to recursive examination without leaving every consequence equally open.
Probability does not stand outside conviction formation and dictate what a person must believe. A mathematically impeccable posterior may fail to convince. An invalid probability model may convince strongly.
Probability can keep several possibilities open while making their relative weights explicit and constraining how those weights change with an observation. The assumptions on which the calculation depends can in turn become objects of inquiry.
We can choose to enter the procedure. We cannot simply choose what follows from it.