Hi there, argmin readers! As the fall semester picks up, posting volume will, too. So I’m going to commit to writing short descriptive headers to help you sort through the different threads. Today’s post is a live blog of Class 5 of my graduate seminar “Forecasting: A Critical Retrospective.” A table of contents is here.
The fun thing about teaching a new class is realizing in week three how your sequencing is already off. I do this every year, so I’m no longer surprised. At this point in my career, I’d put the probability that I’ll be disappointed with my syllabus before the official add-course deadline at 95%.
How did I get to that number?
So yes, though it’s not directly about forecasting, the first mathematical lecture of this class should have been about probability. Because probability is unavoidable in forecasting, and it is used far too casually for my tastes. We saw this in the weather example where outputs of data assimilation were interpreted as “probability distributions” over the state of the atmosphere. Samples from this distribution were used to create a sample of future possibilities. Frequencies of these future possibilities were turned into beliefs about whether it will rain on Sunday.
All of these probabilities are sort of different, right? Some are about counts, some are about beliefs, and somehow we move between the two as if there is a well-specified set of formal rules for doing so.
I think every class on applied probability, including the undergrad ones, should call out this slippery transmutation between frequency and belief, and I’m particularly fond of the development in Paul Meehl’s course on Philosophical Psychology.1 Meehl follows Rudolf Carnap, who explicitly distinguishes the two kinds of probability.
Probability 1, sometimes called logical probability, relates propositions with beliefs. It quantifies how much credence we should give a hypothesis based on the observations and facts laid before us. Probability 1 represents the certainty of a logical proposition being true, the doubt that something in the past happened, the likelihood that future events will occur, or a person’s internal beliefs about the world.
Probability 2 concerns frequencies of events. It more or less just amounts to counting, measuring the relative frequency of some property in a set of objects. Probability 2 might quantify the relative frequencies of occurrences that currently exist in the world, but it can also describe the relative frequencies of hypothetical infinite populations.
Probability 2 is usually the one we start with and teach and never problematize, and then we casually jump to using Probability 1 in our thinking, writing, and experimenting without realizing it. This is because they both obey the same axioms of probability.
Let’s say I have a bunch of marbles in an urn. I want to describe the frequencies with which certain properties of those marbles hold. The following things are true for any subset of marbles:
Any subset has proportion greater than or equal to 0
If I take the entire set, the proportion equals 1.
If I take two nonoverlapping subsets, the proportion of their union is equal to the sum of their proportions.
These are Kolmogorov’s axioms of probability. Item 1 is nonnegativity, Item 2 is unit measure, Item 3 is additivity.
Probability 2 obviously obeys Kolmogorov’s axioms, but what about Probability 1? Well, we can sort of force it to have those properties too.
Any syntactically valid statement has probability greater than or equal to zero.
Any statement that is certain or tautological has probability 1.
If two statements describe mutually exclusive outcomes, then the probability of either one or both of the outcomes is equal to the sum of the individual probabilities.
Lo and behold, statements of belief seem to also obey Kolmogorov’s Axioms. These assertions might feel a bit less obvious and certain than marble counting. You can’t get a tangible grip on why beliefs should obey Kolmogorov’s axioms because beliefs live in your head. You can check the properties of Probability 2 on a set of marbles. You can never check whether the axioms work for Probability 1.
But there are strategic reasons for using probabilities when quantifying beliefs. In last Thursday’s lecture, we saw that if you are scored using the Brier Score, then when your forecasts don’t obey the axioms of probability, there is always another set of forecasts that achieves a higher place on the leaderboard. If instead of being a forecaster, you’re a degenerate gambler, there are arguments about dealing with bookies that motivate the same probabilistic rules. As Dennis Lindley put it, probability is inevitable once you try to quantitatively evaluate belief.
I think we should also distinguish a third kind of probability, Probability 0, to describe the formal language of mathematical probability. This is probability whose referent is neither frequencies nor beliefs but mathematics itself. In this case, we’d consider the following to all be Probability 0:
Any finite list of numbers that is nonnegative and sums to one
Any infinite list of numbers that is nonnegative and whose infinite sum converges to one
Any nonnegative function on the unit interval whose integral is equal to one
These sorts of mathematical structures arise a lot when we’re doing calculations. And whenever people find a convenient mathematical application of Probability 0, they tend to find a convenient application of that structure in Probability 1 or Probability 2. We see Probability 0 everywhere. It only becomes metaphysical when we attach some meaning to it off the chalkboard.
When we build computational forecasts, we have to play with all three kinds of probability. We saw this last time: We need Probability 2 because the best predictions correspond to the rates of outcomes in similar future events. We need Probability 1, or else our forecasts are incoherent. We need Probability 0 to write and reason about algorithms that analyze noisy data. But how we tie the three together is not given. There is no god-given algorithm of prognostication that we can derive from Kolmogorov’s axioms, and techniques and conventions vary between disciplines. Which scoring rule are we using? Are we insisting on building calibrated forecasts? What do our forecasts do? These questions determine how we work with probability.


I've often meant to ask whether you've come across Stephen Toulmin's discussion of probability in ch 2 of his Uses Of Argument, and if so what you think of it. "To say that a statement is a probability-statement is *not* to imply that there is some one thing which it can be said to be about or express. There is no single answer to the questions, ‘What do probability-statements express?...Some express one thing: some another" (p65) "whether backed by mathematical calculations or no, the characteristic function of our particular, practical probability-statements is to present *guarded* or *qualified* assertions and conclusions." (p86) (https://johnnywalters.weebly.com/uploads/1/3/3/5/13358288/toulmin-the-uses-of-argument_1.pdf)
"Yeah man, they call gambling a disease, but it’s the only disease where you can win a bunch of money." -Good ol' Norm