
Hi there, argmin readers! Today’s post is a live blog of Class 11 of my graduate seminar “Forecasting: A Critical Retrospective.” The syllabus and list of past posts are here.
In “Communicating Uncertainty in Policy Analysis,” Charles Manski takes to the Proceedings of the National Academy of Sciences to chastise government officials for their failure to report uncertainty in their forecasts. Now, what would that reporting look like? We need the uncertainty reports to be clear, legible, and objective. Being good bureaucrats, those desiderata tell us we need to quantify our uncertainty.
Sadly, our technocratic imagination for what uncertainty quantification means is woefully narrow. Uncertainty quantification almost exclusively means error bars. In this case of forecasts, an error bar specifically means a prediction interval. A prediction interval is a probabilistic object. It quantizes a forecast of some numerical quantity into a discrete binary event. Rather than saying I think GDP will grow by 3% in the first quarter of 2027, I say, “The chance that GDP will grow between 0 and 7% in the first quarter of 2027 is 95%.” Based on your modeling, you believe some measurement will lie in some band with some probability.
Even our methods to construct these prediction intervals lack creativity. 95% of the time, we build those intervals by assuming the data is Gaussian, estimating the variance, and then setting the bounds to be plus or minus two sigma. That variance estimate might come from direct analysis of the model. For example, if you assume linear dynamics perturbed by Gaussian shocks, you can compute the variance exactly.
For more complicated models, you might get the variance from Monte Carlo simulation or estimate it from past data. For example, in weather forecasting, you could generate error bars by sampling a bunch of likely atmospheric states, propagating them through a complex simulation, and then computing quantiles. Other methods bin past prediction errors, arguing that this isolates aleatoric uncertainty from epistemic uncertainty, then build error bars from those so-called “innovations.”1 If you want to be really fancy and pretend you are avoiding assumptions about distributions, you can just compute 95% quantiles directly from the observed errors.
And that’s more or less all of the tricks we know. How many classes must we spend on it? How many papers must we write about it?2 In this class, we’ll only use one.
I’m always left with the uncomfortable problem that I have no idea what to do with probabilistic error bars. Usually what “probability” means is incredibly sloppily specified in these models. Even in the constructions I described above, the probability is dependent on an unverifiable chain of modeling decisions, and there’s no way to doubly quantify the uncertainty in my modeling. A good engineer will just multiply their error bars by 2 and call it a day.
For better or for worse, however, forecasters are not always engineers. They just might want to use the uncertainty to hedge their bets. If they turn their numerical estimate into a probabilistic one, they can’t be wrong. Conveniently, probabilistic predictions of binary events are more easily scored by our standard proper scoring rules.
Outside narrow forecasting contests, it’s unclear how reporting an interval changes policy decisions. In weather forecasting, you can give people a sense of whether they should bring an umbrella. But error bars on growth are harder to parse from a decision-theoretic standpoint. Yet these are, perhaps unsurprisingly, exactly the sorts of uncertainty reports Manski calls for:
“For example, the CBO could report the 0.10 and 0.90 quantiles of the distribution of potential outcomes that it referenced when scoring the American Health Care Act of 2017. Alternately, it could present a full probabilistic forecast in a graphical fan chart, such as the Bank of England uses to predict GDP growth (see the discussion later in this article).”
Perhaps you can say that all that matters is the sign of GDP growth: negative growth is bad, and positive growth is good. This leads to a self-fulfilling prophecy in macroeconomic planning. When they estimate a negative number, economists declare a recession, and everyone gets mad. The problem, of course, is that many countries have positive GDP growth right now, and people are still pissed off. Making policy where their forecast sign is correct didn’t solve the current administration’s political problems.
Anyway, I’ve written about my disdain for this blindered approach to uncertainty quantification before, and every time I come back to how we always want something else. People want a prefactual analysis for what we should do if our story is wrong. They want to know how we will recover from failures, and articulating the vast complexity of uncertainty might help with such preparation. Halfway through this class on forecasting, my knee-jerk conclusion is that these are what most serious forecasters want too! In Manski’s long list of sources of uncertainty, only a couple can be quantified. However, holistic reporting of uncertainty can prepare us for what the model doesn’t say and help us think about what to do when our forecasts inevitably miss the mark.
Innovation is such a wild name for prediction error.
A Google Scholar search of “conformal prediction” says the answer is thousands.

I'm nowhere near as sophisticated in these matters, but I do have to get some introductory stats students off on a footing that lets me sleep at night, and I'm always trying to emphasize something along the lines of 'every statistic is just a description of the data,' and while it lets you have a relationship with uncertainty, there's ultimately not *anything at all* you can do computationally about *just not knowing something*. You can bound your uncertainty to the extant you can bound the system, and that's easy to do on a roulette wheel and impossible to do in a world full of intelligent, bounded actors keeping secrets and/or nature doing shit behind screens you don't get to peek behind. Show me a five-nines data center and I'll show you a building that got hacked on Thursday, much less in the past billion years or whatever.
~"We only care about the accuracy of whether the decision was right or not. We can use a held-out set as a benchmark."
> Fair enough. Are you going to measure just the overall accuracy, or would it be useful to break that down by the label (incorrect or correct)?
"By label would be good. One of the classes could be noticeably lower frequency or just harder to predict."
> OK. Next, how are you going to control whether the new data you see at test time resembles your benchmark close enough to not be meaningless?
"No idea. Probably we'll just YOLO it."
> In that case, why bother to do the benchmarking in the first place?
"Optics. Also, we're committed to Safety."
> I have so many questions about that. Just use this:
[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics. https://arxiv.org/abs/2509.12760
(And by the way, I agree that it may well be that "most uncertainty can't be quantified". That's the case when you construct the estimator in [1] and the points you're interested in are not assigned to a region, as well as the more extreme case where an SDM estimator can't even be constructed in the first place, because the task is under-specified, under-specifiable, or there is no data.)