I'm nowhere near as sophisticated in these matters, but I do have to get some introductory stats students off on a footing that lets me sleep at night, and I'm always trying to emphasize something along the lines of 'every statistic is just a description of the data,' and while it lets you have a relationship with uncertainty, there's ultimately not *anything at all* you can do computationally about *just not knowing something*. You can bound your uncertainty to the extant you can bound the system, and that's easy to do on a roulette wheel and impossible to do in a world full of intelligent, bounded actors keeping secrets and/or nature doing shit behind screens you don't get to peek behind. Show me a five-nines data center and I'll show you a building that got hacked on Thursday, much less in the past billion years or whatever.
i like this reminder that, given raw data, we indeed don't get anything at all. modeling choices have an inductive hypothesis (or similar assumption) baked in, perhaps that's a better topic of discussion than merely quantifying uncertainty!
~"We only care about the accuracy of whether the decision was right or not. We can use a held-out set as a benchmark."
> Fair enough. Are you going to measure just the overall accuracy, or would it be useful to break that down by the label (incorrect or correct)?
"By label would be good. One of the classes could be noticeably lower frequency or just harder to predict."
> OK. Next, how are you going to control whether the new data you see at test time resembles your benchmark close enough to not be meaningless?
"No idea. Probably we'll just YOLO it."
> In that case, why bother to do the benchmarking in the first place?
"Optics. Also, we're committed to Safety."
> I have so many questions about that. Just use this:
[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics. https://arxiv.org/abs/2509.12760
(And by the way, I agree that it may well be that "most uncertainty can't be quantified". That's the case when you construct the estimator in [1] and the points you're interested in are not assigned to a region, as well as the more extreme case where an SDM estimator can't even be constructed in the first place, because the task is under-specified, under-specifiable, or there is no data.)
If you try to be Bayesian about it, you can see the problem in another way. To have any chance of quantifying a probability distribution, you have to assume that the space of hypotheses you're looking at is exhaustive, which it obviously isn't. i.e., there are always unknown unknowns for which you can't account for in any good way. You can often get away with it in the right scenarios but often you shouldn't be able to.
Think about reporting the result of an opinon poll which comes at 51% A, 49% B. Reporting that "a majority supports A" is wrong. Sampling uncertainty gives you a "confidence interval "of, say ± 3 %,. Because of non-sampling errors, it is also not precisely correct to report 51% ± 3 %, but this isn't as big a problem as spurious certainty.
I'm nowhere near as sophisticated in these matters, but I do have to get some introductory stats students off on a footing that lets me sleep at night, and I'm always trying to emphasize something along the lines of 'every statistic is just a description of the data,' and while it lets you have a relationship with uncertainty, there's ultimately not *anything at all* you can do computationally about *just not knowing something*. You can bound your uncertainty to the extant you can bound the system, and that's easy to do on a roulette wheel and impossible to do in a world full of intelligent, bounded actors keeping secrets and/or nature doing shit behind screens you don't get to peek behind. Show me a five-nines data center and I'll show you a building that got hacked on Thursday, much less in the past billion years or whatever.
i like this reminder that, given raw data, we indeed don't get anything at all. modeling choices have an inductive hypothesis (or similar assumption) baked in, perhaps that's a better topic of discussion than merely quantifying uncertainty!
~"We only care about the accuracy of whether the decision was right or not. We can use a held-out set as a benchmark."
> Fair enough. Are you going to measure just the overall accuracy, or would it be useful to break that down by the label (incorrect or correct)?
"By label would be good. One of the classes could be noticeably lower frequency or just harder to predict."
> OK. Next, how are you going to control whether the new data you see at test time resembles your benchmark close enough to not be meaningless?
"No idea. Probably we'll just YOLO it."
> In that case, why bother to do the benchmarking in the first place?
"Optics. Also, we're committed to Safety."
> I have so many questions about that. Just use this:
[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics. https://arxiv.org/abs/2509.12760
(And by the way, I agree that it may well be that "most uncertainty can't be quantified". That's the case when you construct the estimator in [1] and the points you're interested in are not assigned to a region, as well as the more extreme case where an SDM estimator can't even be constructed in the first place, because the task is under-specified, under-specifiable, or there is no data.)
If you try to be Bayesian about it, you can see the problem in another way. To have any chance of quantifying a probability distribution, you have to assume that the space of hypotheses you're looking at is exhaustive, which it obviously isn't. i.e., there are always unknown unknowns for which you can't account for in any good way. You can often get away with it in the right scenarios but often you shouldn't be able to.
https://www.amazon.com/How-Measure-Anything-Intangibles-Business/dp/1118539273
Reporting uncertainty with spurious precision is better than reporting a single number with entirely spurious precision.
How so?
Think about reporting the result of an opinon poll which comes at 51% A, 49% B. Reporting that "a majority supports A" is wrong. Sampling uncertainty gives you a "confidence interval "of, say ± 3 %,. Because of non-sampling errors, it is also not precisely correct to report 51% ± 3 %, but this isn't as big a problem as spurious certainty.