Discussion about this post

User's avatar
Kalen's avatar

I'm nowhere near as sophisticated in these matters, but I do have to get some introductory stats students off on a footing that lets me sleep at night, and I'm always trying to emphasize something along the lines of 'every statistic is just a description of the data,' and while it lets you have a relationship with uncertainty, there's ultimately not *anything at all* you can do computationally about *just not knowing something*. You can bound your uncertainty to the extant you can bound the system, and that's easy to do on a roulette wheel and impossible to do in a world full of intelligent, bounded actors keeping secrets and/or nature doing shit behind screens you don't get to peek behind. Show me a five-nines data center and I'll show you a building that got hacked on Thursday, much less in the past billion years or whatever.

Allen Schmaltz's avatar

~"We only care about the accuracy of whether the decision was right or not. We can use a held-out set as a benchmark."

> Fair enough. Are you going to measure just the overall accuracy, or would it be useful to break that down by the label (incorrect or correct)?

"By label would be good. One of the classes could be noticeably lower frequency or just harder to predict."

> OK. Next, how are you going to control whether the new data you see at test time resembles your benchmark close enough to not be meaningless?

"No idea. Probably we'll just YOLO it."

> In that case, why bother to do the benchmarking in the first place?

"Optics. Also, we're committed to Safety."

> I have so many questions about that. Just use this:

[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics. https://arxiv.org/abs/2509.12760

(And by the way, I agree that it may well be that "most uncertainty can't be quantified". That's the case when you construct the estimator in [1] and the points you're interested in are not assigned to a region, as well as the more extreme case where an SDM estimator can't even be constructed in the first place, because the task is under-specified, under-specifiable, or there is no data.)

6 more comments...

No posts

Ready for more?