Hi there, argmin readers! As the fall semester picks up, posting volume will, too. So I’m going to commit to writing short descriptive headers to help you sort through the different threads. Today’s post is a live blog of Class 6 of my graduate seminar “Forecasting: A Critical Retrospective.” A table of contents is here.
Given the week’s events, it’s a bit unfortunate that I scheduled our discussion of p(doom) for the last week of class. I predict AI won’t have killed us by then, and the real question is whether we’ll all be bored to tears discussing the topic in November. But the agenda for today, earthquakes, is a good preview for the challenges associated with quantifying uncertainties about catastrophe. Seismologists don’t think an earthquake will lead to human extinction, but it can cause massive casualties and damage. How do we quantify our predictions of whether an earthquake will happen? And then what do we do about it?
Most experts agree that predicting the exact time and location of earthquakes on long time horizons is impossible. The dynamics of the Earth moving, building up stress, and slipping are far too complicated to predict with any reasonable granularity using differential equation models. Earthquake forecasting couldn’t be further removed from weather forecasting in that regard.
At best, we can make coarse predictions based on a mix of temporal and spatial localization. Earthquakes tend to occur near fault lines. Fault lines have a history of previous ruptures of different sizes. Using these data, we can estimate rough statistical models. You might naively estimate an exponential recurrence time: the rate at which earthquakes occur is just the count divided by the observation window. In an exponential model, the expected time to the next earthquake would be the inverse of this number. A slightly more complicated formula then gives you the chance of an earthquake in the next decade.
chance = 1 - np.exp( - rate * time )Such primitive models are not precise, but they are helpful. What do you do with these probabilities? You can turn them into general warnings. If you expect a certain frequency of shaking, you should build infrastructure that can withstand it and teach people how to prepare for the disruption the next one will cause. If you know big earthquakes occur every few decades, that’s enough to inform planning and insurance.
But nailing down the probability of an earthquake, even to one decimal place, is a fool’s errand. Our first reading of the week, Freedman and Stark’s classic paper “What is the Chance of an Earthquake?”, highlights the futility of precise probability models. If you want to validate a probabilistic forecast, you need a lot of events. The law of large numbers needs a lot of numbers! Large earthquakes are rare. Probabilistic models can’t be tested on human time scales. Moreover, when you add more geological reality to your model, you introduce a variety of hard-to-estimate parameters and researcher degrees of freedom into the equations. Every new modeling assumption introduces new unidentifiable parameters. More realistic doesn’t mean better estimates.
If you want to predict really big earthquakes, like those with magnitudes greater than 8.5, then we have an even sparser record. The old-fashioned AI chatbot, Wikipedia, has dozens of tables listing earthquakes by all sorts of characteristics. It lists only 17 of these in the past hundred years. Scientists have developed techniques to infer the occurrence of giant earthquakes thousands of years in the past. These tend to give noisier estimates of recurrence times, but sometimes they yield very ominous predictions.
One of the most ominous is in this week’s reading, “The Really Big One,” a riveting 2015 New Yorker article by Kathryn Schulz. Schulz reports on the Cascadia subduction zone, a thousand-mile fault that runs from Northern California to Vancouver Island. Combining oral history, Japanese tsunami records, and tree rings, seismologists determined that a massive earthquake, with a magnitude pinned between 8.7 and 9.2 on the Richter scale, happened on this fault on the evening of January 26, 1700. It killed coastal forests of the Pacific Northwest and created a massive tsunami in Japan. Oral histories from First Nations tell of entire communities vanishing. Scientists have gone back to geological samples and counted 41 major earthquakes on this fault in the last ten thousand years. Using the rough rule of thumb, we should expect a major, destructive earthquake once every 243 years. It’s been 326 years since the last one.
Now, you could try to guess the probability that an earthquake occurs on this fault before 2050, but that number doesn’t really do much of anything for you. We don’t know when it will occur, but we know an earthquake is inevitable here, and we know it will be catastrophic.
Shulz details some predictive horror stories of what will happen when the next big one hits the Cascadia Subduction Zone. It does seem like a bad idea to put millions of people near such a seismically volatile region. But this is the problem with our slow ape brains. As Shutz writes, “[forty] years ago, no one knew that the Cascadia subduction zone had ever produced a major earthquake. [Fifty-five] years ago, no one even knew it existed.” In 1970, Seattle was already a major city with over half a million people.
So the question is, what do we do now? The low end of state estimates of fatalities from the next major earthquake is in the tens of thousands. One answer would be to move millions of people away from the danger zone. No one is proposing this. The other is to build as much infrastructure as possible to handle the incoming crisis through seismic retrofitting and social infrastructure for tsunami evacuation protocols and earthquake preparedness. The work involves building systems to keep damage as small as possible, even though the damage will be unavoidably large. As Freedman and Stark say, “probabilities are a distraction.”


To play devil's advocate here, because as you know I'm also a numerical probability skeptic: the numerical probability aficionado might say something like "at some point, with limited resources, you have to decide whether to build that infrastructure in Seattle to hedge against Cascadia or in St. Louis to hedge against New Madrid, and what's your basis for decisions about allocation?"