Great piece--I really like this thread you're on. One thing this made me think of is the role of professional ethics and codes of ethics in more subjective approaches like PT. If our standard is "only things that pass RCTs are good interventions, so just do those" then we don't really need ethics. If we're opening up a role for judgment, the ethics of the person applying that judgment become important.
Broadly agree with your points here! I'm curious if you have thoughts on https://www.painscience.com/. I've overall found this site both interesting and helpful. It's "science-ish" but not dogmatic in my experience?
Definitely some good and some bad there. One of these days I'm going to write about chiropractic, and he makes a bunch of good points about it. But he's a bit weirdly dogmatic about stretching!
I'm curious what Paul would say about what Ben has written here. From where I stand, Paul is much better placed than Ben to be weighing in on this topic.
Nice thoughts. Reminded me that the "intention to treat" bucket of folks who opt-out creates effects very similar to "response-rate bias" in polls/surveys results. This phenomenon cuts across pretty much all non-coercive interventions (as dramatically played out during Nate Silver's downfall around predicting the 2016 election for Clinton based on polls with substantial response-rate biases, in that would-be Trump voters had lower response rate, biasing the survey results in favor of Clinton, e.g. in Michigan). Maturing how we handle this response heterogeneity (e.g. in some kind of Bayesian framework that integrates over some plausible range of response-rate-to-outcome correlation), might help across the board. Further, I'm not sure why it's hard to imagine that the suite of outcomes could include measures such as, "is your life better than it was before?" and take the responses seriously, especially in N-of-1 style trials. I tend to push back against the "EBM is flawed so the answer is to reduce the emphasis on trials". Is your take that we need to puncture EBM's cult-like and and somewhat simplistic over-emphasis within administrative/bureaucratic contexts?
Yes, exactly to your last question. For example, I am not against drug trials. But there, the bureaucratic nature is necessary to ensure that drugs are not harmful and potentially helpful before approving them for sale.
Thanks for an interesting stream of posts on this topic.
I know I'm risking heresy in this crowd, but it seems LLMs are better at capturing the kind of poorly-specifiable knowledge you're talking about here than classical ML and stats methods, perhaps because text can be more open-ended and vague.
Of course LLMs don't provide any more of a correctness guarantee than a Reddit thread does. I understand and agree with your criticisms of the bureaucratic use of statistics, but still in general I would trust a randomised trial more than an LLM recommendation.
The thumb rule I have followed is discard statistics + any claim of observational causal inference in these ITT scenarios but look at the purported mechanism of action and see if that mechanism can overwhelm my skepticism. The show-me-the-derivation unfortunately becomes show me a plausible explanation the more out of depth I am in that field. And these chatbots are ridiculously good in generating endless drivel of seemingly plausible explanations.
My impression is that "the purported mechanism" approach is pretty flawed itself, for example, the craze around antioxidant supplementation had a plausible mechanism of action but basically doesn't show any benefit.
Can't agree more - which is why it is best a guardrail for a skeptic. It works at least in one direction: If you show me crazy stats and causal inference but can't explain why, I won't believe you.
And the point above is "plausible" explanations are very cheap to generate now in a way that can fool a lot of people! And I get that your point is that even a plausible argument or better still a scientific argument (follows from premises as a logical argument, with each step in MoA verified) is still insufficient. May be the effect is too small and there is no benefit.
I think a lot of this boils down to a certain kind of 'sciencepilling' that takes statistics as the only good reason for believing something true is true. We saw it during the pandemic with masking- a certain breed of insufferable commentator would not that there weren't any RCTs saying masks would help. Well, the reason it's clear masks would help is that the thing that spread COVID was a little particle, and putting shit in the way of the particles would stop them. If anything the RCT was going to muddle the waters by producing a failure rate that consisted of idiots doing dumb things rather than the failure of the physics of the situation.
A math teacher I read (Michael Pershan) made a comment in one of his books differentiating between 'evidence-based' teaching and 'evidence-informed' teaching, suggesting the first is routinely both impossible and a nightmare and the second is what actually happens in a feedback-and-variable dense environment (which is pretty much anything involving human beings). He compared it to hiking with a map. The evidence of the map makes it clear that some approaches are likely to be better than others, with some ruled out categorically. But you still gotta hike, and your interests, physical capabilities, the weather and vegetation and rockfall, are going to be dictating where your feet go, and that's as it should be.
There's an interesting conflict between the examples in the first and second paragraphs. Masking is about preventative medicine. You can never say whether or not it was "the mask" that prevented or failed to prevent your infection. By contrast, in teaching, just like in therapy, you stitch together a train of signals that things are improving. I want a better scientific/engineering language for those sussing out improvement through iterative feedback.
I agree with your overall message. And yes, opioids/opiates (natural or synthetic) are not universally helpful in musculoskeletal pain. My lumbar pain is an affirmation of that.
But regarding the combinatorial explosion of treatments. I thought Latin Squares were a method to reduce the number of treatment variables. Shouldn't it work to reduce treatment tests even with more than a couple of variables?
In addition, Judea Pearl's "The Book of Why" demonstrates that one can make inferences even with reduced data on variables. (which seems similar to the Latin Squares method but applied to incomplete coverage by the available data). I have no reason to doubt his group's mathematical rigor. Is he wrong?
Great piece--I really like this thread you're on. One thing this made me think of is the role of professional ethics and codes of ethics in more subjective approaches like PT. If our standard is "only things that pass RCTs are good interventions, so just do those" then we don't really need ethics. If we're opening up a role for judgment, the ethics of the person applying that judgment become important.
Broadly agree with your points here! I'm curious if you have thoughts on https://www.painscience.com/. I've overall found this site both interesting and helpful. It's "science-ish" but not dogmatic in my experience?
Definitely some good and some bad there. One of these days I'm going to write about chiropractic, and he makes a bunch of good points about it. But he's a bit weirdly dogmatic about stretching!
I'm curious what Paul would say about what Ben has written here. From where I stand, Paul is much better placed than Ben to be weighing in on this topic.
Nice thoughts. Reminded me that the "intention to treat" bucket of folks who opt-out creates effects very similar to "response-rate bias" in polls/surveys results. This phenomenon cuts across pretty much all non-coercive interventions (as dramatically played out during Nate Silver's downfall around predicting the 2016 election for Clinton based on polls with substantial response-rate biases, in that would-be Trump voters had lower response rate, biasing the survey results in favor of Clinton, e.g. in Michigan). Maturing how we handle this response heterogeneity (e.g. in some kind of Bayesian framework that integrates over some plausible range of response-rate-to-outcome correlation), might help across the board. Further, I'm not sure why it's hard to imagine that the suite of outcomes could include measures such as, "is your life better than it was before?" and take the responses seriously, especially in N-of-1 style trials. I tend to push back against the "EBM is flawed so the answer is to reduce the emphasis on trials". Is your take that we need to puncture EBM's cult-like and and somewhat simplistic over-emphasis within administrative/bureaucratic contexts?
Yes, exactly to your last question. For example, I am not against drug trials. But there, the bureaucratic nature is necessary to ensure that drugs are not harmful and potentially helpful before approving them for sale.
Thanks for an interesting stream of posts on this topic.
I know I'm risking heresy in this crowd, but it seems LLMs are better at capturing the kind of poorly-specifiable knowledge you're talking about here than classical ML and stats methods, perhaps because text can be more open-ended and vague.
Of course LLMs don't provide any more of a correctness guarantee than a Reddit thread does. I understand and agree with your criticisms of the bureaucratic use of statistics, but still in general I would trust a randomised trial more than an LLM recommendation.
The thumb rule I have followed is discard statistics + any claim of observational causal inference in these ITT scenarios but look at the purported mechanism of action and see if that mechanism can overwhelm my skepticism. The show-me-the-derivation unfortunately becomes show me a plausible explanation the more out of depth I am in that field. And these chatbots are ridiculously good in generating endless drivel of seemingly plausible explanations.
My impression is that "the purported mechanism" approach is pretty flawed itself, for example, the craze around antioxidant supplementation had a plausible mechanism of action but basically doesn't show any benefit.
Can't agree more - which is why it is best a guardrail for a skeptic. It works at least in one direction: If you show me crazy stats and causal inference but can't explain why, I won't believe you.
And the point above is "plausible" explanations are very cheap to generate now in a way that can fool a lot of people! And I get that your point is that even a plausible argument or better still a scientific argument (follows from premises as a logical argument, with each step in MoA verified) is still insufficient. May be the effect is too small and there is no benefit.
I think a lot of this boils down to a certain kind of 'sciencepilling' that takes statistics as the only good reason for believing something true is true. We saw it during the pandemic with masking- a certain breed of insufferable commentator would not that there weren't any RCTs saying masks would help. Well, the reason it's clear masks would help is that the thing that spread COVID was a little particle, and putting shit in the way of the particles would stop them. If anything the RCT was going to muddle the waters by producing a failure rate that consisted of idiots doing dumb things rather than the failure of the physics of the situation.
A math teacher I read (Michael Pershan) made a comment in one of his books differentiating between 'evidence-based' teaching and 'evidence-informed' teaching, suggesting the first is routinely both impossible and a nightmare and the second is what actually happens in a feedback-and-variable dense environment (which is pretty much anything involving human beings). He compared it to hiking with a map. The evidence of the map makes it clear that some approaches are likely to be better than others, with some ruled out categorically. But you still gotta hike, and your interests, physical capabilities, the weather and vegetation and rockfall, are going to be dictating where your feet go, and that's as it should be.
There's an interesting conflict between the examples in the first and second paragraphs. Masking is about preventative medicine. You can never say whether or not it was "the mask" that prevented or failed to prevent your infection. By contrast, in teaching, just like in therapy, you stitch together a train of signals that things are improving. I want a better scientific/engineering language for those sussing out improvement through iterative feedback.
I agree with your overall message. And yes, opioids/opiates (natural or synthetic) are not universally helpful in musculoskeletal pain. My lumbar pain is an affirmation of that.
But regarding the combinatorial explosion of treatments. I thought Latin Squares were a method to reduce the number of treatment variables. Shouldn't it work to reduce treatment tests even with more than a couple of variables?
In addition, Judea Pearl's "The Book of Why" demonstrates that one can make inferences even with reduced data on variables. (which seems similar to the Latin Squares method but applied to incomplete coverage by the available data). I have no reason to doubt his group's mathematical rigor. Is he wrong?
The issue is that with each additional treatment or branch, you increase the required sample size to satisfy the statistical bureaucrats.