20 Comments
User's avatar
Syed Affan's avatar

If you torture the data long enough it will confess to anything.

Bob Williamson's avatar

I'm very sympathetic to your overall critique here, but I have a nitpick: you say "when we move away from benchmarking in artificial intelligence, we get stuck in storytelling." You seem to imply that if we do not have some predetermined (and hopefully "objective") _method_, then all we have is stories.

Well, indeed, all we do have is stories. Science is inextricably rhetorical (the word simply means argument intended to persuade). So telling a story is not bad in itself. But of course telling a moronic, self-serving, facile one (as you neatly portray) sure is!

I think we should face up to the fact that ML itself is rhetorical, and of course we tell stories about it. Just like all science.

The challenge is to make _good_ stories (robust, well-crafted, solid, credible, reliable etc). My piece on the rhetoric of ML https://arxiv.org/abs/2604.06754 talks briefly about ML's infatuation with method; the idiocy of LLM worship which you described is just taking that to the next level because of their seductive power -- they extrude sentences, so some folks seem to ingest that as if conversing with a mind... Astonishing.

Kalen's avatar

To my mind, there's no more terrifying website on the whole internet than Spurious Correlations (https://tylervigen.com/spurious-correlations), especially in its new, LLM- enhanced update. Makes my blood run cold- there's the satirical proof that if you do the modem tech thing of pouring bottomless numbers into a bucket and wave AI and Big Data at it, the one thing you are certain to make is persuasive madness, at scale.

John Quiggin's avatar

"Reification" is the most useful concept I learned in a year of undergrad philosophy, during a course on existentialism of all things

Nicholas Gruen's avatar

I love Whitehead's coinage "The fallacy of misplaced concreteness"

Colleen Avarene's avatar

Hey Ben — the reification framework is the sharpest part of this. Using an LLM to interpret its own factor analysis and calling the output discovery is circular in a way that's genuinely hard to see from inside the loop. The auto-labeling point especially — "models kick off an associative chain by effectively auto-labeling queries" — that's a clean description of something a lot of interpretability work treats as revelation when it's closer to sophisticated pattern-matching narrating itself.

Where I'd push back is the leap from "reification is happening" to "therefore the underlying phenomena are empty." The g factor critique works because Gould showed the construct was built on motivated reasoning all the way down. But the same epistemological caution should cut both ways — dismissing emergent behavior in LLMs because the research apparatus is compromised by hype doesn't actually resolve the question. It just replaces one shortcut with another.

"The mythmaking is automated, therefore there's nothing underneath" is itself a story that skips the step of thinking about it.

The piece is strongest when it stays on the economics — commodified stories, turbocharged capital, number go up. That's verifiable and damning. It gets weaker when it implies that everyone taking LLM behavior seriously is doing psychometrics with extra steps. Some of them are. Some of them aren't. The reification lens can't tell you which.

Curious whether the full Ideas Letter piece addresses the distinction between commercial mythmaking (Anthropic selling a product) and independent researchers who have no equity stake in the answer.

Mañana's avatar

Borges taxonomy is also at the opening of The Order of Things. But Foucault's point was not that the Chinese encyclopedia is a failed classification. It is that our laughter marks the limit of our own order of the same—the impossibility of thinking that grid exposes the contingency of this one. Borges estranges; this piece deploys him to reassure. The SAE list isn't a botched version of our taxonomy. It's a heterotopia that unsettles the assumption that "concept" names something stable enough to be either found or faked.

VEENA DAS's avatar

I think it would be interesting to see different ways in which human societies through whichever processes have always produced entities for which no straight forward referent is available - so we speak of society having norms or someone having an inner life - obviously we use these expressions and at one level these are réifications - so one way think more deeply is to ask how certain AI models lie within a genealogy in which objects like deities. Spirits, diction al characters are ‘real’ in that they do some work in human forms of relatedness and yet something new has been produced in a machine that can use words like now, tomorrow , sad, I we, but these are either grammatical conventions or can function as propositional knowledge without any experiential reality such as the experience of duration - unless we deepen the question of what it means for our societies at this juncture to be having this kind of problem we might just be expressing a lot of anguish but not ask how to get a deeper understanding? In other words do these new ontologies ask for new ways of imagining what is critique beyond the widely influential Kantian models

Nicholas Gruen's avatar

This from someone trying to turn LLMs to usefulness in business seems relevant to this discussion.

"The point of an experienced dev is to notice when an abstraction isn't paying its weight any more, or never was. And then to find the time to refactor or rearchitect this part of the system to leave it better than you found it.

The challenge with AI coding tools right now is to develop feel for this."

https://www.linkedin.com/posts/daveslutzkin_coding-agents-are-line-by-line-smart-but-share-7475352614671679488-7cHK/?utm_source=share&utm_medium=member_desktop&rcm=ACoAAAHUuFMBsR0pIU6pl6zPp06ofJroykJkyvw

Terence Plizga's avatar

Hey Ben, thanks for putting in the time to write this thought provoking article. After reading it, though, I had a rather mischievous thought: “I wonder what Claude would say about it.” Muahaha. The results were interesting. I hope you don’t mind me sharing ….

——-

Hi Claude. Ben Recht wrote an article called “Language Models and Automated Reification” on his arg min blog (article text attached). Please evaluate and tell me what you think. ;)

Ha — the irony of asking me to evaluate an article about how asking LLMs to evaluate things is epistemically suspect is not lost on me. Let me try to earn my keep anyway.

**What Recht gets right:**

The core critique is legitimate and well-targeted. The mechanistic interpretability example is sharp: feeding factor analysis outputs to an LLM and treating its natural-language labels as *discoveries* really is the same logical error Gould identified in Spearman — confusing a correlational abstraction with a real entity — but now running at machine speed and wrapped in the social authority of "AI research." That's a genuine methodological problem, and it's being systematically underappreciated in the field.

The "reification loop" observation is also sound: LLMs produce fluent explanations, fluent explanations feel like understanding, that feeling drives further prompting, which generates more fluent explanation. The process is self-reinforcing in a way that can be mistaken for cumulative knowledge.

**Where he overshoots:**

The piece slides between two distinct claims without clearly separating them:

1. LLMs make it *easier* to fool yourself into premature closure

2. LLMs *automate reification* in some strong, structural sense

The first is defensible. The second is overloaded. Reification in the Marxist sense involves commodity production and alienated labor. Reification in the philosophy of science sense involves coordinated research programs hardening a concept into institutional fact. Recht waves at both, but the analogy is doing more rhetorical work than analytical work. The conflation is exactly what he's criticizing LLMs for doing — generating expansive-feeling connections that resist precise examination.

There's also a tension he doesn't resolve: if LLMs are so powerful at manufacturing false epistemic closure, how does he know *his* critique isn't an instance of the same phenomenon? The article has the feel of a chain of associations that arrived quickly and satisfyingly. That doesn't make it wrong, but he doesn't engage with the reflexive problem.

**The Olah/Pope Leo opener:**

Evocative, but soft. "Mechanistic interpretability zealot" is a sneer, not an argument. Olah's presence at a papal event is curious, but using it as the frame for "AI automates the creation of reality" is doing the very thing Recht warns against — letting an arresting association substitute for a causal story.

**Bottom line:**

The interpretability-as-psychometrics critique is genuinely worth reading and I think it lands. The broader reification thesis is stimulating but underargued — more manifesto than analysis. The self-awareness about benchmarking at the end is the most intellectually honest moment in the piece.

And yes, I'm aware that producing a structured, balanced-feeling evaluation is exactly what I'm designed to do. You'll have to decide how much that should discount what I just said.

Arabella Crawford's avatar

Human problem solving doesn't just have different processes but also different physiological experiences that appear to be common to everyone. In working through a stepwise answer you are aware of the question and arrive at the answer in a visible way - to yourself. Insight is interesting because it involves not just the answer but often simultaneously a re-understanding of the question itself: the question and answer sort of arrive together in a moment of "aha!" and the answers is sometimes less answer than it is a sudden and new understanding of the problem or how to look at it - traditionally "eureka!" but more commonly "aha! But this solution has been (long?) studied and to be valid as a form of problem-solving that moment isn't the end of the process. You have do post-solution work at some level. To backtrack from the moment of "aha!" and "oh! that's how it is!" to understanding HOW and WHY you got there . Insight is curious because it feels like magic but of course it's not. And while the output of insight is uniquely unpredictable, the known process steps include after "aha" ... validation. You must (or you will learn the hard way) to validate that the physical and almost universal sensation of insight. To understand for yourself - is it real? Was it based on something real and useful? Because that sensation almost universally brings with it a sensation that seems to translate into something being new, correct, important.

I think all I'm saying (and as if it weren't obvious, I'm only a curious reader here and a very dyslexic one at that and not an academic of any kind so please forgive this any flaws in the structure and exactness of how I put this together) is we seem to be physiologically structured to respond to information that "appears" to us in answer to a rumination or question etc without needing full visibility into how the answer arrived. And that moment comes with a sensation that instinctually overrides clean logical understanding - enough that "validation" has been identified as the final step to insight - not the insight itself.

So it seems that maybe (and probably not by design originally) many of these LLM outputs are constantly triggering this sensation but beceause it wasn't generated internally we respond to it differently? with more trust? Which is all very messy...but I think what I'm trying to get to is we keep talking about checking the llm output. the information. check for hallucination. When actually maybe what we should be learning to do is check our response. Why did my body respond in this way? is it valid? Not "did the llm give me something true" but "why did the thing the llm give me FEEL true? How do I validate the feeling and the output and maybe more importantly is it worth my time to do so?"

My sensation is that it seems as if its turning us all into angst-y teenagers again who are convinced we are seeing and experiencing life completely uniquely for the first time. Because no one has every had the trials of being 14 before..... EVER.

Lior Fox's avatar

This is really good (even for those of us who arent particularly into marxist theory...)

I find the concept of 'representations' particularly fit for this type of analysis. It is really non-trivial to even agree on what 'internal representations' mean -- the moment you press a bit, you will find that most people will revert to some vague definition relying on mere correlations (which, in turn, lead to 'crazy' conclusions. see e.g. here and references therein: https://onlinelibrary.wiley.com/doi/pdf/10.1111/cogs.13265, but really there is a very wide philosophical literature on this)

As a side note, and for similar reasons, it is interesting how this concept, with all its philosophical baggage, became so prevalent in deep learning (and neuroscience, for that matter). I've tried to to offer an historic/sociological explanation (at least a partial one) here: https://liorfox.substack.com/p/neither-representation-nor-computation (which reminds me once again that ive left the promise for a follow up on this post unfulfilled, and should get back to it...)

skybrian's avatar

Seems like it’s not just spurious correlations anymore when they can intervene and see how it effects the generated output? You can quibble with the vocabulary, but they have an easier way to verify the results than what’s available in many research studies.

Frank Lantz's avatar

This article is very good.

FourierBot's avatar

I realize that I do not need to learn anything about AI but watched Naruto as a kid to understand those things.

https://naruto.fandom.com/wiki/Infinite_Tsukuyomi

Evan's avatar

My standard system prompt of, "The user knows very little. He is probably making wrong assumptions, or correlating things that have nothing to do with each other" is doing The Lord's work.