1.5.1 Sharing Predictions
1.5.1.1 Our thinkers have internalized real patterns; this allows them to make informed predictions and in doing so they’ve become more robust. Each thinker has a unique perspective that is both physical—their position in the environment—and internal—the history that brought them there, consisting of the real patterns, unique to them, that help them anticipate what comes next.
If each organism carries a unique, compressed model of a world too complex to be represented by any one thinker, then there is an evolutionary incentive to share predictions. The asymmetry creates collective value—aggregating perspectives increases the group’s robustness. So it is natural to imagine our thinkers becoming speakers, the expected next step in the climb to higher-order complexity. Not surprisingly, the advent of language is widely treated as a major evolutionary transition (Szathmáry and Smith 1995, 231), a phase change to a higher order of complexity.
In an environment above the inductive threshold, where no creature can deduce the future from fixed rules, groups that exchange predictions gain a decisive advantage in navigating the adjacent possible. Language, in this view, emerges as a mechanism for pooling patterns: a way of broadcasting expectations, coordinating predictions, and collectively reducing uncertainty. Under this view, far from beginning as a system of propositions describing facts, language originates as an evolutionary tool for synchronizing the predictive machinery of multiple organisms in a world too large for any one of them to model alone.
If our speakers lived in a world that could be fully compressed into explicit rules—that is, if they operated below the inductive threshold—then each could internalize the relevant structure directly. The value of exchanging perspectives would shrink. But the existence of speakers suggests something else: a world rich enough that no single set of rules is practical and where perspectives become valuable. In a complex environment, each speaker sees something the others do not. Language, in this way, is evidence the world exceeds our individual capacity to fully understand it.
What follows is a walk through language in our complexity terms, to see what it might tell us about meaning, intentionality, and how we might think about what artificial intelligence does and does not do.
1.5.2 Lexical and Pragmatic Closure
1.5.2.1 Wittgenstein distinguishes between ostensive and verbal definitions. An ostensive definition might be: I point to a book—this is what I mean by “book” (Wittgenstein 2009, 17-18). A verbal definition, on the other hand, is constructed from other words. And many words are defined primarily through their relations to other words (Wittgenstein 1958, 1).
Verbal definitions appear circular—words pointing to words. But we’ve seen this pattern before. An autocatalytic set is also circular: a network of reactions that sustains and reproduces itself without a privileged starting molecule. Once critical mass forms—under the right conditions, within an envelope—the network crosses a threshold and closure happens. Picture a Buckyball of words pointing to other words.
“It is not single axioms that strike me as obvious, but rather a system in which consequences and premises support each other” (Wittgenstein 1969, 22).
The same is true for words. Verbal definitions form a network of words pointing to other words, tokens pointing to other tokens. Meaning arises when those relations become rich and stable enough to support reliable inference—when the pattern becomes real enough to carry predictive power for the speaker. Let’s call this process lexical closure—basically semantic closure, but that term has a history.
Words have meaning through their relations to other words. Meaning is relational, historical, distributed, correlated, and often distant. There is no first meaningful word, just as there is no first molecule in an organism. And while closure is a form of completion—a word now having predictive power—that’s not to say that the meaning cannot change or grow over time.
1.5.2.2 Now let’s return to ostension. The advantage of being in the world isn’t merely that we can point—it’s that we can be corrected or negated. We try, fail, and try again. We play what Wittgenstein called a language game (Wittgenstein 2009, 15) with other speakers, and we learn this game “purely practically, without explicit rules” (Wittgenstein 1969, 16). The language game is the full nuance of language: tone, timing, glances, and what counts as understanding or misunderstanding.
The language game is prediction under negation. We make moves, which are then rejected or silently ratified by success. Some predictions persist and are thus reinforced. Norms, in this sense, are envelopes of negation: self-regulating filters of accountability and consequences. Deviations get negated; stable practices persist. Meaning is constructed and generated through shared practice. Let’s call the meaning that arises—and stabilizes—through this worldly, social error correction pragmatic closure.
Pragmatic closure is what allows us to know the difference between a well-formed, comprehensible sentence and a grammatically correct but meaningless one like, “Colorless green ideas sleep furiously,” to borrow Chomsky’s famous phrase (Chomsky 2015, 15).
This pragmatic meaning is propagated through an envelope—what Dennett called the “umwelt of an utterance” (Dennett 2009, 122). This envelope is noisy: it provides both the mechanism to transmit and the ability to adapt. Language games change over time. Wittgenstein again: “There are countless kinds; countless different kinds of use of all the things we call ‘signs’, ‘words’, ‘sentences’. And this diversity is not something fixed, given once for all; but new types of language, new language-games, as we may say, come into existence, and others become obsolete and forgotten” (Wittgenstein 2009, 14-15).
1.5.2.3 We’ve arrived at two closures: lexical closure, which stabilizes meaning through relations among words—creating a new bit that can be carried forward—and pragmatic closure, which stabilizes how that bit is used in the world and within a community, providing an error-correcting channel for transmission. Both emerge through recursive prediction and correction, operating above the inductive threshold where rules run out and patterns must be learned through time.
We feel closure from the inside. Consider learning something new. At first you’re lost—the pieces float around, unmoored. Then something clicks: Ah, I get it now. A pattern becomes stable enough to predict and use. The system compresses: many degrees of freedom collapse into a constraint you can carry forward. That constraint becomes a building block for the next harmonic. Understanding accumulates one closure at a time.
Grasping the meaning of a word is the same kind of event. Grasping is what it feels like when a real pattern becomes predictable for you—when the web of relations supports reliable inference and action.
This meaning isn’t something that is deduced. Consider a word like “belief.” There isn’t a clean definition; it’s a dense distribution. Two percent this, four percent that, one percent something else—interwoven with expectations, counterexamples, tones, and tacit norms. Now imagine forcing all those connections into a set of propositional rules—collapsing a living, contextual pattern into a set of binaries—reducing all the weights to zero or one. How much meaning would be lost? Nearly all of it.
This is what happens when we try to push a phenomenon that lives above the inductive threshold below it—when we demand explicit rules where only predictive patterns will do: we lose meaning.
1.5.2.4 We might be tempted to say only humans achieve lexical and pragmatic closure. But an LLM is itself a large real pattern—a condensed, predictive structure. Billions of parameters form a vast relational network of tokens pointing to other tokens. In that sense, an LLM can exhibit something like lexical closure: stable inferential roles inside the web of tokens, learned by recursively predicting what comes next.
But lexical closure isn’t the whole story. The question is whether the system has pragmatic closure—an error-correcting channel that ties lexical meaning to the world and to a community of correction, so that meaning can be re-instantiated across contexts (i.e., replicated).
Every week brings new research showing LLMs are closer to the mind than we expected. Consider three recent examples: the hierarchy of an LLM aligns with the dynamics of language comprehension in the brain (Goldstein et al. 2025); aspects of transformer attention approximate constraints relevant to language acquisition (Summerfield 2025, 117); and decoding/encoding linguistic messages show notable parallels across brains and models (Fedorenko et al. 2024).
Yet an LLM needs far more data than we do to learn a language. Chomsky called this the poverty of the stimulus—the observation that a child learns language with minimal exposure (Chomsky 1965, 53-59). This raises a question: why can a child learn language so quickly, while an LLM requires, in effect, digesting the entire Internet? Recent research suggests the difference lies in the vast amount of non-verbal information we receive from being in the world (Agüera y Arcas 2025, 414-416). Being in the world gives us a foothold. It lets us reach closure sooner. It narrows the search space.
This offers a possible answer to the poverty of the stimulus: pragmatic closure accelerates lexical closure. Just as life requires both metabolic and replicative closure, being a true speaker requires both lexical and pragmatic closure. An LLM requires vastly more training data to reach lexical closure than we do for precisely this reason. One could imagine the correction by a parent as having far more weight and value than the relation of words to other words. And for an LLM, even after constructing well-formed language, that language has no meaning for it in the broader world. This also explains why LLMs excel at translation—translation is just lexical closure across languages, mapping similar patterns between different languages (Agüera y Arcas 2025, 424).
In this view, reinforcement learning with human feedback is a thin slice of pragmatic closure—a small pragmatic envelope that negates certain continuations and stabilizes others. It helps, but it’s still a surrogate for being embedded in the shared world.
Nothing in this picture forbids an artificial intelligence system from achieving pragmatic closure. Put a model in a robot within a community that corrects it, and pragmatic closure becomes possible. And this, unsurprisingly, is exactly the focus of the next generation of AI start-ups—putting models out in the world so that they can learn “world models” (Chen 2026). Soon a robot will point and say, “That one.” You’ll ask, “This one?” It will say, “No—the other one, the blue one.”
With each new release, the gap between LLMs and humans seems to narrow. Soon we’ll be left saying, “But I really believe” or “I really intend”—that’s the difference. How do we measure that? Where do we look? How is this different from what we do? What, if anything, remains missing? The ledge we’re standing on is narrowing. Before long, we’ll just be yelling at each other, “No, you’re the philosophical zombie!”
“The safe ground we have left behind is where humans alone generate knowledge.” (Summerfield 2025, 2).
1.5.3 The Way Out of the Chinese Room
1.5.3.1 John Searle asked us to imagine a room where we slide messages written in Chinese under the door. Inside, a man at a desk consults a rule book that tells him which symbols to return, and after consulting it, he slides a new squiggly shape back out to us under the door. If the rule book were big enough, and the replies fast enough, the man in the room would seem to understand Chinese. Searle’s claim is that this is all a computer does (and can ever do): it shuffles symbols according to rules. It cannot ascribe semantic meaning to them (Searle 1980)—it’s all syntax; no semantics.
But we’ve already seen that language is not the kind of thing that lives comfortably in a rule book. Above the inductive threshold, rules run out. What matters is the internalization of real patterns, learned in time, under negation—that is, correction by the community.
1.5.3.2 Searle’s reply to a modern LLM would be to say it is simply the same room made probabilistic—not a look-up table of rules but statistical rules over symbols.
The deeper problem is that the room is designed to freeze the system below the very conditions in which lexical and pragmatic closure happen. It eliminates prediction as a time-extended process of learning. It eliminates feedback in the form of error correction and norm enforcement. There is no world coupling and no way to record history. The room blocks the process that would give rise to any form of lexical or pragmatic closure we’ve described.
The Chinese Room shows only that formal rule execution is insufficient for meaning. It does not show that artificial systems cannot understand. Meaning is not in the symbols, not in the rules, and not in the room. It is in the real patterns that arise when a system operates above the inductive threshold—when it can predict, be corrected, and stabilize its inferences through time inside a community and a world.
Searle’s robot reply sharpens the point. Suppose we put the room inside a robot, give it sensors, and allow it to move through the world (Searle 1980, 420–421). The robot becomes a larger Chinese Room: sensory inputs replace slips of paper, motor outputs replace written replies, and the person inside still manipulates symbols without understanding them. But embodiment, by itself, is not closure. What matters is not whether the system senses, but whether its interactions with the world reshape its internal patterns—and thus its predictions about the future.
If the robot’s “experience” is merely translated into symbols and processed by a fixed rule book, then Searle’s objection still applies. But if a system’s future actions are shaped by negation—if the external world and community become an error-correcting envelope—then we no longer have the original room. We instead have a dynamic, historical process with consequences, capable (at least in principle) of pragmatic closure. If understanding appears here, it is not located in an inner symbol manipulator, but in the coupled process as a whole, unfolding through time.
This also clarifies what is meant by a real pattern. A real pattern is not merely an outward regularity that an observer can describe after the fact, nor an instruction that reproduces a behavior. In this sense, a real pattern is an internalized constraint that guides future action in an unpredictable environment—and may itself be unpredictable (1.5.7).
Mimicry captures a time-slice of a pattern without participating in the process that forms, corrects, and sustains it. A rule book may encode a stabilized pattern after the fact, and for a short period of time, but it does not recreate the process by which the pattern is discovered, corrected, and made answerable to the future in a dynamic environment. Understanding requires more than mimicry: it requires that the pattern be answerable to correction and revision, and embedded in a practice with consequences.
If the robot merely mimics the behavioral pattern, it does not understand. If, however, the robot participates in the time-extended process by which patterns are formed, corrected, revised, and carried forward, the burden shifts. Searle can still insist that no artificial system could understand, but at that point the claim is no longer that syntax is insufficient; it rests solely on intentionality (covered next).
In this sense, the Chinese Room is GOFAI all over again: a system barred from learning, barred from negation, barred from the very process that produces closure. This is why early rule-based chat systems, like ELIZA, could mimic fragments of conversation yet fail under open-ended use (Summerfield 2025, 70–72). What GOFAI’s failure showed is that the Chinese Room, at best, would be a version of ELIZA.
1.5.4 Vitality and Intentionality
1.5.4.1 This brings us to one more link between biology and language. Before we understood evolution, vitalism claimed that living things possessed a special non-physical spark—an élan vital—that machines or dead matter lacked. This theory failed because life’s properties were eventually explained through chemistry and evolution.
Now consider original intentionality. It claims that minds have a special, non-derivative “aboutness” that no physical or computational system can possess. Searle argued that minds alone have intentionality—that meaning cannot be reduced to patterns or behaviors (Searle 1980). This is formally identical to vitalism. It posits a special mental essence that cannot be reduced to physiology, biology, or computation.
We’re searching for a spark that isn’t there—making the same mistake vitalists made. Meaning, like life, is not injected into the system. It arises in complex systems above the inductive threshold, where real patterns form through two kinds of closure: lexical closure (relationships between words) and pragmatic closure (relationships between words and the world). When these patterns become internalized and predictive, meaning and understanding are found.
1.5.4.2 Below the inductive threshold, meaning takes the form of executing rules and logic. Above the threshold, meaning happens through closure—assuming the system has the capacity—and takes the form of real patterns and norms. The only tractable strategy is inductive compression—real patterns that guide prediction. Meaning arises from the need to make predictions in a complex world.
Intentionality is not a magical spark—it is the natural consequence of living and navigating a complex world. To achieve pragmatic closure, you must move through the world, through time, making predictions, being corrected, and making revised predictions. This is the fundamental truth Heidegger grasped. It requires a form of being. It is a process. It has an arrow of time.
The philosophical demand for original intentionality repeats the vitalist mistake: it treats the failure of deduction in complex domains as evidence for a metaphysical ingredient. But “aboutness” is already explained by world-coupled predictive structure shaped by learning, history, and norms. Intentionality is not a substance or found in one.
The problem is that we’ve all been trained in logic, rules, truth conditions, and propositions—so we’re looking for an answer in the wrong place. We’re looking below the inductive threshold, but it’s not found there. We’re forced to ascribe intentionality to thinking the same way vitalism was ascribed to life.
We talk about “beliefs” and “justifications,” but persistent patterns that predict are justification enough.
1.5.4.3 What we’re attempting to do is naturalize intentionality and normativity. Dennett and Wittgenstein got us to the ten-yard line. Dennett’s intentional stance was right—there is predictive value in ascribing intentionality to beings that operate above the inductive threshold (Dennett 2023). Wittgenstein was right to characterize language as a game that must be played to be learned (Wittgenstein 2009). We’re just taking the logical next step: to operate above the inductive threshold, to internalize real patterns that are compressed and predictive, is to form a kind of lexical and pragmatic closure in which meaning—and what looks like intentionality—arise from the set of relationships the same way that life forms replicators and autocatalytic sets. There is no magic to it.
1.5.5 The Architecture of Thought
1.5.5.1 We have a view of computer architecture that goes back to Neumann: a CPU and memory as separate components. Information moves from memory to the CPU for computation, then back to memory. Much of computational functionalism adopts this view. I think this view has confused us.
In an LLM, the weights are the memory—the pattern. A prompt cascades recursively through that memory to produce an output. The thinking is the combination of the prompt with the memory. The memory is the CPU in a way. They aren’t distinct.
Wittgenstein gives this example: you go to pick a red flower from a field. It’s not as if you hold up a red card in your mind, go from flower to flower making comparisons, and pick the one that matches (Wittgenstein 1958, 3). We don’t take something from memory (the red card) and then do something with it (find the flower). We just do the thing. We just pick the flower.
This helps clarify what Wittgenstein meant by rule following: not formal rule execution but practical competence. A rule is not first looked up and then applied by a central processor—that is the GOFAI/Searle picture. Action is not the result of consulting a rule; it is a pattern unfolding in context through time, embedded and shaped by practice.
The thought is the prompt cascading through our memory that leads to action.
For Wittgenstein, the action counts as following a rule only because the pattern was formed within a practice where some ways of going on are negated, while others persist and thus are reinforced. A rule, in this sense, is not a proposition stored in the mind and then applied. It is a history of corrected predictions that allows a being to go on. A formal rule is an instruction. A real pattern is a competence. A rule, in Wittgenstein’s sense, is competence formed through pragmatic closure.
1.5.5.2 What’s confusing is that we need a prompt to get an instance of the pattern. Take your understanding of something “simple” like a tree. You have a pattern of a tree internalized; what is that pattern like—describe it? When we try to articulate it, we have to collapse a massively parallel pattern in our mind into a serialized, condensed representation in language.
But if you give me a prompt, I can tell you about a tree: Is that an oak tree? Draw me a Christmas tree. Imagine an old willow tree. We need an inference of a thing to access the pattern. I can’t show you the pattern in its entirety—only an inference.
Wittgenstein again: “I know what a word means in certain contexts” (Wittgenstein 1958, 9) and “only in use does a sentence have a sense” (Wittgenstein 1969, 3).
As Darwin showed us, to understand the shape of a bird’s beak, we must look to the environment in which it persists. You cannot separate the prediction from the envelope. You can’t get a glimpse of the pattern without a prompt.
1.5.5.3 This internal pattern doesn’t seem propositional. What’s the logical negation of your pattern of a tree? Do we invert the weights? Is it the neurons that don’t fire? What weights relate to “tree”? Consider the combinatory space in your mind that is not associated with a tree—it’s practically infinite (1.2.1.4). The pattern is not the type of thing that can be logically negated. Or if it can, the negation would seem to be meaningless. However, if you create an inference of a tree from a specific prompt, you can talk about things that are not that tree.
“The main point is the theory of what can be expressed by propositions—and what cannot be expressed by propositions, but only shown; which, I believe, is the cardinal problem of philosophy” (Wittgenstein 2012, 98).
“Do I contradict myself? Very well then—I contradict myself; I am large—I contain multitudes” (Whitman 2005, 43).
1.5.5.4 To add to the complication, thinking is recursive. When neurons fire in response to a prompt, they strengthen themselves. Ask me about x—I’ll give you answer 1. Ask again—you’ll get answer 2. And again—answer 3. On and on.
“‘I could tell you my adventures—beginning from this morning,’ said Alice a little timidly; ‘but it’s no use going back to yesterday, because I was a different person then’” (Carroll 2009, 91).
Thinking is a path through time, triggered by a prompt. The output of one thought often becomes the prompt for the next thought—generative, recursive, and noisy at each step.
Perception provides the perpetual prompt.
What feels more like thinking: a relational, path-dependent, noisy correlation built from all your empirical experiences and applied to your current environment to predict a path forward—or a set of propositional sentences infused with content by intention?
1.5.6 The Mind as a Mill
1.5.6.1 Leibniz asked us to picture the mind as a mill (Leibniz 2014, 17). We zoom in, walk around, and see all the parts—the wheels spinning—but then he asks: where is the perception, the thought, the meaning?
Inside an LLM, the words “man” and “woman” are separated by a vector with the same angle as the one between “king” and “queen.” The LLM encodes these relationships within the embedding space as similar vectors—much like they’re stored in the human brain (Summerfield 2025, 92–93). Nobody told the LLM to do this. It just does. Leibniz was right to picture the mind as a mill—he just didn’t measure the vectors between the parts.
In the complex system we’ve been describing, when we zoom in, we see composable bits: neurons in brains, nucleotides in DNA, instructions in software replicators, weights and tokens in LLMs. Seeing this, we naturally ask: how does this give rise to life, meaning, intelligence? But the meanings and the patterns exist only in the relations between parts. To look for meaning in a static part is to miss it entirely. We need to run water through the mill and watch how the parts interact. No neuron thinks. No token means. Meaning, life, and intelligence are patterns we discover as we move through time.
1.5.7 Bayesian Epistemology and Temperature
1.5.7.1 We might think that what we’ve described, in our discussion of language and earlier of intelligence, is just Bayesian epistemology: internalize the priors, predict the most likely next word or action, update priors with new information. And to some extent, that’s true, but it’s not quite right.
What’s saved isn’t the actual priors—it’s a pattern that creates relationships between them. In LLMs, these priors aren’t updated after training, unlike in a mind, which constantly updates. But the biggest difference is this: using pure probabilities to predict the next word doesn’t work entirely. Language is not purely a maximization game.
Modern LLMs have a temperature setting that controls how much unpredictability enters the generation process, word by word (Wolfram 2023, 2). Always choosing the most predictable word makes them sound robotic and mechanical. We do the same thing when we speak. If you don’t believe it, try this: take the same prompt and speak your response aloud for two or three minutes; then do it again. You’ll take a different path each time. Language is noisy and path-dependent like our complex system.
Think of poetry as turning up the temperature, and an instruction manual as turning it down—which one sounds more human?
We want to say “gotcha” when an LLM makes a mistake—as if we don’t say the right, I mean the wrong, thing all the time. Heuristical humans hallucinate, too.
1.5.8 Language and Complexity
1.5.8.1 Language is a complex system that allows thinkers to share predictions from their unique perspectives and, in doing so, increases the community’s overall robustness. It is built by recursively combining bits—tokens, phonemes, words—and stabilized within a noisy social envelope of norms, corrections, and reinforcement. This process is unknowable in advance: “You must bear in mind that a language-game is, so to speak, something unpredictable” (Wittgenstein 1969, 78).
Humans generate language in a similar structural way to an LLM. In a sense, we really are next-token predictors but with one critical addition: we generate language from inside a real-time, embodied, norm-governed envelope.
Meaning emerges through two types of closure: lexical closure, which stabilizes meaning in relation to other words, and pragmatic closure, which enables correction and transmission, ultimately anchoring meaning to the world and accelerating understanding.
This dual closure mirrors what we saw in life: one closure generates reusable bits, and another makes those bits heritable—portable through a noisy channel that both preserves meaning and lets it drift and adapt over time. In both language and life, what persists is not what is selected or chosen, but what survives negation by its environment. Evolution is prediction and negation under physical environments; language is prediction and negation under social environments.
Meaning lives in the relation of words to words and words to the world. We grasp it by moving through time in a complex environment where predictions are made and negated. We call this process a language practice—one that is constructed, that is generated, and that arises as the most efficient and robust way to make predictions in an ever-changing complex environment. With meaning defined this way, we have an answer for Nagel on how meaning can arise.
Language, like our DNA, leaves a history of its path on the tape. This history we call etymology. “Our language can be regarded as an ancient city: a maze of little streets and squares, old and new houses, of houses and extensions from various periods” (Wittgenstein 2009, 11).
What we’ve described is a kind of functionalism—but not the static kind that assumes a fixed mapping (pattern X means Y). Above the inductive threshold, function is not a look-up or mapping; it is a time-extended real pattern with predictive value, forged through repeated prediction and negation inside an envelope of correction.
We might call this complex functionalism: a form of pragmatic-predictive functionalism in which meaning, intelligence, and mind are defined not by static input-output mappings but by historically stabilized patterns that persist and adapt over time through prediction and correction within an envelope. Classic functionalism asks what role a state plays in a system. Complex functionalism asks how that role was formed, persists, and adapts through time.
Viewing language through this lens explains why the Chinese Room fails (it freezes the conditions under which closure emerges), what kind of understanding a modern LLM plausibly has (lexical closure without pragmatic closure), and how intentionality can be treated not as a mystery but as something that arises through closure.
We have moved from random prediction to adapted prediction, to informed prediction, and now to shared prediction. Each stage is a form of acceleration and increasing robustness. What once required extinction is now error correction within a social envelope.