1.4 Patterns and Prediction → 💭 Intelligence

1.4.1  The Inductive Threshold

1.4.1.1  Let’s return to the soup we’ve been brewing. Strings have developed into complex shapes. These shapes have begun to undergo metabolic and replicative closures, forming hierarchies—harmonics of complexity that grant the stability needed to find the next harmonic. Let’s assume they’ve developed basic mobility and sensing capabilities.

The environment is now extremely complex. There are many kinds of shapes and countless ways they can combine and interact. Each new interaction changes the likelihood of what comes next. We pulse the clock forward, and the adjacent possible expands again.

Imagine you’re a shape floating around in this world. Every moment, things change. What strategy would make you more robust? Or, put differently: which shapes would persist? Shapes that can anticipate—that can make informed predictions in an ever-changing environment—gain a decisive advantage. Successful prediction increases robustness. To persist is to efficiently and effectively predict.

We could call these shapes organisms or agents, but for lack of a better word, we’ll call them thinkers.

1.4.1.2  Let’s assume that, as an envelope in your own regard, you can encode some information about your environment. But you’re a small thinker with finite memory, floating in a combinatorial space many orders of magnitude larger than a cosmic byte. You cannot possibly encode the entire history. And in a genuinely complex environment, encoding a complete rule book is not just impractical—it’s fragile. As soon as you harden a rule, the environment shifts: the adjacent possible expands, new thinkers appear, new interactions emerge.

Thinkers that maximize prediction under noise with minimal storage persist. Instead of a rule book, they store compressed response patterns—flexible heuristics dense enough to fit inside finite memory.

We’ll borrow Dennett’s term and call these stored patterns of response real patterns (Dennett 1991). A real pattern is the compression of a higher-dimensional input into information with predictive value. The more compressible and predictive it is, the more valuable it becomes to our thinker—cheap to store, yet powerful in a noisy environment. We’ll assume going forward that calling something a “pattern” means a real pattern.

Internalizing one of these patterns is itself a form of closure: a diffuse set of relationships collapses into a usable unit inside the thinker. We’ll call this predictive closure, and it happens when a real pattern with predictive value is internalized. 

In small, closed, stable systems, you can encode the rule book—rules and rigidity. In large, open, dynamic systems, you encode real patterns—heuristics and flexibility.

We’ve now picked up the idea of “sophisticated information processing, and adaptation via learning or evolution” in Mitchell’s definition of complexity (1.1.4.1).

1.4.1.3  Framed this way, for a given thinker with limited memory, there exists a threshold where the environment becomes too dynamic to be captured by a finite set of explicit rules—it must be navigated by real patterns. Let’s call this the inductive threshold.

Below the threshold, the environment is small and stable enough (and memory large enough) for deduction—a world of rules. Above the threshold, the environment is too complex, too dynamic, and memory too limited—the thinker must be inductive.

Importantly, this internalization of patterns happens under constraint (time and energy). We can think of a heuristic as a real pattern, discovered by induction, under the constraints of time and energy.

Good enough allows you to survive where perfect might get you killed.

1.4.1.4  A mind likely evolved as the ideal structure for storing real patterns—those that give thinkers maximum predictive value under constraint in a complex environment. Minds are optimized to store real patterns—just as a honeycomb’s hexagonal structure is optimal for bees to store honey (Chittka 2022, 73–75). 

Real patterns are to minds as honey is to hives.

Additionally, these patterns would be deeply relational because it’s more efficient to store differences between thinkers than to store information about each one separately (i.e., neural networks).

1.4.1.5  In the early stages of a complex system, thinkers almost certainly have tiny memories. Thus, a thinker in a complex environment begins above the inductive threshold. Induction is the only working epistemology in a sufficiently complex environment. And thus, induction would have been the first primary epistemology. 

Wittgenstein writes: “If anyone said that information about the past couldn’t convince him that something would happen in the future, I wouldn’t understand him” (Wittgenstein 2009, 143). 

To predict the possible next state, you must explore by educated guesswork. Novelty cannot be deduced in such a world—it must be generated and tested. Above the inductive threshold, the future is not derivable; it is discoverable.

1.4.1.6  Further, there is a second threshold where even induction fails—where the environment is too complex to encode in real patterns. We’ll call this the knowability threshold, above which a thinker would be left with only a random prediction strategy. Wittgenstein again: “Explanations come to an end somewhere” (Wittgenstein 2009, 6).

1.4.1.7  If we wanted to formalize our thresholds, then being below the inductive threshold would result from the combination of being within the search threshold (1.2.1.6), avoiding the kind of self-reference that gives rise to unknowability (1.2.2), and removing noise from the system.

As a thinker’s capacities grow, we could imagine both their inductive threshold and their knowability threshold rising. Importantly, these thresholds would differ and remain dynamic for each thinker.

We could quantify being above the inductive threshold by comparing predictive success to random trial and error. When the prediction rate falls to the level of random guessing, we have crossed the knowability threshold. This way of measuring prediction works in a bounded environment with a stable answer against which predictions can be tested. But in an open-ended, complex environment, there may be no fixed target. The act of prediction itself changes the environment, expands the adjacent possible, and shifts the next benchmark. In that case, accuracy is not the deepest measure—persistence is. 

A prediction succeeds when it helps preserve the envelope that makes future prediction possible.

In this way, persistence is prediction’s first aim, not truth. Truth, rather, names a higher-order stabilization of persistence: a prediction that survives not merely for one thinker, in one moment, or within one narrow envelope, but across many perspectives and repeated negations in many contexts. What we call truth is persistence under maximum pressure.

1.4.1.8  Wimsatt, in his discussion of the ontology of complex systems, distinguishes between levels, perspectives, and causal thickets (Wimsatt 2007, 193–240). Levels are clearly defined, but as complexity grows, boundaries blur. He describes the transition from levels to perspectives as an “increased richness of ways entities have of interacting with one another due in part to the increasing number of degrees of freedom and of emergent properties” (Wimsatt 2007, 205). At this point, we’re forced to take various perspectives to gain understanding—we’ll return to the idea of perspectives in section 1.5. Above perspectives lie what he calls a causal thicket: the “high entropy … causal structure of the universe—sort of an ontological primal slime” (Wimsatt 2007, 240).

While the exact mapping to our thresholds isn’t necessarily what he meant, it offers a useful way to think about them. Below the inductive threshold, things are well defined and knowable. Above it, each thinker has a perspective, and at best, things can be probabilistically predictable. Above the knowability threshold lies an “ontological primal slime”—our unknowable combinatory space.

In a loose and metaphorical sense, this also mirrors the P vs. NP problem in computer science. Below the inductive threshold are domains more like P: relatively tractable, rule-like, and deductively navigable. Above it are domains more like hard search problems in NP: not necessarily uncomputable, but resistant to straightforward deduction, such that workable solutions depend on heuristics, approximation, and the internalization of real patterns.

1.4.1.9  This framing also aligns with Karl Friston’s free-energy principle, which offers a more technical model for what we’ve been describing (Friston 2010). 

Friston describes systems that minimize free energy—systems that “maintain their states and form in the face of a constantly changing environment” (Friston 2010, 127). They do so through “a constructive process based on internal generative models” (Friston 2010, 128) that uses a “probabilistic model that can generate predictions” (Friston 2010, 129), with the goal to “maximize the accuracy of predictions, under complexity constraints” (Friston 2010, 132). 

In essence, the system makes predictions to maintain the integrity of the envelope that makes those predictions possible in the first place. It survives by avoiding states that would negate its own existence.

Prediction is necessary for persistence in a complex environment, and good predictors persist longer.

1.4.2  Ought Has an Arrow

1.4.2.1  Above the inductive threshold, there is an arrow of time—in the sense that history matters. Complexity arises because the environment is noisy and path-dependent, because it is recursive, and because the space is too large to iterate (1.2). To exist above the inductive threshold is to be in a process through time: carving a path through a noisy, dynamic, near-infinite combinatory space.

To reduce this complexity—to get below our inductive threshold—we must idealize the world. We remove noise, making it predictable. We cap the complexity, limiting possibilities. We shrink the space, making it searchable. In this simplification, the arrow of time fades. Below the inductive threshold, the world becomes a domain where rules and laws dominate and where models can be treated as effectively deterministic.

There is a truth here that the phenomenologists got right about ontology and time. Martin Heidegger saw this: “the central problematic of all ontology is rooted in the phenomenon of time” and “with the problematic of Temporality as our clue” (Heidegger 2008, 19–63).

And again, Nagel points the way here: “some laws of nature would apply directly to the relation between the present and the future, rather than specifying instantaneous functions that hold at all times” (Nagel 2012, 93).

1.4.2.2  We can think of idealizing and simplifying our combinatory space as taking a limit. We move from an indeterminate environment to a determinate one—regularities harden into rules. Think of it as a world where probabilities converge to 0 and 1. Below the threshold, there is no middle; above it, there is almost nothing but middle.

It’s not that the purity of mathematics and the rationality of logic exist in some platonic realm. They are what happens when we remove noise, shrink the search space, and make it timeless. They are reductions and limits, not something different. They feel pure because they allow us to simplify. In that simplified world, we can discover things like laws because they are deterministically knowable. Above the threshold, by contrast, they have to be lived to be known.

Logic, in this view, emerges as a limit of induction—crystallized expectations. It is frozen, idealized induction in a regime where we have narrowed the possible and removed noise. The propositional is just the probabilistic with probabilities pushed to 0 or 1.

We evolved in a world with an arrow. To imagine one without it is an abstraction.

In a complex system, induction likely preceded deduction, and empiricism preceded rationality. Logic is a later cultural achievement, built from stabilized inductive regularities. After modeling higher-level regularities, the simple rules emerge. The limit of a regularity, as complexity decreases, is what we call a rule or law.

To search for proof is to play the wrong game: learning is not propositional; life is not a truth table to be iterated.

And in that sense, you can’t derive ought from is because the ought came first—not as a metaphysical priority but as the original problem of action in a nondeterministic world.

We learn the grammar after we can speak the language. Wittgenstein again: “As if it took a logician to show people at last what a proper sentence looks like” (Wittgenstein 2009, 43).

Dostoevsky expounding on the inductive threshold and the limits of deduction: “Two times two is four—why, in my opinion, it’s sheer impudence, sirs. Two times two is four has a cocky look; it stands across your path, arms akimbo, and spits. I agree that two times two is four is an excellent thing; but if we’re going to start praising everything, then two times two is five is sometimes also a most charming little thing.”

And on what lies above it: “Consciousness, for example, is infinitely higher than two times two. After two times two, there would, of course, be nothing left—not only to do, but even to learn” (Dostoevsky 2021, 33–35).

1.4.2.3  Framed this way, philosophy has been bouncing around our threshold for thousands of years: between process and description; between time-bound means and measurable ends; between the messy normativity of practice and the clean idealizations of proof. Between empiricists above and rationalists below. Between continental above and analytic below. Dennett’s intentional stance versus design stance mirrors it. Wilfred Sellars’s manifest image versus scientific image mirrors it too (Sellars 1962). The list could go on and on. I can hear the philosophers groan as I converge all these ideas. And yes, it is an enormous simplification, but there also seems to be a fundamental truth here—and one that becomes legible in the context of complexity.

As we move below the threshold, the probabilistic becomes the propositional, heuristics become laws, regularities become rules, the manifest becomes the scientific, process becomes description, being becomes is.

“Perhaps that’s their big idea, to keep on saying the same old thing, generation after generation” (Beckett 2009, 376).

To occupy a space above our inductive threshold is to live in the world of oughts—not morality in the narrow sense but normativity in the broad one: more correct predictions, know-how, tacit skill, rules of thumb, expectations. To live below it is to live in the world of laws, propositions, proofs, rules, facts.

Hume told us we cannot derive an ought from an is; perhaps the is is what remains when an ought collapses in a simplified world.

Complexity is the domain of the oughts.

1.4.3  The Imitation Game Revisited 

1.4.3.1  We turn down the temperature, put our soup on simmer, and let it be for now—the analogy can only take us so far. Which brings me to thermostats (I told you this was going to be noisy in the beginning).

Borrowing from Dennett yet again, imagine a simple thermostat. When the temperature drops below a set threshold, it turns on the heat; when it rises above, it turns it off. It reacts deterministically: if A, then B.

Now add features. Connect it to weather forecasts. Let it monitor motion in the house. Allow it to store and analyze its own operating history. Suddenly, it no longer lives in a world where rigid conditionals suffice. It compares the present state—its prompt—to a history, and uses that comparison to predict what should happen next. Sometimes it succeeds, sometimes it fails. And notice, we call a thermostat like this a smart thermostat (Dennett 2023, 209-210).

1.4.3.2  In moving from a basic thermostat to a smart thermostat, we shift from rule-following in a closed space to prediction in an open one—from if A then B to if A then probably B: from a set of instructions to a set of predictable patterns, from deductive-like behavior to inductive prediction. Our intuition treats this transition as a move toward intelligence.

We naturally associate our inductive threshold with intelligence. Intelligence is knowing what to do when there’s no single correct answer—only better and worse ones. To be intelligent is to successfully navigate an environment above the inductive threshold: to persist where a complete rule book cannot guide you.

Agüera y Arcas writes that “prediction underlies intelligence at every scale” (Agüera y Arcas 2025, 324). Decades earlier, Lovelock offered a similar view: “Intelligence is a property of living systems and is concerned with the ability to answer questions correctly” (Lovelock 2016, 137).

Intelligence, learning, and therefore thinking are underpinned by efficient prediction in a complex environment—by operating in the space of ought rather than the comfort of is.

1.4.3.3  This illuminates the brilliance of Turing’s choice of language as a criterion for thinking (Turing 1950). Language is a complex system (we’ll revisit that claim). We’ve already seen that it has an enormous combinatorial space—one where a basic grammar opens into a near-infinite space of possibilities. In choosing language, Turing chose a domain that lies unmistakably above the inductive threshold.

Framed this way, the Turing test is not so much an imitation game as an induction game. It is a stress test for whether a system can keep making good predictions—good moves—across an unbounded range of contexts.

Below the inductive threshold, I can imitate thinking with rules. Above it, my imitation with rules breaks. Only a system that has internalized real patterns can keep its footing as novelty accumulates. As Warren Buffett said, “Only when the tide goes out do you discover who’s been swimming naked” (Buffett 2007, 3)—only above the inductive threshold can we see who is really intelligent.

We demonstrate intelligence by navigating environments that sit above the inductive threshold. Intelligence is what happens when complex systems internalize real patterns and use them to predict the adjacent possible. Turing was more right than he knew.

1.4.3.4  This also points toward an explanation for the development of artificial intelligence. In the early days of good old-fashioned AI (GOFAI), researchers tried to encode the world in rules: expert systems, rules engines, hand-built ontologies (Summerfield 2025, 27-29). Marvin Minsky went as far as to say “once you have the right kind of descriptions and mechanisms, learning isn’t really that important” (Summerfield 2025, 17). GOFAI researchers tried to force the world below the inductive threshold. They made progress in small, well-bounded domains, and then repeatedly hit a wall and an AI winter followed.

Researchers then leaned into connectionist systems—neural networks—and progress began to accelerate: first with deep learning, then transformers, and ultimately the LLMs we see today. GOFAI treated language as if it could be made tractable through explicit rules; neural networks succeeded because they learned to operate above the inductive threshold by internalizing statistical regularities—real patterns—at scale. And yes, on this account, LLMs exhibit a form of intelligence.

1.4.3.5  The progression from GOFAI to LLMs also mirrors the shift from early to late Wittgenstein. Early Wittgenstein aimed to reveal the logical structure of language and, in doing so, the limits of what could be said (Wittgenstein 2024). But to treat language this way is to treat it as though it lives below the inductive threshold—as if meaning could be clarified by exposing its crystalline logical form. When we force language into this ideal, we risk repeating GOFAI’s mistake: treating language as rule-governed all the way down. Depending on how charitably one reads it, the Tractatus can appear either as a warning to GOFAI researchers or as an early version of their temptation.

Late Wittgenstein moves in the other direction. The language games of the Philosophical Investigations live above the inductive threshold: meaning is not deduced from logical form but stabilized through use, correction, and practice over time (Wittgenstein 2009). Language is not a rule book applied to the world—not formal rules all the way down—but patterns of use.

One might extend this point to the idea that thought itself is fundamentally propositional. Are we—or were we—in a kind of propositional winter in how we think about thinking? Have we overestimated the role of propositions?

1.4.4  Predicting Proteins

1.4.4.1  I want to return to proteins and ask a question: why are modern neural networks—especially transformer-like architectures—so good at both predicting protein shapes and generating well-formed language?

Before modern AI, resolving the structure of a single protein could take years of experimental work—sometimes an entire PhD. Brute-force enumeration was never an option; the combinatorial space is too large (1.2.1.5). And yet AlphaFold, an AI model designed to predict protein shapes, has now predicted the structure of virtually all known proteins—a feat that won its creators a Nobel Prize (Hassabis et al. 2025). AlphaFold uses the transformer (Jumper et al. 2021, 585), the same neural network architecture that powers today’s LLMs.

A key 1997 paper on protein folding identifies the core challenge: “An important point that emerged from these studies is that an essential element of the complexity is the presence of long-range interactions.” The paper continues: “The problem would not be NP hard if each amino acid could find its native conformation independently of the others or if only near-neighbor interactions were involved.” NP hard means the time to solve grows exponentially with complexity—in this case, 1052 years for the protein used as an example (Karplus 1997, 70; emphasis added).

The problem becomes tractable only if relationships are independent or correlations involve near-neighbors alone. Predicting protein folding sits above our inductive threshold precisely because it combines a vast combinatorial space with long-range interactions across the shape.

Compare this to the 2017 paper introducing the transformer architecture, which allows “modeling of dependencies without regard to their distance in the input or output sequences” (Vaswani et al. 2017, 2; emphasis added). The transformer solved the problem of correlations and dependencies that span large distances—whether in proteins or language. AlphaFold operates above the inductive threshold. The transformer architecture, in particular, enables the discovery of real patterns in highly complex environments. 

A modern transformer-based model might have anywhere from 10,000 to 100,000 tokens, which are small combinations of a couple of letters (Summerfield 2025, 108-109). Compare that to our minds: we can’t possibly internalize the real patterns that might involve the distant interactions of 10,000—let alone 100,000—tokens. The result is that these models have a higher knowability threshold than we do—what looks, in effect, random to us is a real pattern for them.

Proteins and language each have a shape—carved from a nearly infinite combinatorial space through a noisy process of combination and negation. In each case, if I change the order of bits—and therefore the shape—I lose something: the protein changes function, and the sentence changes meaning. Proteins and language are formed by similar systems.

Both protein folding and language are only tractable above the inductive threshold. They are too large and their relations too distant and too path-dependent for any rule-based system. Modern AI systems are, in effect, the world’s best shape predictors—a result of their incredible ability to discover and model real patterns.

1.4.4.2  The power of an LLM arises from a number of factors that our complex system shares—parallelism, open-endedness, and the ability to scale a large number of bits. Their success largely coincides with the availability of graphical processing units (GPUs), which allow for massive parallelization (1.1.2.7) of the problems (Agüera y Arcas 2025, 44). 

Further, recent papers have associated open-endedness—think of our uncapped complexity (1.1.5.1)—as essential for advances to artificial general intelligence (AGI) and hold that this “open-endedness is fundamentally an experimental process” (Hughes et al. 2024, 6). We can think of the experimental process as being above the inductive threshold. 

Lastly, Richard Sutton describes the “bitter lesson” of artificial intelligence development: scaling learning tends to outperform domain-specific understanding, because the contents of minds—and the world—are “tremendously, irredeemably complex” (Sutton 2019). The answer, then, is to allow these systems to find good approximations on their own (Sutton 2019)—that is, to let them operate above the inductive threshold and internalize real patterns.

Learning and prediction go hand in hand, and it’s a similar mechanism in both minds and LLMs (Summerfield 2025, 161–162).

1.4.5  Intelligence and Complexity

1.4.5.1  This view of intelligence aligns with the relatively new understanding of “biological intelligence,” a view that traces intelligence back to the earliest organisms and is not based on the development of brains or humans (Adee 2025, 75). In this view, “cognition is a relational property between the organism and its environment” and is based on the importance of electrical signalling that enables the integration of information from the outside to the inside (Adee 2025, 79; emphasis added). Intelligence defined this way is a complex process in its own regard. 

Minds, then, are complex systems capable of implementing intelligence. They consist of neurons that connect and prune, combine and negate as we develop—a process called plasticity. They exhibit regions and hierarchies (1.3.3), occupy just one configuration in a near-infinite combinatorial space (1.2.1.4), and are noisy: any given neuron may or may not fire. In this way, minds model the systems they are trying to predict. As Roger Conant and Ross Ashby showed, “Any good regulator must be a model of that system” (Conant and Ashby 1970).

Intelligence is the ability to operate above the inductive threshold—to make informed predictions that grow more robust over time in a changing and ultimately unknowable environment. Intelligence is a kind of predictive closure that happens when we internalize a real pattern with predictive value—when a set of relationships becomes a usable thing unto itself.

Once structure retains information under negation, then you have inductive bias embodied in structure. As a result, a thinker can now make many predictions and internalize their results. Each prediction doesn’t require death; we can make many predictions in a lifetime.

Knowledge is created through conjecture and criticism in the mind, and through variation and natural selection in nature (Marletto 2022, 17). As David Deutsch puts it, “Knowledge creation depends on error correction” (Deutsch 2011, 271). Knowledge is what survives contact with the environment, not what is derivable a priori. Karl Popper called this falsification (Popper 2002); Darwin called it selection (Darwin 2012). We might even think of knowledge harmonics—pockets of dynamic stability, much like Thomas Kuhn’s paradigms (Kuhn 2012). Intelligence and learning—the production of knowledge—mirror our complex process.

Framing intelligence this way helps us to understand various views in philosophy, why the Turing test works as a test for intelligence, and the recent advancements in AI.

1.4.5.2  In response to the Lovelace objection—that machines can only do what we tell them to do—Turing suggested that instead of programming adult-level intelligence directly, we might build “one which simulates a child” and educate it. He imagined systems that learn from experience, shaped by “punishments and rewards,” that might become so complex “its teacher will often be very largely ignorant of what is going on inside” (Turing 1950, 454–460). In these gestures, he anticipated the basic ideas of modern machine learning: training rather than programming, reinforcement functions, and the opacity of large language models—what’s now called the interpretability problem. He also emphasized the need for a “critical size” or scale (think about being above our inductive threshold) before interesting behavior could emerge (Turing 1950, 454). If Turing were alive today and you described neural networks, machine learning, reinforcement learning with human feedback, scaling laws, and the interpretability problem, you could imagine him yelling, “Exactly!”

When Neumann was asked what it would take for a computer to think, he reportedly said, “It would have to grow, not be built; it would have to understand language, to read, to write, to speak; and it would have to play, like a child” (Labatut 2024, 257–265). 

Turing and Neumann—arguably two of the greatest minds of the last century—both pointed toward the same ideas: the intersection of biology, language, and computation; the link between language and intelligence; the primacy of learning over programming; the necessity of operating above the inductive threshold.

← Previous  |  Next →