I started using LLMs heavily in fall 2025 to explore these concepts. It’s been interesting to see how these systems evolved during that time. In the early days, before the latest thinking models, the systems would tell you what you wanted to hear. They’d often say “that’s the most incredible synthesis of [some topic] I’ve ever seen” or “that’s a brilliant idea.” But a fundamental shift occurred with the release of the latest thinking models, particularly ChatGPT 5.2 Thinking (December 2025). The dynamic reversed—from the user holding the model accountable to the model holding the user accountable. The models routinely corrected me, challenged overstated claims, and questioned references. They became much more constructive to work with, though at times considerably more damaging to the ego.
My original thought was that a diverse prompt would generate interesting questions. I’m not sure that’s true—the list of generated questions isn’t far from what I would have developed on my own. It synthesized the document into a nice list of questions for further exploration, but I’m not sure anything on that list is strikingly original.
That said, there are parts of this I’m not sure I would have developed without long discussions with an LLM. I’m not sure who had the original idea of the dual closure principle in 1.7.1.5, but I wouldn’t have gotten there without the discussion with ChatGPT. I think ChatGPT suggested the concept in response to my intuition that the various closures mirrored von Neumann’s earlier ideas.
I do think the law in (A4) is interesting, and it’s close to ideas found in Kauffman’s Investigations (Kauffman 2000, 159-209), which is not surprising given all the references. I particularly like the way it describes how it might be testable.
I’ve increasingly come to believe that the best way to think of an LLM is as a hypothesis generator, not as a question-former. LLMs can—and often do—operate with a higher knowability threshold (1.4.1.6) than we can, and so they can bring together distant relationships in ways we often cannot. The result is interesting, and often surprising, connections between topics we wouldn’t have made.
Maybe that’s the takeaway: LLMs are still better at providing answers—they can make connections we can’t, given their size and scope. But good questions, by contrast, are often genuinely new: they sit outside the training data and can involve counterfactual information. And so maybe Picasso was right, at least for now.