Language Models as a Discovery and a Philosophical Challenge

Large language models (LLMs) appear to be one of the most important discoveries of the twenty-first century. The word discovery is more appropriate here than invention, because even the creators of these models do not fully understand what exactly they have produced. They know the technical details: how the graph of tokens is structured, how linear algebra is applied, and how computations can be efficiently parallelized using graphics processing units. They know how to optimize the training process. But in essence, they cannot give an unambiguous answer to the question: “What exactly is our model a model of?”

Modern language models should not be considered solely through the prism of technical solutions; their philosophical aspects must also be taken into account. In the twenty-first century, we have encountered a situation in which “technology is explained through nineteenth-century metaphors,” creating an enormous gap between scientific and technological achievements and our understanding of their meaning. This is why language models should be approached not only from the perspective of algorithms and architectures, but also from the standpoint of the history of philosophy and epistemology.

From the Cartesian Subject to a Network Conception of Thought

To understand why language models provoke so much controversy, it is useful to recall the origins of modernity. René Descartes, who is often called the “father of modern philosophy,” introduced the proposition: “I think, therefore I am.” In doing so, he made human consciousness the unconditional center of being, implicitly equating it with something divine: it is the only thing of which we can be completely certain. Historically, this gave rise to a close association between the concepts of reason, will, and life—they came to be perceived as inseparable from one another.

The fantastic literature of the nineteenth and early twentieth centuries—for example, Mary Shelley’s Frankenstein—supported the idea that “artificial intelligence” would necessarily have to be an entity resembling a human being, endowed with the capacity to love, suffer, and empathize. This approach implicitly follows the Cartesian line: intelligence was conceived only as a reflection of human consciousness and will.

However, already in the first half of the twentieth century, traditional philosophical ideas began to give way to what is now often called postmodernism. An “attack” was launched against the classical subject–object relationship and, along with it, against the conventional understanding of signs and meaning. If, in modernity, meaning was located inside the “thinking subject,” now culture, language, and mythology themselves were said to “think.” Meaning does not originate from an isolated observing “I,” but emerges from a complex web of interrelations. Thus, the Aristotelian logic of “A or not-A” gives way to the idea of a multidimensional graph of concepts, in which meaning arises through the combination, intersection, and overlapping of different contexts.

Language Models and the Reconsideration of “Intelligence”

Modern large language models are often called “artificial intelligence,” which, in my view, creates additional confusion. People are still inclined to associate intelligence with consciousness and emotional-volitional qualities—an association formed as early as the time of Descartes and Romantic literature. As a result, one often encounters the objection: “How can this be intelligence if it cannot love or suffer?”

Technically, however, LLMs are far removed from modeling the human brain in any literal sense. Terms such as “neuron” and “neural network” are indeed used to describe them, but instead of actual nerve cells, we find ordinary matrices, multidimensional vectors, and operations of linear algebra. “Neural networks” began as an attempt to imitate biological neurons, but today the term is better understood as a practical name for classes of architectures capable of processing vast amounts of data and discovering statistical dependencies within linguistic corpora.

This gives rise to an interesting parallel with postmodern philosophy: LLMs really do “think” not through a rigid subject–object framework, but through the large-scale reading and modeling of relationships within language. Moreover, if the idea of the “world as text” became popular toward the end of the twentieth century, large language models inadvertently embody this proposition. They treat reality—or at least its verbal dimension—as an endless body of “text,” in which meaning is formed through the structure of relationships between elements.

From the “World as Text” to a New Kind of Thinking

This model suggests that human consciousness is not the only “pure source” of reason. Language, culture, and texts may themselves “think”—that is, they may generate new meanings and construct conceptual connections. At the level of LLMs, this becomes visible in the model’s ability to generate meaningful texts, answer questions, and even discover subtle relationships between concepts without necessarily possessing “consciousness” in the Cartesian sense.

We are therefore confronted with a profound philosophical question: might our familiar conception of consciousness and reason be only one possible way of explaining the phenomenon of thought? And might it now be easier to describe thinking as “text”—as a system of relationships within an enormous body of data? Perhaps the centrality of consciousness has indeed been exaggerated—or, more precisely, was specific to the modern era—while the development of language models offers us a new way of looking at what we call “intelligence.”

Conclusion

Large language models are not merely an engineering innovation, but also a serious challenge to our familiar philosophical ideas about reason. Like every major discovery, they force us to reconsider fundamental assumptions about consciousness, knowledge, and the nature of reality. To reduce them either to simple “imitation of the human being” or, conversely, to primitive statistics is to overlook their unique nature. We will probably continue to live for a long time at the intersection of different paradigms, where technology races far ahead while philosophy attempts to provide it with a new and more adequate metaphorical and conceptual framework.

And if the idea of the “world as text” once seemed merely a radical postmodern proposition, the emergence of large language models has made it far more tangible—and perhaps central to further reflection on what intelligence is and what forms it may take.

2 Likes

Perhaps “world as text” was the continental error, “semantics as syntax” the anglo-saxon one?

I’m referring to the Chinese Room debates. LLM’s seem to me a surprisingly vivid manifestation of that thought experiment.

1 Like

Derrida is among those postmodern authors asserting ‘world as text’ Having claimed that there is nothing outside the text, he may be the Continental philosopher most closely associated with the idea. But by text he meant context of relevance, and relevance is intrinsically affective. Relevance cannot be generated by statistical calculations, a reduction of semantics to syntax. While it is not the product of subjective consciousness, neither is it an unconscious arrangement of context-free ‘data’. So his idea of text is far removed from the philosophical assumptions behind LLM’s and ‘semantics as syntax’.

I think what we learn from LLMs is what physicalists have always suggested. Complicated life like behaviour can arise from relatively simple programmes run on hardware.

I don’t wish to detract from Descartes reputation but I think he was largely mistaken on the issue of mind and body. Democritus was clearer about the nature of reality. But I admit Democritus is premodern and a little unreliable. However, a long time before post modernism, Hume also pointed out that the self is constructed from experience and rooted in substance.

Perhaps, we should try to reveal the truth obscured too often by unecessary mystification rather than invent new sources of mystification.

It’s true that LLM begins with differential relations among signs. But postmodern thinkers like Deleuze, Heidegger, Foucault, and Derrida dont just replace the subject with an abstract network of signs. The network itself is not a cold mathematical structure. Relations are not objective connections, they’re relations of relevance, salience, attraction, repulsion, concern, desire, fear, usefulness, danger, etc.

An LLM lacks this dimension. The relationship between “fire” and “danger” in an LLM is ultimately a statistical relation in a learned vector space. The relationship between fire and danger for a living being is an affectively charged significance relation. Fire matters. And it’s not just in living beings that affective mattering is primary. For writers like Deleuze the inorganic world is also structured this way

It might seem that structuralist accounts like Saussurian linguistics share features with llm’s , but even Saussure would give more priority to affective relevance than llm’s do. Saussure was describing a human linguistic institution. The differential structure of language exists because speakers inhabit a social world in which distinctions matter. The sign “tree” means what it does because it functions within the practices, concerns, and forms of life of speakers. The relational system is never completely detached from lived significance.

From this vantage, LLMs don’t overcome Cartesian dualism so much as reproduce it in a new form, treating meaning as objective structure while severing it from and ignoring affective relevance. When we inquire into the philosophical presuppositions animating the kinds of approaches which consider the affective dimension to be superfluous, we are led to pre-Hegelian, and perhaps pre-Kantian thinking.

Yes. Another familiar way of saying this might be: LLMs are neither alive nor conscious. That’s why they don’t exhibit the listed relations.

But that’s the familiar Cartesian model, slightly modified to include other animals within the category of existing things which have affectivity and consciousness. It still presumes a domain of things which are devoid of subjective qualities. Postmodernists, Hegelians and even Kantians would dispute this.

Yes, it does. And @Patterner would also object to such a domain! For my part, it’s less a theoretical issue than a practical one. If we continue to think that “alive” and “conscious” are meaningful terms for constructing truth-apt statements, then LLMs and their ilk present a challenge. I thought your list of relational qualities was a good way to show what is characteristic of living, conscious things, as we currently understand them, and I was interested to see how LLMs are excluded on that understanding.

Hmm? (looks up from comic book) Oh! Yes, indeed! No such thing!

I just found out that you can hide text by putting it between < and >. I hid some below.

In fact, LLMs manifest not the culmination of the “world as text” thesis, but its historical decline. While postmodernism focused on the primacy of the text, LLMs apparently perform the de-centering and displacing of the text. Fragments of language become AI’s training material, incorporated within an enormous computational space. Instead of interpretation and reading, AI entertains a variety of mathematical and statistical processes. The digitized text integrates with various non-linguistic forms of human symbolism and coding procedures. In this way, language becomes a functional part of an automated, digitized ecosystem. A novel, operationalized and automated interface through which human symbolic practices are reorganized displaces the text as a center of the enactment of meaning. The philosophical significance of LLMs lies less in their inaccessible computational substrate than in its interface’s operational logic. Far from being simply informational or interpretive, it transforms the dynamics of collective symbolic activity. Its operative speed blurs the distinctions between the text appearance, its interpretation, commentary, critique, and reception. The topology and temporality of the human-LLM operational milieu ultimately reshape the text’s epistemological and hermeneutical presuppositions.