Topological constraints on self-organization in locally interacting systems

**PHILOSOPHICAL TRANSACTIONS A** [royalsocietypublishing.org/rsta](https://royalsocietypublishing.org/rsta)

Research
Cite this article: Sacco F, Sakthivadivel D, Levin M. 2026 Topological constraints on self-organization in locally interacting systems. Phil. Trans. R. Soc. A 384: 20250011. https://doi.org/10.1098/rsta.2025.0011

Received: 25 January 2025
Accepted: 8 August 2025

One contribution of 18 to a theme issue ‘World models in natural and artificial intelligence’.

Subject Areas:
statistical physics, complexity, artificial intelligence

Keywords:
phase transitions, potts model, self-organization, emergence, diverse intelligence

Author for correspondence:
Dalton Sakthivadivel
e-mail: dsakthivadivel@gc.cuny.edu

Francesco Sacco, Dalton Sakthivadivel and Michael Levin

Tufts University, Allen Discovery Center, Medford, MA, USA
Department of Mathematics, CUNY Graduate Center, New York, NY, USA
Tufts Center for Regenerative and Developmental Biology, Tufts University, Medford, MA, USA

DS, [0000-0002-7907-7611](); ML, [0000-0001-7292-8084]()

GRAY

All intelligence is collective intelligence, in the sense that it is made of parts that must align with respect to system-level goals. Understanding the dynamics that facilitate or limit navigation of problem spaces by aligned parts thus impacts many fields ranging across life sciences and engineering. To that end, consider a system on the vertices of a planar graph, with pairwise interactions prescribed by the edges of the graph. Such systems can sometimes exhibit long-range order, distinguishing one phase of macroscopic behaviour from another. In networks of interacting systems, we may view spontaneous ordering as a form of self-organization, modelling neural and basal forms of cognition. Here, we discuss necessary conditions on the topology of the graph for an ordered phase to exist, with an eye towards finding constraints on the ability of a system with local interactions to maintain an ordered target state. By studying the scaling of free energy under the formation of domain walls in three model systems—the Potts model, autoregressive models and hierarchical networks—we show how the combinatorics of interactions on a graph prevent or allow spontaneous ordering. As an application, we are able to analyse why multiscale systems like those prevalent in biology are capable of organizing into complex patterns, whereas rudimentary language models are challenged by long sequences of outputs.

This article is part of the theme issue ‘World models in natural and artificial intelligence’.

© 2026 The Authors. Published by the Royal Society under the terms of the Creative Commons Attribution License [http://creativecommons.org/licenses/by/4.0/](http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, provided the original author and source are credited.

1. Introduction

Self-organization is a fascinating phenomenon observed across diverse systems in nature and technology. In biology, self-organization of complex structures and functions, from subcellular machinery to multicellular morphogenesis, requires alignment of parts to navigate a problem space towards specific adaptive ends [1–7]. Dynamics of this form can be viewed as a sort of autopoietic cognition, raising interesting questions about how self-organizing systems navigate through configuration space given some target morphology [8–19]. More recently, current approaches to machine learning have involved physics-inspired models such as the Hopfield network [20–22] and spin glasses [23] and energy-based [24] or diffusion models [25]. In these systems, in order to maintain a pattern out of equilibrium or produce sensible data over long time spans, there must exist a ‘condensed’ phase with long-range order. While some systems—like those prevalent in biology—demonstrate remarkable abilities to organize at large scales, other systems that also exhibit some degree of collective intelligence—such as natural language models—have much more limited capabilities [26–30]. Despite this, spin systems have also been used to model cellular morphogenesis [31–35], suggesting the difference goes beyond the substrate of intelligence considered. This article is motivated by the following question: what is the functional distinction between simple language models and multicellular organisms, and can generative AI harness that property to achieve long-range order? In particular, we are interested in placing explicit thermodynamic bounds on the probability of some example systems occupying such a phase, with applications to estimating their capability to self-organize.

The question of how probable ordered configurations are as the number of elements increases is equivalent to asking about the existence of a phase transition, where the ultimate energetic constraints on the existence of an ordered phase arise from the topology of the interactions between subunits of the system. A quick argument due initially to Landau–Lifshitz [36]Landau and Lifshitz, §149 shows why the one-dimensional Ising model, the prototypical spin system, has no ordered phase at any temperature; their study of scaling is the technique we will use in this article, so we will review it here. Recall that a thermodynamically favourable change in state is one which decreases the free energy

delta F equals delta E minus T delta S

Suppose a domain wall of perimeter P forms at an arbitrary location in the chain. Using an effective Hamiltonian, calculating the free energy of the system with zero domain walls and the system with one domain wall shows that the change in free energy scales like minus T log P, meaning that forming a domain wall is always thermodynamically favourable for sufficiently large P. As a result, long chains are unstable under thermal fluctuations, and disorder always propagates through the chain. In the thermodynamic limit, there is no spontaneous magnetization. Ising (1925)Ising proves this by finding the partition function and calculating the mean magnetization analytically in his celebrated 1925 paper, but the advantage of the argument of Landau and Lifshitz is that it applies generically to phase transitions in many different sorts of one-dimensional system (though the details prove somewhat subtle) and uses a simple scaling argument. The argument to the contrary in dimension greater than one is famously due to Peierls [37], where it is shown that the change in internal energy has a factor of P balancing the change in entropy. This is found by computing the number of spins an ‘island’ interacts with—which is highly dependent on the properties of the lattice.

The main results in this article concern when an ordered phase can exist in a self-organizing system based on its interactions. We will focus on the scaling relationship between energy and entropy as it is determined by combinatorial topology. We begin this section by presenting our argument for a large universality class of models and deriving the necessary conditions for the existence of an ordered phase in this class. In §3, we analyse the one-dimensional Potts chain as a simple model system, recovering the argument given by Landau–Lifshitz in a more general setting. §4 maps autoregressive models [38] onto this one-dimensional framework and proves their

A diagram showing a chain of variables labeled s sub minus one through s sub five. A box labeled H sub two groups the variables from s sub zero to s sub four.Figure 1. A local Hamiltonian in dimension one. The Hamiltonian of this chain is a sum of windowed Hamiltonians of length omega equals five. Pictured is the second window; the first begins at s sub minus one.

inability to maintain long-range order. A discussion of transformers follows in §5section five. In §6section six, we apply these results to multiscale systems to evaluate the possibility of order in that context, based on the hierarchies built into the topology of interaction. Finally, we discuss implications of these results and future directions in §§7 and 8sections seven and eight. We conclude that certain systems cannot remain ordered over long ranges for all time and suggest this no-go theorem as a source of fitness pressure for biological phenomena like stigmergy and embodiment. The key insight of our work will be that topology is the critical factor differentiating these systems. Namely, while cells in the human body can coordinate and organize over vast scales, forming coherent tissues and organs, language models struggle to maintain consistency beyond their limited context windows. This disparity stems directly from the underlying topology of interactions in these systems.

2. The main argument

Consider a k-vertex graph with -by-k by k adjacency matrix G and an n-ary variable on each vertex. Denote each such variable as s sub i, indexed by i from one to k. The interactions between any set of spins will be given by the Hamiltonian H.

Definition 1. A windowed Hamiltonian H is a Hamiltonian in that spin–spin interactions are defined only within a finite window of interaction omega.

Remark. Without loss of generality, we will assume that the lowest energy state of any windowed Hamiltonian is zero. We will moreover assume the existence of an upper bound to the highest energy of any possible window Hamiltonian, E max sub i is less than or equal to E max.
We call a sum of such effective Hamiltonians a local Hamiltonian. In particular,

Definition 2. A local Hamiltonian is a Hamiltonian which can be decomposed into a sum of several window Hamiltonians H equals the sum over u of H sub u, all of which have the same window length omega.

See figure 1 for a depiction. For the purpose of intuition, one may notice a local Hamiltonian with window size equal to one is a discrete Markov field.

Our methodology will be given by the following prescription:

  1. Begin with the system in one of the ordered configurations, e.g. one of the stored patterns
  2. Create a domain wall
  3. Estimate the energy gained by the system
  4. Find the asymptotics of the free energy as the number of domain walls increases

Computing the change in F by changing the number of domain walls requires detailed knowledge of the combinatorics of the interactions on the graph. Instead, we can use the structure imposed by the windows. Namely, we have regulated the length of interactions such that we need only count the number of windows containing a domain wall. This does away with the particularities of the lattice and so can be applied to systems on very generic graphs. We will explore this in the present section.

The most general way to consider the interaction between the elements of our system is if the interactions are represented by a graph Hamiltonian. For simplicity, we will assume that the

Two diagrams comparing a graph G and a graph Hamiltonian H. On the left, graph G shows six vertices s1 through s6 connected in a circular pattern with two cross-connections. On the right, graph H shows the same structure but with edges labeled e1 through e7 representing coupling strengths.Figure 2. Illustration of a graph Hamiltonian. Left: the graph G represents the adjacency matrix of an undirected graph with six vertices, where edges indicate connections between vertices s sub i. Right: the graph H represents the graph Hamiltonian, where the coupling strengths e sub i have come from H bar.

edges of the graph are fixed in time; one can easily justify this by appealing to an adiabatic approximation where the birth or death of edges is unlikely for the timescale on which observations are made.

Definition 3. Let G be the adjacency matrix of a graph on k vertices and H bar a -by- matrix. A graph Hamiltonian is the Hamiltonian of a weighted directed graph; namely, a Hamiltonian which can be written as

H equals H bar entry-wise product G

where the symbol is entry-wise multiplication and H bar has as entries the coupling strengths of the system.

See figure 2. We will now give a recipe for computing the scaling behaviour of energy and entropy for a graph Hamiltonian system. For the sake of simplicity, we will work with square planar embeddings of graphs (grid-like lattices) in this article. In these cases, a domain wall and its perimeter can be easily defined (figure 3).

Let H be a graph Hamiltonian and P be the perimeter length of a domain wall. As the perimeter length increases, the number of possible configurations of domain barriers increases, increasing the entropy of the system delta S. We say that the entropy gained scales as if

delta S is of order f S of P

We give a similar definition for energy scaling: as the perimeter length increases, the upper and lower bounds of the energy gained scale as and order f E high of P and order f E low of P, respectively. If f E high equals f E low, defined as f E we say that the energy gained scales as , that is, delta E is of order f E of P.

If we can estimate how the energy and entropy scale, we can verify the existence or non-existence of an ordered phase without knowing the exact formulae for and E of P and S of P, only considering how they scale as P goes to infinity.

Proposition 1. If then there exists an ordered phase.

Proof. Suppose . Then the change in free energy is

............................................................................................................. Downloaded from [Royal Society Publishing](http://royalsocietypublishing.org/rsta/article-pdf/doi/10.1098/rsta.2025.0011/6138032/rsta.2025.0011.pdf) by guest on 05 July 2026

A 5 by 5 grid of circles, with most being red and a small cluster of five being blue. An orange line outlines the boundary between the blue and red regions.Figure 3. Two different domains in a two-dimensional grid. The perimeter that separates the two domains is drawn in orange.

If we now take T to zero, the change in free energy becomes increasingly positive. As such, the creation of a domain wall is unfavourable. ■

In this way, we can rule out details of the system that could complicate the combinatorics—for example, the number of stored patterns. Let

\pi = (\dots, s_{-1}, s_0, s_1, \dots) $$<yap-speak>pi equals the sequence s sub i</yap-speak> denote a ground state of the system. One may think of this state as a stored pattern or morphogenetic configuration. When there is a ground state degeneracy—that is, multiple states with zero energy exist, for instance, if a Hopfield model has $m > 1$<yap-speak>m greater than one</yap-speak> stored patterns—they will be enumerated as

\pi^\alpha = (\dots, s_{-1}^\alpha, s_0^\alpha, s_1^\alpha, \dots), \quad 1 < \alpha \leqslant m.

We can show that the number of such patterns affects neither the scaling of the entropy nor energy. **Lemma 1.** Let $H = \sum_u H_u$<yap-speak>H equal the sum over u of H sub u</yap-speak>. If there exist two energies $E^{\text{max}}, E^{\text{min}}$<yap-speak>E max and E min</yap-speak> which are the greatest and least non-zero energy levels of all the windowed Hamiltonians $H_u$<yap-speak>H sub u</yap-speak>, respectively, then at thermal equilibrium, the ability to converge to an ordered phase is independent of energy levels and window sizes. *Proof.* For any $H_u$<yap-speak>H sub u</yap-speak> let $\omega_1$<yap-speak>omega one</yap-speak> be the size of the smallest window and $\omega_2$<yap-speak>omega two</yap-speak> that of the largest. The energy gained from the creation of a domain wall is bounded by

\omega_1 P E^{\text{min}} \leqslant \Delta E \leqslant \omega_2 P E^{\text{max}}.

In both cases, we have $E = O(P)$<yap-speak>E is of order P</yap-speak> so that the asymptotics of the system depend only on the perimeter length. ■ **Lemma 2.** Let $H$<yap-speak>H</yap-speak> be a graph Hamiltonian with $m > 1$<yap-speak>m greater than one</yap-speak> stored patterns. At thermal equilibrium, the ability to converge to an ordered phase is independent of $m$<yap-speak>m</yap-speak>. <yap-show> Downloaded from [royalsocietypublishing.org](http://royalsocietypublishing.org/rsta/article-pdf/doi/10.1098/rsta.2025.0011/6138032/rsta.2025.0011.pdf) by guest on 05 July 2026 </yap-show> *Proof.* Let $B(P)$<yap-speak>B of P</yap-speak> be the number of possible configurations of domain barriers with perimeter $P$<yap-speak>P</yap-speak>. The change in entropy due to the creation of a domain barrier can always be written as

\Delta S = \log [(m - 1)B(P)] = \log B(P) + \log(m - 1).

<yap-speak>delta S equals the log of m minus one times B of P, which is the log of B of P plus the log of m minus one.</yap-speak> In the thermodynamic limit, the term proportional to the number of barriers increases, while the term proportional to the number of patterns stored stays constant. As such, its contribution can be ignored. ■ Since neither the shape nor number of stored patterns affects the thermodynamics of the problem, and neither does the size of the interaction window, we reduce our analysis to the nearest-neighbour Ising model. Our main technical result is the following topological equivalence theorem: **Theorem 1.** *All local Hamiltonians on lattices with the same combinatorial structure have asymptotically equivalent free energies.* *Proof.* We will show that we can approximate any $H$<yap-speak>H</yap-speak> with the Hamiltonian of some arbitrary other model in such a way that their asymptotics are equivalent, implying that in the thermodynamic limit, the Peierls argument depends only on the perimeter length in the approximate system. Let $H = \sum_u H_u$<yap-speak>H equals the sum over u of H sub u</yap-speak> be a local Hamiltonian on a planar graph $G$<yap-speak>G</yap-speak> with any spin $i$<yap-speak>i</yap-speak> coupled to $\eta_i$<yap-speak>eta sub i</yap-speak> other spins. Let $H_0$<yap-speak>H zero</yap-speak> be the Hamiltonian of an arbitrary system. One knows by the Bogoliubov inequality <yap-show>[39]</yap-show> that

F \le F_0 + \langle H - H_0 \rangle_0

<yap-speak>F is less than or equal to F zero plus the expectation of H minus H zero with respect to the H zero ensemble</yap-speak> and therefore that if $H$<yap-speak>H</yap-speak> and $H_0$<yap-speak>H zero</yap-speak> are asymptotically equivalent then $F \sim F_0$<yap-speak>F is asymptotic to F zero</yap-speak>. As the perimeter length increases, we have $E = O(P)$<yap-speak>E is order P</yap-speak> and $E_0 = O(P_0)$<yap-speak>E zero is order P zero</yap-speak> by Lemma 1. If the combinatorics of the lattices are the same—namely, if the set of $\eta_i$<yap-speak>eta sub i</yap-speak> is equal on both lattices—then we can take $P = P_0$<yap-speak>P equals P zero</yap-speak>. The change in the energy is computed by the Hamiltonian of the new configuration, implying $H - H_0 \sim 0$<yap-speak>H minus H zero is approximately zero</yap-speak>. By Lemma 2, the probability of any disorder occurring is asymptotically the same in both systems, so that the *change* in free energies is also asymptotically equivalent. The claim follows. ■ **Corollary 1.** *The capacity for self-organization in any system on a graph $G$<yap-speak>G</yap-speak>—that is, the existence or non-existence of a phase transition—is equivalent to that of a nearest-neighbour Ising model on the same graph, with*

H_0 = -J \sum_{i,j} s_i s_j

\begin{align} \pi^1 &= (1, 1, 1, \dots, 1) \ \pi^2 &= (-1, -1, -1, \dots, -1). \end{align}

<yap-speak>pi one as a vector of ones and pi two as a vector of minus ones</yap-speak> An important remark is that we do not claim the phase transitions are the same. Naturally, they will differ in general, for instance, if there are more stored patterns in one system than the other. It is the existence of a phase transition from the topology of the lattice that we study. ## 3. Stored patterns in the windowed With this in hand, we can study our first model system: the one-dimensional Potts model. The Potts model is a variant of the Ising model whose state vectors are richer than simply binary numbers. In particular, one can encode data in a Potts chain by choosing different spin configurations, <yap-show> Downloaded from [Royal Society Publishing](http://royalsocietypublishing.org/rsta/article-pdf/doi/10.1098/rsta.2025.0011/6138032/rsta.2025.0011.pdf) </yap-show> recalling the discussion in the Introduction. It is a suitable model to describe stored images or other configurations of multiply-valued variables. **Definition 4.** Let $n$<yap-speak>n</yap-speak> be a finite positive number. A Potts chain $\mathcal{C}$<yap-speak>C</yap-speak> is a $\mathbb{Z}$-indexed set of $n$-ary integer symbols; that is, a chain of spins $s_i$<yap-speak>s sub i</yap-speak> indexed by $i \in \mathbb{Z}$<yap-speak>i in the set of integers</yap-speak>, where each $s_i$<yap-speak>s sub i</yap-speak> can assume any integer value in $\{1, \dots, n\}$. It follows from Theorem 1 that it is sufficient to consider the scaling of $\mathcal{C}$<yap-speak>C</yap-speak> to establish the existence of an ordered phase. In the same fashion as Landau–Lifshitz, it can be shown that local Hamiltonians in dimension one do not converge to a prescribed ordered state, since the formation of a domain wall always 'interrupt' the pattern. **Theorem 2.** *Let $H$<yap-speak>H</yap-speak> be a one-dimensional local Hamiltonian with $m > 1$<yap-speak>m greater than one</yap-speak> stored patterns. At non-zero temperature, the formation of a domain wall is thermodynamically favourable.* *Proof.* Suppose that our Potts chain starts out in the first ground state or pattern, $\mathcal{C} = \pi^1$<yap-speak>C equals pi one</yap-speak>. If

\Delta F = \Delta E - T \Delta S < 0,

<yap-speak>the change in free energy, delta F, equals delta E minus T delta S, which is less than zero</yap-speak> so that the free energy of the system decreases upon the formation of a domain barrier, then the formation of a domain barrier is thermodynamically favourable. It is immediate that any $\pi^\alpha$<yap-speak>pi alpha</yap-speak> differs from some other $\pi^\gamma$<yap-speak>pi gamma</yap-speak> by at least one change in spin. In a sequence of length $L$<yap-speak>L</yap-speak> there are $L - 1$<yap-speak>L minus one</yap-speak> possible places where a domain wall can appear, and at each such place, we may obtain one of the $m - 1$<yap-speak>m minus one</yap-speak> other patterns saved. As such, the change in entropy of the system is

\Delta S = \log[(m - 1)(L - 1)].

<yap-speak>delta S equals the log of, m minus one times L minus one</yap-speak> Upon the formation of a domain barrier, the windowed Hamiltonians that intersect it will have non-zero, positive energy. By assumption, the energy is bounded, and no more than $\omega$<yap-speak>omega</yap-speak> windows can be affected by a domain wall, from which we obtain

0 \leqslant \Delta E \leqslant \omega E^{\text{max}}.

\Delta F \leqslant \omega E^{\text{max}} - T \log[(m - 1)(L - 1)].

<yap-speak>delta F is less than or equal to omega E max minus T log of, m minus one times L minus one</yap-speak> Assume $T > 0$<yap-speak>T is greater than zero</yap-speak>. As $L$<yap-speak>L</yap-speak> increases, the right-hand side of the equation eventually becomes negative. ■ ## 4. Autoregressive models We begin this section by recalling in figure 4 the motivation stated in the introduction: contrasting the generation of textual features with the generation of morphological features, we would like to understand how the self-organizing nature of text is constrained by the topology of the interactions between subunits generating that text. We will now discuss a similar result as the one in the previous subsection, for autoregressive models. An autoregressive model is formed when values of a sequence are regressed against previous values of that sequence. Here, it will be defined as an estimator of some conditional probability distribution; namely, that of a sequence of observations of some data at step $i$<yap-speak>i</yap-speak>, given a history of observations until $i - 1$<yap-speak>i minus one</yap-speak>. **Definition 5.** Let $\{s_1, \dots, s_{i-1}\}$<yap-speak>the set s one through s i minus one</yap-speak> be a random $n$-ary sequence of length $i - 1$<yap-speak>i minus one</yap-speak>. Given a window (also called context) of length $\omega$<yap-speak>omega</yap-speak>, an order $\omega$<yap-speak>omega</yap-speak> autoregressive model computes

P(s_i \mid s_{i-1}, \dots, s_{i-\omega})

<yap-speak>the probability of s i given the sequence from s i minus one to s i minus omega</yap-speak> as a probability vector over $\{1, \dots, n\}$ by generating samples of $s_i$<yap-speak>s i</yap-speak> according to some random process on window states. Such a function is called an AR($\omega$<yap-speak>omega</yap-speak>) model in particular. As shorthand, we will denote the estimator whose output is $v$<yap-speak>v</yap-speak> when given a sequence $\{s_{i-1}, \dots, s_{i-\omega}\}$<yap-speak>s sub i minus one through s sub i minus omega</yap-speak> as $M$<yap-speak>M</yap-speak>:

v := M(s_i \mid s_{i-1}, \dots, s_{i-\omega}).

The function $M$<yap-speak>M</yap-speak> has the type of a conditional probability distribution. We further write

P(s_i = c) = M(s_i = c \mid s_{i-1}, \dots, s_{i-\omega})

for the $c$-th component of $v$<yap-speak>v</yap-speak>. Conceptually, we have the picture in figure 5. To fit this in the discussion of patterns and long-range order in spin chains, we find the following theorem useful. We will set the convention that if $u - \omega < 1$<yap-speak>u minus omega is less than one</yap-speak> then the window is empty at that index, and that conditioning on the empty set is the same as taking unconditional probability. **Theorem 3.** *A unique local Hamiltonian with window length $\omega$<yap-speak>omega</yap-speak> can be associated to any AR($\omega$<yap-speak>omega</yap-speak>) model.* *Proof.* Let $M$<yap-speak>M</yap-speak> be our autoregressive model, and $s_{i-1}, \dots, s_{i-\omega}$<yap-speak>s sub i minus one through s sub i minus omega</yap-speak> our input sequence. The probability that the next observed spin in the sequence is equal to $c$<yap-speak>c</yap-speak> is

P(s_i = c) = M(s_i = c \mid s_{i-1}, \dots, s_{i-\omega}).

Suppose[^1] we assign an energy to each possible $c$<yap-speak>c</yap-speak> with a scalar function $E : \{1, \dots, n\} \to \mathbb{R}$<yap-speak>E from the set one to n into the real numbers</yap-speak> and constant weight $\beta \geqslant 0$<yap-speak>beta greater than or equal to zero</yap-speak>, such that

\beta E_c = -\log P(s_i = c) + \text{const.} \quad \text{with} \quad c \in {1 \dots n}.

H_u(s_u) = -\log M(s_u \mid s_{u-1}, \dots, s_{u-\omega}) + \text{const}

for any $u$<yap-speak>u</yap-speak> in $\{1, \dots, i\}$<yap-speak>the set one to i</yap-speak> so that the probability of any $u$th token occurring before $i$<yap-speak>i</yap-speak> is conditioned on the window preceding it. The full local Hamiltonian is the sum

H = \sum_u H_u(s_u)

in an obvious way. ■ *Remark.* By Theorem 3, the generation of samples by an autoregressive model can now be seen as sampling from a Boltzmann distribution over sequences. We close this subsection with the following conclusion. The free energy can be calculated by $F = E - \beta^{-1}S$<yap-speak>F equals E minus beta inverse S</yap-speak> as usual. By Theorem 3, autoregressive models are one-dimensional systems described by a local Hamiltonian. The following corollary is then a consequence of the scaling argument in Theorem 2. **Corollary 2.** *For any finite $\beta$<yap-speak>beta</yap-speak>, an autoregressive model is unable to converge to a single stored pattern.* ## 5. Transformers and attention An interesting question to ask is what these results imply for the organizational capabilities, and hence intelligence, of large language models. Most state-of-the-art language models of today are [^1]: One could appeal to Gibbs’ equation or more general maximum entropy principles to do so. decoder-only transformer architectures <yap-show>(see [40])</yap-show>, meaning that to predict the next word (or token) they perform autoregression on all the previous elements of text—in other words, they are autoregressive at inference time. Indeed, transformers can be described as spin collectives <yap-show>[41–43]</yap-show>, with training dynamics under asynchronous updates behaving like a spin chain evolving under Glauber dynamics and text generation being sampling from the corresponding equilibrium distribution, just as we have modelled autoregression by a spin chain here. A consequence of the results in §4 is a fundamental limitation on the coherence of a simple version of a large language model for long sequences of outputs—and hence an explanation for limitations observed in papers such as <yap-show>[30]</yap-show>. **Proposition 2.** *Causally masked attention in a decoder-only model has no ordered phase.* *Proof.* Let $s \in \mathbb{R}^k$<yap-speak>s in R to the k</yap-speak> be the configuration of the network, $X$ a matrix of $m$ stored patterns $(\pi^1, \dots, \pi^m)$ with any $(X)_\alpha = \pi^\alpha = (s_1, \dots, s_k)$ and $X_{i\alpha} = \pi_i^\alpha = s_i$, and $\beta$<yap-speak>beta</yap-speak> a constant scalar. Now $X$ is a $k$-by-$m$ matrix of $n$-ary variables. One knows <yap-show>(see [41])</yap-show> the standard attention algorithm with the identity matrix for weights is a modern Hopfield network with an exponential potential function <yap-show>[22]</yap-show>. The asynchronous updates maximize the probability of a Gibbs distribution, where the energy $\beta E(s)$ of a configuration, $-\log P(s) + \text{const}$, is

-\log \left( \sum_{\alpha=1}^m \exp \left( \beta \sum_{i=1}^k X_{i\alpha} s_i \right) \right) + \frac{\beta}{2} \sum_{i=1}^k s_i s_i + \log m + \frac{\beta}{2} \left( \max_\alpha |X_\alpha| \right)^2.

With a causally-masked positional encoding from $i - \omega$ to $i$, the expression at $i$ with context length $\omega$<yap-speak>omega</yap-speak> becomes

-\log \left( \sum_{\alpha=1}^m \exp \left( \beta \sum_{j=i-\omega}^i X_{j\alpha} s_j \right) \right) + \frac{\beta}{2} \sum_{j=i-\omega}^i s_j s_j + \log m + \frac{\beta}{2} \left( \max_\alpha |X_\alpha| \right)^2.

This can easily be seen as an expression for a local Hamiltonian in a particular $\text{AR}(\omega)$<yap-speak>A R omega</yap-speak> model; it follows from Theorem 3 that we can apply Theorem 2 to the transformer, completing the proof of the statement. ■ For simplicity, we have used no mapping in an associative space; however, the result generalizes readily to those cases by the results in <yap-show>[41]</yap-show>. More sophisticated large language models employ multi-headed attention algorithms <yap-show>[44]</yap-show>, where different attention heads in the decoder are attenuated to different classes of relationship. However, even in attention with heads for long-range dependencies, a maximum context length is imposed as a hyperparameter for reasons of computational cost <yap-show>(see the implementations in the review [45])</yap-show>. While the algorithm itself is more sophisticated and allows better performance when applied to two-dimensional data (e.g. images), the coherence of a sequence of text is constrained by context length, which is a realistic and necessary-to-consider constraint given hardware. More discussion on this will be given at the end of §6. ## 6. Large hierarchical systems on graphs In many systems of interest, there are hierarchical interactions, i.e. effective interactions between subgraphs. Many multiscale systems in biology and physics exhibit complex patterns consisting of pockets of regularity assembled into structures with global irregularity, such as tissues with different morphogenetic features assembled out of their constituent cells <yap-show>[46–51]</yap-show>, functional networks in the brain consisting of regions specialized to process certain sorts of information <yap-show>[52–55]</yap-show>, and self-assembling molecules in active matter situations <yap-show>[56–60]</yap-show>. In this section, we will investigate the interplay between the scaling of free energy at one level of a hierarchy and that of another, <yap-show> Downloaded from http://royalsocietypublishing.org/rsta/article-pdf/doi/10.1098/rsta.2025.0011/6138032/rsta.2025.0011.pdf by guest on 05 July 2026 </yap-show> ![Diagram illustrating the analogy between text generation and morphogenesis, showing navigation in morphological space, error correction, and regeneration.](./topological-constraints-on-self-organization-in-locally-interacting-systems-4.png)<yap-cap>**Figure 4.** Text generation in analogy to a morphogenetic process.</yap-cap> when the topology of the graph organizes those levels meaningfully, to quantify when such patterns are possible. We will complement the results in previous sections by showing that systems lacking hierarchical structure may be limited in their ability to form complex patterns. Recall that a clique is a (non-empty) complete induced subgraph (figure 6). Suppose there exist $\ell > 1$<yap-speak>ell greater than one</yap-speak> independent[^2] cliques in $G$<yap-speak>G</yap-speak>, with $n_1, \dots, n_\ell$<yap-speak>n one through n ell</yap-speak> vertices, respectively, and each $n_i > 2$<yap-speak>n i greater than two</yap-speak>. There will be a macrostate associated with each clique giving the magnetization of that subgraph. We will argue that there are ways for local order but global disorder to exist in systems with cliques. Since the effective behaviour of the system, i.e. the properties of the cliques, is often a meaningful experimental variable (and hence a useful macrostate), this means organizing a system hierarchically can create interesting order phenomena. In particular, we want to show that there can exist ‘hierarchical behaviours’, where each clique individually is a coherent phase, but that phase varies from clique to clique. [^2]: By independent, we mean the vertex set of each clique is disjoint from that of each other clique. For brevity, and without loss of generality[^3], we will assume the coupling constants $J$<yap-speak>J</yap-speak> are uniform across the graph. We will need the following observation before we prove our main theorem in this section. Recall that Boltzmann’s formula $S = k \log W$<yap-speak>S equals k log W</yap-speak> denotes by $W$<yap-speak>W</yap-speak> the multiplicity of a macrostate. A clique is positively (negatively) magnetized if all spins in the clique have value $+1$<yap-speak>plus one</yap-speak> ($-1$<yap-speak>minus one</yap-speak>). We will begin with $\ell$<yap-speak>ell</yap-speak> positively magnetized cliques. If the number (denote it $r$<yap-speak>r</yap-speak>) of ‘flipped’ (i.e. uniformly changed) cliques is a macrostate, then the multiplicity is the number of ways to arrange $r$<yap-speak>r</yap-speak> distinguished cliques out of the total $\ell$<yap-speak>ell</yap-speak> cliques. As such, we have

S = k \log \binom{\ell}{r}.

<yap-speak>S equals k times the log of ell choose r</yap-speak> **Theorem 4.** *Take any of the $\ell$<yap-speak>ell</yap-speak> cliques. There exist parameter regimes where individual cliques may change from positive to negative magnetization. For that temperature, take any individual $n_i$<yap-speak>n sub i</yap-speak>-clique and consider a spin within it. If the coupling of every spin in the clique is greater than $\frac{T}{2} \log n_i$ then the clique remains uniformly magnetized.* *Proof.* Flipping $r$<yap-speak>r</yap-speak> cliques from positive to negative magnetization takes an energy of $\sum_{\gamma=1}^{r} 2Jn_{i_\gamma}$, where $\gamma$<yap-speak>gamma</yap-speak> indexes cliques flipped, and has an entropy of the logarithm of $\ell$<yap-speak>ell</yap-speak> choose $r$<yap-speak>r</yap-speak>. The full change in free energy is

\Delta F = \sum_{\gamma=1}^{r} 2Jn_{i_\gamma} - T \log \frac{\ell!}{r!(\ell - r)!}

<yap-speak>the change in free energy delta F equals the sum from gamma equals one to r of two J n sub i gamma, minus T log of the combination formula ell factorial over r factorial times ell minus r factorial</yap-speak> so that the scaling behaviour of $r$<yap-speak>r</yap-speak> flips to

\Delta F \sim O(r) - TO(r \log \ell - r \log r + r),

where we have obtained the approximation of the entropy from $\binom{\ell}{r} \approx \ell^r / r!$ and Stirling’s formula. Clearly, if (up to some constant)

T > \frac{1}{\log \ell - \log r + 1}

then $\Delta F < 0$<yap-speak>delta F is less than zero</yap-speak>. Hence, there exist temperatures for which it is favourable for some number $r$<yap-speak>r</yap-speak> of entire cliques to change. Fix such a $T$<yap-speak>T</yap-speak>. Within an $n_i$<yap-speak>n sub i</yap-speak>-clique, there are $n_i$<yap-speak>n sub i</yap-speak> ways to flip a single spin. This takes energy $2J$<yap-speak>two J</yap-speak>. If we have $2J > T \log n_i$, then there is no domain wall within the clique. By minding constants in the $O(r)$ terms above, we show the hypothesis can be attained—namely, that there exist values of $T$<yap-speak>T</yap-speak> for which the inequality is satisfied. **Proposition 3.** *Let $n_{\text{max}}$<yap-speak>n max</yap-speak> be an integer greater than zero denoting the number of vertices in the largest clique. There exists a non-empty critical temperature range of hierarchical behaviour.* *Proof.* From the theorem above, we have

\frac{2J}{\log n_i} > T > \frac{2J \sum_{\gamma=1}^r n_{i_\gamma}}{r \log \ell - r \log r + r}.

rn_{\text{max}} \geqslant \sum_{\gamma=1}^r n_{i_\gamma}

<yap-speak>r times n max is greater than or equal to the sum of n sub i gamma</yap-speak> [^3]: This is because the $J$<yap-speak>J</yap-speak>’s do not affect the combinatorics used to compute the entropy; as in Lemma 1, we are permitted to simply take $J$<yap-speak>J</yap-speak> large enough to bound all of the true couplings in the clique. ![Diagram showing a sequence of tokens from s one to s thirteen with arrows indicating autoregressive dependencies.](./topological-constraints-on-self-organization-in-locally-interacting-systems-5.png)<yap-cap>Figure 5. An autoregressive model with $\omega = 5$<yap-speak>omega equals five</yap-speak>.</yap-cap> and

\frac{1}{\log n_i} \geqslant \frac{1}{\log n_{\max}},

\frac{2J}{\log n_{\max}} > T > \frac{2Jrn_{\max}}{r\log \ell - r\log r + r}.

\frac{1}{\log n_{\max}} > T > \frac{1}{\log \ell - \log r + 1} n_{\max}.

Since the upper bound is a strictly decreasing function of $n_{\max}$<yap-speak>n max</yap-speak> with a singularity at one, and the lower bound is a linear function of $n_{\max}$<yap-speak>n max</yap-speak>, they intersect for some sufficiently small slope, before which the inequality is satisfied. We conclude that there exist $\ell, r$<yap-speak>ell and r</yap-speak> for which the set of satisfactory $T$<yap-speak>T</yap-speak> is non-empty. In particular, this occurs whenever

\frac{\ell}{r} > \frac{n_{\max}^{n_{\max}}}{e}.

<yap-show>■</yap-show> By looking at the effective behaviour of the graph, we consider the statistical properties of cliques to be like those of individual spins. When those cliques themselves form cliques, we have complete induced subgraphs on an effective graph and can make this argument again (see figure 7 for an example of what is meant). This section demonstrates that whenever we have cliques at some level, there can be interesting hybridized behaviours where local order may exist, but the system may be globally disordered. This recapitulates observations of multiscale phenomena in biology and physics, where for some parameter regimes there is coherence at one level (e.g. the tissues constituting an individual organ) and non-uniformity at another (e.g. the differing organs in a body). It also shows that there can be order within a window of a transformer—if the attention heads have the right relationships—but is consistent with a lack of global coherence for all spins. ## 7. Applications and limitations There are a number of compelling applications and important limitations of these results; we will discuss them here. While a simplification, the fixed window models the practical context length limitations that real systems face. As the context increases, it becomes computationally infeasible to predict the next token—and only the last few words are fed as input to the model. Intuitively, this leads the model to completely forget what is left outside of its input window, placing constraints on how coherent the model is able to be over long periods of time. This will be a useful observation in §8. On the other hand, we have neglected to consider newly emerging strategies for expanding context or making the context window dynamic, such as those highlighted in <yap-show>[61]</yap-show><yap-speak>prior work</yap-speak>. We hope to study these in the future. ![Two diagrams of a graph G. The left shows a graph with six nodes connected in a circular and cross-linked pattern. The right shows the same graph with two three-vertex cliques highlighted: one in blue (top) and one in red (bottom).](./topological-constraints-on-self-organization-in-locally-interacting-systems-6.png)<yap-cap>Figure 6. A graph G with two independent three-vertex cliques and two edges connecting the cliques.</yap-cap> Theorem 1 is best applicable to systems falling under the Landau theory of equilibrium phase transitions. Many interesting non-equilibrium systems fall outside this regime, suggesting future directions which generalize these results to the kind of non-equilibrium free energies considered in stochastic thermodynamics <yap-show>[62–67]</yap-show>. However, in certain situations, the same scaling relationships are respected and simply become dynamic in time or localized in space <yap-show>[68,69]</yap-show>, meaning the results generalize readily to those cases. At criticality and out of equilibrium, finite-size effects become relevant to the computations of critical exponents <yap-show>[39]</yap-show>, but not to reasoning about the existence of phase transitions. Similarly, uniformity in coupling strength (as in Theorem 4) and window length (as in Theorem 3) can be assumed when studying asymptotics, since the only quantities of interest are the upper and lower bounds on the change in energy (as in Lemma 1). ## 8. Conclusions and future directions In this article, we have developed a scaling argument extending Peierls’ argument, which can be used to reason about the possibility of spontaneous order in systems with local interactions. Biological systems are known to exploit the complex, hierarchical structures formed by cells in multicellular organisms to enable self-organization and problem-solving in anatomical, physiological and gene expression spaces, as well as the three-dimensional space of conventional behaviour and linguistics <yap-show>[70,71]</yap-show>. Many recent discussions have focused on the similarities and differences in the dynamics that allow evolved, engineered and hybrid systems to bind their parts towards efficient navigation to large-scale goals—a capability which is ultimately a defining feature of ‘life’ <yap-show>[5,19,72–74]</yap-show>. In stark contrast, we have demonstrated that current natural language processing algorithms lack the necessary topology for self-organization. This limitation explains their inability to generate long and coherent text that matches the complexity and consistency of biological systems. The key insight of our results is that topology is the critical factor differentiating these systems. While cells in the human body can coordinate and organize over large scales, forming coherent tissues and organs, language models struggle to maintain consistency beyond their limited context windows. This disparity stems directly from the underlying topology of interactions in these systems. By understanding the crucial role of topology in enabling self-organization, we can begin to envision new architectures for natural language processing and biologically inspired computing that might better emulate the self-organizing capabilities of biological systems. The inability of autoregressive large language models to maintain states of long-range order resembles the tangential speech or derailment in formal thought disorder, such as that found in schizophrenia and other forms of psychosis <yap-show>[75]</yap-show>, lending a complementary aspect to the ‘hallucinations’ seen in current large language models. Proposed therapeutic approaches to formal thought disorder involve ![A complex graph with nodes labeled A through V showing various connections.](./topological-constraints-on-self-organization-in-locally-interacting-systems-7.png)<yap-cap>(a) Full graph with all interactions</yap-cap> ![The same graph with specific groups of edges highlighted in green, blue, and red to show cliques and supercliques.](./topological-constraints-on-self-organization-in-locally-interacting-systems-8.png)<yap-cap>(b) Cliques and supercliques highlighted</yap-cap> ![A simplified graph with seven numbered nodes representing the consolidated groups from the previous steps.](./topological-constraints-on-self-organization-in-locally-interacting-systems-9.png)<yap-cap>(c) Effective graph obtained from supercliques</yap-cap> **Figure 7.** Consider a large (i.e. having many vertices) graph containing two or more complete induced subgraphs whose vertex sets are disjoint. An example is depicted in (a). In (b) the edges forming those independent cliques are highlighted in green. The edges between the consolidated cliques form two independent supercliques which are also highlighted in (b) in blue and red respectively, and edges participating in neither a clique nor a superclique are left black. Note that supercliques may consist of cliques with differing vertex number (e.g. the red superclique); also note that the supercliques shown in (c) form a further superclique with vertex set $\{\{1, 2, 3\}, \{4, 5, 6\}, \{7\}\}$<yap-speak>consisting of groups one, two, and three; four, five, and six; and seven</yap-speak> and edges coloured in black. refining the articulated thought in conversation or written (i.e. diagrammatic) form under the guidance of a psychotherapist <yap-show>[76]</yap-show>, suggesting that interacting with an environment is a way to enforce coherence of information.[^4] This further suggests that an *embodied* world model, extending the system in space and time by its interactions with an environment, can be leveraged to [^4]: We thank Karl Friston for elaborating on these points. ![A diagram illustrating the hierarchy of biological organization from molecules to the biosphere, linked to a morphospace of different scales—metabolic, physiological, transcriptional, and behavioral. Above this, a linguistic path navigate through pockets of relevant information toward a narrative goal.](./topological-constraints-on-self-organization-in-locally-interacting-systems-10.png)<yap-cap>**Figure 8.** Cognition as goal-directed behaviour, with a focus on producing and processing meaningful pieces of text.</yap-cap> maintain coherence. We hypothesize this explains why stigmergy and other forms of extracellular signalling arise in biological systems, which is known to enhance the ability for a collective system to order itself <yap-show>[77–82]</yap-show> and has been used in robot design and control for this feature <yap-show>[83–85]</yap-show>. Indeed, human brains are also unable to maintain the coherence of generated text past certain lengths of time due to energy constraints on neural processing, but have tools to cope with this, such as memory aids and cues in the environment. Throughout, we have assumed that the free energy is minimized. For some systems, the equilibration time is sufficiently long that, in the interest of practicality, one would be interested in relaxing this assumption—for example, when the energy landscape is very rugged. One calls such systems spin-glasses <yap-show>[23]</yap-show>. Studying the phase transitions of glassy systems can be challenging even for relatively simple glasses, such as Hopfield networks <yap-show>[86,87]</yap-show> or the Edwards–Anderson model <yap-show>[88]</yap-show>. In forthcoming work, we will give similar estimates constraining the existence of phases with frozen disorder by the topology of the underlying lattice. <yap-show> **Data accessibility.** This article has no additional data. **Declaration of AI use.** We have not used AI-assisted technologies in creating this article. **Authors’ contributions.** F.S.: formal analysis, software, visualization, writing—original draft, writing—review and editing; D.S.: conceptualization, formal analysis, supervision, writing—original draft, writing—review </yap-show> and editing; M.L.: conceptualization, funding acquisition, project administration, resources, supervision, writing—review and editing. All authors gave final approval for publication and agreed to be held accountable for the work performed therein. **Authors’ Notes.** An interactive preprint: <yap-show>[89]</yap-show> and video abstract: [YouTube](https://www.youtube.com/watch?v=cGcY-ReeGDU) accompany this paper. We thank Karl Friston for helpful comments and Mikalai Shevko for assistance with code debugging. **Conflict of interest declaration.** We declare we have no competing interests. **Funding.** DARS is grateful for the support of the Einstein Chair programme at the CUNY Graduate Centre. ML gratefully acknowledges support of Karen Fries. # References 1. Couzin ID. 2007 Collective minds. *Nature* **445**, 715–715. ([doi:10.1038/445715a](https://doi.org/10.1038/445715a)) 2. Watson RA, Buckley CL, Mills R. 2011 Optimization in ‘self-modeling’ complex adaptive systems. *Complexity* **16**, 17–26. ([doi:10.1002/cplx.20346](https://doi.org/10.1002/cplx.20346)) 3. Levin M. 2019 The computational boundary of a ‘self’: developmental bioelectricity drives multicellularity and scale-free cognition. *Front. Psychol.* **10**, 2688. ([doi:10.3389/fpsyg.2019.02688/full](https://doi.org/10.3389/fpsyg.2019.02688/full)) 4. Kriegman S, Blackiston D, Levin M, Bongard J. 2020 A scalable pipeline for designing reconfigurable organisms. *Proc. Natl Acad. Sci. USA* **117**, 1853–1859. ([doi:10.1073/pnas.1910837117](https://doi.org/10.1073/pnas.1910837117)) 5. Watson RA, Levin M, Buckley CL. 2022 Design for an individual: connectionist approaches to the evolutionary transitions in individuality. *Front. Ecol. Evol.* **10**, 823588. ([doi:10.3389/fevo.2022.823588](https://doi.org/10.3389/fevo.2022.823588)) 6. Baluška F, Miller WB, Reber AS. 2023 Cellular and evolutionary perspectives on organismal cognition: from unicellular to multicellular organisms. *Biol. J. Linnean Soc.* **139**, 503–513. ([doi:10.1093/biolinnean/blac005](https://doi.org/10.1093/biolinnean/blac005)) 7. Miller WB, Baluška F, Reber AS. 2023 A revised central dogma for the 21st century: all biology is cognitive information processing. *Prog. Biophys. Mol. Biol.* **182**, 34–48. ([doi:10.1016/j.pbiomolbio.2023.05.005](https://doi.org/10.1016/j.pbiomolbio.2023.05.005)) 8. Stone JR. 1997 The spirit of D’arcy Thompson dwells in empirical morphospace. *Math. Biosci.* **142**, 13–30. ([doi:10.1016/s0025-5564(96)00186-1](https://doi.org/10.1016/s0025-5564(96)00186-1)) 9. Deisboeck TS, Couzin ID. 2009 Collective behavior in cancer cell populations. *Bioessays* **31**, 190–197. ([doi:10.1002/bies.200800084](https://doi.org/10.1002/bies.200800084)) 10. Abzhanov A. 2017 The old and new faces of morphology: the legacy of D’Arcy Thompson’s ‘theory of transformations’ and ‘laws of growth’. *Development* **144**, 4284–4297. ([doi:10.1242/dev.137505](https://doi.org/10.1242/dev.137505)) 11. Reber AS, Baluška F. 2021 Cognition in some surprising places. *Biochem. Biophys. Res. Commun.* **564**, 150–157. ([doi:10.1016/j.bbrc.2020.08.115](https://doi.org/10.1016/j.bbrc.2020.08.115)) 12. Lyon P, Keijzer F, Arendt D, Levin M. 2021 Reframing cognition: getting down to biological basics. *Phil. Trans. R. Soc. B* **376**, 20190750. ([doi:10.1098/rstb.2019.0750](https://doi.org/10.1098/rstb.2019.0750)) 13. Levin M. 2022 Technological approach to mind everywhere: an experimentally-grounded framework for understanding diverse bodies and minds. *Front. Syst. Neurosci.* **16**, 768201. ([doi:10.3389/fnsys.2022.768201](https://doi.org/10.3389/fnsys.2022.768201)) 14. Pio-Lopez L, Bischof J, LaPalme JV, Levin M. 2023 The scaling of goals from cellular to anatomical homeostasis: an evolutionary simulation, experiment and analysis. *Interface Focus* **13**, 20220072. ([doi:10.1098/rsfs.2022.0072](https://doi.org/10.1098/rsfs.2022.0072)) 15. Mathews J, Chang AJ, Devlin L, Levin M. 2023 Cellular signaling pathways as plastic, proto-cognitive systems: Implications for biomedicine. *Patterns* **4**, 100737. ([doi:10.1016/j.patter.2023.100737](https://doi.org/10.1016/j.patter.2023.100737)) 16. Levin M. 2023 Bioelectric networks: the cognitive glue enabling evolutionary scaling from physiology to mind. *Anim. Cogn.* **26**, 1865–1891. ([doi:10.1007/s10071-023-01780-3](https://doi.org/10.1007/s10071-023-01780-3)) 17. Levin M. 2023 Collective intelligence of morphogenesis as a teleonomic process. In *Evolution ‘on purpose’: teleonomy in living systems*. Cambridge, MA: The MIT Press. ([doi:10.7551/mitpress/14642.003.0013](https://doi.org/10.7551/mitpress/14642.003.0013)) 18. Lagasse E, Levin M. 2023 Future medicine: from molecular pathways to the collective intelligence of the body. *Trends Mol. Med.* **29**, 687–710. (doi:[10.1016/j.molmed.2023.06.007](https://doi.org/10.1016/j.molmed.2023.06.007)) 19. Zhang T, Goldstein A, Levin M. 2025 Classical sorting algorithms as a model of morphogenesis: self-sorting arrays reveal unexpected competencies in a minimal model of basal intelligence. *Adapt. Behav.* **33**, 25–54. (doi:[10.1177/10597123241269740](https://doi.org/10.1177/10597123241269740)) 20. Hopfield JJ. 1982 Neural networks and physical systems with emergent collective computational abilities. *Proc. Natl Acad. Sci. USA* **79**, 2554–2558. (doi:[10.1073/pnas.79.8.2554](https://doi.org/10.1073/pnas.79.8.2554)) 21. Krotov D, Hopfield JJ. 2016 Dense associative memory for pattern recognition. *Adv. Neural Inf. Process. Syst.* **29**. 22. Demircigil M, Heusel J, Löwe M, Upgang S, Vermet F. 2017 On a model of associative memory with huge storage capacity. *J. Stat. Phys.* **168**, 288–299. (doi:[10.1007/s10955-017-1806-y](https://doi.org/10.1007/s10955-017-1806-y)) 23. Castellani T, Cavagna A. 2005 Spin-glass theory for pedestrians. *J. Stat. Mech.: Theory Exp.* **2005**, 05012. (doi:[10.1088/1742-5468/2005/05/P05012](https://doi.org/10.1088/1742-5468/2005/05/P05012)) 24. Ranzato M, Poultney C, Chopra S, Le Cun Y. 2006 Efficient learning of sparse representations with an energy-based model. *Adv. Neural Inf. Process. Syst.* **19**. (doi:[10.7551/mitpress/7503.003.0147](https://doi.org/10.7551/mitpress/7503.003.0147)) 25. Yang L, Zhang Z, Song Y, Hong S, Xu R, Zhao Y, Zhang W, Cui B, Yang MH. 2023 Diffusion models: a comprehensive survey of methods and applications. *ACM Comput. Surv* **56**, 1–39. (doi:[10.1145/3626235](https://doi.org/10.1145/3626235)) 26. Khandelwal U, He H, Qi P, Jurafsky D. 2018 Sharp nearby, fuzzy far away: how neural language models use context. In *Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics*, pp. 284–294. Stroudsburg, PA, USA: Association for Computational Linguistics. (doi:[10.18653/v1/P18-1027](https://doi.org/10.18653/v1/P18-1027)) 27. Sun S, Krishna K, Mattarella-Micke A, Iyyer M. 2021 Do long-range language models actually use long-range context? In *Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing*, pp. 807–822. Stroudsburg, PA, USA: Association for Computational Linguistics. (doi:[10.18653/v1/2021.emnlp-main.62](https://doi.org/10.18653/v1/2021.emnlp-main.62)) 28. Zhao Z, Wallace E, Feng S, Klein D, Singh S. 2021 Calibrate before use: improving few-shot performance of language models. In *Proceedings of the 38th International Conference on Machine Learning*, vol. 139, pp. 12697–12706, Proceedings of Machine Learning Research. 29. Malkin N, Wang Z, Jojic N. 2022 Coherence boosting: when your pretrained language model is not paying enough attention. In *Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics*, pp. 8214–8236. Stroudsburg, PA, USA: Association for Computational Linguistics. (doi:[10.18653/v1/2022.acl-long.565](https://doi.org/10.18653/v1/2022.acl-long.565)) 30. Kwa T et al. 2026 Measuring AI Ability to Complete Long Software Tasks. In *The Thirty-ninth Annual Conference on Neural Information Processing Systems*. 31. Graner F, Glazier JA. 1992 Simulation of biological cell sorting using a two-dimensional extended Potts model. *Phys. Rev. Lett.* **69**, 2013–2016. (doi:[10.1103/PhysRevLett.69.2013](https://doi.org/10.1103/PhysRevLett.69.2013)) 32. Chaturvedi R et al. 2005 On multiscale approaches to three-dimensional modelling of morphogenesis. *J. R. Soc. Interface* **2**, 237–253. (doi:[10.1098/rsif.2005.0033](https://doi.org/10.1098/rsif.2005.0033)) 33. Torquato S. 2011 Toward an Ising model of cancer and beyond. *Phys. Biol.* **8**, 015017. (doi:[10.1088/1478-3975/8/1/015017](https://doi.org/10.1088/1478-3975/8/1/015017)) 34. Szabó A, Merks RMH. 2013 Cellular potts modeling of tumor growth, tumor invasion, and tumor evolution. *Front. Oncol.* **3**, 87. (doi:[10.3389/fonc.2013.00087](https://doi.org/10.3389/fonc.2013.00087)) 35. Weber M, Buceta J. 2016 The cellular Ising model: a framework for phase transitions in multicellular environments. *J. R. Soc. Interface* **13**, 20151092. (doi:[10.1098/rsif.2015.1092](https://doi.org/10.1098/rsif.2015.1092)) 36. Landau LD, Lifshitz EM. 1958 *Statistical physics*. vol. 5. Oxford, UK: Pergamon Press. 37. Peierls RE. 1936 On Ising’s model of ferromagnetism. *Math. Proc. Camb. Philos. Soc.* **32**, 477–481. (doi:[10.1017/S0305004100019174](https://doi.org/10.1017/S0305004100019174)) 38. Box GEP, Jenkins GM. 1970 *Time series analysis: forecasting and control*. San Francisco, CA: Holden-Day. 39. Sakthivadivel DAR. 2022 Magnetisation and mean field theory in the Ising model. *SciPost Phys. Lect. Notes* **35**. (doi:[10.21468/SciPostPhysLectNotes.35](https://doi.org/10.21468/SciPostPhysLectNotes.35)) 40. Lin T, Wang Y, Liu X, Qiu X. 2022 A survey of transformers. *AI Open* **3**, 111–132. (doi:[10.1016/j.aiopen.2022.10.001](https://doi.org/10.1016/j.aiopen.2022.10.001)) 41. Ramsauer H et al. 2021 Hopfield networks is all you need. In *International conference on learning representations*. <yap-show> Downloaded from http://royalsocietypublishing.org/rsta/article-pdf/doi/10.1098/rsta.2025.0011/6138032/rsta.2025.0011.pdf by guest on 05 July 2026 </yap-show> <yap-show> 42. Millidge B, Salvatori T, Song Y, Lukasiewicz T, Bogacz R. 2022 Universal Hopfield networks: a general framework for single‑shot associative memory models. In *International conference on machine learning*, pp. 15561–15583. 43. Krotov D. 2023 A new frontier for Hopfield networks. *Nat. Rev. Phys.* **5**, 366–367. (doi:10.1038/s42254‑023‑00595‑y) 44. Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. 2017 Attention is all you need. *Adv. Neural Inf. Process. Syst* **30**, 6000–6010. (doi:10.5555/3295222.3295349) 45. Phuong M, Hutter M. 2022 Formal Algorithms for Transformers. *arXiv* 2207.09238. [arXiv](https://arxiv.org/abs/2207.09238) 46. Thomson KS. 1988 *Morphogenesis and evolution*. Oxford, UK: Oxford University Press. 47. Engler AJ, Humbert PO, Wehrle‑Haller B, Weaver VM. 2009 Multiscale modeling of form and function. *Science* **324**, 208–212. (doi:10.1126/science.1170107) 48. Levin M. 2012 Morphogenetic fields in embryogenesis, regeneration, and cancer: non‑local control of complex patterning. *Biosystems* **109**, 243–261. (doi:10.1016/j.biosystems.2012.04.005) 49. Chan CJ, Heisenberg CP, Hiiragi T. 2017 Coordination of morphogenesis and cell‑fate speci‑fication in development. *Curr. Biol* **27**, R1024–R1035. (doi:10.1016/j.cub.2017.07.010) 50. Kuchling F, Friston K, Georgiev G, Levin M. 2020 Morphogenesis as bayesian inference: a variational approach to pattern formation and control in complex biological systems. *Phys. Life Rev.* **33**, 88–108. (doi:10.1016/j.plrev.2019.06.001) 51. McMillen P, Levin M. 2024 Collective intelligence: a unifying concept for integrating biology across scales and substrates. *Commun. Biol.* **7**, 378. (doi:10.1038/s42003‑024‑06037‑4) 52. Balduzzi D, Tononi G. 2008 Integrated information in discrete dynamical systems: motivation and theoretical framework. *PLoS Comput. Biol.* **4**, e1000091. (doi:10.1371/journal.pcbi.1000091) 53. Kanwisher N. 2010 Functional specificity in the human brain: a window into the func‑tional architecture of the mind. *Proc. Natl Acad. Sci. USA* **107**, 11163–11170. (doi:10.1073/pnas.1005062107) 54. Johnson MH. 2011 Interactive specialization: a domain‑general framework for human func‑tional brain development? *Dev. Cogn. Neurosci.* **1**, 7–21. (doi:10.1016/j.dcn.2010.07.003) 55. Ramstead MJD, Hesp C, Tschantz A, Smith R, Constant A, Friston K. 2021 Neural and pheno‑typic representation under the free‑energy principle. *Neuroscience & Biobehavioral Reviews* **120**, 109–122. (doi:10.1016/j.neubiorev.2020.11.024) 56. Sanchez T, Chen DTN, DeCamp SJ, Heymann M, Dogic Z. 2012 Spontaneous motion in hierarchically assembled active matter. *Nature* **491**, 431–434. (doi:10.1038/nature11591) 57. Haxton TK, Whitelam S. 2013 Do hierarchical structures assemble best via hierarchical pathways? *Soft Matter* **9**, 6851. (doi:10.1039/c3sm27637f) 58. Whitelam S. 2015 Hierarchical assembly may be a way to make large information‑rich structures. *Soft Matter* **11**, 8225–8235. (doi:10.1039/c5sm01375e) 59. Doncom KEB, Blackman LD, Wright DB, Gibson MI, O’Reilly RK. 2017 Dispersity effects in polymer self‑assemblies: a matter of hierarchical control. *Chem. Soc. Rev.* **46**, 4119–4134. (doi:10.1039/c6cs00818f) 60. McGivern P. 2020 Active materials: minimal models of cognition? *Adapt. Behav.* **28**, 441–451. (doi:10.1177/1059712319891742) 61. Wang X, Salmani M, Omidi P, Ren X, Rezagholizadeh M, Eshaghi A. 2024 Beyond the limits: a survey of techniques to extend the context length in large language models. In *Proceedings of the thirty‑third international joint conference on artificial intelligence*, pp. 8299–8307. Marina del Rey, CA: IJCAI. 62. Qian H. 2001 Relative entropy: free energy associated with equilibrium fluctuations and nonequilibrium deviations. *Phys. Rev. E* **63**, 042103. (doi:10.1103/PhysRevE.63.042103) 63. Seifert U. 2012 Stochastic thermodynamics, fluctuation theorems and molecular machines. *Rep. Prog. Phys.* **75**, 126001. (doi:10.1088/0034‑4885/75/12/126001) 64. Hohenberg PC, Krekhov AP. 2015 An introduction to the Ginzburg–Landau theory of phase transitions and nonequilibrium patterns. *Phys. Rep.* **572**, 1–42. (doi:10.1016/j.physrep.2015.01.001) 65. Seifert U. 2018 Stochastic thermodynamics: from principles to the cost of precision. *Physica A: Stat. Mech. Appl.* **504**, 176–191. (doi:10.1016/j.physa.2017.10.024) </yap-show> 66. Lucarini V, Pavliotis GA, Zagli N. 2020 Response theory and phase transitions for the thermodynamic limit of interacting identical systems. *Proc. R. Soc. A.* **476**, 20200688. [doi:10.1098/rspa.2020.0688](https://doi.org/10.1098/rspa.2020.0688) 67. Zakine R, Vanden-Eijnden E. 2023 Minimum-action method for nonequilibrium phase transitions. *Phys. Rev. X* **13**, 041044. [doi:10.1103/PhysRevX.13.041044](https://doi.org/10.1103/PhysRevX.13.041044) 68. Hohenberg PC, Halperin BI. 1977 Theory of dynamic critical phenomena. *Rev. Mod. Phys.* **49**, 435–479. [doi:10.1103/RevModPhys.49.435](https://doi.org/10.1103/RevModPhys.49.435) 69. Marro J, Dickman R. 2005 *Nonequilibrium phase transitions in lattice models*. Cambridge, UK: Cambridge University Press. 70. Fields C, Levin M. 2022 Competency in navigating arbitrary spaces as an invariant for analyzing cognition in diverse embodiments. *Entropy* **24**, 819. [doi:10.3390/e24060819](https://doi.org/10.3390/e24060819) 71. Levin M. 2023 Darwin’s agential materials: evolutionary implications of multiscale competency in developmental biology. *Cell. Mol. Life Sci.* **80**, 142. [doi:10.1007/s00018-023-04790-z](https://doi.org/10.1007/s00018-023-04790-z) 72. McShea DW. 2013 Machine wanting. *Stud. Hist. Phil. Sci. Part C: Stud. Hist. Phil. Biol. Biomed. Sci.* **44**, 679–687. [doi:10.1016/j.shpsc.2013.05.015](https://doi.org/10.1016/j.shpsc.2013.05.015) 73. Bongard J, Levin M. 2021 Living things are not (20th Century) machines: updating mechanism metaphors in light of the modern science of machine behavior. *Front. Ecol. Evol.* **9**, 650726. [doi:10.3389/fevo.2021.650726](https://doi.org/10.3389/fevo.2021.650726) 74. Barwich AS, Rodriguez MJ. 2024 Rage against the what? The machine metaphor in biology. *Biol. Phil.* **39**, 14. [doi:10.1007/s10539-024-09950-4](https://doi.org/10.1007/s10539-024-09950-4) 75. Kircher T, Bröhl H, Meier F, Engelen J. 2018 Formal thought disorders: from phenomenology to neurobiology. *Lancet Psychiatry* **5**, 515–526. [doi:10.1016/S2215-0366(18)30059-2](https://doi.org/10.1016/S2215-0366(18)30059-2) 76. Palmier-Claus J *et al.* 2017 Cognitive behavioural therapy for thought disorder in psychosis. *Psychosis* **9**, 347–357. [doi:10.1080/17522439.2017.1363276](https://doi.org/10.1080/17522439.2017.1363276) 77. Theraulaz G, Bonabeau E. 1999 A brief history of stigmergy. *Artif. Life* **5**, 97–116. [doi:10.1162/106454699568700](https://doi.org/10.1162/106454699568700) 78. Marsh L, Onof C. 2008 Stigmergic epistemology, stigmergic cognition. *Cogn. Syst. Res.* **9**, 136–149. [doi:10.1016/j.cogsys.2007.06.009](https://doi.org/10.1016/j.cogsys.2007.06.009) 79. Giuggioli L, Potts JR, Rubenstein DI, Levin SA. 2013 Stigmergy, collective actions, and animal social spacing. *Proc. Natl Acad. Sci. USA* **110**, 16904–16909. [doi:10.1073/pnas.1307071110](https://doi.org/10.1073/pnas.1307071110) 80. Gloag ES, Turnbull L, Whitchurch CB. 2015 Bacterial stigmergy: an organising principle of multicellular collective behaviours of bacteria. *Scientifica* **2015**, 387342. [doi:10.1155/2015/387342](https://doi.org/10.1155/2015/387342) 81. Heylighen F. 2016 Stigmergy as a universal coordination mechanism I: definition and components. *Cogn. Syst. Res.* **38**, 4–13. [doi:10.1016/j.cogsys.2015.12.002](https://doi.org/10.1016/j.cogsys.2015.12.002) 82. Sims R, Yilmaz Ö. 2023 Stigmergic coordination and minimal cognition in plants. *Adapt. Behav.* **31**, 265–280. [doi:10.1177/10597123221150817](https://doi.org/10.1177/10597123221150817) 83. Holland O, Melhuish C. 1999 Stigmergy, self-organization, and sorting in collective robotics. *Artif. Life* **5**, 173–202. [doi:10.1162/106454699568737](https://doi.org/10.1162/106454699568737) 84. Valckenaers P, Germain BS, Verstraete P, Van Brussel H. 2007 MAS coordination and control based on stigmergy. *Comput. Ind.* **58**, 621–629. [doi:10.1016/j.compind.2007.05.003](https://doi.org/10.1016/j.compind.2007.05.003) 85. Boldini A, Civitella M, Porfiri M. 2024 Stigmergy: from mathematical modelling to control. *R. Soc. Open Sci.* **11**, 240845. [doi:10.1098/rsos.240845](https://doi.org/10.1098/rsos.240845) 86. Amit DJ, Gutfreund H, Sompolinsky H. 1987 Statistical mechanics of neural networks near saturation. *Ann. Phys.* **173**, 30–67. 87. Mézard M. 2017 Mean-field message-passing equations in the hopfield model and its generalizations. *Phys. Rev. E* **95**, 022117. [doi:10.1103/PhysRevE.95.022117](https://doi.org/10.1103/PhysRevE.95.022117) 88. Edwards SF, Anderson PW. 1975 Theory of spin glasses. *J. Phys. F: Met. Phys.* **5**, 965–974. [doi:10.1088/0305-4608/5/5/017](https://doi.org/10.1088/0305-4608/5/5/017) 89. Sacco F, Sakthivadivel D, Levin M. 2023 Requirements for self-organization (v1.1). Zenodo. [doi:10.5281/zenodo.8416764](https://doi.org/10.5281/zenodo.8416764)