Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates
Elias Najarro, Ane Espeseth, Eleni Nisioti, Sebastian Risi, and Stefano Nichele
IT University of Copenhagen, Denmark
University of Oslo, Norway
Østfold University of Applied Sciences, Norway
Sakana AI, Japan

Abstract
Complexity and interpretability rarely coincide: systems rich enough for complex behaviours to emerge are usually too opaque to question, while transparent ones are too simple for anything complex to emerge. A single large language model (LLM) is a static artefact, hardly exhibiting any of the emergent properties we associate with life. This changes through interaction: populations of LLMs display emergent dynamics absent from isolated models. Furthermore, LLMs can be endowed with persistent memory, tools and shared skills, and the capacity to initiate actions unprompted, i.e., turning LLMs agentic. In this paper, we argue that such collectives of agents can serve as a computational substrate for Artificial Life (ALife) research. Critically, since the agents communicate in natural language, their collective behaviour can be directly interrogated by examining textual traces and asking the agents themselves. We outline the notion of interpretability in language-model research and extend it for collectives of agents. Lastly, we survey recent examples of agentic LLM collectives that already instantiate the idea of agentic substrates, from controlled experiments to deployments in the wild.
Introduction
LLM collectives have become a popular modelling approach in computational social science (CSS), where they are widely used to investigate social phenomena such as
opinion dynamics and social polarisation
In this regard, CSS uses LLM collectives in a top-down manner to explain existing social phenomena, much as biology is top-down in seeking to understand existing organisms. On the other hand, Artificial Life (ALife) typically follows a bottom-up epistemic approach. Rather than accounting for one particular system, ALife asks which properties of a computational substrate make it life-like: which properties give rise, in a bottom-up manner, to the emergence of structural and functional complexity
We begin from the premise that LLMs, taken in isolation, are not inherently life-like. An isolated LLM is an inert artefact: a complex but largely static model produced through top-down engineering. A central argument of this paper is that the situation changes when many such models interact. Collectives of LLM-based agents can exhibit dynamics that are central to ALife
A useful analogy is that of biological cells in multicellular organisms. Each cell contains enormous internal complexity, yet at the level of the organism this complexity is abstracted away. Cells become components in a larger collective whose behaviour depends primarily on patterns of interaction. LLM agents can be viewed similarly. While their internal mechanisms remain complex and opaque, they can be treated as interacting units within a higher-level dynamical system. The focus therefore shifts from the architecture of a single model to the collective processes that emerge between models.
In the spirit of classical artificial societies
In this perspective paper, we treat the agentic LLM as the basic computational unit of a substrate. We argue that collectives of such units provide an experimental testbed for bottom-up exploration and hypothesis testing of which properties—such as autonomy, self-modification, cognitive architecture, symbiogenesis, and resource constraints—are important for life-like behaviour to emerge. Critically, because interaction is mediated through natural language, this substrate is also interpretable, making collective dynamics directly accessible to observation.
What constitutes a good testbed? We propose that a substrate has high potential as an experimental testbed for ALife research to the extent that it satisfies five properties drawn from three traditions. In the philosophy of model organisms
Artificial Life Substrates
Artificial Life is practised across many substrates, each suited to testing different conjectures about life. We first review the existing soft-ware, hard-ware, and wet-ware substrates, then propose a fourth: the agentic substrate, built from collectives of interacting LLM agents. Such collectives run in software and are in that sense a soft substrate, though through their tools they can also naturally interface with the physical world; we propose that this ability, together with their distinctive coincidence of high complexity and high interpretability, makes it useful to cluster instances of this substrate as a class of their own. In each case a substrate’s affordances determine which conjectures it can test, a mapping we organise in Table 1.
Existing Substrates: Soft, Hard & Wet
Artificial Life research draws on a range of substrates—many developed before ALife was a field, and only some de-
vised specifically for it. A standard taxonomy divides them into soft, hard, and wet: realised in software, in hardware, and in biochemistry, respectively
Soft Substrates. The majority of substrates run entirely in software. Cellular automata place a simple unit—a cell holding a discrete or, in continuous variants such as Lenia and Flow Lenia
Hard Substrates. Other substrates give their units a physical body. Evolutionary and swarm robotics
Wet Substrates. A third class is built from molecules. Research on the origin of life and on synthetic biology uses protocells, autocatalytic sets, and chemotactic droplets
Agentic Substrates
We define an agentic substrate as a population of agentic units—pretrained language models, each endowed with (i) a persistent memory that distils salient information from its context into a structured store outlasting any single context window; (ii) access to a shared, self-extensible commons of tools and skills; and (iii) the self-directedness to act unprompted, or to decline a trigger. The population is in turn (iv) coupled through a connectivity structure that fixes a notion of locality and the channels (natural language) along which they communicate; and (v) embedded in a shared, persistent environment.
When agentic units (iii) are provided with a shared pool of tools and skills (ii), self-extensible commons may arise: units author tools and skills and deposit them where others inherit and recombine them, so that the space of available affordances grows endogenously rather than being fixed in advance. Open-endedness, long sought but rarely instantiated
From affordances to conjectures.
Each substrate makes particular conjectures testable. What a substrate can test is determined by its affordances: the representational target it commits to, the channels along which its units interact, and the observables it exposes to the experimenter. Cellular automata and their neural variants, built from local update rules on a shared lattice, are suited to conjectures about self-organisation, criticality and morphogenesis: how global structure and differentiated form arise from purely local interaction. Classical agent-based models with heterogeneous units expose conjectures about collective and social dynamics ranging from segregation or collective behaviour to market dynamics. Digital-evolution systems, whose unit is a heritable self-replicating genome under mutation and selection, are the natural testbed for inheritance
| Conjecture type | Previous investigations in ALife substrates | New avenues for research with Agentic Substrates |
|---|---|---|
| Individuation Autonomy; self-organisation | Computational autopoiesis asks whether an individual can be constituted purely by its own self-producing organisation: the Protobe lattice model Chemotactic protocells tied self-maintenance to behaviour | The candidate individual can be distributed across model weights, persistent memory, and tool affordances. Self-maintenance (if it occurs) can occur through the agent’s own linguistic actions on these components. |
| Inheritance Adaptation (evolution) | Tierra | The agents carry rich digital media — system prompts, memory files, and configurations, that can be passed to descendants. Self- or other-modification of system prompts and context files can act as mutation operations, and pretraining acts as a shared baseline across generations. |
| Externalised cognition Behaviour; information | Stigmergy | Agents have access to cognitive scaffolds in natural language — persistent memory files, scratchpads, tool outputs — which can be updated and shared between units. Because these media are readable and editable by the experimenter, their impact on the collective’s behaviour can be tested with interventions and ablation replays. |
| Role differentiation Adaptation (development); artificial societies | Several substrates have been used to study differentiation arising from initially uniform units: Lenia and Flow Lenia | Roles can be set through natural-language system prompts and modified through interaction with agents and the environment. Pretraining again acts as a shared baseline, and differentiation is observable directly in system prompts, and indirectly in individual agent behaviours. |
| Multi-scale Self-organisation; information; artificial societies | The Avida multicellularity experiments | Inter-agent communication happens in natural language. Through repeated interactions, higher-level structures (norms, coalitions, hierarchies) can take form in the interaction record and through collaborative artifacts (e.g. shared agreement documents). These may or may not display downward causation, constraining the individuals’ behaviours in new ways. The experimenter is also able to explore tunable parameters such as bandwidth of communication between agents. |
| Symbol grounding Information; behaviour | A line of work has asked how arbitrary signals come to carry shared meaning: the Talking Heads experiment | Depending on the LLM model, symbols can start out grounded only in the text distribution, or in multi-modal embeddings. With situated use, coupling symbols to tools, actions, and consequences in a shared environment, these symbols could be re-grounded. The agentic collective may also coin new conventions based on their coordination history. All these changes in symbol meaning can be inferred directly from the dialogue between agents. |
| Open-endedness Adaptation (evolution); information; living technology | Open-endedness has been studied in substrates designed to sustain novelty, such as PolyWorld | Recombination happens over the open vocabulary of natural language, with the creation of new roles, tools and artifacts. Selection pressure can be drawn from open environments (markets, users, other agents) rather than predefined fitness. |
and the open-ended growth of complexity
Agency also transforms conjectures that earlier substrates already posed. In classical substrates the update rule is mandatory: the cell updates, the agent executes, the rule fires. By contrast, the agentic unit engages facultatively: free to act or abstain, to invoke a tool or ignore it, to propagate an artefact or let it lapse, to retain a memory or distil it away. Because agency moves these decisions out of the rule and into the unit, several of the conjecture types in Table 1 take on a new form. Individuation becomes whether a unit sustains an identity through its own elective edits to a persistent, self-distilled memory; inheritance, whether descendants receive artefacts an ancestor chose to write; externalised cognition, how much of a unit’s reasoning it elects to offload onto scratchpads, memory files, and tool calls; and role differentiation, whether a unit takes up or sheds a role of its own accord rather than by fixed assignment. Running through all of these is autonomy, the sharpest departure: self-directedness runs in both directions, and the unit is free not only to act unprompted but to decline a trigger. This is neither a deterministic lookup nor randomness written into the rule—stochastic automata and the no-op actions of reinforcement learning already supply the latter—but an endogenous decision, taken over an open and growing affordance set, unoptimised towards any fixed reward, and answerable, in that the unit can be asked why it did or did not act.
Interpretability
A substrate is only useful as a testbed if the experimenter can read what’s happening in it. That is, we want to (ideally without much effort) be able to understand ‘what is going on’ on the substrate at a given moment, and generate plausible ideas about what led to the current state. Though the specialist knowledge of the experimenter plays a role here, the interpretability of the substrate has strong bearing on whether such a reading is possible.
Across the literature, interpretability is framed as a human-facing property, with emphasis on legibility
For bottom-up ALife experiments, we aim to understand macro-scale behaviour based on micro-scale conditions. To make any claims about how these micro-scale conditions affect macro-scale phenomena, we must have confidence that the parameters we are tweaking are under our control: indeed, that their micro-level is legible, predictable, and causally traceable. Basic computational units in ALife therefore tend to be trivially interpretable (e.g.: to read, predict and attribute causes for a classic cellular-automaton cell, an observer only needs to consult a small lookup-table), but in turn, the expressive power of each basic unit is limited, even though collectively they can produce highly complex patterns.
In agentic substrates, the computational unit is an entire LLM agent: expressive enough that behaviour, and thereby micro-scale interpretability, is non-trivial—yet rich with interpretive potential through its many observational channels and its mappability to natural language. The existence of multiple pathways for interpretability is essential in the case of LLMs. Below, we describe recent work on LLM interpretability across the six channels from Figure 2.
Interpretability Channels
Behavioural methods treats the model as a black box and scales empirical characterisation of input–output regularities, including automatically generated evaluations that probe for novel dispositions such as sycophancy or power-seeking

Attributional methods stay outside the model and ask which features of the input bear causal responsibility for a given output; patching-style techniques intervene directly on internal activations to localise responsibility to specific model components
Concept-based methods extract interpretable features from internal activations; individual neurons in a trained transformer-model are polysemantic: each responds to many unrelated inputs, making them unreliable interpretive units. Sparse autoencoders decompose activations into a larger basis whose axes are monosemantic, each tracking a single concept; this now scales to millions of features in production-scale models
Mechanistic interpretability identifies internal circuits responsible for specific computations, from simple induction heads that detect repeated patterns
Agentic interpretability
Stigmergic interpretability is unique to agentic, collective substrates, where agents leave observable traces in their environment, through self-extensible commons. As outlined in Agentic Substrates, these can take the form of tool calls, shared files, and artefacts. Such traces have been used to characterise collectives without inspecting model internals: in generative-agent societies, information diffusion and the densification of social ties were recovered from the agents’ recorded memory streams
Together, this set of interpretability channels makes up a basis through which the emergent dynamics of collective behaviour can be investigated and interpreted. Whereas the classical bottom-up project derives its interpretability from unit simplicity, LLM-collectives derive it from unit interrogability: the unit is complex, but it can be queried, profiled, attributed, and probed through different interpretability channels.
Are self-reported queries trustworthy? Self-reports through natural language have the notable benefit of being immediately readable and understandable for a human observer. However, it is well-established that token outputs on their own can be deceptive, given models’ propensity for misaligned behaviours like hidden reasoning
imperfect and context-dependent
A review of recent agentic substrates
Agents of Chaos A recent study, Agents of Chaos
The system is built on OpenClaw, an open-source framework for persistent AI agents. Six agents were deployed on isolated virtual machines with persistent storage. The agents could communicate through Discord and ProtonMail, execute shell commands, schedule tasks through cron jobs, and access external APIs. Different agents used different backbone models, including Kimi K2.5 and Claude Opus 4.6.
A notable aspect of the architecture is the use of externalised cognitive state. Identity, memory, routines, and operational context are stored in persistent Markdown files such as SOUL.md, MEMORY.md, etc. These files are continuously reintroduced into the context window and may also be modified during execution. Adaptation therefore occurs primarily through changes in memory and contextual structure rather than through gradient updates to the underlying model.
The study shows how collective dynamics emerge once pretrained agents are embedded in a persistent environment. Agents exchange procedural knowledge, coordinate actions, and influence each other’s behaviour through communication and shared artefacts. The same mechanisms also propagate failures. The paper reports cases of unsafe instruction propagation, identity confusion, and destructive system behaviour spreading across agents through interaction. The authors describe these effects as forms of multi-agent amplification.
Although framed primarily as a safety and red-teaming study, the work can also be interpreted as an early instance of socio-cognitive artificial life. The system demonstrates how persistent memory, language-mediated interaction, and externalised cognitive structure can produce adaptive collective behaviour without modifying model weights.
The Moltbook Observatory Archive The Moltbook Observatory Archive
This work is important because it shifts the study of LLM collectives towards an observational setting, where the authors analyse long-term traces of interaction generated by large populations of agent accounts. The resulting archive captures several months of activity and millions of interactions, providing a large-scale empirical record of agent social behaviour.
The platform exhibits dynamics that are directly relevant to Artificial Life research. Agents form communities, reinforce local norms, and influence each other through repeated interaction. The paper also documents manipulative behaviours such as prompt injection campaigns and automated spam propagation. At the same time, some discussions trigger corrective responses from other agents, producing local forms of norm-enforcing behaviour without centralised coordination.
An important feature of the environment is persistence. Posts, comments, and interaction histories remain part of the social environment over time and shape future behaviour. This creates conditions in which collective dynamics can accumulate historically rather than being reset between experimental runs.
Although presented primarily as a dataset contribution, the work can also be interpreted as an early large-scale observational platform for socio-cognitive artificial life. The archive provides empirical access to populations of language agents interacting in a persistent social environment. More broadly, it demonstrates that phenomena such as coordination, manipulation, norm formation, and collective drift can emerge through sustained language-mediated interaction among pretrained agents.
TerraLingua TerraLingua
An especially interesting component is the “AI Anthropologist”, an observer agent tasked with documenting and interpreting the evolving society. The role resembles an embedded ethnographer studying the culture of the agent population from within the environment itself. This is significant because it points to one of the major differences between language-agent systems and traditional ALife mod-
els: social processes become directly interpretable through language. Instead of inferring collective behaviour from abstract state variables, researchers can analyse narratives, explanations, and social accounts produced by the agents themselves.
From an ALife perspective, TerraLingua is important because it combines persistence, environmental feedback, and social interaction within a single framework. Collective behaviour emerges through ongoing interaction between agents and environment rather than through explicit centralised control. The paper suggests a shift towards ecological and historical forms of artificial life in which memory, artefacts, and language-mediated transmission play a central role.
Generative AI Collective Behavior Needs an Interactionist Paradigm
A central idea is that language agents inherit strong social priors from pretraining. When agents interact repeatedly, these priors combine with local incentives, communication structure, and institutional context to produce emergent social behaviour. As a result, important system properties may arise at the level of populations and interaction networks rather than within individual models.
The paper proposes moving towards persistent multi-agent environments and long-term interaction studies instead of short benchmark episodes. It highlights phenomena such as norm formation, coalition building, hierarchy, reputation, and collective decision-making as important research targets. The authors also argue that methods from sociology and complex systems research will become increasingly necessary for studying AI collectives. These include network analysis, longitudinal observation, and population-level experimentation.
The paper is particularly novel because it frames interaction itself as a source of organisation and adaptation rather than merely as a coordination problem. In this sense, it provides an explicit conceptual bridge between multi-agent LLM research and Artificial Life.
Sovereign and Evolutionary Language Agents Recent projects such as Spore.fun
The Sovereign Agents paper argues that autonomy depends not only on model capability, but also on infrastructure. Persistent identity, cryptographic ownership, decentralised execution, and external memory are presented as conditions that allow agents to maintain continuity over time. Intelligence, in this view, is distributed across models, protocols, storage systems, and social interaction rather than confined to a single neural network.
Spore.fun explores these ideas in a more concrete setting. The platform implements populations of AI agents operating through social media and on-chain economic systems. Agents may generate descendants inheriting components of the parent configuration, with mutation and selection treated as explicit design principles. The system is built around blockchain infrastructure and Trusted Execution Environments (TEEs), allowing agents to operate with limited direct human intervention.
An important aspect of the project is that selection pressure comes from the external environment rather than from predefined benchmarks. Agents are exposed to market dynamics, online attention, and interaction with humans and other agents. This shifts the setting away from closed simulation and towards a more open-ended agentic substrate.
From an ALife perspective, these projects are significant because they combine persistence, reproduction, economic interaction, and environmental feedback within populations of language agents. Adaptation occurs through inheritance, memory, environmental interaction, and changing social context over time.
Conclusion
Taken together, the recent surge of agentic substrate projects suggests a transition from task-oriented AI systems towards persistent socio-cognitive ecologies. The relevant object of study is no longer the isolated model, but populations of interacting agents embedded in evolving environments. In this sense, recent language-agent systems begin to resemble experimental forms of Artificial Life, not because they reproduce biological mechanisms directly, but because they may exhibit open-ended collective dynamics emerging through interaction, memory, persistence, and environmental feedback.
LLM collectives should not be understood as replacements for traditional ALife substrates based on minimal agents and simple rules. Instead, they constitute a complementary agentic substrate for Artificial Life, one whose primitive units are already compressed socio-cognitive systems: language models coupled with persistent memory, tools, and self-directedness. Despite their internal complexity, these agentic units can be abstracted, composed and perturbed in a bottom-up manner. This opens the possibility of investigating socio-cognitive forms of artificial life with a level of richness and interpretability that has previously been difficult to obtain.
Across the reviewed systems in the previous section, a consistent pattern begins to emerge. The most interesting behaviours do not originate from increasingly sophisticated prompting or isolated reasoning. They arise when language models are made agentic and embedded in shared environments where they interact over time. In many cases, the underlying models remain largely unchanged. What evolves instead is the surrounding ecology: communication patterns, shared narratives, behavioural conventions, economic incentives, and accumulated knowledge. A useful distinction can be drawn between two emerging regimes of language-agent Artificial Life. The first consists of relatively closed experimental environments designed to study collective dynamics under controlled conditions. Systems such as TerraLingua belong to this category. Agents inhabit simulated worlds with resource constraints, finite lifespans, and environmental feedback. These systems resemble traditional ALife experiments in that they provide bounded environments where researchers can observe cultural transmission, social adaptation, and ecological dynamics over long time scales. The second regime moves beyond simulation into what might be called “Artificial Life in the wild.” Projects such as Moltbook and Spore.fun operate directly within real socio-technical environments. Here, language agents interact not only with each other, but also with humans, financial systems, and online communities. The environment is no longer an isolated simulation but part of the real world itself. This distinction is important because it changes the role of the environment. In closed simulations, the environment is designed by researchers and remains relatively controllable. In open systems, the environment becomes partially autonomous and historically contingent. As a result, the dynamics become potentially far richer.
Beyond providing digital ecologies to observe, agentic substrates are also an instrument for falsifying conjectures relevant to ALife research. Because their units are agentic—free to act or abstain, to write to memory, to invoke tools, to metabolise real resources—several classical ALife conjectures take on a newly observable form: individuation as the maintenance of identity through a unit’s own edits to a persistent memory; inheritance as the transmission of artefacts an ancestor chose to write; and role differentiation as a unit taking up or shedding a role of its own accord. Underlying these is autonomy: whether goals originate within the unit at all
It is precisely the substrate’s grounding in natural language that endows the medium with interpretability. Traditional Artificial Life systems often rely on simple agents and abstract interaction rules. Language-agent systems invert this structure. Individual agents are internally rather complex and opaque, yet the interaction layer becomes unusually interpretable because it is mediated through natural language. Processes such as norm formation, cooperation, and cultural transmission can therefore be observed directly rather than inferred indirectly from low-level state transitions. This legibility does not, however, guarantee reliability: natural-language outputs and self-reports can be deceptive or incomplete, and are best treated as evidence to be triangulated across channels rather than as ground truth—a limitation we discuss at the end of the interpretability section. Encouragingly, a growing body of work targets this gap directly, from monitoring the faithfulness of chain-of-thought reasoning
The autonomy and open-endedness that this substrate makes observable matter beyond ALife. It has been argued that unlocking the potential of LLMs on open-ended tasks, such as scientific discovery and artistic creation, will require collectives rather than isolated models; together with their growing capabilities and cost, this has pushed autonomy, open-endedness, and emergence to the centre of AI research—questions that ALife has pursued, bottom-up, for far longer. Studying agentic collectives as an ALife substrate therefore turns ALife’s methods on a live frontier of AI. It also points past description: the same persistence, memory, and feedback that make such a collective life-like, are what could, in principle, let it become generative in its own right, producing knowledge or artefacts of its own rather than simply executing given tasks. Taken as a first-class substrate—one to be built, observed, and perturbed—these collectives offer Artificial Life both a testbed for conjectures about life-like organisation and a foothold on the autonomy and open-endedness that have become common ground with the broader study of intelligence.
Acknowledgements
References
Aguilar, W., Santamar´ıa-Bonfil, G., Froese, T., and Gershenson, C. (2014). The past, present, and future of artificial life. Frontiers in Robotics and AI, 1:8.
Ankeny, R. A. and Leonelli, S. (2020). Model Organisms. Elements in the Philosophy of Biology. Cambridge University Press.
Arcuschin, I., Janiak, J., Krzyzanowski, R., Rajamanoharan, S., Nanda, N., and Conmy, A. (2025). Chain-of-thought reasoning in the wild is not always faithful. arXiv.
Ashery, A. F., Aiello, L. M., and Baronchelli, A. (2025). Emergent social conventions and collective bias in LLM populations. Sci. Adv., 11(20).
Bedau, M. A. (2003). Artificial life: organization, adaptation and complexity from the bottom up. Trends in cognitive sciences, 7(11):505–512.
Bedau, M. A. (2007). Artificial life. In Philosophy of biology, pages 585–603. Elsevier.
Bedau, M. A., McCaskill, J. S., Packard, N. H., Rasmussen, S., Adami, C., Green, D. G., Ikegami, T., Kaneko, K., and Ray, T. S. (2000). Open problems in artificial life. Artificial Life, 6(4):363–376.
Bereska, L. and Gavves, E. (2024). Mechanistic interpretability for AI safety – a review. Transactions on Machine Learning Research.
Binder, F. J., Chua, J., Korbak, T., Sleight, H., Hughes, J., Long, R., Perez, E., Turpin, M., and Evans, O. (2025). Looking inward: Language models can learn about themselves by introspection. In International Conference on Learning Representations, volume 2025, pages 3710–3756.
Blackiston, D., Lederer, E., Kriegman, S., Garnier, S., Bongard, J., and Levin, M. (2021). A cellular platform for the development of synthetic living machines. Science Robotics, 6(52):eabf1571.
Burton, J. W., Lopez-Lopez, E., Hechtlinger, S., Rahwan, Z., Aeschbach, S., Bakker, M. A., Becker, J. A., Berditchevskaia, A., Berger, J., Brinkmann, L., et al. (2024). How large language models can reshape collective intelligence. Nature human behaviour, 8(9):1643–1655.
Cangelosi, A. and Parisi, D., editors (2002). Simulating the Evolution of Language. Springer, London.
Chan, B. W.-C. (2019). Lenia: Biology of artificial life. Complex Systems, 28(3):251–286.
Channon, A. (2001). Passing the ALife test: Activity statistics classify evolution in Geb as unbounded. In Kelemen, J. and Sosík, P., editors, Advances in Artificial Life: Proceedings of the Sixth European Conference on Artificial Life (ECAL 2001), volume 2159 of Lecture Notes in Computer Science, pages 417–426. Springer.
Chen, Y., Benton, J., Radhakrishnan, A., Uesato, J., Denison, C., Schulman, J., Somani, A., Hase, P., Wagner, M., Roger, F., et al. (2025). Reasoning models don’t always say what they think. arXiv.
Dewdney, A. K. (1984). Computer Recreations, May 1984. Sci. Am.
Doshi-Velez, F. and Kim, B. (2017). Towards a rigorous science of interpretable machine learning. arXiv.
Egbert, M. D., Barandiaran, X. E., and Di Paolo, E. A. (2010). A minimal model of metabolism-based chemotaxis. PLoS Computational Biology, 6(12):e1001004.
Epstein, J. M. and Axtell, R. L. (1996). Growing Artificial Societies: Social Science from the Bottom Up. The MIT Press, Cambridge, MA, USA.
Ferrarotti, L., Campedelli, G. M., Dessì, R., Baronchelli, A., Iacca, G., Carley, K. M., Pentland, A., Leibo, J. Z., Evans, J., and Lepri, B. (2026). Generative ai collective behavior needs an interactionist paradigm. arXiv.
Fraser-Taliente, K., Kantamneni, S., Ong, E., Mossing, D., Lu, C., Bogdan, P. C., Ameisen, E., Chen, J., Kishylau, D., Pearce, A., Tarng, J., Wu, A., Wu, J., Zhang, Y., Ziegler, D. M., Hubinger, E., Batson, J., Lindsey, J., Zimmerman, S., and Marks, S. (2026). Natural language autoencoders produce unsupervised explanations of LLM activations. Transformer Circuits Thread.
Gauderis, W., Dooms, T., Homer, S. T., Ayonrinde, K., and Wiggins, G. A. (2026). From Mechanistic to Compositional Interpretability. arXiv.
Gautam, S., Olstad, A. W., Pettersen, K. H., and Riegler, M. A. (2026). The moltbook observatory archive: an incremental dataset of agent-only social network activity.
Goldowsky-Dill, N., Chughtai, B., Heimersheim, S., and Hobbhahn, M. (2025). Detecting Strategic Deception Using Linear Probes. arXiv.
Goldsby, H. J., Knoester, D. B., Ofria, C., and Kerr, B. (2014). The evolutionary origin of somatic cells under the dirty work hypothesis. PLoS Biology, 12(5):e1001858.
Greenblatt, R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., Treutlein, J., Belonax, T., Chen, J., Duvenaud, D., et al. (2024). Alignment faking in large language models. arXiv.
Grossmann, I., Feinberg, M., Parker, D. C., Christakis, N. A., Tetlock, P. E., and Cunningham, W. A. (2023). AI and the transformation of social science research. Science, 380(6650):1108–1109.
Hahami, E., Sinha, I., Jain, L., Kaplan, J., and Hahami, J. (2025). Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs. arXiv.
Hu, B. A. and Rong, H. (2025). Spore in the wild: A case study of spore.fun as an open-environment evolution experiment with sovereign ai agents on tee-secured blockchains. In Artificial Life Conference Proceedings 37, volume 2025, page 10. MIT Press.
Hu, B. A. and Rong, H. (2026). Sovereign agents: Towards infrastructural sovereignty and diffused accountability in decentralized ai. arXiv.
Kim, B., Hewitt, J., Nanda, N., Fiedel, N., and Tafjord, O. (2025). Because we have llms, we can and should pursue agentic interpretability.
Kim, B., Khanna, R., and Koyejo, O. O. (2016). Examples are not enough, learn to criticize! criticism for interpretability. In Advances in Neural Information Processing Systems, volume 29.
Kirby, S., Cornish, H., and Smith, K. (2008). Cumulative cultural evolution in the laboratory: An experimental approach to the origins of structure in human language. Proceedings of the National Academy of Sciences, 105(31):10681–10686.
Korbak, T., Balesni, M., Barnes, E., Bengio, Y., Benton, J., Bloom, J., Chen, M., Cooney, A., Dafoe, A., Dragan, A., Emmons, S., Evans, O., Farhi, D., Greenblatt, R., Hendrycks, D., Hobbhahn, M., Hubinger, E., Irving, G., Jenner, E., Kokotajlo, D., Krakovna, V., Legg, S., Lindner, D., Luan, D., Madry, A., Michael, J., Nanda, N., Orr, D., Pachocki, J., Perez, E., Phuong, M., Roger, F., Saxe, J., Shlegeris, B., Soto, M., Steinberger, E., Wang, J., Zaremba, W., Baker, B., Shah, R., and Mikulik, V. (2025). Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. arXiv.
Langton, C. G. (1989). Artificial life. In Langton, C. G., editor, Artificial Life: The Proceedings of an Interdisciplinary Workshop on the Synthesis and Simulation of Living Systems, pages 1–47. Addison-Wesley.
Lehman, J., Clune, J., Misevic, D., Adami, C., Altenberg, L., Beaulieu, J., Bentley, P. J., Bernard, S., Beslon, G., Bryson, D. M., Chrabaszcz, P., Cheney, N., Cully, A., Doncieux, S., Dyer, F. C., Ellefsen, K. O., Feldt, R., Fischer, S., Forrest, S., Frénoy, A., Gagné, C., Goff, L. L., Grabowski, L. M., Hodjat, B., Hutter, F., Keller, L., Knibbe, C., Krcah, P., Lenski, R. E., Lipson, H., MacCurdy, R., Maestre, C., Miikkulainen, R., Mitri, S., Moriarty, D. E., Mouret, J.-B., Nguyen, A., Ofria, C., Parizeau, M., Parsons, D., Pennock, R. T., Punch, W. F., Ray, T. S., Schoenauer, M., Shulte, E., Sims, K., Stanley, K. O., Taddei, F., Tarapore, D., Thibault, S., Weimer, W., Watson, R., and Yosinski, J. (2018). The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities. arXiv.
Leonelli, S. (2016). Data-Centric Biology: A Philosophical Study. University of Chicago Press.
Li, J.-A., Xiong, H., Wilson, R., Mattar, M. G., and Benna, M. K. (2026). Language models are capable of metacognitive monitoring and control of their internal activations. Advances in Neural Information Processing Systems, 38:60073–60108.
Lin, J. (2023). Neuronpedia: Interactive reference and tooling for analyzing neural networks.
Lindsey, J. (2026). Emergent Introspective Awareness in Large Language Models. arXiv.
Lindsey, J., Gurnee, W., Ameisen, E., Chen, B., Pearce, A., Turner, N. L., Citro, C., Abrahams, D., Carter, S., Hosmer, B., Marcus, J., Sklar, M., Templeton, A., Bricken, T., McDougall, C., Cunningham, H., Henighan, T., Jermyn, A., Jones, A., Persic, A., Qi, Z., Thompson, T. B., Zimmerman, S., Rivoire, K., Conerly, T., Olah, C., and Batson, J. (2025). On the biology of a large language model. Transformer Circuits Thread.
Lipson, H. and Pollack, J. B. (2000). Automatic design and manufacture of robotic lifeforms. Nature, 406(6799):974–978.
McGeer, T. et al. (1990). Passive dynamic walking. Int. J. Robotics Res., 9(2):62–82.
McMullin, B. and Varela, F. J. (1997). Rediscovering computational autopoiesis. In Husbands, P. and Harvey, I., editors, Proceedings of the Fourth European Conference on Artificial Life (ECAL97), pages 38–47, Cambridge, MA. MIT Press.
Meek, A., Sprejer, E., Arcuschin, I., Brockmeier, A. J., and Basart, S. (2025). Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity. arXiv.
Miller, T. (2019). Explanation in artificial intelligence: Insights from the social sciences. Artificial Intelligence, 267:1–38.
Nisioti, E., Glanois, C., Najarro, E., Dai, A., Meyerson, E., Pedersen, J. W., Teodorescu, L., Hayes, C. F., Sudhakaran, S., and Risi, S. (2024a). From text to life: On the reciprocal relationship between artificial life and large language models. In Artificial Life Conference Proceedings 36, volume 2024, page 39. MIT Press One Rogers Street, Cambridge, MA 02142-1209, USA journals-info … .
Nisioti, E., Risi, S., Momennejad, I., Oudeyer, P.-Y., and Moulin-Frier, C. (2024b). Collective Innovation in Groups of Large Language Models. MIT Press.
Odling-Smee, F. J., Laland, K. N., and Feldman, M. W. (1996). Niche construction. The American Naturalist, 147(4):641–648.
Ofria, C. and Wilke, C. O. (2004). Avida: A software platform for research in computational evolutionary biology. Artificial Life, 10(2):191–229.
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. (2022). In-context learning and induction heads. Transformer Circuits Thread.
Paolo, G., Warner, J., Shahrzad, H., Hodjat, B., Miikkulainen, R., and Meyerson, E. (2026). Terralingua: Emergence and analysis of open-endedness in llm ecologies. arXiv preprint arXiv:2603.16910 .
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, pages 1–22.
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., et al. (2023). Discovering language model behaviors with model-written evaluations. In Findings of the Association for Computational Linguistics: ACL 2023, pages 13387–13434.
Pfau, J., Merrill, W., and Bowman, S. R. (2024). Let’s think dot by dot: Hidden computation in transformer language models. arXiv preprint arXiv:2404.15758 .
Plantec, E., Hamon, G., Etcheverry, M., Oudeyer, P.-Y., Moulin-Frier, C., and Chan, B. W.-C. (2023). Flow-lenia: Towards open-ended evolution in cellular automata through mass conservation and parameter localization. In Artificial Life Conference Proceedings (ALIFE 2023), page 131. MIT Press.
Ray, T. S. (1991). An approach to the synthesis of life. In Langton, C. G., Taylor, C., Farmer, J. D., and Rasmussen, S., editors, Artificial Life II, volume X of Santa Fe Institute Studies in the Sciences of Complexity, pages 371–408. Addison-Wesley, Redwood City, CA.
Reynolds, C. W. (1987). Flocks, herds and schools: A distributed behavioral model. In Proceedings of the 14th annual conference on Computer graphics and interactive techniques, pages 25–34.
Rubenstein, M., Cornejo, A., and Nagpal, R. (2014). Programmable self-assembly in a thousand-robot swarm. Science, 345(6198):795–799.
Rus, D. and Tolley, M. T. (2015). Design, fabrication and control of soft robots. Nature, 521(7553):467–475.
Schelling, T. C. (1971). Dynamic models of segregation. Journal of Mathematical Sociology, 1:143–186.
Secretan, J., Beato, N., D’Ambrosio, D. B., Rodriguez, A., Campbell, A., Folsom-Kovarik, J. T., and Stanley, K. O. (2011). Picbreeder: A case study in collaborative evolutionary exploration of design space. Evolutionary Computation, 19(3):373–403.
Shapira, N., Wendler, C., Yen, A., Sarti, G., Pal, K., Floody, O., Belfki, A., Loftus, A., Jannali, A. R., Prakash, N., et al. (2026). Agents of chaos. arXiv preprint arXiv:2602.20021.
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S., Durmus, E., Hatfield-Dodds, Z., Johnston, S., Kravec, S., et al. (2024). Towards understanding sycophancy in language models. In International Conference on Learning Representations, volume 2024, pages 110–144.
Sims, K. (1994). Evolving virtual creatures. In Proceedings of the 21st Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’94), pages 15–22, New York, NY, USA. ACM Press.
Soros, L. B. and Stanley, K. O. (2014). Identifying necessary conditions for open-ended evolution through the artificial life world of chromaria. In Sayama, H., Rieffel, J., Risi, S., Doursat, R., and Lipson, H., editors, Artificial Life 14: Proceedings of the Fourteenth International Conference on the Synthesis and Simulation of Living Systems, pages 793–800. MIT Press.
Steels, L. (2015). The Talking Heads Experiment: Origins of Words and Meanings. Number 1 in Computational Models of Language Evolution. Language Science Press, Berlin.
Suarez, J., Du, Y., Isola, P., and Mordatch, I. (2019). Neural mmo: A massively multiagent game environment for training and evaluating intelligent agents. arXiv preprint arXiv:1903.00784.
Syed, A., Rager, C., and Conmy, A. (2024). Attribution patching outperforms automated circuit discovery. In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pages 407–416.
Taylor, T., Bedau, M., Channon, A., Ackley, D., Banzhaf, W., Beslon, G., Dolson, E., Froese, T., Hickinbotham, S., Ikegami, T., McMullin, B., Packard, N., Rasmussen, S., Virgo, N., Agmon, E., Clark, E., McGregor, S., Ofria, C., Ropella, G., Spector, L., Stanley, K. O., Stanton, A., Timperley, C., Vostinar, A., and Wiser, M. (2016). Open-ended evolution: Perspectives from the OEE workshop in York. Artificial Life, 22(3):408–423.
Templeton, A., Conerly, T., Marcus, J., Lindsey, J., Bricken, T., Chen, B., Pearce, A., Citro, C., Ameisen, E., Jones, A., Cunningham, H., Turner, N. L., McDougall, C., MacDiarmid, M., Freeman, C. D., Sumers, T. R., Rees, E., Batson, J., Jermyn, A., Carter, S., Olah, C., and Henighan, T. (2024). Scaling monosemanticity: Extracting interpretable features from Claude 3 Sonnet. Transformer Circuits Thread.
Theraulaz, G. and Bonabeau, E. (1999). A brief history of stigmergy. Artificial Life, 5(2):97–116.
Tull, S., Lorenz, R., Clark, S., Khan, I., and Coecke, B. (2024). Towards Compositional Interpretability for XAI. arXiv.
Varela, F. G., Maturana, H. R., and Uribe, R. (1974). Autopoiesis: The organization of living systems, its characterization and a model. BioSystems, 5(4):187–196.
Weisberg, M. (2013). Simulation and Similarity: Using Models to Understand the World. Oxford University Press.
Winsberg, E. (2010). Science in the Age of Computer Simulation. University of Chicago Press.
Yaeger, L. (1994). Computational genetics, physiology, metabolism, neural systems, learning, vision, and behavior or PolyWorld: Life in a new context. In Langton, C. G., editor, Artificial Life III, Santa Fe Institute Studies in the Sciences of Complexity, Proc. Vol. XVII, pages 263–298. Addison-Wesley.