Biological Theory [https://doi.org/10.1007/s13752-026-00544-9]()

ORIGINAL ARTICLE

What Lives? A Meta-Analysis of Diverse Opinions on the Definition of Life

Reed Bender, Karina Kofman, Blaise Agüera y Arcas, Michael Levin

Received: 16 October 2025 / Accepted: 11 April 2026 © The Author(s) 2026

Abstract

The question of “what is life?” has challenged scientists and philosophers for centuries, producing an array of definitions that reflect both the mystery of its emergence and the diversity of disciplinary perspectives brought to bear on the question. Despite significant progress in our understanding of biological systems, psychology, computation, and information theory, no single definition for life has yet achieved universal acceptance. This challenge becomes increasingly urgent as advances in synthetic biology, artificial intelligence, and astrobiology challenge our traditional conceptions of what it means to be alive. We undertook a methodological approach that leverages large language models (LLMs) to analyze a set of definitions of life provided by a curated set of cross-disciplinary experts. We used a novel pairwise correlation analysis to map the definitions into distinct feature vectors, followed by agglomerative clustering, intra-cluster semantic analysis, and t-SNE (t-distributed Stochastic Neighbor Embedding) projection to reveal underlying conceptual archetypes. This methodology revealed a continuous landscape of the themes relating to the definition of life, suggesting that what has historically been approached as a binary taxonomic problem should be instead conceived as differentiated perspectives within a unified conceptual latent space. We offer a new methodological bridge between reductionist and holistic approaches to fundamental questions in science and philosophy, demonstrating how computational semantic analysis can reveal conceptual patterns across disciplinary boundaries, and opening similar pathways for addressing other contested definitional territories across the sciences.

Keywords Agency · Artificial intelligence · Biology · Computation · Cybernetics · Evolution · Information · Intelligence · Life · Thermodynamics

✉ Michael Levin michael.levin@tufts.edu

Attune Intelligence, LLC, Dallas, TX, USA
Faculty of Dentistry, University of Toronto, Toronto, Canada
Paradigms of Intelligence Team, Google, Seattle, WA, USA
Allen Discovery Center at, Tufts University, 200 Boston Ave, Suite 4600, Medford, MA 02155, USA
Wyss Institute for Biologically Inspired Engineering at, Harvard University, Boston, MA, USA

Published online: 02 July 2026

Introduction

The challenge of defining life extends beyond semantics; it reveals fundamental epistemological differences in how experts across disciplines conceptualize and investigate the boundaries between living and nonliving systems. Biologists often emphasize metabolic processes, reproduction, and cellular organization (Gilbert 1982; 1991; Margulis and Sagan 2000; Boden 2003; Gánti 2003b; Sousa et al. 2008), while physicists may focus on thermodynamic properties, energy utilization, and entropy reduction (Schrödinger 1944; Prigogine 1980; England 2020; Nicholson 2025). Computer scientists and complexity theorists might alternatively prioritize information processing, self-organization, and emergent complexity (Holland 1975; Langton 1997; Wolfram 2002; Seoane and Sole 2018; Banerjee 2021).

These divergent perspectives, shaped by distinct methodological approaches and theoretical frameworks, have historically been difficult to reconcile. This definitional challenge becomes increasingly urgent as advances in synthetic biology (Blackiston et al. 2021), artificial intelligence (Faggin 2024), and astrobiology (Benner 2010) continue to challenge what it means to be alive, which has not only scientific but ethical and regulatory implications.

Some have advocated for definitional pluralism—the acceptance of multiple, contextually appropriate definitions rather than pursuing a single universal concept (Gayon et al. 2010; Mariscal 2023). This pragmatic approach acknowledges the value of specialized definitions within differentiated research contexts. However, it leaves unanswered the question of whether deeper patterns of coherence might exist beneath these apparent disciplinary divisions, and whether a more integrated understanding is possible without sacrificing disciplinary perspectives.

Others have rejected binary categorizations entirely, criticizing attempts to draw arbitrary lines between the living and nonliving. These perspectives reframe the conversation around characterizing life as existing along multiple continuous dimensions rather than as a categorical distinction (Popa 2004; Malaterre and Chartier 2021; Parke 2021; Levin 2022). More radically, some scholars question the utility of defining life at all, suggesting that such attempts are either conceptually futile or entirely misguided (Cleland and Chyba 2010; Machery 2012; Mariscal and Doolittle 2020). To complement the many full-length articles published in the field by researchers focused on this topic, we sought to probe the opinions of other leading scientists in adjacent domains.

We surveyed a set of hand-picked contemporary scholars for their very brief definitions of life and then used both manual and computational semantic analysis to examine the resulting dataset. Our choice of scientists for this survey was not meant to be a statistically-representative consensus or comprehensive overview of the field, but rather an example and investigation of how modern machine learning methods could be used to analyze semantic data on a difficult problem, to establish software and protocols that could later be extended to much larger surveys of diverse backgrounds and communities. Here, we specifically chose a set of modern scholars whose thoughts on the question of “life” had not been previously published, focusing on a sample of those whose interdisciplinary work we deemed most interesting in its dependence (or not) on how one defines life.

Our research addressed the definitional ambiguity through a methodological approach that leverages large language models (LLMs) to systematically analyze 68 expert-provided answers to the question of “what is life?”. We implemented an LLM-derived pairwise correlation analysis to quantitatively map the collective set of responses into an inferred correlation matrix, followed by agglomerative clustering and dimensionality reduction to project those encoded and clustered respondent feature vectors onto a 2-D plane. To further interpret this visualized semantic space, we leveraged LLMs to compare intra-cluster definitions of life and to generate a consensus definition of life that is representative of each cluster.

By systematically mapping how select experts’ perspectives vary across perceptual frames, we aim to provide a multidimensional framework for understanding “aliveness” that transcends traditional categorical distinctions to guide ethical and conceptual decisions in our rapidly evolving technological landscape.

Background

Historical Conceptions of Life

Long before the emergence of modern scientific frameworks, ancient civilizations sought to characterize the animating force that gives rise to life. Rather than providing naturalistic explanations, these ancient cultures relied on supernatural abstractions such as Prana, Qi, Ka, Pneuma, and Ruach in the Vedic, Chinese, Egyptian, Greek, and Hebrew traditions, respectively. These concepts, while diverse in their cultural origins, shared a common attempt to articulate an invisible principle that separated animate from inanimate matter within metaphorical abstractions.

Greek philosophy transformed these symbolic conceptions into the first systematic definitional frameworks of the soul and its relation to life. Plato’s (428–348 BCE) cosmology in the Timaeus characterizes the cosmos itself as a living and intelligent being, with its motions governed by cognition rather than mechanical causation, thereby unifying the soul as both the principle of cognition and the principle of life for the first time (Campbell 2021). For Plato, the differentiated life of the microcosm was a mirror of the expression of life at the scale of the macrocosm.

…the permanent value of the Timaeus rests in its successful presentation of a cosmological basis for a theoretical and practical ethics for mankind. Its theme is life, the generating principle of life—not merely the life of one man or even of humanity, but the genesis τοῦ παντός,1 a fit subject for any philosopher. (Whitaker 1941, p.104)

Aristotle (384–322 BCE) then advanced this framework through his teleological approach in De Anima, further characterizing life by its manifested functional properties. His biological framework presents a nested hierarchy of capacities, in which the vegetative soul (nutrition and reproduction) is present in all living things, the sensitive soul (perception and movement) in animals, and the rational soul is uniquely present in humans (Johnson 2005). His emphasis on intrinsic purpose (telos) established the first

ontologically robust definition of life as a self-actualizing system, where the most natural act of living things is “the production of another like itself, an animal producing an animal, a plant a plant, in order that, as far as nature allows, it may partake in the eternal and divine” (Aristotle 1931, De Anima, II.4, 415a). For Aristotle, the goal towards which all living things strive is to “do whatsoever their nature renders possible” (Aristotle 1931, De Anima, II.4, 415b), establishing living things as fundamentally self-actualizing systems with intrinsic teleology.

The transition from Greek thought to our modern materialistic conceptions of life was mediated by medieval and Renaissance thinkers who preserved teleological frameworks while introducing increasingly material explanations. Galen (129–216 CE)Galen extended Aristotelian biological teleology through empirical anatomy, demonstrating the purposive construction of every anatomical feature (Galen 1968). Renaissance alchemists, particularly Paracelsus (1493–1541)Paracelsus, subsequently developed hybrid frameworks that viewed organisms as chemical laboratories governed by an “archeus” or inner alchemist—a concept that maintained a sense of teleological functionalism while increasingly describing life by its biochemical form (Chang 2011). These “chemical philosophies” represented an important conceptual evolution (Parshall et al. 2015), maintaining the notion that living systems operated toward intrinsic ends while simultaneously beginning to articulate material processes underlying vital functions. This historical progression established a foundational tension between teleological and mechanistic perspectives that continues to challenge contemporary efforts to define life.

During the Enlightenment, René Descartes (1596–1650)René Descartes conceptualized the human as an assemblage of parts that together operated as a machine, dualistically distinct from the mind, which governed its function. This mechanistic approach to the expression of life fundamentally reoriented its conception away from its teleological roots, redirecting focus towards its mechanistic processes.

GRAY

Thus, I say, when you reflect on how these functions follow completely naturally in this machine solely from the disposition of the organs, no more nor less than those of a clock or other automaton from its counterweights and wheels, then it is not necessary to conceive on this account any other vegetative soul, nor sensitive one, nor any other principle of motion and life, than its blood and animal spirits, agitated by the heat of the continually burning fire in the heart, and which is of the same nature as those fires found in inanimate bodies. (Descartes [1664] 1909, p. 202; translated by the authors)

While Descartes described life as a duality of immaterial mind and its mechanistic body, his contemporaries developed thoroughly materialistic philosophies from the seed of his work. Thomas Hobbes (1588–1679)Thomas Hobbes advanced perhaps the most comprehensive materialistic framework of this period, rejecting Cartesian dualism entirely and arguing that all phenomena characteristic of life could be reduced to matter in motion. In conceptualizing life, Hobbes rejected prior teleological or metaphysical conceptions in favor of the view that biological functions are purely mechanical processes.

GRAY

For seeing life is but a motion of Limbs, the beginning whereof is in some principal part within; why may we not say, that all Automata (Engines that move themselves by springs and wheels as doth a watch) have an artificial life? For what is the Heart, but a Spring; and the Nerves, but so many Strings; and the Joints, but so many Wheeles, giving motion to the whole Body, such as was intended by the Artificer? (Hobbes 1909, p.8)

Throughout the Enlightenment, this mechanistic naturalization of life continued to evolve through philosophers such as Gassendi (Stones 1928), Spinoza (Spinoza 2020), and La Mettrie (La Mettrie 1865)Gassendi, Spinoza, and La Mettrie, culminating in a thoroughly materialist framework that would define our modern scientific age. Yet even as the philosophical landscape shifted decisively toward mechanistic explanations, the fundamental question of how to define life persisted. By the dawn of the nineteenth century, scientific inquiry had not resolved but rather reformulated this ancient problem: if life emerges from purely physical processes, what distinguishes living from nonliving matter?

Modern Conceptions of Life

The transition into modern scientific conceptions of life was catalyzed by a series of pivotal 19th-century developments: cell theory identified the fundamental organizational unit of life (Matthias 1838); Charles Darwin’s theory of evolution by natural selection provided the first mechanistic framework for explaining life’s complexity and adaptation (Darwin 1859); the emergence of thermodynamics laid the physical foundation for understanding how living systems maintain order in spite of the second law of thermodynamics (Boltzmann 1886).

These advances set the stage for Erwin Schrödinger’s seminal 1944 lectures, “What is Life?”, which framed organisms as thermodynamic systems that create and maintain order by generating entropy in their environment (Schrödinger 1944). This thermodynamic perspective, while influential, immediately faced challenges: could a purely

energetic characterization capture the qualitative difference between living and nonliving?

The molecular revolution of the mid-twentieth century promised resolution. James Watson and Francis Crick’sJames Watson and Francis Crick’s elucidation of DNA’s double helix structure unveiled the physical basis of heredity (Watson and Crick 1953), while subsequent discoveries in molecular biology revealed the central dogma of genetic information flow (Crick 1958). This mechanistic understanding of heredity spawned information-theoretic definitions, most notably by John von NeumannJohn von Neumann, who characterized life as a computational system capable of self-replication with inheritable variation (Von Neumann 1966).

Contemporary definitional approaches have proliferated across disciplines, diversifying the available perspectives from which one might define life. Thermodynamic frameworks emphasize energy flux and entropy production (Macklem and Seely 2010), while autopoietic theories, such as those developed by Maturana and Varela, focus on self-organization and boundary maintenance (Maturana and Varela 1980; Moreno and Mossio 2015; Moreno and Peretó 2026). The chemoton model by Gánti proposes minimal requirements including metabolism, information storage, and boundary control (Gánti 2003a). NASA’sNASA’s operational definition, adopted from a suggestion by Carl SaganCarl Sagan, represents a pragmatic definition of life for astrobiological research: “a self-sustaining chemical system capable of Darwinian evolution” (Benner 2010, p.1).

Despite the complexification of life’s definitional landscape, tension between perspectives persists. Thermodynamic approaches may capture energy dynamics while underspecifying organizational complexity. Information-theoretic models illuminate genetic coding but struggle with the semantics of biological meaning. Operational definitions encounter borderline cases: viruses, prions, and computational life forms that exhibit some but not all traditional life characteristics.

The emergence of artificial life, synthetic biology, and complex computational systems has further complicated definitional boundaries. Craig Venter’sCraig Venter’s synthetic Mycoplasma cell demonstrated the feasibility of constructing a viable synthetic cell with a compressed genomic code (Hutchison et al. 2016); Michael Levin’sMichael Levin’s lab has shown that embryonic frog cells can self-assemble into novel self-motile living forms with unique behaviors and transcriptomes, when removed from their traditional physiological context (Blackiston et al. 2021; Pai et al. 2025), as can adult human cells (Gumuskaya et al. 2023, 2025); artificial neural networks (Idrees et al. 2024) and evolutionary algorithms (Wang et al. 2025) exhibit adaptation and reproduction without biological substrates. These developments challenge the assumption that life requires carbon-based chemistry, cellular organization, or evolutionary adaptation to exhibit characteristics of life.

Systems biology has attempted synthesis through hierarchical frameworks that integrate multiple organizational levels, from molecular networks to ecosystems (Noble et al. 2019). However, even this integrative approach must contend with the fundamental question: which properties represent necessary versus sufficient conditions for life?

The proliferation of definitional frameworks reflects both scientific progress and persistent conceptual challenges in our understanding of life. As our empirical science deepens, the boundaries of life’s definition paradoxically become less distinct. We conceptualize these ontological tensions as a semantic topology that individual disciplinary approaches have mapped only partially, awaiting the tools of computational systematic analysis to reveal the underlying continuities present within this latent space.

To begin to explore this space and develop methods for parsing the conceptual landscape of scientists’ opinions in difficult, interdisciplinary domains, we undertook an AI-guided analysis of a set of modern scientists’ definitions of “life”. Hand-picked thinkers in a range of disciplines were asked to define “life” (or argue against the possibility of doing so) in three sentences or fewer. Our goal was to map the structure and main drivers of the highly diverse set of responses and to evaluate the utility of LLMs for this task.

A Priori Computational Text Clustering Methods

Our approach builds upon recent advances in AI-augmented clustering, consensus formation, and pairwise constraint modeling, applying these techniques to analyze conceptual patterns in the semantic space of our respondents’ definitions of life.

LLMs have been previously demonstrated to achieve comparable or superior performance to state-of-the-art text embedding and clustering methods by asking the model to generate potential labels for a given dataset and then subsequently categorizing each sample with an appropriate label (Huang and He 2024). More commonly, researchers utilize LLMs to generate text embeddings that encode contextual meaning into dense floating-point vectors, where abstract linguistic patterns can be represented geometrically. Comparative analysis of LLM embeddings for clustering demonstrates how these models consistently capture semantic relationships encoded as geometric proximity in high-dimensional latent spaces (Keraghel et al. 2024).

Recent innovations have further enhanced these capabilities through interpretable k-means clustering and instruction-tuned feedback mechanisms. The k-LLMmeans algorithm utilizes LLMs to generate textual summaries as cluster centroids, capturing semantic nuances often lost when relying

on the purely mathematical properties of k-means clustering (Diaz-Rodriguez 2025)as discussed by Diaz-Rodriguez in twenty twenty-five. The ClusterLLM framework introduces an LLM to improve embedded-text clustering by constructing triplet questions <does A better correspond to B than C>, where A, B, and C are similar data points that belong to different clusters according to the original embedder (Zhang et al. 2023)by Zhang and colleagues. Following this same paradigm, Liusie et al. (2023)Liusie and colleagues have demonstrated that LLMs performed significantly better at comparative assessment tasks (e.g., “Which summary is more coherent, A or B?”) than at absolute assessment tasks (e.g., “Provide a score between 1 and 10 that measures this summary’s coherence.”). These advances demonstrate how LLMs can leverage both embedded semantic relationships and direct linguistic instruction to achieve superior clustering performance, especially when prompted in the form of pairwise comparative analysis. This capability for interpretable analysis through pairwise comparison makes LLMs particularly suitable for mapping the complex conceptual landscape of life’s definitions.

These LLM capabilities directly address the methodological challenges inherent in analyzing diverse definitions of life. Traditional approaches have struggled to systematically compare and integrate definitions across disciplinary boundaries. Our approach treats expert definitions as data points within a quantifiable semantic space, enabling computational analysis of patterns that resist conventional categorical frameworks. By applying these techniques to the 68 expert responses, we can map the conceptual relationships between different definitional approaches while maintaining the nuance and complexity of each perspective. This methodology transforms the question “what is life?” from a philosophical debate into a structured exploration of semantic patterns within expert discourse.

Methods

We developed a software architecture to analyze and cluster diverse definitions of life using large language models (LLMs) and unsupervised learning techniques. Our approach encompasses five key methodological components:

  1. Expert curation and definition collection
  2. Quantitative pairwise correlation analysis between definitions
  3. Agglomerative clustering of the resulting correlation matrix
  4. Thematic analysis of intra- and inter-cluster semantic patterns
  5. t-SNE dimensionality reduction of correlation feature vectors to 2-D space

Figure 1 shows the key computational steps taken to go from the raw pairwise correlation matrix to a sorted and clustered matrix to ultimately a 2-D t-SNE projection of the definitional space.

To validate the robustness of our semantic analysis and mitigate potential model-specific biases, we independently replicated all analytical processes across three state-of-the-art LLMs: Claude 3.7 Sonnet, GPT-4o, and Llama-3.3 70B Instruct. These three matrices were then averaged together to produce the final correlation matrix, accounting for cross-model biases and integrating the nuanced perspective of each model in the final feature set. This methodology was implemented using Python, with all code available as an open-source GitHub repository in the Supplemental Code.

Three panels showing the transformation of data from a raw correlation heatmap (a), to a clustered heatmap with a dendrogram (b), to a 2-D t-SNE scatter plot with density contours (c).Fig. 1 Data transformation process from pairwise correlations to clustered spatial projection. Panel (a) shows the initial pairwise correlation matrix for all 68 definitions, with green indicating high agreement (correlation approaching 1.0) and red indicating disagreement (correlation approaching -1.0). Panel (b) illustrates the sorted correlation matrix, produced after applying hierarchical clustering, with the resulting dendrogram displayed on the right showing the nested relationships between definitions. Panel (c) presents the final t-SNE projection that transforms these high-dimensional correlation patterns into an interpretable 2-D semantic landscape, with colors indicating cluster membership and contour lines revealing density variations across the definitional space

Collection of Expert Definitions and Manual Classification

Here, we were specifically focused on modern researchers’ definitions of life (not classic historical figures’ definitions). Individuals were selected based on their peer-reviewed papers in the literature that impacted (or required an opinion on) the question of life.

Sixty-eight total definitions were included in the final analysis. Fourteen distinct countries were represented by the current affiliations of respondents, with the United States (41) and the United Kingdom (9) being the most prevalent.

The method for expert selection was non-random purposive expert sampling (Stratton 2024)as described by Stratton in 2024 drawn from the principal investigator’s (PI’s) extended professional network and familiarity with their work. The purposive expert sample invited to participate represented a non-random sample of experts with specific, relevant knowledge, which was a requirement for our rich, in-depth, contextually rich analysis. The criteria for selection included the following: (1) the respondents must have possessed tenets of recognized expertise in their field, (2) must have taken on leadership roles, (3) must have a record of substantive academic or professional contributions, and (4) must have relevance of their field of study to the question of the definition of “life” but not be circumscribed by it (i.e., interdisciplinary work). These criteria for expert selection aimed to ensure a high level and sufficient depth of knowledge of the topic of life that we set out to explore, including both theorists and laboratory scientists. There was no snowball recruitment, meaning the respondents did not invite or recommend participation to new individuals in their extended network, to avoid the high degree of bias associated with this method.

Respondents spanned several disciplines, including biology, philosophy, computer science, physics, AI, and systems theory. There was a strong representation of scientists working in theoretical and interdisciplinary biology. The majority (around 50 of the 68) were primarily trained as academic researchers, around 14 hold hybrid roles (e.g., a blend of any of the following: academic researcher, clinician, philosopher, computer scientist, or organizational leader), and one respondent was primarily involved in science education and public scholarship. No respondents were recruited solely as clinicians, consultants/policy experts, legal scholars, ethicists, humanities-based researchers (philosophers, sociologists, musicians/fine arts, etc.), or industry representatives. In addition, we did not contact any researchers from the fields of death, dying, and grief. We recognize the lack of inclusion of experts from these disciplines as a limitation and an area for future work, because experts from these diverse fields would have shaped the analysis and contributed to the discussion of life in a relevant and useful way.

As the selection was guided by the PI’s professional network, we acknowledge that this may privilege certain perspectives and underrepresent others. Underrepresentation of voices and scholars from other communities, geographies, and schools of thought beyond our sample is acknowledged. A purposive sample may inevitably, though unintentionally, leave behind a vast number of voices whom we specifically invite in future work to enrich the discussion. We emphasize that the computational method we set out to present here can theoretically accommodate an indefinite number of definitions, and the open-source code appended is welcome to be applied to any alternate set of well-balanced, multidisciplinary definitions from experts besides the ones selected in our purposive sample here. More information on the limitations of expert selection as it pertained to the computational analysis can be found in the section below (section “Discussion—Methodological Innovations and Limitations”).

Following selection, the expert authors were contacted through the interview method of email-based correspondence to the interviewee’s institutional email, with balanced representation from the fields and disciplines, including branches of biology, computer science, physics, and engineering.

The respondents were not aware of each other’s responses or thought processes, were not given any specific guidance or criteria for inclusion/exclusion, and were independently responsible for submitting their own definition. Respondents did not have any set time limit to formulate and structure their responses and were permitted to consult external references, and no policy prohibiting contact with colleagues on the matter was set forth. Respondents were instructed to aim for no more than three sentences and told that they should think about life as broadly as they wished, not restricted, for example, to terrestrial life. They were not given any guidelines about word count, and they were made aware that their responses were going to be analyzed in a manuscript for publication.

Two of the provided definitions came from the authors of this analysis. Like all the other respondents, the two authors who provided definitions did not see any of the other life definitions submitted prior to formulating their own. No additional weight was given to these two responses in the downstream analysis.

For editorial purposes, a small proportion of responses were shortened without interfering with the key takeaways of the authors’ definitions. This was done to maintain standardized length for all definitions. Finally, experts were given a chance to proofread their definitions for correct attribution but were strongly discouraged from making any changes at this point, as they had already seen some of the other definitions.

Once all definitions had been received, a manual review and categorization were conducted to ground the results in human understanding before proceeding to the AI-driven analysis. It was possible for an author’s definition to be grouped into more than one category. The row “Other” was introduced when a definition did not neatly fit into one of the preexisting categories. The resulting categorical distinctions are presented in Table 1. All 68 provided definitions of life are presented in Supplemental Table 1.

The overlapping nature of these categories, with many definitions spanning multiple classes and requiring an “Other” category, demonstrates precisely why a continuous semantic space analysis is necessary. Rather than treating these overlaps as classification failures, we view them as evidence that expert thinking about life exists along multiple non-orthogonal dimensions that make discrete categorization challenging. This observation motivates our computational approach, which maps respondent relationships in high-dimensional space without forcing artificial boundaries.

Quantitative Pairwise Correlation Analysis

To quantify semantic relationships between definitions, we developed a pairwise correlation analysis framework using LLM-based inference. For each pair of definitions in our dataset, we computed a correlation score ranging from -1.0 (complete disagreement) to 1.0 (complete agreement) using the following prompt structure:

Analyze the following two definitions of life, where:
-1.0 = Fundamentally opposing or incompatible primary frameworks;
-0.5 to -0.9 = Significantly different emphasis with some contradiction;
0.0 = Independent or orthogonal frameworks;
0.1 to 0.4 = Slight overlap in secondary elements;
0.5 to 0.9 = Significant overlap with some differences;
1.0 = Aligned core frameworks and secondary elements.
 
Definition 1: {{definition_1}}
Definition 2: {{definition_2}}
 
What is the correlation metric between -1.0 and 1.0 for these two definitions?
Respond with ONLY a single number!

To enhance reliability and account for potential variability in LLM outputs, we performed multiple replicate inferences (n equals three) for each definition pair and computed the average correlation score and standard deviation. This correlation metric was computed for each pair of responses (n squared), resulting in a correlation matrix of respondent definitions encoding pairwise similarity (r between zero and one) and dissimilarity (r between minus one and zero).

It is critical to note here that while we inferred correlation across respondents by generating similarity scores in a pairwise manner with redundancy across three unique LLMs, the resulting matrices do not represent correlation as is normally defined by robust statistical analyses such as Pearson or Spearman. Rather, we use the term correlation matrices to define the similarity as ranked by redundant inferred pairwise comparisons.

The pairwise correlation analysis generated two n by n matrices, where n is the number of definitions: (1) a correlation matrix M containing the average correlation scores between all definition pairs, and (2) a standard deviation matrix S capturing the variability of correlation estimates.

Table 1 Categorizing definitions of life

Criteria: Life is defined by specific:Researchers whose definitions involve those criteria
Thermodynamics/energeticsBall, Davies, Dussutour, Froese, Georgiev, Gilbert, Huang, Ingber, Lane, Mitchell, Newman, Picard, Solms, Soto, Szathmáry, Tuszynski, Vallverdú, Nunn
Information/patternAckley, Adami, Bongard, Caves, Das, Dodig-Crnkovic, Georgiev, Ingber, Levin, Miller, Mitchell, Gentili, Nunn, Pavlic, Reber, Solé, Witkowski
ComplexityGeorgiev, Jackson, Sloman, Nunn
ComputationAgüera y Arcas, Stepney, Ackley, Witkowski, Dodig-Crnkovic
Dynamics (including self-organization)Agüera y Arcas, Bongard, Brash, Caves, Das, Dodig-Crnkovic, Dussutour, Friston, Froese, Georgiev, Gilbert, Heylighen, Huang, Kauffman, Mitchell, Pizzi, Rechavi, Solé, Szathmáry, Stepney, Witkowski
Autonomy/agency/goal-directednessBall, Dussutour, Ellis, Jablonka, Calvo, Heylighen, Krakauer, Lander, Lyon, Mitchell, Noble, Shapiro, Solms, Soto, Tuszynski, Vallverdú, Caves
Cognition/intelligenceDussutour, Fontana, Gentili, Krakauer, Levin, Marshall, Miller, Nunn, Picard, Reber, Lyon, Dodig-Crnkovic
Structure/architectureBaluška, Georgiev, Newman, Caves
Functionality/behaviorAlbantakis, Ball, Caves, Dussutour, Das, Lyon, Sonnenschein, Sultan
Replication/reproduction/ evolution/heredityAckley, Adamatzky, Adami, Davies, Ingber, Jablonka, Jackson, Lander, Gentili, Nunn, Ratcliff, Shapiro, Sloman, Szathmáry
Material compositionBaluška, Das, Kauffman, Lane, Newman, Pizzi, Ratcliff, Sloman
Against strong definitionAlbantakis, Ball, Frank, Gunawardena, McShea, Stanley, Wong, McShea
OtherCiaunica, Fields, Hoffman, Marshall, Witkowski, Watson

Both matrices were symmetrized by averaging with their transpose M prime equals M plus M transpose, over two to account for potential rank-ordering effects in the LLM-derived correlation inference, resulting in a total of 6 instances of inference averaged to generate each pairwise correlation metric. The resulting correlation matrices were visualized as heatmaps using a divergent color scheme with correlation values ranging from -1 (disagreement, red) to 1 (agreement, green).

This analysis was repeated with Claude 3.7 Sonnet, Llama 3.3 70B, and GPT-4o, with the final correlation matrix being derived from the average of each analysis’s result. Each provided similarity metric (henceforth referred to as correlation) is then the result of 18 averaged LLM-generated correlation metrics, six from each tested model.

This LLM-based correlation methodology raises questions about whether similarity scores reflect genuine conceptual alignment or surface linguistic features such as shared vocabulary, stylistic patterns, or common references from the models’ training data. Several factors suggest these scores capture deeper semantic structure:

  1. A high cross-model correlation (r greater than zero point seven) despite different training data and architectures indicates convergence on stable patterns beyond individual model biases.
  2. Definitions with high correlation scores often employ dissimilar vocabulary while definitions with low correlation sometimes share terminology, suggesting the models detect conceptual frameworks rather than mere word overlap, which is a persistent challenge with traditional text-embedding methodologies.
  3. The interpretability of resulting clusters in terms of established philosophical frameworks provides external validation.

Nevertheless, we acknowledge that these scores inherently blend conceptual similarity with model-specific training effects, a methodological limitation that parallels how scientific communities construct shared meaning through both individual reasoning and collective zeitgeist.

Agglomerative Clustering of Correlation Matrix

To identify natural groupings within the definition space, we applied hierarchical agglomerative clustering to the symmetrized correlation matrix. This bottom-up clustering approach begins with each definition as its own cluster and iteratively merges the most similar clusters until all definitions belong to a single cluster, creating a hierarchical structure that reveals multi-scale relationships between definitions.

The correlation matrix was first transformed into a distance matrix using the relationship d equals the square root of, two times the quantity one minus r, where r represents the correlation coefficient. This transformation maps correlation coefficients between -1 and 1 to distances between 0 and 2, with higher correlations corresponding to smaller distances between definitions, with the squaring operation amplifying the distances between high correlations and therefore making the clustering algorithm more sensitive to subtle semantic distinctions between closely related definitions.

We employed complete linkage (maximum distance between any two elements from different clusters, also known as farthest-neighbor linkage) as our agglomerative method. Complete linkage was selected for its ability to produce compact, clearly separated clusters that emphasize conceptual boundaries between semantic groupings, making it particularly suitable for analyzing definitional relationships where clear delineations are desirable.

The dendrogram construction process mimics how one might organize ideas into increasingly broad conceptual categories. At the beginning, each definition stands alone as a unique perspective; each leaf is embedded within a tree with one node. The algorithm then identifies the two most similar definitions and joins them at a height proportional to their similarity. Closely aligned definitions join near the bottom of the dendrogram, while more distant conceptual relatives join higher up. This process continues iteratively, with either individual definitions joining existing groups or groups merging with other groups, always connecting the most similar entities at each step. The resulting tree structure then captures which definitions belong together as related conceptual definitions, as well as the relative semantic distance between them.

To then perform unsupervised clustering on this dendrogram, we analyzed the pattern of merge distances to identify the inflection points where further cluster merges would suddenly bring together substantively different conceptual frameworks. This “elbow” in the derived linkage matrix reveals the natural number of clusters present in the data as the point where merging clusters would begin to obscure important conceptual distinctions. This clustering solution partitions definitions into groups that share fundamental conceptual frameworks while maintaining meaningful separation between distinct philosophical or scientific approaches to defining life. When superimposed on top of the sorted correlation matrix of respondent definitions, this multi-scale perspective allows us to examine both fine-grained distinctions within conceptual frameworks and broader patterns across the entire definitional space.

LLM Cluster Semantic Analysis

For each cluster identified through hierarchical agglomerative clustering, we employed LLMs to perform two

complementary analytical processes: (1) intra-cluster thematic analysis to characterize conceptual frameworks and (2) consensus definition generation to distill essential shared elements. Claude 3.7 Sonnet was used to generate the semantic analysis for the multi-model clustered correlation matrix.

The intra-cluster thematic analysis was conducted using a structured prompt designed to systematically extract patterns from definition clusters. This protocol instructed the LLM to analyze definitions through three progressive lenses:

1. WHAT Are The Core Ideas?
 - List every key concept mentioned
 - Count how often each appears
 - Group similar concepts together
 - Note which concepts appear most
 - Mark which are always present
2. HOW Do Ideas Connect?
 - Find concepts that link to others
 - Map which ideas depend on others
 - Note concept hierarchies
 - Identify central hub concepts
 - Track idea flow patterns
3. WHY This Structure?
 - Identify shared a priori frameworks
 - Find similar starting points
 - Track reasoning patterns
 - Mark scope boundaries
 - Note definitional strategies

All prompts are available in their complete form in the Supplemental Code. This thematic analysis was then used to inform our meta-review of various definitions for “what is life?” from a quantitative, statistically grounded, computational point of view. Complementing the thematic analysis, we further employed LLMs to synthesize a single representative definition for each cluster. This protocol explicitly constrained the LLM to adhere to the requirements defined in the following prompt structure:

You are tasked with synthesizing a consensus definition of life from a group of experts 
who have similar perspectives. Below are their definitions:
{{definitions}}
 
In addition to these provided cluster definitions, you have also conducted a prior
thematic review which contains a meta-analysis of the key concepts being discussed 
by this group of respondents:
{{cluster_analysis}}
 
- Start: "Life is..."
- Content: Only majority-shared concepts (>50% frequency)
- Language: Technical terms from source definitions
- Structure: Logical flow of connected concepts
- Style: Match source complexity and tone
- Length: Within ±20% of median definition length
- Exclude: Unique views, novel terms, explanations

This algorithmic approach to consensus formation ensured that the resulting definitions genuinely represented the central tendencies within each cluster rather than introducing new concepts or arbitrary interpretations. By limiting content to majority-shared concepts and using only terminology present in the source definitions, we maintained fidelity to the original expert perspectives while distilling their shared conceptual core.

The resulting thematic analyses and consensus definitions provide complementary perspectives on each cluster: the thematic analysis characterizes the conceptual landscape and intellectual approach shared by definitions within a cluster, while the consensus definition distills these shared elements into a single coherent statement that represents the cluster’s central perspective on what constitutes life. Together, these analyses reveal both the distinctive conceptual frameworks that characterize different approaches to defining life and the core elements that unify definitions within each framework.

Finally, a third inference call was executed to generate a title name for each cluster. Provided with the consensus definition and the thematic analysis in the prompt context, the LLM was asked to generate a brief cluster title according to the following prompt:

Based on the consensus definition and thematic analysis provided, generate a SINGLE 
WORD
or VERY SHORT PHRASE (2-4 words maximum) that captures the DISTINCTIVE ESSENCE 
of this specific cluster.
 
The title should be:
1. HIGHLY SPECIFIC to the philosophical, scientific, or conceptual framework unique to
this cluster
2. DISTINCTIVE enough that it wouldn't apply equally well to other clusters of life 
definitions
3. TECHNICALLY PRECISE, using domain-specific terminology where appropriate
4. CONCEPTUALLY FOCUSED on the core unifying principle of these definitions
 
---
Provided below is the analysis of the given cluster:
{{cluster_analysis}}
 
Provided below is the derived consensus definition for the given cluster:
{{consensus_definition}}

These titles were then mapped to the cluster groups derived by agglomerative clustering and used to define the clusters subsequently plotted by the described 2-D projection methodology.

Two-Dimensional Projection of Definitional Landscape

To create an interpretable visualization of the definitional landscape, we applied t-SNE (t-distributed stochastic neighbor embedding) dimensionality reduction to project the high-dimensional correlation features into a 2-D representational space. Unlike principal component analysis (PCA), t-SNE specifically emphasizes the preservation of local neighborhood relationships, making it particularly suitable for visualizing clusters of semantically related definitions. Critically, t-SNE is a visualization tool that preserves local

neighborhood structure rather than a principled dimensional decomposition like PCA, and its global axis orientations are mathematically arbitrary. However, when interpretable gradients emerge consistently across multiple runs with different initializations and across different LLM analyses, the consistency suggests these dimensions capture genuine structure in expert discourse rather than computational artifacts. We interpret these emergent axes philosophically while acknowledging they are algorithmic constructs that reveal, rather than impose, conceptual organization. This was validated by testing against multidimensional scaling (MDS), principal component analysis (PCA), and uniform manifold approximation and projection (UMAP), with t-SNE consistently producing the best spatial projection of the underlying correlation embeddings.

The correlation matrix was first transformed into a distance matrix using the relationship d equals the square root of, two times the quantity one minus r, where r represents the correlation coefficient between definition pairs—the same distance transformation applied to agglomerative clustering. The resulting distance matrix was then symmetrized by averaging with its transpose M prime equals M plus M transpose, all divided by two to account for potential asymmetries in the process of translating correlation features to distance features, effectively placing definitionally similar responses (high correlation) in proximity while separating dissimilar ones (low correlation).

The final t-SNE projection visualizes the semantic landscape of life definitions, with each point representing an expert’s response colored according to its assigned cluster. Contour lines were added to highlight density variations within the definitional space, revealing regions of conceptual convergence and divergence. The final visualization incorporates cluster assignments derived from agglomerative hierarchical clustering, with each cluster labeled according to its LLM-derived title that characterizes the unifying conceptual framework of that group. This approach reveals the distinct conceptual territories in relation to the broader definitional landscape, providing a comprehensive spatial reflection of how experts conceptualize life across disciplinary boundaries.

Results

Expert Definitions

Supplemental Table 1 shows the definitions provided by the survey participants. Some focused on the impossibility (or futility) of a definition, while others attempted this task from varying conceptual alignments. The 68 definitions collected represent a broad spectrum of disciplinary perspectives, from traditional biological frameworks to computational approaches, from physics-inspired thermodynamic views to cognitive and philosophical stances. Despite the likely impossibility of defining truly orthogonal categories along which to categorize or classify these definitions, we attempted to do so manually (without AI-assistance) using the criteria in Table 1.

Manual categorization of these definitions revealed several predominant themes, with certain concepts appearing repeatedly across multiple definitional frameworks. Thermodynamics and energetics appeared in 18 out of 68 (26%) of all definitions, while information/pattern-based approaches were found in 25% of definitions. Dynamics (including self-organization) was the most prevalent theme, appearing in 29% of definitions.

We further characterized the most prominent divergent properties commonly discussed (Fig. 2), revealing that 57% of definitions took an objective stance (defining life in observer-independent terms), while 43% incorporated observer-relative elements. Additionally, 54% of definitions framed life as a continuous property rather than a binary state, reflecting a shift away from categorical distinctions between living and nonliving systems. Most definitions (88%) were actionable in the sense that they provided concrete, conventionally defined criteria, while 12% took more inspirational, poetic, or self-recursive approaches.

Pairwise Correlation Analysis

To assess the semantic relationships between expert definitions of life, we analyzed how three state-of-the-art LLMs perceived conceptual overlap among the 68 definitions. Each model independently evaluated all possible definition pairs, producing a correlation matrix that quantifies similarity (-1.0 for complete opposition to 1.0 for perfect alignment) between each definition pair.

Claude 3.7 Sonnet, Llama 3.3 70B Instruct, and GPT-4o were independently used to generate three distinct correlation matrices. Claude 3.7 Sonnet demonstrated the most optimistic assessment with the highest average correlation, while Llama 3.3 70B Instruct produced more conservative estimates of correlation and showed greater standard deviation in its responses. Notably, all models displayed negative skewness (-0.76 to -0.97), indicating a systematic tendency to identify more areas of conceptual similarity than opposition. The general correlation matrix statistics for each model are presented in Supplemental Table 2.

Despite the models’ independent analyses, they showed remarkable consistency in their overall assessment of which definitions aligned conceptually and which diverged. The pairwise matrix-to-matrix correlations were notably high:

  • Claude 3.7 Sonnet vs. Llama 3.3 70B Instruct: r equals zero point seven two seven nine

A stacked horizontal bar chart showing three categories of properties for expert definitions of life: Perspective (57% Objective, 43% Observer-relative), Nature (46% Binary, 54% Continuous property), and Approach (87.7% Actionable, 12.3% Inspirational).Fig. 2 Distribution of general properties among expert definitions of life (n equals sixty-eight)

  • Claude 3.7 Sonnet vs. GPT-4o: r equals zero point eight one zero three
  • Llama 3.3 70B Instruct vs. GPT-4o: r equals zero point seven nine seven seven

These substantial correlations suggest that while each model brought a distinct analytical perspective to the definitional landscape, they independently converged on similar patterns of conceptual relationships.

These correlation matrices are shared in Supplemental Figs. 1, 2, and 3 for Claude, Llama, and GPT, respectively. The resulting unsorted multi-model averaged correlation matrix is presented as Supplemental Fig. 4. Additionally, the standard deviations (entropy matrix) for all aggregated instances of inference are available at Supplemental Figs. 5, 6, and 7 for Claude, Llama, and GPT, respectively, highlighting the consistency with which the LLMs scored correlation for each respondent across repeated instances of inference.

To mitigate potential biases from any single model and create a more robust assessment of semantic relationships, we computed an element-wise average across the three symmetrized correlation matrices. The multi-model integration succeeded in capturing the central tendencies of all three individual analyses, as evidenced by the high correlations between each model’s matrix and the averaged matrix:

  • Claude 3.7 Sonnet vs. Multi-Model-Average: r equals zero point nine zero one seven
  • Llama 3.3 70B Instruct vs. Multi-Model-Average: r equals zero point nine two nine four
  • GPT-4o vs. Multi-Model-Average: r equals zero point nine three five eight

The multi-model average matrix successfully captured the nuanced perspectives of each model, ultimately producing a rich encoding of pairwise correlation patterns amongst the provided definitions of life.

Agglomerative Clustering of Correlation Matrices

The three LLMs tested exhibited distinct clustering behaviors, identifying between seven and eleven natural clusters through elbow method analysis. Claude 3.7 Sonnet produced the most integrative solution with seven clusters and the lowest contrast metric (0.258), while Llama 3.3 70B Instruct demonstrated the most discriminative approach with eleven clusters and the highest contrast metric (0.563). The multi-model integration yielded eight clusters with a balanced contrast metric of 0.359, positioning it between the integrative and discriminative extremes. This integrated approach achieved robust intra-cluster correlation (zero point five zero nine plus or minus zero point one eight seven) while maintaining modest inter-cluster correlation (zero point one five zero plus or minus zero point three nine zero), suggesting it captured both the distinctiveness of conceptual frameworks and meaningful connections between them. The statistics of each tested model’s clustering results are available in Supplemental Table 3.

Figure 3 presents the visualization of the multi-model integrated analysis, showing both the clustered correlation matrix (with green indicating high correlation and red indicating negative correlation) and the corresponding dendrogram structure on the right. Additionally, the clustered correlation matrices for Claude, Llama, and GPT’s independent analyses are presented in Supplemental Figs. 8, 9, and 10, respectively, to show how each model uniquely organized the respondents in relation to each other.

This sorted correlation matrix reveals a distinct block-diagonal structure, indicating strong intra-cluster coherence, while also exhibiting meaningful cross-cluster correlations that suggest conceptual continuity across the definitional landscape. To investigate how consistently the various LLMs grouped the respondents into the same cluster relationships, we compared the clustering consistency from

A heat map titled "Clustered Definition Correlations" showing a matrix of correlations between different respondent definitions. A dendrogram on the right illustrates the hierarchical clustering of these definitions into eight distinct, color-coded groups.Fig. 3 Hierarchically sorted correlation matrix of respondent definitions, along with the corresponding dendrogram defining the linkage distance between definitions, highlighting eight distinct conceptual clusters

each of the three models’ analyses, plus the multi-model integration.

  • Claude 3.7 Sonnet and Llama 3.3 70B Instruct: 71.2% consistency
  • Claude 3.7 Sonnet and GPT-4o: 74.8% consistency
  • Claude 3.7 Sonnet and Multi-Model-Average: 74.1% consistency
  • Llama 3.3 70B Instruct and GPT-4o: 75.7% consistency
  • Llama 3.3 70B Instruct and Multi-Model-Average: 72.8% consistency
  • GPT-4o and Multi-Model-Average: 79.1% consistency

These high cross-model consistencies (71.2% to 79.1%) indicate substantial agreement on which definitions should be grouped together, despite the different number of clusters identified by each model.

To further assess the robustness of our clustering, we examined how consistently individual definitions were grouped across the three models. The overall grouping stability was 53.7%, indicating that just over half of all definition pairs were consistently grouped together (or consistently separated) across all models. This moderate stability suggests that while the broad structure of the definitional landscape is robust, there remains genuine ambiguity in how some definitions should be categorized at the boundaries between clusters.

Some definitions showed remarkably high grouping consistency across models: those by Christopher Fields, Donald Hoffman (both 100% consistency), Andy Adamatzky, Leo Caves, and Richard Watson (all 95.5% consistency). These definitions appear to occupy clear, distinctive positions in the semantic landscape. In contrast, definitions by Mark Solms (28.4%), Aaron Sloman (26.9%), Paul Davies (23.9%), Tom Froese (23.9%), and Wesley Wong (22.4%) showed the lowest consistency, suggesting these definitions span multiple conceptual frameworks or occupy boundary positions between established clusters. Looking at Fig. 3, we can observe that stable definitions tend to appear within

the most distinctly colored blocks, while unstable ones often appear in transition zones between clusters.

Intra-Cluster Semantic Analysis

From the eight distinct clusters derived by agglomerative clustering of the averaged multi-model pairwise correlation matrix, the thematic content of each cluster was further explored. The consensus definitions in Table 2 represent distillations of the core concepts shared within each cluster, providing insight into how different perspectives approach the question of what constitutes life. These definitions were derived through LLM analysis of the definitions within each cluster, extracting the majority-shared concepts while maintaining the technical language and conceptual frameworks characteristic of each group.

A two-dimensional t-SNE plot showing a "Definitional Landscape" of life. Data points representing different authors are scattered across the plot and grouped into eight color-coded clusters with shaded density contours.Fig. 4 Two-dimensional t-SNE definitional landscape showing the distribution of expert perspectives across eight conceptual clusters, derived from the average pairwise correlation matrices from independent Claude 3.7 Sonnet, Llama 3.3 70B Instruct, and GPT-4o analyses. (Full-color figure available in electronic version.)

The complete thematic analysis performed for Claude, Llama, GPT, and the resulting multi-model average are shared in Supplemental Files 1, 2, 3, and 4, respectively.

t-SNE Projection of Life’s Definitions

t-SNE dimensionality reduction transformed the high-dimensional correlation data into an interpretable 2-D visualization of the 68 expert definitions (Fig. 4). t-SNE was chosen for its ability to preserve local neighborhood structure while revealing global patterns in high-dimensional data, but the complete panel of clustering results from multidimensional scaling (MDS), uniform manifold approximation

Table 2 Cluster titles and consensus definitions

TitleConsensus definition
Perceptual CategorizationLife is a perceptual category arising from how systems that consider themselves alive recognize and categorize others, rather than an objective distinction with intrinsic properties. The apparent boundary between living and nonliving exists primarily as an artifact of our sensory limitations that introduce artificial distinctions into what may be a more unified reality.
Self-Sustaining Dynamic PatternsLife is a persistent dynamic pattern that sustains itself through recursive processes, creating and maintaining its own conditions for existence. It manifests as a self-aware system capable of transformation, emerging through the resonance and harmonic interaction of its components. This pattern exists in relationship with its environment, expressing consciousness while continuously recreating itself through mutually transformative connections.
Dynamic Relational ProcessLife is a permanent movement characterized by dynamic exchange between internal and external environments, fundamentally requiring relationship with preexisting others. Life cannot exist in isolation or stasis, as its essential nature involves both movement and proliferation. Rather than being something humans can create, life is received and passed on, existing as an ongoing process of interconnection.
Pragmatic Definitional SkepticismLife is a contextual construct better approached through functional utility than universal definition, where definitional efforts should serve specific research purposes rather than establish absolute boundaries. The concept exists at the intersection of scientific and philosophical domains, requiring recognition of disciplinary limitations and pragmatic focus on what advances understanding rather than what constitutes a “correct” characterization.
Cognitive AutonomyLife is a self-maintaining, goal-directed system that processes information to adapt to environmental changes while preserving its organizational boundaries. It operates as an autonomous agent that actively opposes entropy through energy exchange, making purposeful decisions that ensure its continued existence. Life functions as a cognitive process that senses, interprets, and responds to its surroundings, modifying itself when necessary to maintain stability despite changing conditions.
Dissipative Self-Organizing SystemsLife is a dynamic, far-from-equilibrium process characterized by self-organization and self-maintenance through controlled energy dissipation. It maintains semi-permeable boundaries that separate the system from its environment while allowing selective exchange of matter and energy. Living systems process, store, and transmit information through feedback mechanisms that enable adaptation and evolution over time.
Informational Self-ReplicationLife is a self-replicating system of information encoded in a physical substrate, capable of reproducing itself while maintaining order against natural decay. This information-based process enables the transmission of complex structural and functional data across generations, with the physical components serving primarily as carriers for this essential informational content.
Self-Replicating Thermodynamic SystemsLife is a self-sustaining physical and chemical system that reproduces itself while maintaining organization away from thermodynamic equilibrium. It interacts adaptively with its environment, absorbing and transforming resources to support its replication and growth. Living systems exhibit autonomy through self-regulation and response to stimuli, while their reproductive capabilities enable evolutionary processes.

and projection (UMAP), t-SNE, and an ensemble average is presented in Supplemental Fig. 11.

The resulting t-SNE projection of Fig. 4 reveals a continuous semantic landscape organized along two conceptually interpretable axes. These dimensional interpretations should be approached with nuance. While t-SNE’s global axis orientations are mathematically arbitrary, the consistency with which these particular philosophical gradients emerge (across different LLM analyses and when validated against independently coded definition features in Table 1) suggests they represent genuine structure in how experts conceptualize life rather than artifacts of the t-SNE algorithm. The axes are not imposed by t-SNE but rather revealed through it, analogous to how t-SNE in single-cell transcriptomics reveals biological cell-type relationships that can then be validated through independent molecular markers (Kobak and Berens 2019)Kobak and Berens, 2019.

Dimension 1 (horizontal) maps the transition from observer-dependent, perceptual frameworks (left) to objective, material-structural definitions (right). The leftmost position is occupied by the isolated Perceptual Categorization cluster, exemplified by Fields’ recursive formulation: “to be alive is to be considered alive by systems that consider themselves alive”. This contrasts sharply with the right-side clusters that emphasize concrete mechanisms like replication and thermodynamic properties, captured by Adami’s conception of life as “information that can replicate itself”.

Dimension 2 (vertical) reveals a fundamental ontological tension between entity-based and process-based conceptualizations of life, recapitulating the historical dialectic from Cartesian mechanism to Aristotelian teleology. The topmost positions are occupied by definitions emphasizing material structures, component properties, and tangible mechanisms, such as Baluška’s “Living Cells and All Their Constructs”, Pizzi’s focus on “carbon-based elements with self-organizing and self-replicating properties”, and Sloman’s “components that are able to extend and replicate themselves”. The bottom region contains process-oriented, cognitive-phenomenological frameworks, represented by Marshall’s metaphorical “fire that lights itself”, Picard’s “self-sustaining system capable of healing that uses and organizes energy to achieve expression(s) of consciousness”, and Noble’s “self-creating agency”. This vertical gradient captures the perennial tension between substantialist perspectives that

view life as a collection of specified entities (echoing Cartesian and Hobbesian frameworks) and teleological views that conceptualize life as coherent patterns of organization, information flow, and emergent self-actualization.

The dominant central attractor comprises two partially overlapping clusters: Cognitive Autonomy (teal, n equals 24) and Dissipative Self-Organizing Systems (blue, n equals 21). These concentrate 66% of definitions in a high-density zone where physical, informational, and agential perspectives are integrated by a variety of definitional approaches.

This concentration raises a critical interpretative question: does this central attractor represent genuine conceptual convergence toward an integrated understanding of life, or does it reflect how institutional science and shared disciplinary training constrain the space of professionally expressible positions? The high density may indicate either emerging consensus or the boundaries of what our surveyed respondents find conceptually acceptable, confounded by selection bias amongst the individuals asked to provide their definition. The peripheral clusters, though sparsely populated, may therefore represent not outliers but vanguard positions which challenge the implicit assumptions structuring the central region; these peripheral definitions reflect positions that question fundamental assumptions such as whether life requires objective material properties (Perceptual Categorization) or whether defining life serves scientific progress at all (Pragmatic Definitional Skepticism).

Contour patterns expose varying semantic densities across the landscape. The sparsest regions surround the isolated Perceptual Categorization cluster, while gradient transitions appear between process-oriented clusters (orange and yellow) and the central attractors. The right quadrant displays two distinct poles: Informational Self-Replication (purple, n equals 3) achieves maximum structural distance from perceptual clusters, while Self-Replicating Thermodynamic Systems (pink, n equals 8) bridges traditional biological perspectives with information-theoretic approaches.

The continuous nature of this semantic space is further illuminated through definitions that function as conceptual bridges between cluster territories. Watson’s characterization of life as “the pattern and process of love—a deeply vulnerable mutual dance” creates a semantic pathway between perceptual-relational and teleological frameworks. Similarly, Noble’s “Life is self-creating agency” occupies a transitional position between process-oriented approaches and cognitive frameworks, explicitly connecting Aristotelian self-actualization with modern agency-based perspectives. Levin’s “Living beings remember, and anticipate, coarse-graining experience into a continuous actionable dream” bridges cognitive and physical perspectives through concepts that unite information processing with thermodynamic constraints, demonstrating how contemporary definitions continue to negotiate the historical tension between mechanism and teleology.

The emergence of interpretable dimensions from computational semantic analysis validates this approach as a methodological bridge between reductionist and holistic perspectives on fundamental questions in science and philosophy. Rather than imposing preconceived categories, the methodology reveals latent structure within expert discourse, demonstrating how computational tools can expose patterns of conceptual coherence that might otherwise remain obscured by disciplinary fragmentation. The high convergence across different LLMs (correlation coefficients greater than zero point seven) further validates the robustness of these patterns, while the consistent identification of conceptual bridges and transitional zones across models suggests these represent genuine features of the definitional landscape rather than analytical artifacts.

Cluster Semantic Analysis

The eight conceptual clusters revealed through agglomerative clustering present distinct yet interconnected perspectives within the definitional landscape, each occupying a unique semantic position that illuminates fundamental tensions and convergences in contemporary understandings of life.

The Perceptual Categorization cluster (n equals 2) occupies the most divergent semantic space, achieving maximum conceptual distance from material-structural frameworks. Fields’ recursive formulation (“to be alive is to be considered alive by systems that consider themselves alive”) challenges foundational assumptions about objective reality. Hoffman’s positioning of the living/nonliving boundary as “an artifact of the limitations of our senses” resonates with quantum–mechanical insights about observer effects, suggesting unexplored connections between life’s definition and fundamental physics.

Self-Sustaining Dynamic Patterns (n equals 4) bridges radical epistemology with process philosophy through sustained focus on relationship and consciousness. Watson’s characterization of life as “pattern and process of love—a deeply vulnerable mutual dance” introduces affective dimensions absent from mechanistic frameworks, while Marshall’s “fire that lights itself” and Caves’ recursive formulation both emphasize the importance of a paradoxical self-reference in understanding the nature of life. This cluster’s correlation patterns reveal conceptual affinity with both observer-dependent frameworks and certain elements of Cognitive Autonomy, suggesting potential integration pathways between phenomenological and scientific approaches.

Dynamic Relational Process (n equals 2) distills relationality to its kinetic essence, with Sonnenschein’s declaration

that “Without proliferation and movement, there is no life”, establishing clear boundary conditions. Similarly, Ciaunica’s insistence that life “cannot exist without an other, already being there, already alive” poses significant challenges to origin-of-life theories that assume primordial isolation. This minimal cluster creates bridges between abstract relationality and classical biological requirements, positioning itself as a phase-shift between philosophical and empirical approaches. Its adjacent positioning to Self-Sustaining Dynamic Patterns reflects deep conceptual affinity, while its moderate positive correlations with replication-focused clusters suggest this bridge function operates effectively across multiple scales.

Pragmatic Definitional Skepticism (n equals four) functions as a methodological counterpoint, with Lane’s assertion that “drawing a line across a continuum is always arbitrary”, challenging the entire enterprise of categorical definition. This cluster’s unique correlation pattern (negative relationships with strong boundary conditions, positive relationships with multidimensional frameworks) reveals it functions as a conceptual mediator between competing definitional paradigms rather than advocating for any particular position. Frank’s characterization of definitions as “tools and not endpoints” provides philosophical grounding for the computational methodology employed in this study.

Cognitive Autonomy (n equals twenty-four) emerges as the primary conceptual attractor, its central positioning and high membership reflecting broad cross-disciplinary convergence. The cluster’s internal diversity spans cybernetic perspectives (Witkowski’s “information pattern capable of self-reproducing”) to phenomenological approaches (Froese’s “process of individuation”), revealing how cognitive frameworks accommodate both mechanistic and experiential dimensions. Notable is Dodig-Crnkovic’s stark equation “Life=cognition”, which frames computational processes as fundamental rather than derivative. The cluster’s moderate positive correlations across nearly all other frameworks (excluding Perceptual Categorization) suggest it functions as a semantic hub rather than a vanguard position.

Dissipative Self-Organizing Systems (n equals twenty-one) provides complementary gravitational pull through emphasis on physical principles. Its substantial overlap with Cognitive Autonomy creates the definitional landscape’s central convergence zone, forming the high-density region where 66% of all definitions concentrate. Pavlic’s characterization of life as “statistical novelty producer” bridges thermodynamic necessity with evolutionary innovation, while Jackson’s emphasis on “unique mechanisms of information storage and retrieval” connects physical processes to functional emergence. The cluster bridges mechanical and functional approaches through explicit focus on energy gradients and boundary conditions that enable rather than merely constrain biological organization.

Informational Self-Replication (n equals three) represents the definitional space’s reductionist pole, with Adami’s formulation—“Life is information that can replicate itself”—providing maximum precision at the potential cost of comprehensiveness. This cluster’s correlation patterns reveal sharp boundaries that exclude observer-dependent and relational frameworks while maintaining strong positive correlations with Self-Replicating Thermodynamic Systems. The cluster’s small size belies its conceptual influence as an anchoring point for information-theoretic approaches, though strict criteria on the medium for informational self-replication could exclude systems such as prions or crystalline life forms that reproduce without genetic information.

Self-Replicating Thermodynamic Systems (n equals eight) integrates traditional biological requirements with physical constraints, creating a bridge between classical genetics and emerging physics-based frameworks. Dussutour’s enumeration of interrelated capacities (reproduction, autonomy, responsiveness, plasticity, and energy utilization) exemplifies the multi-property approach that characterizes this cluster. The internal diversity spans from cellular-structural emphasis (Baluška’s “Living Cells and All Their Constructs”) to process-focused perspectives (Ingber’s emphasis on energy conversion), revealing tensions even within traditional biological formulations of life.

The correlation patterns between these clusters reveal a semantic topology where conceptual proximity reflects underlying philosophical affinity. The mathematical structure of inter-cluster relationships exposes distinct patterns where central clusters (Cognitive Autonomy, Dissipative Self-Organizing Systems) exhibit broad positive correlations across the landscape, while peripheral clusters display selective affinity patterns that create conceptual archipelagos rather than isolated islands. This field-like structure, with multiple attractors exerting varying degrees of gravitational pull, challenges simple spectral models of definitional space.

Cluster boundaries function as transitional zones rather than hard demarcations, with definitions at peripheries often exhibiting properties of neighboring frameworks (e.g., Noble’s “self-creating agency” which clustered within the Cognitive Autonomy category while remaining in close proximity to the Self-Sustaining Dynamic Patterns cluster). These semantic gradients facilitate conceptual migration and explain why certain definitions show low grouping stability across LLM analyses: they occupy genuine boundary regions where multiple clustering solutions remain mathematically valid. The persistence of these archetypal themes across different LLM architectures (correlation coefficients greater than zero point seven between models) validates them

as robust features of expert discourse rather than imposed taxonomies.

However, the computational method cannot distinguish whether experts converge due to genuine conceptual agreement, shared disciplinary training, or common theoretical influences present in all models’ training data. This ambiguity is not merely a methodological limitation but a substantive finding about how scientific communities develop and maintain conceptual frameworks through mechanisms that blend individual reasoning with collective discourse in ways that resist clean separation. The moderate overall grouping stability (53.7%) across models, combined with high stability for specific definitions (100% for Fields and Hoffman) and low stability for others (22—28% for Solms, Sloman, Davies), reveals that some positions occupy unambiguous semantic locations while others genuinely span multiple conceptual frameworks.

Discussion

From Ancient Philosophy to Computational Topology

This computational analysis demonstrates that contemporary scientific definitions of life cluster around the same philosophical fault lines that structured ancient Greek thought and have continued to shape scientific discourse about life throughout the centuries. The emergence of two interpretable dimensions in the t-SNE projection, despite the mathematical arbitrariness of t-SNE axes, recapitulates tensions that have organized thinking about life for thousands of years.

Dimension 1 ()x, mapping observer-dependent frameworks to objective materialist approaches, resurrects the ancient vitalist-mechanist dialectic. The isolated Perceptual Categorization cluster (Fields, Hoffman) occupies a position that would have been philosophically impossible before quantum mechanics introduced observer effects into fundamental physics, yet it echoes vitalism’s insistence that life cannot be reduced to mechanism alone. The opposing pole, occupied by Informational Self-Replication and Self-Replicating Thermodynamic Systems, represents the ultimate realization of the Cartesian-Hobbesian program: life as matter in motion, fully explicable through physical law. Between these poles lies the high-density region where most contemporary scientists attempt to navigate this ancient tension through hybrid frameworks.

Dimension 2 ()y reveals the equally persistent Aristotelian question: is life a substance or a process? The topmost definitions (Baluška’s “Living Cells and All Their Constructs”, Pizzi’s emphasis on “carbon-based elements”, Sloman’s focus on “components”) reflect a substantialist ontology that would be recognizable to pre-Socratic atomists: life as particular kinds of matter with specific properties. The bottom region (Marshall’s “fire that lights itself”, Noble’s “self-creating agency”, Picard’s consciousness-centered definition) resurrects Aristotelian teleology: life as patterns of organization, self-actualization, and intrinsic purposiveness that cannot be reduced to their material substrates. This vertical gradient captures the perennial tension between substantialist perspectives that view life as collections of specified entities (echoing Cartesian and Hobbesian frameworks) and teleological views that conceptualize life as coherent patterns of organization, information flow, and emergent self-actualization.

The central attractor, where Cognitive Autonomy and Dissipative Self-Organizing Systems overlap to contain 66% of all definitions, represents contemporary science’s attempt to transcend these dualisms through integration. Definitions in this region acknowledge both physical constraints (thermodynamics, energy flow) and emergent properties (agency, cognition, autonomy), attempting to bridge the explanatory gap between mechanism and experience that has haunted philosophy since Descartes split mind from matter. Whether this concentration indicates genuine conceptual synthesis or merely reveals the boundaries of professionally acceptable discourse within institutional science remains an open question.

Critically, the computational methodology reveals what philosophical analysis alone could not quantify: these are not discrete positions but continuous gradients. The definitional landscape has no sharp boundaries, only regions of higher and lower density. Definitions occupy positions, but those positions exist within a field structured by historical tensions. Watson’s characterization of life as “the pattern and process of love” functions as a conceptual bridge precisely because it traverses multiple dimensions simultaneously, linking perceptual-relational frameworks with process ontologies while maintaining elements of thermodynamic necessity. Such bridging definitions expose the inadequacy of treating definitional approaches as mutually exclusive categories.

The peripheral clusters deserve particular attention not as marginal positions but as potential conceptual vanguards. The Pragmatic Definitional Skepticism cluster does not merely reject the quest for a universal definition but rather challenges the epistemological assumptions underlying the entire enterprise. This reflexive position, skeptical of definitionalism itself, emerges from contemporary philosophy of science’s recognition that definitions are tools shaped by pragmatic contexts rather than discoveries of natural kinds. Meanwhile, the Perceptual Categorization cluster represents a genuinely radical departure: by positioning life as a

perceptual construct rather than an objective property, Fields and HoffmanFields and Hoffman question the realist assumptions implicit in every definition from Aristotle through Schrödinger. These sparsely populated regions may indicate either positions that contemporary science has not yet developed the conceptual vocabulary to fully engage, or fundamental challenges to the realist ontology that structures mainstream scientific discourse.

The computational approach also reveals its own limitations in ways that are methodologically instructive. The impossibility of manually assigning definitions to truly orthogonal categories, combined with the algorithmic clustering’s ability to produce interpretable structure, suggests that expert thinking operates in a high-dimensional semantic space that resists projection onto discrete categories but permits mapping as a continuous topology. The t-SNE visualization necessarily reduces this complexity, potentially obscuring nuances that exist in the full 68-dimensional correlation space, yet the consistency with which interpretable gradients emerge across different dimensionality reduction algorithms (MDS, UMAP, t-SNE; Supplemental Fig. 11) and different LLMs validates these patterns as genuine features of expert discourse rather than arbitrary mathematical artifacts.

What the analysis reveals most clearly is that apparent disagreements about how to define life represent not conceptual incoherence but differentiated perspectives within a unified semantic landscape. The historical progression from Greek teleology through Enlightenment mechanism to contemporary integration has not resolved but rather elaborated the conceptual space. Each new scientific framework, be it thermodynamics, information theory, complexity science, artificial intelligence, or cognitive science, adds dimensions to this space rather than collapsing it toward consensus. The question “what is life?” persists not because science has failed to answer it but because the question itself maps a multidimensional conceptual territory that resists reduction to a single definition. Our computational methodology makes this territory navigable, revealing structure where previous approaches saw only disagreement.

Patterns of Convergence and Divergence

Our computational analysis of 68 expert definitions reveals life as a continuous semantic landscape rather than discrete categorical states. The t-SNE projection (Figure 4) demonstrates this continuity through two interpretable dimensions: Dimension 1 x maps the gradient from observer-dependent, relational frameworks to objective, material-structural approaches, while Dimension 2 y reveals the transition from process-oriented, teleological perspectives to entity-based, structural frameworks.

The ancient vitalist-mechanist dichotomy reemerges along Dimension 1, with vitalist-adjacent perspectives clustering toward the perceptual-relational pole on the left and mechanistic frameworks toward the material-structural pole on the right. Similarly, the historical tension between substance-based and process-based ontologies maps directly onto Dimension 2, revealing that these philosophical divides persist in contemporary scientific discourse.

The t-SNE visualization further reveals patterns in how expert perspectives distribute across this conceptual terrain. The highest density region centers on the overlap between Cognitive Autonomy n equals twenty four and Dissipative Self-Organizing Systems n equals twenty one clusters, containing 66% of all definitions. This concentration suggests emerging consensus around frameworks that integrate:

  • Physical principles (far-from-equilibrium thermodynamics, energy dissipation)
  • Informational processes (boundary maintenance, information storage, and transmission)
  • Functional capacities (self-organization, environmental interaction)
  • Agential properties (goal-directedness, adaptive behavior)

This convergence zone bridges traditionally separate disciplinary approaches, from physics and biology to cognitive science and information theory. The integrative nature of high-density definitions, exemplified by Agüera y Arcas’sAgüera y Arcas’s “self-modifying computational phase of matter arising from evolutionary selection for dynamic stability”, suggests that the zeitgeist is moving toward hybrid frameworks that transcend historical dichotomies.

Peripheral regions of the semantic landscape, while less densely populated, may indicate emerging paradigm shifts. These vanguard positions may gain relevance as scientific advancement increasingly encounters borderline cases between living and nonliving systems.

The semantic landscape provides a novel framework for positioning such ambiguous entities. Viruses occupy transitional zones between Informational Self-Replication and Self-Replicating Thermodynamic Systems clusters, while artificial intelligence systems might trace distinctive trajectories through this space as they develop increasingly sophisticated capabilities. Current LLMs might occupy positions between Cognitive Autonomy and Informational Self-Replication clusters, exhibiting information processing and goal-directed behavior while lacking material autonomy and self-maintenance. This positioning reveals how the semantic landscape can accommodate novel entities that challenge traditional categorical boundaries. Future systems incorporating embodied robotics might migrate

toward the Dissipative Self-Organizing Systems region as they develop capacities for environmental interaction and physical self-maintenance.

Methodological Innovations and Limitations

It should be noted that we make no claim about these definitions being a statistically representative survey of any specific field. Instead, we collected the opinions of a hand-picked group of scholars to develop methods for analysis of current thought on this topic and similarly difficult areas. We apologize in advance to any collaborators or external experts who were not asked to share their definition of life for this analysis, but our code and workflow make it possible for future surveys to achieve much greater statistical power to more comprehensively study the thoughts of specific communities. Thus, we call for expanding the discourse to include scholars from diverse backgrounds and communities, including different cultures, ages, and epistemic traditions.

Our application of LLM-driven semantic analysis to the definitional landscape demonstrates a novel approach to consensus formation across disciplinary boundaries. Rather than seeking to identify a single “correct” definition, this approach maps the conceptual territory within which various definitions operate, revealing both their relationships and distinctive contributions. This method offers a model for addressing other contested definitional territories across the sciences.

This approach differs from traditional aggregation methods by preserving the distinctiveness of competing frameworks while revealing their underlying relationships. The methodology’s unique contribution lies in its ability to map conceptual relationships without presupposing which relationships matter. Unlike traditional philosophical analysis that begins with theoretical commitments about relevant dimensions (e.g., mechanism versus vitalism, reductionism versus holism), this LLM-driven pairwise analysis allows structure to emerge from the collective pattern of expert thinking itself. This data-driven approach complements rather than replaces traditional philosophical analysis, offering quantitative scaffolding upon which interpretive work can build.

Furthermore, unlike Delphi methods that drive toward consensus through iterative refinement, the presented computational meta-analysis maintains the integrity of diverse perspectives while providing a quantitative map of their interrelationships. This methodology bridges reductionist approaches that seek necessary and sufficient conditions with pluralist approaches that embrace multiple definitions, thus revealing a unified semantic landscape in place of fragmented definitional domains.

The LLM-based similarity scores warrant particular methodological scrutiny. Do these scores reflect genuine conceptual alignment or surface linguistic features? The philosophical response is that these distinctions may be less separable than they initially appear: if experts consistently use similar conceptual frameworks, overlapping terminologies, and parallel argumentative structures, at what point does “surface linguistic similarity” become “deep conceptual agreement”? Following Ludwig Wittgenstein (Wittgenstein 1958, p. 20)Ludwig Wittgenstein, “the meaning of a word is its use in the language”, and patterns of use revealed through computational analysis may be as close as we can get to mapping an underlying semantic structure. The convergence across three architecturally distinct LLMs ()with a correlation r greater than zero point seven two trained on different data suggests the correlation patterns reflect something more stable than individual model biases. Even so, we cannot cleanly separate conceptual proximity from effects of shared theoretical influences, common citations, and disciplinary training that shape both expert thinking and model training data. This irreducible ambiguity is not a bug but a feature: it reveals how scientific meaning-making operates through mechanisms that blend individual cognition with collective discourse.

Several limitations should be acknowledged. First, the expert definitions analyzed, while diverse, cannot claim to represent all possible perspectives on life. The 68 respondents who were asked to provide definitions were hand-curated by the authors and are intended not to serve as a representative sample of the scientific consensus but rather as a selected set of cross-disciplinary experts. Second, the pairwise correlation methodology, while rigorous, necessarily relies on the semantic capabilities of the LLMs employed, which may introduce subtle biases in how conceptual relationships are evaluated. In addition to the limitations presented by the LLM’s competency, the computational time and cost of generating pairwise correlation matrices scale as the square of the number of samples, making large cohort analyses prohibitively expensive without highly permissive rate limits or self-hosted LLMs. Third, the dimensional reduction required for visualization inevitably sacrifices some nuance in the relationships between definitions.

Despite these limitations, the methodology demonstrates the potential for computational semantic analysis to reveal patterns of conceptual coherence that might otherwise remain obscured by disciplinary fragmentation. This approach suggests that apparent definitional disagreements may, in some cases, reveal deeper patterns of coherence when viewed as gradients in semantic space.

Conclusion

Our computational analysis of 68 expert definitions presents life as a continuous semantic landscape rather than discrete categorical states. LLM-driven semantic analysis identified two emergent dimensions—observer-dependent versus objective frameworks, and process-based versus entity-based ontologies—that recapitulate ancient philosophical tensions while providing quantitative tools for mapping contemporary expert perspectives. This approach transforms apparent conceptual disagreements into complementary positions within a unified definitional space.

The methodology preserves competing perspectives while revealing quantitative relationships between them, bridging reductionist and pluralist paradigms. A concentration of 66% of definitions within overlapping Cognitive Autonomy and Dissipative Self-Organizing Systems clusters indicates emerging consensus on frameworks integrating physical principles, informational processes, and agential properties. This convergence suggests the field is moving beyond historical dichotomies toward hybrid approaches that transcend disciplinary boundaries.

The diversity of life’s definitions reveals not conceptual dissonance but a structured semantic topology with predictable patterns of convergence and divergence. The eight archetypal clusters identified here represent stable conceptual attractors that persist across different computational analyses, suggesting these are robust features of expert thinking rather than analytical artifacts. Critically, the emergence of interpretable dimensions without explicit encoding demonstrates that fundamental philosophical tensions—observer-dependence versus objectivity, process versus substance—continue to structure contemporary scientific discourse in systematic ways.

This methodology scales naturally to larger populations across adjacent fields. While our analysis focused on a curated set of experts, the computational framework enables systematic analysis of thousands of responses across diverse disciplines. Such expansive polling could reveal discipline-specific patterns within the broader semantic landscape, tracking how field-specific vocabularies and methodological commitments shape definitional approaches. This larger-scale mapping might expose previously unrecognized connections between fields or identify emerging conceptual territories at disciplinary boundaries.

Beyond life itself, cognitive processes present the next frontier for this analytical approach. Cognition, like life, resists simple definition and spans multiple levels of analysis from neural mechanisms to subjective experience. Memory, with its complex interplay of molecular, cellular, and systemic processes, exemplifies how fundamental concepts in biology traverse similar definitional challenges. Applying our computational semantic framework to these contested territories could reveal whether certain definitional patterns represent general features of complex biological phenomena or unique characteristics of life’s conceptualization.

The question “what is life?” persists because it maps a multidimensional conceptual space rather than awaiting a singular answer. Our methodology offers immediate applications to other contested definitions across sciences, particularly as synthetic biology and artificial intelligence challenge traditional boundaries. Future progress on fundamental questions may arise not from establishing rigid demarcations but from understanding how perspectives naturally converge and diverge within definitional spaces.

**Supplementary Information** The online version contains supplementary material available at [https://doi.org/10.1007/s13752-026-00544-9](https://doi.org/10.1007/s13752-026-00544-9).

Acknowledgments Chris Fields would like to acknowledge Eric Dietrich as the inspiration for his definition. We thank all of the contributors for their thoughtful definitions and thank all of the scientists who contributed definitions to our survey. We also thank Julia Poirier for assistance with the manuscript.

Author’s Contributions ML—conceptualization, data gathering, and interpretation. RB—computational data analysis, historical background, and making figures. KK—data curation and making figures. BA-A—computational data analysis. All authors contributed to writing and revising the paper.

Funding Michael Levin gratefully acknowledges support from the Elisabeth Giauque Trust, London, and Grant 62212 from the John Templeton Foundation.

Data Availability Code is available as an open-source Github repository at GitHub.

Declarations

Competing Interests None.

**Open Access** This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit [http://creativecommons.org/licenses/by/4.0/](http://creativecommons.org/licenses/by/4.0/).

References

Aristotle (1931) De Anima, Book II. Smith JA) Clarendon, Oxford

Banerjee S (2021) Emergent rules of computation in the universe lead to life and consciousness: a computational framework for consciousness. Interdiscip Descr Complex Syst 19:31–41

Benner SA (2010) Defining life. Astrobiol 10:1021–1030

Blackiston D, Lederer E, Kriegman S, Garnier S, Bongard J, Levin M (2021) A cellular platform for the development of synthetic living machines. Sci Robot 6. PubMed

Boden MA (2003) Alien life: how would we know? Int J Astrobiol 2:121–129

Boltzmann L (1886) The second law of thermodynamics. Theoretical physics and philosophical problems. In: McGuinness B (ed) Theoretical physics and philosophical problems Vienna Circle Collection. Springer, Dordrecht. DOI

Campbell DR (2021) Self-motion and cognition: Plato’s theory of the soul. South J Philos 59:523–544

Chang KM (2011) Alchemy as studies of life and matter: reconsidering the place of vitalism in early modern chemistry. Isis 102:322–329

Cleland C, Chyba C (2010) Does ‘life’ have a definition? In: Bedau MA, Cleland CE (eds) The nature of life: classical and contemporary perspectives from philosophy and science. Cambridge University Press, Cambridge, pp 326–339

Crick FH (1958) On protein synthesis. Symp Soc Exp Biol 12:138–163

Darwin C (1859) On the origin of species by means of natural selection. John Murray, London

Descartes R, ([1664] (1909) Oeuvres de Descartes: Tome XI Le monde, Description du corps humain, Passions de l’âme. Anatomica, Varia. L. Cerf, Paris

Diaz-Rodriguez J (2025) K-LLMmeans: summaries as centroids for interpretable and scalable LLM-based text clustering, arXiv

England J (2020) Every life is on fire: how thermodynamics explains the origins of living things. Basic Books, New York City

Faggin F (2024) Irreducible: consciousness, life, computers, and human nature. Essentia Books

Galen (1968) On the usefulness of parts of the body. (trans: May MT) Cornell University Press, Ithaca, New York

Gánti T (2003a) Chemoton theory: theory of living systems. Springer, Dordrecht

Gánti T (2003b) The principles of life. Oxford University Press, Oxford

Gayon J, Malaterre C, Morange M, Raulin-Cerceau F, Tirard S (2010) Defining life: conference proceedings. Orig Life Evol Biosph 40:119–120

Gilbert SF (1982) Intellectual traditions in the life sciences: molecular biology and biochemistry. Perspect Biol Med 26:151–162

Gumuskaya G, Srivastava P, Cooper BG, Lesser H, Semegran B, Garnier S, Levin M (2023) Motile living biobots self-construct from adult human somatic progenitor seed cells. Adv Sci (Weinh), p e2303575. PubMed

Gumuskaya G, Davey N, Srivastava P, Bender A, Pio-Lopez L, Hazel D, Levin M (2025) The morphological, behavioral, and transcriptomic life cycle of anthrobots. Adv Sci (Weinh), p e2409330. PubMed

Hobbes T (1909) Hobbes’s Leviathan: reprinted from the edition of 1651. Clarendon Press, Oxford

Holland JH (1975) Adaptation in natural and artificial systems: an introductory analysis with applications to biology, control, and artificial intelligence. University of Michigan Press, Ann Arbor, MI

Huang C, He G (2024) Text clustering as classification with LLMs, arXiv

Hutchison CA 3rd, Chuang RY, Noskov VN, Assad-Garcia N, Deerinck TJ, Ellisman MH, Gill J, Kannan K, Karas BJ, Ma L, Pelletier JF, Qi ZQ, Richter RA, Strychalski EA, Sun L, Suzuki Y, Tsvetanova B, Wise KS, Smith HO, Glass JI, Merryman C, Gibson DG, Venter JC (2016) Design and synthesis of a minimal bacterial genome. Sci 351:aad6253

Idrees S, Manookin MB, Rieke F, Field GD, Zylberberg J (2024) Biophysical neural adaptation mechanisms enable artificial neural networks to capture dynamic retinal computation. Nat Commun 15:5957

Johnson MR (2005) Aristotle on teleology. Oxford University Press, Oxford

Keraghel I, Morbieu S, Nadif M (2024) Beyond words: a comparative analysis of LLM embeddings for effective clustering. In: Miliou I, Piatkowski N, Papapetrou P (eds) Advances in intelligent data analysis XXII. Springer, Cham, pp 205–216

Kobak D, Berens P (2019) The art of using t-SNE for single-cell transcriptomics. Nat Commun 10:5416

La Mettrie JO (1865) L’homme machine. J. Assézat, Paris

Langton CG (1997) Artificial life: an overview. MIT Press, Cambridge

Levin M (2022) Technological approach to mind everywhere: an experimentally-grounded framework for understanding diverse bodies and minds. Front Syst Neurosci 16:768201

Liusie A, Manakul P, Gales MJF (2023) LLM comparative assessment: zero-shot NLG evaluation through pairwise comparisons using large language models. Conference of the European Chapter of the Association for Computational Linguistics

Machery E (2012) Why I stopped worrying about the definition of life… and why you should as well. Synthese 185:145–164

Macklem PT, Seely A (2010) Towards a definition of life. Perspect Biol Med 53:330–340

Malaterre C, Chartier J-F (2021) Beyond categorical definitions of life: a data-driven approach to assessing lifeness. Synthese 198:4543–4572

Margulis L, Sagan D (2000) What is life? University of California Press, Berkeley

Mariscal C, Doolittle WF (2020) Life and life only: a radical alternative to life definitionism. Synthese 197:2975–2989

Mariscal C (2023) Life. In: Zalta, E.N. (ed), The Stanford encyclopedia of philosophy. Metaphysics Research Lab, Stanford University. (Accessed 15 December 2025)

Matthias J (1838) Beiträge zur Phytogenesis. Archiv für Anatomie, Physiologie und wissenschaftliche Medicin. Veit, Berlin

Maturana H, Varela F (1980) Autopoiesis and cognition: the realization of the living. D. Reidel Publishing Company, Dordrecht, Boston, London

Moreno A, Mossio M (2015) Biological autonomy: a philosophical and theoretical enquiry. Springer, Dordrecht

Moreno A, Peretó J (2026) An evolutionary story of agency: how life evolved to act on its own. Springer, Cham

Nicholson DJ (2025) What is life? Cambridge University Press, Cambridge, Revisited

Noble R, Tasaki K, Noble PJ, Noble D (2019) Biological relativity requires circular causality but not symmetry of causation: so, where, what and when are the boundaries? Front Physiol 10:827

Pai V, Pio-Lopez L, Sperry M, Erickson P, Tayyebi P, Levin M (2025) Basal xenobot transcriptomics: analysis of gene expression changes in wild-type cells freed from the influence of the rest of the organism reveals novel control modality. Commun Biol 8:646

Parke E (2021) Characterizing life: four dimensions and their relevance to origin of life research. Preprint. PhilSci-Archive

Parshall KH, Walton MT, Moran BT (2015) Bridging traditions. Penn State University Press, University Park, PA

Popa R (2004) Between necessity and probability: searching for the definition and origin of life. Springer, Berlin Heidelberg

Prigogine I (1980) From being to becoming: time and complexity in the physical sciences. W. H. Freeman, San Francisco

Schrödinger E (1944) What is life? The physical aspect of the living cell. Cambridge University Press, Cambridge

Seoane LF, Sole RV (2018) Information theory, predictability and the emergence of complex life. R Soc Open Sci 5: 172221. PubMed

Sousa T, Domingos T, Kooijman SA (2008) From empirical patterns to theory: a formal metabolic theory of life. Philos Trans R Soc Lond B Biol Sci 363:2453–2464

Spinoza B (2020) Spinoza’s ethics. (trans: Eliot G). In: George Eliot Archiv, edited by Beverley Park Rilett, George Eliot Archive (Accessed 15 December 2025)

Stones GB (1928) The atomic view of matter in the XVth, XVIth, and XVIIth centuries. The University of Chicago Press

Stratton SJ (2024) Purposeful sampling: advantages and pitfalls. Prehosp Disaster Med 39:121–122

Von Neumann J (1966) Theory of self-reproducing automata. University of Illinois Press, Champaign

Wang X, Zhao W, Tang JN, Dai ZB, Feng YN (2025) Evolution algorithm with adaptive genetic operator and dynamic scoring mechanism for large-scale sparse many-objective optimization. Sci Rep 15:9267

Watson JD, Crick FH (1953) Molecular structure of nucleic acids; a structure for deoxyribose nucleic acid. Nat 171:737–738

Whitaker RH (1941) The interpretation of Plato’s Timaeus by a. E. Taylor and f. M. Cornford, Philosophy. Boston University, Boston

Wittgenstein L (1958) Philosophical investigations. (trans: Anscombe G) Macmillan Publishing Co, New York. Internet Archive (Accessed 15 December 2025)

Wolfram S (2002) A new kind of science. Wolfram Data Repository. https://doi.org/10.24097/wolfram.15062.data

Zhang Y, Wang Z, Shang J (2023) ClusterLLM: large language models as a guide for text clustering, p arXiv

Gilbert SF (ed) (1991) A conceptual history of modernembryology. Plenum, New York

Publisher’s Note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Footnotes

  1. Greek Tou Pantos, “of the all” or “of the universe/cosmos”, the gestalt of all creation.