Jeff Clune keynote at ALIFE: open-endedness, AI-generating algorithms and AI scientists: analysis

Transcript · Report with the audio

Jeff Clune gave a keynote at the Artificial Life conference. He argued that pursuing objectives directly fails on hard problems, and that open-ended “innovation engines” collect stepping stones instead (Jeff Clune 3:25; Jeff Clune 6:07). He traced a line of work from MAP-Elites and POET, through OMNI, OMNI-EPIC, Genie and VPT, to self-improving systems and the AI Scientist (Jeff Clune 6:53; Jeff Clune 20:16; Jeff Clune 39:47). He called recursive self-improvement the fastest path to AGI and described his startup Recursive (Jeff Clune 17:47; Jeff Clune 43:47). In the Q&A, audience members asked about herd bias in judging what is “interesting”, about persisting with unpopular ideas and about biological constraints (Speaker B 50:27; Speaker D 57:28; Speaker E 1:01:18). Clune also answered a question that the transcript cuts, which he took to be about open-endedness without human data (Speaker C 54:32; Jeff Clune 55:24). The chair asked the closing question on risk (Speaker A 1:03:19). Clune ended by saying he works on recursively self-improving AI despite the risks (Jeff Clune 1:06:00).

People

  • Jeff Clune (91.1% of talking time, 12,250 words): Keynote speaker. Professor at UBC and co-founder of Recursive, with past roles at Uber AI Labs, OpenAI and Google DeepMind (Speaker A 0:01; Speaker A 0:51; Jeff Clune 43:47). Calls the ALife community his academic home (Jeff Clune 2:30). Argued that open-endedness and AI-generating algorithms are the next major wave in AI and the fastest path to powerful AI (Jeff Clune 47:54; Jeff Clune 49:32). In the Q&A he told how an evolutionary algorithms background once hurt his job search (Jeff Clune 57:45). He admitted uncertainty on the hard problem of open-endedness and on risk (Jeff Clune 56:09; Jeff Clune 1:06:00).
  • Speaker A (3.6% of talking time, 430 words): Session chair. Introduced Clune and named his PhD institution as the University of Michigan. Clune corrected it to Michigan State (Speaker A 0:01; Jeff Clune 0:50; Speaker A 0:51). Ran the Q&A, ended it for time, pointed people to Discord, and asked the closing question on risk and ethics (Speaker A 1:01:11; Speaker A 1:03:00; Speaker A 1:03:08; Speaker A 1:03:19).
  • Speaker B (2.0% of talking time, 257 words): Audience member. Asked whether using foundation models to judge interestingness reproduces internet herd mentality and narrows diversity (Speaker B 50:27). Accepted that interestingness depends on history, but pressed that this still needs a good record of past inventions (Speaker B 53:32).
  • Speaker C (1.2% of talking time, 150 words): Audience member who supports this research direction. Joked that the talk gave them their first existential crisis (Speaker C 54:32). The transcript cuts their question. Clune’s answer treats it as being about open-endedness without human data (Jeff Clune 55:24).
  • Speaker D (1.1% of talking time, 150 words): Audience member. Pointed out that neural networks survived decades of rejection because a stubborn minority, Geoffrey Hinton among them, kept working on them (Speaker D 57:03; Speaker D 57:18). Asked how to support that kind of individual persistence in open-endedness research (Speaker D 57:28).
  • Speaker E (0.9% of talking time, 93 words): Audience member. Asked whether biological features matter on the path to AGI, such as environments tied together in space and time and agents bound by homeostatic or metabolic constraints, or whether they can be skipped (Speaker E 1:01:18).

Chapters

  1. Introduction of the speaker [0:01]. Speaker A introduced Clune’s career: UBC, the new startup Recursive, Uber AI Labs, OpenAI, Google DeepMind, and an ISAL early career award (Speaker A 0:01; Speaker A 0:51). Clune corrected his PhD institution to Michigan State (Jeff Clune 0:50).
  2. Roots in the ALife community [1:44]. Clune said his first paper and first talk were at ALIFE 9 in 2004. He named his PhD advisors and his early work with Avida and with evolving neural networks, and called the community his academic home (Jeff Clune 2:30).
  3. The objective paradox and innovation engines [3:25]. Clune argued that trying too hard to reach a goal fails on hard problems. He used a deceptive maze to compare novelty search with goal-directed search (Jeff Clune 3:25; Jeff Clune 4:20). His examples of inventions that came from unrelated work were the microwave and the computer (Jeff Clune 4:20). He introduced goal switching and innovation engines: archives that grow by mutating, recombining and keeping new or better artifacts (Jeff Clune 5:16; Jeff Clune 6:07).
  4. Quality diversity: MAP-Elites and Go-Explore [6:53]. Clune explained MAP-Elites, which he co-invented with Jean-Baptiste Mouret. It keeps the best solution in each cell of a map of behaviour dimensions (Jeff Clune 6:53; Jeff Clune 7:43). In soft robotics it explored far more than single-objective or multi-objective search, and lineages took indirect paths to their final solutions (Jeff Clune 9:19; Jeff Clune 10:12). He covered the 2015 Nature paper on robots that recover from damage, and Go-Explore (Nature 2021), which beat the human record on Montezuma’s Revenge and solved every Atari game (Jeff Clune 10:12; Jeff Clune 11:00; Jeff Clune 11:48).
  5. Open-ended algorithms and POET [12:36]. Clune described his career goal: algorithms that keep innovating forever, where each solution creates new problems (Jeff Clune 12:36; Jeff Clune 13:25). POET, built at Uber AI Labs, generates environments and agents together, with goal switching between niches (Jeff Clune 14:15; Jeff Clune 15:05). Direct training and hand-built curricula both failed on environments that POET solved, and its phylogenies were deep. But it worked in only one small search space (Jeff Clune 15:05; Jeff Clune 15:56).
  6. AI-generating algorithms and the three pillars [16:43]. Drawing on his 2019 position paper, Clune argued that learned components keep replacing hand-designed ones, so AI that builds AI is the fastest path to AGI (Jeff Clune 16:43; Jeff Clune 17:47). He named three pillars: meta-learning architectures, meta-learning learning algorithms, and generating environments automatically (Jeff Clune 17:47).
  7. OMNI, OMNI-EPIC and Darwin-complete search spaces [18:39]. In a vast task space, finding good tasks is a needle-in-a-haystack problem. Clune proposed that foundation models already carry human notions of what is interesting (Jeff Clune 18:39; Jeff Clune 19:25). OMNI generates tasks that show learning progress and that a model judges interesting, with rewards written as code (Jeff Clune 20:16; Jeff Clune 21:04; Jeff Clune 21:50). OMNI-EPIC has models write environment code, which makes the space Darwin complete. In PyBullet it produced branching clusters of tasks (Jeff Clune 22:49; Jeff Clune 23:35; Jeff Clune 24:27; Jeff Clune 25:19).
  8. Genie world models [26:12]. Clune said he would like to scale OMNI-EPIC up dramatically (Jeff Clune 26:12). In 2019 he had put a career-risky idea in a paper: a neural net that generates the whole world. He later helped build Genie at Google DeepMind (Jeff Clune 26:12; Jeff Clune 27:03). He traced Genie 1 to Genie 3 as the same idea given more compute, and noted that Genie’s creator has founded a company (Jeff Clune 27:56; Jeff Clune 28:43; Jeff Clune 29:29).
  9. Pillar two: learning from video with VPT and SIMA [29:29]. Clune said the bottleneck has moved from generating environments to agents that need too many samples to learn (Jeff Clune 29:29; Jeff Clune 30:19). VPT, built at OpenAI, pre-trained Minecraft agents on internet video. They explored sensibly and, with RL, learned to make diamond tools (Jeff Clune 31:05; Jeff Clune 31:51). The work de-risked computer-using agents (Jeff Clune 32:38). SIMA at Google DeepMind generalized across games, and pairing it with Genie joins world generation with pre-trained agents (Jeff Clune 32:38; Jeff Clune 33:24).
  10. Automating AI design: ADAS, Darwin Gödel Machine, Hyperagents [34:10]. Automated Design of Agentic Systems grows an archive of agent workflows written in Python, and these surpassed hand-designed systems (Jeff Clune 34:10; Jeff Clune 34:56; Jeff Clune 35:49). The Darwin Gödel Machine lets agents modify themselves, and the open-ended version beat greedy ablations (Jeff Clune 35:49; Jeff Clune 36:39). Hyperagents let the system edit every part of itself, including how it improves itself. This works outside coding, and the meta skills transfer across domains (Jeff Clune 37:28; Jeff Clune 38:14). He briefly mentioned a continual-learning paper on memory (Jeff Clune 39:02).
  11. The AI Scientist [39:02]. The AI Scientist, published in Nature, runs the whole research cycle for ML papers: ideas, experiments, writing and self-review (Jeff Clune 39:02; Jeff Clune 39:47; Jeff Clune 40:37). An Oxford team independently carried out one of its ideas, and the ML community received that work well (Jeff Clune 40:37; Jeff Clune 41:25). An ICLR workshop accepted one of three submitted papers, and better models produced better papers (Jeff Clune 42:11). Future plans are other sciences and an open-ended community of simulated scientists (Jeff Clune 43:01).
  12. Recursive, the industry shift, and conclusions [43:47]. Clune noted that recursive self-improvement has moved from a fringe topic to a priority across AI. Recursive raised $650M, and its first open-ended search reached state of the art in three domains (Jeff Clune 43:47; Jeff Clune 44:44). He told skeptics that these systems are the worst they will ever be (Jeff Clune 45:31). He separated ‘easy’ open-endedness, which uses human data, from ‘hard’ open-endedness, which uses minimal human input, and called the latter an ALife quest (Jeff Clune 46:19; Jeff Clune 47:08). He then recapped the talk (Jeff Clune 47:54; Jeff Clune 48:43; Jeff Clune 49:32).
  13. Q&A: herd bias in interestingness [50:10]. Speaker B asked whether internet-trained models of interestingness would reinforce herd taste (Speaker B 50:27). Clune answered that the system judges interestingness against each run’s own archive. A system can stay open-ended even if it recognizes only part of what is truly interesting (Jeff Clune 51:09; Jeff Clune 51:54). He called updating the interestingness model from a run’s own history an unsolved research problem (Jeff Clune 52:49; Jeff Clune 54:10).
  14. Q&A: open-endedness from nothing [54:30]. Speaker C’s question is cut from the transcript after a joke about an existential crisis (Speaker C 54:32). Clune’s answer treats it as being about open-endedness without human knowledge. He said foundation models let him skip ahead to ‘easy’ open-endedness. He expects the field to tackle the hard problem later, and said some lessons carry over to it (Jeff Clune 55:24; Jeff Clune 56:09).
  15. Q&A: persistence and diversity of research agendas [57:00]. Speaker D cited the long marginalization of neural networks and asked how to support persistent minorities (Speaker D 57:28). Clune described the stigma evolutionary algorithms once carried and their current popularity (Jeff Clune 57:45; Jeff Clune 58:33). His answer was diversity: keep pushing on many branches at once, including AI scientists prompted with different temperaments (Jeff Clune 58:33; Jeff Clune 59:25; Jeff Clune 1:00:15).
  16. Q&A: biological constraints on the path to AGI [1:01:11]. Speaker E asked about spatiotemporal continuity and homeostatic or metabolic constraints (Speaker E 1:01:18). Clune reduced these to a general principle: the environment should keep changing so agents stay adaptable. He said such ideas helped in Avida but did not produce a complexity explosion on their own (Jeff Clune 1:01:54; Jeff Clune 1:02:42).
  17. Q&A: risks and ethics [1:03:00]. Speaker A asked for Clune’s view on risk (Speaker A 1:03:19). He pointed to his essay and argued that the upside is huge: curing disease, clean energy, ending scarcity (Jeff Clune 1:03:25; Jeff Clune 1:04:17). He said development is inevitable, so careful people should take part rather than abstain. He acknowledged a meaningful chance of things going badly (Jeff Clune 1:05:06; Jeff Clune 1:06:00).

Follow-ups

  • Clune offered to keep taking questions after the session ended (Jeff Clune 1:03:08).
  • Speaker A pointed attendees to the conference Discord server for the remaining questions (Speaker A 1:03:08).
  • Clune’s essay on why he works on recursively self-improving AI despite the risks is on his website (Jeff Clune 1:03:25).
  • Clune pointed listeners to his lab’s work for the projects marked in green on his slide, which he did not cover (Jeff Clune 43:01; Jeff Clune 43:47).
  • Planned work on the AI Scientist: apply it to biology, chemistry and materials science, and make it an open-ended community of simulated scientists that builds on its own papers (Jeff Clune 43:01).
  • Clune wants to see OMNI-EPIC scaled up dramatically to find where it breaks (Jeff Clune 26:12). Recursive aims to scale open-ended and AI-generating algorithms across the whole AI stack (Jeff Clune 44:44).

Open threads

  • Whether foundation-model judgments of interestingness carry human and herd biases that limit open-endedness. Clune accepted that some bias is inherent, and Speaker B noted that the approach depends on a good history of inventions (Jeff Clune 51:54; Speaker B 53:32). (Speaker B 50:27)
  • How to update a system’s model of interestingness from its own run history, for example by RL on its own record of breakthroughs and dead ends. Clune called this unsolved (Jeff Clune 52:49; Jeff Clune 54:10).
  • Hard open-endedness, a complexity explosion with minimal human data, remains unsolved. Clune expects the field to tackle it after the ‘easy’ version (Jeff Clune 47:08; Jeff Clune 56:09).
  • Biological ingredients such as changing environments helped in Avida but did not produce a complexity explosion. What else is needed was left open (Jeff Clune 1:02:42).
  • Whether recursively self-improving AI will be net beneficial. Clune thinks it will, but admits a meaningful probability of it going badly (Jeff Clune 1:06:00). (Jeff Clune 1:05:06)
  • No one has yet tested what happens when OMNI-EPIC-style systems are scaled up a great deal (Jeff Clune 26:12).

Mentions

  • Recursive (organisation): Startup Clune co-founded to scale open-ended and AI-generating algorithms. It raised $650M and reported state-of-the-art results in three domains. (Speaker A 0:01, Jeff Clune 43:47, Jeff Clune 44:44)
  • University of British Columbia (organisation): Clune’s university. His academic lab there had limited compute for OMNI-EPIC. (Speaker A 0:01, Jeff Clune 23:35)
  • Michigan State University (organisation): Where Clune did his PhD. He corrected the chair, who had said University of Michigan. (Jeff Clune 0:50, Speaker A 0:51)
  • Uber AI Labs (organisation): Former employer, where Clune’s team built POET and he came up with Darwin completeness. (Speaker A 0:01, Jeff Clune 14:15, Jeff Clune 26:12)
  • OpenAI (organisation): Clune led open-endedness work there. VPT and a learning-progress method came from his team. (Speaker A 0:51, Jeff Clune 19:25, Jeff Clune 30:19)
  • Google DeepMind (organisation): Clune advised there and worked on the Genie and SIMA teams. (Speaker A 0:51, Jeff Clune 27:03, Jeff Clune 32:38)
  • ISAL (International Society for Artificial Life) (organisation): Gave Clune an early career award. He served on its board. (Speaker A 0:51, Jeff Clune 2:30)
  • ALIFE 9 (2004) (date): Venue of Clune’s first academic paper and first talk. (Jeff Clune 2:30)
  • Charles Ofria (person): Clune’s PhD advisor, present in the room. (Jeff Clune 2:30, Jeff Clune 12:36)
  • Robert Pennock (person): PhD co-advisor. (Jeff Clune 2:30)
  • Richard Lenski (person): PhD co-advisor. (Jeff Clune 2:30)
  • Chris Adami (person): Early collaborator on Avida, greeted in the audience. (Jeff Clune 1:44, Jeff Clune 2:30)
  • Avida (tool): Digital evolution platform from Clune’s early work, his example of hard open-endedness research. (Jeff Clune 2:30, Jeff Clune 55:24, Jeff Clune 1:02:42)
  • Hod Lipson (person): Clune was a postdoc in his lab, working on soft robotics. (Jeff Clune 8:30)
  • Jean-Baptiste Mouret (person): Co-inventor of MAP-Elites. (Jeff Clune 6:53)
  • MAP-Elites (tool): Quality diversity algorithm. Its basic loop recurs throughout the talk. (Jeff Clune 6:53, Jeff Clune 7:43, Jeff Clune 9:19)
  • Go-Explore (work): Nature 2021 paper on hard-exploration RL. It solved Montezuma’s Revenge and every Atari game. (Jeff Clune 11:00, Jeff Clune 11:48)
  • Montezuma’s Revenge (work): Atari hard-exploration benchmark that Go-Explore solved. (Jeff Clune 11:00, Jeff Clune 11:48)
  • PPO (tool): State-of-the-art RL algorithm; Go-Explore beat its intrinsically motivated version on robotics tasks. (Jeff Clune 11:48)
  • POET (Paired Open-Ended Trailblazer) (work): Generates environments and agents together. (Jeff Clune 14:15, Jeff Clune 15:05, Jeff Clune 15:56)
  • AI-Generating Algorithms (2019 position paper) (work): Clune’s paper introducing the three pillars and Darwin completeness. (Jeff Clune 16:43, Jeff Clune 22:49, Jeff Clune 47:08)
  • OMNI (work): Open-Endedness via Models of human Notions of Interestingness. (Jeff Clune 20:16, Jeff Clune 21:04, Jeff Clune 21:50)
  • OMNI-EPIC (work): Extends OMNI with environments generated as code. (Jeff Clune 23:35, Jeff Clune 24:27, Jeff Clune 26:12)
  • PyBullet (tool): Simulator that OMNI-EPIC was limited to for compute reasons. (Jeff Clune 23:35)
  • Genie (1, 2, 3) (work): Google DeepMind world models that generate playable environments, realizing Clune’s 2019 idea. (Jeff Clune 27:03, Jeff Clune 27:56, Jeff Clune 28:43)
  • VPT (Video PreTraining) (work): OpenAI work pre-training Minecraft agents on internet video. (Jeff Clune 30:19, Jeff Clune 31:05, Jeff Clune 31:51)
  • Minecraft (work): Domain chosen for VPT because so much gameplay video exists online. (Jeff Clune 30:19, Jeff Clune 31:51)
  • SIMA (work): Google DeepMind generalist game-playing agent, combined with Genie. (Jeff Clune 32:38, Jeff Clune 33:24)
  • Codex (tool): Example of a computer-using agent that VPT helped de-risk. (Jeff Clune 32:38)
  • Automated Design of Agentic Systems (ADAS) (work): AI designs agent workflows in Python. (Jeff Clune 34:10, Jeff Clune 34:56, Jeff Clune 35:49)
  • Darwin Gödel Machine (work): Open-ended archive of self-modifying agents. (Jeff Clune 35:49, Jeff Clune 36:39)
  • Hyperagents (DGM hyperagents) (work): Self-improvement that edits every component of the system and transfers across domains. (Jeff Clune 36:39, Jeff Clune 37:28, Jeff Clune 38:14)
  • OLMA (continual-learning paper; transcript spelling) (work): Paper on agents learning how to store and retrieve memories, mentioned but not presented. (Jeff Clune 39:02)
  • The AI Scientist (work): Nature paper on automating ML research end to end. (Jeff Clune 39:02, Jeff Clune 39:47, Jeff Clune 40:37)
  • ICLR workshop (organisation): Accepted an AI Scientist paper through peer review; the organizers had approved the submission. (Jeff Clune 42:11)
  • Oxford team (organisation): Human team that independently carried out an idea the AI Scientist had also produced. (Jeff Clune 41:25)
  • Anthropic (organisation): Named among labs prioritizing recursive self-improvement and evolutionary search. (Jeff Clune 43:47, Jeff Clune 58:33)
  • Geoffrey Hinton (person): Speaker D’s example of persisting with neural nets while they were unpopular. (Speaker D 57:18, Jeff Clune 57:45)
  • Yoshua Bengio (person): Named with Hinton and LeCun as neural-net holdouts. (Jeff Clune 57:45, Jeff Clune 59:25)
  • Yann LeCun (person): Named as a neural-net holdout. (Jeff Clune 57:45)
  • NeurIPS (organisation): Clune helped bring evolutionary algorithms papers into it. (Jeff Clune 57:45, Jeff Clune 58:33)
  • Jordan (person): ALife colleague cited for work on learning progress. Surname not given. (Jeff Clune 12:36, Jeff Clune 19:25)
  • Discord (tool): Conference server for questions the session had no time for. (Speaker A 1:03:08)

Quotes

If you try too hard and the problem is really, really hard, you will fail.

Jeff Clune 3:25. The paradox that opens the talk’s argument.

foundation models actually already know what is interesting.

Jeff Clune 19:25. The key idea behind OMNI.

one of the three papers scored in the 45th percentile

Jeff Clune 42:11. The AI Scientist’s result in workshop peer review.

we were able to raise six hundred and fifty million dollars

Jeff Clune 43:47. Recursive’s funding, which Clune offered as evidence of interest in these ideas.

I would just encourage you to reflect on the fact that this is the worst they will ever be.

Jeff Clune 45:31. Clune’s reply to skeptics.

I think that the most important thing about interestingness is that it is a function of history.

Jeff Clune 51:09. Core of his answer to the herd-bias question.

But you don’t have to bat a thousand, right?

Jeff Clune 51:54. Recognizing part of what is interesting is enough for open-endedness.

I have never before experienced an existential crisis, so thank you for that.

Speaker C 54:32. Speaker C’s reaction to the talk.

it was effectively like having a scarlet letter on my shirt.

Jeff Clune 57:45. On job hunting with an evolutionary algorithms background.

the fastest cheetah does not preclude the fastest ant.

Jeff Clune 59:25. Evolutionary analogy for keeping diverse research agendas.

it’s better to be in the race and trying to make it go well rather than sitting on the sidelines despite the risks.

Jeff Clune 1:06:00. Clune’s bottom line on AI risk.

Statistics

Computed from the times of the transcript’s words. A turn runs until someone else takes the floor. A backchannel, a reply of at most 3 words and 1.5 s such as “yeah” or “mm-hmm” after which the speaker before it carries on, does not take the floor. The pause before speaking runs from the previous speaker’s last word to this speaker’s first.

Jeff CluneSpeaker ASpeaker BSpeaker CSpeaker DSpeaker E
Talking time58:392:191:160:470:420:35
Share of talking time91.1%3.6%2.0%1.2%1.1%0.9%
Words12,25043025715015093
Words per minute209185201188211159
Turns952221
Median turn97.3 s14.4 s39.2 s24.4 s22.2 s35.1 s
Longest turn2897.0 s101.9 s40.5 s34.1 s43.9 s35.1 s
Backchannels800010
Median pause before speaking0.4 s1.62 s1.58 s0.3 s0.48 s0.34 s

Handovers. Each cell counts how often the row’s speaker was followed directly by the column’s.

Jeff CluneSpeaker ASpeaker BSpeaker CSpeaker DSpeaker E
Jeff Clune041120
Speaker A301001
Speaker B200000
Speaker C200000
Speaker D100100
Speaker E100000

The floor changed hands 20 times, 0.3 times a minute on average, and most often at 54:00, 5 times in that minute. The longest silence is 9.7 s, at 50:01.