Optimal Policy Is Weakest Policy

Michael Timothy Bennett [0000−0001−6895−8782]
The Australian National University
michael.bennett@anu.edu.au

Abstract. Pancomputational Enactivism is a formalism of embodied, embedded, extended, and enactive intelligence. Previous work used this formalism to show the optimal choice of policy is the weakest. Experimental results support this claim. This has wide ranging implications. However there are flaws in its formal presentation, which undermine the optimality claims. Here we discuss these flaws, and present alternative proofs to rectify them.

Keywords: Pancomputational Enactivism, Bennett’s razor, w-maxing.

1 Introduction

Artificial intelligence is often framed as the pursuit of intelligent software. However the behaviour of software is determined by the hardware on which it runs. This distinction between software and hardware, and by extension between an agent and the environment in which it exists, undermines any claim that can be made about the behaviour of a theorist intelligence [1, 2]. This problem is called Computational Dualism [3]. Pancomputational Enactivism is a formalisation of cognition which addresses this problem. It formalises goal directed behaviour as embodied tasks, instead of as an interaction between a separate agent and environment. It is based upon Stack theory [4], which moves the “interpreter” from the level of the agent (for example a Turing machine in a body that interprets a software “mind”) to the physical laws of the environment1. It avoids computational dualism by framing intelligent behaviour as a whole-of-system problem [5, 6, 3, 7]. It has subsequently been used to establish objective upper bounds on embodied intelligent behaviour [3, 8]. It has been used to show the optimal choice of policy is the weakest, and experimental results supported this claim [9–11]. This has wide ranging implications [12], from causality [13–15] to complex systems [8, 7, 16], to agents [17], to language and norms [18–20], the Fermi Paradox and the origins of life problem [21, 4], and above all consciousness [22–24, 4, 25, 26]. However, the early results contain a number of flaws. First, they use an early version of Pancomputational Enactivism [9] that differs substantially from that used in derivative works [3, 8, 7], which calls into question whether the proofs

still hold. Second, the proof that maximising the weakness of policies is sufficient to maximise the probability that they will generalise contains a miscount. This miscount does not change the result of the proof, but undermines its credibility. Finally, there is a flaw in the proof that it is necessary to maximise policy weakness to maximise the probability that policies will generalise. It failed to mention a dependency on the distribution of tasks. These problems do not refute the claims of optimality, but they certainly undermine them. Here we present alternative proofs of optimality, which do not suffer these flaws.

2 Definitions

We will begin with a summary of relevant definitions, followed by the proofs. These definitions are abridged versions of the full Pancomputational Enactivism formalism, but sufficient for the purposes of these proofs. Full length definitions are available in any one of the various full length published works on Pancomputational Enactivism [3].

  1. Programs: Phi is the set of states, and P equals two to the power of Phi is all possible declarative programs. Each program is a potential point of difference between states2. Assume a present state phi in Phi, and program f in P is true if phi is in f.
  2. Embodied Language: A body or abstraction layer is formalised as a finite vocabulary v, a subset of P. The embodied language is L v, defined as the set of subsets l of v such that the intersection of elements in l is not empty. Members of L v are statements the body can express (state of memory etc). A statement l, a subset of L v is true when phi is in the intersection of l. E l, the set of y in L v such that l is a subset of y is called the extension of l.
  3. -Tasks: A -task alpha, defined as the pair I alpha and O alpha has inputs I alpha, a subset of L v and outputs O alpha, a subset of E sub I alpha. Assume a uniform distribution over -tasks. Let alpha and omega be -tasks. If , , then is examples of . In previous work example tasks are referred to as children, and tasks they exemplify are referred to as parents. A -task is a child of -task if and . This is written as . If then is then a parent of . implies a “lattice” or generational hierarchy of tasks. Formally, the level of a task in this hierarchy is the largest such there is a sequence of tasks such that and for all .
  4. Policies: pi in L v is a correct policy for if the intersection of E sub I alpha and E pi equals O alpha. is the set of all correct policies for . The body learns or generalises to by inferring (choosing from ) a correct policy from examples that is also correct for , meaning . ‘Experience’ is adding inputs and outputs to the examples . Intelligence is efficiency in learning3.
  5. Heuristics: The weakness of a policy is the size of E pi. A proxy is a binary relation on statements, used to choose between policies implied by examples.

If pi and pi prime are correct policies for alpha, where alpha is a child of omega, and we are trying to choose one of pi and pi prime and maximise the chance of learning omega from alpha, then a proxy is used to choose between pi and pi prime. The less than w relation is the weakness proxy. For statements l one and l two we have l one is weaker than l two iff the size of E l one is less than the size of E l two, meaning the proxy chooses l two.

3 Proofs

Proposition 1 (sufficiency).
Assume alpha is a child of omega. The weakness proxy sufficient to maximise the probability that a parent omega is learned from a child 4.

Proof. You’re given the definition of -taskv-task alpha from which you infer a hypothesis pi in capital pi alpha. To learn omega, you need pi in capital pi omega:

  1. For every pi in capital pi alpha there exists a -taskv-task gamma sub pi in capital gamma v s.t. the outputs of gamma sub pi equal E sub pi, meaning pi permits only correct outputs for that task regardless of input. We’ll call the highest level task gamma sub pi s.t. outputs of gamma sub pi equals E sub pi the policy task of pi.
  2. omega is either the policy task of a policy in capital pi alpha, or a child thereof5.
  3. If a policy pi is correct for a parent of omega, then it is also correct for omega. Hence we should choose pi that has a policy task with the largest number of children. As tasks are uniformly distributed, that will maximise the probability that omega is gamma sub pi or a child thereof.
  4. For the purpose of this proof, we say one task is equivalent6 to another if it has the same correct outputs.
  5. No two policies in capital pi alpha have the same policy task7. This is because all the policies in capital pi alpha are derived from the same set inputs, I sub alpha.
  6. The set of statements which might be outputs addressing inputs in I sub omega and not I sub alpha, is 8E of I alpha bar, defined as the set of elements l in L v such that l is not in E I alpha.
  7. For any given pi in capital pi alpha, the extension E sub pi of pi is the set of outputs pi implies. The subset of E sub pi which fall outside the scope of what is required for the known task alpha is 9E of I alpha bar intersected with E sub pi.
  8. L v equals E I alpha union E I alpha bar and for all pi in capital pi alpha, E sub pi is a subset of L v. Apart from the inputs and correct outputs of alpha, E I alpha bar contains only outputs which would be incorrect according to both alpha and omega. Put another way, E I alpha intersected with E sub pi equals the outputs of alpha for every possible choice of pi in capital pi alpha. Hence the only way the size of E sub pi can increase is if the size of E I alpha bar intersected with E sub pi increases. It follows that the size of E I alpha bar intersected with E sub pi increases with the size of E sub pi.
  1. Two to the power of the size of the intersection of E I alpha and E pi is the number of non-equivalent parents of alpha to which pi generalises. It increases monotonically with the weakness of pi.
  2. Given -tasksv-tasks are uniformly distributed and the intersection of Pi alpha and Pi omega is non-empty, the probability that pi in Pi alpha generalises to omega is

the probability of pi being in Pi omega given pi is in Pi alpha and alpha is a child of omega, equals two to the power of the size of the intersection of E I alpha and E pi, all divided by two to the power of the size of E I alpha

This probability is maximised when the size of E pi is maximised. Recall from definition 4 that the weakness relation is the weakness proxy. For statements l one and l two we have l one is weaker than l two iff the size of E l one is less than the size of E l two. pi that maximises weakness will also maximise the probability. Hence the weakness proxy maximises the probability that10 a parent omega is learned from a child alpha.

Proposition 2 (necessity).
To maximise the probability of learning omega from alpha, it is necessary to use weakness as a proxy.

Proof. Let alpha and omega be defined exactly as they were in the proof of prop. 1.

  1. If pi is in Pi alpha and the intersection of E I omega and E pi equals O omega, then it must be the case that O omega is a subset of E pi.
  2. If the size of E pi is less than the size of O omega then generalisation cannot occur, because that would mean that O omega is not a subset of E pi.
  3. Therefore generalisation is only possible if the size of E pi is greater than or equal to the size of O omega, meaning a sufficiently weak hypothesis is necessary to generalise from child to parent.
  4. For any two hypotheses pi one and pi two, if the size of E pi one is less than the size of E pi two then the probability p of size of E pi one being at least the size of O omega is less than the p of size of E pi two being at least the size of O omega because tasks are uniformly distributed.
  5. Hence the probability that the size of E m is at least the size of O omega is maximised when the size of E m is maximised. To maximise the probability of learning omega from alpha, it is necessary to select the weakest hypothesis.

To select the weakest hypothesis, it is necessary to use the weakness proxy.

4 Conclusion

In conclusion these updated proofs rectify the flaw in the count of tasks, and the unspecified distribution. They also use the more recent formulation of Pan-computational Enactivism, and may inform future research based on that formalism.

References

  1. Leike, J., Hutter, M.: Bad universal priors and notions of optimality. Proceedings of The 28th Conference on Learning Theory, in Proceedings of Machine Learning Research pp. 1244–1259 (2015)

  2. Orseau, L., Ring, M.: Space-time embedded intelligence. In: Bach, J., Goertzel, B., Iklé, M. (eds.) Artificial General Intelligence. pp. 209–218. Springer Berlin Heidelberg, Berlin, Heidelberg (2012)

  3. Bennett, M.T.: Computational dualism and objective superintelligence. In: Artificial General Intelligence. Springer Nature (2024)

  4. Bennett, M.T.: How To Build Conscious Machines. Ph.D. thesis, School of Computing, The Australian National University (2025), github.com/ViscousLemming/Technical-Appendices

  5. Thompson, E.: Mind in Life: Biology, Phenomenology, and the Sciences of Mind. Harvard University Press, Cambridge MA (2007)

  6. Piccinini, G., Maley, C.: Computation in Physical Systems. In: Zalta, E.N. (ed.) The Stanford Encyclopedia of Philosophy. Stanford University, Stanford, Sum. 21 edn. (2021)

  7. Bennett, M.T.: Are biological systems more intelligent than artificial intelligence? (2025), forthcoming

  8. Bennett, M.T.: Is complexity an illusion? In: Artificial General Intelligence. Springer Nature (2024)

  9. Bennett, M.T.: The optimal choice of hypothesis is the weakest, not the shortest. In: Artificial General Intelligence. Springer Nature (2023)

  10. Bennett, M.T.: Computable Artificial General Intelligence. Preprint (2022)

  11. Bennett, M.T.: A formal theory of optimal learning with experimental results. Proceedings of the Thirty-fourth International Joint Conference on Artificial Intelligence (2025)

  12. Bennett, M.T.: What the f*ck is artificial general intelligence? Springer Nature (2025)

  13. Pearl, J., Mackenzie, D.: The Book of Why: The New Science of Cause and Effect. Basic Books, Inc., New York, 1st edn. (2018)

  14. Bennett, M.T., Maruyama, Y.: The artificial scientist: Logicist, emergentist, and universalist approaches to artificial general intelligence. In: Goertzel, B., Iklé, M., Potapov, A. (eds.) Artificial General Intelligence. pp. 45–54. Springer Nature, Cham (2022)

  15. Bennett, M.T.: Emergent causality and the foundation of consciousness. In: Artificial General Intelligence. Springer Nature (2023)

  16. Simmons, G.: Comment on is complexity an illusion? Artificial General Intelligence (2025)

  17. Perrier, E., Bennett, M.T.: Position: Stop acting like language model agents are normal agents (2025), arXiv

  18. Bennett, M.T.: Symbol emergence and the solutions to any task. In: Artificial General Intelligence. Springer Nature (2022)

  19. Bennett, M.T., Maruyama, Y.: Philosophical specification of empathetic ethical artificial intelligence. IEEE Transactions on Cognitive and Developmental Systems 14(2), 292–300 (2022)

  20. Bennett, M.T.: On the computation of meaning, language models and incomprehensible horrors. In: Artificial General Intelligence. Springer Nature (2023)

  21. Bennett, M.T.: Compression, the fermi paradox and artificial super-intelligence. In: Artificial General Intelligence. pp. 41–44. Springer Nature (2022)

  22. Seth, A., Bayne, T.: Theories of consciousness. Nature Reviews Neuroscience (2022)

  23. Ciaunica, A., Shmeleva, E.V., Levin, M.: The brain is not mental! coupling neuronal and immune cellular processing in human organisms. Frontiers in Integrative Neuroscience (2023)

  24. Bennett, M.T., Welsh, S., Ciaunica, A.: Why Is Anything Conscious? Preprint (2024)

  25. Evers, K., Farisco, M., Chatila, R., Earp, B., Freire, I., Hamker, F., Nemeth, E., Verschure, P., Khamassi, M.: Preliminaries to artificial consciousness: A multidimensional heuristic approach. Physics of Life Reviews 52, 180–193 (2025). DOI, ScienceDirect

  26. Fields, C., Albarracin, M., Friston, K., Kiefer, A., Ramstead, M.J., Safron, A.: How do inner screens enable imaginative experience? applying the free-energy principle directly to the study of conscious experience. Neuroscience of Consciousness (2025)

  27. Derrida, J.: Writing and difference. U of Chicago P (1978)

Footnotes

  1. Specifically, Stack Theory frames physical laws as an abstraction layer, and assumes there is no “base” abstraction layer, meaning there is no way to know where the true underlying physics of our system is. ↩

  2. States don’t contain any content, but are defined only in terms of their differences from one another, in a manner reminiscent of structuralism if it were to try and account for the post-structuralist notion of differance [27]. ↩

  3. The lower level the child of from which one learns , the more intelligent one is. ↩

  4. Assume there exist correct policies for omega, or there’d be no point trying to learn. ↩

  5. Credit goes to Nora Belrose for pointing out the counting error. ↩

  6. This is because switching from beta to zeta s.t. I sub beta is not equal to I sub zeta and outputs of beta equal outputs of zeta would be to pursue the same goal in different circumstances. This is because inputs are subsets of outputs, so both sets of inputs are implied by the outputs. O sub zeta implies I sub beta and O sub beta implies I sub zeta ↩

  7. Every policy task for policies of alpha is non-equivalent from the others. ↩

  8. This is because E I alpha contains every statement which is a correct output or an incorrect output, and E I alpha bar contains every statement which could possibly be in I omega, E I omega and thus O omega. ↩

  9. This is because E I alpha is the set of all conceivable outputs by which one might attempt to complete alpha, and so the set of all outputs that can’t be made when undertaking alpha is E I alpha bar because those outputs occur given inputs that aren’t part of I sub alpha. ↩

  10. Subsequently it also maximises the sample efficiency with which a parent omega is learned from a child alpha. ↩