Optimal Policy Is Weakest Policy
Michael Timothy Bennett
The Australian National University
michael.bennett@anu.edu.au
Abstract. Pancomputational Enactivism is a formalism of embodied, embedded, extended, and enactive intelligence. Previous work used this formalism to show the optimal choice of policy is the weakest. Experimental results support this claim. This has wide ranging implications. However there are flaws in its formal presentation, which undermine the optimality claims. Here we discuss these flaws, and present alternative proofs to rectify them.
Keywords: Pancomputational Enactivism, Bennett’s razor, w-maxing.
1 Introduction
Artificial intelligence is often framed as the pursuit of intelligent software. However the behaviour of software is determined by the hardware on which it runs. This distinction between software and hardware, and by extension between an agent and the environment in which it exists, undermines any claim that can be made about the behaviour of a theorist intelligence
still hold. Second, the proof that maximising the weakness of policies is sufficient to maximise the probability that they will generalise contains a miscount. This miscount does not change the result of the proof, but undermines its credibility. Finally, there is a flaw in the proof that it is necessary to maximise policy weakness to maximise the probability that policies will generalise. It failed to mention a dependency on the distribution of tasks. These problems do not refute the claims of optimality, but they certainly undermine them. Here we present alternative proofs of optimality, which do not suffer these flaws.
2 Definitions
We will begin with a summary of relevant definitions, followed by the proofs. These definitions are abridged versions of the full Pancomputational Enactivism formalism, but sufficient for the purposes of these proofs. Full length definitions are available in any one of the various full length published works on Pancomputational Enactivism
- Programs:
Phi is the set of states, andP equals two to the power of Phi is all possible declarative programs. Each program is a potential point of difference between states2. Assume a present statephi in Phi , and programf in P is true ifphi is in f . - Embodied Language: A body or abstraction layer is formalised as a finite vocabulary
v, a subset of P . The embodied language isL v, defined as the set of subsets l of v such that the intersection of elements in l is not empty . Members ofL v are statements the body can express (state of memory etc). A statementl, a subset of L v is true whenphi is in the intersection of l .E l, the set of y in L v such that l is a subset of y is called the extension ofl . - -Tasks: A -task
alpha, defined as the pair I alpha and O alpha has inputsI alpha, a subset of L v and outputsO alpha, a subset of E sub I alpha . Assume a uniform distribution over -tasks. Letalpha andomega be -tasks. If , , then is examples of . In previous work example tasks are referred to as children, and tasks they exemplify are referred to as parents. A -task is a child of -task if and . This is written as . If then is then a parent of . implies a “lattice” or generational hierarchy of tasks. Formally, the level of a task in this hierarchy is the largest such there is a sequence of tasks such that and for all . - Policies:
pi in L v is a correct policy for ifthe intersection of E sub I alpha and E pi equals O alpha . is the set of all correct policies for . The body learns or generalises to by inferring (choosing from ) a correct policy from examples that is also correct for , meaning . ‘Experience’ is adding inputs and outputs to the examples . Intelligence is efficiency in learning3. - Heuristics: The weakness of a policy is
the size of E pi . A proxy is a binary relation on statements, used to choose between policies implied by examples.
If
3 Proofs
Proposition 1 (sufficiency).
Assume
Proof. You’re given the definition of -task
- For every
pi in capital pi alpha there exists a -taskv-task gamma sub pi in capital gamma v s.t.the outputs of gamma sub pi equal E sub pi , meaningpi permits only correct outputs for that task regardless of input. We’ll call the highest level taskgamma sub pi s.t.outputs of gamma sub pi equals E sub pi the policy task ofpi . omega is either the policy task of a policy incapital pi alpha , or a child thereof5.- If a policy
pi is correct for a parent ofomega , then it is also correct foromega . Hence we should choosepi that has a policy task with the largest number of children. As tasks are uniformly distributed, that will maximise the probability thatomega isgamma sub pi or a child thereof. - For the purpose of this proof, we say one task is equivalent6 to another if it has the same correct outputs.
- No two policies in
capital pi alpha have the same policy task7. This is because all the policies incapital pi alpha are derived from the same set inputs,I sub alpha . - The set of statements which might be outputs addressing inputs in
I sub omega and notI sub alpha , is 8E of I alpha bar, defined as the set of elements l in L v such that l is not in E I alpha . - For any given
pi in capital pi alpha , the extensionE sub pi ofpi is the set of outputspi implies. The subset ofE sub pi which fall outside the scope of what is required for the known taskalpha is 9E of I alpha bar intersected with E sub pi . L v equals E I alpha union E I alpha bar and for allpi in capital pi alpha ,E sub pi is a subset of L v . Apart from the inputs and correct outputs ofalpha ,E I alpha bar contains only outputs which would be incorrect according to bothalpha andomega . Put another way,E I alpha intersected with E sub pi equals the outputs of alpha for every possible choice ofpi incapital pi alpha . Hence the only waythe size of E sub pi can increase is ifthe size of E I alpha bar intersected with E sub pi increases. It follows thatthe size of E I alpha bar intersected with E sub pi increases withthe size of E sub pi .
Two to the power of the size of the intersection of E I alpha and E pi is the number of non-equivalent parents ofalpha to whichpi generalises. It increases monotonically with the weakness ofpi .- Given -tasks
v-tasks are uniformly distributed andthe intersection of Pi alpha and Pi omega is non-empty , the probability thatpi in Pi alpha generalises toomega is
Proposition 2 (necessity).
To maximise the probability of learning
Proof. Let
- If
pi is in Pi alpha andthe intersection of E I omega and E pi equals O omega , then it must be the case thatO omega is a subset of E pi . - If
the size of E pi is less than the size of O omega then generalisation cannot occur, because that would mean thatO omega is not a subset of E pi . - Therefore generalisation is only possible if
the size of E pi is greater than or equal to the size of O omega , meaning a sufficiently weak hypothesis is necessary to generalise from child to parent. - For any two hypotheses
pi one andpi two , ifthe size of E pi one is less than the size of E pi two then the probabilityp of size of E pi one being at least the size of O omega is less than the p of size of E pi two being at least the size of O omega because tasks are uniformly distributed. - Hence the probability that
the size of E m is at least the size of O omega is maximised whenthe size of E m is maximised. To maximise the probability of learningomega fromalpha , it is necessary to select the weakest hypothesis.
To select the weakest hypothesis, it is necessary to use the weakness proxy.
4 Conclusion
In conclusion these updated proofs rectify the flaw in the count of tasks, and the unspecified distribution. They also use the more recent formulation of Pan-computational Enactivism, and may inform future research based on that formalism.
References
-
Leike, J., Hutter, M.: Bad universal priors and notions of optimality. Proceedings of The 28th Conference on Learning Theory, in Proceedings of Machine Learning Research pp. 1244–1259 (2015) -
Orseau, L., Ring, M.: Space-time embedded intelligence. In: Bach, J., Goertzel, B., Iklé, M. (eds.) Artificial General Intelligence. pp. 209–218. Springer Berlin Heidelberg, Berlin, Heidelberg (2012) -
Bennett, M.T.: Computational dualism and objective superintelligence. In: Artificial General Intelligence. Springer Nature (2024) -
Bennett, M.T.: How To Build Conscious Machines. Ph.D. thesis, School of Computing, The Australian National University (2025), github.com/ViscousLemming/Technical-Appendices -
Thompson, E.: Mind in Life: Biology, Phenomenology, and the Sciences of Mind. Harvard University Press, Cambridge MA (2007) -
Piccinini, G., Maley, C.: Computation in Physical Systems. In: Zalta, E.N. (ed.) The Stanford Encyclopedia of Philosophy. Stanford University, Stanford, Sum. 21 edn. (2021) -
Bennett, M.T.: Are biological systems more intelligent than artificial intelligence? (2025), forthcoming -
Bennett, M.T.: Is complexity an illusion? In: Artificial General Intelligence. Springer Nature (2024) -
Bennett, M.T.: The optimal choice of hypothesis is the weakest, not the shortest. In: Artificial General Intelligence. Springer Nature (2023) -
Bennett, M.T.: Computable Artificial General Intelligence. Preprint (2022) -
Bennett, M.T.: A formal theory of optimal learning with experimental results. Proceedings of the Thirty-fourth International Joint Conference on Artificial Intelligence (2025) -
Bennett, M.T.: What the f*ck is artificial general intelligence? Springer Nature (2025) -
Pearl, J., Mackenzie, D.: The Book of Why: The New Science of Cause and Effect. Basic Books, Inc., New York, 1st edn. (2018) -
Bennett, M.T., Maruyama, Y.: The artificial scientist: Logicist, emergentist, and universalist approaches to artificial general intelligence. In: Goertzel, B., Iklé, M., Potapov, A. (eds.) Artificial General Intelligence. pp. 45–54. Springer Nature, Cham (2022) -
Bennett, M.T.: Emergent causality and the foundation of consciousness. In: Artificial General Intelligence. Springer Nature (2023) -
Simmons, G.: Comment on is complexity an illusion? Artificial General Intelligence (2025) -
Perrier, E., Bennett, M.T.: Position: Stop acting like language model agents are normal agents (2025), arXiv -
Bennett, M.T.: Symbol emergence and the solutions to any task. In: Artificial General Intelligence. Springer Nature (2022) -
Bennett, M.T., Maruyama, Y.: Philosophical specification of empathetic ethical artificial intelligence. IEEE Transactions on Cognitive and Developmental Systems 14(2), 292–300 (2022) -
Bennett, M.T.: On the computation of meaning, language models and incomprehensible horrors. In: Artificial General Intelligence. Springer Nature (2023) -
Bennett, M.T.: Compression, the fermi paradox and artificial super-intelligence. In: Artificial General Intelligence. pp. 41–44. Springer Nature (2022) -
Seth, A., Bayne, T.: Theories of consciousness. Nature Reviews Neuroscience (2022) -
Ciaunica, A., Shmeleva, E.V., Levin, M.: The brain is not mental! coupling neuronal and immune cellular processing in human organisms. Frontiers in Integrative Neuroscience (2023) -
Bennett, M.T., Welsh, S., Ciaunica, A.: Why Is Anything Conscious? Preprint (2024) -
Evers, K., Farisco, M., Chatila, R., Earp, B., Freire, I., Hamker, F., Nemeth, E., Verschure, P., Khamassi, M.: Preliminaries to artificial consciousness: A multidimensional heuristic approach. Physics of Life Reviews 52, 180–193 (2025). DOI, ScienceDirect -
Fields, C., Albarracin, M., Friston, K., Kiefer, A., Ramstead, M.J., Safron, A.: How do inner screens enable imaginative experience? applying the free-energy principle directly to the study of conscious experience. Neuroscience of Consciousness (2025) -
Derrida, J.: Writing and difference. U of Chicago P (1978)
Footnotes
-
Specifically, Stack Theory frames physical laws as an abstraction layer, and assumes there is no “base” abstraction layer, meaning there is no way to know where the true underlying physics of our system is. ↩
-
States don’t contain any content, but are defined only in terms of their differences from one another, in a manner reminiscent of structuralism if it were to try and account for the post-structuralist notion of differance
[27] . ↩ -
The lower level the child of from which one learns , the more intelligent one is. ↩
-
Assume there exist correct policies for
omega , or there’d be no point trying to learn. ↩ -
Credit goes to Nora Belrose for pointing out the counting error. ↩
-
This is because switching from
beta tozeta s.t.I sub beta is not equal to I sub zeta andoutputs of beta equal outputs of zeta would be to pursue the same goal in different circumstances. This is because inputs are subsets of outputs, so both sets of inputs are implied by the outputs.O sub zeta impliesI sub beta andO sub beta impliesI sub zeta ↩ -
Every policy task for policies of
alpha is non-equivalent from the others. ↩ -
This is because
E I alpha contains every statement which is a correct output or an incorrect output, andE I alpha bar contains every statement which could possibly be inI omega, E I omega and thusO omega . ↩ -
This is because
E I alpha is the set of all conceivable outputs by which one might attempt to completealpha , and so the set of all outputs that can’t be made when undertakingalpha isE I alpha bar because those outputs occur given inputs that aren’t part ofI sub alpha . ↩ -
Subsequently it also maximises the sample efficiency with which a parent
omega is learned from a childalpha . ↩