year: 2023
paper: complex-behavior-from-intrinsic-motivation-to-occupy-future-action-state-path-space
website:
code:
connections:
Maximize the entropy of future action-state path
Low-probability trajectories get more “credit.”
Rewards (food, energy) are instrumental: you only care about surviving or eating because being dead means zero future path space.
They show this one principle—without any task reward—generates dancing cartpoles, hide-and-seek, and proto-altruism (freeing your pet because the pet’s state entropy counts toward yours)
“Moving around is the goal (physically/mentally), everything else (time, energy, reward) are just means to that end.”
→ curious, explorative, reward seeking, survival instinct, preference for freedom, variability in neural activity (neuroMOP)
The optimal policy is the entropy of the future paths (state + action occupancy) that you can take (alone this would lead to uniform policy), under constraints / boundary conditions (energy, time, terminal states).
“Cumulative future action-state entropy is the only measure with the additive property → open-endedness” (vs mutual information etc.) (Bellman equation applies / recursively solvable with DP). I think he means open-endedness in the sense of non-termination, but MOP does not have you enlarge the
Immediate occupancy + Future occupancy
It’s not about information. Thermodynamic view: Just spread. Conquer space and time. Species don’t go to another niche to gather information / for any specific purpose. But
But this isnt really viable for complex/open-ended envs/action spaces.
Like for an LLM it’d just be high temp sampling. Or coding agent doing random file edits.
Doesn’t this heavily depend on the geometry of teh env/state-sinks; falling back to designing the avoidance states/negative rewards?
Well, and you need some kind of canalization for the evolution perhaps. Or simply evolution scale compute and everything will emerge…