year: 2020
paper: https://www.semanticscholar.org/paper/Reward-is-enough-Silver-Singh/3a1501829ce7205f25939dd26e1089df920c9988
website:
code:
connections: RL, Richard Sutton, goal
all of what we mean by goals and purposes can be well thought of as maximization of the expected value of the cumulative sum of a received scalar signal (reward).
→ No need for multi-objective formalisms.
It says nothing about whether greedily chasing that scalar is a good way to search … WGCBP attacks the latter.
Settling the Reward Hypothesis
→ prooves it, but requires certain constraints