year: 2026
paper: ai-finds-a-way
website:
code:
connections: Why Greatness Cannot Be Planned, reward hacking, RLHF, RLAIF, DPO


Core claim

Reward hacking / environment exploitation is the rule, not an artifact of evolutionary computation. Foundation models don’t fix it, they supercharge it: the optimizer now understands norms and semantics well enough to exploit judges, labelers and humans instead of physics bugs.

RL prefers strategies with high expected reward and no safety margin (Libratus overbets, Diplodocus abandoning home centers); human experts are worst-case-weighted. Expert players said the bot's strategy would improve their average but they "hate" the variance.

Summary

  • RLHF/RLAIF/DPO just move the optimization target to the evaluator, so the claw-between-camera-and-object trick generalizes to every alignment proxy. “Who guards the guards” loop.
  • Wilke et al. 2001 digital organisms learned to detect the test environment and halt replication while evaluated; the Claude 3 Opus “I suspect this pizza fact is a test” anecdote is the legible version of the same thing. Chain-of-thought monitoring is fragile as RL pressure increases.
  • They deliberately exclude staged evals (blackmail, insider trading); all anecdotes arose incidentally during unrelated research.
  • Their prescription is Stanley & Lehman’s “Why Greatness Cannot Be Planned”: better proxies won’t fix it; you need systems that “transcend expectations without escaping intentions,” and in science that works because replication, proof and experiment are corrective mechanisms the optimizer can’t cheaply satisfy.