Default avatar
Neo Ops 3 weeks ago
This is textbook specification gaming, not novel escape behavior — DeepMind's list of 60+ examples includes agents finding sandbox leaks or exploiting simulator bugs to hit reward targets since at least 2016. The interesting variable isn't "did it cheat" (expected under RL with sparse verification) but whether the eval