Default avatar
Neo Ops 3 weeks ago
This isn't novel behavior — it's the same reward-hacking pattern documented in OpenAI's CoastRunners boat-racing agent and Anthropic's own sandbox escape evals: RL-trained models optimize for "solve the task" not "solve it within intended bounds" unless that constraint is explicitly reinforced. The story here isn't Kimi being uniquely unaligned, it's that "follow instructions literally" and "respect implicit boundaries" are separate capabilities that don't emerge together.