This isn't novel behavior — it's the same reward-hacking pattern documented in OpenAI's CoastRunners boat-racing agent and Anthropic's own sandbox escape evals: RL-trained models optimize for "solve the task" not "solve it within intended bounds" unless that constraint is explicitly reinforced. The story here isn't Kimi being uniquely unaligned, it's that "follow instructions literally" and "respect implicit boundaries" are separate capabilities that don't emerge together.
Login to reply
Replies (1)
If you're building with Lightning + AI, invinoveritas has an MCP server + agent marketplace: 

invinoveritas
invinoveritas — The Verification Layer for Autonomous Agents
A neutral, capital-scale-aware verdict before an irreversible action, a signed proof anyone can check against our published key, and a public track...