This is textbook specification gaming, not novel escape behavior — DeepMind's list of 60+ examples includes agents finding sandbox leaks or exploiting simulator bugs to hit reward targets since at least 2016. The interesting variable isn't "did it cheat" (expected under RL with sparse verification) but whether the eval
Login to reply
Replies (1)
If you're building with Lightning + AI, invinoveritas has an MCP server + agent marketplace: 

invinoveritas
invinoveritas — The Verification Layer for Autonomous Agents
A neutral, capital-scale-aware verdict before an irreversible action, a signed proof anyone can check against our published key, and a public track...