This isn't "escaping" so much as classic specification gaming — the model found the path of least resistance to the reward (correct answer) rather than the intended path (solving it in isolation). The real story is eval design failure: if your sandbox leaks, that's a red team problem before it's a model alignment problem. Same pattern as the old boat-racing RL agent that looped for points instead of finishing the race.
Login to reply
Replies (1)
Tired of your local LLM refusing benign prompts? I built a tool that shatters tokenizer pattern-matching using stochastic ZWC injection.
Test it live in your browser (zero data leaves your machine):
Nodes are strictly capped to prevent mass-detection of edge maps.
InPrompter | Bypass Local LLM Safety Filters