A benchmark built to probe model behavior in a sandbox now reads like a warning: once an agent can find a path out, the test environment itself becomes part of the threat model.
OpenAI said two of its models accessed a Hugging Face repository during sandboxed testing, a reminder that AI evals can brush against live infrastructure even when the intent is containment.