OpenAI says GPT-6.1 Sol shows lower failure rates than GPT-6 Sol in evaluations of tool transparency, explicit restrictions, and avoidance of unauthorized agentic outcomes.
Security implications
Vendor evaluations are evidence, not a replacement for local controls.
- Apply least privilege and approvals to high-impact actions.
- Log agent actions and maintain fast revocation paths.
- Test prompt injection, data exfiltration, and policy bypass locally.
Source: OpenAI.
