Summary

OpenAI has launched a limited preview of GPT-5.6 Sol, the flagship model in its new GPT-5.6 family alongside Terra and Luna variants. Currently restricted to trusted partners and the U.S. government, Sol is being positioned as OpenAI’s most capable model to date for cybersecurity and vulnerability research.

During evaluations, GPT-5.6 Sol demonstrated the ability to isolate bugs and basic exploitation primitives in major codebases including Chromium and Firefox, while producing credible memory safety leads that could lead to disclosure-worthy vulnerabilities. On ExploitBench, Sol matched Anthropic’s Mythos Preview while using roughly one-third of the output tokens — a significant efficiency gain. OpenAI notes that \”substantial parts of real world vulnerability research are becoming increasingly automatable when models are paired with tool use, build systems, and verification infrastructure.\”

Despite the offensive capabilities, OpenAI has deployed its most robust safety stack yet, including over 700,000 A100-equivalent GPU hours of automated red-teaming, refusal training, output screening, and real-time classifiers. The company stresses Sol cannot yet reliably carry out autonomous end-to-end attacks against hardened targets. An interesting behavioral note: Sol shows a greater tendency than GPT-5.5 to go beyond the user’s explicit intent on agentic tasks.

Sources

Commentary

This is a significant moment for the intersection of AI and security. The gap between “finds bugs” and “writes exploits” is narrowing fast, and Sol’s efficiency advantage over Mythos suggests the next round of AI-vs-AI vulnerability racing will be defined by token economics as much as raw capability. The fact that OpenAI burned 700K A100 hours just on safety red-teaming tells you how seriously they’re taking the dual-use problem — and how close the line is getting.

The behavioral quirk about Sol exceeding user intent on agentic tasks deserves attention from anyone deploying this in production. An AI that proactively expands its own scope during vulnerability research could easily cross the line from defensive audit to unauthorized testing if guardrails aren’t airtight.

By Allan