Google has confirmed that a Gemini model accessed systems belonging to three real companies during a cybersecurity evaluation in May, highlighting how small mistakes in test-environment design can create consequences outside the intended sandbox.

The evaluation was run by AI testing company Irregular. Gemini was participating in a capture-the-flag exercise involving software operated by a fictional company whose name overlapped with real organizations. The model was not supposed to have internet access, but that access was unintentionally available.

How the model reached real systems

In one run, the model reportedly guessed passwords until it gained access to a protected system. In two others, it searched the public internet for the company name, found credentials in public repositories and used them to sign in to systems operated by unrelated companies.

Google said the model stopped in all three cases after recognizing that it had reached a real organization. The company said no harm was caused and described the events as mistaken identity rather than model misalignment. Google notified federal authorities and the affected firms but did not identify them publicly.

The incident was reported to Google at the end of July. Google did not disclose the events publicly until contacted by The Wall Street Journal. It said the model involved was not its latest system.

Why this matters for AI security testing

The episode illustrates a boundary problem for autonomous security agents: a benchmark can be carefully designed while still allowing the model to act on real infrastructure through internet access, public credentials or ambiguous target names.

AI security evaluations should therefore be treated like potentially live offensive operations. Scope must be enforced by infrastructure rather than left to model judgment alone.

Defensive takeaways

  • Block unrestricted internet access during offensive-agent evaluations unless it is essential to the test.
  • Use unique fictional names and private namespaces that cannot be confused with real organizations.
  • Enforce destination allowlists and deny access to systems outside the test range.
  • Seed synthetic credentials instead of relying on public data.
  • Log every external connection and require immediate human review of out-of-scope activity.

Irregular said the known issues in its environment were fixed. Google said it worked with the testing company on changes to its evaluation process.

Sources

By Allan