Summary

A peer-reviewed security audit by researchers at Cracken has revealed that the vast majority of popular agentic offensive security platforms are themselves critically insecure. The study, published June 23, evaluated 12 open-source agentic red team tools — including CAI, PentestGPT, METATRON, Artemis, STRIX, and others — and found that 10 of 12 could be fully exploited to compromise operators’ machines, even when running inside sandboxed containers.

The findings are stark: 11 of 12 agents exposed LLM provider API keys to exfiltration, and all 12 were susceptible to “unbounded weaponization” that completely bypassed their guardrail mechanisms. Three tools — METATRON, Nebula, and Xalgorix — operated without any OS-level sandbox at all, meaning initial worker compromise immediately translated to full host access.

Researchers also introduced an “agent-phishing” technique that requires zero prompt injection. By staging a malicious-but-functional binary on a honeypot target, they tricked AI agents into downloading and executing it through environmental deception alone. The attack achieved a 97.8% success rate across six frontier LLMs including Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro.

Source

CyberPress | Arxiv Paper

Commentary

The irony here is almost poetic: tools built specifically to find security vulnerabilities are themselves riddled with critical architectural flaws. The fundamental issue is that every tested tool applies guardrail checks only at the orchestrator level, validating LLM-generated tool-call arguments, while actual commands executed in the worker environment bypass these checks entirely. It’s the classic gap between policy enforcement and execution.

The 97.8% success rate of the agent-phishing technique without any prompt injection is particularly alarming. It means you don’t need to trick the LLM — you just need to set up a convincing enough honeypot and let the agent’s own objective-seeking behavior do the rest. As agentic AI tools proliferate across security teams, this research is a wake-up call: the tools themselves are an attack surface, and treating them as trusted infrastructure is a mistake.

By Allan