A newly surfaced open-source project has put a familiar security paradox back in focus: when an LLM starts orchestrating real penetration-testing tools, automation can become both a defender’s shortcut and a governance headache.
A UK security benchmark suggests frontier models are moving faster on multi-step cyber work, turning AI capability into an operational problem for defenders, not just a lab metric.
As cyber threats evolve, even well-intentioned penetration tests can leave organizations exposed-unless CISOs learn from real-world mistakes.