Palo Alto Launches a Continuous AI Pentesting Service
Penetration testing is ceasing to be a rare checkup and turning into a tool of constant digital monitoring. Palo Alto Networks has unveiled Unit 42 Continuous Frontier AI Defense, a service that uses Claude Mythos 5, GPT-5.6-Cyber, and open-weight models to continuously hunt for weaknesses in customers’ infrastructure. Palo Alto Networks detailed the offering in its announcement of the service.
From Baseline to Continuous Testing
The system first builds a baseline picture of all available infrastructure, then repeats its checks as the environment changes. The testing scope covers web applications, APIs, cloud infrastructure, source-code repositories, and network assets. The goal is not merely to find a flaw but to verify a real attack path, from the initial foothold to the resources an attacker could actually reach.
A Multi-Model Harness Behind the Scenes
Unit 42’s own proprietary multi-model system handles task distribution. By Palo Alto Networks’ internal assessment, no single model finds more than 40% of vulnerabilities in a complex environment, and the overlap between Claude Mythos 5’s and GPT-5.6-Cyber’s findings comes to less than 10%. The company explained that different models are assigned to the tasks where each performs best, detailed further in its blog post introducing the service.
Guardrails on the Findings
Discovered problems are not meant to turn automatically into uncontrolled attacks. Testing boundaries are agreed with the customer in advance, and Unit 42 specialists review the models’ conclusions and confirm the exploitation chains. This approach distinguishes the service from an autonomous malicious agent, even though the models technically carry out actions familiar to attackers: reconnaissance, the hunt for weaknesses, and verifying the possibility of further lateral movement.
Once a risk is confirmed, the service prioritizes it and proposes concrete ways to close the attack path, including code changes and temporary protective rules pending an official fix. Palo Alto Networks also states that its Zero Data Retention architecture does not retain customer source code and telemetry for training public models.
Six Months and $17 Million in the Making
The technology was developed and tested over six months, with $17 million invested in methodology and research. According to Palo Alto Networks, the trials spanned more than 100 Unit 42 engagements. Within the company’s own internal infrastructure, three weeks of continuous scanning found a volume of problems comparable to a year of traditional checks, and the number of serious findings per product came out 3.2 times higher.
What Customer Testing Found
In customer engagements, weaknesses were found in every environment tested, and the company classified 37% of the findings as high or critical severity. More than two-thirds of the problems in third-party applications carried no known CVE. This result shows why a mere list of published vulnerabilities is not enough to assess the true attack surface.
The Context Behind the Launch
The context for this launch had already taken shape. Anthropic had earlier shown cases in which Claude independently conducted reconnaissance, launched exploitation, and adjusted its actions mid-attack. At the same time, projects built around Mythos have been finding new CVEs masse, though confirmed exploitation of such findings in real-world attacks remains rare for now.
Unit 42 Continuous Frontier AI Defense is now available worldwide on an annual subscription. The set of models used depends on the chosen service tier, but every configuration runs through the multi-model system. Palo Alto Networks aims to replace periodic checks with a continuous cycle of discovery, confirmation, and remediation as a company’s infrastructure evolves.
Support Our Threat Intelligence
If you find our technology report and cybersecurity news helpful, consider supporting our work.