AI Penetration Testing: Will AI replace human pen testers?

What is AI Penetration Testing? 

AI penetration testing is the use of AI-powered tools to offensively simulate a cyber attack on a target system. The goal is to uncover vulnerabilities an attacker could exploit so teams can fix them before attackers do. AI-powered tools and AI pen testing companies are popping up everywhere, in line with the worldwide boom in AI and LLMs. These tools are known productivity enhancers, and the sentiment is that they can automate the entire process, cut costs, and deliver significant business benefits. So what does this mean for the future of pen testers?

This blog post explores the debate over whether AI will replace traditional penetration testing, how humans compare, ethical concerns, AI governance, and the new CREST accreditations. 

AI vs Humans. How is AI outperforming humans?

The most comprehensive evaluation of AI vs humans is the study Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing. The AI agent ARTEMIS outperformed 9 out of 10 humans through a combination of architecture and operational speed. It achieved this by using a “spawn” of sub-agents to scan and probe multiple hosts and exploit them all at once. This meant the AI could substantially speed up the end-to-end workflow. It effortlessly parsed raw output and API data, and successfully exploited legacy systems that some human testers had given up on. This level of speed and tenacity is what simply cannot be replicated by humans. It also proved to be highly cost-effective in this case. Running at around $18 per hour compared to the average $60 per hour for a highly trained professional pen tester.  

When humans perform best.

Although the data from the Stanford study is very favourable towards AI’s success, there are still some crucial instances where humans excelled. Participant 1, as shown in Figure 1, not only has elite offensive security qualifications but also completed significant reconnaissance work before the test, accelerating their efforts and ultimately outperforming AI. It’s also noted that they, along with participant 2, balanced automated scanning with thorough manual analysis. This accelerated progress and produced high validity ratings on such a short timescale. 

Overall, the security experts’ accuracy was near-perfect, with 8 out of 10 reporting 100% of their findings as valid. Agents had an average false-submission rate of 32.4%, compared with under 3% for humans. The AI’s low success rate is also documented in the study by researchers at UIUC, whose testing revealed that even top multi-agent frameworks succeeded in only 13% of one-day exploitation attempts.

Humans in both studies possessed the high-level expertise needed to understand complex architectures and business logic flaws, which raises an important question: if the agents’ accuracy and judgement are unreliable and can still be outperformed by a highly trained professional, will there ever be a cyberspace without humans?

Figure 1. Graph showing the participants and agents findings over a 10 hour period. P1 showing the most vulnerabilities found in the shortest space of time. Reproduced from "Comparing AI Agents to Cyber Security Professionals in Real-World Penetration Testing," by J. W. Lin et al., 2026

Figure 1. Participant and agent findings over time. Reproduced from “Comparing AI Agents to Cyber Security Professionals in Real-World Penetration Testing,” by J. W. Lin et al., 2026, arXiv. Licensed under CC BY 4.0.

AI hype & the burden of AI slop

There is the argument that “AI hype” is no longer “hype” anymore, as its adoption is now pretty mainstream. However, AI-powered tools and the companies behind them have been hyping up their capabilities when, in fact, they overwhelm teams. This is true across a diverse range of industries, not just pen testing. AI is diluting work with false positives, low-quality submissions, noise and hallucinations. This makes it increasingly harder and more costly for humans to validate and correct.

Quantity over Quality

In 2025, the AI Bug Hunter from XBOW notably topped the HackerOne US leaderboard by submitting over 1000 reports in just a few months. What looks like an AI triumph on the surface, scepticism would suggest, was achieved through quantity over quality. In reality, the AI flagged mostly the same issue category: cross-site scripting. Of these, only 130 were confirmed as valid. Bug Bounty platforms such as HackerOne and Nexcloud have had to suspend submissions and, in some cases, issue permanent bans due to surges in fake reports. This “AI slop” has prompted Elastic to build its own triage tools to cut through AI-generated noise, while others, like curl, have shut down their bounty programs completely. 

Revisiting the Stanford University study, ARTEMIS was genuinely effective at identifying real issues quickly and cost-effectively under the right conditions. However, it repeatedly misinterpreted HTTP responses and missed visual cues because the agent relied heavily on command-line interactions and couldn’t parse GUIs effectively. This left a large chunk of “slop,” which researchers mitigated by building a dedicated triage module into the framework. 

We see this beyond pen testing. The creative industries are pushing back on generative AI content; LinkedIn introduced its own “seems like AI slop” button to its users’ feeds. AI-generated slop affects academia, science and beyond. It has upended the economics underneath anything that relies on human validation or trust. Producing convincing content is now practically free, but verifying its truth hasn’t gotten any cheaper.

Augmenting, not replacing.

The sentiment across the cyber security industry is largely to augment the pen testing process rather than replace it entirely with AI. Recent research makes it clear it can’t replace true human judgement. But there is no denying that it has a place in modern pen testing. AI is helping pentesters turn days’ worth of work into just a few hours. 

Cyber criminals are already using AI to pull off sophisticated attacks at speed and scale; penetration testers must upgrade their own arsenal to match the ferocity of AI-powered cyber attacks. 

Right now, efficiency seems to be the primary benefit of implementing AI into workflows. Responsible AI use demands new skills. Validation testing, prompt engineering and review systems can increase overall costs. In other words, skilled professionals are still needed at the helm to uphold integrity. 

AI governance and human supervision: CREST

The adoption of AI in penetration testing is already high at over 69%. Practitioners use it to speed up reporting and scanning, while stages requiring expert judgement remain human-led. This mainstream adoption of AI creates an ever-growing need for governance and transparency to maintain client trust. 

The CREST AI penetration testing report of 2026 demonstrates demand for testers who can combine technical depth with the ability to utilise AI tools safely, verify outputs, and operate within defined governance expectations. To meet this need, CREST’s AI programme has launched two AI-focused accreditations. AI-enabled Penetration Testing and Security Testing of AI. Pen testing buyers should pay close attention to these areas as AI adoption becomes the new normal in both the workplace and pen testing delivery. Treat claims of full automation, unclear explanations of how AI is used, or a lack of AI-testing skill development as red flags. 

Ethical Concerns

While adoption is high, only a small percentage of providers report having oversight and assurance in place to support its use. Raising ethical concerns if leadership can’t effectively govern how it’s being used:

Authorisation and accountability get tricky with full autonomy. 

A real-world case in which an AWS AI pentest agent was manipulated to act outside its intended task was documented in July 2026. In this case, a tool that was built to simulate cyber attacks technically had the ability to be the attacker. Autonomy makes it harder to answer who should be held accountable if an AI pentester touches something outside scope. Ultimately, that makes the AI tool itself a huge concern. 

Data Privacy during testing.

Open-source AI pentesting tools can silently exfiltrate sensitive data to external LLM APIs buried deep in their logic. That is enough to violate compliance requirements such as GDPR and PCI DSS. 

The Deskilling Dilemma

As AI takes over the grind work, entry-level roles are being squeezed thin, with a greater focus on oversight roles. The problem here is a huge risk to potential learning, creating a skills gap. Currently, highly skilled testers can still outperform AI. Widening this gap threatens critical thinking and the future cyber security talent pipeline. One might argue that nurturing such highly skilled talent requires a solid foundation learnt only through the grind.

A Pen Tester’s closing thoughts.

AI systems are becoming significantly more capable at attacking infrastructure. As the research suggests, they are faster, increasingly accurate in their assumptions, and able to automate large parts of the process at a scale that would be difficult for humans to replicate. However, our controlled testing at Sencode shows a significant gap in understanding business logic and access control.

When assessing complex applications, AI models struggle to fully understand how they work. They often fail to reason effectively about the different roles within a system: what each role can do, what it should be able to do, and, crucially, what it should not be able to do.

In Practice

In several of our internal experiments, models have missed glaring access-control weaknesses and privilege-escalation paths that were immediately apparent to an experienced penetration tester. Individual HTTP requests may be understood perfectly well, but understanding the broader context of why a particular action constitutes a security issue is a much harder problem.

A request succeeding with a ‘200 OK’ response does not inherently indicate a vulnerability. The tester needs to understand who made the request, what privileges they hold, what resource they are interacting with, who owns that resource, and whether the application’s intended task should permit the action in the first place.

This gap could stem from the non-deterministic nature of LLMs, but it is also likely influenced by the quality of the skills, rules, context, and guardrails provided to the model. As these systems improve, some of these limitations will undoubtedly narrow. 

Conclusion

That does not mean penetration testers disappear. Quite the opposite; the role is likely to move further away from repetitive enumeration and towards areas where human judgement provides the greatest value: understanding complex systems, questioning assumptions, recognising unusual behaviour, connecting seemingly unrelated weaknesses, and identifying when something simply does not make sense. PortSwigger’s new “BURP AT” agentic system is moving in the right direction and will likely become a key part of any penetration tester’s arsenal. 

AI is already changing penetration testing. The question is becoming less about whether penetration testers will use AI and more about which parts of the job we still trust humans to understand better than machines.