Is AI pentesting safe?

September 24, 2026

AI pentesting solution enforces scope control, harm avoidance, and testing traceabilty

There’s a lot to like about AI pentesting.

No more scheduling delays. No waiting until the engagement is over to start receiving findings. No cumbersome and non-standard PDF reports. No “three criticals and we’re done” testing methodologies or low-quality hygiene vulnerabilities.

But it’s important to recognise that security teams do have reservations about using AI models to test real assets, especially in production.

This article examines the main concerns surrounding AI pentest, and explains how we’ve addressed them in our own solution, Agentic Pentest.

The 3 big AI pentesting concerns

There are three main areas of concern:

  1. Scope control. How are testing activities confined to the intended targets, and what prevents testing from going beyond the intended scope?
  2. Harm avoidance. How can we be sure testing won’t damage the integrity or availability of our assets?
  3. Testing traceability. Will we know exactly what has been tested and how? Can we prove it to internal stakeholders, customers, and auditors?

For AI-powered pentests to be considered a safe and effective alternative to traditional penetration testing, all three of these concerns need to be addressed. This is one of the primary functions of an AI harness.

Why the concerns about safety?

Security teams have been content to use vulnerability scanners for years, but raise concerns about AI-powered testing solutions. There’s a simple reason for this, and it comes down to one distinction: the difference between deterministic and non-deterministic algorithms.

Given standard inputs, an automated vulnerability scanner (which is deterministic) will produce the same results every time. It follows a strict process from beginning to end, and there is only one correct solution. This is useful for finding known vulnerabilities, but not for uncovering real-world exploit paths.

By contrast, an AI model is non-deterministic and can identify a wide range of different correct answers. An AI pentesting solution can reason based on provided context and information discovered during testing, enabling it to find complex vulnerabilities and full attack paths, just like a human pentester.

It’s this ability to come up with different solutions and take a wide range of actions that creates concern. Security teams worry that an AI pentesting solution may take unintended actions, or degrade the performance of production assets by flooding them with requests.

Addressing these safety concerns is one of the functions of a well-designed AI harness.

What is an AI harness?

An AI model is a program trained on data to recognise patterns, and make predictions and decisions without requiring human intervention or instructions at every step. It does this using a process called inference, applying its learned patterns to new data and conditions to make predictions and take actions.

An AI harness is the digital environment that turns a general-purpose AI model into a domain-specific solution. It also includes a wide range of protective and safety measures to ensure the solution can’t act beyond its intended scope of targets and testing methods.

The harness ensures the AI model’s inference capabilities are safely channeled into an effective domain-specific solution. It does this by managing and enforcing:

  • Domain expertise and customer-provided context
  • Tools and playbooks used
  • Memory and state management
  • Engagement rules and governance
  • Guardrails and safety mechanisms
  • Agent orchestration

From here, we’ll use our own solution, Agentic Pentest, to illustrate the role of an AI harness in addressing the three major concerns security teams have about AI-powered testing.

Introducing Agentic Pentest

The YesWeHack harness is what codifies our 10+ years of offensive security expertise into a fully-capable solution. It includes:

  • Specialised Reconnaissance, Planning, Hunting, Validation, and Reporting agents.
  • YesWeHack’s proprietary offensive security playbooks and tooling.
  • Comprehensive technical controls that prevent out-of-scope activity.
  • Carefully developed safety guardrails that protect against any possible harmful activity.
  • Exploitability validation workflows that leverage our extensive triage experience.
  • Complete audit trails, so you know exactly what Agentic Pentest does and why.

In future articles, we’ll look at how each of these functions plays a role in turning a general-purpose AI model into a tightly focused pentesting solution. Since this is an article about safety, we’ll now look specifically at the role of guardrails and safety mechanisms.

Addressing the 3 big concerns

Again, the three main concerns are scope control, harm avoidance, and testing traceability. Basically, security teams want to be certain that only the intended assets will be tested, that testing activities won’t harm asset integrity or availability, and that all testing activities will be recorded and provable.

The YesWeHack harness is designed specifically to address these three needs.

  1. Your scope is strictly enforced throughout testing

After you set the scope for an engagement, the first testing activity is conducted by the Reconnaissance agent. It explores and maps the target scope, defining the range of targets for the pentest, and identifying the makeup of the application layer, e.g., any domains and API endpoints interacting with the initial asset.

Once the Reconnaissance agent has mapped your target scope and you have manually confirmed that scope, authorized assets are frozen into an allow-list for the pentest. The YesWeHack harness strictly enforces the allow-list throughout testing to ensure hunting is restricted to that scope.

  1. Safe testing is guaranteed

All actions taken by Agentic Pentest are strictly governed by the harness.

Request rates are strictly limited, automatically adapt to server and WAF responses, and are monitored through response codes, latency, error conditions, and signs of saturation. This ensures testing remains completely safe, including for production assets. Critically, the agents can kill their own processes if they are at risk of stressing the asset.

Destructive actions are excluded, preserving the stability and availability of your target scope. Excluded actions include load testing and denial of service (DoS).

Again, these safeguards are enforced by the harness across all agents throughout testing, so there is no possibility of “scope breaking” or harm to in-scope assets.

  1. All testing actions are recorded and traceable

Every action completed by Agentic Pentest is recorded, so you know exactly what has happened during testing. This includes:

  • Every hypothesis formed
  • All tests completed
  • Attack paths explored
  • Security controls encountered
  • Conclusions reached (including hypotheses discarded due to lack of exploitability)

All of this feeds into the Activity Report.

The Activity Report proves testing coverage and methodology. It details what was tested, how it was tested, and why the pentest produced the findings it did. This is particularly crucial for highly secure scopes, making it easy to demonstrate thorough testing has occurred independent of the number of findings.

The report details:

  • Vulnerability categories
  • Affected assets
  • Methodology and OffSec techniques used
  • Signals investigated an reasons for discarding certain leads
  • Exploitability analysis and reasoning

With granular detail, the report explains how the agents reasoned, all of the actions taken, and why a target was (or wasn’t) considered exploitable.

Try Agentic Pentest for yourself

If you’d like to see Agentic Pentest in action, the next step is to launch a full proof of concept.

  1. Choose a representative scope. This should be something meaningful and appropriate for your typical offensive security testing use cases. For example, an application or API with authenticated journeys will typically demonstrate more value than a simple public website.
  2. Work with our Customer Success team. Together, you can define the scope, accounts, expected depth, context, and evaluation criteria that will be used for testing.
  3. Review findings and activity evidence. This goes beyond the pure vulnerability count to assess relevance, depth, traceability, and operational fit for Agentic Pentest.

To find out more about Agentic Pentest or arrange a proof of concept, contact YesWeHack.