Buyer's reference / updated August 2026
Every vendor in this category describes itself with the same four words: autonomous, contextual, validated, human-grade. The words carry no information. This page compares the five platforms buyers actually shortlist, on six things that can be checked instead.
Every vendor here promises validated findings and continuous coverage. Those claims are not comparable, so this page ignores them and ranks on six things a buyer can check during an evaluation.
| Vendor | Coverage model | Model approach | Platform | Target types | Control granularity | Offline / on-prem capable | Free tier |
|---|---|---|---|---|---|---|---|
| 01VLKN | Structural: 90+ discrete modules, per-class execution record | Owned, pentest-tuned models on controlled infrastructure | Yes | Web, API, network, LLM | Module-level: 90+ modules, plus model, passes, aggression, depth | Yes | Yes, full report |
| 02Novee | Not declared | Proprietary offensive model plus frontier models | Yes | Web, API, LLM | Not published | No | No |
| 03Terra Security | Not declared | No custom model claimed; "swarm" of agents | Yes | Web, API, network, LLM | Operator console for supervising agents | No | No |
| 04Horizon3 | Not declared | Graph reasoning, deterministic attack logic, ML, scoped GenAI | Yes | Network, cloud, identity, web, API | Scope authorization and graduated testing modes | No | No |
| 05XBOW | Not declared | Frontier-model agents | Yes | Web, API | Not published | No | No |
Every value here is drawn from the vendor's own public material. "Not published" means the vendor does not state a position, not that the capability is absent. Corrections welcome at info@vlkn.ai.
VLKN is built around one design decision: coverage is a structure, not an instruction. The system runs 90+ specialized modules across the engagement lifecycle, and each vulnerability class is its own execution path with its own frame of focus. A class either ran against a given endpoint or it did not, and the record says which. That makes coverage enumerable rather than asserted, and it is the reason control is granular: you choose which of those modules run, which model backs them, how many passes, how deep, and how hard it pushes.
Attack surface discovery uses vision models reading rendered screenshots rather than a link spider, on the principle that you can only test what you see. Findings ship with raw request and response, screenshots, and provenance for how the agent arrived. Testing is multi-account and role-aware, which is what makes authorization classes reachable at all.
Because the models are owned and run on infrastructure VLKN controls rather than called from a hosted frontier API, engagements can run where the internet does not reach: internal networks, non-production environments, on-premises, and air-gapped. That is what makes regulated environments addressable at all, and it is the one line in this comparison that is a can-versus-cannot rather than a matter of degree.
Run a free VLKN pentest against a target you own and compare the report to whatever you got last time.
Novee (02). The other vendor here betting on a purpose-trained offensive model rather than a wrapper, and one of the few publishing numbers. Read what the numbers measure: their claimed advantage is on constrained web exploitation challenges, which are solve-the-box tasks that end when the flag is found. A score there says the model is good at getting in. It says nothing about whether a whole surface was tested. Control approach is not published, and there is no offline deployment story.
Terra Security (03). Broad surface coverage and a human signing the report, which is the right answer for buyers whose auditors require a named person. Two things to check. Their current site makes no claim to training or tuning their own models, while older collateral still describes a "swarm" of fine-tuned agents, so ask what the word means before treating it as model ownership. And a design that routes a person in at decision points inherits that person's calendar: continuous runs continuous up to the point where someone has to look at it.
Horizon3 (04). The best-resourced entrant and genuinely strong on infrastructure attack paths, which is what the engine was built for. Web application testing is a July 2026 extension of that engine, not its purpose. The distinction matters because an engine optimized to reach an objective stops when it reaches one, which is the right shape for proving business impact and the wrong shape for answering whether every endpoint on one application was tested.
XBOW (05). The name that made this category credible, and the reference point every buyer has heard of. Scope is web and API. Control approach is not published and there is no offline deployment. The leaderboard results that built the reputation measure the agent against other people's applications under bug bounty conditions, which rewards finding one high-severity issue quickly. That is a different job from testing a surface, and worth separating before treating rank as a proxy for coverage.
There are roughly twenty companies selling AI penetration testing and adjacent offensive validation, plus 39+ open-source agents tracked by independent researchers as of April 2026. The table above covers shipping platforms only. Announced-but-unreleased products are left out until there is something to evaluate. Some of the following are worth your shortlist depending on the requirement.
Vendors here agree on almost everything and disagree on one thing: whether a person has to be in the run. It is worth reading both sides, because both are published and both are coherent.
The bottleneck argument. Novee's position is that AI-enabled testing requiring a human on every run inherits the same periodic constraint as fully manual testing. If a pentester must kick off each engagement and triage each finding, the system cannot re-evaluate attack paths as code ships, and continuous is continuous only in the marketing. Their formulation is that humans in the loop should be an option rather than a bottleneck: autonomy in the execution layer, human review available when a team wants it.
The damage argument. Terra's position is that a fully autonomous approach is not yet possible without accepting a risk that accuracy metrics do not capture. Their argument is not about false positives. It is that an autonomous system in a live production environment has no judgment available to stop it when a test payload meets a business-critical transaction, and that automated tools structurally do not understand the business they are testing.
Both are describing something real. The resolution most buyers actually want is neither pole: an execution layer that runs without a person, with controls fine enough that the person decides in advance how hard it pushes and what it is allowed to touch. That is the practical question to bring to an evaluation, rather than the autonomy label. Ask what the system does when it reaches something destructive, and ask to see where that boundary is configured.
Before comparing anything, check which product category you are actually looking at. The phrase currently describes two opposite services, and most search results mix them together.
AI testing your applications. Software agents perform the pentest against your web apps, APIs and networks. This is VLKN, Novee, Terra, XBOW, and as of July 2026 Bugcrowd Savant Pathseeker. What you are buying is testing capacity that does not depend on human headcount.
Humans testing your AI. Human pentesters assess your LLM applications for prompt injection, excessive agency, training data poisoning and sensitive data disclosure. Bugcrowd's AI Pen Test is scoped this way, using OWASP-based methodology and vetted human testers. What you are buying is expert review of a new attack surface.
Bugcrowd now sells both, under similar names, which is a fair illustration of how confusing the label has become. AI Pen Test and Savant Pathseeker are unrelated products solving unrelated problems.
Neither category substitutes for the other. If your AI features handle sensitive data or call tools, you want the second. If you need your application portfolio tested more often than your budget allows humans to test it, you want the first. Buyers who ask one vendor category for the other's outcome end up disappointed in both.
A finding is easy to verify. You read the request, you reproduce it, it either works or it does not. Absence is the hard part. Nothing in a report distinguishes a surface that was tested and found healthy from one that was never reached, which is why every vendor in this category can claim comprehensive coverage without any of the claims being falsifiable. A report full of findings looks identical whether the tool tested one hundred endpoints or eleven.
There are two ways to arrive at coverage, and the difference is architectural rather than a matter of quality.
Instructed coverage means the system was told to be thorough. Whatever a model chose to pursue on a given run is what got covered. It may well be excellent, and on any single engagement it may beat the alternative. But it cannot be enumerated afterward, because there was never a list. Nothing was skipped, exactly, because nothing was ever scheduled.
Structural coverage means each vulnerability class is a discrete module the system executes as a unit of work. Coverage becomes a property of construction rather than of intent. The consequence is the useful part: afterward, the system can say which classes ran against which endpoints, including every one that ran and returned nothing.
VLKN is built structurally, with 90+ modules, and it is the position we would argue for regardless of who built the product. Given a choice between a tool that is thorough and a tool that can prove what it attempted, security programs should take the second one, because the first cannot be audited, cannot be trended over time, and cannot answer the only question that matters after an incident: was this tested.
The other vendors on this page do not declare which model they use. Several publish adjacent artifacts, including logs of actions taken and documentation of attack paths reached, and those are useful things. They are not the same artifact. A record of what happened is not a record of what was attempted and came back clean.
So this is the question to put to every vendor here, including us: is your coverage structural or instructed, and if it is structural, show me the per-engagement record with the negative results in it. A vendor with a structural system can produce that in an afternoon. A vendor without one will explain why the question is the wrong question.
Broken object level authorization is the class most likely to be missing from an AI pentest report, and the reason is structural rather than a matter of tuning.
To find that user A can read user B's invoice, a tester has to hold credentials for both A and B, know which objects belong to which principal, and try the cross-access deliberately. A tool testing as one identity never authenticates as B, so the vulnerability is not merely missed, it is unreachable. The same is true for privilege escalation across roles and for tenant isolation failures in multi-tenant applications.
This is a fair question to put to every vendor on this page including us, and it has a checkable answer. Ask how many accounts the system is provisioned with, whether it is told which roles those accounts hold, and ask to see a finding where the proof is a request signed as one user returning another user's data. If the answer is one account, the whole authorization family is out of scope regardless of what the coverage summary says.
You do not have to take a competitor's word for where agentic testing currently falls short. Bugcrowd publishes it on their own Savant Pathseeker product page, in the cons section of their FAQ: capabilities are still developing for complex business logic flaws, exploit chaining, and zero-days in dynamic applications. They restate it elsewhere on the same page, noting that autonomous-only tools can miss complex business logic and exploit chains, and that those still require human adversarial creativity and judgment.
That is an unusually honest thing for a vendor to print, and it is the most credible statement available about where this category actually sits. It is also the right question to bring to every vendor evaluation, including ours: not whether the tool is good, but which classes it structurally reaches and which it does not.
VLKN's answer is the module registry. 90+ modules across the engagement lifecycle, each vulnerability class its own execution path, multi-account and role-aware so authorization classes are reachable. Ask us for the coverage record on any engagement and you will get the list of what ran, including everything that ran and found nothing.
A fair number of people comparing vendors on this page do not have an application estate. They have a client book, more testing demand than testers, and an internal argument about whether to build their own tooling.
Building one is not hard. Maintaining one is a company. A single agent that finds a SQL injection is a weekend project. Ninety-plus vulnerability modules maintained against a web that changes underneath them, across every client environment, is a permanent engineering commitment that competes with billable work for the same people.
The alternative is running someone else's engine under your own brand. VLKN supports white-label delivery: your name on the report, your client relationship, your pricing. You review nothing you do not want to review, because reports come out finished rather than as raw findings for someone on your bench to triage.
Worth knowing that we are not the only option here. Terra runs a partner program aimed at service providers wanting to in-source an agentic pentesting program, and Bugcrowd is bringing an agentic product to an existing platform. Evaluate all of them on the same two questions: what does your team have to do after a test finishes, and where can the engine actually run.
That is a different conversation from the one above, and it starts the same way. Send a note with a target off your client book and we will run it, free, so you can judge the output before anything else gets discussed.
LLM-driven agents performing reconnaissance, attack surface discovery, exploitation and reporting against a target. It differs from scanning in that agents attempt exploitation and validate what they find rather than matching signatures. It differs from manual pentesting in that software rather than a person performs the execution.
Auditors accept a point-in-time report that documents scope, methodology, findings and severity. What matters is the deliverable format, not who performed the testing. A pile of ad-hoc bug reports is not usually accepted as a pentest report. VLKN's free tier includes an executive summary and, where nothing is found, an attestation-style no-findings report.
Instructed coverage means the system was told to be thorough, and whatever the model pursued on that run is what got covered. Structural coverage means each vulnerability class is a discrete module the system executes, so afterward it can enumerate which classes ran against which endpoints, including the ones that returned nothing. Both can produce good findings. Only one produces a record you can audit. Ask any vendor which they are, and ask to see the record.
Structurally, yes. An annual pentest describes the application as it existed on one day. Most teams ship weekly, so the report starts decaying the moment it lands. Continuous testing re-runs against changes as they happen. Almost every vendor on this page now offers both, so this is not a point of differentiation between them. It is a point of differentiation against a once-a-year human engagement, and it is worth asking what a vendor's continuous mode actually re-tests: the whole surface again, or only what changed.
Both, depending on the vendor, which is why search results for the phrase are confusing. Vendors like VLKN, Novee, Terra and XBOW use AI agents to test your applications. Bugcrowd's AI Pen Test uses human testers to assess your LLM applications for prompt injection and related issues. Check which one a vendor means before comparing prices.
Some can, and every vendor claims to. The check is the same as for authorization: ask to see a finding where the evidence is a sequence of requests that abuse an intended workflow, with the raw traffic attached. Business logic findings are hard to fake and easy to verify.
It carries the same class of risk a human pentester carries against production, which is real and non-zero. Ask any vendor for their stop conditions, their exclusion list, and whether destructive or state-changing actions require approval. A vendor who says the risk is zero is selling.
Most of this category quotes rather than publishes, typically per test, per application, or as an annual continuous subscription. VLKN publishes a free tier that delivers a complete point-in-time web application pentest report at no cost, no credit card, no company size limit.
Nobody in this category can prove a claim on a page, including us. Point VLKN at something you own and read what comes back. It costs nothing.
GET A FREE PENTEST