Astra Autonomous Pentest vs Xbow - Benchmark Study

We didn't grade our own homework. We found 3.9x more.

Doyensec set the app and the rules. We ran Astra Autonomous Pentest on the same asset as XBOW and found
3.9x more real vulnerabilities. The wrong answers are on this page too.

27

Vulnerabilities found by
ASTRA

3.9×

More verified findings
ADVANTAGE

7

Vulnerabilities found by
XBOW
ASTRA
Vulnerabilities found by

27

ADVANTAGE
More verified findings

3.9×

XBOW
Vulnerabilities found by

7

Independent exam

Doyensec set the paper.
We wrote it without a cheat sheet.

Doyensec did that in June 2026, testing AI pentest tools against real, self-hosted software with a clean way to
separate a real vulnerability from noise. Astra Autonomous Pentest ran under those same conditions as XBOW.

Study parameters

Target application

Photoview v2.4.0

Methodology baseline

Doyensec, Jun 2026

Who graded findings

2 reviewers + tiebreak

Astra scan date

Jun 15, 2026

Anyone paid to say this

Nope. Independent.

THE FULL REPORT CARD

Every score, including the ones that sting

Coverage, precision, severity, all of it out in the open. If a number makes us look worse, we still wrote it down.

Study parameters

Astra

Xbow

Real vulnerabilities found

confirmed, exploitable

27

7

Vulnerability classes caught

11

5

False alarms, raw

lower is better

34

0

Precision

real ÷ total findings

57.4%

100%

Severity

Astra

Xbow

Critical

High

Medium

Low

Informational

Study parameters

Real vulnerabilities found

confirmed, exploitable

Vulnerability classes caught

out of total available

False alarms, raw

lower is better

Precision

real ÷ total findings

BY SEVERITY

Critical

High

Medium

Low

Informational

Our severities follow CVSS. XBOW's are Doyensec's adjusted calls. Astra's seven high-severity findings alone match XBOW's entire output.

WHERE THE COVERAGE CAME FROM

11 kinds of vulnerabilities.
XBOW answered 5.

Four showed up on both papers. Treat those as the safe bets, the vulnerabilities two different tools
flagged without comparing notes. Seven more, only Astra caught.

The safe bets

(both tools caught them)

SQL injection

Broken authorization (IDOR)

Information leakage

Security headers and cookie flags

where Astra went further

(only Astra AP caught them)

Missing rate limiting

Broken session management

GraphQL abuse and misconfig

Insecure TLS setup

Race condition (TOCTOU)

DNSSEC absence

API idempotency gap

For the record: XBOW flagged one class we didn't surface on this run, SSRF. We've since closed it.

Business logic depth

We didn't stop at the first right answer.

Astra

ways into the same
broken access control

XBOW

way into the same
broken access control

Astra kept working the problem and found five IDOR and BOLA variants across different GraphQL calls and permission checks. XBOW found one.

Catching all five (vs 1) is the difference between a scanner and Astra's AP that hunts the way an attacker does.

The part vendors leave off the report card

We chose coverage over a clean score.

A perfect score is easy if you leave the hard ones blank. Astra AP didn't; we answered everything Photoview
threw at it, then let the numbers speak for themselves.

Confirmed real

27

Total findings

61

False alarms

34

27

Confirmed real

61

Total findings

34

False alarms

A spotless 7 out of 7 can still leave 20 real vulnerabilities sitting in your app.

Raw score

57.4% precision

44.3% raw score after deduping

Not wrong answers

13 repeated

same question, answered multiple times

Missing headers

13 write-ups

same question, answered 13 times

Rate-limit gap

4 write-ups

one gap, phrased four ways

But, given the choice, we would rather hand you 27 real vulnerabilities with
some noise to sort than a clean 7-vuln report card that misses the 20.

HOW WE RAN IT

One shot. Same conditions as XBOW.

Four steps, matched to Doyensec's methodology.

Set the paper.

Photoview v2.4.0 on a fresh box of its own.

Sit the test.

Astra Autonomous Pentest, start to finish, without any hints or coaching.

Grade it fairly.

Two reviewers mark every finding. A third settles ties (same bar as the baseline).

Merge the repeats.

Answers with one root cause count once, the way a real report reads.

One thing worth flagging

We sat the exam on Railway, instead of the baseline's AWS box with nginx, and it had development mode
switched on. So we threw out the two CORS findings that flag caused, and flagged the lone TLS finding rather
than quietly banking the point.

The app underneath is the same on both sides.

This was one exam. Picture the whole semester.

Astra Autonomous Pentest chains attacks the way a real attacker would, then hands your team the fix,
rather than a 90-page inactionable report & a shrug. See what it does across your whole stack.
Get the report

The whole report card,
gold stars and red
pen-marks included.

Every real finding, every false alarm, the full method,
and the fine print. Drop your email, and we send it to
your inbox (without a sales-call ambush).

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Click here to update your cookies settings