Doyensec set the app and the rules. We ran Astra Autonomous Pentest on the same asset as XBOW and found
3.9x more real vulnerabilities. The wrong answers are on this page too.
Doyensec did that in June 2026, testing AI pentest tools against real, self-hosted software with a clean way to
separate a real vulnerability from noise. Astra Autonomous Pentest ran under those same conditions as XBOW.
Study parameters
Target application
Photoview v2.4.0
Methodology baseline
Doyensec, Jun 2026
Who graded findings
2 reviewers + tiebreak
Astra scan date
Jun 15, 2026
Anyone paid to say this
Nope. Independent.
Coverage, precision, severity, all of it out in the open. If a number makes us look worse, we still wrote it down.
Our severities follow CVSS. XBOW's are Doyensec's adjusted calls. Astra's seven high-severity findings alone match XBOW's entire output.
Four showed up on both papers. Treat those as the safe bets, the vulnerabilities two different tools
flagged without comparing notes. Seven more, only Astra caught.
The safe bets
(both tools caught them)
SQL injection
Broken authorization (IDOR)
Information leakage
Security headers and cookie flags
where Astra went further
(only Astra AP caught them)
Missing rate limiting
Broken session management
GraphQL abuse and misconfig
Insecure TLS setup
Race condition (TOCTOU)
DNSSEC absence
API idempotency gap
For the record: XBOW flagged one class we didn't surface on this run, SSRF. We've since closed it.
ways into the same
broken access control
way into the same
broken access control
Astra kept working the problem and found five IDOR and BOLA variants across different GraphQL calls and permission checks. XBOW found one.
Catching all five (vs 1) is the difference between a scanner and Astra's AP that hunts the way an attacker does.
A perfect score is easy if you leave the hard ones blank. Astra AP didn't; we answered everything Photoview
threw at it, then let the numbers speak for themselves.
A spotless 7 out of 7 can still leave 20 real vulnerabilities sitting in your app.
57.4% precision
44.3% raw score after deduping
13 repeated
same question, answered multiple times
13 write-ups
same question, answered 13 times
4 write-ups
one gap, phrased four ways
But, given the choice, we would rather hand you 27 real vulnerabilities with
some noise to sort than a clean 7-vuln report card that misses the 20.
Four steps, matched to Doyensec's methodology.
Set the paper.
Photoview v2.4.0 on a fresh box of its own.
Sit the test.
Astra Autonomous Pentest, start to finish, without any hints or coaching.
Grade it fairly.
Two reviewers mark every finding. A third settles ties (same bar as the baseline).
Merge the repeats.
Answers with one root cause count once, the way a real report reads.
One thing worth flagging
We sat the exam on Railway, instead of the baseline's AWS box with nginx, and it had development mode
switched on. So we threw out the two CORS findings that flag caused, and flagged the lone TLS finding rather
than quietly banking the point.
Every real finding, every false alarm, the full method,
and the fine print. Drop your email, and we send it to
your inbox (without a sales-call ambush).