Everyone leads with a big number
Read the front page of any tool in application security and you will find a percentage, or a count of checks, sitting where the evidence should be. It is the first thing a buyer is given and the last thing anybody audits.
We could produce one this week. Pick the test applications, pick the bug classes, tune the engine against that set until the score looks handsome, publish the score, mention neither the set nor the tuning. Nothing in that process is illegal and most of it is normal. It is also worthless, because a number is only evidence if somebody else could have got a different one. When the vendor writes the exam, marks the paper and publishes the grade, the result tells you about the vendor's marketing team and nothing about your application.
Which is a decent reason to be sceptical about everybody else's number too, including any we publish. So the useful question is not what the number is. It is whether the method is written down well enough for you to disagree with it.
What we are doing instead
The benchmark runs against four sets of targets, and none of them are ours.
Juice Shop
The OWASP teaching application. Node, deliberately riddled, thoroughly documented, so every miss is checkable against a published answer key rather than against our word.
WebGoat
OWASP's Java equivalent. A different language and a different framework, to show the method is not quietly tuned to one stack.
DVWA
The PHP standard. Old, blunt and universally understood, which makes it a fair floor and an easy one to argue about.
Real applications
Open-source projects carrying known CVEs, where nobody wrote the bugs in order for them to be found. This is the set that actually matters.
Four numbers, not one
For every application in the set we publish the detection rate, the false-positive rate, the wall-clock duration and the dollars per scan. Alongside them goes the method: the versions, the configuration, the roles we set up, the answer key, and what we counted as a hit. Anyone who wants to run it themselves has what they need, and anyone who thinks we marked ourselves generously can say exactly where.
Those four belong together. A detection rate on its own hides the cost of reading the report: a tool that finds everything and flags four hundred maybes has moved the work rather than done it. Published together, the four let you compare like for like, which is the only comparison worth making.
The misses go in. Where the answer key says there is a bug and we did not break it, that is a line in the results with our name against it. If the numbers are ugly, you will see them anyway. The second run publishes 2 March, diffed against the first, so the direction of travel is public as well as the position.
This is not bravery, it is arithmetic
People read publishing the misses as a moral stance. It is mostly a commercial one, and the maths is not complicated.
A miss found in a pilot is a refund
If we sell you a number and your own application quietly contradicts it in week two, we have bought thirty days of your goodwill with a claim we could not keep. A gap you knew about before you signed is a scoping decision. The same gap found afterwards is a broken promise, and those cost more than lost deals.
The buyer we would disappoint self-selects out
Some teams need signature breadth, network and infrastructure coverage, or three references in their own industry today. We are not that, and finding out early is cheaper for both sides than finding out in month three. A published gap is the cheapest qualification tool we have.
It fixes the number in place
Once a result is public with its method attached, we cannot improve it by rewording it. The only way the March run looks better than the November one is if the product got better. That is a useful thing to do to your own engineering team.
The claim survives contact
Our whole argument is that a finding without proof is a suspicion. A company that makes that argument while publishing an unaudited accuracy figure is not making an argument, it is running an advertisement. Receipts or nothing has to apply to our own marketing first.
Today
What we are missing right now
The benchmark is dated, not done. Until it publishes, here is the honest state of things, and none of it is a surprise we are saving for your third week.
No published proof yet
Every claim on this site is one we can demonstrate live on your application. That is a different kind of evidence from a number, and we are not pretending otherwise.
No reference customers yet
We are early and we say so. The first customers are signing now. If you need three references in your industry today, we are not there.
Young operations
Off-box backups, a timed restore drill and a public status page land 18 December. Time and spend caps per scan land 13 November. Today we run the platform closely by hand.
Gaps with dates
Confirmed-only mode lands 19 October, prove-the-fix 16 November, the multi-client console 27 November, business-logic abuse 4 January. Dates, not adjectives.
Gaps without dates
Ticket sync with Jira and GitHub Issues is not built, and today you paste the reference by hand. Single sign-on beyond Okta, self-serve signup and data residency regions are on the list.
Narrower on purpose
No network or infrastructure scanning, the pilot is paid rather than free, and no attempt to win on how many checks we can count. Those are choices, and they are the wrong choices for some buyers.
If any of those dates slip, we will say so the week it happens, here. The full list of what we have not proved →
Judge us on your own application.
The benchmark is dated 3 November, and a date is not a result. Until it lands, the only evidence worth anything to you is what happens when we point this at your code while you watch.
Or start with a $199 pilot on one application: thirty days, success criteria agreed before day one, credited against the annual if you convert.