Platform

Platform overview How it works Authorization testing Evidence & reports Private scanning Integrations

Solutions

Security agencies Product teams Regulated industries Partner programme

Learn

Blog Knowledge hub Compare

Resources

Pricing Documentation FAQ Security & data What we haven’t proved

Company

About Contact Careers Sign in to the platform Start a $199 pilot

Home / Blog / Publish the misses

The benchmark

Why we publish what we miss.

A number you chose the exam for is not evidence. On 3 November we publish the method, the results and the misses, so you can check our marking or run it yourself.

Cyberlop Labs · 20 September 2026

Everyone leads with a big number

Read the front page of any tool in application security and you will find a percentage, or a count of checks, sitting where the evidence should be. It is the first thing a buyer is given and the last thing anybody audits.

We could produce one this week. Pick the test applications, pick the bug classes, tune the engine against that set until the score looks handsome, publish the score, mention neither the set nor the tuning. Nothing in that process is illegal and most of it is normal. It is also worthless, because a number is only evidence if somebody else could have got a different one. When the vendor writes the exam, marks the paper and publishes the grade, the result tells you about the vendor's marketing team and nothing about your application.

Which is a decent reason to be sceptical about everybody else's number too, including any we publish. So the useful question is not what the number is. It is whether the method is written down well enough for you to disagree with it.

What we are doing instead

The benchmark runs against four sets of targets, and none of them are ours.

Juice Shop

The OWASP teaching application. Node, deliberately riddled, thoroughly documented, so every miss is checkable against a published answer key rather than against our word.

WebGoat

OWASP's Java equivalent. A different language and a different framework, to show the method is not quietly tuned to one stack.

DVWA

The PHP standard. Old, blunt and universally understood, which makes it a fair floor and an easy one to argue about.

Real applications

Open-source projects carrying known CVEs, where nobody wrote the bugs in order for them to be found. This is the set that actually matters.

4Numbers published
3 NovPublication date
YesMisses included
2 MarNext run, diffed

Four numbers, not one

For every application in the set we publish the detection rate, the false-positive rate, the wall-clock duration and the dollars per scan. Alongside them goes the method: the versions, the configuration, the roles we set up, the answer key, and what we counted as a hit. Anyone who wants to run it themselves has what they need, and anyone who thinks we marked ourselves generously can say exactly where.

Those four belong together. A detection rate on its own hides the cost of reading the report: a tool that finds everything and flags four hundred maybes has moved the work rather than done it. Published together, the four let you compare like for like, which is the only comparison worth making.

The misses go in. Where the answer key says there is a bug and we did not break it, that is a line in the results with our name against it. If the numbers are ugly, you will see them anyway. The second run publishes 2 March, diffed against the first, so the direction of travel is public as well as the position.

This is not bravery, it is arithmetic

People read publishing the misses as a moral stance. It is mostly a commercial one, and the maths is not complicated.

A miss found in a pilot is a refund

If we sell you a number and your own application quietly contradicts it in week two, we have bought thirty days of your goodwill with a claim we could not keep. A gap you knew about before you signed is a scoping decision. The same gap found afterwards is a broken promise, and those cost more than lost deals.

The buyer we would disappoint self-selects out

Some teams need signature breadth, network and infrastructure coverage, or three references in their own industry today. We are not that, and finding out early is cheaper for both sides than finding out in month three. A published gap is the cheapest qualification tool we have.

It fixes the number in place

Once a result is public with its method attached, we cannot improve it by rewording it. The only way the March run looks better than the November one is if the product got better. That is a useful thing to do to your own engineering team.

The claim survives contact

Our whole argument is that a finding without proof is a suspicion. A company that makes that argument while publishing an unaudited accuracy figure is not making an argument, it is running an advertisement. Receipts or nothing has to apply to our own marketing first.

Today

What we are missing right now

The benchmark is dated, not done. Until it publishes, here is the honest state of things, and none of it is a surprise we are saving for your third week.

No published proof yet

Every claim on this site is one we can demonstrate live on your application. That is a different kind of evidence from a number, and we are not pretending otherwise.

No reference customers yet

We are early and we say so. The first customers are signing now. If you need three references in your industry today, we are not there.

Young operations

Off-box backups, a timed restore drill and a public status page land 18 December. Time and spend caps per scan land 13 November. Today we run the platform closely by hand.

Gaps with dates

Confirmed-only mode lands 19 October, prove-the-fix 16 November, the multi-client console 27 November, business-logic abuse 4 January. Dates, not adjectives.

Gaps without dates

Ticket sync with Jira and GitHub Issues is not built, and today you paste the reference by hand. Single sign-on beyond Okta, self-serve signup and data residency regions are on the list.

Narrower on purpose

No network or infrastructure scanning, the pilot is paid rather than free, and no attempt to win on how many checks we can count. Those are choices, and they are the wrong choices for some buyers.

If any of those dates slip, we will say so the week it happens, here. The full list of what we have not proved →

Judge us on your own application.

The benchmark is dated 3 November, and a date is not a result. Until it lands, the only evidence worth anything to you is what happens when we point this at your code while you watch.

Or start with a $199 pilot on one application: thirty days, success criteria agreed before day one, credited against the annual if you convert.