What “85% pothole detection” actually means for your audit
How to read a detection accuracy figure: why 85% and 97% measure different things, why the missing 15% matters less than you think, and what it hides.
Every vendor in this space, ourselves included, publishes an accuracy figure. Ours are these:
Those numbers are worth roughly nothing to you until you know what they are measuring. This post is our attempt to make our own figures harder to misread — including in the directions that are unflattering to us.
85% and 97% are not the same kind of number
The two headline figures measure different tasks, which is why one is so much higher. Detection must find a defect and localise its boundary against everything that resembles it. Classification only sorts a known stretch of surface into one of a few broad categories. High nineties is routine for the second and hard for the first.
| Detection (85%) | Classification (97%) | |
|---|---|---|
| The question | Is there a pothole here, and where exactly? | Which condition category is this surface? |
| Answer space | Every possible position and boundary in the frame | A handful of broad categories |
| Confounders | Oil stains, shadow pooling, patched repairs, wet tarmac | Few — the visual differences are broad |
| Ways to be wrong | Miss it, hallucinate it, or bound it badly | Pick the wrong category |
| Expected range | Mid-80s is a strong result | High nineties is routine |
The practical consequence: a classification score must never be read as though it were a detection score, and the two should never be averaged into one headline figure. So when you are comparing tools, the first question is not “what’s your accuracy?” It is “accuracy at what?” A supplier quoting a single number across every defect type is either simplifying for a brochure or has not measured carefully.
The 15% matters more than the 85%
Here is the part that gets glossed over. A detection rate of 85% means roughly one pothole in seven is not flagged. Whether that is acceptable depends entirely on what happens to the ones that are missed — and that depends on a property of the survey that has nothing to do with the model.
A manual inspection is a single pass. Anything the inspector walks past is gone until the next survey, which may be a year away. An automated survey is cheap enough to repeat, and defects do not move. A pothole missed on the June drive is very likely caught in July, and by then it is larger and easier to detect, not harder.
This is the real argument, and it is a frequency argument rather than an accuracy one. Eighty-five per cent monthly finds more of a network’s actual defects than a hypothetical hundred per cent once a year — and finds them earlier, which is when repair is cheap. Any comparison that holds survey frequency constant is measuring the wrong thing.
Severity is where the value concentrates
A flat count of defects is a poor planning instrument. What a works department needs is the shape of the distribution, because that is what determines whether a corridor needs patching, resurfacing, or nothing this quarter.
The sample audit published in our public console illustrates the point. Across a 12.4 km stretch it holds 128 detections spread over eight classes — potholes, alligator cracking, faded lane markings, longitudinal cracks, ravelling, kerb damage, crash barrier damage and signage defects — each carrying a severity grade and a confidence score.
The useful reading is not “128 defects”. It is that the high-severity subset is small, concentrated, and actionable this month, while the low-severity majority is a resurfacing conversation for next financial year. Those are two different budgets, and only a graded list separates them.
On that 128: it is a representative sample dataset built for the public demo, not the result of a specific customer engagement. It exists so the console can be explored without a login. Treat it as an illustration of the output format, not as a field measurement of any corridor.
Confidence scores are a filter, not a grade
Every detection carries a model confidence. The instinct is to treat low confidence as low quality and discard it, which is usually a mistake. Confidence is a threshold you tune to your tolerance: raise it and you get a shorter, cleaner list that misses more; lower it and you get a longer list with more false positives to triage.
Which direction is correct depends on what the list is for. Building a capital plan, you want high confidence and few distractions. Chasing a safety complaint on a specific stretch, you want to see everything and judge for yourself. The value of a per-detection score is that it lets one dataset serve both — which a pre-filtered PDF cannot.
What the numbers cannot tell you
No accuracy figure captures three things: what road network it was measured on, what overlap with the ground truth counted as a correct hit, and what lies below the surface. All three are worth asking any supplier about, because two vendors using different hit thresholds are not comparable and neither usually publishes theirs.
- What it was measured on. A model evaluated on the same kind of roads it was trained on will always flatter itself. Ours is trained on Indian road surfaces, which is exactly why we would expect it to underperform on a road network that looks nothing like them.
- What counts as a hit. Detection scores depend on how much overlap with the ground-truth boundary is required to call it correct. Two vendors using different thresholds are not comparable, and neither usually publishes theirs.
- What happens below the surface. Nothing visual detects subgrade failure or drainage voids. A perfect surface score says nothing about structural condition.
A better question than accuracy
The question that actually separates tools is not the headline percentage but can I check the answer? Every detection should trace to the frame it came from, at a coordinate you can drive to. When that holds, accuracy stops being a claim you must accept and becomes something you can verify on your own network.
Every detection should be traceable to the frame it came from, at a coordinate you can drive to. When that is true, accuracy stops being a claim you have to accept and becomes something you can audit on your own network — which is the only figure that was ever going to matter. That is also why our sample audit is public and unlocked: you can interrogate the output format before anyone asks you for a procurement decision.
Questions this raises
What does 85% pothole detection accuracy mean?
It means roughly one pothole in seven is not flagged on a single pass. Whether that matters depends on survey frequency: an automated survey is cheap enough to repeat, and a pothole missed in June is very likely caught in July, by which point it is larger and easier to detect.
Why is classification accuracy higher than detection accuracy?
They measure different tasks. Detection must find a defect and localise its boundary against oil stains, shadows and patched repairs. Classification only sorts a known stretch of surface into one of a few broad categories. High nineties is routine for the second and hard for the first.
How should I compare road inspection vendors on accuracy?
Ask what the figure measures, what road network it was measured on, and how much overlap with the ground truth counts as a correct hit. Two vendors using different thresholds are not comparable. The better question is whether every detection traces back to a frame you can check.
