Skip to content
Explainer

How AI detects road defects from ordinary dashcam video

How a windscreen camera becomes a defect list: how frames turn into detections, how each one gets a location, and why an ordinary dashcam is enough.

The question we get asked first, almost always in the same words: how can a camera on a windscreen possibly do what an inspection team does on foot? It is a fair thing to be sceptical about. This is the answer, without the marketing layer.

Step one: the camera is genuinely ordinary

There is no survey rig, no laser profiler, no roof-mounted sensor mast. A standard dashcam recording 1080p at 30 frames per second, mounted behind the windscreen of a vehicle already making the trip, is the whole capture apparatus. That constraint is deliberate rather than a compromise: specialist equipment is what makes conventional surveys rare events, booked months apart. Equipment you already own is what makes them routine.

The system also accepts a still image or a live camera feed, not only recorded video. What matters is that each frame carries a timestamp, because that is what lets a detection be tied back to a point on the earth later.

Step two: frames become detections

Video arrives as a sequence of images, and each one is passed through a detection model — an ensemble in the YOLO family, trained on more than 34,540 augmented images of Indian road surfaces specifically. That last word is doing real work. A model trained on European motorways has never seen an unsealed shoulder, a monsoon-scoured edge break, or the particular way an Indian bitumen surface ravels, and it will quietly miss all three.

The models run at 45 to 60 frames per second, which is what allows the vehicle to travel at ordinary traffic speed rather than crawling. Three separate concerns are being handled at once:

The three detection modules
ModuleWhat it looks forWho owns the fix
PavementPotholes, cracking, kerb condition — with severity and measurement for eachWorks / maintenance
InfrastructureLane and edge lines, kerb paint, barriers, roadside assets, road studs, rumble markings, speed breakersSafety furniture
SignageInformative, regulatory and warning signs, whether each is reflective — one asset record per physical signAsset inventory

Splitting the problem this way matters more than it looks. A pothole and a faded lane line are not variations of one thing; they belong to different maintenance budgets, different standards, and often different departments. Detecting them with one undifferentiated model produces a list nobody owns.

Step three: a detection gets measured

Knowing a pothole is present is the easy half. Knowing whether it is a patch job or a section failure is what a maintenance engineer actually needs, and that requires size in real-world units from a single camera with no stereo pair to work from.

Monocular depth estimation fills that gap: a second model infers relative distance across the frame, which lets the pixel extent of a defect be converted into an approximate physical dimension. Combined with the detection class, that is what produces a severity grade rather than a yes/no flag. Two hundred low-severity hairline cracks and six high-severity potholes are very different mornings for a works department, and an audit that cannot tell them apart is not much of an audit.

Step four: a detection gets a location

A defect list without coordinates is a document. A defect list with coordinates is a work order. Because GPS is captured automatically alongside the video, every detection inherits the position of the frame it came from — which means each one lands on a map at a specific chainage, and a crew can be sent to it without anyone writing down a landmark description.

This is also what makes the result auditable. A finding that can be traced back to a timestamped frame at a known coordinate is evidence. A finding recorded as “bad surface near the temple turning” is a memory.

Step five: it becomes something you can act on

The end of the pipeline is not a model output, it is five deliverables: annotated media showing what was found, a detection table carrying class, confidence, measurements and coordinates, an interactive map of the corridor, the processed video itself, and a CSV export for whatever system the findings need to live in next.

Turnaround is typically same-day, because nothing in the chain requires a lane closure, a booking window, or a second site visit. The four-step process page walks the same journey from the operator’s side, and the product overview lists what the platform detects class by class.

The honest limits

A windscreen camera sees the road surface, the roadside furniture, and the signs. It does not see what is underneath — subgrade failure, drainage voids, and structural capacity are not visual properties, and no amount of image data will infer them. Nor does it replace the engineering judgement in a formal safety audit; it replaces the weeks of manual data collection that precede one.

The useful framing is that this changes what an inspection costs, not what an inspector is for. When surveying a corridor stops being an expedition, you stop rationing it — and a network you can look at monthly is a fundamentally different thing to manage than one you look at once a year.

Questions this raises

Can a normal dashcam really detect road defects?

Yes. A standard dashcam recording 1080p at 30 fps is sufficient, because the detection work happens in software rather than in the sensor. The constraint is deliberate: specialist equipment is what makes conventional surveys rare events, while equipment you already own makes them routine.

How does the system measure the size of a pothole?

Monocular depth estimation infers relative distance across the frame, which lets the pixel extent of a defect be converted into an approximate physical dimension. That is what turns a detection into a severity grade, rather than just a flag that something is present.

What can a dashcam-based road survey not detect?

Anything not visible from the windscreen. Subgrade failure, drainage voids and structural capacity are not visual properties, so no amount of image data will infer them. It also does not replace the engineering judgement in a formal safety audit — only the data collection preceding one.

Keep reading