Define the defect and the decision
A vision model may classify visible defects, locate features or flag frames for review. Define what the label means and how it affects an inspection decision. A surface image cannot establish hidden structural capacity or remaining useful life by itself.
Build a representative dataset
- Use labelled images from different assets, materials, lighting conditions and cameras.
- Agree a defect taxonomy with experienced inspectors and measure annotation consistency.
- Separate training and test data by asset or inspection, so near-identical frames do not leak across sets.
- Record image quality and cases that require manual inspection.
Evaluate consequential errors
Measure missed critical defects and false alarms alongside overall accuracy. Evaluate by defect type and severity. Review how uncertainty is presented: an apparently precise confidence score may not be calibrated for unfamiliar assets.
Close the review loop
Start as a screening aid with human confirmation. Feed corrected labels back through a controlled dataset process. Connect observations to the asset register using stable identifiers, and preserve the original image and inspection context for audit.
Define the inspection task and the target label
An image model might identify a visible defect, classify condition or rank assets for review. These are different tasks. A visible coating defect is not necessarily a structural-capacity estimate, and the absence of a visible defect is not proof of sound internal condition. Connect each label to what the inspection method can actually observe.
Create an annotation guide with examples, borderline cases and a method for resolving disagreement. Use qualified reviewers where the label requires engineering interpretation. Record uncertainty instead of forcing every ambiguous image into a confident category.
Prevent near-duplicate images from inflating performance
An inspection can produce many adjacent frames of the same asset. Randomly distributing those frames between training and testing lets the model see almost the same surface on both sides. Group the split by asset or inspection campaign as appropriate to the deployment question.
Also test differences in lighting, camera, access angle, surface condition and environment. A model trained on clean, well-lit images can fail on the difficult images encountered in routine work. Retain a separate category for images whose quality is insufficient for a reliable judgement.
Measure the consequence of each error type
Missing a high-consequence defect and sending an acceptable asset for unnecessary review have different costs. Report sensitivity to important defect classes, false-review workload and performance on poor-quality images. Overall accuracy can look impressive when most assets have no defect.
For object detection or segmentation, define the spatial matching or overlap rule and how multiple predictions on one defect are counted. For asset-level prioritisation, aggregate image findings in a documented way so assets with more photographs do not automatically appear riskier.
Connect observations to asset management
Image output is one input to an engineering assessment alongside material, age, loading, service history, environment and consequence of failure. A prioritisation score should distinguish likelihood evidence from consequence. Do not convert a classification probability directly into a remaining-life estimate without a validated relationship.
Link results to stable asset IDs and inspection dates. Preserve the original images, review status and the model version. If the asset register has duplicate or missing IDs, the best image model can still direct a field crew to the wrong location. Use GIS checks before operational use.
Run an assisted inspection pilot
- Evaluate a representative held-out group of assets, including difficult images.
- Present model suggestions to inspectors with the original evidence visible.
- Record agreement, overrides, missed defects and time taken.
- Assess whether the workflow improves coverage or decision quality at an acceptable review burden.
- Define a fallback for unfamiliar assets, low-confidence output and poor image quality.
Begin with assistance to an accountable inspector. Maintenance decisions should retain the applicable engineering review and evidence requirements. Use the GIS workflow to maintain asset identity and the ML workflow for grouped evaluation and provenance.
Sources & further reading
- AI Risk Management Framework ↗National Institute of Standards and Technology · 2023 framework; 2024 generative AI profile
External source · Checked 24 September 2026
Source findings are distinguished from editorial interpretation. Apply current local criteria and project evidence when making engineering decisions.