Case study · Data Labeling, AI Model Refinement, AI Analytics
Labeling 4,012 vehicle photographs, and publishing the dataset
How Bucepha Triage built the labeled damage dataset behind Bucepha Intelligence, published it on Hugging Face, and keeps refining it from what dealerships correct.
- 4,012
- Vehicle photographs labeled
- 6
- Fields on every label
- 2
- Models reading every photograph
The problem
Bucepha Intelligence reads auction photographs and prices the damage it finds. A model that does that is only as good as what it was shown, and a generic car damage dataset can say a bumper is dented without saying where, how badly, or what the repair came to.
So the client was our own product, and the job was a labeled dataset that a buyer could rely on: every photograph read, every defect located and described in body shop terms, and every label checked by a person.
How the labeling works
Every photograph goes through the same five steps. Two models read it independently, and a person decides.
01
Localize
Bucepha's own detection model, published on Hugging Face, proposes where the damage is on the panel.
02
Describe
A pretrained language model reads the same photograph on its own and writes what the damage is and how bad, in the words a body shop prices from.
03
Review
A person checks the label. Where the two models disagree the disagreement is kept, not resolved by whichever sounded surer.
04
Join
Where a repair was carried out, the label is joined to the repair order and the invoice that settled it.
05
Publish
The labeled photographs are published as a dataset on Hugging Face, so the work can be inspected rather than taken on trust.
What a label carries
A box around a dent is the start of a label, not the whole of one. Each defect is recorded with six fields, so the model learns what a body shop would need to price it.
- Defect type
- Dent, scratch, repaint, crack
- Location
- Panel and side, e.g. front bumper, left
- Severity
- On a scale of 1 to 5
- Component
- The part, e.g. bumper cover
- Repair category
- Refinish, repair or replace
- Confidence
- How sure the reading is
Published on Hugging Face
The labeled photographs and the detection model trained on them are published under Bucepha on Hugging Face, so the work can be inspected rather than taken on trust. What stays private is the other half: the corrections dealerships make and the invoices they are joined to, which are not published, sold, or used to train anybody else's model.
The dataset on Hugging Facehuggingface.co/buckets/shalin-code/bucepha-cardamage-detectionHow we keep refining it
Corrections
When a dealership corrects a finding in the panel, the correction is recorded and becomes a label for the next round.
Disagreement
Photographs where the two models disagree are the first reviewed, because they are where the model is least sure.
Invoices
Each settled repair checks the estimate that came before it, so pricing is answerable to what was actually paid.
The dataset is not finished and is not meant to be. Every correction, disagreement and invoice is a reason to label again, and the count grows with the work.
Need a dataset built like this?
- Labeling against a written guide
- Second-annotator review
- Evaluation sets that match production