A YOLOv8 object detector for hot-rolled steel surface defects. It finds and labels six defect types on the NEU-DET dataset: crazing, inclusion, patches, pitted surface, rolled-in scale and scratches.
This repo holds the training and evaluation pipeline, the trained weights, the exact data split and the results. A web dashboard around the model (inspection logs, defect analytics, PDF reports) is being rebuilt and is not part of this repo yet.
Released model: weights/best.pt (YOLOv8s, 11.1M parameters, 22.5 MB). Numbers below come from re-running yolo val on this exact file.
| Set | Images | Instances | Precision | Recall | mAP50 | mAP50-95 |
|---|---|---|---|---|---|---|
| Validation | 370 | 900 | 0.704 | 0.694 | 0.726 | 0.378 |
| Held-out test | 30 | 53 | 0.800 | 0.597 | 0.771 | 0.468 |
The held-out images were never used for training or for choosing the checkpoint. There are only 30 of them, so treat that row as a rough check, not a stronger result than the validation row.
Speed on an RTX 3050 6GB laptop GPU: about 14 to 15 ms of model inference per image, about 18 ms end to end (batch of 8 during validation, 640 px input).
| Class | Images | Instances | Precision | Recall | mAP50 | mAP50-95 |
|---|---|---|---|---|---|---|
| crazing | 67 | 160 | 0.512 | 0.348 | 0.406 | 0.144 |
| inclusion | 82 | 207 | 0.775 | 0.763 | 0.792 | 0.369 |
| patches | 79 | 218 | 0.772 | 0.844 | 0.884 | 0.527 |
| pitted_surface | 54 | 76 | 0.926 | 0.750 | 0.876 | 0.540 |
| rolled-in_scale | 54 | 120 | 0.495 | 0.579 | 0.534 | 0.227 |
| scratches | 70 | 119 | 0.747 | 0.882 | 0.865 | 0.460 |
| Class | Instances | mAP50 |
|---|---|---|
| crazing | 6 | 0.443 |
| inclusion | 14 | 0.840 |
| patches | 10 | 0.895 |
| pitted_surface | 7 | 0.900 |
| rolled-in_scale | 13 | 0.683 |
| scratches | 3 | 0.863 |
Scratches has 3 instances and crazing has 6, so those two rows say very little. Plots for the training run are in results/train-8/ (confusion matrix, precision/recall/F1 curves, results.png).
- Crazing and rolled-in scale are weak. Validation mAP50 is 0.41 and 0.53. On the held-out set, recall for crazing is 0.17.
- Boxes are loose. mAP50 is 0.726 but mAP50-95 is 0.378, so the detector finds most defects but does not localise them tightly.
- The validation set picked the checkpoint. Early stopping and best-epoch selection used it, so the validation score is slightly optimistic. Only the held-out set is untouched.
- One run per configuration. No repeated seeds and no cross-validation, so I have no estimate of run-to-run variation.
- It does not transfer to real-world photos. I tried three stock photos of heavily scratched metal. At confidence 0.25 the model detected nothing. At confidence 0.05 it produced a few boxes near the image edges, all labelled "patches", none "scratches". Three images is an observation, not a measurement. My guess is that NEU-DET scratches are isolated lines on a clean background, while dense real scratching looks more like the "patches" texture, and that augmenting the same lab images cannot fix that. Testing this needs labelled real-world images.
| Run | Model | Data | mAP50 | Notes |
|---|---|---|---|---|
| train-4 | YOLOv8n | old 800 / 200 split | 0.740 | Earlier baseline on a 200-image validation set. Different split, not comparable to train-8. Weights are not in this repo. |
| train-6 | YOLOv8m | about 800 images | 0.693 | Overfit. |
| train-7 | - | - | - | Crashed in epoch 3 (see Environment). No checkpoint. |
| train-8 | YOLOv8s | 1370 / 370 / 30 | 0.726 val, 0.771 held-out | Released. Stopped early: best epoch 111, 161 epochs total, 3.2 hours. |
Settings for train-8 are in results/train-8/args.yaml. The main ones: imgsz=640, batch=8, patience=50, degrees=5, scale=0.2, mosaic=0.5, hsv_h=0.02, hsv_s=0.5, hsv_v=0.5, workers=4.
I changed model size, split and augmentation between runs, so I can't say which change caused which difference in the table.
Python 3.10, ultralytics 8.4.116, torch 2.5.1 (CUDA 12.1 build), Windows.
pip install -r requirements.txt
Two problems I hit, in case you do too:
opencv-python5.x crashes ultralytics' colour augmentation (cvtColorerror).requirements.txtpins<5.workers=8ran out of memory in the dataloader on Windows.workers=4was stable.
The dataset is not in this repo. Get NEU-DET (Song and Yan, Northeastern University) yourself. The copy I used had 1,770 images (commonly cited as 1,800). Merge its train and validation folders into one IMAGES/ folder and one ANNOTATIONS/ folder (Pascal VOC XML) in the repo root.
python src/convert.py
This converts the labels to YOLO format and rebuilds the same split I trained on, using the file lists in splits/. It writes yolo_dataset/ and held_out_test/, and refuses to run if either already exists.
yolo detect train data=yolo_dataset/dataset.yaml model=yolov8s.pt imgsz=640 epochs=300 patience=50 batch=8 device=0 degrees=5 scale=0.2 mosaic=0.5 hsv_h=0.02 hsv_s=0.5 hsv_v=0.5 workers=4
Training on a GPU is not bit-for-bit repeatable, so expect small differences.
yolo detect val model=weights/best.pt data=yolo_dataset/dataset.yaml split=val
yolo detect val model=weights/best.pt data=yolo_dataset/dataset.yaml split=test
yolo detect predict model=weights/best.pt source=path/to/images conf=0.25
weights/best.pt released model (train-8)
splits/ file lists for train / val / held-out test
src/convert.py VOC XML to YOLO conversion and split
results/train-8/ args, metrics, curves, confusion matrices, sample predictions
requirements.txt
- Rebuild the inspection dashboard (detection logs, defect distribution, severity, report export).
- Collect labelled real-world steel images to test the domain gap properly.
- Repeat training with several seeds to get a spread on the scores.
NEU-DET: K. Song and Y. Yan, "A noise robust method based on completed local binary patterns for hot-rolled steel strip surface defects", Applied Surface Science, 285, 2013.
Code and weights are released under AGPL-3.0, the same license as Ultralytics YOLO, which they build on.
Author: Om Bagal, COEP Technological University, Pune.

