← All projects

Metal Plate Defect Inspection

Deep learning pipeline that replaces manual inspection with 95.5% accuracy

Role
Project Intern, MathWorks
Year
2026
Stack
  • MATLAB
  • ResNet-18
  • Transfer Learning
  • Computer Vision
Metal Plate Defect Inspection preview

Overview

During a project internship with MathWorks, I led development of a metal plate inspection system in MATLAB with teammates Diego Hidalgo and Enoch Ho. It sorts plates from the MPDD dataset into four categories (good, scratches, major rust, and total rust) and turns that label into a final pass/fail decision, replacing manual inspection.

Every photo takes two paths. Image processing shows where a defect might be, and a trained ResNet-18 decides what it is.

Select a stage to jump to its details.

Preprocessing

Turn a raw photo into a readable map of possible defects.

01Clean up the image

Each photo is standardized to 512 × 512 pixels, the plate is separated from the background, and uneven lighting is flattened with imflatfield so shadows and bright spots don’t read as defects.

Two copies move forward: a color copy for finding rust, and a grayscale copy with boosted contrast so scratches and texture stand out.

Flattening the lighting left a glowing halo around each plate’s edges. Instead of fighting the halo, the detectors ignore the background and outer edges and only look inside the plate.

32 color photos of metal plates after lighting correction, eight each of good, major rust, total rust, and scratches(opens full size in a new tab)
Color copies after lighting correction. Rows, top to bottom: good, major rust, total rust, scratches.
The same 32 plates in grayscale with boosted contrast(opens full size in a new tab)
Grayscale copies with boosted contrast, where scratches and surface texture stand out.
Back to the pipeline

02Find defect evidence

Rust and scratches look different, so each gets its own detector:

  • Rust, by color. The image is converted to HSV, and hue and saturation thresholds, tuned by trial and error, pick out rust.
  • Scratches, by shape. The grayscale copy goes through edge detection to find scratch lines.

defectEvidence returns a rust mask, a scratch mask, and a combined evidence mask for the plate.

Black and white defect masks for 32 plates. Good plates are mostly black; total rust plates are mostly white.(opens full size in a new tab)
Flagged pixels for the 32 sample plates. White means flagged: good plates stay mostly black, and totally rusted plates light up.
Back to the pipeline

03Overlay and baseline decision

Flagged pixels are painted red on the photo, so every decision comes with visible evidence. extractMetrics then measures the area ratio, the share of the image that was flagged, and decideRules turns it into a first pass/fail call.

The baseline made the right call on 31 of 32 sample plates. Its one miss was a good plate with harsh glare, which the edge detector read as scratches. More damage means a higher ratio, which makes the area ratio a strong, consistent signal of a plate’s condition.

32 plates with red defect overlays, each titled with its category, area ratio, and pass or fail decision(opens full size in a new tab)
Red overlay, area ratio, and baseline decision for each sample plate.
Area ratio of each sample plateEach dot is one plate. More damage flags more pixels, so the categories separate cleanly. The highlighted dot is the good plate that glare pushed into a FAIL.
00.050.100.150.200.25goodgood plate · area ratio 0.0163 · baseline PASSgood plate · area ratio 0.0137 · baseline PASSgood plate · area ratio 0.0109 · baseline PASSgood plate · area ratio 0.0090 · baseline PASSgood plate · area ratio 0.0083 · baseline PASSgood plate · area ratio 0.0099 · baseline PASSgood plate · area ratio 0.0144 · baseline PASSgood plate · area ratio 0.0852 · baseline FAILscratchesscratches plate · area ratio 0.0559 · baseline FAILscratches plate · area ratio 0.0614 · baseline FAILscratches plate · area ratio 0.0466 · baseline FAILscratches plate · area ratio 0.0399 · baseline FAILscratches plate · area ratio 0.0457 · baseline FAILscratches plate · area ratio 0.0880 · baseline FAILscratches plate · area ratio 0.0290 · baseline FAILscratches plate · area ratio 0.0448 · baseline FAILmajor rustmajor rust plate · area ratio 0.1557 · baseline FAILmajor rust plate · area ratio 0.1178 · baseline FAILmajor rust plate · area ratio 0.1292 · baseline FAILmajor rust plate · area ratio 0.1173 · baseline FAILmajor rust plate · area ratio 0.1988 · baseline FAILmajor rust plate · area ratio 0.2038 · baseline FAILmajor rust plate · area ratio 0.0976 · baseline FAILmajor rust plate · area ratio 0.0897 · baseline FAILtotal rusttotal rust plate · area ratio 0.2393 · baseline FAILtotal rust plate · area ratio 0.2275 · baseline FAILtotal rust plate · area ratio 0.2412 · baseline FAILtotal rust plate · area ratio 0.2419 · baseline FAILtotal rust plate · area ratio 0.2399 · baseline FAILtotal rust plate · area ratio 0.2237 · baseline FAILtotal rust plate · area ratio 0.2313 · baseline FAILtotal rust plate · area ratio 0.2186 · baseline FAILglare, called FAIL
View as table
CategoryArea ratio rangeBaseline calls
good0.0083–0.0163, plus one glare outlier at 0.08527 PASS, 1 FAIL
scratches0.0290–0.08808 FAIL
major rust0.0897–0.20388 FAIL
total rust0.2186–0.24198 FAIL
Back to the pipeline

Training

Teach a pretrained network to recognize the four plate categories.

04Train the AI classifier

The dataset holds only 151 images, so training a network from scratch would overfit. Instead we used transfer learning: ResNet-18, pretrained to recognize 1,000 kinds of everyday objects, with its last two layers replaced by a new four-category head.

  • Split: 70% training, 15% validation, 15% testing.
  • Augmentation: random rotations up to 15°, horizontal and vertical flips, and shifts of up to 10 pixels add variety and simulate parts moving on a factory line.
  • Training: SGDM with a learning rate of 1e-4, mini-batches of 16, and up to 40 epochs, with validation every 5 iterations to stop before overfitting.

Early runs reached only about 80% accuracy, and the model kept confusing scratches with good plates and major rust with total rust. More augmentation and 40 epochs instead of 5 lifted it to 95–100%, depending on the run. We trained the same configuration 10 times and kept the best network.

MATLAB training progress: accuracy rising to near 100 percent and loss falling to near zero over 240 iterations(opens full size in a new tab)
One example training run: accuracy (top) passes 90% within the first few epochs while loss (bottom) falls toward zero. 40 epochs and 240 iterations took about 4 minutes on a single CPU.
Back to the pipeline

Results

Run held-out plates through the full hybrid system, then make the photos harder.

05Test the full inspection

inspectPart runs a plate through both paths. It returns the AI’s category and confidence, the red evidence overlay, the area metrics, and the baseline’s pass/fail call.

On 22 held-out test plates, the AI classified 21 of 22 correctly, for 95.5% accuracy, and it never failed a good plate: a 0% false-reject rate. The one miss was a scratched plate labeled good at 68.2% confidence. The baseline and the AI disagreed on 3 plates, and the AI was right on 2 of them.

The test split was randomized and the saved model came from an earlier run, so some test images may overlap with training data. These numbers are likely slightly optimistic.

AI predictions on the 22 test platesRows are each plate's actual category; columns are what the AI predicted. Everything on the diagonal is correct. The one miss was a scratched plate predicted as good.
Predicted
Actualgoodmajor rustscratchestotal rustCorrect
good12000100%
major rust0200100%
scratches104080%
total rust0003100%
22 test plates with red overlays, each titled with its actual label, the AI prediction, and confidence(opens full size in a new tab)
Every test plate with its actual label, the AI's prediction, and confidence. Green titles are correct; the orange title is the one miss.
Back to the pipeline

06Stress test with distortions

Factory cameras aren’t perfect, so the same 22 test plates were rerun under six distortions: Gaussian blur at σ = 2 and σ = 4, Gaussian and salt-and-pepper noise, and low and high contrast.

  • Noise didn’t hurt. Accuracy held at 100% under both noise conditions.
  • Low contrast did the most damage. Simulating a dark scene dropped accuracy to 86.4% and let 3 of the 10 defective plates pass.
  • No false alarms. No condition caused a good plate to fail.
AI accuracy under each distortionThe same 22 test plates under every condition, on a full 0–100% scale. The highlighted bar is the weakest condition.
View as table
ConditionAccuracyGood plates failedDefective plates passed
Clean (baseline)95.5%0 of 121 of 10
Blur, σ = 295.5%0 of 121 of 10
Blur, σ = 490.9%0 of 122 of 10
Gaussian noise100%0 of 120 of 10
Salt-and-pepper noise100%0 of 120 of 10
Low contrast86.4%0 of 123 of 10
High contrast90.9%0 of 122 of 10
Four sample plates shown clean and under blur, Gaussian noise, salt and pepper noise, low contrast, and high contrast(opens full size in a new tab)
How each distortion changes four sample plates, from clean (top row) to high contrast (bottom row).
Back to the pipeline

Limitations and next steps

  • Small dataset. The metal plate set has only 151 images. Augmentation helped, but a network trained on so little data always risks overfitting.
  • Clean backgrounds. Every photo has a uniform backdrop, which made classification easier than it would be on a real production line.
  • Next: capture plates in real factory conditions to grow the dataset, add defect types like dents and cracks, and deploy the pipeline to an industrial camera to test real-time speed.

Dataset: MPDD (Jezek et al., 2021).