All capabilities · Machine learning & data science

Build image models (detection/classification)

OpenCV + pretrained CNN/YOLO fine-tuning for a practical vision task.

~20 focused hoursintermediate
Explore 3 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Preprocess images with OpenCV (resize, normalize, augment, color space conversions)
  2. Fine-tune a pretrained CNN (ResNet/EfficientNet) for an image classification task
  3. Fine-tune or run inference with a detection model (YOLO) for object detection
  4. Prepare and validate an annotated image dataset (bounding boxes) in the right format
  5. Evaluate detection/classification with the right metrics (mAP, IoU, precision/recall)
  6. Handle real-world image quality issues (lighting, blur, occlusion, low resolution)
  7. Optimize a vision model for reasonable inference latency, not just accuracy

Needs first: Build and train neural networks in PyTorch

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

3 tools to explore

Ultralytics YOLO

Images · Test

Train or run an object detector and examine missed detections and false positives.

Practice

Vehicle-type detector for Indian street scenes

Fine-tune YOLO (Ultralytics) to detect and localise the vehicle classes that actually fill an Indian road — auto-rickshaw, two-wheeler, car, bus — from a dataset you build yourself. Shoot 120–200 photos on your phone from a footpath, a bus stop or a window at different times of day, and label them in YOLO format using Roboflow's free tier or LabelImg. Train, evaluate with mAP@0.5 on a held-out split, then test on fresh photos from a location you never trained on. The lesson is how quickly lighting, occlusion and class imbalance eat a small hand-labelled dataset.

Start from

120–200 street photos you shoot on your own phone, labelled in YOLO format with Roboflow's free tier or LabelImg

Milestones
  1. Shoot and curate the photo set across lighting conditions, and fix the class list · ~3h
  2. Label everything in YOLO format and export a clean train/val split · ~4.5h
  3. Fine-tune YOLO, adjust augmentation, and read the per-class mAP breakdown · ~3h
  4. Test on fresh photos, catalogue the failures, and export the weights · ~3h
Done when
  • Annotated dataset (100+ labeled images minimum) in YOLO format with a documented class list
  • YOLO model fine-tuned and evaluated with mAP@0.5 reported on a held-out validation split
  • Inference demonstrated on at least 5 real-world test images not in the training set, with visualized bounding boxes
  • README documents failure cases observed (e.g. poor lighting, occlusion) and what would improve them
Prove it

Evidence a recruiter can check

  • The per-class mAP@0.5 table with the rare class next to the common one, showing what forty examples actually buys
  • Fresh test photos with predicted boxes drawn, including at least two the model gets wrong, captioned with why
  • Your labelling rules — how you handled partly occluded and overlapping vehicles — since the dataset is the real work here
  • Exported weights plus a one-command inference script a stranger can run on a photo of their own street
Signal it

Hand-labelled a 150-image street-scene dataset and fine-tuned YOLO to detect vehicle types — reported per-class mAP@0.5 on a held-out split and documented the occlusion and low-light cases where it fails.

Interview

Questions you'll get asked

  1. Walk me through fine-tuning a pretrained model for a new image classification task
  2. How is object detection different from classification, and what does mAP measure?
  3. How would you build a system to detect defective products on a manufacturing line from images?
  4. What data augmentation would you use for a small image dataset, and why?
  5. How do you handle class imbalance in an object detection dataset (rare object classes)?
  6. What's IoU, and how does it factor into evaluating a detection model?
  7. How would you reduce inference latency for a vision model running on limited hardware?