All capabilities · Machine learning & data science

Build and train neural networks in PyTorch

Tensors, autograd, training loops, transfer learning, GPU usage.

~40 focused hoursintermediate
Explore 3 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Write a training loop from scratch: forward pass, loss, backward, optimizer step
  2. Use PyTorch tensors, autograd and nn.Module correctly (not just call .fit())
  3. Load data efficiently with Dataset/DataLoader, including batching and augmentation
  4. Do transfer learning: fine-tune a pretrained model (ResNet, ViT) on a new dataset
  5. Debug a network that isn't learning (loss not decreasing, exploding/vanishing gradients)
  6. Use a GPU correctly (device placement, mixed precision) and know when it isn't helping
  7. Track experiments (loss curves, learning rate schedules) and know when to stop training

Needs first: Train and evaluate classical ML models

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

3 tools to explore

Google Colab

Code · Data

Share a runnable notebook for an experiment, tutorial or classroom exercise.

Practice

Devanagari handwriting classifier: from-scratch CNN vs transfer learning

Train a CNN in PyTorch to classify handwritten Devanagari characters using the UCI Devanagari Handwritten Character Dataset — 92,000 labelled 32×32 images across 46 characters — writing the Dataset, DataLoader and training loop yourself rather than calling a trainer. Then fine-tune a pretrained ResNet on the same split. Compare the two on accuracy, training time and data efficiency by retraining both on 10%, 50% and 100% of the training set, logging every run so the curves are comparable.

Start from

UCI Devanagari Handwritten Character Dataset — 92,000 labelled 32×32 grayscale images across 46 characters

Milestones
  1. Write the Dataset/DataLoader with augmentation and visualise a real batch · ~9h
  2. Hand-write the training loop and train the from-scratch CNN to convergence · ~11h
  3. Fine-tune a pretrained ResNet on the same split, logging both runs to TensorBoard · ~9h
  4. Run the data-efficiency sweep (10% / 50% / 100% of train) and write it up · ~7.5h
Done when
  • Custom PyTorch Dataset/DataLoader with augmentation (rotation, noise) implemented
  • Training loop written manually (no high-level trainer library) with visible loss/accuracy curves
  • From-scratch CNN and fine-tuned pretrained model both trained and compared on the same test split
  • README states final test accuracy for both, and explains the tradeoff observed
Prove it

Evidence a recruiter can check

  • Loss and accuracy curves for both runs on shared axes, with the epoch where the from-scratch model starts overfitting marked
  • A data-efficiency table: test accuracy for both models trained on 10%, 50% and 100% of the training set
  • A confusion matrix that names the character pairs the model actually mixes up, not just an overall number
  • A runnable checkpoint plus a ten-line inference snippet that classifies a character you wrote and photographed yourself
Signal it

Trained a Devanagari handwriting classifier in PyTorch with a hand-written training loop — from-scratch CNN benchmarked against a fine-tuned ResNet across 10/50/100% training-data slices to quantify what transfer learning is actually worth.

Interview

Questions you'll get asked

  1. Explain what autograd does and how backpropagation actually computes gradients
  2. Your training loss is decreasing but validation loss is increasing -- what's happening and what do you do?
  3. Why would you use transfer learning instead of training from scratch?
  4. What's the effect of learning rate being too high vs too low? How do you find a good one?
  5. Explain batch normalization and why it helps training
  6. How do you handle a dataset that doesn't fit in GPU memory?
  7. Walk me through the shape of tensors flowing through a simple CNN