All capabilities · Machine learning & data science

Operate an ML pipeline (train → register → serve → monitor)

MLflow/SageMaker/Vertex pipelines, model registry, batch/online serving, drift monitoring.

~25 focused hoursadvanced
Explore 5 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Track experiments (params, metrics, artifacts) with MLflow instead of spreadsheets
  2. Register a trained model in a model registry with versioning and stage transitions
  3. Package a model behind a served API (batch or online/real-time endpoint)
  4. Containerize the training and/or serving code with Docker for reproducibility
  5. Set up a scheduled or triggered retraining pipeline
  6. Monitor a deployed model for data drift and performance decay
  7. Roll back to a previous model version safely when a new one regresses

Needs first: Train and evaluate classical ML models, Containerize an application with Docker

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

5 tools to explore

MLflow

Monitor · Deploy

Track experiments and model artefacts so training runs can be compared and reproduced.

Docker

Deploy · Build

Package a service with its dependencies and run a repeatable local environment.

Apache Airflow

Data · Automate

Schedule dependent data tasks and practise retries, backfills and failure recovery.

Grafana

Monitor · Data

Build dashboards for service health and investigate changes in operational metrics.

Practice

End-to-end pipeline for a fraud-detection model

Take a fraud-detection model — reuse the one from the ML fundamentals project, or train a quick one on the Kaggle credit-card fraud dataset (284,807 transactions, 492 frauds) — and wrap it in the machinery that makes it a system rather than a notebook. Track every training run in MLflow, register the winner and load it for serving by version/stage instead of a file path, serve it from a FastAPI endpoint inside Docker, and add a drift check that compares incoming feature distributions against the training set with a documented threshold. Then prove you can roll it back.

Start from

Kaggle credit-card fraud dataset — 284,807 transactions with 492 labelled frauds

Milestones
  1. Wire MLflow into the training script and log three or more runs with params, metrics and artifacts · ~3h
  2. Register the winning run and load it by stage inside a FastAPI /predict endpoint · ~3h
  3. Containerise training and serving, and get a fresh clone running with one compose command · ~3h
  4. Add the drift check with a documented threshold, then rehearse a rollback · ~3h
Done when
  • MLflow tracks at least 3 training runs with params/metrics/artifacts logged and comparable in the UI
  • Best model registered and loaded for serving by version/stage, not a hardcoded file path
  • Dockerized FastAPI endpoint returns predictions and passes a basic load/smoke test
  • A drift check script flags when incoming data statistically diverges from training data, with a documented threshold
Prove it

Evidence a recruiter can check

  • An MLflow run-comparison view (screenshot or exported CSV) showing the three runs side by side and why the registered one won
  • A terminal transcript from a fresh clone: compose up, then a real prediction returned by /predict
  • The drift report run against deliberately shifted input data, showing the check firing and the threshold that triggered it
  • The rollback you actually performed — endpoint output before and after moving serving back to the previous registered version
  • A one-page train → register → serve → monitor diagram whose boxes match the real service and file names in the repo
Signal it

Shipped a fraud-detection model end to end — MLflow tracking and model registry, a Dockerised FastAPI endpoint that loads by registry stage, plus a drift check with a documented threshold and a rehearsed rollback.

Interview

Questions you'll get asked

  1. Walk me through what happens from 'model trained' to 'model serving live traffic' in your setup
  2. How do you detect that a production model's performance has degraded, and what do you do about it?
  3. What's the difference between data drift and concept drift, and how do you monitor for each?
  4. How would you design a rollback strategy if a newly deployed model performs worse than the previous one?
  5. Batch vs online inference -- how do you decide which one a use case needs?
  6. What goes into a model registry entry beyond the weights file?
  7. How do you version datasets alongside model versions for reproducibility?