All capabilities · Machine learning & data science

Train and evaluate classical ML models

Scikit-learn workflow: features, train/test split, cross-validation, metrics, overfitting.

~40 focused hoursintermediate
Explore 5 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Load and clean a tabular dataset (missing values, outliers, encoding categoricals)
  2. Engineer features and build a scikit-learn Pipeline that doesn't leak test data
  3. Do a correct train/validation/test split and k-fold cross-validation
  4. Train and compare a linear model, a tree, and a boosted model (XGBoost/LightGBM)
  5. Pick and justify the right metric (accuracy vs precision/recall/F1/AUC/RMSE) for the business problem
  6. Diagnose overfitting vs underfitting from learning curves and fix it
  7. Tune hyperparameters with GridSearchCV/RandomizedSearchCV or Optuna
  8. Explain feature importance / SHAP values to a non-technical stakeholder

Needs first: Write production-quality Python for AI work

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

5 tools to explore

Practice

Credit default risk model with a leak-free scikit-learn pipeline

Predict which borrowers default using the UCI 'Default of Credit Card Clients' dataset — 30,000 real consumer credit records where roughly one in five accounts defaults. Build a single scikit-learn Pipeline that does preprocessing and the model together, so nothing from a test fold leaks into fitting, and compare logistic regression against XGBoost with cross-validated scores. Address the class imbalance explicitly and show the before/after. Report ROC-AUC and a cost-weighted metric — a missed default and a wrongly refused loan are not the same size mistake — instead of raw accuracy.

Start from

UCI 'Default of Credit Card Clients' dataset — 30,000 labelled consumer credit records with repayment history and demographics

Milestones
  1. Load and profile the data: missing values, categorical encoding, class balance · ~4.5h
  2. Build the leak-free Pipeline and a logistic-regression baseline with 5-fold CV · ~6.5h
  3. Add XGBoost, tune it, and test imbalance handling (class weights vs SMOTE vs threshold tuning) · ~6h
  4. Pick the operating threshold from a cost matrix and write up the metrics table · ~6h
Done when
  • Pipeline object (not manual steps) handles preprocessing + model, no leakage from test fold
  • At least 2 models compared with cross-validated metrics reported in a table
  • Class imbalance explicitly addressed (class weights, SMOTE, or threshold tuning) with before/after metrics
  • README explains which metric was optimized for and why, with a confusion matrix on the held-out test set
Prove it

Evidence a recruiter can check

  • A cross-validated metrics table (ROC-AUC, precision/recall at the chosen threshold) for logistic regression vs XGBoost, with the imbalance treatment as a third dimension
  • The cost matrix you chose and the decision threshold it implies, defended in the README against the default 0.5
  • A SHAP summary plot with a paragraph translating the top three features into plain lending language
  • The Pipeline definition itself, one object covering preprocessing and model, so a reviewer can see the test fold never touches the fit
Signal it

Built a credit-default risk model on 30k labelled consumer loan records — leak-free scikit-learn pipeline, XGBoost benchmarked against logistic regression with cross-validation, and a decision threshold chosen from a cost matrix rather than accuracy.

Interview

Questions you'll get asked

  1. Walk me through your ML pipeline for [X] end to end -- what would you change if precision mattered more than recall?
  2. How do you detect and prevent data leakage in a training pipeline?
  3. Explain bias-variance tradeoff with an example from a project you built
  4. When would you choose logistic regression over a random forest, and vice versa?
  5. How do you handle a severely imbalanced dataset (e.g. 2% fraud rate)?
  6. What's the difference between L1 and L2 regularization, and when do you use each?
  7. How would you validate a model before it goes to production?