scikit-learn
Data · Test
Build baseline models and evaluate them with consistent train/test splits.
Scikit-learn workflow: features, train/test split, cross-validation, metrics, overfitting.
Explore 5 tools for this projectShare of job postings in India, per role, that name this capability.
Needs first: Write production-quality Python for AI work
Start with one tool for each part of your project. You don’t need to learn them all.
5 tools to explore
Data · Test
Build baseline models and evaluate them with consistent train/test splits.
Data
Clean tabular data, join datasets and produce reproducible summaries in Python.
Data
Work with numerical arrays and vectorised operations for analysis and modelling.
Data
Train a boosted-tree baseline and compare its performance with simpler models.
Data · Code
Keep code, results and explanations together in an exploratory notebook.
Predict which borrowers default using the UCI 'Default of Credit Card Clients' dataset — 30,000 real consumer credit records where roughly one in five accounts defaults. Build a single scikit-learn Pipeline that does preprocessing and the model together, so nothing from a test fold leaks into fitting, and compare logistic regression against XGBoost with cross-validated scores. Address the class imbalance explicitly and show the before/after. Report ROC-AUC and a cost-weighted metric — a missed default and a wrongly refused loan are not the same size mistake — instead of raw accuracy.
Built a credit-default risk model on 30k labelled consumer loan records — leak-free scikit-learn pipeline, XGBoost benchmarked against logistic regression with cross-validation, and a decision threshold chosen from a cost matrix rather than accuracy.