Not a course — a build list
75 practice projects, one per capability employers ask for, and 17 bigger ones — one per role — that stand in for the job. Each says what "done" means, so a recruiter can check it. Build the ones your path puts first; the rest are here when you get to them.
The project that stands in for the job
The last thing on every path. Its acceptance checks are the interview, in advance.
Inbox-to-CRM operations agent with an escalation path
Pick a real repetitive process - inbound sales enquiries, support tickets, or invoice emails - and automate it end to end in self-hosted n8n.
A 30-second AI-generated product ad for a fictional Indian D2C brand, from script to a 9:16 export
Invent a small Indian D2C brand — a skincare serum, a sneaker, a cold-brew — and write a 30-second ad script with a hook in the first three seconds and one recurring on-screen character.
Ops metrics dashboard with an AI-assisted analysis log
Take a real, messy operational dataset (public e-commerce/delivery orders, or your college's fee or attendance exports) and run it end to end the way RentoMojo or Vedantu would: load it into Postgres or DuckDB, write SQL to model orders into daily SLA/TAT, cost-per-order and failure-reason metrics, then build one Power BI or Looker Studio dashboard a manager could open every Monday.
A public annotation project with a written guideline, a measured agreement score, and an LLM rating set
Pick a domain you genuinely know — a language you speak natively, cricket clips, medical leaflets, code, local street imagery — and collect 300 public items.
Enterprise document assistant: grounded RAG plus a tool-using agent, deployed and measured
Build a FastAPI service that ingests a real document set (policy PDFs, product manuals, or a public dataset), answers questions with cited RAG over a vector store, and escalates multi-step requests to a LangGraph agent that can call two or three tools — a database lookup, a ticket creation stub, and a human-approval step.
AI use-case register and a full risk assessment pack for one real organisation
Pick an organisation you can actually see inside — your current employer, a college, an NGO, or a public-sector body with published AI use — and do the work an AI governance analyst does in their first month.
Agent feature: PRD, working prototype and an eval scorecard
Pick a real, narrow workflow you understand (support ticket triage, lead qualification for a local business, contract clause review) and take it through the loop an AI PM actually runs.
An eval harness for an AI feature you did not build
Take a small LLM feature you do not own — a RAG assistant over a public document set works well (an insurer's policy wordings, a state scheme handbook, your own product's help centre) — and treat it exactly as a QA engineer treats a build handed over for testing.
An AI-visibility audit for one brand: a tracked prompt set, a technical check and the schema fixes
Pick one Indian brand with a public website in a category people ask assistants about — an edtech course, a term-insurance plan, a SaaS tool — plus three competitors.
Break a RAG-and-agent app you built, then ship the guardrails, the test harness and the risk report
Build a small enterprise-style AI app - a RAG assistant over a document set plus an agent with two or three real tools, one of which writes something - and then attack your own system: direct and indirect prompt injection, system-prompt leakage, retrieval-based data exfiltration, vector-store poisoning, and an agent tricked into calling a tool it should not have.
A customer-style POC in a box: discovery doc, working demo, RFP response
Pick one vertical you can speak to (insurance claims, hospital discharge summaries, manufacturing QMS) and run the full pre-sales loop on yourself.
An AI support lab: a small LLM API in Docker, five injected production failures, and a runbook, ticket and RCA for each
Run a small open model locally with Ollama inside Docker Compose, put a 60-line FastAPI proxy in front of it that adds request IDs, structured JSON logs and a per-minute rate limit, and ship the logs to a Grafana + Loki dashboard (or a SQLite table and a notebook).
A complete 3-hour 'Generative AI for working professionals' workshop kit: slides, a runnable Colab lab on public movie reviews, a 10-question assessment and a recorded 10-minute demo lecture
Build the whole kit a skilling institute or L&D team would hand a new trainer, for a mixed room of graduates and working professionals with no AI background.
Aerial object detector: train on VisDrone, export to ONNX, benchmark on a budget
Train a YOLO detector on the public VisDrone aerial dataset (10 classes, small objects, harsh lighting - the same problems Big Bang Boom, ArcelorMittal and ShipIn describe) on a free Colab or Kaggle GPU, and do the error analysis a hiring manager will ask about: which classes fail, why small objects are missed, what augmentation fixed.
Deploy an agent into a fake customer's stack — with a KPI, an eval bar and a runbook
Invent a plausible Indian customer (a lender, a D2C brand, an OTT platform) and write a one-page discovery note: their workflow today, the number you will move, and their constraints (data cannot leave their VPC, Hindi + English users, an existing ticketing tool).
Train, ship and keep alive: a drift-monitored ML service
Take one real tabular or image dataset, train a baseline scikit-learn model and a small PyTorch model, and track every run in MLflow so the winning experiment is reproducible.
Prompt workbench: a document-extraction assistant with a versioned prompt library and an eval harness
Pick a messy real-world document set (insurance loss runs, invoices, college transcripts, WhatsApp support transcripts) and build a small Python service that extracts a fixed schema from each document using an LLM.
Practice projects, by what they prove
Programming foundations
Support-ticket triage CLI as a typed, tested Python package
Build a Python package that reads a folder of labelled customer-support tickets, classifies each into a category (refund, fraud, technical) with rules plus one small model call, and writes a summary CSV.
Hinglish conversation summarizer API with validated payloads and key auth
Build a FastAPI service with three endpoints: submit a Hinglish support thread, fetch its auto-generated English summary and priority, and list threads with pagination.
Resilient RAG API design for a growing document corpus
Design, and prototype the risky parts of, a RAG API answering questions over a corpus that keeps growing: an ingestion queue, an embedding and vector-store step, a cache for repeated questions, and a fallback path when the primary LLM provider errors.
AI-assisted refactor with a paper trail
Take a messy script of your own — an old college or hackathon project is ideal — and use an AI coding assistant to refactor it into typed, tested modules over four small PRs.
A reviewable PR trail on a repo a stranger can run
Take a small tool you have already written and turn its remaining work into a real collaboration history.
40-problem DSA interview log with pattern notes
Work at least 40 problems across arrays and hashing, trees, graphs, dynamic programming and intervals from the free NeetCode roadmap over two to three weeks.
LLM application development
Prompt regression harness for a support-ticket classifier
Sample 300 rows from a public customer-support dataset, hand-label 30 of them for category, priority and sentiment, and build a prompt plus an eval script that classifies each ticket.
Deployed FAQ assistant for a local Indian business
Write a 30-question FAQ for a small business you actually know - a kirana chain, a coaching institute, your family's shop - or lift one from a real business's public FAQ page.
Multi-provider LLM chat CLI with cost tracking
Build a command-line chat tool that talks to two different LLM providers behind one shared interface, streams the response token-by-token, and logs the cost of every call in INR (fixed USD-INR rate) to a local SQLite file.
Voice-note-to-structured-ticket pipeline for Hindi support calls
Take short Hindi and Hinglish clips from the Common Voice Hindi set, or record your own on a phone, transcribe them with a speech-to-text API, and have an LLM turn each transcript into a structured support ticket: category, urgency, English summary.
Rolling-memory assistant for 100-turn conversations
Build a chat assistant that holds a 100+ turn conversation without ever hitting the context limit.
Schema-validated field extractor for GST invoices
Write a generator that emits 20 synthetic GST invoices, varying vendor layout, GSTIN format, HSN codes and CGST/SGST splits, and keeping the ground truth for each.
Retrieval & knowledge systems
Grounded RAG assistant over Indian labour-law documents
Build a RAG assistant over India's labour codes and related public labour-law documents, answering questions like 'how many days of casual leave am I entitled to?' with citations down to the section.
Hybrid search over Indian consumer-tech product reviews
Embed a few thousand Flipkart product reviews into pgvector or Chroma, build a BM25 keyword index over the same corpus, and fuse the two with a re-ranking step.
Eval harness for a RAG assistant: chunking-strategy shootout
Take the corpus and pipeline from your RAG project and hand-label 25 question / answer / source triples over it.
OCR-fallback ingestion pipeline for RBI policy PDFs
Point an ingestion pipeline at a folder of RBI master directions and circulars - public PDFs, a mix of clean digital text and older scanned ones - and produce clean, chunked, metadata-tagged text ready for embedding.
Agents & workflows
Multi-agent document-triage workflow with human approval
Build a LangGraph workflow that triages loan-application packs through an extraction agent, a policy-check agent and a summarizer agent, pausing for human approval before any application is marked approved.
Transaction-support bot that answers from tools, not memory
Build a support chatbot that answers questions about payment transactions (status, refund eligibility, dispute filing) by calling four tools against a local transactions database instead of guessing.
MCP server over India's company-registry open data
Build an MCP server that wraps a slice of the Company Master Data registry published on data.gov.in and exposes 'search_company' and 'get_filing_details' tools plus a 'filings://recent' resource.
Trajectory eval harness and CI gate for a triage agent
Take the document-triage agent from the agent-workflow project and build a harness that replays recorded runs and scores the whole trajectory — which tools were called, in what order, and whether any unsafe action slipped through — not just the final answer.
Evaluation, safety & observability
Guardrailed loan-support assistant with a red-team report
Build a customer-support assistant that answers loan and EMI questions from a public lending-policy PDF, then wrap it in layered guardrails.
Eval harness and CI gate for a Hindi/English support-ticket summarizer
Build a summarizer for bilingual customer-support tickets, then build the eval harness that keeps it honest.
Per-tenant cost and latency dashboard for a RAG support bot
Stand up a small RAG support bot serving two tenants over two different sets of public policy/FAQ PDFs, then instrument every request with Langfuse tracing — nested spans for retrieval, prompt construction, the model call and any tool call.
Cost and latency optimization pass on a support RAG chatbot
Take a working RAG support chatbot and cut its cost per request and p95 latency without losing answer quality.
Cloud, deployment & production
Tenant isolation and secret-hygiene pass on the ticket API
Harden your FastAPI ticket service for multiple tenants. Add JWT auth with a tenant claim, filter every data query by the tenant in the token, and seed two tenants' worth of rows so isolation is testable rather than assumed.
AI support lab: break a local LLM API five ways and support it back to health
Run a small model on Ollama behind a thin FastAPI wrapper in Docker Compose, with request ids and structured JSON logs on every call.
Deploy a public AI endpoint with logs and a practised rollback
Take your container image and put it behind a public HTTPS URL on a free tier that needs no card — a Render free web service or a Hugging Face Docker Space both deploy an image directly.
Test-to-deploy pipeline for the ticket API
Add a GitHub Actions workflow to your containerized ticket-API repo that runs pytest and a linter on every pull request, builds and pushes a Docker image to GHCR tagged with the commit SHA on merge to main, and deploys only after a manual approval.
Multi-stage container and Compose stack for the ticket API
Containerize the FastAPI ticket service with a multi-stage Dockerfile: build deps in one stage, ship a slim runtime image in the next, and run the process as a non-root user.
Run the containerized API on Kubernetes with autoscaling
Deploy your container image to a local kind cluster — no cloud account, no card — with a Deployment, a Service, a ConfigMap for plain config and a Secret for API keys.
Managed RAG assistant on one hyperscaler
Pick one platform — Bedrock, Vertex AI or Azure AI Foundry, all of which start on free trial credit — and build a small assistant over a set of public policy PDFs using the platform's managed knowledge-base/RAG feature rather than your own pipeline.
Machine learning & data science
Credit default risk model with a leak-free scikit-learn pipeline
Predict which borrowers default using the UCI 'Default of Credit Card Clients' dataset — 30,000 real consumer credit records where roughly one in five accounts defaults.
Vehicle-type detector for Indian street scenes
Fine-tune YOLO (Ultralytics) to detect and localise the vehicle classes that actually fill an Indian road — auto-rickshaw, two-wheeler, car, bus — from a dataset you build yourself.
Devanagari handwriting classifier: from-scratch CNN vs transfer learning
Train a CNN in PyTorch to classify handwritten Devanagari characters using the UCI Devanagari Handwritten Character Dataset — 92,000 labelled 32×32 images across 46 characters — writing the Dataset, DataLoader and training loop yourself rather than calling a trainer.
End-to-end pipeline for a fraud-detection model
Take a fraud-detection model — reuse the one from the ML fundamentals project, or train a quick one on the Kaggle credit-card fraud dataset (284,807 transactions, 492 frauds) — and wrap it in the machinery that makes it a system rather than a notebook.
Fine-tune a small LLM for Hindi/Hinglish support triage
Start from the Bitext customer-support intent dataset on Hugging Face — 27k utterances labelled across 27 intents — sample a few hundred rows and rewrite them into Hindi/Hinglish with an LLM, so you get code-mixed tickets whose labels you can still trust.
Hindi support-message intent classifier and entity extractor
Use the hi-IN split of the MASSIVE dataset on Hugging Face — 16.5k Hindi utterances labelled with intents and slot spans — as a stand-in for an incoming support queue.
Two-stage recommender with cold-start fallback on MovieLens
Using MovieLens ml-latest-small — 100,836 ratings from 610 users across 9,742 titles — build the two halves of a real recommender rather than one flat model: a collaborative-filtering candidate generator, then a ranking step over its output.
Festival-season demand forecast with Prophet and XGBoost backtesting
Use the public Kaggle 'Power consumption in India (2019–2020)' series — daily state-wise electricity demand spanning two festival seasons — as a demand series with strong, genuinely Indian seasonality.
Data engineering & analytics
Star-schema retail dashboard with city and category drill-down
Take the Superstore sales export — a single flat sheet of orders with city, region, category, ship mode and order/ship dates — and model it properly into fact and dimension tables before a single visual is drawn.
One-page decision memo from a churn analysis
Take the Telco customer churn dataset, run the analysis properly, then throw away everything except one page.
Multi-table SQL analysis of a 100k-order e-commerce dataset
Load the Olist e-commerce dataset (nine related CSVs: orders, order_items, payments, reviews, customers, sellers, products) into Postgres and answer 10 business questions purely in SQL — daily order trend, top delivery-delay causes, repeat customers, and a window-function query flagging customers who place 5+ orders inside a 10-minute window.
Clean and reconcile a messy startup-funding export
Take the Indian Startup Funding dataset — a genuinely dirty real export with amounts stored as strings with commas, four different date formats, city names spelled several ways (Bangalore/Bengaluru/Banglore), investor lists crammed into one column, and duplicate rows from repeated scrapes.
Daily incremental warehouse pipeline with Airflow, dbt and data tests
Use the NYC TLC trip-record parquet files as a stand-in daily feed: split one month into day-sized partitions and land them one at a time so the pipeline sees a real arriving feed.
A/B test readout with sample-size check and a guardrail metric
Analyse a real mobile-game A/B test: 90k players split between two versions of a progression gate, with day-1 retention, day-7 retention and rounds played.
LLM triage pipeline for a support-ticket backlog (English + Hinglish)
Sample 200 real support conversations from the Customer Support on Twitter dataset and write 10-15 Hinglish tickets yourself for the code-mixed cases.
Creative production & media
Sixty-second AI-generated product spot with a reusable prompt library
Invent a small D2C brand (a chai, skincare or sneaker label) and write a 60-second script with one recurring presenter character and one hero product.
Reels, Shorts and ad cutdowns finished from AI-generated clips
Take the clips, voiceover and brand sheet from your generative-media project (or any 8–12 clips you generate on free tiers) and finish them into a publish-ready campaign: a 30–45 second 9:16 Reel with a three-second hook, a 1:1 feed version, a 16:9 YouTube version and a 15-second ad cutdown.
Product, business & communication
AI feature PRD: automated GST-invoice extraction for an SMB accounting tool
Talk to three people who re-key invoices by hand — a shop owner, a freelancer's accountant, a friend in accounts payable — or role-play them from a written persona if you cannot reach three.
Responsible-AI checklist and model card for a resume-screening assistant
Take a hypothetical feature that screens resumes for a recruiting team, and write it up against the actual law rather than a summary of it: MeitY publishes the Digital Personal Data Protection Act 2023 as a free PDF.
POC: AI-assisted GST reconciliation for a mid-size retailer
Role-play a discovery call with a retailer's finance lead who wants purchase invoices reconciled against GST returns. Write a one-page POC scope doc defining what will and will not be built in two weeks.
Explainer deck and cost model for a Hindi/Hinglish support bot
Pick a real Indian scenario — a lender's helpline fielding Hindi and Hinglish loan-repayment questions — and write a one-page, jargon-free explainer of how an LLM-based bot would answer it, where it goes wrong, and how grounding it in the lender's published policy documents changes the answer.
Three-hour 'Intro to Generative AI' workshop kit with a runnable lab, assessment and recorded teach-back
Design a complete 3-hour beginner workshop on Generative AI for a mixed cohort of students and working professionals, then prove you can deliver it.
BRD/FRD and solution design for automated loan-document verification
Role-play discovery with an operations lead whose team hand-checks the documents on every loan application.
Trade-off memo and recorded leadership update for a delayed AI launch
Write yourself a one-page scenario brief first: a document-checking feature is three weeks late because accuracy on regional-language documents sits under the bar, and you have the per-language numbers.
Eval harness for a Hindi/English support-intent classifier
Pull Hindi and English utterances from the MASSIVE dataset — real user requests already labelled with intent — and cut a 100-example eval set across the categories a support desk cares about.
Demo script and RFP response for an AI document-search feature
Index a real public document set — the RBI's Master Directions run to dozens of long policy PDFs, free to download — and build a small, reliable semantic-search demo over them.
Deployed prototype: a payment-dispute triage assistant
Build a Streamlit or Gradio app where someone pastes a payment complaint in their own words and gets back a structured classification — fraud, failed-but-debited, merchant issue — plus a drafted complaint message, powered by an LLM API.
AI automation & no-code
Payment-failure follow-up workflow with branching and retries
Build an n8n (or Zapier) workflow that reacts to failed-payment events, branches by failure reason (insufficient balance, bank timeout, wrong PIN), and sends a tailored follow-up message for each branch.
AI triage automation for Hindi and English support tickets
Build a workflow (n8n or Zapier) that takes a support message in Hindi or English, uses an LLM step to classify category and urgency and pull key fields into structured JSON, drafts a suggested reply, and routes high-urgency tickets to a human-approval step before anything is sent.
ROI case for automating vendor-invoice approval
Find someone who actually approves invoices — a contact in finance ops, or a written persona you role-play if you cannot reach one — and get a time estimate for every step of the current process: receipt, matching to the PO, approval routing, payment.
RPA bot that logs in, downloads and files monthly reports
Build a UiPath (or Power Automate Desktop) bot against ACME System 1 — the free practice web app UiPath uses in its own academy, with a real login and downloadable monthly reports.
WhatsApp support bot for a D2C brand's order-status and returns queries
Build a bot (Botpress or Voiceflow) on the WhatsApp Cloud API test number — a free Meta developer account gives you one that messages up to five verified test numbers with no business verification.
Annotation, quality & human feedback
Apply and stress-test a labeling guideline on real app-store reviews
Pull about 150 recent reviews of a popular Indian app with the google-play-scraper package — no API key, and the text is genuinely Hindi, English and Hinglish as users write it.
Label-quality audit and inter-annotator agreement report
Use the GoEmotions raw release, where every Reddit comment carries labels from several named raters — real disagreement between real people, not simulated.
Rate 50 model answers in your language and write the rubric
Pick a language you are native in. Write 50 everyday Indian questions — a PF withdrawal, a train booking, a school admission, a recipe — and put them to a free chat model.
Build an expert eval set in the field you already know
Take the domain you have professional experience in — nursing, accounting, law, civil engineering, software.
Pairwise LLM response ratings with a written rubric and rationales
Sample 40 prompts from the Anthropic hh-rlhf dataset covering factual Q&A, coding and open-ended help, and add a few Hindi-language customer-service prompts of your own.
Rubric and golden set for grading AI customer-support replies
Take real customer-support questions from the Bitext support dataset and generate AI replies to them, deliberately including some weak ones.
Bounding-box labeling project in Label Studio, exported to COCO
Set up a Label Studio (or CVAT) project to bounding-box-label 100 street-scene images from the public COCO val2017 set for three object classes.
Career signal & proof
Portfolio audit and rebuild: 2 deployed AI/automation projects
Take two projects you have already built — earlier practice projects on this path count — and bring them to a bar a recruiter can read in two minutes. Deploy both to live URLs.
Interview-readiness sprint: STAR stories + 2 mock interviews
Write STAR-format stories for your three strongest projects, covering at least one technical deep-dive, one stakeholder-communication story, and one failure with the lesson attached.
Search & AI visibility
AI visibility audit and share-of-voice tracker for a consumer brand
Pick an Indian D2C brand and three competitors in one category, and write a 30-prompt set of real buyer questions grouped by intent.
Technical SEO audit of a small business website, with a verified fix log
Pick a public website of an Indian small business or D2C brand with under 500 pages and crawl it with Screaming Frog's free version, test its key templates in PageSpeed Insights, and read its robots.txt and sitemap.
Schema.org JSON-LD pack for a small business website, validated and monitored
Build a small multi-page site on GitHub Pages for your own or an imagined Indian small business (a home-baker, a D2C skincare label, a coaching centre): home, about, two product or service pages and an article.
Finished one? Mark the step done on your path and paste the link — it stays in your browser, we never upload or share it. Shares are per role, from the job postings behind each role page; the rules are on the methodology page.