All capabilities · Machine learning & data science

Solve NLP tasks (classification, NER, similarity)

Tokenization, transformers via Hugging Face, fine-tune small encoders for classification/NER.

~20 focused hoursintermediate
Explore 3 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

No job posting we have read names this yet. It stays on the map so the picture is complete; it will not appear on a path until a role asks for it.
What employers mean

You should be able to…

  1. Tokenize and preprocess text correctly for a transformer model (truncation, padding, special tokens)
  2. Fine-tune a small pretrained encoder (BERT/DistilBERT) for text classification
  3. Build a Named Entity Recognition (NER) model for a domain-specific entity set
  4. Compute text similarity/embeddings for search or deduplication
  5. Handle multilingual or code-mixed text (e.g. Hindi, Hinglish) preprocessing challenges
  6. Evaluate classification/NER with the right metrics (F1, precision/recall per class, not just accuracy)
  7. Know when a classic ML baseline (TF-IDF + logistic regression) beats a transformer for a small dataset

Needs first: Train and evaluate classical ML models

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

3 tools to explore

Hugging Face Transformers

Build · Data

Load pretrained models and adapt them to a text, image or multimodal task.

Practice

Hindi support-message intent classifier and entity extractor

Use the hi-IN split of the MASSIVE dataset on Hugging Face — 16.5k Hindi utterances labelled with intents and slot spans — as a stand-in for an incoming support queue. Build a TF-IDF + logistic-regression baseline first, then fine-tune a multilingual encoder (mBERT, IndicBERT or MuRIL) on the same label set and compare per-class F1. Add the slot/entity head so the system extracts the useful fields, not just a category. Finally, rewrite a sample of the test set into romanised Hinglish and measure how much accuracy you lose — that gap is the real finding.

Start from

MASSIVE dataset, hi-IN split, on Hugging Face — 16.5k Hindi utterances labelled with 60 intents and slot spans

Milestones
  1. Load the hi-IN split, collapse it to a working label set, build the TF-IDF baseline · ~3.5h
  2. Fine-tune a multilingual encoder on the same intent task · ~5.5h
  3. Train the slot/NER head and inspect extractions on real utterances · ~5.5h
  4. Probe with romanised Hinglish rewrites and document what breaks · ~3.5h
Done when
  • Labeled dataset (100+ examples) with a documented labeling scheme and train/test split
  • Both TF-IDF baseline and fine-tuned transformer trained, with F1 per class compared in a table
  • NER or entity extraction component demonstrated on at least 10 example tickets with correct/incorrect cases shown
  • README honestly states which approach won and why, including any code-mixing failure cases found
Prove it

Evidence a recruiter can check

  • A per-class F1 table for TF-IDF vs the fine-tuned encoder, calling out the classes where the cheap baseline still wins
  • Ten annotated slot-extraction examples with correct and wrong spans shown side by side
  • The script-gap measurement: the same messages in Devanagari and in romanised Hinglish, with the accuracy drop between them
  • A tokenizer inspection — a few Hindi sentences printed as token ids — explaining why subword splits hurt specific classes
Signal it

Built a Hindi support-message triage model on 16.5k labelled utterances — a fine-tuned multilingual encoder beat a TF-IDF baseline on per-class F1, with slot extraction and a measured accuracy drop on romanised Hinglish input.

Interview

Questions you'll get asked

  1. How would you build a classifier to route support tickets into 5 categories, and what would you try first?
  2. Walk me through fine-tuning BERT for a classification task -- what changes vs training from scratch?
  3. How do you evaluate a NER model, and why isn't accuracy the right metric?
  4. What challenges come up with Hindi or Hinglish (code-mixed) text that don't come up in English?
  5. When would a TF-IDF + logistic regression baseline be a better choice than fine-tuning a transformer?
  6. How do you handle class imbalance in a multi-class text classification problem?
  7. What's the difference between a tokenizer's vocabulary and the model's embedding layer?