Hugging Face Transformers
Build · Data
Load pretrained models and adapt them to a text, image or multimodal task.
Tokenization, transformers via Hugging Face, fine-tune small encoders for classification/NER.
Explore 3 tools for this projectShare of job postings in India, per role, that name this capability.
Needs first: Train and evaluate classical ML models
Start with one tool for each part of your project. You don’t need to learn them all.
3 tools to explore
Build · Data
Load pretrained models and adapt them to a text, image or multimodal task.
Data
Build a text-processing pipeline and inspect entities, tokens and annotations.
Data · Test
Build baseline models and evaluate them with consistent train/test splits.
Use the hi-IN split of the MASSIVE dataset on Hugging Face — 16.5k Hindi utterances labelled with intents and slot spans — as a stand-in for an incoming support queue. Build a TF-IDF + logistic-regression baseline first, then fine-tune a multilingual encoder (mBERT, IndicBERT or MuRIL) on the same label set and compare per-class F1. Add the slot/entity head so the system extracts the useful fields, not just a category. Finally, rewrite a sample of the test set into romanised Hinglish and measure how much accuracy you lose — that gap is the real finding.
Built a Hindi support-message triage model on 16.5k labelled utterances — a fine-tuned multilingual encoder beat a TF-IDF baseline on per-class F1, with slot extraction and a measured accuracy drop on romanised Hinglish input.