All capabilities · Machine learning & data science

Fine-tune an open LLM (LoRA/QLoRA)

Prepare instruction data, run PEFT fine-tuning on a small model, evaluate against the base.

~20 focused hoursadvanced
Explore 3 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Decide when fine-tuning is actually needed vs prompting or RAG
  2. Prepare an instruction dataset in the right format (prompt/response pairs, chat template)
  3. Run a LoRA or QLoRA fine-tune of a small open model (e.g. Llama 3 8B, Qwen, Phi) on limited GPU
  4. Configure PEFT hyperparameters (rank, alpha, target modules) and justify choices
  5. Evaluate the fine-tuned model against the base model on a held-out task, not just eyeballing outputs
  6. Merge/export LoRA adapters and quantize for cheaper inference
  7. Avoid catastrophic forgetting and overfitting on a small fine-tuning set

Needs first: Build and train neural networks in PyTorch, Build an LLM evaluation harness

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

3 tools to explore

Hugging Face Transformers

Build · Data

Load pretrained models and adapt them to a text, image or multimodal task.

Practices & references

  • LoRA
  • Quantisation
  • Held-out evaluation
Practice

Fine-tune a small LLM for Hindi/Hinglish support triage

Start from the Bitext customer-support intent dataset on Hugging Face — 27k utterances labelled across 27 intents — sample a few hundred rows and rewrite them into Hindi/Hinglish with an LLM, so you get code-mixed tickets whose labels you can still trust. QLoRA fine-tune a small open model (Llama 3 8B, Qwen2.5 7B or Phi) on a free Colab/Kaggle GPU to emit the intent plus a short reply. Score the base model and the fine-tune on the same held-out split, per intent, and report where fine-tuning helped and where a good prompt on the base model was already enough.

Start from

Bitext customer-support intent dataset on Hugging Face — 27k utterances labelled across 27 intents; a sampled slice rewritten into Hinglish with an LLM

Milestones
  1. Build the Hinglish instruction set: sample, rewrite, spot-check labels, split train/eval · ~4.5h
  2. Get a QLoRA run training on a free GPU with the chat template applied correctly · ~5.5h
  3. Sweep rank / alpha / target modules across two or three short runs and keep the loss curves · ~4.5h
  4. Score base vs fine-tuned on the held-out split and write the honest eval · ~3.5h
Done when
  • Instruction dataset (50-200+ examples minimum) in chat/instruction format, with a clear train/eval split
  • QLoRA fine-tuning run completed and logged (loss curve, hyperparameters recorded)
  • Quantitative eval comparing base vs fine-tuned model (accuracy on intent classification or a scored rubric on reply quality)
  • README documents the eval methodology and honestly reports where fine-tuning did and didn't help
Prove it

Evidence a recruiter can check

  • A base-vs-fine-tuned accuracy table on the same held-out Hinglish tickets, broken down per intent so the wins aren't hidden inside an average
  • The adapter published on the Hugging Face Hub with a model card naming the base model, the LoRA config and the eval split
  • The dataset-generation script, so a reviewer can see how the Hinglish examples were produced and how the labels survived translation
  • A short list of intents where fine-tuning made things worse, with the example outputs that show it
Signal it

Fine-tuned a small open LLM with QLoRA for Hindi/Hinglish support-ticket triage on a free GPU — measured intent accuracy against the base model per intent on a held-out split, reporting the regressions alongside the gains.

Interview

Questions you'll get asked

  1. When would you fine-tune a model instead of just improving the prompt or adding RAG?
  2. Explain LoRA -- what does it actually change in the model, and why is it cheaper than full fine-tuning?
  3. How is QLoRA different from LoRA, and what tradeoff does quantization introduce?
  4. How would you build an eval set to prove your fine-tuned model is actually better than the base model?
  5. What data would you need to fine-tune a support-ticket triage model for a Hindi-speaking customer base?
  6. What's catastrophic forgetting and how do you mitigate it during fine-tuning?
  7. How do you decide the right LoRA rank for a task?