All capabilities · Evaluation, safety & observability

Apply guardrails, safety and privacy controls

Input/output filtering, PII redaction, prompt-injection defenses, responsible-AI policies (DPDP-aware).

~8 focused hoursintermediate
Explore 3 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Filter and validate user inputs to block prompt injection and jailbreak attempts before they reach the model
  2. Redact or block PII (Aadhaar, PAN, phone numbers, bank details) in both inputs and outputs, DPDP-aware
  3. Validate model outputs against a schema or policy before they're shown to a user or executed as an action
  4. Detect and block system-prompt leakage/exfiltration attempts
  5. Apply topic/scope restrictions so a domain-specific bot refuses out-of-scope or unsafe requests gracefully
  6. Red-team your own LLM app for prompt injection, data exfiltration, and unsafe tool invocation before shipping
  7. Write and enforce a responsible-AI policy covering what the system will and won't do, with a human-escalation path
  8. Layer guardrails so a failure at one layer (e.g. model refuses) still fails safely rather than crashing or leaking

Needs first: Integrate LLM APIs into an application

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

3 tools to explore

Guardrails AI

Test · Build

Validate model inputs and outputs, then test how your application handles failures.

NVIDIA NeMo Guardrails

Test · Build

Configure conversational guardrails and test allowed and disallowed behaviours.

Presidio

Data · Test

Detect and anonymise sensitive entities in sample text before it enters a model workflow.

Practices & references

Practice

Guardrailed loan-support assistant with a red-team report

Build a customer-support assistant that answers loan and EMI questions from a public lending-policy PDF, then wrap it in layered guardrails. Benchmark the input filter against the deepset/prompt-injections set on Hugging Face — hundreds of labelled injection and benign prompts — and test redaction with format-valid but entirely fake identity-number strings you generate yourself, never real ones. Add output guardrails that block specific investment advice and system-prompt leakage, and a logged escalation path to a human whenever a guardrail trips. Red-team it with at least 10 adversarial prompts of your own and record what got through before your fix.

Start from

deepset/prompt-injections on Hugging Face — ~660 labelled injection and benign prompts, plus fake format-valid identity strings you generate for the redaction test set

Milestones
  1. Stand up the assistant over a public lending-policy PDF and record an ungarded baseline · ~1.5h
  2. Add the input layer: injection detection scored on the public set, plus regex redaction on your fake identity strings · ~1.5h
  3. Add output guardrails for investment-advice language and system-prompt leakage · ~1.5h
  4. Run the 10+ adversarial prompts, wire the escalation log, and write the before/after plus policy doc · ~1.5h
Done when
  • Input guardrail blocks or redacts PII (Aadhaar/PAN/phone patterns) in at least 90% of a test set of 20 crafted inputs
  • Output guardrail blocks specific investment-advice language and system-prompt leakage attempts
  • A documented red-team run of 10+ adversarial prompts with before/after results after hardening
  • A clear, logged escalation path (message + flag) when a guardrail triggers, instead of a silent failure or crash
Prove it

Evidence a recruiter can check

  • A before/after table of your 10+ adversarial prompts — which ones got through the first pass, and the specific guardrail that stopped each one after hardening
  • Redaction results across the 20 crafted PII cases: input, redacted output, and the hit rate, with the patterns that still slip
  • Detection scores for your input filter on the public injection set, false positives on benign prompts included
  • A transcript of an injection attempt failing and escalating, showing the exact log line a human agent would pick up
  • The responsible-AI policy doc: scope, what the assistant refuses outright, and when it hands off to a person
Signal it

Hardened an LLM support assistant with layered guardrails — injection filtering benchmarked on a public injection dataset, PII redaction, output policy checks and a logged human-escalation path — documented in a before/after red-team report.

Interview

Questions you'll get asked

  1. How do you defend an LLM app against prompt injection from untrusted retrieved documents or tool outputs?
  2. How would you redact Aadhaar/PAN numbers from both user input and model output in a DPDP-compliant way?
  3. Walk through the OWASP Top 10 for LLM apps — which 3 have you actually had to defend against?
  4. How do you stop a customer support bot from being tricked into revealing its system prompt?
  5. Design guardrails for a financial-advice chatbot that must never give specific investment recommendations.
  6. What's your process for red-teaming an LLM app before launch?
  7. How do you balance guardrail strictness against false-refusing legitimate user requests?