All capabilities · Data engineering & analytics

Design and read A/B tests

Hypotheses, sample size, significance, guardrail metrics, reporting results honestly.

~11 focused hoursintermediate
Explore 4 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Write a testable hypothesis with a clear success metric before launching an experiment
  2. Calculate required sample size and expected run duration given baseline rate and MDE
  3. Choose and report the right significance test (and correct for multiple comparisons when needed)
  4. Define guardrail metrics so a 'winning' variant that harms retention/revenue is still caught
  5. Detect and explain novelty effects, weekday/weekend seasonality, and sample ratio mismatch
  6. Report a null or negative result honestly instead of p-hacking toward significance
  7. Translate a statistical result into a clear ship/no-ship recommendation for a PM

Needs first: Query and model data with SQL

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

PostgreSQL

Data · Build

Practise SQL joins and aggregations, or store application records in a relational database.

Practices & references

  • Sample-size planning
  • Statistical uncertainty
Practice

A/B test readout with sample-size check and a guardrail metric

Analyse a real mobile-game A/B test: 90k players split between two versions of a progression gate, with day-1 retention, day-7 retention and rounds played. Do it in the right order — state the hypothesis and work out what sample size the effect you care about would need, before you look at the result. Run the significance test, report a confidence interval rather than a bare p-value, check the second retention metric as a guardrail, and end with a ship/no-ship call that says what you would do if the answer were inconclusive.

Start from

Cookie Cats mobile-game A/B test on Kaggle — 90,189 players, gate-30 vs gate-40, with day-1/day-7 retention and rounds played

Milestones
  1. Write the hypothesis and the sample-size calculation for your minimum detectable effect, before looking at outcomes · ~1h
  2. Run the significance test on the primary metric and compute the confidence interval · ~1h
  3. Check the second retention metric as a guardrail and bootstrap the difference · ~1h
  4. Write the one-page readout with the ship/no-ship call · ~1h
Done when
  • Pre-registers a hypothesis and required sample size calculation before 'running' the simulated test
  • Uses an appropriate significance test and reports confidence interval, not just a p-value
  • Checks at least one guardrail metric and explicitly discusses trade-offs if it moved the wrong way
  • Ends with a one-paragraph honest recommendation, including what you'd do if the result were inconclusive
Prove it

Evidence a recruiter can check

  • The pre-registered hypothesis and sample-size calculation, committed before the analysis notebook — visible in the commit history
  • The effect reported as a confidence interval with the practical-significance threshold marked, not a p-value verdict
  • The guardrail metric's result stated even though it complicates the story, with the trade-off argued explicitly
  • A one-page readout ending in a ship/no-ship call and a stated plan for the inconclusive case
Signal it

Analysed a 90k-user A/B test end to end — pre-registered the hypothesis and sample size, reported the effect as a confidence interval, checked a retention guardrail, and wrote the ship/no-ship recommendation.

Interview

Questions you'll get asked

  1. How do you calculate the sample size needed for an A/B test given a baseline conversion rate and MDE?
  2. What's a guardrail metric and why do you need one even if the primary metric wins?
  3. How would you detect sample ratio mismatch, and why does it matter?
  4. A test hits significance on day 2 — do you call it? Why or why not?
  5. Explain p-value and statistical significance to a non-technical PM in one minute.
  6. How do you handle novelty effects when a new feature initially spikes engagement?
  7. Walk me through how you'd design an experiment to test a new UPI checkout flow.
  8. What would make you recommend NOT shipping a variant that won on the primary metric?