All capabilities · Cloud, deployment & production

Run workloads on Kubernetes

Deployments, services, scaling, and GPU scheduling basics for model serving.

~15 focused hoursadvanced
Explore 3 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Write a Deployment + Service manifest to run a containerized API on a cluster
  2. Scale a deployment up/down manually and via a HorizontalPodAutoscaler
  3. Debug a pod stuck in CrashLoopBackOff using kubectl logs/describe
  4. Set resource requests/limits so a pod doesn't starve or get OOM-killed
  5. Roll out a new image version with zero downtime and roll back if it fails
  6. Understand namespaces and basic GPU/node-pool scheduling for model-serving pods
  7. Use a ConfigMap/Secret to inject config into a pod without rebuilding the image

Needs first: Containerize an application with Docker

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

3 tools to explore

Kubernetes

Deploy

Deploy and inspect containers using declarative workloads, services and scaling controls.

Docker

Deploy · Build

Package a service with its dependencies and run a repeatable local environment.

Practices & references

  • Deployments and Services
  • Autoscaling
Practice

Run the containerized API on Kubernetes with autoscaling

Deploy your container image to a local kind cluster — no cloud account, no card — with a Deployment, a Service, a ConfigMap for plain config and a Secret for API keys. Set resource requests and limits, add a HorizontalPodAutoscaler, and load-test until you can watch it scale up and back down. Then break a pod on purpose and work the CrashLoopBackOff back to its cause with logs and describe.

Start from

Your container image from the containerize project, running on a local kind cluster (Docker only, no cloud account)

Milestones
  1. Stand up kind and get the image running as a Deployment behind a Service · ~2h
  2. Move plain config to a ConfigMap and keys to a Secret, and prove the pod reads both · ~2h
  3. Set resource requests and limits, then add the HorizontalPodAutoscaler · ~2h
  4. Load-test until the HPA scales, recording pod count over time · ~2h
  5. Break a pod deliberately and write the CrashLoopBackOff incident note · ~1h
Done when
  • `kubectl get pods` shows the Deployment running with the correct replica count
  • Service exposes the API and it's reachable (via port-forward or LoadBalancer/Ingress)
  • HPA is configured and demonstrated scaling under a simple load test
  • A documented CrashLoopBackOff incident: cause found via `kubectl logs`/`describe`, then fixed
Prove it

Evidence a recruiter can check

  • The CrashLoopBackOff incident note: the `kubectl describe` output, the real cause, and the manifest line that fixed it
  • Pod count over time during the load test, showing the HPA scaling up and then back down
  • The manifest set — Deployment, Service, ConfigMap, Secret, HPA — with resource requests and limits set and justified in comments
  • A terminal transcript of `kubectl port-forward` reaching the API running inside the cluster
Signal it

Ran a containerized API on Kubernetes with config and secrets split out, resource limits set, and an HPA that scaled pods under load — and traced a CrashLoopBackOff down to the manifest line that caused it.

Interview

Questions you'll get asked

  1. Walk me through the YAML for a Deployment and Service exposing a Python API.
  2. A pod is stuck in CrashLoopBackOff — what commands do you run to find out why?
  3. How does a HorizontalPodAutoscaler decide when to add more pods?
  4. What happens during a rolling update, and how do you roll it back if the new version is broken?
  5. How do resource requests/limits affect scheduling and OOM kills?
  6. How would you schedule a GPU-heavy model-serving pod onto the right node pool?
  7. What's the difference between a ConfigMap and a Secret, and when do you use each?