All capabilities · Cloud, deployment & production

Deploy an AI service to the cloud

Ship an LLM/ML API on Cloud Run / AWS Lambda-ECS / Azure with HTTPS, env config and logs.

~12 focused hoursintermediate
Explore 4 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Take a containerized API and deploy it to Cloud Run/Lambda/ECS with a public HTTPS URL
  2. Configure environment variables and secrets for a deployed service (not hardcoded)
  3. Set up autoscaling limits so a traffic spike doesn't blow the cloud bill
  4. Read and act on deployed service logs to debug a production error
  5. Roll back a bad deploy to the previous working version
  6. Set up a health-check endpoint the platform uses to know the service is alive
  7. Choose between serverless and container-based deploy for a given latency/cost profile

Needs first: Containerize an application with Docker

Learn — free, link-checked

The few resources that matter

Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

4 tools to explore

Google Cloud Run

Deploy

Deploy a containerised service and inspect its configuration, logs and scaling behaviour.

Amazon ECS

Deploy

Run containerised services and practise task definitions, networking and rollout checks.

Azure Container Apps

Deploy

Deploy a container application and configure its revisions, ingress and scaling.

Docker

Deploy · Build

Package a service with its dependencies and run a repeatable local environment.

Practices & references

  • Configuration and secrets
  • HTTPS/TLS
Practice

Deploy a public AI endpoint with logs and a practised rollback

Take your container image and put it behind a public HTTPS URL on a free tier that needs no card — a Render free web service or a Hugging Face Docker Space both deploy an image directly. Inject secrets through the platform's env/secret config rather than baking them into the image, add a /health endpoint that actually checks its dependencies, and turn on structured JSON logging. Then ship a deliberately broken revision, roll it back, and write down the exact steps and what the logs said on each side.

Start from

Your GHCR image from the containerize project, deployed to a Render free web service (no card required)

Milestones
  1. Get the image serving over public HTTPS with env-injected config · ~2.5h
  2. Add a /health endpoint that genuinely checks its dependencies · ~2h
  3. Switch to structured JSON logs and trace one request through the platform's log viewer · ~2.5h
  4. Deploy a broken revision on purpose, roll back, and capture logs from both sides · ~3.5h
Done when
  • Service is reachable over public HTTPS and returns correct responses
  • Secrets are injected via the platform's secret manager/env config, not the image
  • /health endpoint returns 200 when dependencies are up and a non-200 when a dependency is down
  • A documented rollback: deploy a broken revision on purpose, then roll back, with before/after logs
Prove it

Evidence a recruiter can check

  • The rollback in three shots: the broken revision's error logs, the rollback action, and healthy logs after
  • A live HTTPS URL (or a recording of one) returning real responses, with /health flipping non-200 when you stop its database
  • The deploy config in the repo, showing every secret referenced by name and none committed
  • One structured log line pulled from the platform's viewer, carrying request id, latency and status
Signal it

Deployed a containerized AI service behind public HTTPS with env-injected secrets, dependency-aware health checks and structured JSON logs — and rehearsed a rollback from a deliberately broken revision.

Interview

Questions you'll get asked

  1. Walk me through deploying a FastAPI service to Cloud Run from a Dockerfile.
  2. How do you manage secrets for a production deployment without putting them in code?
  3. A deployed service is returning 500s intermittently — how do you find out why?
  4. What's the difference between deploying to Lambda vs ECS vs a VM, and when would you pick each?
  5. How would you set up autoscaling so a viral traffic spike doesn't 10x your cloud bill?
  6. How do you roll back a bad deployment in under 5 minutes?
  7. What does your health-check endpoint verify, and why does that matter for uptime?