All capabilities · Creative production & media

Produce video and images with generative AI tools

Turn a brief or script into finished visuals: shot-level prompts for image and video models (Midjourney, Stable Diffusion, Runway, Kling, Veo / Google Flow, Higgsfield), voice and avatars (ElevenLabs, HeyGen), and keep characters, style and brand consistent across scenes with a prompt library you maintain.

~25 focused hoursbeginner
Explore 7 tools for this project
Market relevance

Which roles ask for this — and how often

Share of job postings in India, per role, that name this capability.

What employers mean

You should be able to…

  1. Turn a script or storyboard into shot-level prompts that name subject, action, camera move, lens, lighting and mood for each shot
  2. Generate a key still per shot first, then drive image-to-video (Flow, Kling, Runway, Higgsfield) from it instead of gambling on text-to-video
  3. Keep one character's face, costume and environment consistent across 8+ scenes using reference images and seeds, and show the retakes it took
  4. Build and maintain a prompt library: what worked, what failed, model, seed, reference, and the brand style references each entry belongs to
  5. Voice a script in ElevenLabs and produce a talking-avatar segment in HeyGen with lip-sync that survives a close look
  6. Review every generation for AI artifacts — extra fingers, broken text, impossible geometry, drifting product shape — and log the fix or the retake
  7. Match a brand's look (colour, typography, product packaging, realistic Indian faces) rather than the model's default 'AI look'
  8. Decide when to abandon a generation and switch model, reframe the shot, or edit around it — and hit a turnaround of several finished clips a day

Needs first: Design and version prompts systematically

Learn — free, link-checked

The few resources that matter

Read · beginner · 20 min · help.heygen.com

HeyGen Avatar IV Complete Guide

Official walkthrough of turning one photo plus a script or audio file into a talking-avatar video. Check the current plan's video and premium-feature limits before planning your intro. — HeyGen Help Center
Read · beginner · 20 min · elevenlabs.io

ElevenLabs Text to Speech product guide

Explains voice choice, model choice, and the stability, similarity and speed settings for a natural voiceover. Practise a short passage first to judge delivery and credit use before voicing the full script. — ElevenLabs docs
Read · beginner · 25 min · academy.runwayml.com

Runway Academy Prompting Guide

Runway's own guide to shot-level prompting — camera vocabulary and why image-to-video prompts should describe motion, not the image. The guide is free to read; generation uses credits, so start with a short test shot. — Runway Academy
Read · beginner · 30 min · docs.cloud.google.com

Veo video generation prompt guide

Google's breakdown of a video prompt into subject, context, action, style, camera motion, composition and ambiance for Veo. The guide is free to read; check Flow's current subscription and account requirements before generating clips. — Google Cloud docs
Watch · beginner · 31 min · youtube.com

How to Make Your First AI Short Film (Full Tutorial)

One end-to-end pass from script to stills to image-to-video clips to voiceover and a cut — the exact loop Indian AI-content ads describe as 'raw script to finished output', using tools with free tiers. — ElevenLabs
Read · intermediate · 40 min · docs.comfy.org

ComfyUI text-to-image tutorial

The open-source route employers name alongside Stable Diffusion and Flux: loads a checkpoint, explains seed, steps, CFG and negative prompts node by node, and runs locally with no credits at all — the deep dive once free tiers run out. — ComfyUI docs
Tools for practice

Choose a tool for the job

Start with one tool for each part of your project. You don’t need to learn them all.

Go to the practice brief

7 tools to explore

Google Flow

Video · Images

Build connected shots from prompts and reference images using Google’s creative models, including Veo.

Subscription access

Google AI plan or eligible Workspace account; availability varies.

Krea

Video · Images

Try different image and video models in one workspace and compare how they interpret your brief.

Limited free plan

Model access and generation limits depend on your plan.

Kling AI

Video

Animate a key still into a short clip to practise subject movement and shot continuity.

Check current access

Review the available models and credits in the app.

Practice

Sixty-second AI-generated product spot with a reusable prompt library

Invent a small D2C brand (a chai, skincare or sneaker label) and write a 60-second script with one recurring presenter character and one hero product. Storyboard it into 8–10 shots, generate a key still per shot, and choose an image-to-video tool from the toolkit. Start with a short test shot and check your available credits before generating the full sequence. Add a voiceover and an avatar intro using tools you can access. Log every prompt, seed, reference and retake in a prompt library so the same character and product look identical from the first shot to the last.

Start from

A 60-second script and a one-page brand sheet (name, colours, product shape, presenter description) you write yourself, with or without an LLM — no footage, no client

Milestones
  1. Write the script and brand sheet, then storyboard 8–10 shots with a shot-level prompt for each (subject, action, camera, lens, light, mood) · ~4h
  2. Lock the presenter and product: generate reference stills, pick seeds, and start the prompt library with working and rejected prompts · ~6h
  3. Test one shot in your chosen image-to-video tool, then generate the sequence within your available credits; switch tools only if needed to keep character and product consistent · ~8h
  4. Voice the script and produce an avatar intro with your chosen tools; check lip-sync and pacing against the script · ~3h
  5. Run an artifact QA pass on every clip (hands, text, geometry, physics), regenerate the failures and write the retake notes · ~2.5h
Done when
  • 8–10 finished clips exist in which the presenter's face, costume and the product's shape and colours match across every shot, with a contact sheet proving it
  • The prompt library lists, for every shot, the model, prompt, seed or reference image, number of retakes, and why each rejected take was rejected
  • A voiceover of the full script and an avatar segment of at least 10 seconds exist with no visible lip-sync drift
  • A QA log names at least five AI artifacts you caught (garbled text, extra fingers, warped packaging, etc.) and how each was fixed or retaken
Prove it

Evidence a recruiter can check

  • A contact sheet of all 8–10 shots showing the same presenter and product from first shot to last, with the reference stills they were locked from
  • The prompt library as a spreadsheet or markdown table — model, prompt, seed/reference, retake count and rejection reason per shot
  • A QA log with before/after frames for at least five artifacts you caught and fixed
  • The raw clips, voiceover and avatar segment, with a note on the tool, model and access plan used for each
Signal it

Produced a 60-second AI-generated product spot with one consistent presenter and product — backed by a prompt library that logged the tools, models, seeds, references and retakes.

Interview

Questions you'll get asked

  1. Walk me through how you would turn this 30-second script into shots and prompts for Google Flow or Kling. What goes in each prompt?
  2. How do you keep the same character consistent across ten scenes? What has actually worked for you and what hasn't?
  3. Show me your prompt library. How is it organised, and how does a teammate reuse it for a new brand?
  4. You've generated a product shot and the packaging text is garbled. What do you do next?
  5. When do you go image-to-video instead of text-to-video? When do you give up on a generation altogether?
  6. How would you produce an ElevenLabs voiceover and a HeyGen avatar segment for this script in Hindi and English, and where does lip-sync usually break?
  7. How do you tell a client's brand look from the model's default look, and how do you push the model towards the brand?
  8. We need 4–6 finished AI clips a day. How do you organise your day and your credits across free and paid tools to hit that?