Computer Vision Engineer
Also posted as: Computer Vision Engineer I · CV Engineer · Vision AI Engineer · AI & Image Processing Engineer · Machine Learning and Computer Vision Engineer · Deep Learning Engineer (Vision)
A Computer Vision Engineer in India trains detection, segmentation and OCR models on images and video, then squeezes them onto the hardware where the camera actually lives - an NVIDIA Jetson on a factory line, a drone, a CCTV box or an Android phone. The postings come from industrial inspection (Ripik.AI, ArcelorMittal, Frinks AI, Mindsprint), drones and robotics (Big Bang Boom, UMA Robotic, Skild AI), medical devices (Infosys, Meril), document and KYC checks (Airtel Africa, Digitap.ai) and IT services (TCS, Infosys, Quest Global), because cameras and cheap edge GPUs are now everywhere and 18 of the 26 job postings collected want the model optimised for that hardware, not just trained. It is more senior-heavy than ML engineering - 12 of 26 ask for 5+ years and 14 of 26 want C++ next to Python - but 5 startup and robotics postings take 0-2 years, so a fresher with a deployed, benchmarked detector has a real, if narrow, door.
- Across 83 Computer Vision Engineer job postings in India, the most-requested capabilities are Build image models (detection/classification) (99%), Write production-quality Python for AI work (89%) and Build and train neural networks in PyTorch (83%).
- Pay at 2–5 yrs averages about ₹10.7 LPA (verified across 2 salary sites: AmbitionBox, Glassdoor); employers offer ₹11–15.5 LPA (median of 8 job postings that state pay, 2–5 yrs · Naukri, LinkedIn, Wellfound).
- 200+ open roles in India — one portal's count, not cross-checked yet, checked 04-10-2026.
- Hiring is concentrated in Bengaluru, Chennai and Hyderabad.
- Postings read from LinkedIn 87%, Wellfound 7%, other portals 4% and Naukri 2% — one portal supplies most of this sample, so the shares lean to the employers that post there.
Hire for this role? Add your read to this page — what decides the offer, what it closes at. An email to Ajeet, ten minutes, credited or not as you choose. How that read is shown.
Capabilities employers ask for
How often employers ask for each capability, measured across the job descriptions behind this page. Click one to see what "knowing it" means, how to learn it, and how to prove it.
Have a posting open? Check it against this map →
Build image models (detection/classification)
This is the job, and nearly every posting says so plainly. Valeo lists object detection, semantic and instance segmentation, tracking and lane detection, Jumio wants deep expertise in face recognition, and Walmart Global Tech names PyTorch, OpenCV and Torchvision. You must be able to train a detector or segmenter on your own images and prove, with the right metrics, that it works on pictures it hasn't seen.
Explore 3 practice tools →Write production-quality Python for AI work
Python is the default, often with C++ sitting next to it. Jumio wants expert Python across Pillow, OpenCV and PyTorch, KoiReader wants "fault-tolerant, multi-threaded/async code for long-running processes", and Netradyne makes Python required with C++ desired. You should write clean, tested Python that survives running all day on a camera feed, and be ready to read some C++.
Explore 4 practice tools →Build and train neural networks in PyTorch
Classical image processing alone won't get you hired; the networks are expected. Ripik.AI wants "deep proficiency in Python and PyTorch", Infosys's medical imaging team wants CNNs, GANs, VAEs and diffusion models, and Valeo simply writes "Deep Learning Frameworks: PyTorch". You need to build, train and debug a convolutional or transformer network yourself, not just call a pretrained one.
Explore 3 practice tools →Train and evaluate classical ML models
Under the vision work sits ordinary machine learning and maths. Entrupy wants probability, statistics, geometry and linear algebra even from interns, KoiReader wants "command over geometry and statistics" for logic on top of detections, and Walmart Global Tech wants scikit-learn and statsmodels. You must know how to split data, pick a loss and read a confusion matrix before the fancy architectures make sense.
Explore 5 practice tools →Deploy an AI service to the cloud
Models have to get off the training machine. KoiReader calls hands-on Docker "mandatory for containerizing complex vision applications", NewCold wants deployment on cloud platforms with Azure preferred, and Katomaran mentions edge platforms like NVIDIA Jetson alongside Docker and REST APIs. You should be able to deploy a vision model as a service, on a cloud or an edge device.
Explore 4 practice tools →Operate an ML pipeline (train → register → serve → monitor)
Vision teams want the whole training-to-production loop, not one-off runs. Jumio wants end-to-end pipelines "Data to Train to Deploy" with orchestrators like Airflow, NGXP Technologies names MLflow, Weights & Biases and DVC, and Walmart Global Tech wants models deployed and monitored in production. You should be able to version data and models, retrain on new images, and track how the deployed model behaves.
Explore 5 practice tools →Optimize inference cost and latency
Vision models often run on a box in a factory or a car, so speed is a requirement. Jumio wants low-latency inference with quantization, distillation and TensorRT or ONNX, Hansa Aerospace wants you to "own the compute and latency budget", and Qualcomm optimises for image quality, power and latency on target hardware. You should be able to shrink and speed up a model with quantization or pruning and show what it cost in accuracy.
Explore 3 practice tools →Containerize an application with Docker
Docker is the standard wrapper for shipping vision code. KoiReader makes hands-on Docker mandatory, ACV Auctions runs Docker on GCP and Kubernetes, and Infosys lists Flask, FastAPI and Docker for model deployment. You should be able to package a model with its OpenCV and CUDA dependencies into an image that runs the same everywhere.
Explore 2 practice tools →Automate builds, tests and deploys with CI/CD
Automated builds and tests show up in the more production-minded postings. Mindsprint wants "CI/CD for ML, model versioning, monitoring, drift detection", InovarTech wants CI/CD pipelines and testing frameworks for vision models, and Qualcomm and Suitable AI name Jenkins. You should be able to set up a pipeline that tests and ships a vision model on every merge.
Explore 3 practice tools →Build voice or vision LLM features
Vision is starting to meet language. NGXP Technologies wants hands-on experience with vision-language models like CLIP and LLaVA, Auxo AI wants visual grounding and captioning, and GE HealthCare wants VLMs and RAG for healthcare. You should be able to use a vision-language model to search, caption or answer questions about images, and know where it gets things wrong.
Explore 4 practice tools →Build and consume REST APIs
A model nobody can call is a demo. ArcelorMittal wants REST API integrations "connecting CV pipelines to plant systems", Difinity Digital wants models exposed via APIs with backend teams, and Entrupy wants interns helping maintain model-serving APIs. You should be able to wrap a model in a small REST service with FastAPI or Flask and handle image uploads cleanly.
Explore 3 practice tools →Work in annotation tools (Label Studio / CVAT / Labelbox)
Only some employers name labeling tools, but they are specific about it. NewCold wants exposure to CVAT, Label Studio or SuperAnnotate, Ometaura AI wants a dataset pipeline with "annotation tooling (CVAT / Label Studio)" and active learning, and Brivo lists annotation tools beside model versioning. Knowing how to set up a labeling project and check label quality matters, because a vision model is only as good as its boxes.
Explore 3 practice tools →Integrate LLM APIs into an application
Few computer vision postings ask you to call an LLM API; the job is still mostly pixels, not prompts. The ones that do are specific: Miracle Eye wants trainees 'integrating ChatGPT and Claude for enhancing our computer vision solutions', Ripik.AI names GPT-4o vision and Gemini, and NGXP Technologies lists Gemini. Treat it as a useful extra: be able to send an image to a multimodal model API and turn its answer into something your vision pipeline can use.
Explore 4 practice tools →Families: Machine learning & data science · Programming foundations · Cloud, deployment & production · Evaluation, safety & observability · LLM application development · Annotation, quality & human feedback
"AWS" on a JD is not "learn AWS"
The words employers write, translated into what they want you to be able to do for this role.
Skip, for now
- Kubernetes — Named in 5 of 26 job postings, and in each case buried in a list of 30-50 tools (Infosys, Ripik.AI, Mindsprint, ACV Auctions, Fractal). Edge boxes do not run Kubernetes. Learn Docker plus TensorRT/ONNX export first; Kubernetes can wait until a platform team hands it to you.
- Visual SLAM, 3D reconstruction and ROS — In 7 of 26 job postings (drones at Big Bang Boom, AR at Meril and Ctruh, robotics at UMA and Skild AI, medical 3D at Infosys, research at Siemens Energy) - real, but every one except Skild AI's is a 2+ years specialist role. Get a 2D detector shipped first; 3D geometry is a second specialisation, not the entry ticket.
- GANs, diffusion models and generative image models — Appear in 4 of 26 job postings (Infosys medical imaging, Sahana, Siemens Energy, GE Vernova) and only as research or augmentation, never as the product. Ten hours on augmentation with Albumentations pays back more than a month on diffusion.
- RAG, agents and LLM fine-tuning — RAG and agentic workflows appear in exactly 1 of 26 job postings (Difinity) and LoRA fine-tuning in 1 (Ctruh, for multimodal models). If you want that work, look at the AI Engineer role instead - here it is a distraction from FPS and mAP.
- Research publications and patents — Only 3 of 26 job postings reward papers (Siemens Energy, GE Vernova, Infosys medical imaging) and GE Vernova wants a PhD with 7-10 years. The other 23 want a model that runs on a Jetson without dropping frames.
Aerial object detector: train on VisDrone, export to ONNX, benchmark on a budget
Train a YOLO detector on the public VisDrone aerial dataset (10 classes, small objects, harsh lighting - the same problems Big Bang Boom, ArcelorMittal and ShipIn describe) on a free Colab or Kaggle GPU, and do the error analysis a hiring manager will ask about: which classes fail, why small objects are missed, what augmentation fixed. Then do the half of the job that separates candidates here - export the model to ONNX, quantise it to FP16/INT8, and publish a latency-versus-accuracy table for PyTorch, ONNX Runtime and the quantised variant at batch size 1. Wrap the ONNX model in a FastAPI service inside a Docker image and run it over a short drone video with ByteTrack tracking, reporting measured FPS. Write it up so a reader can rerun every number.
- A YOLO model trained on VisDrone with mAP50 and per-class AP on the validation split committed to the README, plus an error-analysis section naming the two worst classes and one augmentation or resolution change that measurably helped
- The model exported to ONNX and quantised (FP16 and INT8), with a table of p50/p95 latency at batch size 1 and the mAP delta for PyTorch vs ONNX Runtime vs the quantised model, on named hardware (Colab T4, Kaggle P100 or a laptop CPU)
- A Dockerised FastAPI endpoint that accepts an image and returns boxes as JSON, and a script that runs the detector with ByteTrack over a 30-second video and reports measured FPS
- A GitHub repo with tests for the pre/post-processing (letterboxing, NMS) and a short 'failure modes' note covering low light, occlusion and tiny objects - the questions ArcelorMittal and ShipIn ask in the JD
What the interviews look like
The rounds you'll actually face, in the order they usually come.
- 1
Screening
Recruiter or lead checks years in computer vision specifically (12 of 26 job postings want 5+; Ripik.AI takes 1-3, Digitap.ai and SeeWise.AI take freshers), PyTorch plus OpenCV, whether you have deployed to a Jetson or other edge device, and C++ comfort for robotics, drone and industrial roles. Startups ask for a repo link; a benchmark table gets you past this round faster than a certificate.
- 2
Technical / take-home
Coding in Python (and C++ where the JD names it) plus CV fundamentals: how YOLO differs from Faster R-CNN, IoU, NMS, mAP, anchor-free detection, class imbalance, camera calibration and homography, how you would detect small objects in aerial imagery. Take-homes are usually 'train a detector on this dataset and report metrics' or 'make this model run at X FPS on CPU'. Expect questions on quantisation, INT8 accuracy loss and TensorRT/ONNX export gotchas.
- 3
System/product design
Design a real-time pipeline: 100+ RTSP camera streams (Infosys), a plant safety zone monitor that must degrade gracefully when cameras drop (ArcelorMittal), or a drone perception stack under a 50 ms budget (Big Bang Boom). They probe frame batching, GStreamer/DeepStream, tracking across cameras, where inference runs (edge vs cloud), model versioning, drift monitoring and false-alarm handling.
- 4
Culture / stakeholder
Working with hardware, firmware and field teams, and often on-site: UMA Robotic wants comfort on an industrial floor, SeeWise.AI and Quest Global want you supporting demos and sales bids, Frinks AI wants you troubleshooting live production-line deployments. Medical and KYC employers (Meril, Infosys, Airtel Africa, Digitap.ai) probe how you handle sensitive images, anonymisation and false-reject costs.
What people ask before choosing this role
Can a fresher get a Computer Vision Engineer job in India?
Yes, this is one of the more reachable AI-era roles. 18 of the 83 job postings behind this page accept 0–2 years of experience. The rest want more, so expect the fresher-friendly openings to be competitive.
What is the salary of a Computer Vision Engineer in India?
Pay at 2–5 yrs averages about ₹10.7 LPA (verified across 2 salary sites: AmbitionBox, Glassdoor); employers offer ₹11–15.5 LPA (median of 8 job postings that state pay, 2–5 yrs · Naukri, LinkedIn, Wellfound). Not every posting states pay, and pay varies widely by city and by whether the employer is an IT-services firm, a global capability centre or a product startup.
How long does it take to become a Computer Vision Engineer?
The six capabilities employers ask for most add up to roughly 167 focused hours — about 21 weeks at 8 hours a week, if you are starting from zero on all of them. Most people are not: the self-check on this page works out what you can skip, which is usually a large part of it.
What skills do you need for a Computer Vision Engineer role?
Across the 83 job postings behind this page, the most-requested capabilities are Build image models (detection/classification) (99% of postings), Write production-quality Python for AI work (89% of postings) and Build and train neural networks in PyTorch (83% of postings). Note these are capabilities, not tools — employers write tool names, but what they are buying is the ability to do the work.
Which cities in India post the most Computer Vision Engineer jobs?
Bengaluru (36), Chennai (8), Hyderabad (8) and Delhi NCR (6) — counted across the 83 job postings behind this page. Remote-India roles are counted separately where the posting said so.
Is demand for Computer Vision Engineer roles in India growing?
Naukri JobSpeak for June 2026 put AI and machine learning roles up 25% year on year while overall IT jobs fell 3%, and computer vision demand is spreading into new sectors: defence and UAV autonomy accounted for 4 of 21 new job postings in one recent research cycle (Hansa Aerospace, NewSpace, Green Aero Propulsion, Auric AI Labs). The role stays senior-heavy — 33 of the 77 job postings collected ask for 5+ years and 14 take 0-2 — while AmbitionBox reports average Computer Vision Engineer pay up 11% over the last two years (1.1k salaries).
Do I need a degree or a paid certificate for this?
Nothing on this page requires a paid certificate, and none of the 83 job postings behind it asked for one by name. What they ask for is evidence you can do the work — a public repo, a deployed project, something a hiring manager can open. That is what the path on this page is built to produce.
Companies with this role open in India
A sample of employers we saw hiring for this role — IT services, global capability centres, product companies and startups. Each links to one of the company's postings for this role, checked open on 04-10-2026.
What it pays
most earn ₹5–10 LPA (base pay) · verified across 2 salary sites: AmbitionBox, Glassdoor
verified across 2 salary sites: AmbitionBox, Glassdoor
Employers offer ₹11–15.5 LPA: median of 8 job postings that state pay, 2–5 yrs · Naukri, LinkedIn, Wellfound
Not enough salary data yet for Senior · 5+ yrs.
Across all levels: the middle half earns ₹5.4–14 LPA · Glassdoor
How much demand
What each job portal shows for this role's title — the readings behind the openings figure above.
Where this role is heading
- Naukri JobSpeak for June 2026 put AI and machine learning roles up 25% year on year while overall IT jobs fell 3%, and computer vision demand is spreading into new sectors: defence and UAV autonomy accounted for 4 of 21 new job postings in one recent research cycle (Hansa Aerospace, NewSpace, Green Aero Propulsion, Auric AI Labs).
- The role stays senior-heavy — 33 of the 77 job postings collected ask for 5+ years and 14 take 0-2 — while AmbitionBox reports average Computer Vision Engineer pay up 11% over the last two years (1.1k salaries).
Capability percentages come from 83 job descriptions read in full on 03-10-2026. How we do this