Computer Vision definition
Computer vision is a field of artificial intelligence that enables software to interpret images and video. Using models trained on labeled visual data, computer vision systems can classify pictures, detect and locate objects, read text, measure dimensions and track movement, powering applications such as quality inspection, medical imaging, retail analytics and driver assistance.
How does computer vision work?
A digital image is a grid of pixel values. A computer vision model, usually a convolutional neural network or a vision transformer, passes that grid through layers that detect edges, then textures, then parts, then whole objects. The model is trained on labeled images: boxes drawn around cars, masks traced over tumors, or a single label per photo. After training, it applies the same learned patterns to new images in milliseconds.
Before the model sees an image, a pipeline usually resizes it, normalizes colors and sometimes crops a region of interest. After inference, post-processing filters weak detections, merges overlapping boxes and converts outputs into business events such as "defect found on unit 4471". This surrounding code often takes as much engineering effort as the model itself.
Core computer vision tasks
- Image classification: one label per image, such as ripe or unripe fruit.
- Object detection: finds and boxes each object, using model families such as YOLO or DETR.
- Segmentation: labels every pixel, used for medical scans and satellite imagery.
- Optical character recognition (OCR): reads printed or handwritten text.
- Pose estimation and tracking: follows people or objects across video frames.
- Visual similarity search: matches products or parts using image embeddings.
Computer vision vs image processing
Image processing transforms pixels: resizing, sharpening, adjusting contrast or removing noise, often with libraries such as OpenCV. Computer vision extracts meaning: what is in the picture, where it is and what it is doing. Most vision systems use both, with image processing preparing frames and a learned model interpreting them. Classic rule-based techniques still work well in tightly controlled scenes, such as measuring parts against a fixed backdrop where lighting and position never change.
Example: quality inspection on a production line
A packaging plant mounts a camera over its conveyor and collects images of good and faulty cartons: crushed corners, misprinted labels, missing seals. Engineers label a few thousand images, fine-tune a pretrained detection model and deploy it on an edge device such as an NVIDIA Jetson next to the line. Faulty cartons trigger a reject arm, and borderline cases are saved for human review and later retraining. The plant also tracks how often the reject arm fires, which gives early warning of drift.
Challenges and how to start
Vision models are sensitive to conditions they did not see in training: new lighting, a different camera angle, a product redesign or dust on the lens. Collecting images from real operating conditions matters more than collecting many images in a lab. Privacy also needs attention when cameras capture people, and laws such as GDPR apply to faces and vehicle number plates.
Start with a narrow, high-value task, a fixed camera position and a clear accuracy target. Nexzem's computer vision team usually runs a short pilot on the client's own footage before choosing between cloud inference and edge hardware, because latency, bandwidth and privacy needs often decide that choice more than raw accuracy does.