Computer Vision Development: How to Build AI Applications That Understand Images
How to build computer vision applications: choosing the task (classification, detection, segmentation, OCR), data and labelling, pretrained models versus training, evaluation, cloud and edge deployment, licensing and monitoring.
Quick answer
Computer vision development starts by choosing the right task (classification, detection, segmentation, OCR or tracking) for the business question, then gathering images that reflect real conditions, labelling them consistently and trying pretrained models or vision APIs before custom training. Evaluate per class on a held-out set from real conditions, deploy to cloud or edge depending on latency, connectivity and privacy, check model and data licences, and monitor drift as cameras, lighting and products change.
Where This Fits
Building classifiers and detectors in depth is covered in AI image recognition. Combining images with text and audio is multimodal AI, and production tracking is AI model monitoring. For a sector view, see AI agents in manufacturing.
Choosing the Task
Approach Options
| Approach | Fits | Trade-offs |
|---|---|---|
| Cloud vision APIs | Common tasks: labels, OCR, faces, landmarks | Fast start; limited customization; data leaves your environment |
| Multimodal LLMs | Open-ended questions about images | Flexible; slower and costlier; harder to evaluate |
| Pretrained models fine-tuned | Specific objects or defects | Needs labelled data; strong accuracy |
| Custom models from scratch | Unusual imagery at scale | Most data and expertise |
Data and Labelling
Collect images from the real cameras, angles, lighting and conditions the system will see, including difficult cases. Write labelling guidelines with examples of edge cases, use consistent annotation tools, check agreement between labellers and keep a held-out test set that is never used for training. Confirm you have rights to use the images and handle people in images according to privacy law.
Taxonomies, agreement and quality control for labelling are covered in AI data annotation.
Have an inspection, counting or recognition problem?
ZSpace Labs builds computer vision applications from data collection and model selection to edge or cloud deployment.
Deployment: Cloud, Edge or Hybrid
Edge deployment runs models on devices, cameras or local servers, with frameworks such as LiteRT (formerly TensorFlow Lite), Core ML or ONNX Runtime, giving low latency, offline operation and less data leaving the site. Cloud deployment simplifies updates and supports larger models. Hybrid designs run fast checks at the edge and send uncertain cases to the cloud or to people. See AI-powered mobile apps for on-device options.
Device hardware, compression, offline sync and signed updates are covered in AI edge deployment.
Licensing and Compliance
For example, Ultralytics publishes its licensing options for YOLO models.
- Check model licences: some popular detection libraries, such as Ultralytics YOLO, use AGPL-3.0 with a separate enterprise licence for proprietary use
- Check dataset licences and rights to customer or partner images
- Assess privacy for people in images; avoid biometric processing unless lawful and necessary
- Some uses, such as biometric identification, are heavily regulated (for example under the EU AI Act)
Monitoring and Drift
Vision models degrade when the world changes: new packaging, camera replacement, seasonal light, dirty lenses. Monitor confidence distributions, sample predictions for human review, track downstream outcomes and retrain with new labelled data when performance drops.
Advantages and Limitations
Computer vision automates visual inspection and data capture at a scale and consistency people cannot match. It is sensitive to data conditions, needs labelled examples for specialised tasks, raises privacy questions when people are in view and requires ongoing monitoring.
How to Build a Vision Application Step by Step
- 1. Define the decision the system supports and the cost of errors
- 2. Choose the task type
- 3. Collect and label realistic images
- 4. Test APIs and pretrained models, then fine-tune if needed
- 5. Evaluate per class on a held-out set
- 6. Choose deployment and integrate with workflows
- 7. Monitor and retrain
Video Analytics
Many vision applications process video: counting people or vehicles, monitoring production lines, checking safety equipment. Process on edge devices near cameras where bandwidth or privacy matters, sample frames at the rate the task needs, track objects across frames to avoid double counting and store events rather than raw footage where possible. Monitoring people raises privacy and employment law questions; assess them before deployment.
Cameras, Lighting and Hardware
Many vision projects are won or lost on capture. Fixed mounting, consistent lighting, appropriate resolution and lens choice can matter more than model choice. For edge deployment, choose hardware that supports your framework and accelerators, plan for heat, dust and updates in the field and monitor device health alongside model performance. Building recognition models specifically is covered in AI image recognition.
Pretrained Models, APIs and Custom Training
Cloud vision APIs handle common tasks such as label detection, OCR and face detection without training. Multimodal language models can answer open questions about images. Open-source detection and segmentation models can be fine-tuned on your data. Each step along this path gives more control and accuracy for specific tasks, at the cost of more data and engineering.
Start with the simplest option that meets accuracy and cost needs on your own images, measured properly. Move to custom training when you need specific classes, consistent latency, edge deployment or lower cost at high volume. Check model licences carefully: some popular detection frameworks use copyleft licences that affect commercial distribution.
People in Images
Vision systems that capture people raise privacy, consent and discrimination issues. Biometric identification is heavily restricted in many jurisdictions, and the EU AI Act prohibits some uses, such as untargeted scraping of facial images to build recognition databases and emotion recognition in workplaces and education, with limited exceptions.
Where people appear incidentally, minimize: blur faces, avoid storing raw footage, process on the edge and keep only events or counts. Run a privacy impact assessment before deployment and inform people where required. More on privacy design in AI data privacy.
The Commission has published guidelines on prohibited AI practices.
Typical Business Applications
| Domain | Task | Typical approach |
|---|---|---|
| Manufacturing | Defect detection | Custom detection or classification, edge deployment |
| Retail | Shelf monitoring | Detection plus planogram comparison |
| Logistics | Label and damage checks | OCR plus classification |
| Insurance | Damage assessment support | Multimodal model with human review |
| Agriculture | Crop and pest identification | Classification on mobile or drone imagery |
| Construction | Progress and safety checks | Detection with privacy safeguards |
Worked Example
An illustrative scenario, not a client case: a packaging line wants to detect misprinted labels. A general vision API misses subtle print defects, so the team fine-tunes a detection model on a few thousand labelled images captured from the line camera, runs it on an edge device next to the line and sends uncertain items to an operator screen. Weekly samples are reviewed to catch drift after packaging changes.
Common Mistakes
- Training on images unlike production conditions
- Overall accuracy hiding poor performance on rare classes
- Ignoring model and dataset licences
- No plan for drift
- Processing people's images without a privacy assessment
Planning a computer vision project?
Talk to ZSpace Labs about computer vision development and on-device AI apps.
Conclusion
Computer vision works when the task is well chosen, data reflects reality, evaluation is per class and deployment fits the environment. Related: AI image recognition and multimodal AI.
Common questions
Building software that interprets images or video, for tasks such as classifying images, detecting and counting objects, segmenting regions, reading text and tracking movement, and integrating those results into applications.