What Is Computer Vision and How Can You Use It for 3D Reconstruction?
Computer vision is the field of making software extract meaning from images and video: recognizing objects, tracking motion, and reconstructing 3D structure from 2D pixels. For 3D reconstruction specifically, the practical branch is photogrammetric computer vision — using many overlapping photos of the same subject to recover its shape, camera positions, and surface texture. AliceVision describes itself as exactly this: a "Photogrammetric Computer Vision framework for 3D Reconstruction and Camera Tracking." If you have a camera and a subject you can walk around, you can try this workflow today; if your subject is moving, transparent, or textureless, expect problems.
What computer vision actually does
At a high level, computer vision turns pixel arrays into structured information. The tasks most relevant to 3D work are:
- Feature detection and matching — finding distinctive points (corners, blobs) in each image and matching them across images.
- Camera tracking (structure from motion) — estimating where each camera was positioned and oriented, and building a sparse point cloud of the scene.
- Dense reconstruction — turning the sparse cloud into a detailed surface (depth maps or a mesh).
- Texture mapping — projecting the original photos back onto the geometry so the model looks real.
Other computer vision tasks — classification, object detection, segmentation — are related but separate. They answer "what is in this image," while photogrammetric computer vision answers "what shape is this, and where was the camera."
Photogrammetric vs. general computer vision
| Dimension | General computer vision | Photogrammetric computer vision |
|---|---|---|
| Input | Often a single image or video frame | Many overlapping photos of one static scene |
| Output | Labels, boxes, masks, tracks | 3D points, camera poses, mesh, texture |
| Key assumption | Learned patterns generalize | The scene is rigid and visible from multiple angles |
| Typical tools | Neural network frameworks | AliceVision, Meshroom, COLMAP, OpenMVG |
The distinction matters because photogrammetry fails for reasons general CV does not: not enough overlap, motion between shots, or surfaces with no distinguishable features.
The standard 3D reconstruction workflow
- Capture — photograph the subject from many angles with heavy overlap (commonly cited guidance is roughly 60–80% overlap between neighboring shots). Keep lighting consistent and the subject still.
- Feature extraction and matching — the software finds keypoints in each image and links them across views.
- Camera tracking / structure from motion — it solves for camera positions and a sparse 3D point cloud.
- Dense reconstruction — depth maps are computed and fused into a dense surface.
- Meshing and texturing — the surface is converted to a mesh and the photos are projected onto it.
Each stage depends on the previous one. If camera tracking fails, everything downstream fails too.
Getting started with AliceVision and Meshroom
AliceVision is the underlying framework; Meshroom is the graphical front end that runs AliceVision's pipeline without requiring you to script each step. The practical entry point for a beginner is Meshroom: you load a folder of photos, start the pipeline, and inspect the resulting 3D model. AliceVision itself is the library you would use if you want to integrate the algorithms into your own software.
A minimal first attempt looks like this:
- Shoot 30–100 photos of a small object, walking a full circle with overlap.
- Import the folder into the graph-based interface.
- Run the default pipeline and watch which nodes complete and which fail.
- If camera tracking fails, the problem is almost always the photos, not the settings.
Why reconstructions fail
- Insufficient overlap — gaps between viewpoints leave no matched features.
- Lighting changes — moving shadows or changing exposure break feature matching.
- Textureless surfaces — plain walls, glass, and shiny metal give the matcher nothing to lock onto.
- Motion — anything that moves between shots violates the rigid-scene assumption.
- Scale ambiguity — a reconstruction from photos alone has no absolute scale unless you add measurements or known reference points.
Where this is actually used
- Cultural heritage digitization — turning artifacts and sites into archival 3D records.
- Architecture and surveying — capturing buildings and terrain from drone or ground photos.
- Game and film asset creation — generating textured models from real objects.
- Camera tracking for VFX — recovering camera motion to composite CGI into live footage.
If your goal is to learn the pipeline, start with a small, well-lit, textured object and a full circle of overlapping photos. If your goal is a production asset, plan the capture carefully — the reconstruction quality is decided before the software ever runs.