What Is Photogrammetric Computer Vision and How Does It Reconstruct 3D Scenes?
Photogrammetric computer vision is the branch of computer vision that recovers camera positions and 3D scene structure from multiple overlapping photographs. It applies when you can move a camera around a static subject and capture enough views; it does not work from a single image or from photos with little overlap. AliceVision is an open-source framework in this space, providing the underlying photogrammetric and computer vision algorithms, while Meshroom exposes those algorithms through a graphical pipeline.
How it differs from general computer vision
General computer vision covers tasks like classification, detection, and segmentation, often on single images. Photogrammetric computer vision is narrower and more geometric: its goal is to estimate where each camera was in space and what the scene looks like in 3D.
The key input is multiple views of the same scene from different positions. The key output is a camera pose per image plus a 3D representation of the scene.
The core reconstruction pipeline
A typical photogrammetric pipeline runs through several stages. Each stage feeds the next, and errors early on propagate forward.
1. Feature extraction and matching
The system finds distinctive points in each image (corners, blobs, textured patches) and matches them across images. A point seen in several photos becomes a tie point linking those views.
2. Structure from Motion (SfM)
SfM solves for camera positions, orientations, and a sparse set of 3D points simultaneously. This is the step that turns a pile of photos into a geometrically consistent camera network. The result is a sparse point cloud and calibrated camera poses.
3. Multi-View Stereo (MVS)
MVS takes the calibrated cameras and dense-matches pixels to produce a dense point cloud covering surfaces, not just isolated features.
4. Meshing and texturing
The dense cloud is converted into a mesh, and the original photos are projected onto it to create a textured 3D model.
Why camera tracking and calibration matter
Camera tracking (pose estimation) and calibration are not optional extras—they are the backbone. Without accurate camera poses, dense matching has no consistent geometry to work from. Calibration handles lens distortion and intrinsic parameters so that straight lines in the world project correctly.
This is also why overlapping photos are required. Each surface point should appear in several images from different angles. More overlap means more constraints and a more reliable reconstruction.
Where AliceVision and Meshroom fit
AliceVision is described as a Photogrammetric Computer Vision framework for 3D Reconstruction and Camera Tracking. In practice it serves as the library and algorithm layer: the photogrammetry, reconstruction, and tracking components.
Meshroom is the graphical front end built on AliceVision. It lets you run the pipeline as a node graph rather than calling libraries directly, which makes the same underlying reconstruction accessible without writing code.
The framework's stated scope also includes HDR and panorama handling alongside 3D reconstruction and camera tracking, so it is not limited to a single output type.
Typical uses and where it fails
Common applications include:
- Cultural heritage and artifact scanning — capturing objects or sites as textured 3D models
- Architectural and site surveying — reconstructing buildings or terrain from photo sets
- Object modeling — turning a photographed object into a mesh
Known failure conditions:
- Reflective surfaces — mirrors, glass, and polished metal break matching because the appearance changes with viewpoint
- Textureless surfaces — plain walls or uniform objects give few reliable features to match
- Insufficient overlap or motion blur — weak or inconsistent tie points lead to failed or distorted reconstruction
If your subject is shiny, featureless, or you cannot get overlapping views, photogrammetric reconstruction is the wrong tool; structured light or laser scanning is usually the alternative.
Choosing this approach
Use photogrammetric computer vision when you have a static subject, a camera you can move around it, and enough surface texture for matching. Use an open-source framework like AliceVision when you want the algorithms as a library, or Meshroom when you want the same pipeline through a visual graph. Expect the method to struggle on reflective or textureless subjects regardless of which front end you choose.