What Is Photogrammetry and How Does It Turn Photos Into 3D Models?

Photogrammetry is the process of measuring and reconstructing 3D geometry from overlapping 2D photographs. You take many photos of an object or scene from different positions, and software finds matching points across those images to calculate where the camera was and where each point sits in 3D space. The result is typically a point cloud, a mesh, and a textured model. It works best when your photos overlap heavily, are sharp, and are lit consistently — and it struggles when images are blurry, poorly lit, or shot with too little overlap.

The core idea in plain terms

Every photo is a 2D projection of a 3D scene. If you photograph the same physical point from two or more positions, that point appears at different pixel locations in each image. The software identifies these corresponding points, then solves for two things at once:

  • Camera pose — where each camera was and which way it was pointing.
  • 3D point position — where the matched point actually sits in space.

This is the same principle your eyes use for depth perception: two slightly different views let you triangulate distance. Photogrammetry scales that up to dozens or hundreds of views.

The typical pipeline

Most photogrammetry tools follow a similar sequence, though the names and automation level differ.

1. Image capture

You photograph the subject from many angles with consistent exposure. Overlap between neighboring shots is what makes matching possible.

2. Feature detection and matching

The software finds distinctive points (corners, edges, textures) in each image and matches them across photos. Flat, textureless surfaces give it little to work with.

3. Camera tracking (sparse reconstruction)

Using the matches, it estimates camera positions and orientations and builds a sparse point cloud — a rough skeleton of the scene.

4. Dense reconstruction

From the sparse result and the original pixels, it computes a dense point cloud covering the surfaces.

5. Meshing and texturing

The dense cloud is converted into a mesh (triangles), and the original photos are projected onto it to produce a textured model.

AliceVision describes itself as a "Photogrammetric Computer Vision framework for 3D Reconstruction and Camera Tracking," and Meshroom is its graphical front end — so the pipeline above maps directly onto what that toolchain does.

What makes a good photo set

The quality of your input determines the quality of your output more than any setting.

Factor Good practice Why it matters
Overlap 60–80% between neighboring shots Matching needs the same points in multiple images
Sharpness Use a tripod or fast shutter; avoid motion blur Blurry features can't be matched reliably
Lighting Even, diffuse light; avoid harsh shadows and glare Shadows move between shots and confuse matching
Angles Circle the subject; vary height Depth needs views from genuinely different positions
Texture Include surfaces with visible detail Blank walls and shiny objects give few features
Coverage Photograph all sides you need modeled You can't reconstruct what you never captured

Typical outputs

  • Sparse point cloud — camera positions plus a rough set of 3D points.
  • Dense point cloud — a detailed set of surface points.
  • Mesh — a triangulated surface, often cleaned and decimated.
  • Textured model — the mesh with photographic texture applied, ready for viewing or export.

Choosing a tool

AliceVision/Meshroom is an open-source photogrammetric computer vision framework with a node-based workflow, suited to users who want control over each pipeline stage and don't mind a learning curve. Other photogrammetry tools range from fully automated consumer apps to professional suites; the right pick depends on how much control you want, whether you need open-source licensing, and how large your datasets are. If you want to understand the pipeline stage by stage, a node-based framework like Meshroom makes each step visible; if you just want a model quickly, a more automated tool will get you there with less setup.

Common failure points

  • Blurry or low-overlap images — the most frequent cause of failed or patchy reconstructions.
  • Reflective, transparent, or textureless surfaces — matching finds too few reliable features.
  • Inconsistent lighting — moving shadows break matches between shots.
  • Too few angles — gaps in coverage become holes in the model.
  • Scale ambiguity — reconstruction gives relative shape; absolute size needs a known reference measurement in the scene.

If a reconstruction fails, the fix is usually in the photos, not the settings: add more overlap, improve lighting, and remove blurry shots before reprocessing.

alicevision.org
AliceVision is a Photogrammetric Computer Vision framework for 3D Reconstruction and Camera Tracking.
ccilt.com
Cornforth Consultants - Landslide Technology