Pablo García Ruiz logo
Back to blog
5 min read

PnP Solvers Explained: Why Four Points, and Why Corners Matter So Much

What a Perspective-n-Point solve is actually computing under the hood, why three points isn't enough, and why every post on this blog about marker accuracy keeps coming back to corner localization.

Computer VisionPose EstimationExplainer

Last post ended on a claim I made without unpacking it: that ArUco markers use four corners because that’s exactly what a PnP solver needs to compute a pose. This post is that unpacking, what a Perspective-n-Point solve is actually doing, why three points isn’t enough, and why corner accuracy ends up mattering more than almost anything else in the pipeline.

The problem PnP solves

Perspective-n-Point, PnP, takes camera projection and runs it backwards. Forward projection is easy to state: given a camera’s pose (where it is, which way it’s pointing) and a set of known 3D points, predict where each point lands in the 2D image. PnP inverts that: given the camera’s known intrinsics (focal length, principal point, lens distortion, all fixed by calibration) and a set of correspondences, each a known 3D point paired with its observed 2D position in the image, solve for the unknown pose that would have produced those observations.

A camera and a world coordinate system side by side. A point P on a tree in the world projects through the camera's image plane to a pixel p. The transformation between the two coordinate systems is labelled R and t.
The whole problem in one picture: find the rotation and translation that carry a world point into the camera's frame. From my thesis.

That’s it. A marker’s four corners have known positions in the marker’s own coordinate frame (it’s a flat square, you designed it), and a detector finds their pixel positions in the image. PnP is the step that turns those eight numbers (four 2D points) into a full 6DoF camera pose.

A square marker with its four corners labelled P1 to P4 in the marker's own coordinate system, and the same four corners labelled p1 to p4 where they land in the camera's image. The pose relating the two is labelled R and t.
The four corners, in the marker's own frame and in the image. Those two sets of four are the entire input. From my thesis.
What a PnP solver is given: four known 3D points and their observed 2D positions Two panels. On the left, a square marker lying on a ground plane with its four corners numbered, and a camera drawn as a view frustum with a sight line running from its centre through each corner. On the right, the image that camera records: the same four corners as 2D points, numbered to match. Each numbered pair is one correspondence. The marker's corner positions are known by design and the image positions are measured by the detector; the camera's position and orientation are the unknowns the solver recovers from those four pairs. The four correspondences a square marker providesKnown in 3D, observed in 2D; the pose is what's missing world 1234 camera, pose unknown image 1 2 3 4 the four correspondences: known in 3D, observed in 2Dsight lines: the projection PnP runs backwards
Four known 3D points and their observed 2D positions. The pose that produced them is the unknown.

Why three points isn’t enough

A rigid 3D pose has six degrees of freedom: three for rotation, three for translation. Each point correspondence contributes two constraints, its observed x and y image coordinates. Three points, naively, give six equations for six unknowns, which looks like exactly enough.

It isn’t, because those equations aren’t linear. The relationship between a 3D point, a camera pose, and its 2D projection involves perspective division, and solving the minimal three-point case (known as the P3P problem) reduces to a quartic polynomial. A quartic can have up to four valid real roots. Geometrically: given only three points and the angles between them as seen from the camera, there can be as many as four distinct camera positions that are all equally consistent with those observations. Three points narrow the space of possible poses down to a small set of candidates, not down to one answer.

A fourth point is what breaks the tie. You take each of the (up to four) P3P candidate poses, check which one correctly predicts where the fourth point actually landed in the image, and discard the rest. That’s why OpenCV’s solvePnP refuses to run on fewer than four points: below four, the geometry is fundamentally ambiguous, no amount of numerical cleverness resolves it, because the ambiguity isn’t noise, it’s baked into the equations themselves.

This is also, not coincidentally, why a square marker has four corners and not three. It’s the smallest configuration that yields one unambiguous pose. OpenCV even ships a solver mode built specifically around it, SOLVEPNP_IPPE_SQUARE, tuned for exactly this case: four coplanar points in a known square arrangement.

import cv2

# object_points: the marker's 4 corners in its own flat coordinate frame
# image_points: their detected pixel positions in the current frame
success, rvec, tvec = cv2.solvePnP(
    object_points,
    image_points,
    camera_matrix,     # intrinsics from calibration
    dist_coeffs,        # lens distortion coefficients
    flags=cv2.SOLVEPNP_IPPE_SQUARE,
)
Three points admit four different camera poses; the fourth point picks one Two panels. On the left, a square marker with three of its corners marked, and four different camera poses drawn as view frustums. All four are genuine solutions of the three-point pose problem: each one reprojects those three corners to exactly the same pixels, so the three observations cannot distinguish between them. On the right, the image those cameras record. All four agree about corners 1, 2 and 3, but predict the position of the fourth corner up to 112 pixels apart. Only one prediction matches the corner actually observed, which is how a fourth point resolves an ambiguity that is built into the equations rather than caused by noise. Why a square marker needs its fourth cornerGive the solver only corners 1-3 and four different camera poses fit equally well worldimage corners 1-3: given to the solvercorner 4: held back, to test the answers 1 2 3 corner 4, observed 1234predicts corner 4 correctlycandidate pose 1 of 4 1234misses corner 4 by 57 pxcandidate pose 2 of 4 1234misses corner 4 by 97 pxcandidate pose 3 of 4 1234misses corner 4 by 109 pxcandidate pose 4 of 4 1234all four candidate posesone prediction matches; three miss the three given pointscorner 4, and the pose that predicts itposes that fit three points but not the fourth
All four poses reproject corners 1–3 to the same pixels, to within 1e-6 px. Only one predicts where corner 4 actually landed.

What more points actually buy you

Anything beyond a single flat marker, a full 3D fiducial object, a checkerboard, a calibration rig, hands the solver more than the minimal four points. Past the minimum, extra points stop being about disambiguation and start being about averaging out noise. A common approach (EPnP) expresses every 3D point as a weighted sum of four virtual control points, then solves a much smaller linear system for those control points’ positions in the camera frame, sidestepping the combinatorial blow-up of the exact minimal case entirely. An iterative refinement pass afterward (typically Levenberg-Marquardt, minimizing total reprojection error across every point) then polishes the estimate.

Every additional accurate correspondence pulls that least-squares solution closer to the true pose. It’s the same reason combining detections from multiple faces of a fiducial object measurably outperforms any single flat marker in my own results, more independent, accurate points feeding the same underlying solve.

Computer VisionCamera Pose Estimation

Why corner accuracy matters so much

With a minimal four-point marker, there’s no averaging cushion. Every one of those four corner detections has direct, outsized leverage on the final pose, there’s no fifth or fiftieth point around to dilute a bad one. A 1-pixel corner localization error doesn’t stay a 1-pixel error once it propagates through the solve: because PnP is inverting a perspective projection, that same pixel error maps to a larger real-world positional error the farther the marker sits from the camera (fewer pixels per real-world centimeter at range), and a larger rotational error the more oblique the viewing angle (foreshortening compresses the corners together in the image, so a fixed pixel error represents a bigger fraction of the observed geometry).

That’s the actual reason marker-detection pipelines bother with subpixel corner refinement after the initial coarse detection, fitting a local edge or gradient model instead of trusting the first estimate. And it’s the same reason the calibration step in my Fiducial Objects work exists at all: a 3D-printed marker’s real physical corners never perfectly match its digital design, and that mismatch feeds into the exact same PnP solve as if it were ordinary detection noise. It has to be measured and corrected before the pose estimate can be trusted.

Corner uncertainty grows relative to a marker as the viewing angle steepens A square fiducial marker projected by a camera orbiting from straight-on to 72 degrees off the marker's normal. Four discs of fixed radius, representing 3 pixels of corner-detection uncertainty, stay the same size throughout. The projected marker foreshortens until its shortest edge is about a third of its straight-on length, so the same 3-pixel uncertainty grows from roughly 1.5 percent of that edge to over 4 percent. The detector has not got worse; the geometry available to the pose solver has shrunk. Corner uncertainty vs. viewing angleSame detector, same 3 px uncertainty, shrinking geometryimage plane VIEWING ANGLE SHORTEST EDGE 3 PX AS SHARE OF IT 0.0°200.0 px1.50 % 0.9°199.6 px1.50 % 3.6°198.5 px1.51 % 7.9°196.8 px1.52 % 13.6°194.6 px1.54 % 20.4°188.0 px1.60 % 28.0°177.5 px1.69 % 36.0°163.2 px1.84 % 44.0°145.8 px2.06 % 51.6°126.7 px2.37 % 58.4°107.7 px2.79 % 64.1° 90.9 px3.30 % 68.4° 77.7 px3.86 % 71.1° 69.5 px4.32 % 72.0° 66.7 px4.50 % 71.1° 69.5 px4.32 % 68.4° 77.7 px3.86 % 64.1° 90.9 px3.30 % 58.4°107.7 px2.79 % 51.6°126.7 px2.37 % 44.0°145.8 px2.06 % 36.0°163.2 px1.84 % 28.0°177.5 px1.69 % 20.4°188.0 px1.60 % 13.6°194.6 px1.54 % 7.9°196.8 px1.52 % 3.6°198.5 px1.51 % 0.9°199.6 px1.50 % projected marker3 px corner uncertaintystraight-on reference
The uncertainty discs never change size. The marker shrinks around them, so the same 3 px becomes 4.5% of the shortest edge instead of 1.5%.
One pixel of corner error costs far more range than it does lateral position A chart of the real-world position error produced by a single pixel of corner-detection error, against the distance from camera to marker, for a 10 centimetre marker and a 1000 pixel focal length. Lateral error grows linearly and reaches about 5 millimetres at 5 metres. Range error, inferred from the marker's apparent size, grows quadratically and reaches about 250 millimetres at the same distance -- roughly fifty times worse. The gap widens continuously with distance. What one pixel of corner error costs100 mm marker, 1000 px focal length, 1 px corner error05010015020025012345distance to marker (m)position error (mm) 250 mm along the axis5 mm sidewaysrange error (towards / away) grows as distance squaredlateral error (sideways) grows linearly
The same pixel of error, priced in millimetres. Sideways it grows linearly; along the optical axis it grows with the square of distance.

The takeaway

Every fiducial-marker system I’ve written about on this blog, ArUco, the custom fiducial objects from my thesis, comes down to the same core operation: turn a handful of accurately-localized 2D-3D correspondences into a 6DoF pose via PnP. What actually differs between them is how many points they hand the solver, how precisely those points can be localized, and from how many viewpoints at once, not what happens once those points reach the solver itself.