QR Codes vs. ArUco vs. Fiducial Objects: What's Actually Different
Three black-and-white square patterns that get lumped together constantly, and are not the same tool: what QR codes, ArUco markers, and the custom fiducial objects from my own research actually optimize for.
People ask me some version of this a lot: “so a fiducial marker is basically a QR code, right?” It’s a fair guess, they’re both black-and-white squares you point a camera at, and both eventually spit out an ID. But a QR code, an ArUco marker, and the fiducial objects I spent a good chunk of my Ph.D. on are optimized for three genuinely different jobs. Confusing them usually means picking the wrong tool.
QR codes: built to carry data
A QR code is a data-storage format first. Its dense grid of modules encodes bytes, a URL, a string, whatever, using Reed-Solomon error correction so it can still be read with up to 30% of the code damaged or obscured. The three big squares in the corners (the finder patterns) tell a scanner where the code is and how it’s rotated; the smaller alignment patterns correct for perspective distortion at read time.
What QR codes are not designed for is precise geometry. A phone’s QR scanner needs to know roughly where the square is so it can decode the modules inside it, not the sub-pixel position of each corner. You can hack a 6DoF pose out of a QR code’s finder-pattern corners if you have to, but you’re fighting the format, not using it as intended.
ArUco markers: built for precise pose
ArUco markers (from Garrido-Jurado et al., developed at the same university group I later joined) flip the priority. The interior grid encodes a short binary ID, usually just enough bits to distinguish a few hundred markers, nowhere near QR-code payload size, because carrying data was never the point. The point is the four outer corners: their positions are what a detector optimizes for, because those four 2D-3D correspondences are exactly what a PnP (Perspective-n-Point) solver needs to compute a camera’s full 6DoF pose relative to the marker.
That’s why ArUco (and its cousins, AprilTag, ArUco’s own successor formats) shows up everywhere in robotics, camera calibration, and AR: a single flat marker, correctly detected, tells you exactly where the camera is standing and which way it’s looking. OpenCV ships an aruco module for exactly this. It’s the workhorse marker, and it’s the one I use as the teaching example in my fiducial markers workshop for anyone learning this from scratch.
The catch: it’s still one flat square. Viewing angle and distance have hard limits, past a certain angle the corners get too foreshortened to localize accurately, and past a certain distance the whole marker is a handful of pixels.
Fiducial objects: built for robustness across viewpoints
This is where my own research picked up. Fiducial Objects, the paper that became a chunk of my thesis, asks: what if instead of one flat marker, you wrap custom-shaped markers around an entire 3D solid, a cube, a dodecahedron, an icosahedron?
The naive version of that idea already exists (people glue square ArUco markers onto the faces of a cube), but a square marker wastes most of a non-square face’s usable area. So instead, each face gets a marker shape generated specifically to fill it edge to edge, maximizing how much detectable pattern is visible from a distance or at a steep angle. A calibration pass then corrects for the fact that no 3D-printed object is manufactured with perfect geometric precision, snapshotting the physical object from several viewpoints to estimate its true vertex configuration. At run time, detections from every visible face combine into one pose estimate that’s more accurate and far more robust than any single flat marker could give you.
The results held up: across seven solid geometries, the best-performing shapes hit roughly double the pose-estimation precision of standard square markers, and the gap widens under harsh conditions, at heavy blur the dodecahedron kept a 100% detection rate against 73% for the icosahedron. The idea later became the basis for Markie, a fiducial-object input device built for intuitive, accessible interaction.
Side by side
| Property | QR Codes | ArUco Markers | Fiducial Objects |
|---|---|---|---|
| Optimized for | Carrying data | Precise single-view pose | Robust multi-view pose |
| Typical payload | Up to ~3KB of text/bytes | A short binary ID (no real data) | An ID per face, same idea as ArUco |
| Geometry | Flat, fixed square grid | Flat square, corners optimized for localization | Custom shape per face of a 3D solid |
| Corner precision | Not a design goal | Core design goal, feeds a PnP solve | Same, extended across many faces/angles |
| Robustness to viewing angle | N/A, not the point | Limited, one flat plane | High, multiple faces cover more of the sphere of viewpoints |
| Typical use | Menus, links, payments, general data | Robotics, camera calibration, single-marker AR pose | Handheld devices, HCI, close-range tracking from many angles |
Which one do you actually need?
- Need to encode a URL, text, or a payload someone scans with their phone? QR code. Nothing else here is built for that job.
- Need a single, well-understood 6DoF pose from one flat surface, robotics, camera calibration, a tabletop AR anchor? ArUco (or AprilTag). It’s the standard for a reason, and it’s the one I’d point a beginner toward first, my Tiny Introduction to Fiducial Markers workshop walks through exactly this, from detection to a live 3D render, with runnable code.
- Need a physical object tracked reliably from many angles and distances, a handheld controller, a close-range HCI device, something a camera has to recognize no matter how it’s rotated? That’s the fiducial-object case, and it’s genuinely a harder problem than either of the other two, which is why it took a Ph.D. thesis to work through it properly.
They look similar because they share a common ancestor (square markers detected via computer vision), but “which black-and-white square do I need” really means “what am I actually trying to measure.” Once that’s clear, the right answer usually is too.
