Pablo García Ruiz logo
Back to blog
4 min read

Sparse vs. Large-Scale Indoor Camera Positioning: Where My Two Papers Actually Differ

Two papers, one author, both about positioning networks of indoor cameras with fiducial markers, similar enough that people conflate them constantly. Here's the actual distinction, and the result that surprised me.

Computer VisionPose EstimationResearch

I get this mix-up more than almost any other question about my work: “wait, is the sparse one the big one, or the other way around?” Fair, honestly, they’re two papers by the same author, from the same lab, about positioning networks of fixed indoor cameras using the same core idea, printed fiducial markers, physically relocated through the space over time. It’s an easy pair to blur together. But they differ on one specific assumption, and the way they differ turned out to matter more than I expected going in.

The shared setup

Both papers solve the same underlying problem: you have a large number of fixed cameras installed around an indoor space, and you need each one’s precise 6DoF pose without surveying every one by hand and without requiring every camera to see every other camera. Both use the same mechanism to get there: fiducial markers, printed on regular paper, placed and then progressively relocated deeper into the space, feeding a final bundle-adjustment-style optimization that jointly refines every camera pose by minimizing reprojection error across all observed markers, while enforcing real-world physical constraints like marker and camera coplanarity. Same core toolchain, too:

C++OpenCVCMakeQt Creator

Where they actually part ways is a single assumption: do nearby fixed cameras already share an overlapping view or not?

Large-Scale Indoor Camera Positioning (2024): the assumption that stayed

Large-Scale Indoor Camera Positioning, published in Sensors, still leans on local overlap. Markers are placed wherever nearby fixed cameras already share a view, establishing a pairwise spatial relationship directly, then progressively repositioned deeper into the space to extend that chain across cameras that never directly see each other. The scaling result is real and non-trivial: the same 9 physical markers, reused throughout, were enough to jointly position 51 cameras across a 214 m² floor to roughly 15 cm accuracy. Earlier methods either needed a fresh marker per camera pair or broke down entirely once the overlap ran out; this one didn’t.

Sparse Camera Positioning (2025): dropping the assumption entirely

Sparse Camera Positioning, in Applied Sciences, removes the overlap requirement outright, including between neighboring cameras. Instead of relying on any pair of fixed cameras sharing a view, a mobile camera physically carries the marker relationships from one fixed camera to the next: it follows the markers as they relocate, recording each local configuration as it goes, and those recordings chain together into one connected graph spanning the whole network. Two fixed cameras that never see each other, even ones standing right next to each other with no shared sightline, still end up positioned in the same coordinate frame.

That’s a strictly harder problem, and the paper validates it on networks deliberately built with zero overlapping pairs: 6 cameras through a corridor, 22 across a full floor, 42 across a two-story building, holding translation error in the 13-42 cm range as the network’s physical span grows.

The overlap assumption that separates the 2024 and 2025 camera positioning methods Two three-dimensional views of the same room, each with five wall-mounted cameras. On the left, the 2024 method: wide lenses, so the patches of floor the cameras cover overlap their neighbours. A printed marker is placed in each shared patch, which links that pair of cameras directly. Nine reused markers positioned 51 cameras over 214 square metres to about 15 centimetres. On the right, the 2025 method: narrow lenses, so no two floor patches touch and no direct pair exists anywhere in the network. A mobile camera walks a route across the floor carrying marker configurations between them, and that route supplies every link instead. It handles networks with zero overlapping pairs, up to 42 cameras across two storeys, at 13 to 42 centimetres. Do nearby cameras already share a view?The single assumption the two papers differ on 2024 · SensorsWide lenses, so neighbours already share floor. A marker ineach shared patch links that pair. 9 reused markers, 51 cameras,214 m², ~15 cm. 2025 · Applied SciencesNarrow lenses, zero overlapping pairs. A mobile camera carriesthe relationships between them: up to 42 cameras, two storeys,13-42 cm. mobile camera fixed camera and the floor it coversprinted marker, reused and relocatedmobile camera route (2025 only)
The 2024 method needs an overlap for every link. The 2025 method has a mobile camera carry the relationships instead.

The part that actually surprised me

Here’s the detail that gets lost in a “sparse vs. large-scale” framing, as if they’re just two different use cases. When the 2025 method’s refinements were applied back to the exact same dense, overlapping-view floor dataset from the 2024 paper, the error dropped from 14.82 / 15.72 cm down to 5.22 / 4.88 cm. Roughly a 3x improvement, on the dataset the earlier method was already built for and already solved.

That means the 2025 paper isn’t a niche variant that only helps once you’re stuck without overlap. It’s a strict generalization: everything the 2024 method could do, it does more accurately, and it also handles an entire class of networks the earlier method couldn’t touch at all.

The 2025 method beats the 2024 one on the 2024 paper's own dataset A bar chart of positioning error on the same dense, overlapping-view floor dataset. The 2024 method reported 14.82 and 15.72 centimetres, averaging 15.3. The 2025 method's refinements, applied to that same data, reported 5.22 and 4.88 centimetres, averaging 5.1 -- about three times better. Because the 2025 method also handles networks with no camera overlap at all, which the 2024 method could not address, this makes it a strict generalisation rather than a specialised alternative. Same dataset, three times more accuratePositioning error on the 2024 paper's own overlapping-view floor0510152024 method2025 refinementstranslation error (cm) 15.3 cm 5.0 cm 3.0x better mean of the two reported figuresthe individual figures
The same overlapping-view floor dataset, measured both ways.

So why does the 2024 paper still matter?

Mostly because the piece both papers actually depend on, reusing a small, fixed set of physical markers to cover an arbitrarily large camera network instead of surveying each camera individually, is what the 2024 paper established first. The 2025 paper’s real contribution is specifically about how those markers get relayed across cameras with no shared sightline at all, not about reuse itself. Read together, they’re less “two competing methods” and more one continuous idea worked through in two steps: 2024 showed that a handful of reused markers could scale to a large camera network at all; 2025 removed the last assumption, local overlap, standing between that idea and a genuinely general-purpose method.

The actual distinction, in one line

If you only take one thing from this: the 2024 paper assumes nearby cameras already share a view and shows that reused markers scale from there; the 2025 paper assumes nothing about camera overlap at all, and gets there with a mobile camera relaying marker configurations between every pair, more accurately, even on the case the 2024 paper was built for.

Both live under the same thesis, alongside Fiducial Objects and Markie, if you want the fuller picture of how one problem, camera pose estimation with fiducial markers, kept branching into new directions over four years.