Pumpire: Unified Benchmark for Metric Distance Estimation
Siyu Chen, Zehan Wang, Jiayang Xu, Yihan Wu, Jialei Wang, Junming Chen, Ziang Zhang, Yutong Ying, Zhou Zhao
cs.CV
2026-10-09
Pumpire scores 29 3D setups on tape-measured pairs from 100 scenes (6,400 frames). Normal-surface RealSense δ1.05 is 79.3%; best RGB estimator Metric3D v2 reaches 31.0%.
Grasping and collision checks need the distance in meters between two points. Depth leaderboards and intrinsics leaderboards score those pieces apart. Point-cloud benchmarks report Chamfer Distance and F1. Spatial-reasoning benchmarks stop at object-level questions. None of that is the Euclidean length you get after back-projecting two pixels. Single-image and multi-frame models, and RGB-only estimation versus depth-prior completion, have also lived on separate datasets. Pumpire scores all four settings with one point-pair protocol.
Groups at Zhejiang University and collaborating labs collected pumpire-6k: 100 real scenes, 40 indoor and 60 outdoor. Each scene was filmed for 200 to 300 frames and subsampled to 64, for 6,400 frames at 1280×720, plus aligned sensor depth and intrinsics. The ground-truth length is a tape measurement, not a depth map. Every scene has one pair of circular stickers, 10 mm in radius, and that single center-to-center length is shared by the whole sequence. Pixels start from a hand click on an anchor frame, propagate with CoTracker3, and are corrected frame by frame. Surfaces that break active depth are included, and labels overlap: 12 transparent, 10 reflective, 13 repetitive-texture, 5 textureless, 25 slender-structure, and 11 sharp-edge scenes. Ordinary surfaces are still the majority.
The two annotated pixels are back-projected with depth D and intrinsics K, and the score is their Euclidean distance. Estimation models use predicted intrinsics. Metric3D and Metric3D v2 use the canonical intrinsics tied to how they define depth. Completion models conditioned on sensor depth use ground-truth intrinsics. Absolute error (AE, in meters), relative error (RE), log error (LE), threshold accuracy δ, and the relative standard deviation of the predicted length (RSD) are averaged over frames in a scene, then averaged across scenes with equal weight. Difficulty is borrowed from the sensor: if the point distance implied by raw depth has RSD above 0.1, the scene is hard. That cut yields 24 hard scenes and 76 normal ones. Video models see a fixed 32 groups of 10 frames per scene. All 29 configurations predict absolute scale: 8 image estimation, 10 image completion, 6 video estimation, and 5 video completion.
On ordinary surfaces a RealSense D435 is a strong baseline. On the hard split it is close to unusable. Mean AE over all 100 scenes is 0.991 m, mostly from those 24 hard scenes.
| Method | Normal scenes | Hard scenes |
| RealSense D435 | AE 0.026 m, δ1.05 79.32% | AE 4.047 m, RE 8.390 |
| Metric3D v2 | AE 0.073 m, δ1.05 31.00% | mean rank 1.00 to 4.00 |
| Depth Anything 3 | AE 0.078 m, δ1.05 30.52% | still first, AE 0.167 m |
| Lingbot-Depth | AE 0.022 m, δ1.05 84.52% | δ1.05 47.92%, still first |
| CAPA-VGGT | AE 0.016 m, δ1.05 91.44% | AE 2.352 m, δ1.05 31.48% |
Of 18 image-level configurations, only five completion models approach or beat the sensor δ1.05 on normal scenes: Lingbot-Depth, PriorDA, Omni-dc, LDCM, and InfiniDepth. The best image estimator stops at 31.00%, under half of 79.32%. On hard scenes the order flips: 16 of the 18 have lower RE than the sensor's 8.390, and Depth Pro takes the image-estimation lead. Depth Anything 3 is first on both video-estimation splits. Its hard-scene δ1.05 is 21.16%, almost the same as the sensor's 21.03%, while AE is 0.167 m against 4.047 m. The top three video-completion ranks are CAPA variants. On the hard split CAPA-MoGe2 has a slightly lower AE, 2.351 m, and CAPA-VGGT has the higher threshold accuracy.
A depth board plus an intrinsics board does not predict point-pair error. Among five image estimators, MoGe-2 has the best combined rank, 1.92, but on Pumpire's normal split its mean rank is 4.71, behind Unidepthv2 at 2.43. Replacing predicted intrinsics with ground truth at back-projection hurts 47 of 56 model-metric pairs. The depth map is coupled to the intrinsics it was predicted with.
Hard-scene RE divided by normal-scene RE lands around 1 to 20 times for estimation and 20 to 210 times for completion. Completion posts the higher δ1.05 and the heavier outliers. Estimation sits lower, with errors that stay more bounded. Feeding image completion 100, 1,000, or 10,000 random depth points, or the full sensor map, does not improve accuracy monotonically. On the video side only MapAnything and pi3x accept a partial set of depth views. Raising that fraction from 0.1 to 1.0 lowers RE for both, and MapAnything drops more.
After aligning every setting to the same multi-view sample, normal-split RE/RSD is 0.106/0.067 for Metric3D v2, 0.109/0.029 for Depth Anything 3, 0.030/0.035 for Lingbot-Depth, and 0.023/0.028 for CAPA-VGGT. Extra views barely move RE. They move RSD, from 0.067 down to 0.029. A depth prior moves both.
Pick the model by the surface, not by a depth leaderboard. Where the sensor is reliable, raw depth is still the baseline to beat: RGB-only estimation is more than a factor of two behind on δ1.05, and a newer monocular network does not close that. On transparent, reflective, and thin structure the sensor returns meter-scale nonsense, and learned models are less likely to follow it off the cliff. Use completion when the peak matters. CAPA-VGGT reaches 91.44% δ1.05 on the normal split. Use estimation when the worst frames matter, because its error ratio stays more predictable. More depth points are not free accuracy. A full depth map also keeps the full noise pattern.
There is no new architecture here. The measurement is pulled out of depth error, and the older boards do not explain it.
The paper has no limitations section. The reading of the numbers depends on a few choices that are easy to miss.
Hard scenes are exactly those with sensor RSD above 0.1. Learned models posting lower RE on these 24 scenes means they do not copy the failure of active depth on glare and thin structure. The split itself puts the sensor at a disadvantage. It is not a general definition of hard geometry.
Each scene contributes one sticker pair. The stickers are 10 mm across and high contrast so they can be clicked at any range. The introduction talks about arbitrary point pairs. The protocol scores one fixed chord. Geometry can be wrong everywhere else and still look fine on AE and δ if that one length is close to the tape. Repeatability of the tape measurement and error in the pixel labels are not reported.
Estimation back-projects with predicted intrinsics, completion with ground-truth intrinsics, and video completion also receives sensor depth on every frame. CAPA-VGGT at 91.44% and Metric3D v2 at 31.00% are not under the same inputs. With 24 hard scenes out of 100, mean rank moves with a few outliers, and 0.1 is an empirical cutoff. Relative-depth models are outside the protocol until scale is restored some other way.