DirtyMoCap: Robust Motion Capture from Unconstrained Markers
Long Wang, Shuting Zhao, Shen Yan, Siyuan Yu, Xiaoben Li, Zeyu Cai, Yumeng Hou, Yuliang Xiu
cs.CV
2026-09-17
DirtyMoCap maps unordered, noisy markers with unknown layouts to 113 anchors and fits SMPL-H. Solver MPJPE is 8.8 mm under occlusion and outliers, down from 16.2 mm at init.
Commercial optical systems such as Vicon, Qualisys, and OptiTrack are accurate when the marker layout is known, IDs stay fixed, and trajectories are clean. Studio captures often break all of that at once: the layout was never logged or was swapped, occlusion drops points, points that reappear get new IDs, and noise or bad tracks insert outliers. The paper calls these inputs unconstrained markers.
SOMA labels unordered points and rejects ghosts. LocalMoCap, RoMo, and OpenMoCap complete markers with marker-joint graphs or chains, but the networks are tied to one layout. DAMO accepts arbitrary layouts and regresses body parameters directly, which caps accuracy. A known layout still needs heavy manual cleanup.
DirtyMoCap does not name individual markers. It maps the cloud to proxy anchors, 113 points with fixed identity: 52 joints of the parametric body SMPL-H, plus 61 surface points locked to mesh vertices. Anchors sit in the same coordinates as the markers, so a distance is meaningful, and each anchor has a preset link to the body. Fitting does not need labels or a declared layout.
Joints cannot pin twist. A rotation around the bone barely moves the joint. Surface points on the limbs and torso lock that degree of freedom and give the shape vector a signal for limb thickness.
Training is staged: initialization, tracking on 24-frame clips, then 56-frame clips, then the solver end to end. Motions are CMU bodies plus GRAB hands, 2,068 sequences. Each batch uses one layout of 38 to 175 markers, with shuffle, dropout, and outliers. The 75 real layouts come from SOMA.
Synthetic markers are sampled on SMPL-H vertices, bodies from CMU and hands from GRAB. Of 2,068 sequences, mean length 1,687 frames, the last 20% is the test set. SOMA, LocalMoCap, and OpenMoCap are scored only on their own layouts, and the latter two need ordered markers. DirtyMoCap uses one model on every layout. The experiments section reports lower joint error than SOMA, matching or better joint error on the other two layouts, and lower vertex error than SOMA.
On layout 2 with occlusion and outliers, MPJPE over 52 joints is 16.2 mm after initialization, 7.6 mm after tracking, and 8.8 mm after the solver. The solver gives some of that joint fit back to smoothness and the pose prior. The stages take 3.08, 12.22, and 23.03 ms per frame, 38.33 ms total.
On real SSM data, 67 markers are paired with synced scans. MoSh++ receives marker-to-body correspondences and a 9.5 mm mean surface offset. DirtyMoCap gets the same coordinates and neither prior.
| Method | Metric | Result |
| MoSh++ (with correspondences) | Symmetric Chamfer | 11.3 mm |
| DirtyMoCap | Symmetric Chamfer | 15.3 mm |
| MoSh++ | Valid marker-to-mesh | 10.0 mm |
| DirtyMoCap | Valid marker-to-mesh | 7.5 mm |
| Uniform-weight LM | MPJPE, sparse outliers | 31.0 mm |
| Learned confidence | MPJPE, sparse outliers | 1.5 mm |
| LM, 10 steps | MPJPE, occlusion + outliers | 38.63 mm |
| Learnable solver, 10 steps | MPJPE, occlusion + outliers | 8.76 mm |
| LM, 80 steps | MPJPE, occlusion + outliers | 7.64 mm |
Chamfer, the mean two-way nearest distance between surface and scan, is 4.0 mm higher than MoSh++. Valid marker-to-mesh distance falls from 10.0 mm to 7.5 mm. Valid markers lie within 1 cm of the scan and are 97% of the set. SSM has no finger markers, so the hands fall back to a prior mean pose.
When sparse outliers shift 10% of anchors by 0.86 to 1.68 m, uniform weights sit at 31.0 mm and learned confidence at 1.5 mm. If 25% of frames also corrupt a random body part, all three weights reach 1.9 mm, against 11.5 mm for confidence alone. On the same predicted anchors, 10 solver steps reach 8.76 mm MPJPE and 10.12 mm vertex error (MPVPE), better than 30-step LM at 8.91 mm and close to 40-step LM at 8.34 mm. Ten-step LM lands at 38.63 mm.
Held-out layout config.d scores 12.93 mm MPJPE, hands 14.70 mm, with no adaptation, and 7.57 mm after fine-tuning on the exact layout. A voted proxy layout reaches 10.28 mm overall, while hands stay at 12.14 mm. The same model runs on martial-arts captures with no layout metadata and produces HKMALA-Motion: 134 sequences, about 8,500 frames each, about 180 minutes, 22 styles. The parameters are model output, not hand-checked ground truth. Weapon forms keep the body only.
The costly part of optical capture is cleanup, not the cameras. One set of weights covers layouts seen in training, 38 to 175 markers, so a new configuration in that range does not need its own network. At 38.33 ms per frame the pipeline is an offline batch tool.
Learned confidence is the large term on outliers: error falls from 31.0 mm to 1.5 mm. Against MoSh++, scan error is 4 mm worse, and what disappears is the need for marker-to-body correspondence.
When the layout and the correspondences are already known, MoSh++ still fits the scan more tightly. DirtyMoCap replaces a family of layout-specific models with one model that can start without that metadata. It does not raise the accuracy ceiling of optical capture.
The paper states three limits. Unseen layouts lose accuracy. Voting a proxy layout is unreliable on dense hand vertices and sometimes needs manual picks. Leaving SMPL-H means new anchors, residuals, Jacobians, and a retrain; initialization and tracking can stay.
The synthetic comparison is uneven. Baselines are tested only on their own layouts, and LocalMoCap and OpenMoCap receive ordered markers. One model across layouts is a real result. The abstract says the method consistently outperforms configuration-specific baselines; the experiments section is narrower, lower error than SOMA and matching or better on the other two. Solver MPJPE of 8.8 mm is higher than tracking at 7.6 mm. HKMALA-Motion was not manually checked, and there is no measurement of whether the pose prior flattened the martial-arts motion. Up to 100x is the CUDA solver kernel against PyTorch. End-to-end time is the 38.33 ms per frame above.