SSMB: Self-Supervised Local Feature Detection under Motion Blur
Zhenjun Zhao, Fabio Bellavia, Wenting Wang, Fan Zhu, Jiajun Wu, Suryansh Kumar, Mingqiang Wei, Haoang Li, Javier Civera
cs.CV
2026-08-27
SSMB detects keypoints on motion-blurred images without deblurring or SIFT labels. Repeatability holds at 77.16% on Blur-HPatches Tough; 3-pixel MMA reaches 40.10%.
Motion blur smears local structure across neighboring pixels. Keypoint detectors then lose repeatability, and every downstream task that needs correspondences (SLAM, SfM, visual localization) takes the hit.
Two existing fixes both come with a catch. Deblur-then-detect is slow and injects restoration artifacts into the detector. BALF skips deblurring and runs in real time, but it is trained to regress SIFT keypoints from the paired sharp image. SIFT was designed for sharp photos. The student copies SIFT's tastes instead of finding structures that actually survive blur.
SSMB is a third option: detect directly on the blurred image, with no deblurring stage, no handcrafted detector, and no external pseudo-labels.
The backbone is the same MAXIM-style gated MLP encoder BALF used: four stages down to 1/8 resolution and 256 channels. Each stage splits features into a global GridGmlp path and a local BlockGmlp path. That split fits blur, where intensity is already spread over a neighborhood. Global mixing, though, washes out the fine local contrast that localization needs.
LDE (Local Discriminability Enhancement) sits between those two paths. After LayerNorm, a 3x3 depthwise convolution picks up local gradients, a tiny pointwise bottleneck gates them by channel, and the product is added back with residuals. About 50K extra parameters. Remove it and the probability map collapses toward zero almost everywhere.
The head follows SuperPoint: 64 pixel bins plus a dustbin per 8x8 cell, plus a sub-pixel offset tower.
Training has two stages.
All of this fits on one RTX 3090.
On Blur-HPatches with 1,000 keypoints, Tough blur-to-sharp repeatability is 77.16% for SSMB, 71.84% for BALF, and 52.84% for SuperPoint. Tough blur-to-blur: 74.48% versus BALF 67.71%. SSMB barely moves from Easy to Tough (77.24 to 77.16). Most baselines drop as the kernel gets worse.
Deblurring first does not close the gap. The best deblur-then-detect number on Tough deblur-to-sharp is DeblurGAN-v2 plus SuperPoint at 58.22%. SSMB on the raw blur is 77.16%.
Matching on Blur-HPatches Tough, blur-to-blur, up to 2,048 points:
| Method | MMA@3px | MMA@5px | MMA@10px |
| SSMB+HardNet | 40.10 | 51.43 | 77.22 |
| BALF+HardNet | 25.03 | 44.71 | 72.64 |
| RoMa v2 (detector-free) | 28.51 | 54.58 | 88.54 |
At the strict 3-pixel cutoff, the sparse detector beats every detector-free matcher, RoMa v2 included. Under illumination change the 3-pixel MMA jumps from BALF's 34.58% to 62.02%.
Relative pose on ArchViz, blur-to-sharp AUC@30°: SSMB+HyNet 66.61 versus BALF+HardNet 59.15. On Aachen Day-Night with synthetic blur on both queries and the map, daytime localization at (0.25m,2°)/(0.5m,5°)/(1.0m,10°) is 58.5/76.2/93.0 against BALF+HyNet 51.2/68.9/89.4. Night is 31.6/59.2/86.7 against BALF+HardNet 21.4/36.7/80.6. Detection at VGA is 30.63 ms, about 33 FPS, 3.6 ms slower than BALF.
Ablation on Tough blur-to-sharp: no LDE 33.83%; no synthetic pretraining 59.58%; no spatial diversity 64.07%; full model 77.16%.
For anyone running localization or SLAM on shaky cameras, this removes both the deblurring tax and the SIFT-imitation shortcut. The detector is real-time, and the keypoints can be precomputed and reused. Pairwise matchers such as LoFTR and RoMa cannot do that. RoMa v2 ran out of memory during bundle adjustment on Aachen's 4,328 database images. Sparse detection still scales.
SSMB does not learn a descriptor. Downstream matching still needs HardNet or HyNet. Treat it as a drop-in detector head if the rest of the pipeline already exists and only the points are unstable under blur.
There is no dedicated limitations section. The gaps are still visible.
Training uses GoPro, which makes blur by averaging frames. The main numbers on Blur-HPatches and Aachen also use synthetic PSFs. Real camera blur appears only as qualitative figures on RWBI and RealBlur. Whether shutter-drag and object-motion blur behave the same is untested here.
On clean HPatches, overall repeatability is 75.17%, second place, but viewpoint repeatability is 67.50%, behind REKD at 80.94% and ALIKED at 72.85%. Homographic adaptation does not model parallax or occlusion. On ArchViz blur-to-blur at 5°, BALF+HardNet edges SSMB 12.63 to 11.07. At the loose 10-pixel matching threshold RoMa v2 still leads 88.54% to 77.22%.
The self-supervised objective is brittle. Drop synthetic pretraining or the coverage term and it collapses. LDE mostly patches the damage global MLP mixing does to local gradients. Code and weights are promised after acceptance and are not public yet.