Stu-tU: RANSAC scoring that survives 128x scale miscalibration
ducha_aiki · x · 2026-09-09
- Stu-tU is presented as "RANSAC Scoring Done Right," based on a miscalibration test
- The score is tuned once on a validation set, then remains robust to scale miscalibration up to 128x, where threshold-based scores fall apart and Stu-tU keeps selecting the correct model
- The author teases an upcoming new RANSAC method that will beat SoTA
More from Research
- Math benchmark success shows log-linear diminishing returns with test-time compute, says ramez — sebkrier · 2026-09-09
- Must-Read Papers of the Week: MoE Scaling Laws, World Model Physics Benchmarks, VLA Training — TheTuringPost · 2026-09-09
- Fresh Memory, Stale Plans: PLANFENCE Blocks All Invalid Agent Actions Across 30 Workflows — rohanpaul_ai · 2026-09-09
- 0.54M-parameter robot policy hits 95% on LIBERO, 7700x smaller than π0.5 — ChongZzZhang · 2026-09-09
- Shanghai AI Lab proposes first systematic taxonomy of cognition-induced AI agent risks — jiqizhixin · 2026-09-09
- Test-time compute could be 10,000 model instances coordinating on a message board — mertdumenci · 2026-09-09