SuperNav: An Agentic Navigation System for Any Task in Any Scene
Jinkai Zhang, Jingyi Xu, Yuanhong Yu, Jiarui Guo, Ruizhen Hu, Hujun Bao, Xiaowei Zhou, Sida Peng
cs.RO, cs.CV
2026-10-08
SuperNav does not fine-tune the MLLM: it points at goals in four views, and a tool walks. Single-object success is 78.00% on 150 tasks, vs 34.00% for UniNaVid.
A service robot is asked for more than a category lookup. A specific sofa, an ordered list of objects, and a request to clear a spot for work differ in what to seek, how to explore, and when to stop. The home is also one the system has not trained on. Both new requests and new scenes have to work.
Modular zero-shot stacks such as VLFM and ApexNav use a vision-language model as a scorer. Relevance between the current view and the goal updates exploration, but the handoff from search to confirmation to stop stays hardcoded. Change the request from a category to an instance, a sequence, or an abstract need, and that workflow has to be rewritten. End-to-end models take the other route. NaVid, UniNaVid, and StreamVLN fine-tune a vision-language model on navigation trajectories so it emits actions. The mapping they learn is tied to the training houses. New layouts break it.
SuperNav does not fine-tune the MLLM. Decisions come from GPT-5.6 Terra at high reasoning effort, run through Codex CLI. Tools reach the simulator or the robot over MCP. The harness adds navigation tools, Navigation Skills the model can open on demand, and a record of goal progress plus the running context.
An observation is four RGB images labeled front, right, back, and left. To move, the model names a view and a point, PointNav(view, [u, v]), with u and v normalized from the top-left corner into the unit square. The point may sit on an object or on visible floor. Fresh images come back, and the model decides again.
Two executors share that interface. The geometric one back-projects the pixel with depth and camera pose, then tracks a path on a navigation mesh. Simulation provides those signals directly. On a Unitree Go2 the same role is a LiDAR cost map at 0.05 m voxels plus wheel odometry, with no loop-closure SLAM. The learned executor is RGB-only. A mark fixes the goal on a reference image. DINOv2 encodes the current frame and up to 12 history frames. A flow-matching head predicts eight displacement steps of 0.25 m. The robot follows a short prefix and replans. The policy imitates a geodesic expert. Depth is used while building the training set, not at test time.
Skills are Markdown procedures for search, rechecking, recovery, and completion. The model chooses whether to read them. They do not force an action sequence or replace its semantic judgment. Once an image is superseded, its pixels drop out of later prompts. Text, task state, and file paths stay, so an old view can be loaded again. The close tool stores the model's own achieved or blocked call. Closing the session only means the loop ended. Success is scored separately, from distance.
The decision model is GPT-5.6 Terra. Baselines are NaVid, UniNaVid, StreamVLN, and OmniNav's Action Former, each with its original camera setup. The instance benchmark is built in Habitat-GS on InteriorGS: 150 single-object tasks and 150 ordered multi-object tasks. Arrival means geodesic distance under 1 m to an approved viewpoint, plus an explicit STOP. Demand-driven evaluation uses 200 AI2-THOR tasks from Demand-Bench, rescored as one to eight ordered stages. Each stage requires entering a 2 m box around an admissible object, without reusing an object, and a final STOP. Those scores are not comparable to DemandAgent's published numbers.
| Method | Single-object SR / SPL | Multi-object SR / SPL | Demand-driven SR / SPL |
| NaVid | 24.67% / 0.1599 | 2.67% / 0.0216 | 17.00% / 0.0920 |
| UniNaVid | 34.00% / 0.1833 | 1.33% / 0.0131 | 25.00% / 0.0943 |
| StreamVLN | 13.33% / 0.1036 | 0.00% / 0.0000 | 25.50% / 0.0775 |
| OmniNav Action Former | 27.33% / 0.2133 | 4.00% / 0.0275 | 37.50% / 0.1213 |
| SuperNav geometric | 78.00% / 0.4127 | 34.00% / 0.1388 | 59.50% / 0.1801 |
| SuperNav learned | 68.00% / 0.2059 | 34.00% / 0.0862 | 45.00% / 0.1603 |
With the geometric executor, single-object success is 78.00% and SPL is 0.4127, against 34.00% and 0.1833 for UniNaVid. The learned executor, which does not see depth or a mesh at test time, still reaches 68.00% success, but SPL falls to 0.2059. On multi-object navigation both executors land at 34.00% success, and the geometric paths are shorter. The best baseline there, OmniNav, is at 4.00%. On demand-driven tasks the geometric executor reaches 59.50%, the learned one 45.00%, and OmniNav 37.50%.
Stage-wise progress ignores STOP. On multi-object tasks the geometric executor hits the first target on 85.33% of 150 episodes and the fifth on 39.47% of the 38 episodes that have five targets. On demand-driven tasks it reaches stage one on 72.00% and stage five on 9.09% of 11 episodes. From stage six on, every method scores 0. Only three, two, and one tasks are even eligible.
Category-level tests are separate. On 120 HM3D-OVON val-unseen episodes, with no OVON-specific training, the geometric executor records 73.33% SR and 0.4105 SPL at 1 m, and 68.33% SR and 0.3837 SPL at 0.25 m. The learned executor keeps 70.83% SR at 1 m while SPL drops to 0.1443. On 1,000 HM3Dv2 validation episodes the geometric executor scores 80.30% SR and 0.3338 SPL at 0.2 m, and 86.50% SR and 0.3638 SPL at 1 m.
Ablations fix the geometric executor and those same 120 episodes. Removing Navigation Skills drops 0.25 m SR from 68.33% to 35.00%, and 1 m SR from 73.33% to 54.17%. A front-only interface scores 52.50% SR and 0.2230 SPL at 0.25 m. Routing the point through LocateAnything, so the model names an object in words and an external grounder picks the pixel, scores 55.83% at 0.25 m, below direct pointing. Swapping the decision model for GPT-6 Astra at medium effort, harness unchanged, raises 0.25 m SR to 71.67% (86 of 120) and SPL to 0.4579. Runs on a Unitree Go2 demonstrate search for a basketball and a printer followed by a trash bin. No success rate is reported.
Splitting the decision from the walking makes both pieces replaceable. The MLLM is not trained on navigation actions. A geometric planner and a learned RGB policy accept the same image point. Search and stopping policy live in Markdown, so a procedure change does not require a new training run.
Against the NaVid line and OmniNav's fast action model, single-object success moves from 34.00% to 78.00%, demand-driven success from 37.50% to 59.50%, and multi-object success from 4.00% to 34.00%. The learned backend still clears 68.00% on single-object tasks, so a simulator mesh is not the whole story. Stripping Skills cuts 0.25 m success from 68.33% to 35.00%.
Multi-object success stops at 34.00%. Where depth and localization already exist, the harness is ready to attach. Where the robot has only RGB and path length matters, the learned executor's SPL is not yet enough.
The paper states four limits. Semantic calls are only as good as the MLLM. Geometric execution needs a map or scene geometry. Inference latency varies, and the paper never reports latency or cost, so episode time is hard to budget. Demand-driven scores measure ordered travel to relevant objects, not whether the underlying chore was finished.
The baseline comparison is misaligned. Most action-prediction baselines see a front camera. OmniNav's Action Former sees three views plus pose. SuperNav sees four labeled views. The geometric executor also consumes simulator depth, pose, and a navigation mesh those baselines never receive. The learned executor is the fairer control: single-object success falls from 78.00% to 68.00%, and SPL from 0.4127 to 0.2059. That is still above UniNaVid. Most of the path-efficiency gap sits in the planner.
Published category-level numbers do not share one protocol. The OVON run is 120 episodes, and two of them are counted as successes under an unreachable-goal rule. Instance tasks require STOP. The HM3D numbers here do not. At 0.2 m, ApexNav's SPL of 0.3800 beats SuperNav's 0.3338, while success is 76.20% against 80.30%.
Ordered multi-object success, which requires every target and a STOP, is 34.00%. Demand-driven runs almost all fail after stage five. Progress entries are the model's own reports. The real robot has no success count. Skills are handwritten, so skipping fine-tuning still leaves navigation knowledge in the prompt. The decision model is closed, GPT-5.6 Terra or GPT-6 Astra, with a 3,600 s cap per episode. The instance benchmark was built for this paper. Generative models drafted object annotations that the authors then checked.