ME-World jointly denoises two ego streams, reaching real-world Senv of 0.468

Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction

Dahyun Chung, Siyoon Jin, Hyunwook Choi, Honggyu An, Junyoung Seo, Hyunsung Kim, Seung Wook Kim, Seungryong Kim

cs.CV, cs.AI

2026-10-09

ME-World denoises two ego videos in one sequence with shared poses and scene memory, reaching real-data Senv 0.468 and Supdate 0.466.

What problem this solves

Egocentric world models predict the next first-person frames from an agent's actions. Most of them model a single self: one pair of hands, one head, one future. Multi-person VR and embodied simulation put two people in one changing scene. When one person reaches for a pan, the hand has to land in their own view, and the other view has to show the same body and the pan's new place at the same time.

Recent multi-agent world models compress actions into locomotion, camera paths, or discrete commands. Those controls can line up viewpoints. They do not carry hands, arms, body pose, and head turns. The paper names three failures, each drawn in Fig. 2. The action must match across views. Layout, appearance, and objects must not diverge. A state change caused by the interaction has to show up in every stream.

Method

ME-World fine-tunes the Cosmos-Predict2.5 2B multiview video DiT. The video VAE and text encoder stay frozen. Real and synthetic data are fine-tuned separately. The formulation allows N agents; the experiments use two. Inputs are both agents' future head and body motion, text, the first frame, and a shared observation history. The output is two future ego videos.

Training is rectified flow with AdamW, learning rate 3e-5, and a global batch of 4 on four A100s. Inference uses 35 denoising steps at guidance scale 7.0. Actions are given. The model does not choose what either person does next.

Results

Real video is CoMind, synchronized two-person egocentric recordings. The synthetic set retargets 1,822 two-person motions from Inter-X and InterHuman onto 100 avatars, 7 scenes, and 53 placements, giving 4,736 paired 77-frame clips at 504×504 and 10 FPS.

Consistency is DINOv3 cosine similarity. Co-visible points come from metric depth and ground-truth cameras, and are kept only when reprojected depth agrees within 10 cm. Senv scores regions the history warp does not cover. Supdate scores regions that changed relative to a valid warp. Sid, synthetic only, masks the other person with SAM3 and takes the max similarity to reference renders from 12 azimuths.

Released checkpoints in the first comparison are not fine-tuned on this data. Multi-view baselines also see the other stream's ground-truth video. Lower rotation error and FVD are better.

MethodDataSenvSupdateRot. err.Full-body PCK@10FVD
ME-Worldreal0.4680.4662.2560.881637.9
GEN3Creal0.4170.41212.3760.4561358
LingBot-Worldreal0.3290.28413.1730.243689.9
ME-Worldsynthetic0.4410.4369.0570.644465.5
LingBot-Worldsynthetic0.2490.21029.9620.090841.5

On real data, ME-World also reports translation error 0.042, Hand-F1 0.888, Hand-mIoU 0.531, full-body F1 0.822, PSNR 19.849, and LPIPS 0.289. GEN3C still receives a ground-truth reference stream, and its hand mIoU is 0.087. LingBot-World's FVD of 689.9 sits close to 637.9, with a weaker room and partner pose. Synthetic Sid is 0.561, next to 0.518 for JointControlVideo. Synthetic Senv is 0.441 against GEN3C at 0.414, a narrower gap than on real video. Synthetic rotation error is 9.057, well above the real-data 2.256.

With the same backbone, training data, and shared action conditioning, adapted MetaWorld reaches Senv 0.421, Supdate 0.413, rotation error 4.086, Hand-mIoU 0.379, full-body PCK 0.787, and FVD 704.4. ME-World's matching numbers are 0.468, 0.466, 2.256, 0.531, 0.881, and 637.9. Full-body F1 is 0.816 versus 0.822. Solaris and γ-World fall to 0.146 and 0.188.

On the real-data ablation, fine-tuning Cosmos alone leaves Senv at 0.290. Adding the ego hand skeleton and history warp raises Senv to 0.401 and drops FVD to 558.6. Shared action conditioning moves partner full-body PCK from 0.416 to 0.809, while FVD rises from 554.7 to 651.9. The full model's FVD is 637.9, worse than 541.7 for joint generation plus environment memory without shared action conditioning.

Three-agent clips and a 221-frame rollout are qualitative only. A third agent is another token stream. Long videos are chunked, and the next chunk is conditioned on one or two frames from the end of the previous one. No consistency numbers are reported there.

Why it matters

Given both people's future body and head poses, one sample produces ego videos that agree on the scene. The backbone is 2B, and the fine-tune runs on four A100s at batch 4.

Against untuned GEN3C, full-body PCK moves from 0.456 to 0.881. GEN3C was not trained on CoMind, so domain shift sits inside that gap. Against MetaWorld adapted to the same data and the same action interface, Senv is 0.468 versus 0.421 and rotation error is 2.256 versus 4.086. That is a clear step on the same track. Actions still come from outside the model, so this is not a closed-loop planner.

Limitations

The scored benchmark is two people and 77-frame clips. The three-person and 221-frame demos have no consistency numbers; they show that another stream of tokens can be appended. Joint attention grows with head count and duration. The paper defers larger groups to more efficient cross-agent communication and memory.

Frame count and conditioning in the public-checkpoint comparison do not match. GEN3C pads the camera path to 121 frames and cuts it back to 77. EgoSim emits 61 frames. Every adapted baseline in the controlled comparison inherits ME-World's shared action conditioning, so that table compares communication and scene memory, not the original systems. MetaWorld has no public code and was reimplemented from the paper.

Senv and Supdate ask whether co-visible regions look like the same place, and the correspondence uses ground-truth depth and cameras. Neither score knows whether a hand actually grasped the object. Sid does not cover real people. There is no curve for what happens when history coverage shrinks. The full model also loses on FVD to a thinner ablation: tighter cross-view agreement is not the same thing as a closer video distribution. Thirty-five denoising steps plus reprojection remain far from the real-time multi-agent simulator mentioned in the conclusion.

Terms

Source

Related papers

All paper explainers