Solo dev trains 1B world model running realtime on RTX 5090 with prompt-switching

lucidml_lover · reddit · 2026-10-07

A final-year student in Bangalore (solo, student-incubator funded) unveiled part 2 of his local realtime world model: a 960M-parameter pure transformer trained with block-causal masking and diffusion forcing (per-frame independent noising). Inference runs 2-5 diffusion steps per frame, writes denoised frames to a KV cache, and uses a sliding window of 80 frames of context. Unlike the earlier MMDiT version, it restored text cross-attention with heavy text-video pretraining, so it accepts live keyboard actions (via an adaLN term guiding WASD) and supports mid-rollout prompt switching — "add a pond to the desert," "change environment to icy." It peaks at 50-60fps on an RTX 5090 (throttled to 12fps, 30% GPU utilization), should work on RTX 30/40 cards, and hit 30fps on an M5 MacBook. Trained on 8x H100 for 3-4 weeks. He promises everything stays local-only, no datacenter, and hopes to release something tryable by year-end.

Original post →

More from Multimodal

Multimodal channel →