Solo dev trains 1B world model running realtime on RTX 5090 with prompt-switching
lucidml_lover · reddit · 2026-10-07
A final-year student in Bangalore (solo, student-incubator funded) unveiled part 2 of his local realtime world model: a 960M-parameter pure transformer trained with block-causal masking and diffusion forcing (per-frame independent noising). Inference runs 2-5 diffusion steps per frame, writes denoised frames to a KV cache, and uses a sliding window of 80 frames of context. Unlike the earlier MMDiT version, it restored text cross-attention with heavy text-video pretraining, so it accepts live keyboard actions (via an adaLN term guiding WASD) and supports mid-rollout prompt switching — "add a pond to the desert," "change environment to icy." It peaks at 50-60fps on an RTX 5090 (throttled to 12fps, 30% GPU utilization), should work on RTX 30/40 cards, and hit 30fps on an M5 MacBook. Trained on 8x H100 for 3-4 weeks. He promises everything stays local-only, no datacenter, and hopes to release something tryable by year-end.
More from Multimodal
- Free open-source LoRA Trainer Studio now covers 10 model families, adds ERNIE-Image and RunPod — AcademiaSD · 2026-10-07
- Gallery of 226 Claude Opus motion graphics with the prompts behind them — dotey · 2026-10-07
- Nano Banana 2.1 lands on Magnific with 4K output and better prompt adherence — aziz4ai · 2026-10-07
- Open-Source Project With 2,000+ Stars Adds 9 AI Animation Explainer Styles and Full Workflow — dotey · 2026-10-07
- Paper shows diffusion transformer tokens encode lots of image info before it's interpretable — kwangmoo_yi · 2026-10-07
- Claude Fable builds a Chladni-plate music visualizer that dances to any song — creatoroff · 2026-10-07