Vison: open-source Windows app runs FLUX, Wan, SDXL locally on 6GB GPUs, no account, no telemetry
Decent-Manager-5373 · reddit · 2026-10-02
A developer released Vison, an MIT-licensed open-source Windows desktop app for running local image and video generation models — no account, no credits, no telemetry, backend listening only on 127.0.0.1.
Features:
- Text-to-image: SDXL Turbo, Z-Image Turbo, FLUX.1 Schnell, Qwen-Image
- Text-to-video: Wan 2.1/2.2, HunyuanVideo 1.5, LTX 2.5 (larger models wired up but untested by the author)
- Image-to-image plus ESRGAN image/video upscaling, including GGUF upscalers
- Paste any Hugging Face repo or file link and it auto-matches the model family, reusing its text encoder and VAE
- Reads free VRAM and greys out models that won't fit
- Searchable generation history
Stack: stable-diffusion.cpp + Vulkan under an Electron UI, running on NVIDIA/AMD/Intel without CUDA; falls back to hf-mirror.com if Hugging Face is blocked.
Real numbers: SDXL Turbo q80 GGUF at 768×768, 4 steps, CFG 1.0 produced a first image in 45.7s on a 6GB RTX 4050 laptop (including cold start). Honest caveats: one developer, tested on a single 6GB GPU; Windows only; unsigned installer.
Related event: Open-source app Vison brings local image/video generation to 6GB GPUs(2 posts)→
More from Multimodal
- EgoTools: 100-hour egocentric video dataset teaches AI tool-centric reasoning — liuziwei7 · 2026-10-03
- Runway AI Summit closes with Valenzuela reflection, Labs unveils Continuum — runwayml · 2026-10-03
- Experimenting with AI outpainting to revive and extend old photos — rufusd · 2026-10-03
- PixVerse R2 launches as a real-time steerable world model with persistent memory — lmoroney · 2026-10-03
- Experimental Real-ESRGAN Anime6B fine-tune targets manga screentones and linework — Rapipago123 · 2026-10-03
- Full music video generated locally with ComfyUI and LTX 2.5 on 16GB VRAM — sokmech · 2026-10-02