300M-param image model trained for €5000 claims SD 1.5-level performance
incorporo · reddit · 2026-09-29
Logolabs released Agate-003-preview, a <300M parameter (including text encoder & VAE) image generation model claiming SD 1.5-level performance at 3x fewer parameters, trained for under €5000 in 630 GH-200 GPU-hours (319 GPU-h for the core run).
Key techniques:
- Planner + renderer architecture: a transformer planner works at low resolution with cross-attention to text tokens, feeding a fully convolutional renderer with strong image priors
- Weighted loss using P-Dino/LPIPS and saliency, prioritizing faces and text over backgrounds
Lineage: Agate-001 → 002 (256×256) → 003 (native 512×512 with learned down/up-samplers keeping planning resolution fixed). The tiny model supports cheap, accessible LoRA training; the team plans to use it for font/glyph/symbol generation. Available on Hugging Face, with a WebGPU demo and ComfyUI nodes.
Related event: Logolabs Unveils Agate, a 260M-Parameter Text-to-Image Model Rivaling SD1.5(2 posts)→
More from Multimodal
- "A brief history of bunnies": one-shot short film via Opus 5.5 + Runway MCP — tlakomy · 2026-09-29
- An Eagle Flying Through Space and Time: Striking AI Video Demo — bilawalsidhu · 2026-09-29
- Hyper3D's new Agentic Mode builds, edits, animates 3D models from scattered references — SarahAnnabels · 2026-09-29
- Turning a Reviewed Shot List into Multishot Video Prompts with a VL Model and a Field Template — kim_deadja4951 · 2026-09-29
- Prompt Template: Duotone Cosmic Vision for Striking Two-Color Space Imagery — LudovicCreator · 2026-09-29
- One reference, a hundred styles: Recraft V4 Style demo with prompts — OVolosin82152 · 2026-09-29