Agate-001-preview: 260M-param open-weights image model nears SD 1.5 with thinker-renderer architecture
incorporo · reddit · 2026-09-26
Logolabs released Agate-001-preview, a 260M-parameter (including text encoder) image generation model under MIT license, claiming Pareto-frontier results that approach SD 1.5—a model with roughly 3x the total parameters.
Architecture highlights:
- A hybrid "thinker-renderer" design: a small recurrent transformer "thinker" reads the prompt and plans a 16×16 region map via cross-attention to a text encoder (based on Ettin-68M, frozen for 10 epochs then jointly trained), while an FDCM convolutional renderer converts latent thinking tokens into images.
- Trained 26 epochs on Flux-Reason-6M and not yet converged; currently limited to 256×256 with the SD-1.5 VAE.
Use cases: cheap synthetic data generation (adaptable to generative denoising/editing), training small models on proprietary datasets, thumbnails/search loaders, and on-mobile generation.
Weights and a WebGPU demo are live on Hugging Face.
More from Multimodal
- Entire motion design video is pure code, zero After Effects — dev open-sources the Claude prompt template — EricBuess · 2026-09-26
- viggle-turbo matches original quality on most prompts, now with ComfyUI support — init-5 · 2026-09-26
- Blender agents meet fal's H3 Max: 3D previs turns into photoreal video — gorkem · 2026-09-26
- A short film 'From Rocks to AI' made entirely with Claude Opus 5.5 — coherence · 2026-09-26
- CMU and Meta Unveil TrackEverything, a 3D Point Tracker for 1000+ Frame Long Videos — AdamWHarley · 2026-09-26
- Casual Snapshot Realism LoRA released for Qwen 2.1 with full ComfyUI workflow — kosovack · 2026-09-26