Supra2-Medium: a 25M model trained from scratch on two RTX 5060s beats its 50M predecessor
LH-Tech_AI · reddit · 2026-08-20
- SupraLabs releases Supra2-Medium-Base, a 25M-parameter model built on the Qwen3 architecture and trained entirely from scratch on consumer hardware: an RTX 5060 Ti 16GB plus an RTX 5060 8GB.
- Benchmark results show it competing heavily with the team's previous Supra-50M-Base, which is twice its size.
- This is a base model only; an instruction-tuned version may come later. Model available at huggingface.co/SupraLabs/Supra2-Medium-Base.
- The team teases Supra3 with four tiers: Flash-Lite 25M, Flash 50M, Pro 75M, and Ultra 100M.
More from Models
- Grok 4.6 leads in legal/GDP benchmarks, lags in coding — ChrisGPT · 2026-08-20
- Leaked System Prompt: Domestic Giant's Client Uses 3-Layer Memory & MCP Routing — vista8 · 2026-08-20
- Ant Group opensources Ling-3.0 Base models with training checkpoints — 智东西 · 2026-08-20
- Test finds base model completions inherit AI/human writing traits from prompts — AaronBergman18 · 2026-08-20
- Polymarket: 10% chance a Chinese model tops the leaderboard by year-end — Polymarket · 2026-08-20
- User Reports Excessive GPT Refusals Blocking Tasks — cantrell · 2026-08-20