Silia revision adds benchmarks for a 0.5M-parameter model trained on 1B tokens
SrijSriv211 · reddit · 2026-07-23
A developer published a revised paper and model release for Silia, a tiny 0.5M-parameter transformer trained on 1B FineWeb-Edu tokens for 3 epochs.
What changed in the revision
- Added benchmarks against Quark-v2, Spark-v4 and SupraMini-v6 on HellaSwag, PIQA and LAMBADA.
- Reduced compute and memory overhead using DeepSeek MLA (without decoupled RoPE), Qwen HydraHead, and Apple’s Attention Free Transformer.
- Kept the architecture deliberately simple by dropping residual-connection changes and not pursuing Kimi-style attention residuals.
What the author provided
- Paper on Zenodo and Hugging Face
- Model on Hugging Face
- Full code on GitHub
- Benchmark charts, training-loss plots, and an architecture diagram
More from Models
- Flux 3 is pitched as one multimodal AI for images, video, and reasoning — mark_k · 2026-07-24
- Microsoft says MAI models are now routing traffic in Copilot, Excel, and Outlook — satyanadella · 2026-07-24
- A user says OpenAI is in a different league on usefulness than Opus 4.8 — FlorianGallwitz · 2026-07-24
- OpenAI and Anthropic logos frame a “the labs never sleep” snapshot of frontier rivalry — MeetPatelTech · 2026-07-24
- GPT-5.6 Benchmarked on Slay the Spire, Shows Surprising Gaming Prowess — Jsevillamol · 2026-07-24
- Laguna S-2.1 GGUF fixes its chat template and thinking traces — fragment_me · 2026-07-24