Limite 1B on M4 Pro via experimental MLX port: ~57 tok/s, 4/8 on sampled AIME 2025
GiuseppeSaccar3 · x · 2026-09-23
Mathematician GiuseppeSaccar3 ran Paradigma's Limite 1B (Violetto) on an M4 Pro with an experimental MLX port: 57 tokens/s, 3.04 GiB peak allocation.
Results: 6/6 on elementary diagnostics (2m19s); 4/8 on selected AIME 2025 problems (14m), with the other four attempts hitting the 8,192-token cap — one derived the right answer but never finished. The unofficial port passes 23 local tests but logits weren't compared with the official CUDA runtime; official serving requires Linux + NVIDIA GPU. Code open-sourced as gsaccardi/limite-1b-violetto-experiments.
More from Infra
- Cost of intelligence falling 50% per quarter, 4x faster than DNA sequencing, says Epoch AI — Terminator857 · 2026-09-23
- DiffusionGemma-Jev lands in vLLM: single-step structured answers with confidence scores — bodonoghue85 · 2026-09-23
- Toby Ord questions chart claiming AI prices fell 100,000x since 2021 — tobyordoxford · 2026-09-23
- Four Mac Studios Over Thunderbolt Run Trillion-Parameter Model on One Wall Outlet as Apple Pitches Local AI — mark_k · 2026-09-23
- Archgen Labs, aiming to make chip design 1,000x faster, lands YC backing — retr0jirachi · 2026-09-23
- Running Omarchy and a 90M-parameter LLM on a PSP — Kyrannio · 2026-09-23