Limite 1B on M4 Pro via experimental MLX port: ~57 tok/s, 4/8 on sampled AIME 2025

GiuseppeSaccar3 · x · 2026-09-23

Mathematician GiuseppeSaccar3 ran Paradigma's Limite 1B (Violetto) on an M4 Pro with an experimental MLX port: 57 tokens/s, 3.04 GiB peak allocation.

Results: 6/6 on elementary diagnostics (2m19s); 4/8 on selected AIME 2025 problems (14m), with the other four attempts hitting the 8,192-token cap — one derived the right answer but never finished. The unofficial port passes 23 local tests but logits weren't compared with the official CUDA runtime; official serving requires Linux + NVIDIA GPU. Code open-sourced as gsaccardi/limite-1b-violetto-experiments.

Related event: Limite 1B runs locally on M4 Pro at ~57 token/s, scoring 4/8 on AIME sampling(2 posts)→

Original post →

More from Infra

Infra channel →