Jev-inspired inference makes 350M-parameter LFM2.5 63x faster on L40S, code and weights released

JosephJacks_ · x · 2026-09-20

A developer applied Jev-style inference to LFM2.5-350M: no training, just parallel decisions. The result is a 63x speedup on an NVIDIA L40S and 8x on Apple MPS. Code and weights are published on Hugging Face for anyone to reproduce.

Related event: Parallel Structured Reasoning Speeds Up 350M Small Model by 63x, Code Open-Sourced(2 posts)→

Original post →

More from Models

Models channel →