Jev-style parallel structured inference makes 350M model 63× faster, code released
helloiamleonie · x · 2026-09-20
Inspired by Jev, a developer applied parallel structured decision inference to LFM2.5-350M: 63× faster on an L40S and 8× on MPS, with no training required. Code and weights are available on Hugging Face.
More from Infra
- NVIDIA engineer breaks down why DeepSeek re-engineered V4.1 Flash for speed — thursdai_pod · 2026-09-20
- $140 Radeon MI50 paired with GTX-1080Ti boosts local 27B-35B LLM speeds up to 9x — tabletuser_blogspot · 2026-09-20
- halogen 0.12.0 hits 38 tok/s decode at 1M context for Qwen on Strix Halo — peonist-ai · 2026-09-20
- GPU Programming Diary: Revisiting the Classic CUDA Matmul Optimization Worklog and MIT's Sparsity Lecture — NandoDF · 2026-09-20
- Dev builds SLO-aware inference router with Jev to pick the optimal LLM per request — ai · 2026-09-20
- Memristive Networks Learn by Reorganizing Themselves: When Material Is the Model — bravo_abad · 2026-09-20