Benchmarking 4 open decision models on one RTX 4090: Laya fastest, Lev most accurate at 13x the latency
Fun-Meaning-6474 · reddit · 2026-10-09
The author benchmarked four recently released open decision models (Laya, Liquid's d1 3B, Cloudflare's Clef-Flash 9B, Interfaze's Lev 4B) on a single RTX 4090, having each flag centipede names word-by-word across 9,534 Wikipedia words.
Key results
- Speed: Laya (BF16 GGUF on llama.cpp) leads at 3.9ms p50 per word (7,980 words in 32s); d1 3B follows at 6.0ms.
- Accuracy: Lev 4B catches the most centipede names (83%) with only 4 wrong picks, but is 13x slower than Laya (51ms/word), partly due to running its own PyTorch server and CPU sensitivity. Clef-Flash catches only 36% but almost never misfires.
- Caveat: headline accuracy numbers (95%+) are inflated — answering 'no' also counts as correct; look at catch rate and wrong picks instead.
Verdict: Laya is the fastest overall and easily fine-tuned, making it the go-to pick. Full reproducible setup (llama.cpp b11495, CUDA 12.8, -ngl 99, quant files) included.
More from Infra
- OpenAI nears $50B annualized revenue, ~$20B short — but the cut matters more than the headline — inductionheads · 2026-10-09
- DGX Spark prices skyrocket as resale markups soar — natesiggard · 2026-10-09
- OpenAI bots hit 160K fetches for nonexistent URLs in a week, sparking RL-run speculation — gaganghotra_ · 2026-10-09
- Modal's LLM Engine Advisor picks engine, model and config for your inference workload — charles_irl · 2026-10-09
- LFM 2.5 5.4B seen as better laptop pick; 8B A1B lags 2.6B dense — Aggravating-Push-207 · 2026-10-09
- Architect Fi launches Liquid Inference, a router where providers bid per-prompt across 700+ models — markjeffrey · 2026-10-09