Benchmarking 4 open decision models on one RTX 4090: Laya fastest, Lev most accurate at 13x the latency

Fun-Meaning-6474 · reddit · 2026-10-09

The author benchmarked four recently released open decision models (Laya, Liquid's d1 3B, Cloudflare's Clef-Flash 9B, Interfaze's Lev 4B) on a single RTX 4090, having each flag centipede names word-by-word across 9,534 Wikipedia words.

Key results

Verdict: Laya is the fastest overall and easily fine-tuned, making it the go-to pick. Full reproducible setup (llama.cpp b11495, CUDA 12.8, -ngl 99, quant files) included.

Original post →

More from Infra

Infra channel →