Untrained LFM2.5 logprob readout hits 0.710, trailing fine-tuned lev-350M's 0.725
helloiamleonie · x · 2026-09-24
After a developer ported jaredpalmer's kev to Liquid AI's LFM2.5-350M (dubbed lev), another experiment tested whether raw base-model probabilities alone could work: using llama.cpp 1-token logprob readout with no fine-tuning or decision head, LFM2.5-350M scored 0.638 and LFM2.5-2.6B scored 0.710, versus 0.725 for lev-350M. The gap between untrained models and lev is notable, showing how far raw probability readout already goes.
More from Models
- Opus 5.5 at 'low' reasoning effort matches 'max' on task completion at 12x lower cost — daniel_mac8 · 2026-09-24
- Legal case-hallucination benchmark shows Gemma4:26b slightly beating Claude — MRGWONK · 2026-09-24
- GPC-1 classifier launches with bounding boxes, poses, and precise numeric outputs in milliseconds — anselm · 2026-09-24
- Reddit user shares a basic explainer on which AI model type to use when — fuckme · 2026-09-24
- 29-model coding benchmark: DeepSeek hits 93% of top score for 2% of the cost — PieceKey2640 · 2026-09-24
- Sarvam Vision 2.1 launches with SOTA OCR scores and Indic handwriting recognition across 22 languages — itsOmSarraf_ · 2026-09-24