NeoHorse-1 4B beats its Qwen3.5 base on all ten tests, average score up 58.94 to 64.87
PrajwalTomar_ · x · 2026-09-16
In his TokenRhythm thread, Prajwal Tomar cites early numbers: across ten tests by TokenRhythm, NeoHorse-1 4B improved on its Qwen3.5 base in every single one, with the average moving from 58.94 to 64.87. He notes these are TokenRhythm's own results and he hasn't independently reproduced them.
He also describes testing OpenSquilla on an AI-built MVP before client handoff: it surfaced a production risk with the exact code, then self-corrected under challenge and passed a rerun.
Related event: TokenRhythm Open-Sources NeoHorse-1, a Recursive Self-Improvement Prototype(3 posts)→
More from Models
- all-MiniLM-L6-v2 still trending on Hugging Face, its creator says stop using it — tomaarsen · 2026-09-16
- HF researcher publicly questions Claude limits: monthly exhausted but weekly 84% left — NielsRogge · 2026-09-16
- Meta's promised Muse Spark open weights are over a month late and counting — RishiFurfox · 2026-09-16
- Ex-Huawei researcher: AI is great at small-step optimization, but taste and system design remain out of reach — yangyi · 2026-09-16
- First Jev Test Shows Only 80% Agreement with Verified Gemini 3.5 Flash Workflow — mayfer · 2026-09-16
- New model Jev plays Super Mario Bros in real time on fast inference — hardimanjames · 2026-09-16