NeoHorse-1 4B beats its Qwen3.5 base on all ten tests, average score up 58.94 to 64.87

PrajwalTomar_ · x · 2026-09-16

In his TokenRhythm thread, Prajwal Tomar cites early numbers: across ten tests by TokenRhythm, NeoHorse-1 4B improved on its Qwen3.5 base in every single one, with the average moving from 58.94 to 64.87. He notes these are TokenRhythm's own results and he hasn't independently reproduced them.

He also describes testing OpenSquilla on an AI-built MVP before client handoff: it surfaced a production risk with the exact code, then self-corrected under challenge and passed a rerun.

Related event: TokenRhythm Open-Sources NeoHorse-1, a Recursive Self-Improvement Prototype(3 posts)→

Original post →

More from Models

Models channel →