JevBench v1.5: Cygnet and Winnow-12B tie for the top spot in a first for the leaderboard
airesearch12 · x · 2026-09-29
Benchmark Heaven released JevBench v1.5, and for the first time there is no clear solo #1: Cygnet and Winnow-12B are essentially tied at the top, while the original frontrunner "Jev" slips to third.
The update brings several methodological improvements: filtering by price basis (fixed I/O blends, output-only pricing, etc.), task workload, region (hosting and company location), data confidentiality policies, and optionally folding in a "Benchmaxxing" signal into the composite score. Sub-rankings cover coding agents, full-stack Elo, Epoch ECI, and agentic/tool-use capabilities.
Related event: JevBench v1.5 Released: Cygnet and Winnow-12B Tie at the Top(2 posts)→
More from Models
- Four of the five smartest models on the AA Intelligence Index are now Claude — Hesamation · 2026-09-29
- Databricks evals: Opus 5.5 cuts task costs 20%, GPT-6 Luna is 20x cheaper — pwendell · 2026-09-29
- Sonnet 5.5 live in Claude Code; Pro/Max/Team usage resets available until Oct 22 — ClaudeDevs · 2026-09-29
- Anthropic's preserved thinking blocks account-switching distillation attacks — ClaudeDevs · 2026-09-29
- Claude Sonnet 5.5 Efforts Span 18x Price Range: $0.41 to $7.60 Per Task on AA Index — ArtificialAnlys · 2026-09-29
- Claude Sonnet 5.5 Burns ~193k Output Tokens Per Task, 7x More Than GPT-6 Astra — ArtificialAnlys · 2026-09-29