Third-party eval: Cloudflare's clef beats Jev on coding but is too pricey
airesearch12 · x · 2026-10-02
An independent evaluator published clef and clef-flash eval results, answering whether Cloudflare's models beat Jev: mostly no.
- Both are strong; in model routing they even beat original Jev.
- clef-flash is generally weaker than clef or Jev.
- clef wins on coding and safety/security but is equal or weaker elsewhere.
- Biggest issue: pricing. clef-flash ranks #12 on Capability Score, while clef is disqualified as non-Jev-class for cost; on the cost-inclusive Composite Score they rank #25 and #62.
A tunable public leaderboard is available for experimentation.
More from Models
- Gemini 4 Argon posts lowest hallucination rate (15%) on AA-Omniscience benchmark — import_jmr · 2026-10-02
- User says Opus 5.5 shows no token savings, hits 5-hour limit in 3-4 messages — MarsupialFirst8617 · 2026-10-02
- llama.cpp adds Decision Models, expanding local inference capabilities — paf1138 · 2026-10-02
- Early Argon impressions: dev vibes-codes with it, says it shows no signs of benchmaxxing — cgarciae88 · 2026-10-02
- NVIDIA's Kumo-Tabular Tabular Foundation Model Trends on Hugging Face — nvidia · 2026-10-02
- Claude Sonnet 5.5, Grok 4.7 and GPT-6.1 Sol go live on Runware's OpenAI-compatible endpoint — aziz4ai · 2026-10-02