Jev vs Luna Benchmarked: 139/140 vs 138/140 Labels, 4.6x Faster and 83% Cheaper
TheMoonMidas · x · 2026-09-17
@ajfedor tested Typesafe's Jev against OpenAI's Luna (reasoning off) on 20 synthetic supplier replies with 140 labels under the same rubric: Jev scored 139/140 vs Luna's 138/140, while being 4.6x faster and 83% cheaper at published rates. Jev did once give maximum confidence to an answer missing currency info. Separately, a developer built an open-source code review tool on Jev that screens diffs and surfaces findings in a local dashboard.
More from coding & agent
- ScienceIDE Turns the World's Scientific Codebases into Agent-Learnable Environments — Hejia Geng · 2026-09-17
- EvolveTrade: Self-Evolving LLM Trading Agents Refine Their Own Tool-Use Policy — kaist-ai · 2026-09-17
- Misplaced build cache in AI coding tool cost developer a week of waiting — RileyRalmuto · 2026-09-17
- Agent retried a payment after timeout? Give write ops unique IDs and a way to check — gethackteam · 2026-09-17
- synathic: an MIT-licensed SDK that verifies agent writes in Postgres instead of trusting the 200 OK — Gallegos_Daniel · 2026-09-17
- Jitsu launches Observability Exports, streams pipeline logs to Datadog, Grafana and more via OTLP — nikola_mr64990 · 2026-09-17