5 decision models tested on 1,000 real transactions — forcing yes/no beats abstention
purealgo · reddit · 2026-10-03
A practical eval: 5 decision models (Jev, D1, Solar Decide, Kev 4b, Span-01) were run against 1,000 real transactions, 3 passes each.
- Vague restaurant names broke model confidence, triggering abstentions.
- Span-01 can't abstain and only answers yes/no, which likely explains its much lower score.
- Counterintuitively, Jev forced into yes/no mode outperformed both its own abstain-allowed baseline and Span-01.
- Bank transfers were the worst tax category across all models; Span-01 scored zero on charitable transactions.
Tests ran via the author's system-one MCP; an open-source connector hooks decision models into Claude Code, Codex and pi.
More from coding & agent
- Cloudflare Agents SDK adds first-class support for long-running Pi agents on Durable Objects — irvinebroque · 2026-10-03
- From Writing Code to Reviewing Code to Reviewing the Review Process — ykdojo · 2026-10-03
- Spring AI 2.1.0-M1 ships MessagePart model and OpenAI Responses API support — therealdanvega · 2026-10-03
- Coding Agents Love Decision Records: Why ADRs Are the Missing Project Memory — rseroter · 2026-10-03
- xAI Launches Experimental TypeScript SDK Unifying Text, Voice, Image and Video via Grok Models — mattyp · 2026-10-03
- Dev argues pass@1 success rate is the real test for AI coding tools — madhavsinghal_ · 2026-10-03