Critique: Decision Models Leave 10-15% Performance on the Table and Can't Return Evidence
joecole · x · 2026-10-02
Developer alexdong offers two critiques of the much-hyped decision models, jev in particular:
- Weak grounding: a general model without domain/enterprise-specific grounding is hard to steer via prompting alone, leaving roughly 10-15% of performance on the table.
- No evidence in responses: enterprise clients need quotes and citations to justify decisions for compliance and risk management — a capability current decision models lack.
He does credit the Pointer Head architecture as an efficient way to isolate multiple questions and extract strong performance from SFT, noting that adding a bit of KL regularization largely avoids catastrophic forgetting. But without evidence outputs, adoption in enterprise, high-judgment domains remains limited.
More from Models
- Frontier Learning: LLM reasoners only learn from problems at the edge of capability under GRPO — _rockt · 2026-10-02
- Opus runs non-stop for 12 hours on one 'simple' prompt, user jokes it's doom-scrolling — jaivinwylde · 2026-10-02
- 14 AI models race to rebuild photos in Blender: GPT-6 Astra sweeps with 66/100 — smith2008 · 2026-10-02
- Blogger weighs returning to Claude's $200 plan as Codex infra stumbles — StewartalsopIII · 2026-10-02
- 'Burn the Weights Into a Chip' Praise for Opus 5.5 Gets the '640K Ought to Be Enough' Treatment — kalomaze · 2026-10-02
- New ChatGPT sign-in looks great — until users hit the 'too many devices' limit — jasonkneen · 2026-10-02