Aplomb 1: open-weights 5.3B decision model with 1M context, tops 4B class on Decision Index
empiriolabsai · reddit · 2026-10-07
EmpirioLabs released Aplomb 1, an open-weights 5.3B decision model built on Qwen3.5-4B with the Qwen3-Omni audio encoder:
- Capability: reads up to 1M tokens of text/JSON/images/video/audio in one request and returns probabilities for tool selection and every enum/boolean argument (e.g. issuerefund 0.969), letting agents act on confident calls and escalate the rest; it can also output the probability the input lacks an answer.
- Benchmarks: 44.86 on Decision Index 0.2.1 (#1 among 4B models); top score up to 5.3B on 8 of 38 benchmarks incl. MMLU-Pro, GPQA Diamond, BBH; 77.5% on JevBench Hard; 75% zero-shot intent accuracy across 51 languages on MASSIVE.
- Speed/price: a custom inference runtime cuts full 1M-token decision reads to 3s from 111s (525/525 correct in long-context tests); API costs $0.02/1M input tokens with free output, ZDR by default, OpenAI/Anthropic/Gemini-compatible; 15ms model time for short queries.
- License: runs in bf16 on 12GB GPU; free for research, personal, and internal use at companies under $1M revenue. Note: training data included public train splits of two of the 38 index benchmarks (WinoGrande, ContractNLI).
More from Models
- "Astra Pause Syndrome": steering may be making models go silent, OpenAI has a workaround — thursdai_pod · 2026-10-07
- Unverified rumor suggests Qwen4 Flash is a 400B parameter model, comparable to GLM 5.3 Flash — EAccelerate_42 · 2026-10-07
- Anthropic Expands Cyber Verification Program With Three Tiers, Opens Door to Authorized Offensive Work — EricBuess · 2026-10-07
- Mistral Large 4.0 weights reportedly landing at end of October — cpldcpu · 2026-10-07
- Claude's "reasoning extraction" guardrail blocks users from seeing its thinking, and they're not happy — StewartalsopIII · 2026-10-07
- Grok's quirk: it says 'No.' then argues your point better than you did — gandamu_ml · 2026-10-07