Researchers tried many open-weight setups, none accurate and feasible over 20M cases
jon_mellon · x · 2026-09-19
jonmellon adds detail: they tried many open-weight models and settings but none looked both good and feasible to run over 20 million cases, though they may not have found the optimal configuration.
More from Models
- 16-model calibration test: open-weight models almost never admit uncertainty — AlexKim · 2026-09-19
- Jev answers in 455ms — a gate you can afford to run on everything — AlexKim · 2026-09-19
- I ran 16 models to vet one tool: one task is not a benchmark — AlexKim · 2026-09-19
- Dev tests 16 models to evaluate TypeSafe's Jev — it ranked 10th on accuracy — AlexKim · 2026-09-19
- Mystery Model Jev Launches Claiming 200x Speed and 400x Cost Cuts, Devs Impressed — multiply_matrix · 2026-09-19
- GLM 5.3 Flash leads quality, Qwen 3.8 Flash Next wins speed in open small-model comparison — HankYeomans · 2026-09-19