TEAS benchmark: measures inference at natural lengths, reporting cost, accuracy, and energy
PontiEdoardo · x · 2026-09-01
Edoardo Ponti introduces TEAS: instead of distorting fixed-length measurements, it uses each dataset's natural lengths, handling chunked prefill, refusals, and steps that mix prefill with decode—capturing real sparse activations and bandwidth utilisation. It reports cost, accuracy, performance (user experience and system output), and energy estimates; accuracy confirms run validity and reveals available trade-offs.
More from Infra
- HUMAIN and Together AI partner on 250MW data center in Saudi Arabia — HUMAIN · 2026-09-01
- LITE plans VCSEL products for AI interconnects, delayed by 1-2 years — zephyr_z9 · 2026-09-01
- Xiaohongshu & NVIDIA build GR-Inference engine, doubling throughput for Beam Search — 小红书技术REDtech · 2026-09-01
- antirez shows DeepSeek v4 Flash vision running fast locally on an M5 Max; Metal/CUDA/ROCm support nearly ready — antirez · 2026-09-01
- Nvidia Earnings: Avoiding Consolidation and Dollars per Gigawatt — Stratechery · 2026-09-01
- Sats4AI Offers Bitcoin-Powered AI Tools via Lightning Network — modelcontextprotocol · 2026-09-01