Comparing tokens across different models is meaningless, metrics need refinement in the reasoning era
adamdangelo · x · 2026-08-27
Adam D'Angelo argues that aggregating tokens from multiple models is meaningless because tokens cannot be compared across models. This has always been true but becomes increasingly critical as models diversify and reasoning tokens constitute a larger portion of usage.
More from Models
- Gemini 3.5 Transcribe launches; Enterprise Agent Platform enters public preview — Saboo_Shubham_ · 2026-08-27
- Benchmarking Qwen3.8 27B Quantizations: 4-bit Holds Up, 1-bit Collapses — pmigdal · 2026-08-27
- GLM-5.3-Flash Review: 10% Cost, Pareto Frontier Performance — ArtificialAnlys · 2026-08-27
- Google announces pricing details for Gemini 3.7 Flash — OfficialLoganK · 2026-08-27
- Unsloth releases GGUF quantization of GLM-5.3-Flash model — unsloth · 2026-08-27
- Will N-Gram tables revolutionize the local AI race? — AcreMakeover · 2026-08-27