Discussion: Tokens Are the Wrong Unit to Measure LLM Costs
willccbb · x · 2026-07-19
The author points out that while it's inaccurate to simply claim "smaller models are better," large models like GPT-4.5 and Llama-405B are actually less efficient than smaller versions within their respective series.
They emphasize that cost is a key factor when evaluating models, and the industry's current standard of pricing by token is flawed. This approach fails to offer a fair and intuitive apples-to-apples comparison of actual efficiency across models with varying parameter sizes.
Related event: Evaluating Token as a Flawed Metric for LLM Costs(2 posts)→
More from Models
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22
- Moonshot’s Kimi K3 sets a new open-weights ECI record at 156 — scaling01 · 2026-07-22
- Nanbeige4.2-3B launches as a 3B Looped Transformer model that beats larger baselines — Wooden-Deer-1276 · 2026-07-22
- A post says six companies now beat Google’s best LLM, including two open-source models — soham_btw · 2026-07-22
- Google says Gemini 3.5 Pro is in testing and Gemini 4 is already pre-training — Wide-Ad1564 · 2026-07-22
- Gemini 3.5 Flash Lite Tested: Not Frontier-Optimal, but Hits 350 tok/s — brandon_galang · 2026-07-22