Harvard Analyzes 6.12B AI Requests on Chutes: Agents Use More Tokens, Answer Shorter
markjeffrey · x · 2026-09-22
Harvard researchers analyzed 6.12 billion AI requests on Chutes (Bittensor subnet SN64), a decentralized inference platform. Key findings:
- AI agents are processing more tokens while producing shorter answers — a hallmark of agentic workloads with long reasoning but concise outputs.
- Reusing previously processed information can make AI faster and cheaper, making caching and information reuse key cost levers.
- Smarter GPU routing can reduce wasted computing power, so scheduling strategy directly affects inference economics.
Beyond the research, the Chutes team is building Parallax, a project to train AI models across distributed GPUs, strengthening its position as an open-source AI lab.
More from Infra
- python-build-standalone enables full LTO for CPython 3.12+, modestly boosting runtime — charliermarsh · 2026-09-22
- Measured trade-offs of three REAP-pruned Qwen3.8-Flash-Next MLX builds on Apple Silicon — MensaProdigy · 2026-09-22
- Dev claims further-optimized DeepSeek V4 NVFP4 uses 190GB of 192GB VRAM — HankYeomans · 2026-09-22
- 'AWS made the industry soft': AI infra isn't mature enough to outsource the hard parts — mgill25 · 2026-09-22
- Grass network audited: 3M+ users, $32.1M revenue serving AI training data — Ronangmi · 2026-09-22
- SemiAnalysis: mapping MoE models onto inference hardware — zephyr_z9 · 2026-09-22