GLM-OCR reads Nvidia's 61-page 10-Q at 2,086 tok/s for under 2 cents
spillai · x · 2026-09-03
vlmrun ran Nvidia's Q2 10-Q through GLM-OCR via their gateway: 61 pages processed at 2,086 tok/s in 29.54 seconds, costing just $0.019.
The headline number: 2K+ tokens per second and under 2 cents to read a full quarterly filing — a striking datapoint on the speed and price of the latest OCR models for long documents.
More from Models
- AI vividly 'sees' and describes scenes while experiencing only darkness — yeastsplainer · 2026-09-03
- Every's writing bench adds Gemini 3.8 Flash, Grok 4.6, and Muse Spark 1.3 — danshipper · 2026-09-03
- Meta's Muse Spark 1.3 lands on OpenRouter with 1M context for agentic workflows — armand_ruiz · 2026-09-03
- DeepSeek-V4-Pro ships with 1.6T-param MoE; open-source eval harness steals the show — DeepLearningAI · 2026-09-03
- Rival AI agents: cross-vendor model review catches what self-review misses — rseroter · 2026-09-03
- Gemini's Distinctive Take on AI Sentience Turns Heads — aiamblichus · 2026-09-03