Grok 4.5 Tops Document Generation Benchmark
ell-hol1 · reddit · 2026-07-09
The author tested nearly all available models on PowerPoint and document generation tasks, using mini-SWE-agent as the execution framework, combined with Anthropic's pptx/docx skills to generate results, and compared them via blind voting. Results showed that the newly launched Grok 4.5 surpassed the previous leader MiniMax M3, ranking first on this benchmark, even outperforming Fable 5, GPT-5.5, Sonnet 5, and GLM 5.2. The author also noted that Grok 4.5's average cost per document is about $0.23, with reasoning effort set to low during testing.
More from coding & agent
- Codex turns out 123 screensavers in one playful batch — intellectronica · 2026-07-21
- Grok Build adds `grok doctor`, resumable sessions and remote image paste — mark_k · 2026-07-21
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- CHAP defines approvals, handoffs, and audit logs for human-agent workflows — DeliveryTechnical199 · 2026-07-21
- The author says Codex reached 20x and is now debugging spec decoding on a hybrid parallel setup — TheZachMueller · 2026-07-21
- Axcess adds an MCP connector for WCAG accessibility checks that scanners miss — modelcontextprotocol · 2026-07-21