Beyond Benchmarks: How Do You Test LLMs in Production?
WhatererBlah555 · reddit · 2026-08-25
The post discusses the limitations of mainstream benchmarks, noting models might be optimized for them, while gut feeling is unreliable and simple tasks aren't representative of actual work. Asking the community how they benchmark LLMs in their own working environments to get practical insights.
More from coding & agent
- Replit CEO says Replit Agent fully replaced Claude CoWork in his workflow — amasad · 2026-08-25
- rakyll: There's a Huge Spectrum Between Full Auto-Generative Builders and Coding-Agent Users — rakyll · 2026-08-25
- Tip: Visualize Any Response in Codex with '/visualize' Command — daniel_mac8 · 2026-08-25
- AI Agent Architecture: Too Good as Background Agent, Not Good Enough as Assistant — herbiebradley · 2026-08-25
- Using Agents to Assess Exploitability of Dependency Vulnerabilities — andreamichi · 2026-08-25
- HuggingNews: AI Agent Aggregates Daily AI News — ivan_bezdomny · 2026-08-25