How do you benchmark specific prompts across local and API models?

CptMarvelIsDead · reddit · 2026-09-07

A developer argues general benchmarks are useless for specific use cases as new models ship weekly, and asks: how to score subjective outputs, how to keep LLM judges from favoring their own style, and what tools can fire one prompt at both cloud APIs and local models for side-by-side cost/quality comparison.

Related event: Developers Seek Ways to Benchmark Models on Their Own Prompts as Generic Rankings Lose Value(2 posts)→

Original post →

More from coding & agent

coding & agent channel →