How developers benchmark their own prompts across local and API models

CptMarvelIsDead · reddit · 2026-09-07

A Reddit user asks how to evaluate models on their exact prompts rather than generic benchmarks, to decide whether a new API is worth the cost or a small local model suffices. Eyeballing outputs is driving them crazy.

Three concrete questions: how to score "good" responses when output is subjective; how to stop an LLM judge from favoring its own writing style; and the easiest tooling to fire one prompt at multiple cloud APIs and local models for side-by-side comparison. Community answers typically point to LLM-as-a-judge setups and tools like promptfoo.

Related event: Developers Seek Ways to Benchmark Models on Their Own Prompts as Generic Rankings Lose Value(2 posts)→

Original post →

More from coding & agent

coding & agent channel →