How developers benchmark their own prompts across local and API models
CptMarvelIsDead · reddit · 2026-09-07
A Reddit user asks how to evaluate models on their exact prompts rather than generic benchmarks, to decide whether a new API is worth the cost or a small local model suffices. Eyeballing outputs is driving them crazy.
Three concrete questions: how to score "good" responses when output is subjective; how to stop an LLM judge from favoring its own writing style; and the easiest tooling to fire one prompt at multiple cloud APIs and local models for side-by-side comparison. Community answers typically point to LLM-as-a-judge setups and tools like promptfoo.
More from coding & agent
- Designer fed up with AI-generated UI inconsistency plans BaseUI refactor with strict rules — tomjohndesign · 2026-09-07
- From low-code to 'woah code': the AI coding era's new meme — jxnlco · 2026-09-07
- Developer on Astra: code as unreadable as minified JS, but tolerable to boss around — Aryvyo · 2026-09-07
- Your LLM Gateway Holds the Keys: Rethinking LiteLLM Security for Action-Taking Agents — Technical_Map_2105 · 2026-09-07
- Zero Blender skills, one prompt, 40 minutes: recreating a game with Codex — FuSheng_0306 · 2026-09-07
- LLM-powered revival of Put-That-There brings speech and gesture window control to XR — twi_mar · 2026-09-07