How do you benchmark specific prompts across local and API models?
CptMarvelIsDead · reddit · 2026-09-07
A developer argues general benchmarks are useless for specific use cases as new models ship weekly, and asks: how to score subjective outputs, how to keep LLM judges from favoring their own style, and what tools can fire one prompt at both cloud APIs and local models for side-by-side cost/quality comparison.
More from coding & agent
- Designer fed up with AI-generated UI inconsistency plans BaseUI refactor with strict rules — tomjohndesign · 2026-09-07
- From low-code to 'woah code': the AI coding era's new meme — jxnlco · 2026-09-07
- Developer on Astra: code as unreadable as minified JS, but tolerable to boss around — Aryvyo · 2026-09-07
- Your LLM Gateway Holds the Keys: Rethinking LiteLLM Security for Action-Taking Agents — Technical_Map_2105 · 2026-09-07
- Zero Blender skills, one prompt, 40 minutes: recreating a game with Codex — FuSheng_0306 · 2026-09-07
- LLM-powered revival of Put-That-There brings speech and gesture window control to XR — twi_mar · 2026-09-07