Local LLM benchmark collection site launches with 30+ domain-specific benchmarks

EmilPi · reddit · 2026-08-02

Reddit user EmilPi shares their new website (beta.locallm.top) for creating and running domain-specific benchmarks for local LLMs. The platform allows creating benchmarks with query sets, defining model/system prompt combinations ('Pipelines'), evaluating responses, and viewing comparison tables. Currently includes 10+ narrow-domain benchmarks (5-18 questions each, created by domain experts in fields like food safety, laser physics, wireless networking, English literature, psychology) and 25+ small 2-4 query benchmarks of varying quality. The author discusses future improvements such as multi-turn evaluation, weighted scoring, additional pipeline types (RAG, structured output, agentic), and whether to support audio/image understanding benchmarks.

Original post →

More from Research

Research channel →