Reddit: vLLM and llama.cpp already ship training-free zero-shot classifiers via grammar-constrained output
Altruistic_Heat_9531 · reddit · 2026-10-02
- Reacting to the JEV-style classifier hype, the author points out the capability already exists in local inference stacks: structured output / grammar-constrained generation, with no server-side changes or custom models needed.
- How it works: the decoder constrains the LLM to only emit tokens valid under a given grammar or JSON schema. vLLM supports backends like XGrammar and lm-format-enforcer; llama.cpp uses GBNF.
- Includes a working vLLM example: a JSON schema for sentiment classification (negative/neutral/positive + score + value), using responseformat: schema with temperature: 0.
- The author, who works on lakehouse platforms, uses 1B–4B models to extract structured fields (place, time, sentiment, entities) from unstructured images/documents, then queries via SQL instead of repeatedly hitting Elastic/OpenSearch.
More from coding & agent
- Indie dev: all our projects run on TanStarter — Cloudflare costs + agent-friendly setup — yihui_indie · 2026-10-02
- Three ChatGPT coding integration workflows worth comparing: Firecrawl, Figma, Drive — EstablishmentSea4024 · 2026-10-02
- Indie dev ships AI-built games to all platforms at once, from iOS to WeChat mini-programs — ezshine · 2026-10-02
- Prompt2Skill Builds LLM Skills From a Single Prompt, +10.8 Average Across Four Domains — Bo Ni · 2026-10-02
- InFlowOp: Label-Free In-Flow Multi-Agent Workflow Optimization Gains up to +11.97% — Xuehang Guo · 2026-10-02
- Open-source Harness fork moves coding agents out of the app into orca — dee_hw · 2026-10-02