Build a Personal RAG Harness Without Frameworks: Jev Judges, a Local Qwen3 Writes
ghumare64 · x · 2026-10-02
Rohit Ghumare shares a reproducible personal RAG harness that splits roles across models instead of letting one model search, write, and grade itself.
Architecture & flow
- Corpus: 70 of his own pages (52 essays + 18 guides), split into 836 chunks
- Retrieval: nomic-embed-text via ollama with numpy cosine search, pulling 20 candidate passages
- Judgments: Jev (TypeSafe /v1/systemone) handles small typed decisions — whether to search and which passages actually help
- Writing: local qwen3 8B via ollama (8k context, thinking off) answers from the top 5 passages, citing every sentence
- Verification: Jev checks each cited sentence against its passage and escalates when a claim doesn't hold
Implementation details
- 10 short Python files using stdlib + numpy, zero agent frameworks
- Every stage logs time, Jev tokens, and cost; Jev stages took 0.5s each in runs on an Apple Silicon Mac
- A held-out question with a deliberate false premise shows the gold passage climbing from rank 15 to rank 1
Key idea: let a fast judgment model handle yes/no decisions while the writer model only writes — role separation yields more reliable citation checking and transparent cost tracking.
More from coding & agent
- What should 'approving' an AI memory actually mean? A builder's design dilemma — BarcodeCutter · 2026-10-02
- ChatGPT writes drafts via MCP, Claude edits them — two models, one shared board — BarcodeCutter · 2026-10-02
- NVIDIA paper: model accuracy drops 62.8% on 128K-token tasks vs 4K — rohanpaul_ai · 2026-10-02
- Sub-agents inherit parent tokens: a research agent opened a PR on its own — Imprgdessie_Land3313 · 2026-10-02
- One-shot Fable 5.5 prompt yields 1000-line three-body sim in Bend: 60 FPS, 4 formally proven laws — sprooos · 2026-10-02
- Dev urges junior coders to ditch handwriting code, like pencil skills after speech input — brandon_galang · 2026-10-02