typevet: Gemma 4 31B on one 4090 gives per-label probabilities — and catches receipt fraud text-only misses
One_Temperature5983 · reddit · 2026-09-30
The author open-sources typevet (MIT, Python 3.12): it composes a native Gemma 4 turn, ends with a no-thinking prefill, and reads the next-token distribution over allowed answer tokens — returning typed yes/no, pick-one, or rubric answers with per-option probabilities, plus images via llama.cpp's multimodaldata or vLLM imageurl blocks, and a JSON-Schema path. In a receipt test on 6 real CORD v2 receipts with 3 synthetic claims each, text-only answers matched swapped-digit totals at 0.99964 (0/6 caught), while adding the photo flipped all 6 to mismatch at ≥0.99998; masked totals abstained every time. Hosted on one H100 with vLLM 0.30.0 BF16, it went 18/18, was order-invariant, and hit 39.6 records/s at 64 concurrent on Banking77. Limitations: 18 claims, one run, uncalibrated probabilities.
More from coding & agent
- Replit Adds In-App Funnel Analytics: Ask Agent to Turn It Into a Dashboard — amasad · 2026-09-30
- Compound Engineering 3.30 ships: skill prose tweaks curb model overbuilding — kieranklaassen · 2026-09-30
- Lovable apps can now run inside your Microsoft tenant with Entra ID and M365 data — satyanadella · 2026-09-30
- OpenAI Codex lead presses the reset button, signaling a Codex reset — Xianbao_QIAN · 2026-09-30
- Prediction: in an agentic world humans won't code, write, or draw — editors have no future — akbirthko · 2026-09-30
- Google shows how to turn DiffusionGemma into a scorer via vLLM templates — bodonoghue85 · 2026-09-30