Why this team stopped letting LLMs do compliance math and added episodic memory
Brilliant_Tension_53 · reddit · 2026-09-29
- Author's RAG pipeline for grant-compliance checking failed: LLMs hallucinate math, can't handle rolling time windows, and lose track of cumulative caps; hardcoded SQL lacks institutional memory of waivers and overrun history.
- Hybrid fix: a deterministic SQLite/Python core with absolute veto over hard limits and currency math, plus a semantic episodic-memory layer (Hindsight) for context.
- In practice: memory surfaces the exact waiver code used for a vendor three months ago to clear a blocked foreign invoice in one click; a $1,800 flight gets flagged because that user's invoices historically overrun estimates by 68%.
- Three lessons: never let LLMs evaluate financial logic; tag memories with structured keys to cut vector-search noise; use circuit breakers/outbox pattern so AI latency can't hang the app. Open source at ronaksarda/microsoft-hindsight.
More from coding & agent
- Scraping Xiaohongshu hit posts with Codex + a wired Android phone — huangyun_122 · 2026-09-29
- Relic turns multi-agent collaboration failures into executable org protocols, +5.7pp delivery — ulab-ai · 2026-09-29
- Xiaomi's GAGAR: quality-aware advantage redistribution improves code agent RL — XiaomiMiMo · 2026-09-29
- Same cyberpunk Oregon Trail prompt, wildly different build times on two AI tools — BertMacklenF8I · 2026-09-29
- Do separate verification agents actually fix AI coding's false 'it works'? — T_hompson · 2026-09-29
- ITIS MCP Server Brings the Taxonomic ITIS Database to LLMs via Model Context Protocol — modelcontextprotocol · 2026-09-29