Multi-Agent System With 4-Layer Memory Learns Procedural Rules From Its Own Resolutions

yashghotekar · reddit · 2026-09-30

A team built a self-learning multi-agent architecture to fix memoryless RAG, converting successful resolutions into persistent procedural rules. Stack: plain Python orchestrator, Hindsight 4-layer memory (episodic/semantic/procedural/preference) to keep chat logs from polluting semantic search, Qdrant static docs index, a Critic Agent with a hard 2-redraft cap, and async reflection off the request path — ReflectionAgent scores outcomes and writes new rules at ≥0.8 confidence. Live example: an HTTP 429 during bulk DB sync becomes a stored "cap batch at 50" rule applied automatically next time. Open issues: rule drift/dedup, org-level memory scoping, rule summarization.

Original post →

More from coding & agent

coding & agent channel →