8 of 9 LLMs Fail 'Letter D Days of the Week' — With a RLM Harness All 9 Pass

PawarBI · x · 2026-09-27

Microsoft's Sandeep Pawar shows how a harness beats model swaps: asked how many weekdays contain the letter 'd', 8 of 9 small/mid-size models from OpenAI, Mistral, Google and Qwen answered wrong directly (1/9 correct), yet all 9 got it right via Fabric-RLM. The cause: LLMs read tokens, not letters. Fabric-RLM adds no knowledge — just a Python workspace and loop — proving that how you use a model can matter more than which model.

Original post →

More from coding & agent

coding & agent channel →