8 of 9 LLMs Fail 'Letter D Days of the Week' — With a RLM Harness All 9 Pass
PawarBI · x · 2026-09-27
Microsoft's Sandeep Pawar shows how a harness beats model swaps: asked how many weekdays contain the letter 'd', 8 of 9 small/mid-size models from OpenAI, Mistral, Google and Qwen answered wrong directly (1/9 correct), yet all 9 got it right via Fabric-RLM. The cause: LLMs read tokens, not letters. Fabric-RLM adds no knowledge — just a Python workspace and loop — proving that how you use a model can matter more than which model.
More from coding & agent
- Dev hails Claude Code's Dynamic Workflows with Opus 5.5 as best multi-agent experience yet — daniel_mac8 · 2026-09-27
- One Opus 5.5 prompt builds a virtual office for managing GitHub agents — minchoi · 2026-09-27
- Microsoft open-sources SkillOpt: training skill docs lifts GPT-5.5 accuracy 23.5 points — sanjaykalra · 2026-09-27
- Hacker uses NVIDIA DGX Station to run local models for CAD, invites SF folks to join — hudzah · 2026-09-27
- Agent control plane: model proposes, plane decides — no standing write credentials for the model — brucemacv · 2026-09-27
- SlopOps: mistaking AI capability for competitive advantage — generativist · 2026-09-27