Sapience System Reads Only ~15K Tokens Yet Beats GPT-6 Astra on MRCR, 97.7 vs 96.3

gordic_aleksa · x · 2026-09-15

A surprising long-context result: Sapience, a custom system that reads only 15K tokens and delegates answering to open model DeepSeek V4 Pro, scores 97.7 on OpenAI's MRCR benchmark — beating GPT-6 Astra's 96.3, which must ingest the full 500K–1M token context.

The author notes similar gains on RULER and other long-context benchmarks, suggesting it's not overfit to a single benchmark — hinting that smart retrieval may beat brute-force full-context reading.

Related event: Sapience tops GPT-6 Astra on MRCR reading just 15K tokens(2 posts)→

Original post →

More from Models

Models channel →