The real frontier: models grinding for days on deterministic, verifier-proof tasks

mike64_t · x · 2026-09-05

The author argues that determinism will grow in value as AI evolves, and deterministic processes will increasingly look like grindable, harness-like workloads.

His proposed yardstick for the true frontier of model capability usage: a model that, from a single prompt (no explicit /goal), grinds for multiple days without giving up—sustaining consistent progress simply because it cannot cheat the verifier implied to it.

An opinion thread, but the core claim is notable: strictly verifiable deterministic task structures may reveal and measure long-horizon agent capability better than open-ended tasks.

Related event: Deterministic verifiable tasks and multi-day grinding seen as agent frontier(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →