AI Competition Shifts to Hidden Data Layer: Expert Judgment and High-Value Signals Become New Moat

Recently, author @darian314 published a series of posts discussing the hidden reality of the AI training data market, arguing that the focus of AI competition has shifted from public compute and benchmarks to an invisible data layer. Public leaderboards are saturated, and the consensus that "open source has caught up" is overly optimistic; true frontier capabilities lie beyond gamifiable benchmarks.

Expensive and Hidden Expert Data

Labs are pivoting towards "paid created data" by purchasing expert judgment at high prices. For instance, they pay doctors by the hour or day to annotate clinical reasoning, or have litigation lawyers revise legal documents. These expenditures are rarely disclosed like compute costs. The scale of this market is astonishing: one data provider reached $1.2 billion in revenue, and another grew from $1 million to a $500 million annualized run rate in just 17 months. Meta invested about $15 billion in a specific data asset, and top labs can spend up to $1 billion annually on training data.

Enterprise Data Sovereignty and Vendor Risks

When enterprises feed their proprietary processes into closed-source models, they are continuously handing over the high-value signals labs want most, effectively "teaching" competing systems to replace them. Furthermore, the boundaries of model vendors are blurring. For example, Anthropic has launched its own drug discovery program, meaning today's AI vendors could become direct competitors in their clients' markets tomorrow. The author noted that an upcoming full article will discuss the rationale behind investing in open-weight labs, data sovereignty tools, and orchestration layers.

2026-07-09 ~ 2026-07-09 · 7 related posts