Why Crippled Public Models Don't Reflect True Internal Capabilities

teortaxesTex · x · 2026-07-20

The author points out that a lack of certain capabilities in public models (especially compute-constrained Chinese open-source models) does not mean their internal versions are equally limited. The core reason is the ability to inexpensively assemble "teacher models."

In the referenced supplementary content, the author further explains that companies like Anthropic and OpenAI deliberately weaken the cybersecurity capabilities of their models before public release. Anthropic described this mechanism in the Fable system card, which can hinder specific chains of thought during the training phase. For Chinese vendors, they can train cybersecurity experts via reinforcement learning in dedicated environments and then exclude them during the MOPD (Multi-Objective Preference Learning) phase. While this doesn't significantly degrade their software writing ability, it restricts their agentic thinking regarding vulnerability exploitation.

Original post →

More from Models

Models channel →