U.S. Policies Unintentionally Accelerated China's Open AI Ecosystems
Wang Jin, Nadav Kunievsky, Bowen Lou, Tianshu Sun, James Evans
econ.GN, cs.CY
2026-06-15
Across six data sources, Chinese developers raised open-LLM-repo forking at roughly 12x the U.S. rate after the 2022 chip controls, yet Chinese-origin models are nearly absent from U.S. patents.
Starting in 2022, U.S. export controls on AI chips were designed to squeeze the compute chokepoint and slow China's march toward frontier models. This paper asks a question nobody had measured: did those controls have a side effect, pushing China toward open-source AI?
The authors' argument is that once compute got scarce, "downloadable, modifiable, locally runnable" open models became strategically cheap for China, because they bypass foreign-controlled closed platforms and high-end hardware. Containment did not simply slow a rival; it nurtured an open ecosystem as a hedge. The paper tests that intuition across six data sources.
The study stitches together six sources: Chinese AI and open-source policy documents, a hand-curated open-weight model release database, GitHub event records, arXiv metadata with author-country attribution, company-affiliation evidence for arXiv papers, and full-text U.S. patent records.
The centerpiece is a GitHub event study. The authors pick four U.S. policy shocks (the CHIPS and Science Act of August 2022; the October 7, 2022 export controls; the October 17, 2023 expansion; the December 2, 2024 revision), arrange each LLM-related repository's fork counts into a panel indexed by weeks relative to the shock, and watch activity jump at the moment of impact.
Country attribution uses a time-zone proxy: GitHub events at hour 17 or earlier are assigned to China, later than 17 to the U.S. It is crude but functional, because developer activity tracks local clock time. The authors stress repeatedly that this is descriptive and quasi-experimental, capturing "timing and relative response," not a structural causal model.
A second panel tracks efficiency research: arXiv AI papers from 2022 through 2025 are classified by title and abstract into six compute-restriction topics (compression, parameter-efficient fine-tuning, inference efficiency, training compute, memory efficiency, edge deployment), to see whether China's research emphasis shifted toward compute-saving methods after the controls.
The GitHub numbers are the most direct. Across the pooled shock window, China-associated LLM repositories gained 0.143 forks per repository-week, against 0.012 for U.S.-associated ones, a gap of 0.131, roughly 12x the U.S. rate. The non-LLM control group shows no comparable split, so this is not a general tide lifting everything.
The efficiency panel agrees: after the 2022 controls, China's relative research weight rose in parameter-efficient fine-tuning, inference optimization, compression, and edge deployment, consistent with "compute got expensive, so pivot to efficiency."
The diffusion results tell an asymmetric story:
| Domain | Finding |
| arXiv science | Chinese models (Qwen, DeepSeek) adopted by Chinese and U.S. researchers at comparable rates |
| Company research | U.S. and Chinese firms' LLM-use distributions are highly similar, with slight home bias |
| U.S. patents | Heavy references to LLaMA, ChatGPT, Claude; Qwen and DeepSeek nearly absent |
| GitHub repos | Chinese models diffused fast; Chinese developers also engaged strongly with U.S. open models |
In science and open-source communities, Chinese models move freely. The moment innovation enters formal commercialization through the patent system, they largely vanish.
For practitioners, this gives an explanation for why China's open models (Qwen, DeepSeek) surged that goes beyond "they can just build them." Under compute constraints, openness doubles as innovation community and resilience infrastructure: locally deployable, modifiable, free of foreign platforms. That pins down the strategic position of the Qwen and DeepSeek line.
The patent gap is the finding to remember. U.S. companies genuinely use Chinese models in open research, yet almost never disclose them in patents. If patents are the official gauge of commercial dependence, that dependence is systematically understated.
For policymakers, the conclusion resists closure: restricting frontier compute may not merely slow a rival; it can reshape the rival's innovation architecture toward something more open, more distributed, and harder to choke.
The biggest one, which the authors state plainly: this is descriptive and quasi-experimental evidence, not a structural causal model. The "accelerated" claim rests mostly on timing correlation, with no clean counterfactual to peel the control effect away from the general 2022-2024 LLM boom.
The country proxy is coarse. Splitting China and the U.S. by event hour misclassifies multinational teams, remote workers, and off-hours contributors; the authors do not separately bound that error.
The absolute fork numbers are small (0.143 per repository-week). Statistically visible, but a thin basis if read as a measure of whole-ecosystem scale.
The patent asymmetry has an unexcluded alternative: patents lag by years from filing to publication, and Chinese models only scaled broadly in 2024. Their absence from patents now could be a disclosure-lag artifact rather than institutional filtering. The authors lean toward the filtering reading but do not close off the lag explanation.