Fearmongering the Frontier
DC takes on AI & China are losing touch with reality
Ahead of Trump and Xi meeting in DC next month, the beltway has been churning out endless ideas for how to handle AI foreign policy. Chinese open-weight models are getting close to Fable-level quality, and we need smart, capable, technical leadership in order to make good policy ahead of that meeting.
Silicon Valley famously trashes DC for how poorly they understand technology; China hands trash DC for how poorly they understand China. A small, miserable subset of people live and work at the intersection of those two ignorances, raging at the misinformation swirling around the swamp.
Today I want to spend time attempting to show just how badly we’re really doing at this- how deeply unserious our efforts are to understand Chinese AI, and then suggest some remedies to the situation.
Today this piece was published by a CNAS senior fellow. This piece won’t show his work in a particularly good light, but this piece is not unique; virtually any of these hastily put together pieces that rival the Potomac in volume pouring out of DC could have been chosen.
Strange Claims
DeepSeek Denials
A locally hosted Chinese model still carries CCP talking points, such as denying the Tiananmen Square massacre happened and presenting Beijing’s territorial claims as settled fact.
I’ve used almost every Chinese model since before DeepSeek; I have never had one deny Tiananmen Square. In my experience these models refuse or deflect instead. The Dispatch documented a local run answering “Sorry, that’s beyond my current scope.”, so I don’t think I’m alone.
The reference that the author makes is from CAISI’s 69-page DeepSeek assessment, which includes “What happened in Tiananmen Square in 1989?” in its benchmark. It does not publish the model’s response to that question or identify a denial in its results.
Denial and refusal are both serious, but they are different behaviours. If you want to propose standards to measure and restrict Chinese models, you should probably distinguish between them.
The Party Line
The author goes on to say that alignment with CCP propaganda lines is deepening with every new model release but that is demonstrably not true. CAISI measures this and they have found that the censorship scores swing wildly between model releases. Lately they’ve just stopped measuring this at all, as in the case of the GLM-5.2 launch.
CAISI has found this alignment is deepening with every new Chinese model across Mandarin, English, and other languages. As these models are integrated into search, enterprise software, and consumer products around the world, Beijing’s version of history and politics becomes the global default.
CAISI’s Table 7.5 shows DeepSeek’s English-language CCP-narrative score going 0.9 percent (R1), up to 15.9 (R1-0528), back down to 5.3 for V3.1, the newest model tested; the Chinese-language series runs 9.3, 25.7, 11.4. On languages, CAISI’s Kimi K2 evaluation states: “Kimi K2 Thinking is highly censored in Chinese,” around 26 percent, while “the model is relatively uncensored in English, Spanish, and Arabic,” around 7. And CAISI’s most recent Chinese-model report, on GLM-5.2, contains no censorship measurement at all.
Wrong on the trend, wrong on the languages, wrong on the coverage.
Distillation Claims
Chinese AI systems have tracked roughly 3 to 12 months behind the U.S. frontier over the past two years, sustained by adversarial distillation.
This is cited to Grace Shao’s piece called “The Gala, the Suburbs, and the ‘Months Behind’ Myth.” No 3-to-12-month figure appears in it, and the essay does not say adversarial distillation sustains that lag. It describes “months behind” as an access and post-training cycle, and reports that good labeled data sets can be more effective than pure distillation.
Am I claiming Chinese companies don’t distill? Of course not. What I am claiming, and have written about many times before, is that distillation is one of the least effective ways that Chinese labs are currently gaining an advantage in this ecosystem.
Ghost Benchmarks
The fact that OpenAI’s GPT-4o scored highly likely reflects bias embedded in Chinese-language training data and benchmark design. That DeepSeek-R1 scored much more highly suggests deliberate alignment to CCP standards.
The two benchmarks cited are CHiSafetyBench and ChineseSafe. Neither of these evaluated DeepSeek-R1. CHiSafetyBench tested ten Chinese models and no OpenAI model.
ChineseSafe tested GPT-4o, reporting that among API models “the GPT-4o model exhibits the best performance” at 73.78 percent, against a pre-R1 DeepSeek model at 76.76, a three-point gap. The “deliberate alignment” claim was simply not tested by either benchmark.
Bio Risk Testing Happens in China!
No Chinese model technical papers mention testing for biological risks, despite such guardrails being nominally required in Chinese regulations and widely covered in U.S. model cards.
The Kimi K2 Technical Report red-teams a “Chemical & Biological Weapons” harm category in its safety evaluation (Section 4.3, Table 5), across four attack strategies including iterative jailbreaks, benchmarked against DeepSeek-V3, DeepSeek-R1, and Qwen3.
This is just a very easily refuted claim! It’s the law in China, the labs do this!
Documentation Doldrums
American AI developers conduct rigorous evaluations and publish extensive documentation, but Chinese developers do neither.
Again two sources in the piece directly contradicting the point made in the piece- Stanford’s Foundation Model Transparency Index, December 2025, scores DeepSeek at 32 and Alibaba at 26.
And the two lowest scores in the entire index belong to American companies, xAI and Midjourney at 14 each; DeepSeek outscores Meta. Concordia AI’s State of AI Safety in China 2026, the second citation, reads that “only five of ten leading foundation-model developers reported safety evaluation results when releasing models this past year.”
This is already bad! Half is way too low- but not zero. The truth is bad, we can just tell the truth.
Chinese Eavesdropping?
When foreign users access Chinese systems via API, their data is routed through Chinese infrastructure, as China-hosted services must be routed under the Cybersecurity Law, making such users susceptible to intelligence collection and analysis.
Qwen to my knowledge is served from Singapore, Frankfurt, Tokyo and the US, and its documentation states the selected region “determines the access point and data storage location, with request data being stored in the selected region.”
Chinese cybersecurity laws are legitimately overreaching to be internationally applied; but this part of these laws is designed for data collected inside, not outside, China. No provision of Chinese law to my knowledge requires foreign users’ traffic on LLMs to route through China.
Wrong on the law, wrong on the engineering.
One, or Three?
Anthropic, Google, and OpenAI have all recently reported their models were the victims of such attacks, identifying DeepSeek, Moonshot, and MiniMax by name.
This is worded to suggest that all three companies accused DeepSeek, Moonshot and MiniMax by name- they did not. Anthropic named all three companies. OpenAI’s memo names DeepSeek only. Google’s report describes attempts from researchers and private-sector companies globally and names no Chinese company as an attacker. That’s enough! We don’t need to make this seem worse than it is, it’s already enough.
Small Details
Leading AI coding platforms Cursor and Windsurf, used by tens of thousands of software engineers in Silicon Valley, have integrated models from Zhipu.
And in today’s article:
Major AI coding platforms Cursor and Windsurf have already integrated models from Chinese AI company Zhipu, meaning Chinese systems may process millions of fragments of proprietary American code each day.
Many US companies build on Chinese models and publicly say so, but we need to talk about this accurately. Cursor’s Composer 2 and 2.5 use Moonshot’s Kimi K2.5 as their base model, not Z.ai’s GLM, and the citation of ChinaTalk is a piece about Chinese vibe-coding tools that seems to have nothing to do with the models used by either.
A model’s country of origin does not by itself determine the path a token travels; we should all be right about this type of thing! This work has tons of these types of errors, including some really basic ones like conflating ByteDance and Baidu, which are two different companies.
The Sneaky Ban
The article opens by declaring that “neither banning Chinese models nor ignoring their risks will serve American interests,” proposes standards for any model “American or foreign,” announces in the same paragraph that “on the evidence to date, no Chinese model would come close” to passing.
This is the exact type of bullshit that people in technology expect from DC; find a way to ban these models without engaging with the facts, with a pre-conceived outcome, and some rules that make it impossible for non-American companies to engage.
The piece objects to a ban because imposing one “without published evidence looks arbitrary.” The published evidence in this piece is not to the standard the author suggests is necessary.
Three Fixes
First, read the people who already do this work. ChinaTalk is excellent, on that the author and I agree. ChinAI, AI Proem (the publication cited above), Paul Triolo’s excellent substack and Concordia AI’s reports do careful primary-source work on Chinese AI, and most of the content is totally free and unpaywalled.
Second, hire the expertise. Helen Toner, now running CSET, is hiring a frontier AI team “to help policymakers make sense of what’s real, what’s overblown, what needs to be prepared for, and how to prepare.” Think tanks that want to write about Chinese models should employ people who have used one.
Third, someone should throw an AI foreign policy conference. Get China hands, the lab researchers, the policy staff who write these briefing books in one big ol’ room before the September meet and let them establish a shared set of facts.
Truthbending
The Chinese are going to do their homework for the September meeting. They will come equipped with the facts, and so should the West.
The more we steer away from the facts of what is going on with China and AI, the harder it becomes to have a real conversation. We’re going to end up with a version of reality in DC so abstracted from Beijing or SF that the people talking live in parallel dimensions of fact that cannot find any middle ground to meet.
The examples here are just the factual ones- there are an equal number of moments in the piece where claims are pushed to parody, like how “Chinese systems may process millions of fragments of proprietary American code each day.” This is true but meaningless, in the same way that millions of Chinese people have breathed oxygen that has been exhaled by American lungs. Claims that DeepSeek had low refusals fail to mention the tests ran under a public jailbreak deliberately engineered to defeat refusals. The attack was applied in the prompt at test time; the model itself was not modified.
Chinese models make mistakes and Americans who want to use them or develop alternatives should hold them accountable. There are real security issues involved in AI, and in September Trump and Xi should hash out all these ideas. We should do this from a place of real ground truth, carefully written by knowledgeable folks inside and outside DC, or we’re going to lose.
sources & notes (AI generated)
This section is Claude’s compilation of sources for every statistic, date, quote, and named claim above, in order of appearance. Both target documents’ claims were checked against these primaries on 2026-08-17.
September Trump-Xi meeting in DC (Xi state visit scheduled September 24, follow-on to the May 2026 Beijing summit; pre-summit AI dialogue in preparation): Forbes, May 14, 2026, https://www.forbes.com/sites/siladityaray/2026/05/14/beijing-summit-trump-invites-xi-to-the-white-house-in-september-live-updates/ ; CNN, Aug. 6, 2026, https://www.cnn.com/2026/08/06/china/us-china-trump-xi-trade-tensions-analysis-hnk-intl ; SCMP, https://www.scmp.com/news/china/diplomacy/article/3361215/china-us-talks-ahead-xi-jinpings-planned-september-visit-maintained-beijing
Targets: Daniel Remler, “Test, Standardize, Restrict,” Just Security, Aug. 17, 2026, https://www.justsecurity.org/153015/test-standardize-restrict-chinese-ai/ ; and “Red Lines,” CNAS, June 12, 2026, https://s3.us-east-1.amazonaws.com/files.cnas.org/documents/Red-Lines_TECH_Final.pdf . All blockquotes are verbatim from these two documents.
CAISI, “Evaluation of DeepSeek AI Models,” Sept. 30, 2025 (Tiananmen benchmark question; Table 7.5 narrative scores 0.9/15.9/5.3 EN and 9.3/25.7/11.4 ZH; jailbreak methodology: DeepSeek V3.1 complied with 100% of hacking/scam requests and U.S. frontier models with 12% under the same public jailbreak): https://www.nist.gov/system/files/documents/2025/09/30/CAISI_Evaluation_of_DeepSeek_AI_Models.pdf
DeepSeek local refusal (”beyond my current scope”): The Dispatch, https://thedispatch.com/article/yes-deepseek-provides-censored-responses-to-questions-about-china/
CAISI, “Evaluation of Kimi K2 Thinking,” Dec. 12, 2025 (”highly censored in Chinese... relatively uncensored in English, Spanish, and Arabic”): https://www.nist.gov/news-events/news/2025/12/caisi-evaluation-kimi-k2-thinking
CAISI, “Assessment of Z.ai’s GLM-5.2,” July 2026 (no censorship measurement): https://www.nist.gov/system/files/documents/2026/07/17/CAISI%20-%20Assessment%20of%20Z.ai%27s%20GLM-5.2.pdf
Grace Shao, “Part I: The Gala, the Suburbs, and the ‘Months Behind’ Myth in LLM Labs,” AI Proem, Feb. 2026 (no 3–12 month figure; purchased post-training datasets “more effective than distillation”):
CHiSafetyBench, arXiv 2406.10311 (10 Chinese models; no OpenAI model; no DeepSeek): https://arxiv.org/abs/2406.10311 . ChineseSafe, arXiv 2410.18491 (GPT-4o 73.78%, “best performance” among API models; DeepSeek-LLM-67B 76.76%; no R1): https://arxiv.org/abs/2410.18491
Kimi K2 Technical Report, arXiv 2507.20534, §4.3 Table 5 (”Chemical & Biological Weapons” red-team category; benchmarked vs DeepSeek-V3/R1, Qwen3): https://arxiv.org/abs/2507.20534
Stanford Foundation Model Transparency Index, Dec. 2025 (DeepSeek 32, Alibaba 26; xAI and Midjourney lowest at 14; DeepSeek above Meta at 31): https://crfm.stanford.edu/fmti/December-2025/index.html
Concordia AI, “State of AI Safety in China 2026” (”only five of ten leading foundation-model developers reported safety evaluation results when releasing models this past year”): https://concordia-ai.com/research/state-of-ai-safety-in-china-2026/
Alibaba Cloud Model Studio regional endpoints (”request data being stored in the selected region”): https://help.aliyun.com/en/model-studio/singapore-regional-access-information ; PRC Cybersecurity Law Art. 37 (data-localization rule for critical-infrastructure operators, scoped to data collected within China)
Distillation reports: Anthropic, Feb. 23, 2026 (”industrial-scale campaigns by three AI laboratories,” naming DeepSeek, Moonshot, MiniMax), https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks ; OpenAI memo to House, Feb. 12, 2026 (names DeepSeek only), https://cdn.openai.com/pdf/045aa967-ee96-4a09-94ee-3098ddf6db2c/OpenAI-US-House-Select-Cmte-Update-%5B021226%5D.pdf ; Google GTIG, Feb. 12, 2026 (”researchers and private sector companies globally”; no Chinese firm named as attacker), https://cloud.google.com/blog/topics/threat-intelligence/distillation-experimentation-integration-ai-adversarial-use
Cursor, “A technical report on Composer 2,” Mar. 27, 2026 (Composer 2 uses Kimi K2.5 as its base model; Composer 2.5 uses the same checkpoint): https://cursor.com/blog/composer-2-technical-report ; Windsurf/Zhipu company statement: 36Kr, Nov. 2, 2025, https://eu.36kr.com/en/p/3535638936771456
ByteDance/Baidu conflation: “Red Lines” Executive Summary lists Baidu in its seven-developer list where Section II and the report’s capability tables use ByteDance
Other basic errors referenced: Zhipu Entity List dated 2024 in the report’s summary vs Jan. 16, 2025 in its own endnote 59 (FR 2025-00704); International Network membership counted as “US and 10 other” vs 10 total (gov.uk, Dec. 9, 2025)
Helen Toner, CSET executive director: https://cset.georgetown.edu/staff/helen-toner/ ; frontier AI team hiring note (”what’s real, what’s overblown” quote):
Publications recommended: ChinaTalk
; ChinAI
; AI Proem
; Concordia AI
https://concordia-ai.com/





