LLM Writing Quality by Language — Chinese

Raw research notes for Chinese, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-chinese.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)


Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)


Summary verdict (4-8 sentences)

Native Chinese assessments split sharply by register. For logic-heavy non-fiction (reports, academic prose, analysis), reviewers consistently rate Claude’s Chinese output at the top tier — “逻辑清晰、结构完整、论证严密” with minimal 翻译腔 — while noting GPT-series Chinese still carries a “塑料感” (plastic feel) and that all Western models lack 地道感 (native groundedness): weak on internet slang, dialects, and China-domestic knowledge. For fiction, the picture inverted somewhat during 2025-2026: domestic models (DeepSeek, Kimi, Qwen, Doubao) are credited with more native-feeling colloquial Chinese and web-novel fluency, with reviewers claiming they “已达国际一线水平” (reached international first tier), while one influential 2026 web-novel guide still scored Claude highest for literary sentence craft (“最接近’人写’的感觉”). The most credentialed literary voices — Mao Dun Prize winner 麦家, web-fiction editors, 番茄小说 writers — are broadly dismissive of unassisted AI fiction: technically fluent but “寡淡无味,套路化…没有情感” (bland, formulaic, emotionless), with long-form logic and character-consistency collapse as the dominant complaint. The canonical failure modes natives name are “AI味” (首先/其次 scaffolding, piled-up 排比 parallelism, over-ornamented but empty prose, hedged fake-objectivity, formulaic endings) rather than grammatical error. Notably, the sharpest literary critiques date to the early-2025 DeepSeek-R1 moment; 2026-era commentary on frontier models comes mostly from tech media, vendor blogs, and 知乎, not literary critics — a real evidence gap.

Sources

  1. 麦家谈DeepSeek对文学创作的冲击https://finance.sina.com.cn/tech/digi/2025-03-09/doc-inenzies2017417.shtml

    • Author: 麦家 (Mai Jia) — novelist, Mao Dun Literature Prize winner, Vice Chairman of the China Writers Association. Outlet: 新浪科技. Date: 2025-03-09. Model: DeepSeek-R1 (vintage: early 2025).
    • Quote: “DeepSeek可能比95%的人写得好,但问题在于没法暴露人的局限性,而人的局限性恰恰是很多经典的灵感来源。” (“DeepSeek may write better than 95% of people, but it cannot expose human limitations — and human limitations are precisely the source of inspiration for many classics.”) Also: “正因为没有人类的局限,甚至缺陷,机器在文学创作上永远无法超越人。”
    • Claim: AI prose is competent beyond most humans but structurally incapable of the flawed particularity that makes literature.
  2. AI闯入网文赛道,网文编辑:工作量多了50%https://www.chinawriter.com.cn/n1/2025/0226/c404027-40426506.html

    • Authors quoted: working web-fiction editors and a writer/screenwriter (小白, 缦彩笺). Outlet: 封面新闻 via 中国作家网 (China Writers Association’s official site). Date: 2025-02-26. Model: DeepSeek-R1 (vintage: early 2025).
    • Quotes: “直接全部丢给AI进行生成的文稿,出来的东西很缺乏逻辑性和代入感” (“Manuscripts generated wholesale by AI severely lack logic and immersion”); “寡淡无味,套路化,最重要是没有情感,没有创新,全是老梗” (“Bland, formulaic; above all no emotion, no originality — all stale tropes”).
    • Claim: AI slush flooded submission queues (+50% editor workload) and editors can identify and reject it on quality grounds.
  3. 番茄小说的AI难题https://36kr.com/p/3501232140474501

    • Author: 何旭, 海克财经 (business/tech analysis outlet), republished on 36氪. Date: 2025-10-09. Models: DeepSeek V3/R1, ChatGPT, 阅文妙笔, 文心一言 (vintage: mid/late 2025).
    • Quotes from working web-novelists: “人物关系前后矛盾” (character relationships contradict across chapters); “喜欢乱加戏,导致人设走偏” (adds gratuitous scenes, derailing characterization); “文笔漂亮但没有’时间’的概念” (pretty prose but no sense of time); DeepSeek openings “堆满精密数字和专业名词” (stuffed with precise numbers and jargon).
    • Claim: on China’s biggest free-reading platform, AI assistance boosts output volume but produces recognizable structural/temporal incoherence; one author nonetheless reported viable read-completion rates (~23%) with 1 hr/day of AI-assisted work.
  4. 2026 AI写小说用哪个模型?实测排名https://maliangwriter.com/blog/ai-model-selection-guide/

    • Author: 马良写作 team (commercial Chinese AI-writing platform — vendor interest; flag). Date: 2026-03-04, updated 2026-06. Models: Claude Sonnet/Opus “4.6”, Gemini 2.5 Pro, GPT o3, DeepSeek V3/R1, Qwen3, Llama 4.
    • Quotes: Claude — “句式富于变化,文学性强,最接近’人写’的感觉” (varied sentence patterns, strong literariness, closest to human-written feel; 9.2/10); DeepSeek — “中文理解深厚,偶有’教科书感’” (deep Chinese comprehension, occasional textbook feel; 8.8); GPT — “英文训练为主,中文文学性相对弱” (English-centric training, weaker Chinese literariness; 7.5); trend note: 国产模型 “已达国际一线水平”.
    • Claim: for Chinese fiction, Claude leads on prose craft, DeepSeek/Qwen are close and far cheaper — the recommended workflow mixes DeepSeek for volume and Claude for quality chapters.
  5. 2026高考作文AI大测试:豆包可能需要复读了https://www.163.com/dy/article/KURO27BH0556MPIL.html

    • Author: 网易号 “AI唱反调” (independent Chinese AI commentator; informal; essays were graded by ChatGPT-5.5, not human graders — flag). Date: 2026-06-07. Models: DeepSeek, Kimi, 通义千问, 豆包, 文心一言, 智谱清言, 腾讯元宝 (current 2026 domestic frontier).
    • Quotes: Kimi (56.5, 1st) — “最接近真实考生” (closest to a real examinee); 文心一言 praised for sensory-detail density (“暖黄的灯光落在她发梢”); 豆包 (49.5, last) — “AI腔明显” (obvious AI voice), formulaic three-part structure and 排比 without lived detail.
    • Claim: even among 2026 domestic models, the spread between “reads like a person” and “reads like AI” remains wide on the gaokao-essay register.
  6. 写中文论文 Claude / DeepSeek / GPT 哪个更好?2026深度实测https://paper.inkfount.com/blog/chinese-academic-writing-ai-comparison-2026-inkfount

    • Author: InkFount (research-writing SaaS — vendor blog; flag). Date: 2026-04-26. Models: Claude 3.5/4-series, DeepSeek R1/V3, GPT-4o/5 (mixed vintage labels).
    • Quotes: Claude captures “委婉但严谨” academic rhetoric with minimal 翻译腔; DeepSeek shows “比国外模型更高的’原生感’” (more native feel than foreign models) for Chinese citations/derivations; GPT sometimes has “语感’塑料感’” (plastic-feeling Chinese).
    • Claim: for Chinese academic non-fiction, DeepSeek is the most natively idiomatic, Claude the most rhetorically refined, GPT the most structurally balanced but least natural-sounding.
  7. Claude中文能力深度测评(2026版)https://www.claude-anthropic.com/guide/289.html

    • Author: anonymous; third-party SEO-style “Claude 中文” site, not Anthropic; low editorial credibility, but synthesizes 知乎 evaluations. Date: 2026-03-24. Models: Claude Sonnet/Opus 4.6, GPT-5.4, DeepSeek, 文心, 通义, 豆包.
    • Quotes: strengths “逻辑清晰、结构完整、论证严密”; a 知乎 workplace-writing test found Claude’s Chinese weekly reports “最难被识别为AI写作” (hardest to identify as AI-written); weaknesses — 网络用语/方言 “不够接地气”, and limited “中文创意表达的’地道感’”.
    • Claim: Claude’s Chinese is first-tier for reasoning-heavy prose but domestic models win on culturally grounded creative expression.
  8. Kimi K2.5 深度实测:变强了,但尚待「封神」https://www.geekpark.net/news/359902

    • Author: 徐珊, 极客公园 (GeekPark — established Chinese tech outlet). Date: 2026-02-04. Models: Kimi K2.5 vs Qwen3-Max, Gemini, Claude (current vintage).
    • Quote: on novel comprehension, Kimi K2.5 “从小说内容理解上,比 Qwen3-Max 要更深一步” (goes a step deeper than Qwen3-Max in understanding novel content); design/aesthetic output still “PPT感非常浓” (strong PowerPoint feel).
    • Claim: 2026 domestic frontier (Kimi) leads peers on long-fiction comprehension depth, though direct prose-generation quality wasn’t the review’s focus.
  9. Kimi K2 短篇小说创意写作夺冠https://caip.org.cn/news/detail?id=34164 (also https://www.letsclouds.com/news/kimi-k2-short-story-creative-writing)

    • Author: aggregator coverage of EQ-Bench/creative-writing benchmark results (secondary reporting; benchmark judged by LLM, flag). Date: ~2025-07. Model: Kimi K2 (mid-2025 vintage).
    • Claim: Kimi K2 topped a short-story creative-writing leaderboard over o3-Pro, praised for literary compression and metaphor, but Chinese commentary noted a “匠气” (crafted/manufactured) quality that “doesn’t fully touch readers’ hearts.”
  10. 知乎/36氪 “去AI味” discussions — e.g. https://www.zhihu.com/question/966797856, https://eu.36kr.com/zh/p/3824601267196037 (2026), https://www.53ai.com/news/neirongchuangzuo/2025071418539.html

    • Authors: anonymous practitioners (informal but native and numerous). Date: 2024-2026 rolling.
    • Representative diagnostics: “首先、其次、再次、然后、最后…一上来可以判断大概率是AI写的” (the first/second/then/finally scaffolding instantly marks AI text); AI味是风格问题 — “过于书面化、对仗工整、面面俱到” (over-bookish, over-parallel, exhaustively even-handed); AI “很少表达强烈的情感和鲜明的个人色彩” (rarely expresses strong emotion or distinct personal color).
    • Claim: an entire cottage industry of “de-AI-flavoring” prompts/tools exists because native readers reliably detect LLM Chinese by style, not grammar.

Failure modes observed

Praise / strengths noted

Evidence quality & gaps


Addendum: GLM (Zhipu / Z.ai) — supplementary sweep, 2026-07-23

Added after the main sweep, which listed GLM in the search brief but returned no sources on it. Raw output of a focused follow-up agent; unedited except for this note.

Release facts (versions + dates as of July 2026, with sources)

Summary verdict on writing quality

GLM-5.x is a coding/agent line whose Chinese prose is regarded by native users as competent mid-pack, not first-rank — behind DeepSeek and Kimi for fiction, and its own community frequently redirects writers elsewhere (“如果写中文小说,你不如试试别家,比如ds”). The line’s distinctive trait is a deliberate Claude-like flavor (“GLM有很浓郁的Claude味道”; 酒馆/RP users call it “克味”), inherited from the well-liked GLM-4.6, which was praised as “像claude的感觉,文字很流畅,细节很丰富” and less florid than DeepSeek. Native commentary describes a regression across the 5.x line for writing: GLM-5 was dismissed by some as “除了代码都挺烂” (bad at everything but code) and “文字水平不如deepseek3.2,” and GLM-5.2’s writing is judged to have “下滑了” (declined) versus GLM-5/5.1 — strong coherence and detail from its heavy reasoning, but weak 文笔 and emotional register. In the June 2026 gaokao-essay test (钛媒体/AI唱反调), 智谱清言 sat mid-pack: credited with “社会学深度” but “时代感强但个人经历弱” — conceptually sophisticated, humanly thin. This is fully consistent with the study’s broader landscape: Claude as the stylist benchmark (GLM is praised precisely to the extent it imitates Claude), DeepSeek/Kimi as the idiomatic native leaders, and GLM’s weaknesses framed in AI味 terms (emotional flatness, ornament needing prompt-engineering) rather than grammar.

Sources

  1. linux.do thread「GLM-5水平怎么样?」https://linux.do/t/topic/1805750 — anonymous power users (酒馆/writing community), linux.do forum, 2026-03-24, GLM-5. OP scope: “指的不是代码水平,是文字可读性与文笔” (not asking about code — text readability and prose style). Replies: “用了一段时间感觉国模里glm5文字水平不如deepseek3.2,仅限文字水平” (after a while, GLM-5’s text quality is below DeepSeek 3.2 among domestic models — text only); “除了代码都挺烂的吧…上一个版本4.7还好用点,5挺烂的,只能code” (bad at everything except code; 4.7 was better, 5 is only good for code); “在酒馆里用有佬反应有点克味” (RP users report a Claude-ish flavor). Claim: GLM-5’s writing regressed relative to its coding gains; below DeepSeek for prose.
  2. linux.do thread「佬们有试过glm5.2写作吗」https://linux.do/t/topic/2434692 — user wzxwhxcz, linux.do, 2026-06-19, GLM-5.2. OP (pro-GLM minority view): “glm5.2的写作由于强度思考,有一种非常强的连贯性,及细节的思考 除了文笔情感挑不出毛病” (thanks to intensive reasoning it has very strong coherence and detail-thinking — apart from prose style and emotion, nothing to fault). Reply (inndeex): GLM-5.2’s “中文水平还是比较差的,” recommending Qwen 3.7 Max for 逻辑和简明性. Claim: even fans concede 文笔/emotion are GLM-5.2’s weak axes; strengths are coherence and detail.
  3. linux.do thread「glm5.2写作能力怎么样」https://linux.do/t/topic/2454966 — users quantum41/lucienfc/alpacamo, linux.do, 2026-06-23, GLM-5.2. “能力强指的是编程,如果写中文小说,你不如试试别家,比如ds” (its strength means coding; for Chinese fiction try DeepSeek instead); “写作能力和GLM5和5.1比下滑了,你可以用ds v4 pro去写,仿文风挺强的” (writing declined vs GLM-5/5.1; use DeepSeek V4 Pro — its style imitation is strong). Claim: community consensus that GLM-5.2 writing regressed; DeepSeek V4 Pro preferred for fiction.
  4. linux.do vintage/flavor snippetshttps://linux.do/t/topic/1003033 and https://linux.do/t/topic/2180534 (via search snippets; threads themselves Cloudflare-gated), GLM-4.6 / GLM line. “glm4.6很聪明,特别是拿来写作,他不是deepseek那种花团锦簇的,像claude的感觉,文字很流畅,细节很丰富” (GLM-4.6 is smart, especially for writing — not DeepSeek’s flowery excess; feels like Claude, fluent, rich in detail); “GLM有很浓郁的Claude味道,就是色色容易截断” (GLM has a heavy Claude flavor, but NSFW gets truncated). Claim: GLM-4.6 earned a Claude-adjacent writing reputation that anchors expectations for 5.x.
  5. 钛媒体 / 网易号-syndicated「AI唱反调」— 七模型2026高考作文大测试https://news.qq.com/rain/a/20260608A03GCT00 — AI唱反调 (钛媒体APP), 2026-06-08, 智谱清言 (GLM-5.1-era product). Graded by human teachers; Kimi 1st (56.5/60), 文心一言 2nd (55.5), 豆包 last (49.5); 智谱清言 unranked-middle. “清言选’附近’的社会学深度,都拿到了高分” (Qingyan’s sociologically deep choice of “nearby” scored high); “清言时代感强但个人经历弱” (strong era-consciousness, weak personal experience). Claim: mid-pack in native human grading — intellectually strong, personally/emotionally thin. (Confirms at its 钛媒体 source the gaokao test cited in the main Chinese notes.)
  6. Dr. Jackei Wong,「DeepSeek V4 vs Kimi K3 vs GLM 5.2 深度比較」https://drjackeiwong.com/2026/07/21/deepseek-v4-kimi-k3-glm-5-2-comparison/ — HK-based tech consultant/blogger (Traditional Chinese), 2026-07-21, GLM-5.2. On writing, crowns Kimi: “K3 的另一個明顯強項是中文寫作的「語感」…不像 DeepSeek 那樣有明顯的翻譯腔或機器味,更接近一位受過訓練的中文編輯” (K3’s strength is Chinese prose “language-feel”… unlike DeepSeek’s translation-tone/machine flavor, closer to a trained Chinese editor); GLM-5.2 positioned as the enterprise/private-deployment/structured-output play, not the writing pick. Claim: in July 2026 三巨头 framing, GLM-5.2 is the enterprise model; Kimi K3 owns literary 语感.
  7. XSCT Bench (洛小山, independent, LLM-as-judge)https://xsct.ai/testcase/l_write_007_v2/glm-5 and https://xsct.ai/testcase/l_creative_035/glm-5.2 — June–July 2026, GLM-5 / GLM-5.2. GLM-5 argumentative essay: 89.47, rank 26/79; judges praised “较强的写作功底”/“典雅” language but “中心论点略显宽泛…焦点稍显分散.” GLM-5.2 散文文风迁移: 90.3 basic / 69.6 hardest tier; Kimi-judge: “高质量” mimicry but “缺少原文的生活温度” (lacks the original’s warmth of lived life). Overall composite puts GLM-5.x mid-table behind Claude and top DeepSeek/Kimi entries. Claim: benchmark-mid-pack; recurring judge note is missing human warmth, not broken mechanics.
  8. 虎嗅「GLM5.2、Kimi2.7、DeepSeek V4最佳搭配清单」https://m.huxiu.com/article/4868348.html — 丸美小沐, 虎嗅, 2026-06-18, GLM-5.2. For writing/copy work: “推荐DeepSeek V4 Pro,直接用免费的网页版即可,而且做文案非常适合” — GLM-5.2 assigned to coding, flagged for higher hallucination rate. Claim: mainstream scenario guides route writing away from GLM-5.2.

Supporting release-fact sources: 腾讯新闻 GLM-5发布, 搜狐 GLM-5开源, 智谱官方文档, 腾讯云 GLM-5.2开源. (Note: glm-5.org is an unofficial SEO/fan site; not relied upon.)

Failure modes / strengths vs DeepSeek, Kimi, Qwen, Claude

Evidence quality & gaps