LLM Writing Quality by Language — Chinese
Raw research notes for Chinese, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-chinese.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)
Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)
Summary verdict (4-8 sentences)
Native Chinese assessments split sharply by register. For logic-heavy non-fiction (reports, academic prose, analysis), reviewers consistently rate Claude’s Chinese output at the top tier — “逻辑清晰、结构完整、论证严密” with minimal 翻译腔 — while noting GPT-series Chinese still carries a “塑料感” (plastic feel) and that all Western models lack 地道感 (native groundedness): weak on internet slang, dialects, and China-domestic knowledge. For fiction, the picture inverted somewhat during 2025-2026: domestic models (DeepSeek, Kimi, Qwen, Doubao) are credited with more native-feeling colloquial Chinese and web-novel fluency, with reviewers claiming they “已达国际一线水平” (reached international first tier), while one influential 2026 web-novel guide still scored Claude highest for literary sentence craft (“最接近’人写’的感觉”). The most credentialed literary voices — Mao Dun Prize winner 麦家, web-fiction editors, 番茄小说 writers — are broadly dismissive of unassisted AI fiction: technically fluent but “寡淡无味,套路化…没有情感” (bland, formulaic, emotionless), with long-form logic and character-consistency collapse as the dominant complaint. The canonical failure modes natives name are “AI味” (首先/其次 scaffolding, piled-up 排比 parallelism, over-ornamented but empty prose, hedged fake-objectivity, formulaic endings) rather than grammatical error. Notably, the sharpest literary critiques date to the early-2025 DeepSeek-R1 moment; 2026-era commentary on frontier models comes mostly from tech media, vendor blogs, and 知乎, not literary critics — a real evidence gap.
Sources
-
麦家谈DeepSeek对文学创作的冲击 — https://finance.sina.com.cn/tech/digi/2025-03-09/doc-inenzies2017417.shtml
- Author: 麦家 (Mai Jia) — novelist, Mao Dun Literature Prize winner, Vice Chairman of the China Writers Association. Outlet: 新浪科技. Date: 2025-03-09. Model: DeepSeek-R1 (vintage: early 2025).
- Quote: “DeepSeek可能比95%的人写得好,但问题在于没法暴露人的局限性,而人的局限性恰恰是很多经典的灵感来源。” (“DeepSeek may write better than 95% of people, but it cannot expose human limitations — and human limitations are precisely the source of inspiration for many classics.”) Also: “正因为没有人类的局限,甚至缺陷,机器在文学创作上永远无法超越人。”
- Claim: AI prose is competent beyond most humans but structurally incapable of the flawed particularity that makes literature.
-
AI闯入网文赛道,网文编辑:工作量多了50% — https://www.chinawriter.com.cn/n1/2025/0226/c404027-40426506.html
- Authors quoted: working web-fiction editors and a writer/screenwriter (小白, 缦彩笺). Outlet: 封面新闻 via 中国作家网 (China Writers Association’s official site). Date: 2025-02-26. Model: DeepSeek-R1 (vintage: early 2025).
- Quotes: “直接全部丢给AI进行生成的文稿,出来的东西很缺乏逻辑性和代入感” (“Manuscripts generated wholesale by AI severely lack logic and immersion”); “寡淡无味,套路化,最重要是没有情感,没有创新,全是老梗” (“Bland, formulaic; above all no emotion, no originality — all stale tropes”).
- Claim: AI slush flooded submission queues (+50% editor workload) and editors can identify and reject it on quality grounds.
-
番茄小说的AI难题 — https://36kr.com/p/3501232140474501
- Author: 何旭, 海克财经 (business/tech analysis outlet), republished on 36氪. Date: 2025-10-09. Models: DeepSeek V3/R1, ChatGPT, 阅文妙笔, 文心一言 (vintage: mid/late 2025).
- Quotes from working web-novelists: “人物关系前后矛盾” (character relationships contradict across chapters); “喜欢乱加戏,导致人设走偏” (adds gratuitous scenes, derailing characterization); “文笔漂亮但没有’时间’的概念” (pretty prose but no sense of time); DeepSeek openings “堆满精密数字和专业名词” (stuffed with precise numbers and jargon).
- Claim: on China’s biggest free-reading platform, AI assistance boosts output volume but produces recognizable structural/temporal incoherence; one author nonetheless reported viable read-completion rates (~23%) with 1 hr/day of AI-assisted work.
-
2026 AI写小说用哪个模型?实测排名 — https://maliangwriter.com/blog/ai-model-selection-guide/
- Author: 马良写作 team (commercial Chinese AI-writing platform — vendor interest; flag). Date: 2026-03-04, updated 2026-06. Models: Claude Sonnet/Opus “4.6”, Gemini 2.5 Pro, GPT o3, DeepSeek V3/R1, Qwen3, Llama 4.
- Quotes: Claude — “句式富于变化,文学性强,最接近’人写’的感觉” (varied sentence patterns, strong literariness, closest to human-written feel; 9.2/10); DeepSeek — “中文理解深厚,偶有’教科书感’” (deep Chinese comprehension, occasional textbook feel; 8.8); GPT — “英文训练为主,中文文学性相对弱” (English-centric training, weaker Chinese literariness; 7.5); trend note: 国产模型 “已达国际一线水平”.
- Claim: for Chinese fiction, Claude leads on prose craft, DeepSeek/Qwen are close and far cheaper — the recommended workflow mixes DeepSeek for volume and Claude for quality chapters.
-
2026高考作文AI大测试:豆包可能需要复读了 — https://www.163.com/dy/article/KURO27BH0556MPIL.html
- Author: 网易号 “AI唱反调” (independent Chinese AI commentator; informal; essays were graded by ChatGPT-5.5, not human graders — flag). Date: 2026-06-07. Models: DeepSeek, Kimi, 通义千问, 豆包, 文心一言, 智谱清言, 腾讯元宝 (current 2026 domestic frontier).
- Quotes: Kimi (56.5, 1st) — “最接近真实考生” (closest to a real examinee); 文心一言 praised for sensory-detail density (“暖黄的灯光落在她发梢”); 豆包 (49.5, last) — “AI腔明显” (obvious AI voice), formulaic three-part structure and 排比 without lived detail.
- Claim: even among 2026 domestic models, the spread between “reads like a person” and “reads like AI” remains wide on the gaokao-essay register.
-
写中文论文 Claude / DeepSeek / GPT 哪个更好?2026深度实测 — https://paper.inkfount.com/blog/chinese-academic-writing-ai-comparison-2026-inkfount
- Author: InkFount (research-writing SaaS — vendor blog; flag). Date: 2026-04-26. Models: Claude 3.5/4-series, DeepSeek R1/V3, GPT-4o/5 (mixed vintage labels).
- Quotes: Claude captures “委婉但严谨” academic rhetoric with minimal 翻译腔; DeepSeek shows “比国外模型更高的’原生感’” (more native feel than foreign models) for Chinese citations/derivations; GPT sometimes has “语感’塑料感’” (plastic-feeling Chinese).
- Claim: for Chinese academic non-fiction, DeepSeek is the most natively idiomatic, Claude the most rhetorically refined, GPT the most structurally balanced but least natural-sounding.
-
Claude中文能力深度测评(2026版) — https://www.claude-anthropic.com/guide/289.html
- Author: anonymous; third-party SEO-style “Claude 中文” site, not Anthropic; low editorial credibility, but synthesizes 知乎 evaluations. Date: 2026-03-24. Models: Claude Sonnet/Opus 4.6, GPT-5.4, DeepSeek, 文心, 通义, 豆包.
- Quotes: strengths “逻辑清晰、结构完整、论证严密”; a 知乎 workplace-writing test found Claude’s Chinese weekly reports “最难被识别为AI写作” (hardest to identify as AI-written); weaknesses — 网络用语/方言 “不够接地气”, and limited “中文创意表达的’地道感’”.
- Claim: Claude’s Chinese is first-tier for reasoning-heavy prose but domestic models win on culturally grounded creative expression.
-
Kimi K2.5 深度实测:变强了,但尚待「封神」 — https://www.geekpark.net/news/359902
- Author: 徐珊, 极客公园 (GeekPark — established Chinese tech outlet). Date: 2026-02-04. Models: Kimi K2.5 vs Qwen3-Max, Gemini, Claude (current vintage).
- Quote: on novel comprehension, Kimi K2.5 “从小说内容理解上,比 Qwen3-Max 要更深一步” (goes a step deeper than Qwen3-Max in understanding novel content); design/aesthetic output still “PPT感非常浓” (strong PowerPoint feel).
- Claim: 2026 domestic frontier (Kimi) leads peers on long-fiction comprehension depth, though direct prose-generation quality wasn’t the review’s focus.
-
Kimi K2 短篇小说创意写作夺冠 — https://caip.org.cn/news/detail?id=34164 (also https://www.letsclouds.com/news/kimi-k2-short-story-creative-writing)
- Author: aggregator coverage of EQ-Bench/creative-writing benchmark results (secondary reporting; benchmark judged by LLM, flag). Date: ~2025-07. Model: Kimi K2 (mid-2025 vintage).
- Claim: Kimi K2 topped a short-story creative-writing leaderboard over o3-Pro, praised for literary compression and metaphor, but Chinese commentary noted a “匠气” (crafted/manufactured) quality that “doesn’t fully touch readers’ hearts.”
-
知乎/36氪 “去AI味” discussions — e.g. https://www.zhihu.com/question/966797856, https://eu.36kr.com/zh/p/3824601267196037 (2026), https://www.53ai.com/news/neirongchuangzuo/2025071418539.html
- Authors: anonymous practitioners (informal but native and numerous). Date: 2024-2026 rolling.
- Representative diagnostics: “首先、其次、再次、然后、最后…一上来可以判断大概率是AI写的” (the first/second/then/finally scaffolding instantly marks AI text); AI味是风格问题 — “过于书面化、对仗工整、面面俱到” (over-bookish, over-parallel, exhaustively even-handed); AI “很少表达强烈的情感和鲜明的个人色彩” (rarely expresses strong emotion or distinct personal color).
- Claim: an entire cottage industry of “de-AI-flavoring” prompts/tools exists because native readers reliably detect LLM Chinese by style, not grammar.
Failure modes observed
- “AI味” stylistic signature (the dominant native complaint, displacing 翻译腔): mechanical connectives (首先/其次/最后), piled 排比 parallelism, over-ornamented 对仗 prose that is “用力过猛” (trying too hard), exhaustive even-handedness/和稀泥 hedging, formulaic three-part structure and formulaic endings, emotional flatness.
- 翻译腔 / 塑料感 — now attributed mainly to Western models’ Chinese (GPT most often; “语感塑料感”), and reported as largely fixed in top-tier 2026 outputs for formal registers.
- Lack of 地道感: Western models weak on internet slang, memes, dialects (粤语/河南话/上海话), and China-domestic knowledge — a creative-writing handicap.
- Long-form fiction collapse: character relationships contradicting across chapters, 人设走偏 (characterization drift), no coherent sense of story time, gratuitous added scenes, inexplicable psychological asides at chapter ends.
- Register/genre missteps: DeepSeek’s “教科书感” and argumentative/expository drift inside fiction; unprompted insertion of sci-fi elements; openings stuffed with precise numbers and technical jargon; 豆包’s safe templated gaokao essays.
- Emotional and experiential hollowness: “没有情感…全是老梗”; fluency without rhythm or 代入感 (immersion); 麦家’s structural critique — no human limitation, therefore no classic.
- (成语 misuse specifically was not a prominent complaint in 2025-26 sources — the critique has moved up from lexical errors to style and structure.)
Praise / strengths noted
- Claude: repeatedly rated the most “human-feeling” literary Chinese among Western models (“最接近’人写’的感觉”, 9.2/10 in the 马良 test); near-zero 翻译腔 in academic register; best-in-class logical/argumentative Chinese (“明显优于大多数国产模型” for analysis-heavy writing); Chinese reports hardest to detect as AI.
- Domestic models at native idiom: DeepSeek “原生感” beats foreign models for Chinese academic writing; 通义千问 called the most natural, translation-artifact-free Chinese for official/business copy; Qwen3 fluent in web-fiction conventions; 阅文 reported DeepSeek-R1 “尤其擅长打斗场面和对话补充” (especially good at fight scenes and dialogue) for web novels.
- Kimi: strongest 2026 domestic showing on creative registers — won a short-story benchmark over o3-Pro, “最接近真实考生” in the gaokao test, deepest long-novel comprehension in GeekPark’s review.
- Industry adoption as tacit praise: claims (vendor-sourced, unverified) that AI touched >35% of 2025 web novels and >60% of professional web writers use AI tools; one 番茄 author sustained a commercially viable serial on ~1 hr/day.
- Consensus division of labor: domestic models for volume, colloquial and China-grounded content; Claude/frontier Western for prose craft and rigor — natives now describe the top of both camps as one tier for Chinese.
Evidence quality & gaps
- Credentialed literary judgment is mostly early-2025 vintage. The strongest voices (麦家, 莫言/刘慈欣 remarks, working editors via 中国作家网) reacted to DeepSeek-R1, not to Opus 4.5/4.6, GPT-5.x, Gemini 3, or Fable/Mythos-class models. No 2026 assessment by a named literary critic in 澎湃 or 新京报书评周刊 of a specific frontier model’s Chinese prose was found.
- 2026-era evaluations come from interested or informal parties: AI-writing SaaS blogs (马良写作, InkFount), an unofficial “Claude中文” SEO site, 知乎 posts, and tech media. Scores like “9.2/10” are unaudited; several sources use garbled or nonexistent model names (“GPT o3”, “GPT-5.5”, “Claude 4.6”, “Gemini 3.5”), a general reliability caveat for Chinese third-party eval content.
- AI-judged evaluations: the gaokao test and the EQ-Bench short-story result were scored by LLMs, so they document native commentary around results more than native human judgment of prose.
- No frontier-Anthropic/OpenAI-specific fiction criticism in Chinese literary venues was located for late-2025/2026 — Chinese critics discuss “AI” generically or DeepSeek specifically; Western frontier models are mostly assessed by tool-reviewers, not writers.
- Fiction vs non-fiction evidence is asymmetric: non-fiction assessments (academic, workplace) are more systematic and more favorable; fiction assessments are richer in failure detail but anchored in the web-novel industry, with essentially no coverage of literary fiction quality from 纯文学 critics against 2026 models.
- Search-side gaps: Zhihu article pages 403 on direct fetch (quotes came via search snippets/secondary syntheses); WeChat public-account critic essays are effectively unsearchable from outside and are likely where the best native critic commentary lives.
Addendum: GLM (Zhipu / Z.ai) — supplementary sweep, 2026-07-23
Added after the main sweep, which listed GLM in the search brief but returned no sources on it. Raw output of a focused follow-up agent; unedited except for this note.
Release facts (versions + dates as of July 2026, with sources)
- GLM-5 — released 2026-02-11 (腾讯新闻), open-weights release 2026-02-12 (搜狐/AtomGit首发). ~744–745B total params MoE, ~40–44B active. Official positioning: “面向 Agentic Engineering 打造,能够在复杂系统工程与长程 Agent 任务中提供可靠生产力” — built for agentic engineering / long-horizon agent tasks (智谱开放文档). Coding/agent-first, not writing-focused.
- GLM-5.1 — released 2026-04-08 (dev early access 2026-03-27); pitched as “first open-source model for long-program tasks,” rolled out to Coding Plan tiers without a launch event.
- GLM-5.2 — current flagship as of July 2026, released/open-sourced 2026-06-13, MIT license, 744B MoE, 1M-token context; positioning again long-task/coding (“长任务时代的旗舰模型”) (腾讯云开发者社区, 2026-06-22; tahou.com claims it ranks “third globally after GPT-5.5 and Claude Opus 4.8” on aggregate benchmarks). Landed in the same June window as Kimi K2.7-Code and DeepSeek V4 Pro.
- Vintage context: GLM-4.5/4.6 (2025) were the coding-open-weights breakout; notably, GLM-4.6 had a genuine reputation as a writing model in Chinese communities (see below), which GLM-5.x has partially lost.
- Consumer chat product remains 智谱清言 (Zhipu Qingyan); z.ai is the international brand.
Summary verdict on writing quality
GLM-5.x is a coding/agent line whose Chinese prose is regarded by native users as competent mid-pack, not first-rank — behind DeepSeek and Kimi for fiction, and its own community frequently redirects writers elsewhere (“如果写中文小说,你不如试试别家,比如ds”). The line’s distinctive trait is a deliberate Claude-like flavor (“GLM有很浓郁的Claude味道”; 酒馆/RP users call it “克味”), inherited from the well-liked GLM-4.6, which was praised as “像claude的感觉,文字很流畅,细节很丰富” and less florid than DeepSeek. Native commentary describes a regression across the 5.x line for writing: GLM-5 was dismissed by some as “除了代码都挺烂” (bad at everything but code) and “文字水平不如deepseek3.2,” and GLM-5.2’s writing is judged to have “下滑了” (declined) versus GLM-5/5.1 — strong coherence and detail from its heavy reasoning, but weak 文笔 and emotional register. In the June 2026 gaokao-essay test (钛媒体/AI唱反调), 智谱清言 sat mid-pack: credited with “社会学深度” but “时代感强但个人经历弱” — conceptually sophisticated, humanly thin. This is fully consistent with the study’s broader landscape: Claude as the stylist benchmark (GLM is praised precisely to the extent it imitates Claude), DeepSeek/Kimi as the idiomatic native leaders, and GLM’s weaknesses framed in AI味 terms (emotional flatness, ornament needing prompt-engineering) rather than grammar.
Sources
- linux.do thread「GLM-5水平怎么样?」 — https://linux.do/t/topic/1805750 — anonymous power users (酒馆/writing community), linux.do forum, 2026-03-24, GLM-5. OP scope: “指的不是代码水平,是文字可读性与文笔” (not asking about code — text readability and prose style). Replies: “用了一段时间感觉国模里glm5文字水平不如deepseek3.2,仅限文字水平” (after a while, GLM-5’s text quality is below DeepSeek 3.2 among domestic models — text only); “除了代码都挺烂的吧…上一个版本4.7还好用点,5挺烂的,只能code” (bad at everything except code; 4.7 was better, 5 is only good for code); “在酒馆里用有佬反应有点克味” (RP users report a Claude-ish flavor). Claim: GLM-5’s writing regressed relative to its coding gains; below DeepSeek for prose.
- linux.do thread「佬们有试过glm5.2写作吗」 — https://linux.do/t/topic/2434692 — user wzxwhxcz, linux.do, 2026-06-19, GLM-5.2. OP (pro-GLM minority view): “glm5.2的写作由于强度思考,有一种非常强的连贯性,及细节的思考 除了文笔情感挑不出毛病” (thanks to intensive reasoning it has very strong coherence and detail-thinking — apart from prose style and emotion, nothing to fault). Reply (inndeex): GLM-5.2’s “中文水平还是比较差的,” recommending Qwen 3.7 Max for 逻辑和简明性. Claim: even fans concede 文笔/emotion are GLM-5.2’s weak axes; strengths are coherence and detail.
- linux.do thread「glm5.2写作能力怎么样」 — https://linux.do/t/topic/2454966 — users quantum41/lucienfc/alpacamo, linux.do, 2026-06-23, GLM-5.2. “能力强指的是编程,如果写中文小说,你不如试试别家,比如ds” (its strength means coding; for Chinese fiction try DeepSeek instead); “写作能力和GLM5和5.1比下滑了,你可以用ds v4 pro去写,仿文风挺强的” (writing declined vs GLM-5/5.1; use DeepSeek V4 Pro — its style imitation is strong). Claim: community consensus that GLM-5.2 writing regressed; DeepSeek V4 Pro preferred for fiction.
- linux.do vintage/flavor snippets — https://linux.do/t/topic/1003033 and https://linux.do/t/topic/2180534 (via search snippets; threads themselves Cloudflare-gated), GLM-4.6 / GLM line. “glm4.6很聪明,特别是拿来写作,他不是deepseek那种花团锦簇的,像claude的感觉,文字很流畅,细节很丰富” (GLM-4.6 is smart, especially for writing — not DeepSeek’s flowery excess; feels like Claude, fluent, rich in detail); “GLM有很浓郁的Claude味道,就是色色容易截断” (GLM has a heavy Claude flavor, but NSFW gets truncated). Claim: GLM-4.6 earned a Claude-adjacent writing reputation that anchors expectations for 5.x.
- 钛媒体 / 网易号-syndicated「AI唱反调」— 七模型2026高考作文大测试 — https://news.qq.com/rain/a/20260608A03GCT00 — AI唱反调 (钛媒体APP), 2026-06-08, 智谱清言 (GLM-5.1-era product). Graded by human teachers; Kimi 1st (56.5/60), 文心一言 2nd (55.5), 豆包 last (49.5); 智谱清言 unranked-middle. “清言选’附近’的社会学深度,都拿到了高分” (Qingyan’s sociologically deep choice of “nearby” scored high); “清言时代感强但个人经历弱” (strong era-consciousness, weak personal experience). Claim: mid-pack in native human grading — intellectually strong, personally/emotionally thin. (Confirms at its 钛媒体 source the gaokao test cited in the main Chinese notes.)
- Dr. Jackei Wong,「DeepSeek V4 vs Kimi K3 vs GLM 5.2 深度比較」 — https://drjackeiwong.com/2026/07/21/deepseek-v4-kimi-k3-glm-5-2-comparison/ — HK-based tech consultant/blogger (Traditional Chinese), 2026-07-21, GLM-5.2. On writing, crowns Kimi: “K3 的另一個明顯強項是中文寫作的「語感」…不像 DeepSeek 那樣有明顯的翻譯腔或機器味,更接近一位受過訓練的中文編輯” (K3’s strength is Chinese prose “language-feel”… unlike DeepSeek’s translation-tone/machine flavor, closer to a trained Chinese editor); GLM-5.2 positioned as the enterprise/private-deployment/structured-output play, not the writing pick. Claim: in July 2026 三巨头 framing, GLM-5.2 is the enterprise model; Kimi K3 owns literary 语感.
- XSCT Bench (洛小山, independent, LLM-as-judge) — https://xsct.ai/testcase/l_write_007_v2/glm-5 and https://xsct.ai/testcase/l_creative_035/glm-5.2 — June–July 2026, GLM-5 / GLM-5.2. GLM-5 argumentative essay: 89.47, rank 26/79; judges praised “较强的写作功底”/“典雅” language but “中心论点略显宽泛…焦点稍显分散.” GLM-5.2 散文文风迁移: 90.3 basic / 69.6 hardest tier; Kimi-judge: “高质量” mimicry but “缺少原文的生活温度” (lacks the original’s warmth of lived life). Overall composite puts GLM-5.x mid-table behind Claude and top DeepSeek/Kimi entries. Claim: benchmark-mid-pack; recurring judge note is missing human warmth, not broken mechanics.
- 虎嗅「GLM5.2、Kimi2.7、DeepSeek V4最佳搭配清单」 — https://m.huxiu.com/article/4868348.html — 丸美小沐, 虎嗅, 2026-06-18, GLM-5.2. For writing/copy work: “推荐DeepSeek V4 Pro,直接用免费的网页版即可,而且做文案非常适合” — GLM-5.2 assigned to coding, flagged for higher hallucination rate. Claim: mainstream scenario guides route writing away from GLM-5.2.
Supporting release-fact sources: 腾讯新闻 GLM-5发布, 搜狐 GLM-5开源, 智谱官方文档, 腾讯云 GLM-5.2开源. (Note: glm-5.org is an unofficial SEO/fan site; not relied upon.)
Failure modes / strengths vs DeepSeek, Kimi, Qwen, Claude
- vs Claude: GLM’s writing identity is Claude-imitation — “很浓郁的Claude味道,” “像claude的感觉” — praised when the imitation lands (4.6, partly 5.2’s coherence), but no one claims it matches Claude’s 质感; it is the copy, Claude the referent. Confirms the study’s Claude-as-stylist-benchmark finding from the reverse direction.
- vs DeepSeek: split verdicts. GLM-4.6 was preferred by some for avoiding DeepSeek’s “花团锦簇” over-ornamentation; by GLM-5.2’s era the flow reversed — DeepSeek V4 Pro is the community’s default fiction/style-imitation recommendation over GLM. One tieba snippet (thread inaccessible) had GLM-5.1 beating DeepSeek on descriptive quality but losing coherence over long outputs.
- vs Kimi: Kimi K3 is credited with the best Chinese 语感 (“受过训练的中文编辑”); Kimi also won the human-graded gaokao test where 智谱清言 was mid-pack. No source puts GLM above Kimi for prose.
- vs Qwen: less discussed; one GLM-5.2 thread recommends Qwen 3.7 Max for 逻辑和简明性 over GLM’s Chinese.
- GLM-specific failure modes: (1) writing regression across 5.x as training budget went to code/agents — “只能code”; (2) emotional flatness / missing “生活温度” (the human-warmth deficit, an AI味-family complaint); (3) aggressive content filtering — NSFW/romance “容易截断” on official channels; (4) florid tics requiring prompt engineering (tieba snippet); (5) over-long thinking time (“思考过久成最大槽点”) and higher hallucination rate.
- GLM-specific strengths: long-output coherence and detail-tracking from heavy reasoning + 1M context; strong structured/formal writing (议论文 89.47); style-transfer competence at basic difficulty.
Evidence quality & gaps
- Quality: The linux.do quotes are verbatim native-speaker text — high credibility, hobbyist-writer demographic (酒馆/网文 users, exactly the population that detects AI味). The gaokao test is human-teacher-graded (strong), but gives 智谱清言 no explicit numeric score — only qualitative mid-pack placement. XSCT Bench is LLM-as-judge, one independent operator — useful for relative placement, not a native-human assessment. Some page reads were mediated by a summarizer; quotes marked verbatim came through in Chinese and are internally consistent across independent fetches, but line-level fidelity of paraphrased material is medium.
- Gaps: (1) Baidu Tieba threads and a Zhihu question on GLM-5.2 writing are hard-blocked — only snippet-level content recovered. (2) No professional author/editor (as opposed to hobbyist) assessment of GLM-5.x prose was found. (3) The exact 智谱清言 gaokao score is unpublished in the accessible version of the piece. (4) Search discovery for this sweep skewed toward DuckDuckGo-indexed pages; WeChat 公众号 coverage is under-sampled. (5) One CSDN long-novel stress test repeatedly returned 521 and could not be verified.