LLM Writing Quality by Language — Korean
Raw research notes for Korean, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-korean.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)
Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)
Summary verdict
Native Korean assessment of frontier-LLM Korean prose is split by genre and model. For non-fiction/business prose, the 2026 Korean consensus (tech bloggers, practitioner comparisons, Threads/Clien power users) is that Claude produces the most natural Korean — repeatedly described as free of 번역체 (translationese) — while GPT-5’s launch was met with sustained native criticism that its Korean regressed versus GPT-4-era models: stiff 직역체, awkward collocations, even garbled transliterations of proper nouns. For fiction, working novelists who have tested the models (김연수, in a June 2026 Kyunghyang piece) find AI useful for editing and ideation but of little help in actual draft prose, describing the workflow as “repetitive labor”; the literary establishment (김보영, 전성태) largely treats AI output as not-authorship, and major Korean literary contests now ban AI submissions. The webnovel market supplies the harshest reader-level verdict: works where AI 특유의 문체 (AI-typical style) is detected get 별점 테러 (rating bombing), and authors now peer-review manuscripts to scrub AI-sounding expressions — evidence that AI Korean prose remains detectable and stigmatized at scale. Well-documented failure modes are stable across sources: hedging constructions (~라고 할 수 있습니다), calqued English idioms (핵심을 찔렀어), double passives (되어지다), particle misplacement, over-commaing, and formulaic enumeration. A counterpoint: a widely covered US academic study (including 한강’s style, fine-tuned models) found readers preferred fine-tuned AI imitations — but that finding is about style-mimicry under fine-tuning, not off-the-shelf frontier Korean output.
Sources
-
URL: https://www.khan.co.kr/article/202606180600051/ Author: 고희진 (Go Hee-jin), culture reporter, quoting novelist 김연수 (Kim Yeon-su) — one of Korea’s most acclaimed literary novelists. Outlet: 경향신문 (Kyunghyang Shinmun). Date: 2026-06-18. Models: ChatGPT (paid, since 2024), Claude Sonnet 4.6, Gemini 3 (current-vintage). Quotes: “프롬프트를 작성하고 답변을 읽는데 드는 수고는 창의적이라기보다는 반복된 노동에 가까웠다” (“The effort of writing prompts and reading responses felt closer to repetitive labor than creativity.”); “챗지피티와의 대화는 머릿속의 아이디어를 문자화할 수 있는 좋은 방법” (“Conversation with ChatGPT is a good way to turn ideas in my head into text.”); “인간의 경험과 AI의 합리가 결합할 때 새로운 인간선언이 가능” (“A new declaration of humanity becomes possible when human experience combines with AI’s rationality.”) Claim: A top-tier Korean novelist finds frontier models genuinely useful for editing/ideation but of marginal value for producing literary Korean prose itself.
-
URL: https://www.khan.co.kr/article/202602030700001 Author: 조해람 (Jo Hae-ram), reporter, quoting SF author 김보영 and novelist 전성태, professor 노대원 (Jeju Nat’l Univ., digital-humanities/AI-literature scholar), novelist 장강명. Outlet: 경향신문. Date: 2026-02-03. Models: general frontier (Grok/xAI editorial hiring cited); current-vintage. Quotes: 김보영: “누군가가 나에게 상과 저작권료를 준다는 것은 그것이 나의 온전한 창작물이기 때문” (“Someone gives me awards and royalties because it is entirely my own creation.”). Article documents AI-mass-produced e-books (Lumineri, ~9,000 titles in 2025) containing anachronisms like internet slang “머선129” inside an Odyssey edition. Claim: The Korean literary establishment sees current AI fiction output as low-quality mass production; major Korean literary contests have added AI-exclusion clauses, while some authors (장강명) note edited AI text is becoming undetectable.
-
URL: https://www.mt.co.kr/tech/2026/03/11/2026022508052980108 Author: 김평화 (Kim Pyeong-hwa), tech reporter, quoting webnovel industry insiders and readers. Outlet: 머니투데이 (Money Today). Date: 2026-03-11. Models: unnamed (current frontier, 2026). Quotes: “웹소설은 텍스트라서 파악할 방법이 거의 없다” (“Webnovels are text, so there’s almost no way to detect [AI use].”); readers feel they “창의력이 아닌 데이터 조합에 돈을 냈다” (“paid for data combination, not creativity”) — detected AI works get “별점 테러” (rating bombing). Claim: In Korea’s commercial webnovel market, AI-written Korean prose is stigmatized and, when its telltale style leaks through, punished by readers — implying quality/style remains distinguishable in practice.
-
URL: https://gist.github.com/woonjangahn/3ad4d8fe1804aed2e7cafc9493ec566f Author: woonjangahn (Korean developer/writer; GitHub gist “한국어 AI 글쓰기에서 피해야 할 상투적 패턴”, cites the KatFishNet Korean AI-text-detection paper). Outlet: GitHub Gist. Date: 2026-03-09. Models: frontier LLMs generally; GPT-4o referenced (mixed vintage). Quotes: Flags hedging “이것은 매우 효과적인 접근법이라고 할 수 있습니다” (evasive “it can be said that…” instead of direct assertion); calques like “핵심을 찔렀어” (from “You hit the nail on the head”) and “깊은 질문이야” (from “That’s deep”); “되어지다” double passives; “~에 의해 만들어진” (by-passive translationese); AI text uses commas in 61% of sentences vs 26% for human Korean (KatFishNet). Claim: A native-speaker catalog of concrete, recurring stylistic tells of AI-generated Korean — mostly English-interference patterns.
-
URL: https://namu.wiki/w/GPT-5 (fetch blocked; content via search excerpt) and https://www.etoday.co.kr/news/view/2497424 Author: namu.wiki community (crowd-edited, native Korean users); 이투데이 article by 정상원. Outlets: 나무위키 / 이투데이. Dates: 2025-08 onward (GPT-5 launch vintage — pre-GPT-5.x fixes; flag: early-vintage for the 5-series). Models: GPT-5. Quotes: namu.wiki criticism section: “작문과 대화의 자연스러움이 이전 GPT-4 계열 모델보다 크게 떨어졌으며, 직역체와 딱딱한 말투, 어색한 표현이 잦고 대화 흐름도 매끄럽지 않다” (“Naturalness of writing and conversation dropped sharply vs the GPT-4 line; literal-translation style, stiff tone, awkward expressions are frequent and conversation flow isn’t smooth.”). 이투데이 documents garbled Korean renderings: “Tennessee” → “토네시주”, “George Washington” → “기어지 워싱지언”. Claim: Korean users judged GPT-5’s Korean output at launch a regression — translationese, stiffness, and even corrupted proper-noun transliteration.
-
URL: https://my-blog.org/ai-edu/post/ai-korean-language-performance-comparison Author: anonymous Korean AI-education blogger (practitioner-level, not credentialed critic). Outlet: AI 교육 블로그. Date: updated 2026-03-11. Models: GPT-4o, Claude and Gemini (versions unspecified; mixed vintage). Quotes: “Claude에서 한국어로 보고서를 써달라고 했더니 번역체 느낌 없이 자연스러운 문장이 나왔다” (“When I asked Claude to write a report in Korean, natural sentences came out with no translationese feel.”); “AI를 쓸 때 영어를 억지로 쓸 필요는 없습니다” (“You don’t need to force yourself to use English with AI.”). Claim: For non-fiction Korean, all majors are now serviceable and Claude is the naturalness leader.
-
URL: https://www.threads.com/@yoonkwon_ai/post/DYik0Cak60v/ (and https://www.threads.com/@btobooks/post/DXYkt1IEXM1/) Author: Korean AI-practitioner accounts (native speakers; social-media grade). Outlet: Threads. Date: 2026. Models: Claude (current), ChatGPT, Gemini. Quotes: “Claude로 글 써본 사람은 다시 ChatGPT 못 돌아감… 문장이 사람이 쓴 것처럼 자연스럽게 나옴. AI 티 안 나는 글” (“Anyone who’s written with Claude can’t go back to ChatGPT… sentences come out natural, like a person wrote them — writing that doesn’t smell like AI.”); on Gemini: “제미나이는 아첨이 너무 심하다” (“Gemini is far too sycophantic”). Claim: Korean power-user folk consensus ranks Claude first for Korean prose naturalness, with an explicit final step of human polishing still assumed.
-
URL: https://aimatters.co.kr/news-report/ai-report/32847/ Author: AI매터스 staff, covering a Stony Brook/Columbia/Michigan study. Outlet: AI 매터스 (Korean AI news). Date: 2025-10-20. Models: ChatGPT, Claude, Gemini (late-2025 vintage), fine-tuned on 50 authors incl. 한강. Quotes: “파인튜닝이 AI 특유의 문체적 흔적을 단순히 가린 것이 아니라 근본적으로 제거했다” (“Fine-tuning didn’t merely mask AI’s characteristic stylistic traces — it fundamentally removed them.”); 159 evaluators (28 experts) preferred fine-tuned AI style-imitations ~8:1. Claim: Korean media relayed evidence that fine-tuned frontier models can beat human prose in blind stylistic preference — a caveat against assuming permanent AI-style detectability (note: study prose likely English; Korean relevance is via 한강 and Korean-media framing).
-
URL: https://www.etnews.com/20260706000183 (also https://v.daum.net/v/20260706170226818) Author: 전자신문 tech-China correspondent. Outlet: 전자신문 (ET News). Date: 2026-07-06. Models: unnamed current frontier/Chinese models (comparison market: China webnovels, covered for Korean industry readers). Quotes/points: Fanqie Novel rejected 100,000+ low-quality works in June 2026; authors now cross-check manuscripts for “AI 특유의 표현” (“AI-characteristic expressions”); detection is confounded because “인간의 글조차 AI 스타일로 오인될 가능성” (“even human writing risks being mistaken for AI style”). Claim: At webnovel-industry scale, AI fiction is overwhelmingly low-quality slop, and the “AI-sounding style” category is now policed — including false positives against humans.
Failure modes observed
- 번역투 / calqued English idioms: “핵심을 찔렀어”, “깊은 질문이야”, “~에 의해 만들어진” by-passives, “되어지다” double passives (gist; KatFishNet).
- Hedging/indirection: wrapping assertions in “~라고 할 수 있습니다”, inflating sentences and dodging commitment (gist).
- Stiff 직역체 and awkward collocations, broken conversational rhythm — GPT-5 specifically, judged worse than GPT-4-era (namu.wiki community).
- Proper-noun/transliteration corruption in GPT-5 Korean output at launch (“토네시주”, “기어지 워싱지언”) (이투데이).
- Mechanical texture: over-commaing (61% vs 26% human), formulaic 첫째/둘째/셋째 enumeration, repeated closing endings, overused vocabulary (다양한, 혁신적인, 핵심) (gist).
- Fiction-specific: short-sentence repetition, stock connectives, excessive psychological narration; miscounting concrete details; anachronistic register (internet slang in classical text) (전자신문, 경향 2026-02).
- Register/tone: ChatGPT prone to “이상한 번역체” and rigid formal phrasing in Korean; Gemini criticized for sycophancy and context loss rather than sentence-level errors (Clien comment, Threads, my-blog).
Praise / strengths noted
- Claude repeatedly singled out by native users as producing Korean “번역체 느낌 없이” (without translationese), best tone-keeping, best knowledge of Korean 문체 varieties and standard usage; described as least “AI-smelling” (Threads, my-blog, nerdlog/sireal-type 2026 comparisons; Anthropic’s own multilingual evals cited Opus 4.5 top across tested languages incl. Korean).
- Editing/correction and ideation rated genuinely strong even by a skeptical literary novelist (김연수: good for “머릿속의 아이디어를 문자화”).
- Fine-tuned style mimicry (incl. 한강) preferred by readers 8:1 and undetectable 97% of the time in the US study relayed by Korean media.
- Domestic models (HyperCLOVA X, SKT A.X, EXAONE) still credited with superior Korean cultural/local knowledge (KMMLU-Korean subsets), though the writing-quality crown in 2026 rankings goes to Claude, not domestic models.
Evidence quality & gaps
- Strongest: two 경향신문 features (Feb & Jun 2026) with named, credentialed literary figures (김연수, 김보영, 전성태, 장강명, 노대원) — genuine native literary assessment, current-vintage models named in the 김연수 piece (Claude Sonnet 4.6, Gemini 3).
- Medium: journalism on webnovel-market reception (머니투데이, 전자신문) — credible industry signal, but models unnamed; the woonjangahn gist — concrete linguistic detail from a native practitioner, citing an academic detection paper, but a personal document.
- Weak/flagged: Threads/blog comparisons are anecdotal power-user opinion with fuzzy version attribution; namu.wiki is crowd-edited (though it is a genuine aggregate of native-user sentiment); the 2026 “TOP 5” ranking blog mixes stale model versions (Claude 3.5/3.7, GPT-4o) despite its date; the aimatters-relayed study concerns fine-tuned models and likely English prose.
- Gaps: No located review by a professional 문학평론가 in 문학동네/창비 that sentence-level-critiques a named frontier model’s Korean fiction since late 2025; no rigorous head-to-head of Opus 4.5/4.6, GPT-5.x, Gemini 3 on Korean creative prose by native judges; honorific/register-error evidence is mostly folk-level, not systematically documented; arca.live’s webnovel-focused LLM evaluation and wikidocs’ “LLM 한국어 사용성 순위” were 403-blocked to fetching and could not be verified; web-search budget was exhausted before a “literary-contest judges on AI submissions” (심사평) search could run.
Sources: 경향신문 김연수 실험, 경향신문 점선면, 머니투데이 웹소설 별점테러, woonjangahn gist, 나무위키 GPT-5, 이투데이 GPT-5 오류, AI 교육 블로그 비교, Threads @yoonkwon_ai, Threads @btobooks, AI매터스 연구 보도, 전자신문 테크차이나, 클리앙 4대천왕 비교, allaboutlog TOP5