LLM Writing Quality by Language — Japanese

Raw research notes for Japanese, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-japanese.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)


Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)


Summary verdict (4-8 sentences)

Native Japanese assessments have shifted markedly between 2024 and 2026: the dominant complaint has moved from “翻訳調” (translationese) to “indistinguishability.” The single strongest data point is the 13th Nikkei Hoshi Shinichi Award (Feb 2026), where 3 of 4 general-category winners used generative AI and judges repeatedly said they could not tell human from AI prose — prompting nonfiction writer 最相葉月 to resign-in-protest terms (“AIの執筆した文章は、もう読みたくない”). Practitioner reviewers consistently rank the Claude Opus line as the best Japanese stylist among frontier models — tech journalist 新清士 says Opus 4.6’s 150,000-character novel “matches the highest quality humans can write” — while GPT-5.x is judged natural and readable but emotionally flat and forgettable, and Gemini is praised for rhythm and sentence-ending nuance. Interestingly, native reviewers track regressions between versions: Opus 4.8 is judged below the writers’ favorite Opus 4.6, and GPT-5.5 no better than 5.4, indicating evaluation has become fine-grained enough to detect intra-vendor style drift. Remaining criticisms concentrate on rhythm-level tells (monotone 文末 endings, ruler-even paragraph cadence, connective overuse), absence of lived specificity, and over-polite register — not grammatical error. Literary-establishment voices (九段理江, 最相葉月) concede competence while denying AI produces ideas or dignity-bearing language beyond human level. The web-novel industry response (Narou mandatory AI disclosure June 2026, full-AI-generation ban; Kakuyomu AI tags after a 2025 contest-winner scandal) confirms output is now good enough to pass slush-pile screening at scale. Caveat: the most granular model-vs-model assessments come from expert but informal note.com practitioners, not refereed criticism.

Sources

  1. 日本経済新聞 — 日経「星新一賞」受賞作を味わう 受賞3作がAI使用

    • URL: https://www.nikkei.com/article/DGXZQOLM072920X00C26A2000000/ (also https://www.nikkei.com/article/DGXZQOGB097AH0Z00C26A2000000/)
    • Author/outlet: Nikkei staff; major national newspaper (award sponsor). Feb 2026.
    • Models: unspecified frontier models used by entrants (2025–2026 vintage).
    • Key quote: 「人間の作品かAIの作品か、区別がつかない」— “We cannot tell whether a work is by a human or by AI” (final-round judges, reported repeatedly).
    • Claim: At a major national SF/short-fiction award, professional judges found AI-assisted Japanese fiction indistinguishable from human work; 3 of 4 winners used AI.
  2. ピンズバNEWS via Yahoo!ニュース — 『星新一賞』でAI小説が上位独占も…審査員「AIの文章は読みたくない」

    • URL: https://news.yahoo.co.jp/articles/56d0c7a277c6ceb9799da2f9ff0f750d50afb430
    • Author quoted: 最相葉月 (Saishō Hazuki), acclaimed nonfiction writer (『絶対音感』), award judge. 2026.
    • Key quote: 「たとえAI小説がどれほど面白かったとしても、私は人間の内から生まれた言葉こそが尊いと思う。人間の尊厳を守りたい。AIの執筆した文章は、もう読みたくない」— “However interesting an AI novel may be, I believe words born from within a human being are what is precious. I want to protect human dignity. I no longer want to read AI-written text.”
    • Claim: A top-tier professional reader concedes AI fiction can be “interesting” (quality no longer the objection) and objects on ethical/dignity grounds instead.
  3. ASCII.jp — 新清士「AIが15万字の小説を1週間で執筆──『Claude Opus 4.6』が示した創作の未来」

    • URL: https://ascii.jp/elem/000/004/381/4381354/
    • Author: 新清士 (Shin Kiyoshi), veteran game/AI tech journalist. 2026-03-16. Models: Claude Opus 4.6 vs Gemini 3.0 (both current).
    • Key quotes: 「人間が書ける最高品質に匹敵しているとさえ感じます」— “I even feel it rivals the highest quality a human can write.” On Opus 4.6’s continuation chapter: 「豊かで自然な表現に成功しており」— “it succeeds with rich and natural expression” (surpassing Gemini).
    • Claim: A professional journalist judges Opus 4.6’s long-form Japanese fiction (150k characters, consistent across chapters) as near-top-human quality.
  4. note.com — IT navi「Claude Opus 4.8の文章執筆性能は以前のモデルより向上したのか?」 (flag: expert-informal — anonymous but widely-read AI evaluator who runs controlled same-prompt comparisons)

    • URL: https://note.com/it_navi/n/n4bf0e8158f83
    • Date: 2026-06-01. Models: Opus 4.8 / 4.7 / 4.6, GPT-5.5 Thinking (all current).
    • Key quotes: 「Opus 4.8の文章執筆性能は、前バージョンのOpus 4.7から大きく改善しました。しかし、物書きの間で評価の高いOpus 4.6の水準には、まだ追い付いていない」— “Opus 4.8’s writing greatly improved over 4.7, but still hasn’t caught up to Opus 4.6, which writers rate highly.” 「文学的な文章の執筆については、感情表現があまり得意ではないGPT-5.5 Thinkingよりも優れている」— “For literary prose it is better than GPT-5.5 Thinking, which is not very good at emotional expression.”
    • Claim: Japanese writers track per-version literary quality closely; Opus 4.6 is the community favorite, and Claude beats GPT-5.5 on literary emotion.
  5. note.com — IT navi「GPT-5.5の日本語文章生成能力は向上したのか?」 (same informal-expert flag)

    • URL: https://note.com/it_navi/n/n93304be6bea8
    • Date: 2026-04-24. Models: GPT-5.5, GPT-5.4, Claude Opus 4.7.
    • Key quotes: 「GPT-5.5の日本語文章生成能力は、前バージョンのGPT-5.4より向上しているとは感じられませんでした」— “I could not feel that GPT-5.5’s Japanese writing improved over GPT-5.4.” Praise: 「自然で読みやすい文章になっています」— “the prose is natural and easy to read” — but it lacks narrative depth and memorable lines; for expository non-fiction the naturalness is a real advantage.
    • Claim: GPT-5.x Japanese is fluent and well-suited to non-fiction, but plateauing and comparatively weak in fiction depth.
  6. nippon.com — 九段理江が語る “共作者”AIとの距離と「言葉」「リズム」へのこだわり

    • URL: https://www.nippon.com/ja/japan-topics/e00231/
    • Author: 九段理江 (Rie Kudan), 170th Akutagawa Prize winner (『東京都同情塔』, famously “5% AI”). Interview published 2025-07-18. Models: ChatGPT (2023–24 vintage — flag).
    • Key quote: 「AIから人間の知性をはるかに超えるアイデアは出てこなかった」— “No ideas ever came out of the AI that far surpassed human intelligence.” She also clarifies the “5%” was minimal, off-the-cuff, and misunderstood; her craft priority is linguistic rhythm.
    • Claim: The most prominent AI-adjacent literary author judges AI competent as an interlocutor but not a source of above-human creative language. (See also 東京新聞: https://www.tokyo-np.co.jp/article/310036, and her later “95% AI” experiment 『影の雨』 with co-author credit “CraiQ”: https://ledge.ai/articles/ai_95_percent_akutagawa_writer_5_percent_kagenoame)
  7. Real Sound Tech — 「AIが書いた小説」は作品と呼べるか?

    • URL: https://realsound.jp/tech/2026/07/post-2447680.html
    • Culture-media outlet; 2026-07-02. Covers Hoshi Shinichi Award 2026 and Kudan’s usage; references 樋口恭介『AI先生のSF小説教室』(2025) and its “vibe writing” method (translating atmospheres like 「ノスタルジック・フューチャリズム」 into prompts).
    • Claim: Mainstream criticism has moved past “is the Japanese good enough?” to authorship and judging-burden questions — judges cannot verify AI use in submissions.
  8. note.com — Key君(シナリオライター)「Gemini 2.5 Pro、小説本文執筆に最強かも──AIに”自然な日本語”は書けるのか?」 (flag: informal platform, professional scenario writer; earlier-vintage models — April 2025)

    • URL: https://note.com/ktworks/n/n5549a3d8beae
    • Models: Gemini 2.5 Pro vs GPT-4o. Praises Gemini’s 「自然さ、文末のニュアンス、間の取り方」— “naturalness, sentence-ending nuance, control of pauses/ma” — citing authentically hedged workplace dialogue 「まあ、なんだ。そういうことになったんだよ」; criticizes GPT-4o’s overwrought metaphors (e.g. 「薄桃色の空を穿ち」 as too purple for narrative prose).
    • Claim: For a working scenario writer, the differentiator in Japanese fiction is rhythm and sentence-ending nuance, where Gemini led GPT-4o in early 2025.
  9. Togetter — 「最近AIに書かせたなって文章が分かるようになってきた…」

    • URL: https://togetter.com/li/2658815
    • Crowd-sourced thread of Japanese readers/writers, 2026-02-02 (flag: informal but large-N native intuition).
    • Key quote: 「改行が微妙に不自然・急に箇条書きが始まる・体験談が無い」— “line breaks subtly unnatural, bullet lists start abruptly, no personal anecdotes.” Other cited tells: 「〜が重要です」 boilerplate, excessive politeness, zero typos/too-perfect structure, bold-marker habits.
    • Claim: Native readers’ AI-detection heuristics for non-fiction are now structural/rhythmic, not grammatical — and people parody the style deliberately.
  10. Web-novel industry policy — 小説家になろう / カクヨム AI rules

  1. Domestic-model comparison — Swallow LLM eval project / 国産LLM roundups (flag: benchmark/blog quality)

Failure modes observed

Praise / strengths noted

Evidence quality & gaps