How Well Do Frontier LLMs Write in Languages Other Than English? What Native-Speaker Critics Say

~4,071 words. Edition 2026-07-23. The main report of the LLM Writing Quality by Language project.
Written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)


Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)


Method: Native-language web research across 15 languages, synthesized from published assessments by native-speaking critics, authors, literary translators, editors, and academics. Raw per-language research notes, with original-language quotations and English translations, are in the per-language research files linked throughout. Scope: Frontier models released since late 2025 — Claude Opus 4.5/4.6/4.8, Claude Fable 5 / Mythos 5, GPT-5.x, Gemini 3.x — plus regional/domestic models (DeepSeek, Qwen, Kimi, GLM/Zhipu, Mistral, YandexGPT, HyperCLOVA X, Sarvam, Maritaca/Sabiá, Jais/ALLaM, Bielik, Kumru, SEA-LION, Amália, ALIA) as comparison points. Earlier-model assessments are included only where flagged by vintage.

Disclosure: This research was performed by an AI agent (Claude Fable 5) directing web searches in each target language. An Anthropic model is thus both the researcher and one of the subjects; where sources rank Claude favorably, that ranking comes from the cited sources (often practitioner-tier, see Limitations), not from the researcher’s own judgment. Quotations were machine-extracted from web sources and should be verified verbatim before formal publication.


Executive summary

1. Grammar is essentially solved; style is not. In every major language surveyed, native experts no longer complain about grammatical errors from frontier models — Spanish critics note even the subjunctive, “long the tell of machine Spanish,” is now handled correctly; an Italian reviewer found Claude Opus 4.5 could write a grammatically correct story without the letter a; Russian critics have stopped mentioning case and aspect entirely. The criticism has moved up the stack, from correctness to nativeness: register, idiom, rhythm, typography, and voice.

2. Every language has independently coined a word for the same disease. Chinese 「AI味」 (AI flavor) and 「翻译腔」 (translation tone); Japanese 「翻訳調」; Korean 번역투; Russian канцелярит (bureaucratese) and «переводной привкус» (translated aftertaste); Arabic ركاكة (clumsy stiffness) — “superficial, resembling language clumsily translated from English”; Turkish çeviri kokusu (translation smell) — “it can translate into Turkish but cannot make it Turkish”; Spanish “lenguaje maquínico”; Italian “plasticoso”; German “Textroboter-Deutsch.” The convergence is striking: native readers in unrelated languages describe frontier-LLM output as their language wearing English clothes — calqued idioms, English sentence rhythm, English typography (the em-dash tell has gone global), and a flattened neutral register.

3. English is a tier above everything else — but the gap is about ceiling, not floor. Only in English (and, notably, Japanese) does the 2026 discourse include serious claims of near-professional literary quality: contest judges unable to distinguish AI from human fiction, named testers calling Claude Fable 5’s constrained verse “hair-raisingly good,” a Japanese tech journalist judging Opus 4.6’s 150,000-character novel as rivaling “the highest quality a human can write.” No such claims exist in any other language surveyed. But the floor is high everywhere in the top ~10 languages: fluent, coherent, structurally competent prose that most casual readers cannot identify as machine-made.

4. No major language is an outlier — the differences are of degree, not kind. The same syndrome recurs in every language surveyed: solved grammar, marked style, imported typography, flattened register, absent voice. What varies between languages is how much residual mechanical error sits underneath the style problem, and how well regional varieties are served — which is why the languages can be ranked into tiers (see “Ranking the languages” below) even though no language escapes the syndrome itself. The literary-register verdict is the same everywhere: “correct but dead” — as one Spanish-language editor put it, “the machine has read about grief, but has never lost anyone.”

5. The quality hierarchy tracks training data, with a twist. The full five-tier ranking is below (“Ranking the languages”), but in silhouette: English alone at the top; Japanese a surprising second; the big European and East Asian languages in a broad middle tier (grammar solved, style marked); Polish and Turkish below them with residual mechanical errors; Hindi, Indonesian, and Arabic at the bottom, where the translated feel still dominates — deficits native testers attribute to training-data share (e.g. “Polish is <0.5% of training data”). The twist: in Chinese, domestic models (DeepSeek, Kimi, Qwen) are now judged more natively idiomatic than Western frontier models, and several countries (Portugal’s Amália, Spain’s ALIA, Turkey’s Kumru, Poland’s Bielik, the Gulf’s Jais/ALLaM/Fanar) have funded sovereign models explicitly because frontier models default to a foreign or wrong-variety register.

6. Fiction and non-fiction have decisively split. Non-fiction — reports, memos, academic and journalistic prose — is rated serviceable-to-excellent by native professionals in every Tier 1–3 language, often “needing minimal editing.” Fiction at a literary standard is rated publishable nowhere, including English: the consensus ceiling is “competent pastiche” — Henry Oliver’s English verdict, “not slop, but it’s also not especially good,” generalizes almost perfectly across languages.

7. Among models, a consistent pattern in native rankings. Across Spanish, French, Italian, German, Polish, Russian, Chinese, Japanese, Korean, and Arabic sources, Claude is most often ranked the most natural, least “AI-smelling” stylist in the local language (with Mistral the French favorite for short non-fiction, Gemini praised for Hindi colloquial register and Japanese rhythm); GPT-5.x is repeatedly described as correct but generic/anglicized, with a widely noted writing-quality regression at the GPT-5/5.2 launches (harshest in Korean, where GPT-5’s launch Korean was called a regression to stiff translationese with garbled transliterations); Gemini reads formal and “schoolbook.” These rankings come disproportionately from practitioner/tech-blog sources, not literary critics — see Limitations.

8. The biggest evidence gap is the one this report cares most about. In nearly every language, the cell {newest frontier models} × {credentialed literary critics} is almost empty. Literary establishments critique “AI” generically or react to 2024–early-2025 models; the people who name Opus 4.6 or GPT-5.5 are tech bloggers. English and Japanese are the exceptions. This gap is itself a finding: serious per-model literary criticism of non-English output barely exists yet.


Research questions and verdicts

Three questions this project set out to answer:

Q1: Do frontier models (Opus 4.6+, Fable 5, GPT-5.x, Gemini 3.x) only write English well? Verdict: overstated, but directionally half-true at the top end. English is the only language where 2026 native experts debate whether output approaches professional literary quality (with Japanese a partial second). But “well” by any ordinary standard — fluent, correct, coherent, useful — is now true of at least the top 8–10 languages. The accurate statement: only in English does the ceiling approach literary; everywhere else (and arguably in English too) the product is correct-but-denatured prose.

Q2: Ranking the languages — how does quality vary across the fifteen? Verdict: a defensible ordinal ranking exists, in five tiers. Two caveats before the list. First, this ranks how well each language is served by the frontier as a whole, not any single model. Second, the ranking is confounded by critical culture: Japanese ranks high partly because its evaluators are the most engaged and fine-grained anywhere; Arabic ranks low partly because its critics test poetry — the hardest register in any language. Treat the tiers as robust and the within-tier order as soft.

Q3: Why do native judgments diverge so sharply — from “writes my language well” to “barely usable”? Verdict: because there are two calibrated instruments measuring different things, and both are honest. At the literary-professional standard, harsh verdicts are the norm in every language: a Polish reportage author called 2026 ChatGPT prose in her style “grafomania in its purest form”; Italy’s leading sociolinguist calls AI Italian “plasticky, unfragrant, standardized”; Spanish novelists call it “dehumanized style.” At the utility standard (email, reports, summaries), the same corpus says frontier output in Tier 2–3 languages is good and improving. Highly trained readers detect tells that ordinary readers demonstrably miss — in a Turkish university study, 160 literature students rated GPT-4.5’s pastiches more readable and “more human” than the canonical authors being imitated, while working authors simultaneously judged the same technology “flat.” The instruments disagree because they measure different layers of the same object; neither is over-sensitive, and neither is wrong.


Cross-language failure modes (the universal syndrome)

Ordered roughly by how universally they were reported:

  1. Calques and anglicisms — idioms translated literally, English argument scaffolding (“not only X but also Y” appears as a named tic in Spanish, Portuguese, French, Italian, German, and Chinese), English SVO rhythm in languages with freer word order.
  2. Typography and punctuation imported from English — the em-dash tell (documented as a public phenomenon in French — a Le Monde column and a prime-ministerial mini-scandal — and in Spanish, Portuguese, Italian, Turkish, Indonesian, Russian); English quotation marks where the language’s fiction demands dialogue dashes (raya, travessão, pauza dialogowa, tiret); wrong quote glyphs in German.
  3. Register flattening toward formal-neutral — español neutro / dubbing Spanish; Indonesian defaulting to stiff bahasa baku; Arabic defaulting to stiff MSA and unable to produce dialect; Hindi drifting Sanskritized/textbook; German drifting toward Beamtendeutsch; Korean/Japanese honorifics pitched one notch too formal.
  4. Formulaic rhetoric — enumerative scaffolding (首先/其次; 첫째/둘째; “en outre/par ailleurs”; “além disso”), tricolons, hollow intensifiers (“crucial”, “fascinante”, “혁신적인”, “atemberaubend”), hedged both-sidesing, uniform sentence-length rhythm.
  5. Voice and interiority deficit in fiction — the deepest, most consistent literary judgment everywhere: no personal voice, no lived specificity, obligatory tidy endings, “every sentence a banger” uniform intensity (English), emotional “room temperature” (Russian), 「感情表現があまり得意ではない」(Japanese, of GPT-5.5), “un bel testo non coincide con la letteratura” (Italian).
  6. Variety/dialect failure — European vs Brazilian Portuguese confusion (state-funded Amália exists because models sound “brasileiro por defeito”); voseo and regional Spanish lexicon failing at 80–90% rates in academic testing; Arabic dialects understood but not produced; Chinese internet slang and dialects weak in Western models.
  7. Residual mechanical errors in morphology-heavy, data-thin languages — Polish declension and quantifier agreement; Turkish suffix over-regularity; Hindi gender agreement and chandrabindu; Arabic poetic meter collapse (كسر الوزن); Korean particle misplacement and double passives.

And the strengths, equally consistent: long-form structural coherence (the flagship 2026 jump — Fable 5 “can hold a book”), canonical style pastiche (startlingly good even in Turkish divan poetry, fooling expert judges 88% of the time), editing/analysis (several literary professionals who dismiss LLM drafting praise LLM editorial judgment), and non-fiction registers approaching parity with competent human writing.


Per-language overviews

Each heading links to that language’s full research file — sources, credibility notes, and original-language quotations with translations.

English (baseline) — research notes

The 2026 ceiling: competent-MFA-workshop prose with reliable book-length coherence. The Hyperstition “Unslop” contest (judged by Gwern Branwen, Alexander Wales, et al.) found best-effort unedited frontier fiction “more literary than expected, but still recognizably AI”; Henry Oliver: “not slop, but it’s also not especially good… the same is true of much human fiction.” A documented capability jump around Claude Fable 5 (June 2026): “hair-raisingly good” constrained verse, dramatically reduced AI register, near-professional editorial judgment — with self-taste (“ranking the goodness of its own ideas”) the surviving weakness. Non-fiction judged near-done. Failure modes: cliché density, metaphor pile-ups, hollow profundity, uniform intensity, mode collapse.

Spanish — research notes

Grammar now near-flawless — Edmundo Paz Soldán (Cornell): “el subjuntivo ya es perfecto, pero a costa de perder la voz personal” (“the subjunctive is now perfect, but at the cost of losing the personal voice”). The consensus failure is denatured prose: “lenguaje maquínico,” anglicized calques, RAE orthotypography violations, dialect flattening to a dubbing-Spanish neutral, and formulaic rhetoric. Literary translators (ACE Traductores) judge output unfit for demanding literary text; Jorge Carrión concedes models “ya redactan mejor que muchos autores de libros superventas” (“already draft better than many bestseller authors”). Claude ranked most natural in Spanish by trade press (practitioner-tier sources). Academic testing shows voseo recognized but regional lexicon failing badly. A Spanish-specific irony: models overuse the em-dash in expository prose while failing to use the raya correctly where Spanish fiction actually requires it (dialogue) — an error an English reader structurally cannot see. No serious critic has yet close-read Opus 4.5+/Fable 5/GPT-5.x Spanish specifically — the literary commentary is largely model-agnostic or GPT-4-era.

Chinese (Mandarin) — research notes

Split verdict: for reasoning-heavy non-fiction Claude’s Chinese is rated top-tier with minimal 翻译腔; for creative and colloquial registers domestic models (DeepSeek, Kimi, Qwen) are judged more natively grounded — 地道感 is the axis Western models lose on. Mao Dun laureate 麦家: “DeepSeek may write better than 95% of people,” yet machines “can never surpass humans in literature” because they lack human limitation. Web-fiction editors report AI slush is identifiable and rejected (“bland, formulaic… no emotion, all stale tropes”); the native tell is now 「AI味」 — style, not grammar. Zhipu’s GLM-5.x line (GLM-5.2, June 2026) confirms the stylist hierarchy from the reverse direction: its writing identity is Claude-imitation — community shorthand 「克味」, “Claude flavor” — rated mid-pack for prose and judged to have declined across the 5.x releases as training focus went to coding/agent tasks; Chinese writers are routed to DeepSeek V4 Pro or Kimi K3 instead.

Hindi — research notes

Thinner coverage; consistent verdict: grammatical but translated — “planned in English, rendered in Hindi,” register drifting Sanskritized/textbook, subtle gender-agreement errors, cultural displacement (Western herbs where tulsi/pudina expected). Gemini’s Hindi repeatedly rated the most natural of the frontier models (“the way people actually speak… not textbook Hindi”). Literary commentary judges AI verse formally capable (meter, figures) but 「यांत्रिक」 (mechanical). Indic-model builders frame it structurally: global models “effectively translate rather than generate natively.”

Arabic — research notes

The harshest tier-3 profile. Signature critique: “سطحية وتشبه لغة مترجمة من الإنجليزية إلى العربية بركاكة” (“superficial, like language clumsily translated from English”). A pan-Arab poets’ panel (Al Bayan, Dec 2025) documents broken meter, defective rhyme, and school-essay إنشائية in AI verse. MSA correctness has improved markedly; producing dialect, idiom, and sarcasm has not. Gulf practitioner tests rate Claude best for dialect/cultural nuance, GPT-5 best for fusha. Critic Nadia Hanawi’s book-length verdict: texts that “imitate rather than create.”

French — research notes

Tier-2 exemplar. 2026 practitioner benchmarks rank Claude and Mistral as writing the most natural French; GPT/Gemini retain “subtle anglicisms.” The em-dash became a national AI tell (Le Monde column, April 2026). The startling counter-example: a Claude-written noir story judged against Goncourt winner Hervé Le Tellier (2025) — Le Tellier conceded it was better written than half of what gets published in France. Linguists still call AI French fiction “fades” (bland) and cliché-aggregating; the Ministry of Culture’s 2026 report frames LLM French as structurally anglo-centric — a sovereignty problem.

Portuguese — research notes

“Fluent but flat.” pt-BR grammatical fluency treated as solved; Brazilian editors catalogue 12–13 named “vícios de linguagem de IA.” Distinctive Lusophone axis: variety confusion — models default Brazilian for European users (“brasileiro por defeito”), motivating Portugal’s state-funded Amália (July 2026); conversely Brazil’s Maritaca claims Sabiá-4 beats Opus 4.8/GPT-5.4/Gemini 3.1 on Brazilian legal writing and slang benchmarks. The travessão panic: Brazilian authors falsely accused of AI use for normal Portuguese dialogue punctuation.

Russian — research notes

Grammar solved; style condemned. The canonical complaint is канцелярит plus calques (“стоит отметить, что…”), overused “является,” missing discourse particles (же, ведь), “room-temperature” affect. Tolstaya: asked it for a story in her own style — “nothing in common.” Shargunov (mid-2026): AI prose “over-sugared… the machine does not notice its own clichés.” Claude rated the most stylistically precise Western model in Russian practitioner bake-offs; domestic models (YandexGPT, GigaChat) trail on creative tasks. Critics concede AI is peerless at bureaucratic register — backhanded praise.

German — research notes

“Sehr gutes Deutsch” for functional prose; literary German recognizably machine-flavored. Wolfgang Tischer (literaturcafe.de, testing Opus 4.8, June 2026) catalogues stable fiction tells: semantic triads, stacked negations, café-in-the-rain openings, contradictory metaphors. GPT-5’s German judged “hölzerner, gestelzter” (more wooden, more stilted) than GPT-4o by Austrian practitioners. Claude Opus 4.8/Fable 5 rated most stylistically assured for long German texts. Translators’ association (VdÜ, re-affirmed July 2026) maintains literary post-editing isn’t worth it; translator Olga Radetzkaja’s question stands: “Formfleisch oder Universalpoesie?” (processed meat-paste or universal poetry?).

Japanese — research notes

The strongest non-English showing. At the 13th Nikkei Hoshi Shinichi Award (Feb 2026), 3 of 4 general-category winners used AI and judges “could not tell whether a work is by a human or by AI” — prompting judge/nonfiction writer 最相葉月’s protest: “I no longer want to read AI-written text” (a dignity objection, notably not a quality objection). Tech journalist 新清士 judges Opus 4.6’s 150k-character novel as rivaling top human quality; Japanese writers track per-version style drift (Opus 4.6 is the community’s favorite stylist; 4.8 not yet caught up — the finest-grained native evaluation culture found in any language). Remaining tells are rhythm-level: monotone sentence endings, metronomic paragraph cadence. Platforms (Narou, Kakuyomu) instituted mandatory AI disclosure after AI works won contests undetected.

Korean — research notes

Claude widely judged the naturalness leader (“writing that doesn’t smell like AI”); GPT-5’s launch Korean was received as a regression — stiff 직역체, awkward collocations, garbled transliterations (“George Washington” → “기어지 워싱지언”). Novelist 김연수 (June 2026, testing Claude Sonnet 4.6 and Gemini 3): useful for editing/ideation, “closer to repetitive labor than creativity” for actual drafting. Webnovel readers rating-bomb detected AI style; authors peer-review manuscripts to scrub AI-sounding expressions. Documented tells: calqued idioms, ~라고 할 수 있습니다 hedging, double passives, over-commaing (61% of AI sentences vs 26% human).

Italian — research notes

Grammatically solid (lipogram-passing) but “plasticoso, poco fragrante, standardizzato” (sociolinguist Vera Gheno). Practitioner catalogues document calques (“navigare le complessità”), em-dash abuse, congiuntivo slips, rhetorical inflation. Claude ranked most natural (“più umana, briosa”); GPT-5.2 seen as a writing regression (conceded by Altman, partially fixed in 5.3). Literary establishment concedes “bei testi” while denying they constitute literature: “un bel testo non coincide con la letteratura” (Lipperini).

Turkish — research notes

The memorable formulation: AI “Türkçeye çevirebiliyor ama Türkçeleştiremiyor” — it can translate into Turkish but cannot make it Turkish. Core diagnosis: üslupsuzluk (stylelessness), literal idiom transfer, em-dash alien to Turkish typography, uniform rhythm. Yet the strongest empirical result cuts the other way: GPT-4.5 pastiches of canonical authors fooled majorities of literature students and expert judges (Anadolu University studies), with AI prose rated more readable and even “more human” than the modernist masters. 2026 frontier Turkish rated “workable” for utilitarian registers; the frontier × literary-critic cell is empty.

Indonesian — research notes

Grammatically clean, “terlalu rapi” (too tidy); defaults to stiff bahasa baku, misses the formal→gaul register spectrum and regional coloring. A June 2026 peer-reviewed study: ChatGPT poems beat humans on structure and coherence, lose on diction originality, symbolic complexity, existential grounding. Establishment verdicts range from categorical rejection (“secara semiotik cacat dan rendah” — semiotically defective and low) to pragmatic collaboration (Hasan Aspahani’s AI-assisted story cycle). Regional models (SEA-LION, Sahabat-AI) outperform general models on localized tasks.

Polish — research notes

Two disjoint evidence pools: version-precise tech benchmarks (GPT-5.2 topped one 20-task Polish benchmark; Claude “makes inflection errors most rarely — 1-2 per article, not 5-10 like the competition”; domestic Bielik wins Polish-specific homonym traps) vs literary judgment without version numbers. Reportage author Olga Gitkiewicz on ChatGPT imitating her style (July 2026): “grafomania w czystej postaci… nieznośnie drętwe, sztuczne” (“graphomania in its purest form… unbearably wooden, artificial”). Root cause named by Polish testers: Polish is <0.5% of training data. Genre-formula fiction conceded; literary Polish denied.


Methodology and limitations

Language selection. Fifteen languages: English (baseline) plus Spanish, Mandarin Chinese, Hindi, Arabic, French, Portuguese, Russian, German, Japanese, Korean, Italian, Turkish, Indonesian, Polish — chosen by crossing speaker population with publishing/digital presence. Bengali, Urdu, Vietnamese, and others with larger speaker counts than some included languages were cut because this method needs a findable native literary-critical internet; Italian and Polish punch far above their speaker rank on that axis. This selection bias is real: the languages most likely to be worst served by LLMs are also the ones whose critical discourse this method cannot reach.

Search method. One research agent per language, searching in the target language (plus English where useful), instructed to find assessments by credentialed native speakers — critics, authors, literary translators, editors, linguists, academics — of models released since late 2025, with earlier-model critiques admitted only when vintage-flagged. Each agent returned source-by-source citations with original-language quotations and translations; these are preserved unedited in research/.

Limitations, in rough order of severity:

  1. The frontier × literary-critic cell is nearly empty outside English and Japanese. Most credentialed literary judgment reacts to “AI” generically or to 2024–early-2025 models; most named-version 2026 testing comes from SEO-adjacent tech blogs and vendor-conflicted practitioners. Verdicts about the newest models’ literary quality in most languages are therefore inference, not observation.
  2. The researcher is an LLM. Extraction and translation errors are possible; several sources were characterized from search snippets when pages were paywalled or bot-walled (each research file flags which). Verify quotes before republication.
  3. Conflict of interest. Claude both conducted this research and is a subject of it. The recurring “Claude most natural” finding is faithfully reported from sources but those sources skew practitioner-tier; a skeptic should weight it accordingly.
  4. Detection discourse contaminates quality discourse. Several languages now show false-positive panics (human writers accused of AI style — documented in English, Portuguese, Korean); “I can tell it’s AI” claims are becoming unreliable in both directions.
  5. AI-judged benchmarks disagree with human judges (documented in English: Fable 5 under-scored by LLM judges relative to named human testers), so leaderboard results were treated as context, not evidence.
  6. Survivorship of complaint. Experts who find AI prose unremarkable write columns; experts who quietly find it fine mostly don’t. The corpus likely over-represents negative judgments at every quality level.

Bottom line

The 2026 answer to “does frontier AI write X language well?” is the same sentence almost everywhere, with the dial set differently: grammatically excellent, stylistically identifiable, literarily insufficient — and less native the further the language sits from the center of the training distribution. English gets the real ceiling raise; the big European and East Asian languages get correct-but-denatured; data-thin and morphology-heavy languages still get mechanical errors on top. A reader who judges by utility will say the machines write their language well. A reader who judges by literature will say they don’t write it at all — in any language, including English. Both are reporting the same object.