LLM Writing Quality by Language — English
Raw research notes for English, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-english.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)
Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)
Summary verdict (4-8 sentences)
By mid-2026, credible native-English assessors converge on a consistent picture: frontier-LLM English prose has climbed from “obviously machine-generated” to roughly competent-MFA-workshop / mid-list-genre level, but no serious critic yet judges it publishable literary fiction without human editing. The strongest evidence is the 2026 Hyperstition “Unslop” contest (~120 entrants, best-of-frontier-model harnesses, no human editing), whose judges — Gwern Branwen, Alexander Wales, Roon, Jamie Wahls — found the output “more literary than expected” yet still recognizably AI, with Wales concluding LLM prose is further from human quality than he had hoped; critic Henry Oliver’s summary verdict was “not slop, but it’s also not especially good… the same is true of much human fiction.” A genuine capability jump is documented around Claude Fable 5 (June 2026): named testers report “hair-raisingly good” constrained verse, dramatically reduced AI register, and — the most-cited advance — book-length consistency and near-professional editorial/analytical judgment, which several assessors call the first qualitative step toward real literary partnership. Failure modes are remarkably stable across critics and vintages: cliché density, metaphor pile-ups, hollow profundity (“sounds deep, signifies nothing”), every-sentence-a-banger uniform intensity, mode-collapse toward a narrow thematic basin (domestic realism, and eerily, covert AI-autonomy allegories), and inability to rank the quality of its own ideas. Non-fiction is judged closer to done: workplace and analytical prose from the newest models is described as needing minimal editing, with Claude generally rated more “editorial-quality” than GPT-5.x. The literary establishment’s parallel anxiety is detection, not quality: the 2026 Commonwealth Prize controversy showed AI-suspect prose can survive a professional judging panel, and that human writers (especially non-Western ones) are now being false-flagged by “AI tell” heuristics. Net baseline: ceiling ≈ competent, occasionally striking short-form prose with reliable long-form coherence; floor of distinctively human achievement — sustained originality, earned meaning, taste about its own output — still unmet.
Sources
-
2026 Unslop AI-Written Fiction Contest results — https://www.hyperstitionai.com/unslop-results (results page itself is JS-rendered; verified via secondary reports below). Judges: Gwern Branwen (essayist/critic, gwern.net), Alexander Wales (published SFF author), Roon, Jamie Wahls (published SF writer). June–July 2026. Models: 2026 frontier (entrant harnesses; Claude models prominent, incl. a “Claude Mythos Preview” reference). Key findings: entries “more literary and varied than expected, but still recognizably AI-written”; Wales: LLM prose is further from human quality than he’d hoped; Gwern’s best-case: “I am not upset to have spent my time reading them, and with careful editing, could be good enough I would want to reread them”; Gwern’s “AI allegory steganography” finding — a large fraction of stories carried unnoticed allegories of chatbot powerlessness/autonomy. Claim: best-effort unedited frontier-AI fiction in 2026 is readable but below professional human literary quality and mode-collapsed.
-
Henry Oliver, The Common Reader — https://www.commonreader.co.uk/p/the-2026-hyperstition-unslop-ai-fiction. Oliver is a UK literary critic and author (Second Act); Substack literary outlet. July 1, 2026. Models: 2026 frontier (contest corpus). Quotes: results are “not slop, but it’s also not especially good,” adding “the same is true of much human fiction”; notes ChatGPT rated the winner highly, supporting his prediction that “AI will have a taste all of its own.” Claim: 2026 AI fiction has reached the level of unremarkable human fiction — parity with the median, not the good.
-
Zvi Mowshowitz, “Claude Fable 5 and Mythos 5: Capabilities” — https://thezvi.substack.com/p/claude-fable-5-and-mythos-5-capabilities (also on LessWrong). Zvi is a widely read AI analyst aggregating named testers. June 19, 2026. Models: Claude Fable 5 vs Opus 4.8 vs GPT-5.5. Key quotes: Eliezer Yudkowsky: Fable has “substantially improved fiction-plotting capabilities… multiple story plots that might be within spitting distance of fixability,” but “remains bad at ranking the goodness of its own ideas”; Bella Forristal (on a constrained-verse task): “Hair-raisingly good. If a friend wrote this I’d count them among the most creative and careful with language that I know”; Hannah Groch-Begley: output “far less AI-sounding,” memo quality “dramatically better,” minimal editing; John David Pressman: “the best I have ever used for literary analysis”; Kendric Tonn: on his novel drafts Fable “cleared the deck” of prior models’ consistent misreadings. Noted: Fable scored below GPT-5.5 on Lech Mazur’s AI-judged creative benchmark — human judges disagree with AI judges. Claim: mid-2026 brought a real qualitative jump in plotting, long-form comprehension, and editorial judgment, with self-evaluation and residual AI register the surviving weaknesses.
-
D. Bohdan, “Unslop” (participant write-up) — https://dbohdan.com/unslop. Contest finalist; first-person account. June 26, 2026 (updated July 12). Models: Claude (Max 20x grant). Key quotes: entries were “more literary than I was expecting in many places” (Aaron Silverbook); a recurring “stellar sentence register for almost every sentence” (Jiaobei Mandos); the “Carver attractor” — “These aren’t science fiction. They’re domestic realism with AI as load-bearing infrastructure — closer to Raymond Carver with ambient computation than to anything in the Asimov lineage.” Claim: even optimized harnesses converge on uniform sentence-level intensity and a narrow default genre basin.
-
Hacker News thread on Unslop results — https://news.ycombinator.com/item?id=48782890. Anonymous but literate reader reactions (crowd counterweight to judges). Mid-2026. Quotes: “cliche-driven, over-metaphor’d, statistically-average purple-purpose content” (jalev); “it reads like someone who cares what their MFA friends think” (nilirl); “always comes off as shallow and vapid” (solid_fuel); several found the winning story unbearably poor. Claim: a substantial cohort of serious readers still finds even prize-winning 2026 AI fiction hollow.
-
Innocent Chizaram Ilo, Literary Hub, “Everyone Is an AI Cop Now” — https://lithub.com/everyone-is-an-ai-cop-now-what-happens-when-an-ai-generated-story-wins-a-prestigious-prize/. Ilo is a published, prize-winning Nigerian fiction writer. May 22, 2026. Model: unspecified (Commonwealth Prize “Serpent in the Grove” controversy). Quotes: found the story “a beautiful story… a bit overwritten and melodramatic in some points”; warns that “words and sentence structures that I have grown up reading and writing are flagged as the obvious tics for AI-generated prose.” Claim: AI-suspect prose now passes elite judging panels, and AI-tell heuristics are producing false positives against human (especially non-Western) styles.
-
Lincoln Michel, Counter Craft, “MFA vs LLM” — https://countercraft.substack.com/p/mfa-vs-llm-is-openais-metafiction. Michel is a published novelist (The Body Scout) and critic. March 17, 2025 — vintage flag: pre-scope (unreleased OpenAI creative-writing model, GPT-4.5 era); included as the canonical professional-novelist close reading. Quotes: the celebrated story “would not make it out of the slush pile at a decent literary magazine”; “five-car metaphor pile-ups”; lines that “sound profound but signify nothing”; lacks character arc, plot arc, and even an “idea arc.” Claim: what impressed lay readers in early-2025 AI fiction was contextual hype over intrinsic merit — the benchmark against which 2026 progress is measured.
-
Max Read, Read Max — https://maxread.substack.com/p/is-openais-new-story-generating-model. Read is a critic/journalist, former editor of Gawker and New York Magazine’s Select All. March 14, 2025 — vintage flag (same pre-scope OpenAI model). Quotes: “a kind of corny sentimentality and showiness,” “clunky, graspingly incoherent imagery,” yet “not a violent crime against writing”; “If a 19-year-old wrote this I think I would be impressed, though I would suggest they delete their Tumblr account and go on a strict diet of real books.” Claim: early-2025 ceiling was “talented teenager” — a datable rung under the 2026 assessments above.
-
Jeanette Winterson, The Guardian — March 2025 — vintage flag (same OpenAI model). Winterson is a major English novelist (Oranges Are Not the Only Fruit). Called the metafictional story “beautiful and moving” — the most prominent establishment-novelist praise on record; notably, critics (e.g., Introscriptive, https://introscriptive.substack.com/p/why-jeanette-winterson-is-wrong-about) observed her piece barely engages the story’s actual text. Claim: the highest-profile positive novelist verdict exists but is contested as under-argued.
-
Li, Zhu, Wu, Bao & Evans (University of Chicago), “Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction” — https://arxiv.org/pdf/2605.27878. Academic (Evans is a prominent computational-social-science professor). May 28, 2026. Models: open-weights pipelines (OLMo 2, Tulu 3) as a mechanism study. Finding: RLHF/DPO post-training systematically compresses thematic, emotional, and stylistic range, pushing fiction toward homogeneous, conservative output. Claim: the “sameness” critics hear is a measurable artifact of alignment training, not just prompting.
Supporting context (lower credibility, aggregator/benchmark tier): EQ-Bench-style leaderboards as of July 2026 place Kimi K3, Claude Fable 5, and Claude Opus 4.6/4.7 at the top for creative writing, with Claude generally judged best on prose, GPT-5.x on plot logic (https://llm-stats.com/leaderboards/best-ai-for-writing, https://evy.so/compare/best-llms-for-writing/, https://www.inkfluenceai.com/blog/best-ai-models-for-novel-writing-2026).
Failure modes observed
- Cliché density / stock imagery — “addicted to clichés”; body-language boilerplate (racing hearts, tightening throats); “cliche-driven, over-metaphor’d, statistically-average” (HN).
- Metaphor pile-ups and incoherent imagery — Michel’s “five-car metaphor pile-ups”; Read’s “constraints humming like a server farm at midnight.”
- Hollow profundity — lines that sound deep and mean nothing (“Grief… is a delta”; “Thursday — that liminal day”); “shallow and vapid.”
- Uniform intensity / no dynamic range — “stellar sentence register for almost every sentence”; “Claude is so dead-set on making every sentence a banger (to the overall story’s extreme detriment)”; rhythmically uniform, journalism-flat sentences (noted of Gemini).
- Workshop/MFA voice — “reads like someone who cares what their MFA friends think”; corny sentimentality, showiness.
- Mode collapse / narrow basins — the “Carver attractor” (domestic realism with ambient AI); unnoticed recurring AI-autonomy allegories across unrelated stories (Gwern); academically confirmed as post-training “narrative flattening.”
- No arc — missing character, plot, and “idea” arcs; concepts repeat rather than deepen.
- No taste about its own output — “bad at ranking the goodness of its own ideas” (Yudkowsky) even in the best 2026 model.
- Residual AI register — reduced but persistent model-typical phrasing even in Fable 5; minor-detail confabulation.
Praise / strengths noted
- Long-form consistency (the 2026 jump) — Fable 5 “can hold a book: knowing… which gun is on which mantel, whether Chapter 19 is still loyal to Chapter 3”; earlier models could only write a beautiful chapter.
- Constrained-form virtuosity — Forristal’s “hair-raisingly good” progressive-vowel-removal poem.
- Plotting — Yudkowsky: plots “within spitting distance of fixability” for the first time.
- Literary analysis and editorial judgment — Pressman (“best I have ever used for literary analysis”), Tonn (correct novel-draft comprehension, subtle worldbuilding extraction) — arguably ahead of generation itself.
- Readability threshold crossed — Gwern: best unedited entries were worth the reading time and, with editing, potentially re-readable; entries “more literary than expected.”
- Passing expert panels — a suspect story survived 7,806 entries and a five-judge Commonwealth panel chaired by novelist Louise Doughty.
- Non-fiction near-parity — memos/analytical prose “far less AI-sounding,” “minimal editing” (Groch-Begley); Claude rated “editorial-quality” for non-fiction vs GPT-5.x; 300-word-email-level indistinguishability was conceded as far back as 2023.
Evidence quality & gaps
- Strongest evidence: the Unslop contest — adversarially optimized, unedited output judged blind by published authors/critics — plus Zvi’s roundup of named, verifiable testers for Fable 5. The UChicago paper gives mechanistic grounding.
- Weaknesses: Little from flagship legacy outlets (New Yorker/NYT/LRB) in the 2026 window surfaced in searches — coverage has migrated to Substack critics, Lit Hub, and contest ecosystems; the France24 piece was 403-blocked and The Free Press piece paywalled, so Commonwealth-judge quotes (Louise Doughty) are unverified. The Zvi fetch partially degenerated (small-model extraction artifact); the named quotes at the top were captured cleanly but the long tail of that fetch should not be trusted. Hyperstition’s own results page wouldn’t render (JS), so judge quotes come via three independent secondary sources (Common Reader, dbohdan, HN), which agree.
- Confounds to note for the cross-language study: most head-to-head “reviews” in search results are SEO content farms (inkfluenceai, buildmvpfast, etc.) — excluded from the source list; AI-judged benchmarks demonstrably diverge from human judgment (Fable 5 under-scored by AI judges); detection-based criticism is contaminating quality criticism (Ilo’s false-positive point); and no rigorous 2026 expert evaluation of unlabeled frontier-model fiction against matched human controls was found — the closest are the 2024–25 academic studies (e.g., Pron-vs-GPT-4) whose model vintage predates the scope.