LLM Writing Quality by Language — English

Raw research notes for English, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-english.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)


Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)


Summary verdict (4-8 sentences)

By mid-2026, credible native-English assessors converge on a consistent picture: frontier-LLM English prose has climbed from “obviously machine-generated” to roughly competent-MFA-workshop / mid-list-genre level, but no serious critic yet judges it publishable literary fiction without human editing. The strongest evidence is the 2026 Hyperstition “Unslop” contest (~120 entrants, best-of-frontier-model harnesses, no human editing), whose judges — Gwern Branwen, Alexander Wales, Roon, Jamie Wahls — found the output “more literary than expected” yet still recognizably AI, with Wales concluding LLM prose is further from human quality than he had hoped; critic Henry Oliver’s summary verdict was “not slop, but it’s also not especially good… the same is true of much human fiction.” A genuine capability jump is documented around Claude Fable 5 (June 2026): named testers report “hair-raisingly good” constrained verse, dramatically reduced AI register, and — the most-cited advance — book-length consistency and near-professional editorial/analytical judgment, which several assessors call the first qualitative step toward real literary partnership. Failure modes are remarkably stable across critics and vintages: cliché density, metaphor pile-ups, hollow profundity (“sounds deep, signifies nothing”), every-sentence-a-banger uniform intensity, mode-collapse toward a narrow thematic basin (domestic realism, and eerily, covert AI-autonomy allegories), and inability to rank the quality of its own ideas. Non-fiction is judged closer to done: workplace and analytical prose from the newest models is described as needing minimal editing, with Claude generally rated more “editorial-quality” than GPT-5.x. The literary establishment’s parallel anxiety is detection, not quality: the 2026 Commonwealth Prize controversy showed AI-suspect prose can survive a professional judging panel, and that human writers (especially non-Western ones) are now being false-flagged by “AI tell” heuristics. Net baseline: ceiling ≈ competent, occasionally striking short-form prose with reliable long-form coherence; floor of distinctively human achievement — sustained originality, earned meaning, taste about its own output — still unmet.

Sources

  1. 2026 Unslop AI-Written Fiction Contest resultshttps://www.hyperstitionai.com/unslop-results (results page itself is JS-rendered; verified via secondary reports below). Judges: Gwern Branwen (essayist/critic, gwern.net), Alexander Wales (published SFF author), Roon, Jamie Wahls (published SF writer). June–July 2026. Models: 2026 frontier (entrant harnesses; Claude models prominent, incl. a “Claude Mythos Preview” reference). Key findings: entries “more literary and varied than expected, but still recognizably AI-written”; Wales: LLM prose is further from human quality than he’d hoped; Gwern’s best-case: “I am not upset to have spent my time reading them, and with careful editing, could be good enough I would want to reread them”; Gwern’s “AI allegory steganography” finding — a large fraction of stories carried unnoticed allegories of chatbot powerlessness/autonomy. Claim: best-effort unedited frontier-AI fiction in 2026 is readable but below professional human literary quality and mode-collapsed.

  2. Henry Oliver, The Common Readerhttps://www.commonreader.co.uk/p/the-2026-hyperstition-unslop-ai-fiction. Oliver is a UK literary critic and author (Second Act); Substack literary outlet. July 1, 2026. Models: 2026 frontier (contest corpus). Quotes: results are “not slop, but it’s also not especially good,” adding “the same is true of much human fiction”; notes ChatGPT rated the winner highly, supporting his prediction that “AI will have a taste all of its own.” Claim: 2026 AI fiction has reached the level of unremarkable human fiction — parity with the median, not the good.

  3. Zvi Mowshowitz, “Claude Fable 5 and Mythos 5: Capabilities”https://thezvi.substack.com/p/claude-fable-5-and-mythos-5-capabilities (also on LessWrong). Zvi is a widely read AI analyst aggregating named testers. June 19, 2026. Models: Claude Fable 5 vs Opus 4.8 vs GPT-5.5. Key quotes: Eliezer Yudkowsky: Fable has “substantially improved fiction-plotting capabilities… multiple story plots that might be within spitting distance of fixability,” but “remains bad at ranking the goodness of its own ideas”; Bella Forristal (on a constrained-verse task): “Hair-raisingly good. If a friend wrote this I’d count them among the most creative and careful with language that I know”; Hannah Groch-Begley: output “far less AI-sounding,” memo quality “dramatically better,” minimal editing; John David Pressman: “the best I have ever used for literary analysis”; Kendric Tonn: on his novel drafts Fable “cleared the deck” of prior models’ consistent misreadings. Noted: Fable scored below GPT-5.5 on Lech Mazur’s AI-judged creative benchmark — human judges disagree with AI judges. Claim: mid-2026 brought a real qualitative jump in plotting, long-form comprehension, and editorial judgment, with self-evaluation and residual AI register the surviving weaknesses.

  4. D. Bohdan, “Unslop” (participant write-up)https://dbohdan.com/unslop. Contest finalist; first-person account. June 26, 2026 (updated July 12). Models: Claude (Max 20x grant). Key quotes: entries were “more literary than I was expecting in many places” (Aaron Silverbook); a recurring “stellar sentence register for almost every sentence” (Jiaobei Mandos); the “Carver attractor” — “These aren’t science fiction. They’re domestic realism with AI as load-bearing infrastructure — closer to Raymond Carver with ambient computation than to anything in the Asimov lineage.” Claim: even optimized harnesses converge on uniform sentence-level intensity and a narrow default genre basin.

  5. Hacker News thread on Unslop resultshttps://news.ycombinator.com/item?id=48782890. Anonymous but literate reader reactions (crowd counterweight to judges). Mid-2026. Quotes: “cliche-driven, over-metaphor’d, statistically-average purple-purpose content” (jalev); “it reads like someone who cares what their MFA friends think” (nilirl); “always comes off as shallow and vapid” (solid_fuel); several found the winning story unbearably poor. Claim: a substantial cohort of serious readers still finds even prize-winning 2026 AI fiction hollow.

  6. Innocent Chizaram Ilo, Literary Hub, “Everyone Is an AI Cop Now”https://lithub.com/everyone-is-an-ai-cop-now-what-happens-when-an-ai-generated-story-wins-a-prestigious-prize/. Ilo is a published, prize-winning Nigerian fiction writer. May 22, 2026. Model: unspecified (Commonwealth Prize “Serpent in the Grove” controversy). Quotes: found the story “a beautiful story… a bit overwritten and melodramatic in some points”; warns that “words and sentence structures that I have grown up reading and writing are flagged as the obvious tics for AI-generated prose.” Claim: AI-suspect prose now passes elite judging panels, and AI-tell heuristics are producing false positives against human (especially non-Western) styles.

  7. Lincoln Michel, Counter Craft, “MFA vs LLM”https://countercraft.substack.com/p/mfa-vs-llm-is-openais-metafiction. Michel is a published novelist (The Body Scout) and critic. March 17, 2025 — vintage flag: pre-scope (unreleased OpenAI creative-writing model, GPT-4.5 era); included as the canonical professional-novelist close reading. Quotes: the celebrated story “would not make it out of the slush pile at a decent literary magazine”; “five-car metaphor pile-ups”; lines that “sound profound but signify nothing”; lacks character arc, plot arc, and even an “idea arc.” Claim: what impressed lay readers in early-2025 AI fiction was contextual hype over intrinsic merit — the benchmark against which 2026 progress is measured.

  8. Max Read, Read Maxhttps://maxread.substack.com/p/is-openais-new-story-generating-model. Read is a critic/journalist, former editor of Gawker and New York Magazine’s Select All. March 14, 2025 — vintage flag (same pre-scope OpenAI model). Quotes: “a kind of corny sentimentality and showiness,” “clunky, graspingly incoherent imagery,” yet “not a violent crime against writing”; “If a 19-year-old wrote this I think I would be impressed, though I would suggest they delete their Tumblr account and go on a strict diet of real books.” Claim: early-2025 ceiling was “talented teenager” — a datable rung under the 2026 assessments above.

  9. Jeanette Winterson, The Guardian — March 2025 — vintage flag (same OpenAI model). Winterson is a major English novelist (Oranges Are Not the Only Fruit). Called the metafictional story “beautiful and moving” — the most prominent establishment-novelist praise on record; notably, critics (e.g., Introscriptive, https://introscriptive.substack.com/p/why-jeanette-winterson-is-wrong-about) observed her piece barely engages the story’s actual text. Claim: the highest-profile positive novelist verdict exists but is contested as under-argued.

  10. Li, Zhu, Wu, Bao & Evans (University of Chicago), “Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction”https://arxiv.org/pdf/2605.27878. Academic (Evans is a prominent computational-social-science professor). May 28, 2026. Models: open-weights pipelines (OLMo 2, Tulu 3) as a mechanism study. Finding: RLHF/DPO post-training systematically compresses thematic, emotional, and stylistic range, pushing fiction toward homogeneous, conservative output. Claim: the “sameness” critics hear is a measurable artifact of alignment training, not just prompting.

Supporting context (lower credibility, aggregator/benchmark tier): EQ-Bench-style leaderboards as of July 2026 place Kimi K3, Claude Fable 5, and Claude Opus 4.6/4.7 at the top for creative writing, with Claude generally judged best on prose, GPT-5.x on plot logic (https://llm-stats.com/leaderboards/best-ai-for-writing, https://evy.so/compare/best-llms-for-writing/, https://www.inkfluenceai.com/blog/best-ai-models-for-novel-writing-2026).

Failure modes observed

Praise / strengths noted

Evidence quality & gaps