LLM Writing Quality by Language — Russian
Raw research notes for Russian, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-russian.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)
Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)
Summary verdict
Native Russian assessments converge on a consistent picture: frontier LLMs now produce grammatically clean, fluent Russian — case/aspect errors have largely disappeared from the conversation — but the texts fail at the level of style and voice. The dominant complaints are канцелярит (bureaucratese), calques from English (“стоит отметить, что…”), overuse of “является”, missing discourse particles (же, ведь, вот), “room-temperature” affect, over-structured list-heavy formatting, and a translated feel. Literary professionals (Tolstaya, Shargunov, critic Lidia Maslova) judge AI Russian fiction as saccharine, cliché-blind, and emotionally hollow — good only for “тупая черновая работа” — while conceding it is unmatched at official/formal registers. In head-to-head tests by Russian-speaking practitioners, Claude is repeatedly rated the most stylistically precise Western model in Russian, though a professional translator found its Russian translations “неуклюжий местами”; domestic models (YandexGPT, GigaChat) trail Western frontier models on creative tasks but win on local idiom and availability. A significant caveat: much of the 2026-dated Russian-language comparison content is SEO/aggregator-driven, and very little rigorous native criticism targets the newest frontier tier (Opus 4.5+, GPT-5.x, Gemini 3.x) on fiction specifically — the strongest literary-world critiques still date to the YandexGPT-era experiments of 2024.
Sources
-
URL: https://habr.com/ru/articles/1022426/ Author: stas-clear (Станислав), working editor/author on tech and AI topics. Outlet: Habr (flag: informal-expert). Date: April 12 (Habr ID sequence places it 2026; the fetched page shows only “12 апр”). Models: unnamed, general frontier-LLM era. Quotes: “нейросеть слишком часто пишет текст, который выглядит как текст, но не ощущается как мысль” (“a neural net too often writes text that looks like text but doesn’t feel like thought”); “аккуратно собранная имитация мысли” (“a neatly assembled imitation of thought”); notes “переводной привкус” (“a translated aftertaste”). Claim: An editor’s verdict that AI Russian non-fiction irritates readers because it is “artificially correct yet empty” — false confidence, template structure, translated phrasing.
-
URL: https://vc.ru/id490208/2847189-plagin-dlya-russkogo-teksta-ubirayushchiy-sledy-neyroseti Author: Илья Утов, developer, daily Claude (Claude Code) user; built open-source (MIT) “humanizer-ru” plugin. Outlet: vc.ru (informal-expert). Date: April 3, 2026. Models: Claude (current 2026 vintage). Quotes: “В русском совершенно другой набор [маркеров]: канцелярит… кальки с английского… злоупотребление словом ‘является’, отсутствие частиц ‘же’” (“Russian has a completely different set [of markers]: bureaucratese… calques from English… abuse of the word ‘yavlyaetsya’ [is], absence of particles like ‘zhe’”). Catalogues 37 patterns in 8 categories, incl. “стоит отметить что” as a literal calque of “it’s worth noting that”, empty openers (“в современном мире”), dash overuse, bold-spam, “perfect typography as a marker.” Claim: Even current Claude output in Russian carries a stable, enumerable fingerprint of canceleritis and Anglicism requiring systematic de-AI-ification.
-
URL: https://vc.ru/ai/3006515-chatgpt-claude-gigachat-sravnenie Author: “НейроВед” (anonymous team; 200+ test prompts over two weeks). Outlet: vc.ru (informal-expert; some SEO flavor). Date: July 1, 2026. Models: GPT-4o, Claude 3.5 Sonnet, GigaChat — vintage flag: tests pre-frontier models despite 2026 date. Quotes: On ChatGPT in Russian: “в чисто русскоязычных задачах мы стабильно фиксировали характерные шаблоны” (“in purely Russian-language tasks we consistently recorded characteristic templates”) — clichés “в заключение”, “таким образом”. On Claude: “Claude 3.5 Sonnet показал себя как самый стилистически точный ИИ из тройки” (“Claude 3.5 Sonnet proved the most stylistically precise AI of the three”) — 92% style-match vs 74%/61%. On GigaChat: “Творческие задачи — слабое место” (“Creative tasks are its weak spot”). Claim: In a Russian practitioner bake-off, Claude leads on Russian style fidelity, ChatGPT defaults to bureaucratic clichés, GigaChat is strong on domestic domain text but weak creatively.
-
URL: https://news.mail.ru/society/62600936/ (syndicated from Известия) Author: Лидия Маслова, professional literary critic (Известия, ex-Kommersant). Outlet: Известия. Date: September 1, 2024 — vintage flag: YandexGPT era. Models: YandexGPT/Алиса (anthology “Механическое вмешательство”). Quotes: “когда речь идет о тонких движениях души и переливающихся оттенках смысла… результат потуг ИИ в лучшем случае можно отнести к жанру «непреднамеренной комедии»” (“when it comes to subtle movements of the soul and shimmering shades of meaning… the AI’s strainings at best belong to the genre of ‘unintentional comedy’”); “словесный шлак” (“verbal slag”); but “когда надо написать какой-то официальный, казенный текст на суконном канцелярите… тут нейросети равных нет” (“when an official, drab text in stiff bureaucratese is needed… the neural net has no equal”). Claim: A professional critic’s verdict that (domestic-model) AI fiction is unintentional comedy, fit only for rough drafting, while AI is peerless at bureaucratic register.
-
URL: https://moskvichmag.ru/lyudi/tatyana-tolstaya-hamit-alise-kak-proshla-vstrecha-s-avtorami-mehanicheskogo-vmeshatelstva Author: Ксения Василашвили (reporting); assessments by Татьяна Толстая (major Russian prose writer) and Ксения Буржская (novelist, Yandex Alisa AI trainer). Outlet: Москвич Mag. Date: November 12, 2024 — vintage flag: YandexGPT era. Quotes: Tolstaya: “Попросила ее написать рассказ в стиле Татьяны Толстой — ничего общего” (“I asked it to write a story in the style of Tatyana Tolstaya — nothing in common”); “Все рассказики у нее с добрым концом… Большая литература всегда про драму, на разрыв души и сердца” (“All its little stories have happy endings… Great literature is always about drama, tearing at soul and heart”). Burzhskaya (praise): “Я поручила ей сгенерировать монолог свекрови… она с ней справилась блестяще” (“I had it generate a mother-in-law’s monologue… it handled it brilliantly”). Claim: A canonical living stylist finds the model incapable of imitating literary voice or sustaining dramatic stakes; an AI-trainer novelist grants it competence at stock comic registers.
-
URL: https://snob.ru/literature/dopustimo-li-ispolzovat-neiroseti-v-literature-bolshoi-razgovor-s-kseniei-burzhskoi-i-sergeem-shargunovym/ Author/participants: Сергей Шаргунов (novelist, editor-in-chief of «Юность») and Ксения Буржская. Outlet: Сноб. Date: May 21, 2026. Models: unnamed current models + Алиса. Quotes: Shargunov: “они всегда какие-то переслащённые… слова подобраны приторно” (“they’re always somehow over-sugared… the words are chosen cloyingly”); “Машина не замечает собственные штампы” (“The machine does not notice its own clichés”). Burzhskaya: “у машины этой боли нет и не будет” (“the machine doesn’t have this pain and never will”). Claim: As of mid-2026 established Russian writers still find LLM prose saccharine and cliché-blind, with the deficit framed as experiential rather than grammatical.
-
URL: https://kinzhal.media/ask-kto-napisal/ Author: Кинжал editorial team (журнал Яндекс Практикума; Ilyakhov-school editing culture). Outlet: Кинжал (professional-editor guidance; informal-expert). Date: November 28, 2025. Models: LLMs generically (late-2025 vintage). Quotes: “Идеальная структура без личности — текст как из учебника, ни одной шероховатости” (“Perfect structure without personality — textbook-like text, not a single rough edge”); “Общая температура — комнатная: нет эмоции, нет позиции, нет удивления” (“Overall temperature — room temperature: no emotion, no stance, no surprise”); “Нелепые метафоры — «маяк в океане возможностей»” (“Absurd metaphors — ‘a lighthouse in the ocean of opportunities’”); “Очень часто — списки, везде и всегда” (“Very often — lists, everywhere and always”). Claim: Professional Russian editors codify AI Russian’s tells as flawless-but-lifeless structure, hedging formulas, kitsch metaphors, and compulsive listing.
-
URL: https://englishgeeks.ru/blog/test-kachestva-perevoda-15-neyrosetey Author: Мария Коржакова, professional written translator (письменный переводчик). Outlet: English Geeks blog. Date: February 2026. Models: 15 incl. ChatGPT, Claude, Gemini, DeepSeek, YandexGPT (specific sub-versions not frontier-flagged — treat as late-2025/early-2026 consumer tiers). Quotes: Claude: “неуклюжий местами перевод” (“translation awkward in places”), untranslated Anglicisms (“центр для бэкпекеров”); ChatGPT: “довольно выразительный” (“fairly expressive”) but with bureaucratic tone; DeepSeek: reads “довольно гладко” (“quite smoothly”), handled the idiom “cup of tea”; YandexGPT: “самый странный результат” (“the strangest result”) — failed the task outright. Claim: A working translator finds every model’s EN→RU output still needs a human pass — Western frontier models are serviceable but calque-prone, and YandexGPT is far behind.
-
URL: https://gorky.media/reviews/nejroset-horoshij-soyuznik-v-vojne-so-vzrosloj-kulturoj/ Author: Иван Напреенко, sociologist/critic at «Горький». Outlet: Горький (leading Russian literary review site). Date: June 29, 2022 — vintage flag: ruGPT-3, pre-frontier; included as literary-critical baseline. Models: ruGPT-3 (Pepperstein’s “Пытаясь проснуться”). Quotes: the prose’s defining trait is “незапоминаемость” (“unmemorability”) — it “разворачивает свой текст на плато” (“unrolls its text on a plateau”) with no accentuation; strengths: a hypnotic, séance-like “лярвический характер” prized only as avant-garde texture. Claim: The earliest serious Russian literary-critical framing: machine prose lacks the salience-hierarchy of embodied consciousness — flat, unmemorable, interesting only as found-object aesthetics.
-
Supporting/context: https://habr.com/ru/articles/1008120/ (Prometeriy/Влад Орлинскас, Habr, March 9, 2026 — “без хорошей редактуры генерация читается ужасно и сразу бросается в глаза” / “without good editing, generation reads terribly and is immediately conspicuous”, but edited AI text is now unidentifiable; “battle for text is lost”); https://habr.com/ru/articles/888614/ (Chad_AI, Habr, March 6, 2025 — six markers incl. канцелярит example “Рациональное распределение времени между трудовой деятельностью и отдыхом” vs human “Поработали — надо отдохнуть!”).
Failure modes observed
- Канцелярит / bureaucratese — the single most-cited flaw across editors, critics, and tool-builders (Maslova, Utov, Chad_AI, НейроВед): “рациональное распределение времени…”, “таким образом”, “в заключение”.
- Calques from English — “стоит отметить, что” (< “it’s worth noting that”), untranslated Anglicisms (“бэкпекер”, “центр для бэкпекеров”); translated feel (“переводной привкус”).
- Overuse of “является” and copula-heavy syntax mirroring English.
- Missing discourse particles (же, ведь, вот) — absence of the connective tissue “of any living Russian text” (Utov).
- Flat affect / “room temperature” — no stance, no surprise, no emphasis hierarchy; Napreenko’s “незапоминаемость”, Kinzhal’s “комнатная температура”.
- Oversweetness and cliché-blindness in fiction — Shargunov’s “переслащённые… приторно”; “машина не замечает собственные штампы”; Tolstaya: obligatory happy endings, inability to hold a named author’s style.
- Kitsch metaphors — “маяк в океане возможностей”.
- Structural compulsions — lists everywhere, bold-spam, dash (тире) overuse, symmetrical paragraphs, “PowerPoint” texts; padding/“резиновый текст” (repeating one idea to fill volume).
- Register failures — inability to be colloquial-informal convincingly; GigaChat weakest on informal synonyms and poetry.
- Notably absent from 2025–26 native complaints: case, aspect, and agreement errors — grammar is treated as solved; all criticism has moved up to style, register, and voice.
Praise / strengths noted
- Official/formal register mastery: Maslova — “тут нейросети равных нет” for казённый text; consensus that AI is excellent for business letters, descriptions, FAQs.
- Claude rated best Western model for Russian style in practitioner comparisons (92% style-match in НейроВед test; vc.ru commentary that Claude’s Russian business text is “без характерных «машинных» оборотов”).
- DeepSeek reads “довольно гладко” in RU translation and handles idioms (Korzhakova).
- Domestic models’ local grounding: YandexGPT knows Russian trends/memes/VK-Telegram formats; GigaChat Max strong for banking/legal/medical Russian.
- Stock-register competence in fiction: Burzhskaya’s mother-in-law monologue “блестяще”; Pepperstein found ruGPT-3’s flatness aesthetically usable as hypnotic texture.
- Edited AI Russian prose is now judged indistinguishable from human text (Habr 1008120) — praise of fluency framed as a societal problem.
Evidence quality & gaps
- Strongest evidence: professional critics and writers (Maslova/Известия, Tolstaya, Shargunov/Snob), a working translator’s structured 15-model test (Korzhakova), and editor-culture codifications (Kinzhal, Utov’s 37-pattern catalogue). These are genuinely native, credentialed, and specific.
- Vintage skew: the deepest literary critiques (Механическое вмешательство reviews, Gorky/Pepperstein) target YandexGPT-2024 and ruGPT-3-2022, not the frontier tier. The May-2026 Snob conversation is current but names no models.
- Frontier gap: no credible native assessment found that specifically evaluates Claude Opus 4.5/4.6, Fable 5, GPT-5.x, or Gemini 3.x on Russian fiction. 2026 Habr/vc.ru coverage of these models is access-guide/SEO material (aggregator ads: chadgpt, mashagpt, study24) that asserts Russian quality without evidence; even the best 2026 bake-off (НейроВед) still tested GPT-4o/Claude 3.5.
- SEO contamination is severe in Russian-language search results for model comparisons — most “ТОП нейросетей 2026” listicles are affiliate content and were excluded.
- Non-native corroboration (excluded from Sources per scope, noted here): Western MT-industry blogs claim editors found Claude/Gemini Russian drafts “mechanical and stiff”, mirroring English syntax — directionally consistent with Korzhakova but not native testimony.
- Fiction vs non-fiction split is clear in the evidence: non-fiction criticism centers on канцелярит/structure (fixable by editing); fiction criticism centers on affect, cliché, and voice (framed by writers as categorical). No native source yet documents the dialogue-punctuation (тире) convention as a specific frontier-model failure — dash overuse is cited, but as typography, not dialogue formatting.