LLM Writing Quality by Language — Polish
Raw research notes for Polish, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-polish.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)
Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)
Summary verdict
Coverage of frontier-model (late-2025/2026) Polish writing quality by native speakers exists but is split between two communities that barely overlap: tech reviewers who run systematic model-vs-model Polish tests (naming GPT-5.2/5.5, Claude 4.6/Sonnet 5, Gemini 3.1 Pro, Bielik, PLLuM), and literary critics/writers who assess “AI” prose seriously but rarely name model versions. The tech-test consensus (2026): frontier models now write largely correct Polish — Claude is repeatedly singled out as making the fewest inflection/syntax errors and “sounding like a Pole rather than an American translating into Polish” — but all models still show systematic declension errors, calques from English, and stylistic drift, attributed to Polish being <0.5% of training data. The literary-community verdict is harsher: reporter Olga Gitkiewicz (July 2026) called ChatGPT’s attempt at her style “grafomania w czystej postaci” — wooden, artificial, logic-free; critics Barbara Rojek and others hold that AI still cannot write good Polish poems or literary prose on its own, while conceding it handles formula genre fiction. The one high-profile fiction data point is Justyna Bargielska’s AI-collaboration poetry volume Kubek na tsunami (Sept 2025), which was reviewed on its merits — with the AI treated as raw material, not author. A genuine gap: no substantive native-critic assessment of Claude Opus 4.5+/Fable 5/Gemini 3-specific literary Polish was found; academic linguistic error analyses (Mazur/UJ) are pre-frontier vintage.
Sources
-
URL: https://www.egospodarka.pl/196675,ChatGPT-Gemini-Claude-czy-Bielik-ktore-modele-AI-najlepiej-radza-sobie-z-jezykiem-polskim,1,14,1.html Author: Test by Marek Jeleśniański, founder of Oxido (Polish digital agency); tech/business outlet eGospodarka.pl. Date: 2026-03-19. Models: GPT-5.2, Claude 4.6, Gemini 3.1 Pro, Grok 4.2, Qwen 3.5, DeepSeek V3.2, EuroLLM, Mistral, Bielik, PLLuM 8x7B — current frontier, no vintage flag needed. Quotes: “model skupia się na tym, że słowo oznacza istniejący region geograficzny, i ‘dopowiada’ sobie resztę, ignorując błąd ortograficzny” — “the model fixates on the word denoting a real geographic region and ‘fills in’ the rest, ignoring the spelling error” (on the pomoże/Pomorze homonym trap, which Bielik passed and Grok 4.2, Gemini 3.1 Pro, DeepSeek failed). Claim: In a 20-task Polish-language benchmark, US/Chinese frontier models beat Polish models overall (GPT-5.2: 7.66 vs Bielik: 6.38), but all still stumble on proper-name declension, stylistic subtlety, and Polish-specific homonyms where Bielik excels.
-
URL: https://aiport.pl/testy-i-rankingi/chatgpt-vs-claude-vs-gemini-ktory-najlepiej-pisze-po-polsku/ Author: Piotr Wolniewicz, AIPORT.pl (Polish AI-tools review site — practitioner, not literary critic). Date: 2026-02-11, updated 2026-07-02. Models: GPT-5.5 (Instant/Thinking), Claude Sonnet 5 and Sonnet 4.6, Gemini 3.1 Pro / 3.5 Flash — current frontier. Quotes: “Claude najrzadziej popełnia błędy fleksyjne i składniowe — jego pomyłki to zazwyczaj 1-2 na cały artykuł, a nie 5-10 jak u konkurencji” — “Claude makes inflectional and syntactic errors most rarely — usually 1-2 per article, not 5-10 like the competition.” “Polski stanowi mniej niż 0,5% danych treningowych” — “Polish makes up less than 0.5% of training data.” Claim: Systematic non-fiction writing tests find Claude best on Polish grammatical correctness and long-form stylistic consistency, ChatGPT best on dialogue naturalness and tone (~4% inflection error rate), Gemini prone to tonal “dryfowanie” (drift).
-
URL: https://spidersweb.pl/plus/2026/07/potwor-ai-zywi-sie-polska-literatura Author: Marek Szymaniak & Rafał Pikuła, Spider’s Web+ (major Polish tech/longform outlet); quality assessment from Olga Gitkiewicz, award-winning Polish reportage author. Date: 2026-07-13. Models: ChatGPT (2026-era, so GPT-5.x — exact version not named; mild vintage-precision flag), plus Claude/Gemini/Grok/Llama discussed re: training data. Quotes: Gitkiewicz on ChatGPT imitating her style: “To, co wypluł, okazało się grafomanią w czystej postaci. Było tak nieznośnie drętwe, sztuczne, dalekie od czyjegokolwiek stylu, a do tego kompletnie pozbawione logiki” — “What it spat out turned out to be graphomania in its purest form. It was so unbearably wooden, artificial, far from anyone’s style, and on top of that completely devoid of logic.” Also: “Tekst… był pełen niedopowiedzeń, błędów logicznych, a nawet językowych” — “The text… was full of vagueness, logical errors, even linguistic ones.” Claim: A professional Polish prose writer who tested a 2026 frontier chatbot on imitating her own reportage style judged the output worthless as literature — stylistically dead and logically incoherent.
-
URL: https://www.tygodnikpowszechny.pl/sztuczna-inteligencja-juz-pomaga-pisarzom-czy-faktycznie-zle-192223 Author: Marcin Wilkowski; quoted: poet/essayist Tadeusz Dąbrowski, novelist Piotr Siemion, poet Justyna Bargielska. Outlet: Tygodnik Powszechny (leading Polish intellectual weekly). Date: 2025-11-18. Models: ChatGPT, Claude (versions unnamed — late-2025 vintage, borderline in-scope). Quotes (as reported): Dąbrowski — AI lacks the “courage” to deliberately spoil too-perfect writing, which essays require. Siemion — AI can churn out genre fiction fast, but humans keep “fundamental superiority” grounded in “unrepeatable experience and sensitivity”. Bargielska on her ChatGPT collaboration: a relationship “between something backed by knowledge but not feelings, and something backed by feelings but only a fraction of knowledge.” Claim: Established Polish writers concede frontier AI competence at formulaic prose while denying it literary quality in Polish — its output is too smooth, uninhabited by experience.
-
URL: https://malyformat.com/2025/12/kocur-literacki-7-o-czym-justyna-bargielska-nie-porozmawiala-ze-sztuczna-inteligencja/ (plus review at https://www.rp.pl/plus-minus/art43298201-kubek-na-tsunami-zywotna-zaloba-przegadana-ze-sztuczna-inteligencja) Author: Barbara Rojek, literary critic, editor at “Mały Format” and “Twórczość”; Rzeczpospolita review by Damian Piwowarczyk. Date: Mały Format issue 09-10/2025 (publ. Dec 2025); rp.pl 2025-11-07. Models: Bargielska used ChatGPT (late-2025 vintage). Quotes: Rojek: “na razie sztuczna inteligencja nie pisze sama dobrych wierszy” — “for now, artificial intelligence does not write good poems on its own.” Piwowarczyk credits the book’s power to Bargielska’s own “hipnotyczne i drapieżne obrazki, zderzeń języka” (“hypnotic, predatory images, collisions of language”) confronting lived experience — not to the AI. Claim: In the highest-profile Polish AI-fiction case (Kubek na tsunami, Biuro Literackie 2025), critics judged the AI incapable of good Polish poetry unaided; whatever succeeds in the book is attributed to the human poet’s curation.
-
URL: https://forsal.pl/lifestyle/technologie/artykuly/9504993,polski-byc-trudna-jezyk-dla-ai-jak-chatgpt-radzi-sobie-z-polszczyzna.html Author: Jolanta Nabiałek (philologist/journalist), reporting research by Rafał Mazur, Polish Philology, Jagiellonian University (related academic paper in LingVaria, UJ). Date: 2024-05-13 — VINTAGE FLAG: pre-frontier ChatGPT (GPT-4 era); useful as a linguistic baseline only. Quotes: Agreement error example: “Wielu pisarzy i poetów… starali się zgłębić” — should be “starało się” (“wielu” requires singular neuter verb); lexical clunkers “popada w mroczne czyny” (“falls into dark deeds”), “namiętnie ambitny” (“passionately ambitious”). Claim: Academic error-count of a ChatGPT matura essay: 76 syntax, 54 lexical, 48 punctuation errors — establishing the pre-frontier failure-mode taxonomy (agreement, punctuation, awkward collocations, repetitions) that 2026 tests measure improvement against.
-
URL: https://www.dwutygodnik.com/artykul/11956-sztuczna-inteligencja-jest-literatura.html Author: Jerzy Stachowicz, cultural studies scholar, Institute of Polish Culture, University of Warsaw. Outlet: Dwutygodnik. Date: June 2025 — mild vintage flag (pre-late-2025 models). Quote: “Pisanie z czatem to całkiem żmudna praca z tekstem. Nikt tu nie popracuje za nas, a zwłaszcza nie przejmie za nas myślenia.” — “Writing with a chatbot is quite tedious text-work. Nobody here will do the work for us, least of all take over the thinking.” Claim: A UW academic who tested LLM writing in Polish found unedited output generic; acceptable Polish prose emerges only through heavy human prompting and revision.
-
URL: https://xyz.pl/burza-wokol-bielika-i-pllum-to-jak-porownywac-f-16-do-furgonetki-czy-polskie-ai-naprawde-przegrywa-wywiad/ (not fully fetched; from search results) Author: Interview with Sebastian Kondracki, head of the Bielik project (SpeakLeash). Date: 2026 (responding to the Jeleśniański test). Models: Bielik vs GPT-5.2-class frontier. Claim: Kondracki disputes the benchmark methodology (“like comparing an F-16 to a delivery van”), arguing a Polish-native-trained model of modest size can beat far larger frontier models on Polish-specific reasoning and instruction tasks — the domestic counterpoint in the quality debate.
Failure modes observed
- Inflection/declension errors — still present in all frontier models (aiport: Claude ~1-2 per article, competitors 5-10, ChatGPT ~4% inflection error rate); proper-name declension singled out (eGospodarka).
- Subject-verb agreement with quantifiers (“wielu… starali się” for “starało się”) — documented in the pre-frontier UJ study; the canonical Polish AI error class.
- Calques from English — repeatedly named (Forsal, LingVaria paper, general AI-copywriting commentary): constructions translated structurally from English that “brzmią po prostu dziwnie” (just sound weird).
- Awkward collocations / idiom poverty — “popada w mroczne czyny”, “namiętnie ambitny” (Mazur/UJ); unjustified repetitions and weak pronoun use.
- Homonym/orthography blindness — pomoże vs Pomorze: frontier models (Grok 4.2, Gemini 3.1 Pro, DeepSeek V3.2) “fill in” meaning past a spelling error (Jeleśniański).
- Stylistic deadness in fiction — “nieznośnie drętwe, sztuczne” wooden/artificial prose, no individual style, logical incoherence when imitating a real author’s voice (Gitkiewicz); too-perfect smoothness, no “courage” to break register (Dąbrowski).
- Tonal drift across long texts — Gemini’s “dryfowanie” (aiport).
- Punctuation — high error counts in the academic baseline; note: no source found specifically discussing the pauza dialogowa (—) dialogue convention — a real coverage gap.
- Root cause cited by Polish testers: Polish is <0.5% of training data vs >50% English.
Praise / strengths noted
- Claude repeatedly rated best for Polish grammatical correctness and long-form consistency; described in the Polish tools press as sounding “more like a Pole writing in Polish than an American translating into Polish” (aiport/webyjuice-type comparisons, 2026).
- ChatGPT (GPT-5.x) praised for “naturalność dialogu i wyczucie tonu” — dialogue naturalness and tonal sensitivity (aiport); GPT-5.2 topped Jeleśniański’s overall Polish benchmark.
- Genre fiction competence conceded: AI “radzi sobie z książkami pisanymi według formuły” — handles formula-driven fantasy/crime writing (Siemion, Tygodnik Powszechny; lubimyczytac commentary).
- Bielik (domestic 11B model) beats frontier models on Polish-specific homonyms and linguistic nuance despite losing on general tasks; Polish achieved 88% accuracy in one international multi-language prompting study (GazetaPrawna, 2025) — Polish comprehension is no longer the bottleneck; production quality is.
- Bargielska (poet) found the AI’s output genuinely “wartościowy” (valuable) as collaborative raw material — the strongest pro-quality claim from a literary professional, though for co-creation, not autonomous writing.
Evidence quality & gaps
- Two disjoint evidence pools. Version-precise quality tests (GPT-5.2/5.5, Claude 4.6/Sonnet 5, Gemini 3.1) come from tech/SEO-adjacent outlets (eGospodarka, aiport.pl) testing non-fiction/business writing — competent but not literary authorities. Literary authorities (Gitkiewicz, Rojek, Dąbrowski, Siemion, Piwowarczyk — genuinely credentialed critics and writers in Spider’s Web+, Tygodnik Powszechny, Mały Format, Rzeczpospolita) assess fiction/poetry seriously but say only “ChatGPT”/“AI” without versions.
- No native-critic assessment found of Claude Opus 4.5+, Fable 5, or Gemini 3.x literary Polish specifically. Claude appears in Polish coverage almost exclusively via productivity reviews, not literary criticism.
- Fiction evidence is thin and anecdotal: one author’s style-imitation test (Gitkiewicz) and one AI-collab poetry volume (Kubek na tsunami). No systematic native evaluation of frontier-model Polish fiction (e.g., no STL — Stowarzyszenie Tłumaczy Literatury — statement on frontier-model output quality surfaced; STL-adjacent discussion found is 2024-era and about machine translation generally).
- Academic linguistic error analysis is stale (Mazur/LingVaria, 2024, GPT-4-era); no post-2025 replication found, so “how much better did GPT-5/Claude 4.6 get” rests on commercial testers’ counts.
- Nothing surfaced on dialogue punctuation (pauza dialogowa) or aspect errors specifically; calque and declension complaints are well-attested.
- Paywalls limited depth on Spider’s Web+ and Rzeczpospolita; quotes extracted from accessible portions.