LLM Writing Quality by Language — Polish

Raw research notes for Polish, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-polish.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)


Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)


Summary verdict

Coverage of frontier-model (late-2025/2026) Polish writing quality by native speakers exists but is split between two communities that barely overlap: tech reviewers who run systematic model-vs-model Polish tests (naming GPT-5.2/5.5, Claude 4.6/Sonnet 5, Gemini 3.1 Pro, Bielik, PLLuM), and literary critics/writers who assess “AI” prose seriously but rarely name model versions. The tech-test consensus (2026): frontier models now write largely correct Polish — Claude is repeatedly singled out as making the fewest inflection/syntax errors and “sounding like a Pole rather than an American translating into Polish” — but all models still show systematic declension errors, calques from English, and stylistic drift, attributed to Polish being <0.5% of training data. The literary-community verdict is harsher: reporter Olga Gitkiewicz (July 2026) called ChatGPT’s attempt at her style “grafomania w czystej postaci” — wooden, artificial, logic-free; critics Barbara Rojek and others hold that AI still cannot write good Polish poems or literary prose on its own, while conceding it handles formula genre fiction. The one high-profile fiction data point is Justyna Bargielska’s AI-collaboration poetry volume Kubek na tsunami (Sept 2025), which was reviewed on its merits — with the AI treated as raw material, not author. A genuine gap: no substantive native-critic assessment of Claude Opus 4.5+/Fable 5/Gemini 3-specific literary Polish was found; academic linguistic error analyses (Mazur/UJ) are pre-frontier vintage.

Sources

  1. URL: https://www.egospodarka.pl/196675,ChatGPT-Gemini-Claude-czy-Bielik-ktore-modele-AI-najlepiej-radza-sobie-z-jezykiem-polskim,1,14,1.html Author: Test by Marek Jeleśniański, founder of Oxido (Polish digital agency); tech/business outlet eGospodarka.pl. Date: 2026-03-19. Models: GPT-5.2, Claude 4.6, Gemini 3.1 Pro, Grok 4.2, Qwen 3.5, DeepSeek V3.2, EuroLLM, Mistral, Bielik, PLLuM 8x7B — current frontier, no vintage flag needed. Quotes: “model skupia się na tym, że słowo oznacza istniejący region geograficzny, i ‘dopowiada’ sobie resztę, ignorując błąd ortograficzny” — “the model fixates on the word denoting a real geographic region and ‘fills in’ the rest, ignoring the spelling error” (on the pomoże/Pomorze homonym trap, which Bielik passed and Grok 4.2, Gemini 3.1 Pro, DeepSeek failed). Claim: In a 20-task Polish-language benchmark, US/Chinese frontier models beat Polish models overall (GPT-5.2: 7.66 vs Bielik: 6.38), but all still stumble on proper-name declension, stylistic subtlety, and Polish-specific homonyms where Bielik excels.

  2. URL: https://aiport.pl/testy-i-rankingi/chatgpt-vs-claude-vs-gemini-ktory-najlepiej-pisze-po-polsku/ Author: Piotr Wolniewicz, AIPORT.pl (Polish AI-tools review site — practitioner, not literary critic). Date: 2026-02-11, updated 2026-07-02. Models: GPT-5.5 (Instant/Thinking), Claude Sonnet 5 and Sonnet 4.6, Gemini 3.1 Pro / 3.5 Flash — current frontier. Quotes: “Claude najrzadziej popełnia błędy fleksyjne i składniowe — jego pomyłki to zazwyczaj 1-2 na cały artykuł, a nie 5-10 jak u konkurencji” — “Claude makes inflectional and syntactic errors most rarely — usually 1-2 per article, not 5-10 like the competition.” “Polski stanowi mniej niż 0,5% danych treningowych” — “Polish makes up less than 0.5% of training data.” Claim: Systematic non-fiction writing tests find Claude best on Polish grammatical correctness and long-form stylistic consistency, ChatGPT best on dialogue naturalness and tone (~4% inflection error rate), Gemini prone to tonal “dryfowanie” (drift).

  3. URL: https://spidersweb.pl/plus/2026/07/potwor-ai-zywi-sie-polska-literatura Author: Marek Szymaniak & Rafał Pikuła, Spider’s Web+ (major Polish tech/longform outlet); quality assessment from Olga Gitkiewicz, award-winning Polish reportage author. Date: 2026-07-13. Models: ChatGPT (2026-era, so GPT-5.x — exact version not named; mild vintage-precision flag), plus Claude/Gemini/Grok/Llama discussed re: training data. Quotes: Gitkiewicz on ChatGPT imitating her style: “To, co wypluł, okazało się grafomanią w czystej postaci. Było tak nieznośnie drętwe, sztuczne, dalekie od czyjegokolwiek stylu, a do tego kompletnie pozbawione logiki” — “What it spat out turned out to be graphomania in its purest form. It was so unbearably wooden, artificial, far from anyone’s style, and on top of that completely devoid of logic.” Also: “Tekst… był pełen niedopowiedzeń, błędów logicznych, a nawet językowych” — “The text… was full of vagueness, logical errors, even linguistic ones.” Claim: A professional Polish prose writer who tested a 2026 frontier chatbot on imitating her own reportage style judged the output worthless as literature — stylistically dead and logically incoherent.

  4. URL: https://www.tygodnikpowszechny.pl/sztuczna-inteligencja-juz-pomaga-pisarzom-czy-faktycznie-zle-192223 Author: Marcin Wilkowski; quoted: poet/essayist Tadeusz Dąbrowski, novelist Piotr Siemion, poet Justyna Bargielska. Outlet: Tygodnik Powszechny (leading Polish intellectual weekly). Date: 2025-11-18. Models: ChatGPT, Claude (versions unnamed — late-2025 vintage, borderline in-scope). Quotes (as reported): Dąbrowski — AI lacks the “courage” to deliberately spoil too-perfect writing, which essays require. Siemion — AI can churn out genre fiction fast, but humans keep “fundamental superiority” grounded in “unrepeatable experience and sensitivity”. Bargielska on her ChatGPT collaboration: a relationship “between something backed by knowledge but not feelings, and something backed by feelings but only a fraction of knowledge.” Claim: Established Polish writers concede frontier AI competence at formulaic prose while denying it literary quality in Polish — its output is too smooth, uninhabited by experience.

  5. URL: https://malyformat.com/2025/12/kocur-literacki-7-o-czym-justyna-bargielska-nie-porozmawiala-ze-sztuczna-inteligencja/ (plus review at https://www.rp.pl/plus-minus/art43298201-kubek-na-tsunami-zywotna-zaloba-przegadana-ze-sztuczna-inteligencja) Author: Barbara Rojek, literary critic, editor at “Mały Format” and “Twórczość”; Rzeczpospolita review by Damian Piwowarczyk. Date: Mały Format issue 09-10/2025 (publ. Dec 2025); rp.pl 2025-11-07. Models: Bargielska used ChatGPT (late-2025 vintage). Quotes: Rojek: “na razie sztuczna inteligencja nie pisze sama dobrych wierszy” — “for now, artificial intelligence does not write good poems on its own.” Piwowarczyk credits the book’s power to Bargielska’s own “hipnotyczne i drapieżne obrazki, zderzeń języka” (“hypnotic, predatory images, collisions of language”) confronting lived experience — not to the AI. Claim: In the highest-profile Polish AI-fiction case (Kubek na tsunami, Biuro Literackie 2025), critics judged the AI incapable of good Polish poetry unaided; whatever succeeds in the book is attributed to the human poet’s curation.

  6. URL: https://forsal.pl/lifestyle/technologie/artykuly/9504993,polski-byc-trudna-jezyk-dla-ai-jak-chatgpt-radzi-sobie-z-polszczyzna.html Author: Jolanta Nabiałek (philologist/journalist), reporting research by Rafał Mazur, Polish Philology, Jagiellonian University (related academic paper in LingVaria, UJ). Date: 2024-05-13 — VINTAGE FLAG: pre-frontier ChatGPT (GPT-4 era); useful as a linguistic baseline only. Quotes: Agreement error example: “Wielu pisarzy i poetów… starali się zgłębić” — should be “starało się” (“wielu” requires singular neuter verb); lexical clunkers “popada w mroczne czyny” (“falls into dark deeds”), “namiętnie ambitny” (“passionately ambitious”). Claim: Academic error-count of a ChatGPT matura essay: 76 syntax, 54 lexical, 48 punctuation errors — establishing the pre-frontier failure-mode taxonomy (agreement, punctuation, awkward collocations, repetitions) that 2026 tests measure improvement against.

  7. URL: https://www.dwutygodnik.com/artykul/11956-sztuczna-inteligencja-jest-literatura.html Author: Jerzy Stachowicz, cultural studies scholar, Institute of Polish Culture, University of Warsaw. Outlet: Dwutygodnik. Date: June 2025 — mild vintage flag (pre-late-2025 models). Quote: “Pisanie z czatem to całkiem żmudna praca z tekstem. Nikt tu nie popracuje za nas, a zwłaszcza nie przejmie za nas myślenia.” — “Writing with a chatbot is quite tedious text-work. Nobody here will do the work for us, least of all take over the thinking.” Claim: A UW academic who tested LLM writing in Polish found unedited output generic; acceptable Polish prose emerges only through heavy human prompting and revision.

  8. URL: https://xyz.pl/burza-wokol-bielika-i-pllum-to-jak-porownywac-f-16-do-furgonetki-czy-polskie-ai-naprawde-przegrywa-wywiad/ (not fully fetched; from search results) Author: Interview with Sebastian Kondracki, head of the Bielik project (SpeakLeash). Date: 2026 (responding to the Jeleśniański test). Models: Bielik vs GPT-5.2-class frontier. Claim: Kondracki disputes the benchmark methodology (“like comparing an F-16 to a delivery van”), arguing a Polish-native-trained model of modest size can beat far larger frontier models on Polish-specific reasoning and instruction tasks — the domestic counterpoint in the quality debate.

Failure modes observed

Praise / strengths noted

Evidence quality & gaps