LLM Writing Quality by Language — Italian
Raw research notes for Italian, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-italian.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)
Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)
Summary verdict (4-8 sentences)
Native-Italian assessments converge on a two-level verdict: frontier models now produce grammatically solid, coherent Italian — reviewers testing Claude Opus 4.5 found even a constrained lipogram narrative “coerente e grammaticalmente corretto” — but the prose is widely judged stylistically recognizable and inferior to the models’ English output. The dominant complaint is a “translated-from-English” feel: calques like navigare le complessità, em-dash abuse (an Anglo convention far rarer in Italian typography), rhetorical inflation, and formulaic scaffolding phrases (è importante sottolineare che), catalogued in detail by Italian practitioners. Sociolinguist Vera Gheno (formerly of the Accademia della Crusca) characterizes AI Italian as “plasticoso, poco fragrante, standardizzato.” Among models, Italian reviewers fairly consistently rank Claude first for Italian prose — “più umana, briosa,” less translated-sounding, fewer agreement errors — while GPT-5.2 was seen as a writing regression (acknowledged by Altman, partially fixed in GPT-5.3’s “anti-cringe update”) and Gemini as “consistente ma formulaico.” For fiction, editors and literary translators (Crepaldi, Pareschi, Wu Ming 1, Lipperini) hold that the deficit is not grammatical but embodied: flat conflict, no lived experience, rhythm “massacrata” in literary translation, texts readers recognize as “senz’anima.” Academic corpus work confirms Italian output quality correlates with the thinner Italian share of training data relative to English. Assessments of the very latest models (Fable 5, Gemini 3.x, Opus 4.6) by Italian literary critics specifically are still scarce.
Sources
-
URL: https://massimilianopioli.com/claude-vs-chatgpt-vs-gemini-scrivere/ Author: Massimiliano Pioli — Italian AI strategist and educator (synthesis of community + press assessments). Outlet: personal professional site. Date: 2026-03-30. Models: Claude Sonnet 4.6, GPT-5.2, GPT-5.3 Instant, Gemini (current) — current-vintage. Quotes: Claude output described as “meno AI-smelly” e “più umano” (“less AI-smelly and more human”); Gemini “consistente ma formulaico” (“consistent but formulaic”); community on GPT-5: “Boring. No spark. Feels like a corporate bot.” Claim: Among current frontier models, Claude produces the most natural long-form Italian prose; GPT-5.2 regressed on writing; Gemini is least refined for pure prose.
-
URL: https://github.com/mario-montanari/italiano-scrittura-anti-ai Author: Mario Montanari — Italian practitioner; skill aimed at copywriters, journalists, authors, translators; grounded in stylometry literature. Outlet: GitHub (Claude skill, v1.2.0). Date: updated 2026-07-22. Models: Claude (current) — current-vintage; catalogues LLM-Italian patterns generally. Quotes: catalogued patterns include calchi like “navigare le complessità” and “elevare il tuo business” (calques of “navigate the complexities,” “elevate your business”); “abuso dell’em-dash (—)” (“abuse of the em-dash”); inflation like “si configura come un punto di svolta” (“configures itself as a turning point”); meta-scaffolding “è importante sottolineare che” (“it is important to underline that”). Claim: AI-generated Italian is so consistently marked by identifiable calques, congiuntivo errors, punctuation and register tics that a dedicated countermeasure skill is needed for professional Italian writing.
-
URL: https://www.heraldo.it/2026/04/07/le-parole-sono-atti-vera-gheno-e-il-peso-di-cio-che-diciamo/ (also https://www.seozoom.it/linguaggio-comunicazione-consigli-vera-gheno/) Author: Vera Gheno — sociolinguist, University of Florence; ~20 years collaborator of the Accademia della Crusca. Outlet: Heraldo (interview/lecture coverage). Date: 2026-04-07. Models: generative AI generally (ChatGPT named) — current-era commentary. Quotes: AI language is “plasticoso, poco fragrante, standardizzato” (“plasticky, unfragrant, standardized”); it lacks, via Calvino, the spark “dallo scontro delle parole con circostanze nuove” (“from the clash of words with new circumstances”); AI produces “testi apparentemente sensati ma privi di significato” (“texts apparently sensible but devoid of meaning”). Claim: A leading Italian sociolinguist judges AI Italian competent but impoverished — standardized, plastic prose lacking contextual vitality.
-
URL: https://www.iltascabile.com/scienze/traduzione-e-intelligenza-artificiale/ Author: Riccardo Rinaldi (philosopher, translator) interviewing Silvia Pareschi — premier literary translator into Italian (Franzen, DeLillo, Hemingway, Zadie Smith). Outlet: Il Tascabile (Treccani). Date: 2024-01-18 — vintage flag: pre-frontier (DeepL, early ChatGPT), but Pareschi’s three-tier market thesis is still the reference frame in 2026 Italian debate. Quotes: “Le macchine non pensano, dunque non capiscono il testo, dunque non possono tradurlo” (“Machines don’t think, therefore don’t understand the text, therefore cannot translate it”); machine rendering of Woolf’s rhythm described as “massacrata” (“massacred”); machines “non sanno cogliere la polisemia” (“can’t grasp polysemy”). Claim: Italy’s best-known literary translator holds that machine output fails at polysemy and rhythm, predicting a quality-segmented market (raw MT / post-edited / fully human).
-
URL: https://www.internazionale.it/notizie/alberto-puliafito/2026/01/15/letteratura-intelligenza-artificiale Author: Alberto Puliafito — Italian journalist/director, editor of Slow News, covers AI for Internazionale. Outlet: Internazionale. Date: 2026-01-15. Models: general frontier LLMs — current-vintage. Quotes: Wu Ming 1 (novelist): when AI writes physical sensation it only lines up “probabilità statistiche” (“statistical probabilities”) rather than shared biological reality; Loredana Lipperini (writer, RAI): “un bel testo non coincide con la letteratura” (“a fine text does not coincide with literature”) — literature needs a “corpo… fatto di incontri, scambi, liti, amori” (“a body… made of encounters, exchanges, quarrels, loves”). Claim: Prominent Italian novelists/critics concede AI produces “bei testi” in Italian while denying that fluency amounts to literature.
-
URL: https://www.editorromanzi.it/scrivere-un-romanzo-con-li-a-e-quali-errori-evitare/ Author: Stefania Crepaldi — professional fiction editor (10+ yrs), Salani-published novelist, co-founder of Editor Romanzi and LabScrittore. Outlet: Editor Romanzi blog. Date: 2025-07-11 — vintage flag: mid-2025 models (ChatGPT, Gemini, Claude). Quotes: AI texts “non hanno quel ‘calore’… dettato dalle emozioni di chi scrive” (“lack that ‘warmth’… dictated by the writer’s emotions”); “il pubblico non è ingenuo. Sa riconoscere un testo senz’anima” (“the public isn’t naive. It can recognize a soulless text”); “non è sufficiente mescolare un protagonista timido, un trauma d’infanzia e un colpo di scena” (“it’s not enough to mix a shy protagonist, a childhood trauma and a plot twist”). Claim: A working Italian fiction editor finds AI drafts structurally competent but marked by bland conflict, POV confusion, flattened “positive writing,” and restricted vocabulary.
-
URL: https://riviste.unimi.it/index.php/promoitals/article/view/21990 Author: Francesco Cicero — Università di Napoli “L’Orientale”. Outlet: Italiano LinguaDue (peer-reviewed, Univ. of Milan). Date: 2023-12-15 — vintage flag: ChatGPT/Bard/Bing Chat era; foundational for the data-disparity point. Quotes: “sono emersi anche importanti limiti dal punto di vista informativo e comunicativo” (“important informational and communicative limits also emerged”); Italian-vs-English comparison shows disparities tied to dataset completeness. Claim: Peer-reviewed corpus analysis shows AI Italian reproduces genre traits but with quality deficits correlated to the thinner Italian training data versus English.
-
URL: https://www.lucarosati.it/blog/claude-vs-chatgpt Author: Luca Rosati — Italian information architect and university lecturer, author on content design. Outlet: personal professional blog. Date: 2025-01-08 — vintage flag: Claude 3.5 Sonnet vs ChatGPT o1. Quotes: “la scrittura di Claude è più concisa ma anche più umana, briosa” (“Claude’s writing is more concise but also more human, lively”); “ChatGPT risulta viceversa enciclopedica ma piuttosto piatta” (“ChatGPT is conversely encyclopedic but rather flat”), with “abuso dei punti elenco” (“abuse of bullet points”). Claim: In direct Italian-language comparison, Claude’s prose reads more human and incisive while ChatGPT’s reads flat and listy.
-
URL: https://www.giacomobruno.it/come-usare-claude-in-italiano/ (supporting, promotional-leaning) Author: Giacomo Bruno — Italian publisher (Bruno Editore), 37 bestsellers, AI-in-publishing figure. Outlet: personal site. Date: 2025 (updated; discusses Sonnet 4.5/Opus 4.1). Models: Claude Sonnet 4.5, Opus 4.1 — late-2025 vintage. Quotes: Claude’s Italian “suona meno ‘tradotta dall’inglese’” (“sounds less ‘translated from English’”), with “prosa più naturale, meno errori, registri gestiti meglio” (“more natural prose, fewer errors, better-handled registers”) and fewer “errori di concordanza” (“agreement errors”). Claim: An Italian publishing insider (with commercial interest — weigh accordingly) rates Claude’s Italian as the least translated-sounding among major models.
-
URL: https://iapertutti.it/claude-opus-4-5-recensione/ (supporting) Author: IA per tutti (Italian AI review site, unbylined). Date: late 2025. Models: Claude Opus 4.5 — current-vintage. Quotes: under a no-letter-”a” lipogram constraint the model produced “un racconto coerente e grammaticalmente corretto” (“a coherent and grammatically correct story”). Claim: Opus 4.5 passes demanding formal-constraint tests in Italian at the grammatical level.
Failure modes observed
- Anglicisms and calques — navigare le complessità, elevare il tuo business; syntax that “suona tradotta dall’inglese” (Montanari; Bruno; general consensus).
- Punctuation/typography tics — “abuso dell’em-dash,” a convention imported from English and rare in edited Italian; excess bullet points; anomalous capitalization (Montanari; Rosati).
- Congiuntivo errors and agreement (concordanza) slips — explicitly catalogued as an LLM stumbling block in Italian (Montanari; Bruno on older competitors).
- Rhetorical inflation and scaffolding — si configura come un punto di svolta, è importante sottolineare che, false contrasts (non solo X ma anche Y), forced synonym rotation (Montanari).
- Standardized, “plastic” register — “plasticoso, poco fragrante, standardizzato”; loss of the Calvinian “spark” (Gheno).
- Fiction-specific: bland/absent conflict, flattened “positive writing,” POV confusion, template-assembled plots, “testo senz’anima” (Crepaldi); sensation rendered as “probabilità statistiche” rather than lived experience (Wu Ming 1).
- Literary translation: polysemy missed, rhythm “massacrata” (Pareschi — older-model evidence).
- Structural cause: Italian output quality tracks the smaller Italian share of training corpora vs English (Cicero, Italiano LinguaDue).
- Model-specific: GPT-5.2 writing-quality regression (“testi più piatti e meccanici”), publicly conceded by Altman Jan 2026, partially remedied in GPT-5.3 (Pioli; Italian trade press).
Praise / strengths noted
- Grammatical solidity is now largely a solved problem at the frontier: coherent, correct Italian even under lipogram constraints (Opus 4.5 test).
- Claude repeatedly singled out by Italian reviewers as most natural: “più umana, briosa,” “meno AI-smelly,” better register handling, least translated feel (Rosati, Pioli, Bruno).
- Useful as tool, not author: brainstorming, character sheets, plot hypotheses, synonymy, editing support (Crepaldi); accessibility and participatory-creativity upsides (Puliafito’s sources).
- Even skeptics (Lipperini) concede models can produce “un bel testo” — the dispute is whether that constitutes literature.
Evidence quality & gaps
- Strongest evidence: the Montanari failure-mode catalog (practitioner-built, stylometry-informed, July 2026), Gheno’s 2026 commentary (top-tier sociolinguistic credentials), and Pioli’s March-2026 model comparison (current versions, but a synthesis of others’ judgments rather than original testing).
- Vintage caveats: the most rigorous linguistic analysis (Cicero, peer-reviewed) is 2023-era models; Pareschi’s authoritative translation critique is Jan 2024 (DeepL/early ChatGPT). Their structural arguments (data disparity, embodiment) likely persist but per-model claims are dated.
- Gaps: no located assessment by a named Italian literary critic of Fable 5, Claude Opus 4.6, or Gemini 3.x Italian prose specifically; nothing found on Italian narrative-dialogue punctuation conventions (caporali «» / lineetta) as an AI failure mode; Accademia della Crusca has an explicitly on-point 2025-26 piece (“Sui testi generati dall’intelligenza artificiale: verso un nuovo rapporto tra norma e uso?”, accademiadellacrusca.it/it/contenuti/…/46422) and a tornata on “quanto è naturale l’italiano sintetico,” but both pages returned 403 to fetch — content unverified, flagged as the single most valuable follow-up. The Linkiesta May-2026 piece (“Freddo stil novo”) also 403’d. STRADE/AITI position statements exist on economics/rights but no fetched text assessing frontier-model prose quality per se.
- Bias notes: Bruno has commercial interest in AI-publishing evangelism; iapertutti and Pioli are AI-enthusiast outlets; the literary establishment sources (Wu Ming 1, Lipperini, Pareschi, Gheno) carry the opposite prior. The convergence across both camps on the “correct but standardized/translated-feeling” verdict is therefore fairly robust.