LLM Writing Quality by Language — Spanish
Raw research notes for Spanish, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-spanish.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)
Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)
Summary verdict
Native Spanish-speaking experts converge on a consistent picture: frontier LLMs now produce grammatically near-flawless Spanish — Edmundo Paz Soldán notes that even the subjunctive, long the tell of machine Spanish, is now handled correctly — but the prose reads as denatured: a “lenguaje maquínico” that is correct without being alive, and that systematically drifts toward a neutral, Iberian-or-dubbing-Spanish register flattened of regional variety. The most-cited failure modes are anglicisms and syntactic calques (“translated-from-English” feel), orthotypographic errors against RAE norms (accentuation of “guión”/“sólo”, English-style punctuation, English quotation marks where Spanish fiction demands the raya), lexical monotony and formulaic rhetoric (tricolons, antithetic parallelisms, muletillas like “crucial”, “en el mundo actual”), and poor handling of voseo and dialectal lexicon — academic testing shows models recognize voseo but fail regional vocabulary badly. Literary figures (Paz Soldán, Ortuño, Saldaña París, translators’ associations) frame the deeper failure as loss of personal voice, humor, double meanings and intentional ambiguity, while conceding fluency and structural competence; Jorge Carrión’s concession that GPT-class models “ya redactan mejor que muchos autores de libros superventas” is about the most praise on record. Model comparisons in Spanish-language trade press consistently rank Claude highest for natural literary Spanish (fewest anglicisms, best register control), GPT as slightly anglicized, and Gemini as generic. A significant gap: serious literary critics have not yet published careful evaluations of the newest frontier models (Opus 4.5/4.6, GPT-5.x, Gemini 3) specifically in Spanish — the literary commentary is largely model-agnostic or ChatGPT-4-era, while the model-specific Spanish-quality claims come from lower-credibility tech/trade blogs.
Sources
-
Milenio — “Escritores ante la IA detectan sesgo de género y deshumanización”
- URL: https://www.milenio.com/cultura/escritores-ante-la-ia-detectan-sesgo-de-genero-y-deshumanizacion
- Author: Vicente Gutiérrez (culture journalist); quoting Rosa Montero (Spanish novelist), Edmundo Paz Soldán (Bolivian novelist, professor of Latin American literature, Cornell), Daniel Saldaña París (Mexican writer, Anthropic lawsuit plaintiff), Antonio Ortuño (Mexican novelist), Luna Miguel (Spanish writer)
- Outlet: Milenio (major Mexican daily), 2026-01-30
- Models: ChatGPT; Claude (in lawsuit/training context). Current vintage.
- Key quotes: Paz Soldán on student writing: “el subjuntivo ya es perfecto, pero a costa de perder la voz personal y volverse indistinguibles” (“the subjunctive is now perfect, but at the cost of losing the personal voice and becoming indistinguishable”); “El uso de la IA no solo facilita el plagio, sino que también deshumaniza el estilo” (“AI use not only facilitates plagiarism, it also dehumanizes style”); the flattening is called “lenguaje maquínico” (“machinic language”). Saldaña París reports gender bias in AI translation of his own work.
- Claim: Grammatical Spanish is now solved, but AI-influenced prose homogenizes voice; multiple front-rank Spanish-language authors treat the output as stylistically dead. High credibility.
-
ENTER.CO — “Cuando la IA escribe mal y nosotros le creemos: los errores de ChatGPT según la RAE”
- URL: https://www.enter.co/cultura-digital/el-popurri/cuando-la-ia-escribe-mal-y-nosotros-le-creemos-los-errores-de-chatgpt-segun-la-rae/
- Author: Digna Irene Urrea, tech-culture journalist
- Outlet: ENTER.CO (Colombian tech outlet), 2025-12-23
- Models: ChatGPT (GPT-5.x era). Current vintage.
- Key quotes: “La IA escribe para sonar bien, no siempre para ajustarse a la norma” (“AI writes to sound good, not always to conform to the norm”). Documented RAE-norm violations: writing “guión” (should be “guion”), accenting “solo” and “súper” wrongly, malformed laughter (“jajaja” vs. normative “ja, ja, ja”), period placement inside quotation marks (English convention), dropped accents/question marks in interrogative titles.
- Claim: ChatGPT systematically violates RAE orthography and punctuation norms, and mass exposure is normalizing the errors among Spanish speakers. Medium-high credibility (journalism grounded in RAE norms).
-
El Salto — “La IA no sustituirá a los autores humanos, pero ya está precarizando la industria editorial”
- URL: https://www.elsaltodiario.com/inteligencia-artificial/ia-no-sustituira-autores-humanos-precarizando-industria-editorial
- Author: Jose A. Cano; quoting Marta Sánchez-Nieves (president of ACE Traductores, Spain’s literary translators’ association) and Jorge Carrión (Spanish writer/critic, author of Membrana and Los campos electromagnéticos)
- Outlet: El Salto (Spanish independent daily), 2024-02-17. Vintage flag: GPT-4 era.
- Key quotes: Sánchez-Nieves: “La máquina no capta los dobles sentidos, el contexto o no tiene sentido del humor” (“The machine doesn’t catch double meanings or context, and has no sense of humor”); AI is weaker in languages with “menos datos de los que aprender” (“less data to learn from”) — an explicit Spanish-worse-than-English claim; bridge-translation through English degrades quality. Carrión: “Ya redactan mejor que muchos autores de libros superventas” (“They already draft better than many bestseller authors”).
- Claim: Top-tier literary translation remains beyond AI, but commercial-grade prose is already matched. High credibility.
-
arXiv — “It’s the same but not the same: Do LLMs distinguish Spanish varieties?”
- URL: https://arxiv.org/pdf/2504.20049
- Authors: researchers at Universidad Autónoma de Madrid (Spanish-native academic team), 2025
- Models: GPT-4o, Llama, Gemma, Mistral, Occiglot. Vintage flag: pre-late-2025 models.
- Findings: Models generally identify voseo (the signature Rioplatense morphosyntax), but fail dialect-specific lexical questions at rates up to 80–90% (Mistral, Occiglot); Rioplatense shows the highest inter-model variance; GPT-4o strongest overall.
- Claim: LLMs have shallow, uneven command of Spanish’s regional varieties — recognition without productive mastery. High credibility (peer-track academic).
-
Vasos Comunicantes (ACE Traductores) — CEATL positioning + nº 70 essays
- URLs: https://vasoscomunicantes.ace-traductores.org/2024/09/30/17491/ ; https://vasoscomunicantes.ace-traductores.org/2025/03/11/vasos-comunicantes-70/
- Authors: CEATL (European council of literary translators’ associations) via ACE Traductores; essays incl. Ilya U. Topper, “El código caníbal. Sobre el círculo vicioso de la Inteligencia Artificial”
- Outlet: Vasos Comunicantes, journal of ACE Traductores, 2024–2025
- Substance: professional literary translators’ collective position — post-editing AI literary translation approaches full retranslation in effort; “el código caníbal” argues AI trained on AI output degrades recursively. Related expert voices in the same community: Isabel García Adánez (translator of Thomas Mann into Spanish): “no hay máquina que traduzca bien textos literarios exigentes” (“no machine translates demanding literary texts well”), with risk conceded for entertainment fiction.
- Claim: The professional literary-translation community judges AI Spanish output not fit for demanding literary text without human rewriting. High credibility; partly rights-motivated — noted.
-
Mariana Eguaras editorial consultancy — “Corrección de textos generados por la inteligencia artificial”
- URL: https://marianaeguaras.com/correccion-de-textos-generados-por-la-inteligencia-artificial/
- Author: Víctor J. Sanz (professional corrector/writing coach), on the blog of Mariana Eguaras (veteran Spanish-language editorial consultant)
- Date: 2025-01-07. Models: ChatGPT. Vintage flag: GPT-4o era.
- Key quotes: “los textos generados por la inteligencia artificial no suenan humanos, creíbles, genuinos” (“AI-generated texts don’t sound human, credible, genuine”); sample output showed stilted over-formal phrasing (“Utilidad definida en el marco de funciones avanzadas”); AI creates text “sin tener en cuenta circunstancias que son imprescindibles” (“without accounting for indispensable circumstances” — audience, linguistic variant, context). Professional correction of AI text “often approaches rewriting.”
- Claim: From a working Spanish editor’s chair, AI copy needs near-total rewriting to pass as natural Spanish. Medium-high credibility (practitioner).
-
Genbeta — “No necesitas un detector… 11 señales para descubrir que un escrito está hecho con IA”
- URL: https://www.genbeta.com/a-fondo/no-necesitas-detector-textos-para-ia-11-senales-para-descubrir-que-escrito-esta-hecho-inteligencia-artificial
- Author: Eva R. de Luis; Outlet: Genbeta (Webedia Spain tech journalism), 2024-07-20. Vintage flag: GPT-4o era.
- Substance: catalog of Spanish-language AI tells: muletillas (“crucial”, “en el mundo actual”, “si eres como yo”), marketing verbs (“profundizar”, “descubrir”, “transformar”, “desbloquear”, “dominar”), excessive both-sidesing (“por un lado…, por otro”), tricolon addiction, antithetic parallelisms, absence of anastrophe/catachresis (no stylistic risk-taking), suspiciously perfect orthography.
- Claim: AI Spanish is identifiable by formulaic rhetoric and rhythmic/lexical monotony rather than by errors. Medium credibility (tech journalism, not literary).
-
El Economista — “ChatGPT pone fin al problema de la raya en los textos escritos por IA”
- URL: https://www.eleconomista.es/tecnologia/noticias/13647451/11/25/chatgpt-pone-fin-al-problema-de-la-raya-en-los-textos-escritos-por-ia-las-cosas-se-salieron-de-control.html
- Outlet: El Economista (Spanish business daily), Nov 2025; also El Español: https://www.elespanol.com/elandroidelibre/noticias-y-novedades/20251117/adios-guiones-largos-chatgpt-generar-texto-openai-anuncia-puedes-pedir-ia-no-use/1003744017447_0.html
- Models: ChatGPT / GPT-5.1 era. Current vintage.
- Substance: the em-dash (“guion largo”) became the canonical AI tell in Spanish text too; OpenAI shipped controllability in Nov 2025. Note the Spanish-specific irony documented in writing communities: in Spanish fiction the raya (—) is required for dialogue, yet models default to English quotation-mark dialogue conventions — the failure runs opposite to English (raya overused in expository prose, underused/miscased in novelistic dialogue).
- Claim: Punctuation conventions are a live, model-specific failure surface in Spanish. Medium credibility.
-
Letra Minúscula — “Escritores frente a la IA en 2026”
- URL: https://www.letraminuscula.com/escritores-frente-a-la-ia-en-2026/
- Author: Roberto Augusto (Spanish editor/publisher, Letra Minúscula press), 2026-06-27. Current-vintage, model-agnostic. Informal flag (self-publishing services blog).
- Key quotes: “Corrección no es lo mismo que verdad ni que profundidad” (“Correctness is not the same as truth or depth”); “Le cuesta la ambigüedad intencional, ese terreno incómodo donde habita la mejor literatura” (“It struggles with intentional ambiguity, that uncomfortable terrain where the best literature lives”); “La máquina ha leído sobre el duelo, pero no ha perdido a nadie” (“The machine has read about grief but has never lost anyone”).
- Claim: 2026 models write correct, fast Spanish fiction but default to stereotype, over-resolution and emotional flatness.
-
Revista Inteligencia Artificial — “Modelos LLM en español: comparativa actualizada (2026)”
- URL: https://www.revistainteligenciaartificial.com/modelos-llm-espanol-comparativa/
- Author: staff (“Claude Guerra”), 2026-07-23. Models: GPT-4o/4.1, Claude Opus/Sonnet/Haiku, Gemini 2.5, Llama 4, Mistral Large, ALIA (Spanish government model). Mixed vintage; trade-press credibility (low-medium), informal flag.
- Key quotes: Claude “destaca especialmente en la calidad literaria y naturalidad del texto en español” (“stands out especially in literary quality and naturalness of Spanish text”), fewer anglicisms, better register control; GPT-4o “ocasionalmente produce textos que suenan ligeramente anglicizados” (“occasionally produces slightly anglicized-sounding texts”); Gemini “el estilo tiende a ser más genérico” (“style tends to be more generic”); Mistral strong on European languages.
- Claim: Among-model ranking for Spanish prose: Claude ≥ Mistral > GPT > Gemini. Corroborated by similar informal comparisons at donweb.com, mentoraia.com, aprender21.com.
-
TIC’s en la Web — “ChatGPT en español no es el mismo chatbot”
- URL: https://www.ticweb.es/chatgpt-en-espanol-no-es-el-mismo-chatbot-los-prejuicios-culturales-que-ignoras-antes-de-montarlo-en-tu-web/
- Outlet: Spanish web-tech blog, 2026-07-10. Models: ChatGPT, Claude, Google models. Informal flag.
- Key quotes: “los LLM actuales se entrenan con datos sesgados hacia contextos occidentales” (“current LLMs are trained on data biased toward Western contexts”); “Le pides al modelo textos «optimizados para SEO» en español y obtienes calcos del inglés” (“Ask the model for SEO-optimized text in Spanish and you get calques from English”); with “el voseo rioplatense, el español andino o coloquialismos muy locales, ChatGPT a veces cae en formas neutras o directamente erróneas” (“…sometimes falls into neutral or outright wrong forms”); registers can come out “condescendiente”.
- Claim: The same chatbot is measurably worse in Spanish than English — anglicized syntax, dialect flattening, mis-scaled register — and Spanish benchmarks are mostly translated English benchmarks.
-
Cuadernos Hispanoamericanos — “Crear literatura es más que escribir…”
- URL: https://cuadernoshispanoamericanos.com/crear-literatura-es-mas-que-escribir-o-los-problemas-de-que-la-inteligencia-artificial-no-pasee-ni-sienta-ni-padezca/
- Outlet: Cuadernos Hispanoamericanos (venerable literary review, AECID/Spanish cultural establishment). Fetch blocked by bot-wall; characterization from search excerpts only — verify before citing verbatim.
- Key idea: AI as “zombi literario: parece que escribe, pero lo que hace es redactar” (“a literary zombie: it seems to write, but what it does is draft/compose”) — the redactar/escribir distinction (competent text production vs. literary writing) is the essay’s core.
- Claim: The Spanish literary establishment’s canonical framing: fluent redacción, zero escritura. High-credibility outlet; access-limited.
-
WMagazín — “La inteligencia artificial en el mundo del libro” (context/praise datapoint)
- URL: https://wmagazin.com/relatos/la-inteligencia-artificial-en-el-mundo-del-libro-y-la-literatura-mitos-verdades-temores-fantasmas-dudas-preguntas/
- Author: Winston Manrique Sabogal (ex-El País/Babelia literary journalist), 2023-03-23. Vintage flag: GPT-3.5/4 era.
- Quote: novelist Jean-Baptiste del Amo on style-mimicry tests: “Ha sido asombroso. Textos como si los hubiera escrito yo mismo” (“Astonishing. Texts as if I had written them myself”).
- Claim: Even early models could impress a literary novelist at surface-style imitation.
-
elcastellano.org — “Cómo enfrentar los errores en la traducción de textos por inteligencia artificial”
- URL: https://www.elcastellano.org/news/c%C3%B3mo-enfrentar-los-errores-en-la-traducci%C3%B3n-de-textos-por-inteligencia-artificial
- Outlet: El Castellano (La Página del Idioma Español, long-running language-norm site). Characterization from search excerpt — verify quotes before reuse.
- Substance: recurring calques in AI-translated Spanish: comma before “y” (serial comma), capitalized month names, English-style adjective anteposition — “rasgos que… delatan el origen automatizado del texto” (“traits that give away the text’s automated origin”).
- Claim: A stable checklist of English-interference errors marks AI Spanish.
Related institutional context: RAE’s LEIA project (https://www.rae.es/leia-lengua-espanola-e-inteligencia-artificial) exists precisely because the RAE judges AI-mediated Spanish a live threat — a degradation of Spanish via AI would be “una lesión cultural de primer orden” (“a cultural injury of the first order,” per RAE framing reported by Cervantes Virtual/Milenio) — though LEIA publishes norms/agreements with tech companies rather than model-by-model quality reviews.
Failure modes observed
- Anglicisms and syntactic calques / “translated-from-English” feel — “hacer una decisión”, “está siendo usado por” passives, redundant pronouns, imported metaphors, rigid SVO order, serial comma, capitalized months, adjective anteposition (ENTER.CO/RAE reporting; elcastellano.org; ticweb.es; Revista IA comparativa on GPT: “ligeramente anglicizados”).
- Orthotypography against RAE norms — “guión”/“sólo”/“súper” accent errors, period inside quotes, English quotation marks and mishandled raya in fiction dialogue, “jajaja” formatting (ENTER.CO; El Economista/El Español on the em-dash tell; fanfic/writing-community complaints about quotation-mark dialogue).
- Register mis-scaling and condescension — over-formal or “condescendiente” tone; over-formal stilted copy per working editors (ticweb.es; Mariana Eguaras blog).
- Dialect flattening — “español neutro” — defaults to a dubbing-style neutral or peninsular standard; voseo and Andean/Rioplatense colloquialisms come out neutral or wrong; lexical dialect knowledge fails 80–90% in some models (UAM arXiv study; ticweb.es; Simi/holasimi informal reports).
- Loss of personal voice / homogenization — Paz Soldán’s “el subjuntivo ya es perfecto, pero a costa de perder la voz personal”; “lenguaje maquínico”; “deshumaniza el estilo” (Milenio).
- Formulaic rhetoric and rhythmic/lexical monotony — muletillas (“crucial”, “en el mundo actual”), marketing verbs, tricolons (“la regla de tres”), antithetic parallelisms, “por un lado… por otro” neutrality, repeated adjective pairs (“afilado e inflexible… fría e inflexible”) (Genbeta; Infobae/NYT-sourced detection reporting; Substack practitioner posts).
- No stylistic risk — absence of anastrophe, catachresis, intentional ambiguity; over-resolution (“da demasiadas respuestas”), stereotyped characters (Genbeta; Letra Minúscula).
- Humor, double meanings, cultural context lost — especially in literary translation (Sánchez-Nieves/ACE Traductores; García Adánez; Morábito on sound/rhythm in translation).
- Gender bias in translation into Spanish — grammatical gender choices skew (Saldaña París, Milenio).
- Emotional flatness despite correctness — “frío, sin vida”; “corrección no es verdad ni profundidad”; “ha leído sobre el duelo, pero no ha perdido a nadie” (Letra Minúscula; editor blogs).
Praise / strengths noted
- Grammatical mastery is now essentially complete, including the subjunctive — historically the hardest tell (Paz Soldán, Milenio, 2026).
- “Ya redactan mejor que muchos autores de libros superventas” — commercial-grade prose competence conceded by Jorge Carrión (El Salto, GPT-4 era).
- Style mimicry can astonish professional novelists (del Amo, WMagazín; Pérez-Reverte “Proyecto Maquet” experiments, Zenda/Fundación Telefónica).
- Claude repeatedly singled out in Spanish-language trade press for the most natural, least anglicized literary Spanish and best register control; also praised as a developmental editor for fiction (“puede señalar cuándo una escena explica demasiado”) (Revista IA comparativa; donweb/mentoraia/aprender21 — all informal).
- Speed, structure, coherence over long projects, and usefulness for brainstorming/unblocking acknowledged even by skeptics (Letra Minúscula; Literautas; WMagazín).
- Newer models are “seriously competing with human writers” in publishers’ slush piles — grudging quality acknowledgment implicit in the detection panic (Infobae, Mar 2026 — but note this is a translated NYT piece, not a native assessment).
Evidence quality & gaps
- Recency: Good on the journalism side (Milenio Jan 2026, ENTER.CO Dec 2025, letraminuscula Jun 2026, ticweb Jul 2026, El Economista Nov 2025, Revista IA Jul 2026). The most literary-credentialed commentary (Carrión, ACE Traductores, Cuadernos Hispanoamericanos, WMagazín) is largely 2023–2024, i.e., GPT-3.5/4 vintage — flagged above.
- Expertise: Strong at the top (Paz Soldán, Montero, Ortuño, Saldaña París, ACE Traductores presidency, RAE norms, UAM academics); the model-by-model Spanish-quality rankings, by contrast, come from SEO-flavored trade blogs (donweb, mentoraia, revistainteligenciaartificial) — directionally consistent (Claude > GPT > Gemini for Spanish naturalness) but individually weak; treat as informal consensus, not evidence.
- Gaps: (1) No serious literary critic has yet published a close reading of Opus 4.5/4.6, Fable 5/Mythos 5, GPT-5.x or Gemini 3 prose in Spanish — nothing comparable to the English-language “can Claude write fiction now” discourse; Fable/Mythos 5 and DeepSeek/Qwen/Grok are essentially absent from Spanish sources found. (2) Rigorous same-model Spanish-vs-English comparisons are anecdotal only (Sánchez-Nieves’s data-scarcity point; ticweb’s benchmark-translation complaint) — no controlled study surfaced. (3) Fiction vs. non-fiction is rarely separated explicitly; the implicit pattern is that experts rate AI non-fiction/redacción as serviceable-to-good and AI fiction as correct-but-dead. (4) RAE/LEIA produces norms and agreements, not model quality evaluations. (5) Two sources (Cuadernos Hispanoamericanos, elcastellano.org) were characterized from search excerpts due to fetch blocks, and all quotes passed through an extraction layer — verbatim verification recommended before quoting in a published deliverable.