LLM Writing Quality by Language — Portuguese
Raw research notes for Portuguese, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-portuguese.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)
Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)
Summary verdict
Native-speaker assessment of frontier-LLM Portuguese writing in 2026 is broadly “fluent but flat”: grammatical competence is no longer questioned for Brazilian Portuguese, and Brazilian tech reviewers rank GPT-5-class, Claude 4.x-class, and Gemini 3.x-class models as near-parity in pt-BR fluency, but literary practitioners (novelists, poets, translators, and critics like Paulo Franchetti of Unicamp) consistently deny the output literary value — the recurring charges are genericity, absence of authorial voice, over-explanation, and an unmistakable set of “vícios de linguagem de IA” (formulaic connectors, empty adjectives, symmetrical structure, em-dash overuse) that Brazilian editors now catalogue explicitly. A distinctly Lusophone failure axis is variety confusion: European Portuguese speakers complain that frontier models sound “brasileiro por defeito,” a grievance strong enough that Portugal’s government funded Amália (launched July 2026) specifically to get “maior rigor em todas as dimensões da língua”; conversely, Brazil’s Maritaca claims its Sabiá-4 Thinking beats Opus 4.8, GPT-5.4 and Gemini 3.1 Pro on Brazilian legal writing and slang/irony benchmarks. The em-dash (travessão) panic has jumped languages: Brazilian writers report being falsely accused of AI use for employing normal Portuguese dialogue punctuation, a cultural side-effect documented at length by O Povo. Literary translators in Portugal (2025, vintage-flagged) call AI translation “a morte da literatura” for flattening voice. Direct fiction-quality evaluations of named late-2025/2026 models in Portuguese by literary critics remain scarce; most rigorous comparisons are non-fiction/professional writing, and much of the fiction criticism is either ChatGPT-vintage or model-unspecified.
Sources
-
- Author/credentials: Reported feature quoting Natércia Pontes (novelist, Vida Doçura, Companhia das Letras), Giovana Madalosso (novelist, Todavia), Clarice Freire (poet; Creative Writing professor, Cesar School), Sidney Rocha (novelist; Creative Writing professor, Cesar School)
- Outlet: O Povo (Fortaleza, Brazil) — special report; Date: 2026-05-26
- Models: ChatGPT, Gemini, Claude (current generation, unspecified versions)
- Quotes: Clarice Freire: “A IA roubou meu travessão… me senti incomodada, depois chateada e, por fim, roubada” (“AI stole my em-dash… I felt bothered, then upset, and finally robbed”). Natércia Pontes: “Um travessão é um respiro dentro de uma narrativa que a IA jamais poderá emular” (“An em-dash is a breath within a narrative that AI will never be able to emulate”).
- Claim: Frontier models’ heavy travessão use has stigmatized a core Portuguese dialogue/punctuation resource, causing false AI accusations against human writers, while the writers quoted maintain AI prose lacks the human “respiro” of literature.
-
- Author/credentials: Envox (Brazilian digital-marketing agency; professional pt-BR content editors); Outlet: Envox blog; Date: 2026-02-23
- Models: unspecified current LLMs (2026 output)
- Quotes: Example of redundancy tic: “A IA automatiza tarefas. Ela automatiza processos…” (“AI automates tasks. It automates processes…”); catalogued tics include “além disso” repetition, empty adjectives “fascinante, incrível, essencial,” the “Não é X, é Y” pattern, and “excesso de travessões.”
- Claim: A 12-item native-editor catalogue of recognizable AI tics in Portuguese non-fiction: formulaic connectors, empty intensifiers, symmetric structure, excessive lists, and personality-free neutrality.
-
URL: https://seohappyhour.substack.com/p/vicios-de-linguagem-de-ia-que-podem
- Author/credentials: Rafael Simões, Brazilian SEO/content professional; Outlet: SEO Happy Hour newsletter; Date: 2026-07-15
- Models: Claude, ChatGPT, Gemini, DeepSeek, Kimi (2026 generation)
- Quotes: “um texto totalmente feito por IA fica genérico” (“a text made entirely by AI comes out generic”); tic example: “Essa estratégia não só aumenta o tráfego, mas também melhora as taxas de conversão” (the “não só… mas também” pattern — “not only… but also”).
- Claim: Even across the newest frontier models, fully AI-written pt-BR text is identifiably generic, marked by 13 recurring stylistic tics including “não só… mas também,” em-dash excess, three-item lists, and clichés like “o pulo do gato.”
-
URL: https://jornal.unicamp.br/audio/2026/06/03/a-inteligencia-artificial-vai-substituir-os-escritores/ (companion blog: http://paulofranchetti.blogspot.com/, posts “Excesso de IA” and “Ménage à trois — sobre IA,” June 2026)
- Author/credentials: Paulo Franchetti — literary critic, writer, retired full professor, IEL-Unicamp (PhD USP); Outlet: Jornal da Unicamp (TV interview) + personal blog; Date: 2026-06
- Models: Claude explicitly (he and critic Alcir Pécora ran an experiment feeding Pécora’s essay “A máquina de gêneros” to Claude and having it write a personal response); ChatGPT implicitly
- Quotes: Blog: “Há tempos tenho sofrido de excesso de IA. Não da IA que convoco no meu notebook ou celular…” (“For a while now I’ve been suffering from an excess of AI. Not the AI I summon on my notebook or phone…”). Interview addresses “o fim da autoria?” (“the end of authorship?”).
- Claim: Brazil’s senior literary-critical establishment engages frontier models (including Claude) experimentally and finds the authorship/quality question serious but distinguishes utilitarian competence (translation, revision) from literary authorship, which he does not concede to the machine. (Video interview; full verbatim assessment not extractable from page text.)
-
URL: https://rr.pt/noticia/vida/2025/02/21/da-morte-da-literatura-ao-fim-dos-tradutores-inteligencia-artificial-faz-soar-alarmes/414526/ — VINTAGE FLAG: Feb 2025, pre-late-2025 frontier models
- Author/credentials: Reported piece quoting Sara Veiga (literary translator, Coletivo de Tradutores Literários), Guilherme Pires (translator/editor, Caixa Alta), Clara Capitão (editorial director, Penguin Random House Portugal); Outlet: Renascença (Portugal); Date: 2025-02-21
- Models: none named (generic MT/LLM)
- Quotes: Veiga: “a Inteligência Artificial já está a tomar o lugar de nós, profissionais. A literatura nunca será traduzida da mesma forma” (“AI is already taking the place of us professionals. Literature will never be translated the same way”). Pires: “a máquina não domina essas nuances da voz literária… é a morte da literatura” (“the machine does not master those nuances of literary voice… it is the death of literature”).
- Claim: Portuguese literary translators judge AI-produced literary Portuguese as voice-flattening and artistically non-credible, while conceding the economic displacement is already underway.
-
URL: https://www.alura.com.br/artigos/maritaca-ai-lanca-sabia-4-thinking-modelo-raciocinio-brasileiro (see also https://www.maritaca.ai/en/blog/sabia-4/ and Canaltech coverage)
- Author/credentials: Alura (major Brazilian tech-education platform) reporting Maritaca AI (Unicamp-spun Brazilian lab, Rodrigo Nogueira) benchmarks; Outlet: Alura/Canaltech/Maritaca; Date: 2026 (Sabiá-4 Thinking launch)
- Models: Sabiá-4 Thinking vs Gemini 3.1 Pro, Claude Opus 4.8, GPT-5.4 (current frontier)
- Quotes/data: “No teste de redação jurídica, o Sabiá-4 Thinking alcança 77,7% de acurácia, acima do Gemini 3.1 Pro (75,9%), do Opus 4.8 (74,8%) e do GPT-5.4 (72,8%)” (“On the legal-writing test, Sabiá-4 Thinking reaches 77.7% accuracy, above Gemini 3.1 Pro (75.9%), Opus 4.8 (74.8%) and GPT-5.4 (72.8%)”); also cites the “Sotaques Digitais” benchmark for “gírias, ironias e regionalismos” (slang, irony, regionalisms).
- Claim: A Brazilian lab’s benchmarks show frontier US models remain beatable by a locally-tuned model on Brazilian legal writing and colloquial/regional comprehension — quantitative evidence of a residual pt-BR gap in frontier models.
-
URL: https://androidgeek.pt/portugues-ganha-forca-no-chatgpt-mas-portugal-ainda-exige-nuances (context: https://observador.pt/2026/07/01/governo-lanca-amalia-ia-em-portugues-vai-ser-reforcada-com-15-milhoes-de-euros/ and https://www.publico.pt/2024/11/19/tecnologia/noticia/modelo-linguagem-ia-portugues-chamase-amalia-versao-final-lancada-2026-2112426)
- Author/credentials: Portuguese tech press (AndroidGeek.pt); Observador/Público reporting Portuguese government + FCT/IST Amália project; Date: 2026 (Amália launched 2026-07-01)
- Models: ChatGPT (GPT-5 era), Amália (pt-PT sovereign model)
- Quotes: users want models that don’t sound “brasileiro por defeito” (“Brazilian by default”); Amália promises “maior rigor em todas as dimensões da língua e da cultura” (“greater rigor in all dimensions of language and culture”).
- Claim: European Portuguese speakers judge frontier models’ pt-PT as unreliable — defaulting to Brazilian vocabulary/register — a deficiency significant enough to justify a state-funded European-Portuguese LLM.
-
URL: https://www.metropoles.com/entretenimento/literatura/escritores-nao-temem-chat-gpt-na-literatura-falta-de-sentimento — VINTAGE FLAG: ChatGPT-3.5/4 era (~2023)
- Author/credentials: Reported piece quoting Andréa Del Fuego (novelist, Jabuti winner) among others; Outlet: Metrópoles (Brasília); Date: ~2023
- Models: ChatGPT (early)
- Quotes: paraphrased in coverage: “a máquina ainda não sabe lidar com ambiguidades profundas… o robô não tem angústia” (“the machine still can’t handle deep ambiguity… the robot has no anguish”).
- Claim: Early-vintage baseline: Brazilian literary novelists dismissed ChatGPT-era fiction as emotionally hollow and prone to “textos floreados” (over-flowery, adjective-stuffed prose with forced sonority).
-
- Author/credentials: Crusoé magazine staff (Brazilian newsweekly); Date: 2026 era
- Models: current generation, unspecified
- Claim (one sentence): Brazilian press consensus that AI Portuguese text keeps improving yet remains detectable by experienced human readers and detectors (92–98% claimed rates) — i.e., still stylistically marked.
Failure modes observed
- Catalogued “vícios de linguagem de IA” (Envox, SEO Happy Hour — both pt-BR, 2026): repetitive connectors (“além disso,” “por outro lado,” “portanto,” “no entanto”), the “não só… mas também” frame, “Não é X, é Y” antithesis tic, empty adjectives (“fascinante,” “incrível,” “essencial”), vague grandiose vocabulary (“jornada,” “essência,” “universo”), exactly-three-item lists, excessive bolding/lists over flowing prose, identical paragraph openings, redundant restatement, personality-free neutrality.
- Travessão/em-dash overuse — notable because in Portuguese the travessão is standard dialogue punctuation; AI overuse in expository text has stigmatized it and triggered false accusations against human authors (O Povo, Coletiva.net, 2026).
- Variety confusion (pt-PT vs pt-BR): frontier models default to Brazilian lexicon/register for European users (“configurações” vs “definições”; “brasileiro por defeito”) — the core motivation for Portugal’s Amália and startups like Anália.pt; symmetrically, Maritaca’s “Sotaques Digitais” benchmark shows frontier models trail on Brazilian slang, irony, regionalisms.
- Literary voice deficit: flattening of authorial voice in translation (Pires: “só identifica símbolos” — “it only identifies symbols”); no handling of “ambiguidades profundas”; over-explaining rather than letting readers infer; less narrative diversity (fewer subplots, scenes, dialogue); “textos floreados” — forced-sounding adjective/metaphor pileups (early-vintage but repeated in 2026 criticism).
- Genericity even in newest models: “um texto totalmente feito por IA fica genérico” (Simões, July 2026, across Claude/GPT/Gemini/DeepSeek/Kimi).
- Style imitation failure: ChatGPT-lineage models can’t sustain a requested author style in Portuguese revision work (PUC-Minas Cadernos CESPUC study — academic, revision-focused).
Praise / strengths noted
- pt-BR grammatical fluency at frontier level is treated as solved; TechTudo (Mar 2026) rates Claude “texto bem escrito e bem estruturado, mais informativo,” ChatGPT “boa redação” with strong fluency and regional expression handling, Gemini strong on creative headlines.
- Brazilian comparative testing (2026) puts GPT-5.x marginally first on pt-BR fluency/regionalisms with Claude 4.x noted as “mais natural com nuances da língua”; honors roughly split across the big three on real pt-BR tasks.
- Academic/professional non-fiction: Claude praised as best for formal academic Portuguese (coherence over long sections, low repetition — tesify.pt, though it cites older Claude 3.5 Sonnet, vintage flag).
- LLMs seen as a “ponto de virada” for literary translation because they accept customization (rhyme scheme, form) — UNESP repository study; utility for revision/translation conceded even by hostile critics (Franchetti uses AI for “tradução, revisão de texto”).
- Sabiá-4 Thinking (Brazilian) matching/beating Opus 4.8, GPT-5.4, Gemini 3.1 Pro on Brazilian legal writing shows locally-tuned Portuguese quality is achievable and now competitive at far lower cost (R$206 vs R$590 benchmark-suite cost cited).
Evidence quality & gaps
- Strongest evidence: O Povo’s May 2026 feature (multiple credentialed Brazilian literary voices, current models); the two 2026 pt-BR “AI tics” catalogues with concrete linguistic examples; Maritaca’s quantitative benchmarks naming Opus 4.8/GPT-5.4/Gemini 3.1 Pro; the Amália project as institutional evidence of the pt-PT deficiency.
- Gaps: Almost no published native-critic evaluation of named late-2025/2026 frontier models writing original fiction in Portuguese — literary criticism is mostly model-agnostic or ChatGPT-vintage, while model-specific 2026 comparisons are non-fiction/marketing/academic. Franchetti’s Claude experiment is the closest critic-plus-named-frontier-model datapoint, but the substantive verdict sits in video/blog posts not fully extractable here. Nexo Jornal’s June 2026 piece (“Autoria e qualidade…”) is paywalled. No mesóclise-specific commentary surfaced. Several comparison sites (tesify, swen.ia.br, jenova.ai, chatgptbrasil.com.br) are SEO-adjacent and of low evidentiary weight; one (chatgptbrasil) 404’d on fetch.
- Vintage flags: Renascença translators piece (Feb 2025), Metrópoles writers piece (~2023), tesify’s Claude 3.5 Sonnet claim — all pre-date the late-2025 frontier and are used as baseline only.