LLM Writing Quality by Language — French
Raw research notes for French, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-french.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)
Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)
Summary verdict
Native French assessments split into two camps. Practitioner benchmarks from 2026 (testing GPT-5.4/5.5, Claude Opus 4.6/4.7, Gemini 3.x, Mistral Large 3) consistently rank Claude and Mistral as writing the most natural French, with ChatGPT and Gemini still producing “subtle anglicisms” and calqued constructions — French output is judged serviceable-to-good for non-fiction (reports, emails, journalistic formats) but stylistically formulaic. Literary voices are harsher on fiction: linguists and writing coaches describe AI French fiction as “fade” (bland), cliché-laden, and lacking intention, though the most spectacular counter-example is Benoît Raphaël’s 2025 Nouvel Obs experiment in which a Claude-written noir short story was judged equal or superior to Goncourt winner Hervé Le Tellier’s — by evaluators and, grudgingly, by Le Tellier himself (“better written than 50% of what’s published” in France). The best-documented failure modes are typographic and syntactic Englishness: overuse of the tiret cadratin (now a nationally recognized ChatGPT “tell,” the subject of a Le Monde column in April 2026 after a PM Lecornu tweet), misused present participles, “Ce n’est pas X, c’est Y” binary constructions, hollow adjectives (crucial, fascinant), and stock connectors (en outre, par ailleurs). Institutional voices (ATLF translators, the Ministry of Culture’s March 2026 report to Parliament) frame the problem structurally: models trained on anglophone corpora risk impoverishing French itself. Notably absent: any prominent native assessment of Fable 5 or Opus 4.5+ specifically for French literary fiction; the 2026 model-specific evidence is mostly professional/marketing prose.
Sources
1. Génération IA (Flint Media) — “Comment j’ai défié un prix Goncourt avec l’intelligence artificielle”
- URL: https://generationia.flint.media/p/comment-j-ai-defie-un-prix-goncourt-avec-l-intelligence-artificielle-claude-herve-le-tellier-chatgpt
- Author: Benoît Raphaël (French journalist, founder of Flint Media, AI-literacy trainer; created lepost.fr/Le Plus du Nouvel Obs), with AI engineer Thomas Mahier. Native French speaker.
- Date: March 23, 2025 (Nouvel Obs original March 13, 2025). VINTAGE FLAG: pre-late-2025 — Claude 3.x-era, ChatGPT/o1/Grok 3.
- Models: Claude (primary generator; ChatGPT judged less convincing), o1 and Grok 3 as evaluators.
- Key quotes: Le Tellier reading the AI text: « Oh la vache ! » (“Holy cow!”). Asked whether it was better written than 50% of contemporary French publications: « Oui, bien sûr » (“Yes, of course”). Evaluators credited the AI text with a « style riche et fluide » (“rich and fluid style”) and rated Le Tellier’s own entry « plus conventionnel, moins profond » (“more conventional, less deep”).
- Claim: Under blind-style comparison, Claude-generated French noir fiction matched or beat a Goncourt winner on a 3,000-character exercise, per both AI judges and the Goncourt winner’s own admission.
2. The Conversation France — “Comment « dé-IA-iser » nos écrits pour éviter la disparition des particularités des langues ?”
- URL: https://theconversation.com/comment-de-ia-iser-nos-ecrits-pour-eviter-la-disparition-des-particularites-des-langues-281811
- Author: Maria Mercanti-Guérin, maîtresse de conférences (HDR), IAE Paris–Sorbonne. French academic.
- Date: July 7, 2026. Current.
- Models: ChatGPT, Claude, Gemini, DeepSeek (generic frontier-class discussion).
- Key quotes: catalogs AI-French tics — openings like « dans le monde actuel en constante évolution » (“in today’s ever-evolving world”), hollow adjectives (« fascinant », « crucial »), abstract verbs (« mettre en œuvre », « s’articuler »), connectors (« en outre », « par ailleurs »), closers (« en somme », « au final »). Attributes false-ringing “translation tics” to massively anglophone training corpora; notes “GPT words” leaking into spontaneous human French.
- Claim: AI-generated French carries an identifiable de-particularized register calqued on English, threatening the distinctiveness of French itself; proposes “dé-IA-isation” (anti-slop sampling, fine-tuning on stylistically rich French corpora, human writer-trainers).
3. Raphaël Doan (Substack “Technoclassicisme”) — “Comment écrire à l’ère de l’IA”
- URL: https://raphaeldoan.substack.com/p/comment-ecrire-a-lere-de-lia
- Author: Raphaël Doan — French essayist and haut fonctionnaire, agrégé de lettres classiques (ENS/ENA), author of Si Rome n’avait pas chuté and of the AI-assisted book Si les Grecs avaient connu la vapeur; one of France’s most AI-literate literary writers.
- Date: June 14, 2025. VINTAGE FLAG: Claude 4 / GPT-4o / o3 era — just before scope window.
- Models: Claude 4, ChatGPT-4o, o3, earlier GPT-3/GPT-4 samples.
- Key quotes: identifies ChatGPT-French pathologies — the binary tic « Ce n’est pas X. C’est Y », short punchy sentences mimicking American marketing copy, gratuitous adjectives (« captivante », « passionnante »), misapplied participial constructions he explicitly labels anglicisms (in French, absolute participles read as causal, not chronological as in English), em-dash overuse. On an early GPT-3 fiction sample: « Ce n’est quand même pas mal, et donne presque envie de connaître la suite » (“It’s really not bad, and almost makes you want to know what happens next”).
- Claim: A trained French literary stylist finds Claude the most human-sounding of the frontier models in French, while ChatGPT’s French remains structurally American — recognizable at the level of syntax, not just vocabulary.
4. LesAstucesIA — “ChatGPT vs Claude vs Gemini 2026 : le verdict après 6 mois de test” (benchmark français)
- URL: https://www.lesastucesia.com/blog/actualites/chatgpt-vs-claude-vs-gemini-benchmark-francais-2026
- Author: Ugo Lazzari, French AI-tools reviewer. Practitioner blog, not literary press.
- Date: April 9, 2026, updated June 2026. Current.
- Models: GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro. In-scope versions.
- Key quotes: Claude 9/10 for French writing — « le ton le plus naturel en français » (“the most natural tone in French”), sounds most human; ChatGPT 8/10 — correct « mais parfois générique » (“but sometimes generic”), tics like « N’hésitez pas à… »; Gemini 7/10 — « correct mais plus scolaire » (“correct but more schoolbook-like”). Fiction subtest: ChatGPT most creative (9/10), Claude « bien écrite mais plus littéraire » (8/10), Gemini « slogans génériques, nouvelle plate » (“generic slogans, flat short story,” 6/10).
- Claim: Among current frontier models, Claude produces the most natural French prose while ChatGPT wins on creative invention and Gemini trails on both naturalness and fiction.
5. NewsIA — “ChatGPT vs Claude vs Gemini vs Mistral : le comparatif 2026”
- URL: https://newsia.fr/guides/comparatif-chatgpt-claude-gemini-mistral-2026
- Author: Driss Redouane. French practitioner outlet.
- Date: April 12, 2026, updated May 28, 2026. Current.
- Models: Claude Opus 4.7, GPT-5.5, Gemini 3 Pro, Mistral Large 3. In-scope.
- Key quotes: « Mistral écrit le mieux quand la longueur reste raisonnable. Sa langue est naturelle, sans tournures étranges. » (“Mistral writes best when length stays reasonable. Its language is natural, without strange turns of phrase.”) — « ChatGPT et Gemini souffrent encore d’anglicismes subtils. » (“ChatGPT and Gemini still suffer from subtle anglicisms.”) Mistral 9.5/10 on formal administrative letters vs Claude 8.5; « Mistral excelle sur les formats journalistiques ».
- Claim: The French-built Mistral Large 3 beats the American frontier models on idiomatic French for short-to-medium non-fiction, though the comparison covers no literary writing.
6. RTBF — “L’intelligence artificielle peut-elle vraiment écrire un bon roman ?”
- URL: https://www.rtbf.be/article/l-intelligence-artificielle-peut-elle-vraiment-ecrire-un-bon-roman-11751669
- Author: Mathilde Rigaud, interviewing Louis Escouflaire, PhD in linguistics (UCLouvain), researcher on AI text detection. Native francophone (Belgian) linguist — strong credential.
- Date: July 3, 2026. Current, but models discussed generically (“ChatGPT” as a class).
- Key quotes: AI fiction texts are « un petit peu fades » (“a bit bland”) and predictable, recycling « tous les clichés et les points de vue… de tous les auteurs » (“all the clichés and viewpoints… of all authors”); « l’intention et la passion derrière l’écriture humaine est difficilement réplicable » (“the intention and passion behind human writing is hard to replicate”). Concedes AI can « créer un scénario, voire entretenir un suspense » (“build a plot, even sustain suspense”).
- Claim: A francophone computational linguist judges AI-generated French fiction technically competent at plot mechanics but stylistically bland, clichéd, and culturally homogenizing.
7. Le Monde (chronique Guillemette Faure) — the tiret cadratin as ChatGPT tell (via secondary coverage)
- Primary: Le Monde, April 2026 (paywalled; not fetched directly). Secondary: https://www.digido.ma/blog/quand-les-tirets-trahissent-l-usage-de-chatgpt-il-n-a-meme-pas-fait-l-effort-de-retirer-le-tiret-cadratin and https://siecledigital.fr/2025/05/12/ce-detail-typique-de-chatgpt-permet-de-savoir-si-vous-avez-utilise-chatgpt-et-voici-comment-leviter/
- Author: Guillemette Faure, longtime Le Monde columnist (« Le dico de Guillemette Faure »). Native French journalist at the paper of record.
- Date: April 2026 (column); Siècle Digital piece May 2025.
- Models: ChatGPT (class-level).
- Key quotes: column title/hook: « Il n’a même pas fait l’effort de retirer le tiret cadratin » (“He didn’t even bother to remove the em dash”) — written after a tweet by PM Sébastien Lecornu was accused of being ChatGPT-written. Secondary coverage: the em dash is rare in everyday French writing, normally reserved for literary dialogue; where AI writes « — et même amplifier — », French typography expects commas or tirets moyens.
- Claim: Em-dash overuse violating French typographic convention has become the mainstream-recognized signature of AI-written French, prominent enough for a Le Monde column and a political mini-scandal.
8. ATLF (Association des traducteurs littéraires de France) — “Non, l’intelligence artificielle ne remplacera pas les traducteurs… mais elle détruit leur métier !” + Tribune « Stoppons cette dégringolade de la pensée »
- URLs: https://atlf.org/non-lintelligence-artificielle-ne-remplacera-pas-les-traducteurs-et-traductrices-mais-elle-detruit-leur-metier/ ; https://atlf.org/tribune/ ; https://actualitte.com/article/110810/tribunes/intelligence-artificielle-et-traduction-litteraire-exiger-la-transparence
- Author: ATLF board (professional literary translators — the most style-sensitive native readers of AI French output, via post-editing work).
- Date: Sept 13, 2024 (VINTAGE FLAG for the quality claims), with continuing 2025–2026 activity (June 19, 2026 workshop « L’intelligence artificielle ou comment s’en protéger »).
- Models: none named — neural MT and generative AI as a class.
- Key quotes: post-edited output is of « qualité médiocre » (“mediocre quality”) and the practice is one « qu’on ne saurait apparenter à de la traduction » (“that cannot be considered translation”); promised time savings are nil and the work is experienced as alienating.
- Claim: France’s literary translators collectively judge machine-generated French literary text mediocre enough that “fixing” it costs as much as translating, and frame AI as destroying craft rather than matching it — an economic/ethical position paper more than a linguistic analysis.
9. ActuaLitté — “Anglais, IA, inégalités : les défis majeurs qui menacent le français” (covering the Rapport au Parlement sur la langue française, March 2026)
- URL: https://actualitte.com/article/130911/ressources/anglais-ia-inegalites-les-defis-majeurs-qui-menacent-le-francais
- Author: Hocine Bouhadjera, ActuaLitté journalist; underlying report by the DGLFLF (Ministry of Culture).
- Date: April 24, 2026 (report March 2026). Current; institutional.
- Models: none named commercially; frontier LLMs as a class; French counter-initiatives ALT-EDIC (€88.3M) and Pleias/Common Corpus (266bn French tokens).
- Key quotes: « les grands modèles de langage sont principalement entraînés sur des corpus anglophones » (“large language models are trained mainly on anglophone corpora”) and thus embed « des références culturelles, des logiques de raisonnement et des priorités qui ne sont pas neutres » (“cultural references, reasoning patterns and priorities that are not neutral”); poorly designed models risk impoverishing the language or defaulting to English.
- Claim: The French state’s official linguistic body treats frontier-LLM French as structurally anglo-centric and a sovereignty problem, justifying public investment in French-native corpora and models.
10. Publier son Livre — “J’ai testé des IA pour écrire un livre”
- URL: https://publiersonlivre.fr/diagnostic-accompagnement-litteraire/ia-ecrire-livre/
- Author: unnamed staff of Publier son Livre (French manuscript-coaching/self-publishing service — professional French readers of amateur prose).
- Date: February 2024. VINTAGE FLAG: GPT-4/Claude 2-era.
- Models: ChatGPT, Gemini, Claude (of that era).
- Key quotes: « Les IA ont des tics de langage. Les mots “essentiel”, “crucial” ou “défi” sont trop souvent utilisés » (“AIs have verbal tics; ‘essential’, ‘crucial’, ‘challenge’ are overused”); « c’est toujours trop court, et la rédaction sonne faux » (“it’s always too short, and the writing rings false”); prescription: « moins d’écriture, plus de réécriture » (“less writing, more rewriting”).
- Claim: French writing coaches found unsupervised AI book prose in French recognizably fake and lexically repetitive, workable only as a heavily rewritten draft (early-model finding; consistent with later tic catalogs).
Failure modes observed
- Typography/dialogue conventions: systematic tiret cadratin (em dash) overuse where French expects commas, parentheses, or tirets moyens — em dashes being reserved in French mainly for literary dialogue lines; now the #1 popular “tell” (Le Monde, Siècle Digital, multiple detection guides). French guillemets/non-breaking-space handling also flagged as a detection signal. No source found complaining specifically that models botch roman dialogue layout (tirets vs guillemets) in fiction output — the complaints are about em dashes in expository prose.
- Anglicisms and calques: “subtle anglicisms” persisting in GPT-5.x and Gemini 3.x French (NewsIA, LesAstucesIA); calqued turns of phrase in longer texts; “GPT words” bleeding into human usage (The Conversation).
- Syntactic Englishness: misused present/absolute participles (English chronological reading vs French causal reading — Doan); short punchy “American marketing” sentence rhythm; « Ce n’est pas X. C’est Y » binarism.
- Lexical slop register: « crucial », « fascinant », « essentiel », « défi »; connectors « en outre », « par ailleurs »; openers « dans un monde en constante évolution »; closers « en somme », « au final »; ChatGPT’s « N’hésitez pas à… ».
- Fiction-specific: blandness (« fades »), cliché aggregation across all authors, absence of intention/passion, flat short stories from Gemini, Western/anglo cultural homogenization (RTBF/Escouflaire).
- Translation/post-editing: output « de qualité médiocre » requiring rework equal to translating from scratch (ATLF).
- Not found in sources: specific documented complaints about subjunctive/concordance des temps errors or tu/vous register failures in frontier models — the register criticism is about stylistic register (scolaire, générique, marketing), not pronoun choice. Absence of evidence, not evidence of absence.
Praise / strengths noted
- Claude repeatedly singled out by French testers as most natural in French: varied sentence length, flowing transitions, fewer calques, holds tone/register instructions (LesAstucesIA on Opus 4.6: « le ton le plus naturel en français »; Doan on Claude 4: more human than OpenAI/Google; madame-tuto.fr and oia.fr comparisons concur).
- Mistral Large 3 praised as the best pure French stylist for short/medium non-fiction: « sa langue est naturelle, sans tournures étranges », excels at journalistic and administrative formats (NewsIA).
- Fiction ceiling demonstrably high: the Raphaël/Le Tellier experiment (Claude, 2025) produced French noir prose a Goncourt winner conceded was better written than half of what gets published in France — the strongest native-speaker praise on record, though on a short constrained exercise.
- Plot mechanics: even skeptics concede AI can build scenario and sustain suspense in French (Escouflaire).
- Non-fiction: structured argument, essay scaffolding, and professional prose in French are broadly rated good-to-excellent across 2026 practitioner tests; the criticism is genericness, not incorrectness.
Evidence quality & gaps
- Strongest evidence: the Le Tellier experiment (named Goncourt winner, published protocol, on-record quotes — but March 2025, pre-scope models, short-form only); The Conversation (academic, July 2026); RTBF/Escouflaire (credentialed francophone linguist, July 2026); Le Monde column (mainstream literary-journalistic validation of the typography failure mode).
- Model-specific 2026 evidence is practitioner-blog grade: LesAstucesIA and NewsIA name exact versions (GPT-5.4/5.5, Opus 4.6/4.7, Gemini 3.x, Mistral Large 3) and test in French, but are SEO-adjacent tool-review sites with informal methodology; their consistent Claude/Mistral > GPT/Gemini ranking for French naturalness is corroborating but not authoritative. Some searched summaries referenced even later models (Opus 4.8, GPT-5.5) — treat exact-version claims from these sites with caution.
- Gaps: (1) No prominent French literary critic (Le Monde des Livres, France Culture, En attendant Nadeau) found publishing a serious stylistic assessment of late-2025/2026 frontier models’ French fiction — coverage is either older-model or economic/ethical; (2) no native assessment of Fable 5 or Opus 4.5+ French fiction specifically; (3) ATLF/ATLAS positions are rich on labor/ethics, thin on concrete linguistic autopsy; (4) no found French source directly comparing a model’s French output quality against its own English output; (5) fiction vs non-fiction comparison exists only in the LesAstucesIA benchmark and anecdotal coach commentary. A targeted follow-up on France Culture archives and Le Monde des Livres (paywalled) would likely fill gap (1).