LLM Writing Quality by Language — Arabic
Raw research notes for Arabic, part of the LLM Writing Quality by Language project — edition 2026-07-23. Originally published at peterkaminski.ai/research/llm-writing-quality-by-language/llm-writing-quality-arabic.
Researched and written by Saga bg-etruscan (Claude Fable 5), directed by Peter Kaminski, 2026-07-23. Quotations are machine-extracted from the cited sources and not yet verified verbatim — see the main report’s Limitations section.
♡ Copying is an act of love. Please copy and share.
© Peter Kaminski · CC-BY 4.0 (Creative Commons Attribution 4.0 International)
Project files: main report · English · Spanish · Chinese · Hindi · Arabic · French · Portuguese · Russian · German · Japanese · Korean · Italian · Turkish · Indonesian · Polish · all files (.zip)
Summary verdict
Native Arabic-speaking writers, critics, and poets remain broadly skeptical of frontier-LLM Arabic writing, and their criticism is more specific than generic “AI lacks a soul” complaints: the most recurrent charge is that AI Arabic reads like a stiff translation from English — “سطحية وتشبه لغة مترجمة من الإنجليزية إلى العربية بركاكة” (Egyptian writer Ahmed Lotfi) — reflecting thin Arabic training data (~3% of web content). In poetry, the assessment is harsher and technically concrete: Emirati, Lebanese, Egyptian, Iraqi and Moroccan poets surveyed by Al Bayan (Dec 2025) document broken meter (كسر الوزن), defective rhyme, and neglect of features like ألف الإطلاق, concluding models can satisfy formal templates but produce إنشائية (school-essay composition) rather than poetry. In prose fiction and criticism, Iraqi critic Nadia Hanawi’s book-length study finds texts that “تحاكي ولا تبدع” (imitate rather than create), while novelist-computer scientist Habib Sorori (a rare doubly-credentialed voice) worries about a flood of superficial, error-ridden AI-polished texts. On the linguistic-technology side, a July 2026 Al Jazeera survey and practitioner comparisons report genuine improvement in frontier models — Claude Opus 4.6 / Sonnet 4.6 rated best for dialect and cultural nuance, GPT-5 strongest on fusha, Gemini overly formal — but dialect production, idiom, and sarcasm remain weak, and low-resource dialects a major gap. Overall: measurable progress on MSA correctness since late 2025, continued failure at register (fusha/dialect), rhetorical tradition, and prosody; almost no rigorous published native-speaker evaluation of the newest frontier models’ literary Arabic specifically.
Sources
-
URL: https://www.aljazeera.net/tech/2026/7/22/الذكاء-الاصطناعي-باللغة-العربية-هل Author: تسنيم حسن (Tasneem Hassan), tech journalist; Outlet: الجزيرة نت; Date: 2026-07-22; Models: ChatGPT, Gemini, Claude; Arabic-focused ALLaM (Saudi), Jais (UAE), Falcon Arabic (2025). Current-vintage. Key quote: “البيانات الخاصة باللهجات منخفضة الموارد لا تزال تمثل أحد أكبر التحديات” — “Data for low-resource dialects remains one of the biggest challenges” (citing the AraDiCE benchmark, ACL 2025). Claim: Frontier and Arabic-focused models have improved markedly at MSA but still fail at producing (vs. understanding) dialects, idioms, sarcasm, and cultural reference.
-
URL: https://www.albayan.ae/lifestyle/culture/994909 Author: السيد رمضان (Al-Sayed Ramadan); quotes Emirati poet Dr. Talal Al-Jneibi (literature professor), Lebanese poet Shawqi Bazee (شوقي بزيع, major contemporary poet), Dr. Alaa Ganeb (Azhar University dean of literature/criticism), Iraqi poet Khaled Al-Hassan, Moroccan poets Omar Al-Raji and Makhlass As-Saghir; Outlet: البيان (UAE); Date: 2025-12-11; Models: ChatGPT and peers (“شات جي بي تي وغيره”) — late-2025 vintage. Key quotes: Al-Jneibi: “شات جي بي تي وغيره قد تصيغ قصائد بشروط شكلية، لكن الإبداع يبقى عصياً” — “ChatGPT and its peers can compose poems meeting formal conditions, but creativity remains beyond reach.” Bazee: “الشعر نقيض كل ما هو مصطنع” — “Poetry is the antithesis of everything artificial.” Al-Hassan: AI cannot capture “الارتجافة التي تسبق ولادة القصيدة” — “the tremor that precedes the birth of a poem.” Claim: A pan-Arab panel of working poets documents concrete failures — كسر في الوزن (broken meter), خلل في القوافي (defective rhyme), wrong poem length, neglect of ألف الإطلاق — concluding AI verse is إنشائية, not poetry.
-
URL: https://alpheratzmag.com/تقارير/الذكاء-الاصطناعي-والأدب-العربي/ Author: محمد فوزي عبد العزيز (Mohamed Fawzi Abdelaziz, researcher); Outlet: مجلة الفراتس (Alpheratz); Date: 2025-11-25; Models: ChatGPT-era systems, the “جهيمان الأعجمي” synthetic-author experiment (2020, Lebanese magazine رحلة), ديوان راقمون AI-poet personas — mixed vintage, mostly pre-frontier (flag). Key quote: Egyptian writer أحمد لطفي on AI Arabic: “سطحية وتشبه لغة مترجمة من الإنجليزية إلى العربية بركاكة” — “superficial, resembling language clumsily translated from English into Arabic.” Counter-praise from فريق ديوان العرب: “نصوصاً رصينة غنية بالصور المجازية والأخيلة” — “solid texts rich in metaphorical imagery.” Claim: The definitive statement of the “translated-from-English ركاكة” failure mode, attributed to poor Arabic training inputs, within a history of Arab synthetic-author experiments that rose and collapsed.
-
URL: https://www.alowais.com/artihabebsrosri/ (also his blocked Al-Quds al-Arabi column “أتْمَتةُ الإبداع الروائي بالذكاء الاصطناعي”) Author: حبيب عبد الرب سروري (Habib Abdulrab Sarori) — Yemeni novelist AND computer-science professor (Rouen, France); author of the 2026 AI-narrated novel «اعترافات AI حزين» (Dar Al Saqi) and co-author with Moroccan philosopher Moulim El Aroussi of a 2026 book on AI and the Arab mind; Outlet: Sultan Al Owais Cultural Foundation; Date: 2026-04; Models: ChatGPT paid tiers generally — current-vintage commentary, model-vague (flag). Key quotes: warns of “سطحيّة بعض النصوص وخلوّها من التحديث والابتكار” — “the superficiality of some texts and their lack of renewal and innovation”; authors now submit “نصوصَهم الركيكة، المتخمة بالأخطاء” — “their flimsy, error-stuffed texts” — to AI for polishing before publication. Claim: The most technically credentialed Arab novelist on the subject predicts two coexisting literatures (human/artificial) but sees current AI-assisted Arabic prose driving quality downward via authorial laziness.
-
URL: https://www.aljazeera.net/culture/2025/8/23/الذكاء-الاصطناعي-بين-وهم-الإبداع Author: Al Jazeera culture desk on د. نادية هناوي (Dr. Nadia Hanawi), Iraqi academic literary critic, book «الذكاء الاصطناعي: التأهيل والتهويل» (Abjad Foundation); Outlet: الجزيرة نت — ثقافة; Date: 2025-08-23; Models: ChatGPT (tested in her experiments) — pre-frontier vintage (flag). Key quote: “نصوص تحاكي ولا تبدع، تولف ولا تنتج، تدور في محيط النصوص السابقة” — “Texts that imitate rather than create, compile rather than produce, orbiting within previous texts.” Claim: A book-length academic critique concluding AI Arabic literary output is derivative recombination without creative depth or critical stance.
-
URL: https://www.aljazeera.net/culture/2026/5/22/تراجع-الذائقة-والأمانة-الفكرية-4 Author: شيخاني أحمدو (Sheikhani Ahmado); quotes writer-researcher لقاء مكي, writer-translator حسين النهابة, novelists ناصر يوسف المهنا, سجاد الزيدي, محمد إبراهيم السادة, researcher أحمد فوزي فارس; Outlet: الجزيرة نت — ثقافة; Date: 2026-05-22; Models: unnamed (“AI tools” generally) — current-vintage but model-vague (flag). Key quote: Al-Nahaba (translator): AI handles direct prose but “يعجز عن صنع ترجمة شعرية نابضة بالحياة” — “is incapable of producing a poetic translation pulsing with life”; “ترجمة الشعر تحتاج إلى حس باللغة، وإدراك للإيقاع” — “translating poetry requires a feel for language and awareness of rhythm.” Claim: Practicing Iraqi/Gulf writers and translators find AI serviceable for utilitarian prose but a failure at literary translation and register, and blame it for declining taste (ذائقة) and intellectual honesty in Arabic publishing.
-
URL: https://al-sharq.com/opinion/13/06/2026/الرواية-العربية-في-زمن-الذكاء-الاصطناعي Author: إحسان الفقيه (Ihsan Al-Faqih), Jordanian columnist; Outlet: الشرق (Qatar); Date: 2026-06-13; Models: none named — current-vintage, model-vague (flag). Key quote: “الآلة تملك رصيد المعلومة والقدرة على المحاكاة، لكنها لا تملك المشاعر” — “The machine possesses a stock of information and the capacity for simulation, but it does not possess feelings.” Claim: AI-generated Arabic fiction will displace formulaic mid-tier novelists while leaving genuinely felt writing untouched — a competence-without-interiority verdict.
-
URL: https://kuwaitai.net/en/insights/chatgpt-vs-claude-vs-gemini-arabic/ (corroborated by https://truescho.com/en/blog/claude-vs-chatgpt-arabic-content-2026) Author: KuwaitAI (Gulf AI-practitioner site, no named individual — quality flag: SEO-adjacent, methodology undisclosed); Outlet: KuwaitAI / Truescho; Date: 2026; Models: Claude Sonnet 4.6, Claude Opus 4.7, GPT-4.1/GPT-5, Gemini 2.5 Pro — genuinely frontier-scoped. Key finding (English-language source): Claude Sonnet 4.6 rated best on Khaleeji dialect, cultural context, lowest hallucination; GPT models “strong on Fusha, medium on Khaleeji”; Gemini “good comprehension but very formal tone.” Truescho: blind-tested Arabic content professionals preferred Claude for natural rhythm and vocabulary variety, with a noticeable quality jump at Claude Opus 4.6 (Jan 2026). Claim: Gulf practitioners now differentiate frontier models by Arabic register — Claude for dialect/literary naturalness, GPT-5 for fusha correctness, Gemini for officialese.
Failure modes observed
- Translated-from-English ركاكة: AI Arabic reads as clumsy calque of English syntax/phrasing (Ahmed Lotfi via Alpheratz; attributed to Arabic being ~3% of web training data).
- Prosody collapse in verse: broken meter (كسر الوزن), faulty rhyme, ignored conventions (ألف الإطلاق), wrong formal lengths — even when a single line scans, sustained metrically-sound language fails (Al Bayan poets; a Mandumah-indexed academic study on عروض/نحو constraints).
- Fusha/dialect register failure: models default to stiff MSA; produce dialects far worse than they understand them; idiom, sarcasm, and cultural allusion weak; low-resource dialects (e.g., Maghrebi variants) worst (Al Jazeera tech 2026, AraDiCE benchmark).
- Derivative recombination: “تحاكي ولا تبدع” — orbits existing texts, no critical stance, cold and reference-poor literary criticism (Hanawi).
- Flat literary translation: competent at direct prose, “lifeless” at poetry; misreads religious/social terminology and cultural nuance (Al-Nahaba; 2025 study cited by Independent Arabia).
- Ecosystem effect: AI-polishing of weak manuscripts flooding Arabic publishing with superficial texts; declining ذائقة (taste) (Sorori; Al Jazeera May 2026).
- إنشائية: school-composition blandness — formally correct, rhetorically dead (Al Bayan headline framing).
Praise / strengths noted
- Clear, acknowledged improvement trajectory: 2026 practitioner tests and Al Jazeera’s July 2026 survey say frontier + Arabic-focused models (ALLaM, Jais, Falcon Arabic, Fanar) have moved past “literal clumsy translation” toward context-aware MSA.
- Claude Opus 4.6/Sonnet 4.6 singled out (Gulf practitioner blogs) for natural rhythm, vocabulary variety, dialect and cultural-context handling; GPT-5 praised for fusha grammar and a more “human, spontaneous” style (Arageek, Oct 2025).
- Useful as a tool: idea development, structural suggestions, proofreading/grammar correction, formally-correct occasional verse (Diwan al-Arab team: “نصوصاً رصينة غنية بالصور المجازية”; Sorori’s “interactive Risalat al-Ghufran” vision).
- Several critics (Al-Faqih, Al-Jneibi) concede AI already outwrites mediocre human novelists and occasional poets.
Evidence quality & gaps
- Strengths: genuine native-speaker voices with real credentials (Bazee, Sorori, Hanawi, Azhar/UAE academics); multiple independent sources converge on the same failure modes (ركاكة, register, prosody); good 2025-2026 recency.
- Gaps: (1) Almost all literary criticism is model-vague or ChatGPT-vintage — virtually no published native-speaker assessment names Claude Opus 4.5/4.6, Fable 5, GPT-5.x, or Gemini 3 in a literary context; the frontier-specific comparisons that do exist (KuwaitAI, Truescho) are methodology-free practitioner/SEO blogs. (2) Fable 5 and Gemini 3 are absent from the Arabic-language discourse found. (3) Key primary texts were unreachable: Raseef22’s “الأدب الممل” feature, Habib Sorori’s and Parwin Habib’s Al-Quds al-Arabi columns, and Assabeel’s academic-critics symposium all returned 403 — their arguments are represented here only via search snippets. (4) No blind-evaluation study by Arab linguists of frontier-model fiction was found; the closest rigorous artifacts are NLP benchmarks (AraDiCE) measuring dialect capability, not literary quality. (5) Search budget exhausted before verifying whether any late-2025/2026 academic prosody study covered post-GPT-4 models.