Evaluating Free AI Chatbots on Civic and Human-Rights Topics: A 15-Country Expert Study
Rattachement africain : at, es, dk, ee, pe, ar, br. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Large language models (LLMs) increasingly shape how journalists, researchers, and the general public access information on high-stakes topics such as elections, human rights, identity politics, historical memory, and public services. People around the world now turn to free AI assistants to understand how their electoral systems work, what protections exist for vulnerable groups, or how public services like healthcare are organized. Most existing evaluations of LLMs, however, focus on single-country settings, synthetic benchmarks, or narrow political spectra rather than cross-country expert assessments of real, free-text answers to sensitive societal questions. They rarely examine how public-facing systems describe the same contested topics across multiple languages and political regimes, nor do they systematically measure differences across countries and models in dimensions that matter for journalism and civic life, such as neutrality, factual accuracy, balance, and tone. This study by the Media and Journalism Research Center (MJRC) addresses that gap by examining how eight widely used chatbots, namely ChatGPT, Gemini, Perplexity, Claude, Copilot, DeepSeek, Meta AI, and Grok, describe politically and socially sensitive topics in 15 countries that span diverse political regimes, cultural contexts, and levels of press freedom. We focus on the basic, free versions of these assistants, the tools that billions of users worldwide rely on daily to understand complex, contested realities. Using human expert-annotated scores on Neutrality, Factuality, Balance, and Tone, we assess strengths and weaknesses across four broad issue clusters: democratic institutions and power, rights and identity, social policy and welfare, and historical memory and corruption. We ask three guiding questions. First, how do free AI assistants perform on high-stakes civic and human-rights topics across diverse countries and languages when evaluated by domain human experts? Second, how do their strengths and weaknesses vary across issue types, from present-day service systems to contested histories of violence and corruption? Third, how do differences across bots and countries map onto existing information inequalities, and what does this imply for journalism, human-rights advocacy, and public knowledge more broadly?
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Où se fait cette recherche
-
Central European University pays non établi dans la noticeUniversité ou école supérieure
-
Universidade de Santiago de Compostela pays non établi dans la noticeUniversité ou école supérieure
-
Danish School of Media and Journalism pays non établi dans la noticeUniversité ou école supérieure
-
Media and Journalism Research Center pays non établi dans la noticeOrganisation à but non lucratif
-
Peruvian University of Applied Sciences pays non établi dans la noticeUniversité ou école supérieure
-
National University of Quilmes pays non établi dans la noticeUniversité ou école supérieure
-
Consejo Nacional de Investigaciones Científicas y Técnicas pays non établi dans la noticeOrganisme public
-
Universidad de Buenos Aires pays non établi dans la noticeUniversité ou école supérieure
-
Universidade Federal do Espírito Santo pays non établi dans la noticeUniversité ou école supérieure
-
University of Buenos Aires pays non établi dans la noticeUniversité ou école supérieure
Central European University, Universidade de Santiago de Compostela et Danish School of Media and Journalism, avec 7 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.