WSIS Forum 2026
Rapport généré par l'IA

Teaching AI to Speak our Language: A Showcase of Global Efforts to Bridge the ‘Last Mile’ of AI Inclusion for Local Impact

6 intervenants
Résumé

Résumé

Cette discussion a porté sur le défi que représente le développement d'outils linguistiques basés sur l'IA pour les langues sous-représentées, avec un accent particulier sur les contextes humanitaires et de maintien de la paix. Trois projets principaux ont été présentés, ainsi qu'une initiative de coordination internationale plus large. James D'Ercole a décrit un projet de maintien de la paix de l'ONU basé à Djouba, au Soudan du Sud, visant à développer un outil de traduction en temps réel pour combler les lacunes de communication entre les casques bleus, issus d'une centaine de vingtaine de pays . L'outil est conçu pour fonctionner entièrement hors ligne dans des environnements à faible connectivité , avec une approche intégrant l'humain dans la boucle afin que l'IA assiste plutôt qu'elle ne remplace le jugement humain . Une ambition clé est de construire non seulement un ensemble de données linguistiques, mais un corpus culturel intégrant le sens et le contexte locaux , et de préserver l'arabe de Djouba en tant que langue vivante qui pourrait éventuellement être accessible en ligne . Aimee Ansari, de Clear Global, a mis en lumière la disparité mondiale plus large dans le soutien apporté par l'IA aux langues, en soulignant que la plupart des langues parlées dans les pays à faibles revenus restent mal représentées dans les modèles de pointe . Elle a illustré ce constat avec le nord-est du Nigeria, où seul le haoussa bénéficie d'une reconnaissance automatique de la parole véritablement fonctionnelle, excluant potentiellement environ 70 % de la population des flux de communication assistés par l'IA . Elle a souligné la nécessité de données vocales de haute qualité et diversifiées, collectées de manière systématique auprès de locuteurs d'âges, de genres et de dialectes différents , et a insisté sur l'importance du consentement éclairé ainsi que d'une licence de données ouverte et non commerciale . Barbora Bromová a présenté l'outil Loria, développé avec la Bibliothèque nationale de Serbie et le PNUD, qui automatise la numérisation de documents historiques grâce à la reconnaissance optique de caractères et au post-traitement OCR . Dafna Feinholz a présenté la coalition de l'UNESCO pour la diversité linguistique dans l'IA, lancée en juin 2025, qui réunit plus de 30 experts issus des gouvernements, du monde académique, des communautés et du secteur privé pour partager les bonnes pratiques et constituer un répertoire d'approches portées par les communautés . Une discussion de clôture a mis en évidence deux défis persistants : le coût élevé d'une collecte de données éthique et centrée sur les communautés , et la difficulté de coordonner des efforts fragmentés entre organisations . Aimee Ansari a noté que si l'implication du secteur privé est précieuse, les communautés linguistiques sans intérêt commercial risquent d'être laissées pour compte, soulignant la nécessité pour les organisations internationales de contribuer à garantir une représentation équitable dans l'infrastructure mondiale de l'IA .

Points clés

Objectif général

La discussion a porté sur le défi que représente le développement d'outils linguistiques et de traduction basés sur l'IA pour les langues sous-représentées, en particulier dans les contextes humanitaires et de maintien de la paix. Des intervenants du Département des opérations de paix de l'ONU, de Clear Global, du PNUD et de l'UNESCO ont partagé leurs projets respectifs et ont exploré comment la collaboration, l'implication des communautés et une gouvernance éthique des données pouvaient contribuer à combler le fossé mondial en matière de technologie linguistique. ---

Principaux points de discussion

- Le fossé de communication critique dans les opérations de maintien de la paix et humanitaires : Les missions de maintien de la paix de l'ONU impliquent environ 50 000 personnels en uniforme provenant d'une centaine de vingtaine de pays fournisseurs de troupes et de police, mais la communication avec les communautés locales et entre les Casques bleus eux-mêmes reste un problème persistant et de longue date.

Un outil de traduction en temps réel et hors ligne est en cours d'expérimentation à Djouba, au Soudan du Sud, spécifiquement conçu pour fonctionner dans des environnements à faible connectivité et aux conditions difficiles, avec une approche intégrant l'humain dans la boucle pour garantir que l'IA assiste plutôt qu'elle ne remplace le jugement humain. - L'inquiétante sous-représentation des langues peu dotées dans les modèles d'IA : La plupart des langues parlées par de larges populations dans les pays du Sud restent mal prises en charge ou totalement absentes des modèles d'IA de pointe, les données textuelles et vocales étant fortement biaisées par l'intervention de pays plus riches à dominante anglophone. Par exemple, dans le nord-est du Nigeria, parmi les dix langues couramment parlées, seul le haoussa bénéficie d'une reconnaissance automatique de la parole (ASR) véritablement fonctionnelle, risquant d'exclure environ 70 % de la population des flux de communication assistés par l'IA. Les défis incluent la nature principalement orale de nombreuses langues, l'alternance codique, l'absence d'orthographes standardisées, et le fait que les indicateurs de performance en laboratoire ne reflètent pas la précision dans le monde réel. - La nécessité d'une collecte de données centrée sur les communautés et qui préserve les cultures: Plusieurs intervenants ont souligné que le développement d'une IA linguistique efficace nécessite une implication profonde des communautés dès le début du processus de conception, et non comme une réflexion après coup.

Cela inclut la co-conception d'approches d'enregistrement pour les langues principalement orales, l'obtention du consentement éclairé, la protection des données vocales, et la garantie que le contexte culturel - et pas seulement les données linguistiques - soit intégré dans les modèles.

Le projet déployé au Soudan du Sud, par exemple, s'emploie à constituer un corpus culturel avec des collègues nationaux afin de rendre le modèle linguistique moins centré sur l'Occident. - La gouvernance des données, la propriété et le risque d'exploitation : Une préoccupation récurrente dans toutes les présentations était de garantir que les communautés dont les données linguistiques sont collectées conservent une propriété réelle et ne soient pas exploitées.

Aimee Ansari a mis en évidence la tension entre le coût élevé d'une collecte de données éthique - nécessitant des salaires équitables et un consentement approprié - et les ressources limitées des organisations de terrain, suggérant que la mutualisation des ressources entre plusieurs organisations pourrait rendre la collecte de données plus viable financièrement. La question de l'implication du secteur privé a également été soulevée, avec la reconnaissance que si certains partenaires sont ouverts en matière de gouvernance des données, l'intérêt commercial pour les langues très peu dotées reste limité, car leurs locuteurs ne constituent pas de grands consommateurs sur les marchés numériques. - La coordination internationale et la coalition pour la diversité linguistique dans l'IA : L'UNESCO, en partenariat avec l'Islande, a créé la coalition pour la diversité linguistique dans l'intelligence artificielle, lancée en juin 2025, qui réunit plus de 30 experts issus des gouvernements, du monde académique, des communautés, d'experts techniques, d'organisations internationales et du secteur privé. La coalition vise à documenter les bonnes pratiques dans un répertoire partagé, à identifier les lacunes en matière de renforcement des capacités et à éclairer les orientations politiques sur la diversité linguistique dans l'IA - en s'éloignant des efforts cloisonnés vers un modèle collaboratif et multipartite. Dafna Feinholz a souligné que la préservation linguistique est indissociable de la préservation culturelle, et que les communautés doivent être impliquées tout au long du cycle de développement de l'IA. ---

Ton général

Les intervenants ont été collaboratifs, sincères et axés sur les solutions, tout en exprimant une certaine urgence. Ils ont exprimé une véritable passion pour leur travail et un engagement partagé en faveur de l'équité et de l'inclusion dans le développement de l'IA. Ils ont été optimistes - en particulier lorsqu'ils ont décrit le succès de l'engagement communautaire et la coalition croissante de partenaires - mais ils ont reconnu des défis structurels importants, notamment les contraintes de financement, les complexités de la gouvernance des données et les incitations commerciales limitées pour les acteurs du secteur privé à investir dans les langues peu dotées.

Vers la fin de la discussion, lors des questions-réponses avec le public, le ton est devenu légèrement plus direct et informel. De plus, les réponses directes des intervenants ont permis d'ajouter des exemples concrets à la présentation plus formelle. Les remarques de clôture ont maintenu une atmosphère chaleureuse, avec une invitation claire à poursuivre la collaboration.

Intervenants

- James D'Ercole - Rôle/Titre : Non explicitement mentionné, mais travaille dans le maintien de la paix de l'ONU - Domaines d'expertise : Opérations de maintien de la paix de l'ONU, communication de dernier kilomètre, outils de traduction en temps réel pour les langues peu dotées, outils linguistiques assistés par l'IA pour les missions de terrain, gouvernance des données dans les contextes humanitaires - Contexte supplémentaire : Ancien responsable administratif régional pour la mission de l'ONU au Timor oriental ; plus de 20 ans d'expérience dans le maintien de la paix ; dirige un projet pilote à Juba, au Soudan du Sud, axé sur la traduction en temps réel pour le personnel en uniforme - Aimee Ansari - Rôle/Titre : Directrice générale de Clear Global (anciennement Traducteurs sans frontières) - Domaines d'expertise : Technologie linguistique humanitaire, reconnaissance automatique de la parole (ASR) pour les langues peu dotées, collecte de données centrée sur les communautés, diversité linguistique dans l'IA, consentement éclairé dans la collecte de données vocales, réponse humanitaire multilingue - Dafna Feinholz - Rôle/Titre : Représentante de l'UNESCO - Domaines d'expertise : Éthique de l'intelligence artificielle, diversité linguistique et culturelle dans l'IA, gouvernance multipartite, Recommandation de l'UNESCO sur l'éthique de l'IA, Coalition pour la diversité linguistique dans l'intelligence artificielle - Barbora Bromová - Rôle/Titre : Représentante du PNUD (modératrice de la session) - Domaines d'expertise : Diversité linguistique et IA, numérisation de matériaux d'archives, reconnaissance optique de caractères pour les langues peu dotées, outils linguistiques open source, collaboration du PNUD avec des institutions nationales sur la technologie linguistique (par exemple, les projets Librarify et Loria en Serbie) - Public - Rôle/Titre : Traducteur ; travaille avec Meta sur l'internationalisation (I18N) des grands modèles de langage (LLM) - Domaines d'expertise : Traduction, développement de modèles d'IA multilingues, nuances culturelles et linguistiques dans les modèles de langage, défis des langues peu dotées dans les systèmes d'IA commerciaux --- Intervenants supplémentaires : - Intervenant (non identifié) - Domaines d'expertise : Semble avoir une expertise dans le développement de l'IA pour les langues peu dotées, l'ASR pour les langues orales, les structures des langues créoles et des langues véhiculaires, la collecte de données linguistiques en milieu communautaire ; s'est exprimé en réponse à la question du public sur le problème du double coût

Intervenants
JD
James D'Ercole
135 wpm · 12 min
AA
Aimee Ansari
133 wpm · 16 min
BB
Barbora Bromová
137 wpm · 13 min
DF
Dafna Feinholz
138 wpm · 8 min
A
Audience
167 wpm · 3 min
S
Speaker
172 wpm · 1 min

Outils linguistiques basés sur l'IA pour les langues à faibles ressources dans les contextes humanitaires et de maintien de la paix

#

Aperçu général et cadrage

Cette discussion a réuni des membres du Département des opérations des paix (DPO) des Nations Unies, de Clear Global, du PNUD et de l'UNESCO afin d'examiner les défis liés au développement d'outils linguistiques et de traduction basés sur l'IA pour les langues à faibles ressources et sous-représentées, en particulier dans les contextes humanitaires et de maintien de la paix. La séance s'est articulée autour de trois présentations de projets, suivies d'une introduction à une initiative de coordination internationale et d'une discussion ouverte avec le public. Barbora Bromová a joué un double rôle tout au long de la session, à la fois modératrice et présentatrice du projet de numérisation serbe. James D'Ercole, qui s'est appuyé sur son expérience d'administrateur régional au Timor oriental, a introduit un fil conducteur à la discussion : le principe consistant à se concentrer sur le « dernier kilomètre » . Son argument était que si un système fonctionne correctement au point le plus éloigné et le plus contraint en ressources par rapport au siège - que ce soit dans une mission, à New York ou à Genève - alors tout ce qui se trouve en amont de la chaîne fonctionne nécessairement . Cette intervention a permis d'ancrer le thème de la discussion dans la réalité du terrain plutôt que dans des discussions institutionnelles.

#

Le fossé communicationnel dans le maintien de la paix de l'ONU

James D'Ercole a ouvert la session en décrivant l'ampleur et la persistance du problème de communication auquel font face les opérations de maintien de la paix de l'ONU. Fort d'un peu plus de 20 ans d'expérience dans ce domaine, il a qualifié ce problème de défi systémique et durable. Environ 50 000 personnels en uniforme ou plus sont déployés quotidiennement dans 11 missions, issus d'environ 120 pays fournisseurs de troupes et de police . La communication - tant entre les Casques bleus et les communautés qu'ils servent qu'entre les Casques bleus eux-mêmes - est fondamentale pour la mission de maintien de la paix, qui repose sur l'écoute et la compréhension . Pourtant, ce problème est demeuré persistant et systémique tout au long de la carrière de James D'Ercole, documenté à maintes reprises dans des rapports de terrain, des soumissions C-34 de gouvernements, des visites de comités consultatifs militaires et policiers, et des rapports sur les droits de l'homme . Les interprètes et traducteurs n'ont jamais été disponibles en nombre suffisant, même durant les périodes mieux dotées en ressources , et la diversité des langues nationales des Casques bleus aggrave encore le défi .

Le projet pilote mené à Djouba, au Soudan du Sud, est spécifiquement conçu pour combler ce fossé grâce à la traduction en temps réel . Il repose de manière cruciale sur une approche de supervision centrée sur l'humain : l'IA est conçue pour assister le processus, une personne étant toujours présente pour valider les résultats, sans jamais remplacer le jugement humain ni éliminer les rôles des assistants linguistiques et des interprètes . James D'Ercole a soutenu que cette approche amplifierait, si quoi que ce soit, la portée et l'influence des professionnels des langues existants plutôt que de les diminuer . Une exigence technique centrale est que l'outil doit fonctionner entièrement hors ligne, compte tenu de la faible connectivité ou de l'absence de connectivité dans les environnements où opèrent les missions de maintien de la paix . Un dispositif de périphérie - décrit comme un mini-serveur capable de fonctionner sur batterie - est actuellement en phase de test .

Le modèle de traduction au cœur du projet a été développé avec des contributions d'une initiative liée à Harvard et d'un projet de fin d'études de l'NYU impliquant des étudiants de master, ce qui reflète les partenariats académiques collaboratifs qui ont façonné l'initiative dès ses débuts. James D'Ercole a décrit un cycle permettant au système de constamment s'améliorer : les tests hors ligne alimentent la constitution et l'enrichissement de jeux de données, auxquels participent des Volontaires des Nations Unies travaillant en ligne, dont les contributions sont intégrées dans le modèle de traduction, qui est ensuite déployé sur le dispositif de périphérie. La phase de preuve de concept fonctionne dans ce que James D'Ercole a appelé le « mode fantôme », dans lequel le dispositif écoute un interprète en train de travailler, puis l'interprète examine les résultats du dispositif afin de les valider, de les corriger, et contribuant ainsi à établir les garde-fous du système. Cette boucle de validation alimente ensuite le cycle d'amélioration suivant, créant un modèle continuellement affiné, ancré dans l'expérience réelle.

#

Construire un corpus culturel, pas seulement un jeu de données linguistiques

Selon James d'Ercole, l'un des objectifs du projet est de créer un corpus culturel qui ne se limite pas à collecter des données linguistiques . Le projet a fait appel à un expert en intégration de la culture dans les modèles de langage IA, spécifiquement pour rendre le modèle moins occidentalo-centré et plus adapté au contexte de la communauté arabophone de Djouba . Cela implique de valider les données collectées par des volontaires des Nations Unies en ligne, puis de faire appel à des collègues nationaux au sein de la mission pour valider et enrichir davantage ces données avec un contexte culturel . Des assistants linguistiques ont déjà contribué volontairement sur leur temps personnel, motivés par une conviction sincère dans le projet et dans l'importance de la numérisation de l'arabe de Djouba .

James D'Ercole a identifié la dimension de préservation de ce travail comme l'aspect qu'il valorise le plus , décrivant l'objectif comme la création d'une « langue vivante » - préservant non seulement la langue elle-même, mais aussi la culture qu'elle convoque . Il a également noté qu'une fois une langue numérisée et mise à disposition en tant que données en source ouverte, des partenaires commerciaux pourraient éventuellement développer d'autres outils permettant aux communautés d'accéder à Internet dans leur propre langue . James D'Ercole a articulé trois bénéfices fondamentaux du projet : une meilleure communication, une confiance plus profonde et une langue vivante. Le projet est construit à travers des partenariats collaboratifs avec des institutions locales, des universités internationales, des organisations de la société civile et des ONG, avec une coopérative linguistique destinée à garantir que les communautés contribuant aux données puissent éventuellement en bénéficier . La gouvernance des données, la sécurité de l'IA et la protection des données vocales sont abordées dans le cadre d'une approche qui vise à ne pas nuire. C'est pourquoi la propriété communautaire et les garde-fous sont non négociables . Ce travail est mené sous la supervision du Bureau de la protection des données et de la vie privée de l'ONU, un nouveau bureau que James D'Ercole a décrit comme précieux et, compte tenu de sa nouveauté, parfois difficile à appréhender .

#

La disparité mondiale dans la couverture de l'IA linguistique

Aimee Ansari de Clear Global a situé le projet de Djouba dans un schéma mondial beaucoup plus large d'inégalité dans la couverture de l'IA linguistique. S'appuyant sur des recherches menées en 2025, elle a décrit comment la quantité de données textuelles disponibles en ligne dans les différentes langues - un déterminant clé de la performance des grands modèles de langage - est fortement biaisée en faveur des langues parlées dans les pays les plus riches, tandis que les langues comptant des dizaines ou des centaines de millions de locuteurs dans le Sud global restent mal prises en charge ou totalement absentes des modèles de pointe . Elle a illustré cette disparité par une comparaison frappante : le breton, parlé par environ 200 000 personnes en France, dispose de plus de 50 modèles d'IA, tandis que le pidgin nigérian, avec environ 85 millions de locuteurs, n'en compte que 10 . Elle a également cité le saraiki, dans le nord-est du Nigeria, comme autre exemple d'une langue bénéficiant d'une couverture IA négligeable. Ce fossé n'est pas déterminé par le besoin communicatif ou la population de locuteurs, mais par le pouvoir économique et géopolitique.

La situation est encore plus critique pour les modèles de parole. Aimee Ansari a noté que seule une fraction des langues dans le monde est véritablement prise en charge par la technologie de reconnaissance automatique de la parole (automatic speech recognition ou ASR) . Dans le nord-est du Nigeria, parmi les dix langues les plus couramment parlées dans les États de Borno, Adamawa et Yobe, seul le haoussa bénéficie d'un support ASR, y compris une API commerciale . Même pour le haoussa, les performances n'ont été évaluées qu'en conditions de laboratoire, et les performances réelles sur le terrain restent largement inconnues . La conséquence pratique est que les flux de travail assistés par ASR reposant sur le haoussa ou l'anglais risquent d'exclure environ 70 % de la population du nord-est du Nigeria de la communication assistée par IA .

#

Défis spécifiques aux langues à faibles ressources et aux langues orales

Aimee Ansari a identifié plusieurs défis qui rendent particulièrement difficile le développement de l'ASR pour les langues à faibles ressources. Beaucoup de ces langues sont principalement orales, ce qui signifie que la transcription écrite, qui est une condition préalable à la construction de modèles ASR, constitue une entreprise extrêmement complexe . La terminologie et les façons de s'exprimer varient considérablement d'un locuteur à l'autre, et le mélange de codes entre deux langues est courant, créant une complexité supplémentaire pour les technologies numériques du langage . Un point méthodologique crucial qu'elle a soulevé est que les métriques de taux d'erreur sur les mots en laboratoire ne permettent pas de prédire de manière fiable les performances dans le monde réel : un faible taux d'erreur sur les mots signifie généralement qu'un chercheur a testé le modèle dans un environnement contrôlé, ce qui nous apprend peu sur son fonctionnement avec de vrais locuteurs dans de vrais environnements . Au cœur de ce fossé se trouve une pénurie de données vocales de haute qualité et diversifiées : la construction d'un modèle ASR fiable nécessite des centaines d'heures de données enregistrées, collectées systématiquement auprès de locuteurs de différents âges, genres, niveaux d'éducation et dialectes .

La réponse de Clear Global à ces défis est une approche de collecte de données centrée sur la communauté. L'organisation collabore avec les membres de la communauté dès le début du processus de conception - avant toute collecte de données - pour comprendre comment les modèles seront utilisés, saisir les complexités linguistiques et identifier les obstacles . Cela inclut la formation de membres de la communauté en tant que linguistes et la co-conception d'approches pour mettre par écrit des langues principalement orales, car il peut n'exister ni alphabet standard ni forme écrite convenue . Aimee Ansari a donné l'exemple d'un enregistrement en canari où dix façons différentes d'écrire le même mot existaient et où les contributeurs ne pouvaient pas s'accorder sur une seule version correcte, étant donné qu'elles étaient toutes acceptées socialement . Des approches de co-conception sont utilisées pour gérer cette complexité et atteindre une cohérence suffisante pour le développement du modèle . Le consentement éclairé est un principe fondateur : des ateliers communautaires sont organisés pour s'assurer que les locuteurs comprennent comment leurs enregistrements seront utilisés et stockés, quels sont les risques et comment ils peuvent retirer leur consentement à tout moment . Tous les jeux de données sont publiés en accès libre sous des licences non commerciales afin de pouvoir être utilisés dans l'ensemble du secteur .

#

Numérisation de documents historiques : le projet Loria en Serbie

Barbora Bromová a présenté un projet complémentaire abordant une dimension différente du fossé dans les données linguistiques : la numérisation de documents historiques qui existent au format papier. Dans le cadre du projet Librarify, l'équipe a développé un flux de travail en collaboration avec la Bibliothèque nationale de Serbie et le PNUD pour traiter les numérisations des archives de la bibliothèque à l'aide de la reconnaissance optique de caractères, de l'identification de la mise en page, de l'amélioration des images et de la correction post-OCR . Le serbe présente des défis particuliers : malgré un corpus important de sources écrites historiques, il se situe en bas des distributions de données linguistiques , en partie parce que la langue utilise indifféremment les alphabets cyrillique et latin, ce qui perturbe les modèles OCR standard . La phase initiale du projet a numérisé plus de 400 gigaoctets et plus de 16 000 documents .

À partir de cette expérience, l'équipe a développé Loria, un outil plus généralisable et en source ouverte, publié sur GitHub pour être réutilisé et adapté à différentes langues et institutions . Loria suit le même flux de travail en quatre étapes, à savoir l'amélioration de l'image, l'identification de la mise en page, l'OCR et la correction post-OCR, mais est conçu pour la réutilisation et comprend un module d'extension pour les modèles de pointe afin de permettre une correction post-OCR supplémentaire et un traitement par lots . Il a été développé principalement avec les archivistes au centre du processus de supervision humaine , reflétant le consensus général de la session selon lequel l'expertise du domaine et les connaissances institutionnelles sont aussi importantes que la capacité technique . Le projet a impliqué une équipe multidisciplinaire regroupant des experts en apprentissage automatique, des développeurs full stack, des designers, des chefs de produit, la Bibliothèque nationale de Serbie, l'Institut mathématique de l'Académie des sciences et le PNUD .

#

La coalition de l'UNESCO pour la diversité linguistique dans l'IA

Dafna Feinholz de l'UNESCO a présenté l'initiative de coordination internationale de la coalition pour la diversité linguistique dans l'intelligence artificielle, établie en partenariat avec le ministère islandais de la Culture, de l'Innovation et de l'Enseignement supérieur et le Centre islandais pour la langue et la technologie . Dafna Feinholz a décrit la coalition comme ayant débuté environ un an auparavant et ayant été officiellement lancée lors du Forum mondial de l'UNESCO sur l'éthique de l'intelligence artificielle à Bangkok, le prochain forum de ce type étant prévu en Arabie saoudite du 14 au 17 septembre . La coalition compte déjà plus de 30 experts travaillant sur l'inclusion numérique des langues portée par les communautés, provenant de presque toutes les régions du monde . Dafna Feinholz a ancré les travaux de la coalition dans un argument philosophique : la langue n'est pas simplement un outil de communication, mais le medium à travers lequel les personnes expriment leur vision du monde, se comprennent elles-mêmes et transmettent leur patrimoine . La préservation de la langue est donc indissociable de la préservation de la culture, et ce principe est inscrit dans la Recommandation de l'UNESCO sur l'éthique de l'IA, qui inclut la protection de la diversité linguistique .

La coalition adopte délibérément une approche multipartite, réunissant dans un même espace des gouvernements, des milieux académiques, des communautés, des experts techniques, des organisations internationales et le secteur privé . Dafna Feinholz a souligné que ce modèle collaboratif s'est avéré précieux parce que les connaissances qui fonctionnent pour une communauté fonctionnent souvent pour d'autres, et parce que le fait d'avoir de grandes entreprises et de petites start-ups dans le même groupe permet aux acteurs plus modestes de bénéficier de capacités qu'ils ne pourraient pas développer de manière indépendante . Un résultat clé de la coalition est un référentiel de bonnes pratiques, construit autour d'un questionnaire co-créé en collaboration entre toutes les parties prenantes et l'UNESCO afin de systématiser et de rendre accessibles les connaissances sur les approches portées par les communautés en matière de diversité linguistique et culturelle dans l'IA . Ce référentiel est conçu comme une ressource vivante qui alimentera les orientations politiques sur la diversité linguistique dans l'IA, qui constitue un domaine prioritaire selon de nombreux États membres . Dafna Feinholz a insisté sur le fait qu'une inclusion linguistique et culturelle significative exige que les communautés soient impliquées tout au long du cycle de vie du développement de l'IA, de la conception initiale jusqu'au déploiement, plutôt que d'être intégrées à des stades tardifs .

#

Discussion de clôture : financement, coordination et secteur privé

Barbora Bromová a introduit la phase finale de la discussion en parlant de ce que les organisations internationales pourraient faire de plus utile pour soutenir des organisations comme Clear Global . Aimee Ansari a identifié deux domaines principaux dans lesquels les organisations internationales pourraient faire une différence. Le premier est financier : la collecte de données éthique et centrée sur la communauté est coûteuse, nécessitant des salaires équitables pour les contributeurs plutôt que de recourir à des bénévoles, et la plupart des organisations de base ne peuvent généralement pas se permettrede payer 50 000 à 100 000 dollars pour façonner un modèle de langage . Cependant, si cinq ou six organisations travaillant sur des langues apparentées pouvaient être réunies, chacune contribuant environ 10 000 dollars, le coût collectif devient gérable . Les organisations internationales, avec leur vue d'ensemble systémique, sont particulièrement bien placées pour remplir cette fonction de rassemblement, alors même que des organisations comme Clear Global ne sont pas en mesure de le faire .

Le deuxième domaine identifié par Aimee Ansari concernait l'établissement de normes autour des modèles d'IA souverains. Elle a soutenu que si le dialogue mondial sur la gouvernance de l'IA est un bon début, il existe une insuffisance de normes concernant la manière dont toutes les communautés et toutes les langues sont représentées dans les modèles d'IA construits au niveau national . Elle a utilisé le Soudan du Sud comme illustration percutante : dans un gouvernement composé principalement d'un groupe ethnique à la suite d'une guerre civile, la probabilité de construire un modèle souverain qui représente équitablement la langue et la culture des autres groupes est faible . Elle a soutenu que l'un des rôles de l'ONU doit être de veiller à ce que l'infrastructure publique mondiale soit construite de manière juste et équitable .

Barbora Bromová, précisant qu'elle ne s'exprimait pas au nom du PNUD, a observé que si certains partenaires du secteur privé sont véritablement ouverts d'esprit en matière de gouvernance des données, beaucoup ne le sont pas, et que le défi de la propriété des données reste non résolu . Elle a également mentionné que son équipe explore des plateformes décentralisées de partage de données comme moyen de naviguer dans ces défis liés à la propriété.

Un membre du public travaillant avec Meta sur la construction de grands modèles de langage s'est présenté avec un humour en disant qu'il travaillait pour l'ennemi , avant de recadrer le défi comme un problème technique et de coût partagé. Il a noté que même au sein des grandes entreprises technologiques, les équipes axées sur l'internationalisation se heurtent aux mêmes problèmes de nuance culturelle, de mélange de codes et de formes écrites non canoniques . Il a posé ce qu'il a appelé le « problème des deux coûts » : la double difficulté consistant soit à se battre pour faire construire un modèle dans une langue donnée, soit à se battre pour trouver des personnes qui connaissent suffisamment bien la langue pour y contribuer de manière significative . Aimee Ansari a répondu en déclarant explicitement que « Meta n'est pas du tout l'ennemi » et en décrivant la collaboration de Clear Global avec Meta pour tester la sécurité des modèles , tout en maintenant que les locuteurs de langues comme le dinka ne sont « pas commercialement intéressants » pour les grandes entreprises parce qu'ils ne constituent pas des consommateurs significatifs sur les marchés numériques . Elle a également noté que Clear Global travaille abondamment avec des organisations telles que Viamo, la GSMA et Lelapa AI - des acteurs privés de taille petite à moyenne qui occupent une position différente dans l'écosystème par rapport aux grandes entreprises de modèles de pointe.

Un autre participant présent dans la salle a proposé que la construction de structures orthographiques standardisées pour les langues principalement orales - potentiellement approuvées par des organismes tels que l'IEEE ou l'UNESCO - pourrait contribuer à préserver la nuance et le contexte dans la transcription et le développement de l'ASR . Cependant, comme l'a souligné Aimee Ansari, même les membres des communautés eux-mêmes ne peuvent parfois pas s'accorder sur une norme écrite unique. Ainsi a standardisation descendante n'est peut-être ni facilement réalisable ni toujours appropriée .

#

Défis non résolus et agenda prospectif

La séance s'est conclue par l'idée partagée que, malgré un fort consensus conceptuel entre tous les intervenants sur les principes de conception centrée sur la communauté, de supervision humaine des résultats de l'IA, de préservation culturelle et de gouvernance éthique des données, d'importants défis structurels restent non résolus. Ceux-ci incluent le financement durable du développement de modèles de langage pour les langues commercialement peu attractives , la gouvernance des données vocales et le risque d'extractivisme , les dimensions politiques de la construction de modèles d'IA souverains dans des sociétés divisées , et la difficulté pratique de coordonner des efforts fragmentés à travers l'écosystème . La coalition de l'UNESCO pour la diversité linguistique dans l'IA, l'outil en source ouverte Loria et la coopérative linguistique proposée à Djouba représentent tous des étapes concrètes vers la résolution de ces défis, mais tous les intervenants ont reconnu que le travail en est à un stade précoce et que les obstacles à venir sont autant politiques et économiques que techniques . La session s'est clôturée par une invitation ouverte à la collaboration, Barbora Bromová encourageant toute organisation du secteur privé intéressée à prendre contact, et tous les intervenants exprimant leur enthousiasme pour la poursuite des partenariats à travers le secteur .

James D'Ercole
I think it's okay. I think we're okay. Yeah, we're fine. Okay. All right. So Pilot is in Juba, South Sudan, and it's specifically built with the community and focusing on the last mile. Now, for me, the last mile kind of means a lot because in my past, I was a regional administrative officer for the mission in East Timor. So it was the last mile. I was out there, and I knew that if you get it right at the furthest distance away from headquarters, whether that's headquarters in a mission or headquarters in New York or Geneva, then everything along the way has to be working. Focus on the last mile, and this is what we're focusing on specifically on this project. So the project. Problem. What we're dealing with is we've got about 50 ,000 or more uniformed personnel out there every day in 11 missions. And this represents approximately 120 or so troop and police contributing countries. Now, it might not be as much as maybe UNDP collectively, but it's quite significant for us. And all of this peacekeeping, it depends on communication. I mean, this is a of peacekeeping is listening and understanding to the communities that we're working in. Now, since I've been in peacekeeping a little over 20 years now, it's always been a problem. And every year you get countless field reports, whether it's C -34 coming from governments, whether it's a military police advisors committee is going out to mission, whether it's human rights reports coming back. And it's always been a problem. we're just, we're not communicating as well as we could be. Interpreters, translators, and even if we did have the money in the better days of yesteryear, you still can't have enough, right? So this project is trying to look at that. But also, not only that, we've got 120 different countries of actual peacekeepers. So the communication between peacekeepers is also a major challenge. And the people that we serve, the peacekeepers themselves, are coming from low -resource languages. So the solution. So what we're doing is we're focusing on this real -time translation. It's a tool to help close this gap that I was telling you about. What we're doing is specifically with the human in the loop. we're looking at it as AI as assisting the process person will always validate should always be there it never replaces human judgment or it doesn't remove the jobs if anything it kind of amplifies and elevates by language assistance or interpreters because the focus now will turn into the individual having a larger influence on several but I'll get back to that later on we're also focusing on cultural sensitivity we're the United Nations so of course we need to build the device or the tool making sure that local meaning and register and context are all built into the whole process and specifically in the environments that we work there's little to no connectivity. So no signal. The device, the tool, has to work completely offline and in kind of austere environments. I don't know if I'm going Yeah, you may pick it up. Okay. So, all right. So I'll pick it up a little bit. Sorry about that. So the benefits of what we're trying to do are better communications, deeper trust, trust, and a living language. So improving communications, it's obvious, right? The more we have the better communications, the better delivery, we can deliver the mandate better. With more communications and understanding with the people that we're serving, it builds trust, listening, inclusion, and hopefully a durable peace. And the part I like the most is that we're actually, by digitizing these languages, we're preserving the language. When we preserve the language, we preserve the culture. All right. And that to me is really important. And when we do that, it has the language in theory, then commercial partners can one day pick up that language and you may be able to actually surf the web in that in that language itself. OK, so. Real time translation, the whole cycle of this is first we try it offline, again, human in the loop, so we're testing it offline. We're kind of building enhancing data sets. We've had UNVs do this online UNVs. They worked very well. A translation model that we've built in basically Gartner. Harvard was originally a part of that. We had a UN cap NYU capstone project. So graduate students working on that. Then we have the edge device. So we have a sample edge device on loan right now. That is like a mini. server that can go out there and as long as we have the battery pack and can plug it in, it can actually do this remotely without any connectivity. And then the proof of concept. So what we want to do is we want to take that, we want to get it out there and put it into shadow mode where it sits and it listens to an interpreter do what it does and then the interpreter itself will now go through it, say yes, no, help build the guardrails and we'll go through it like that and then with that improvement cycle. Alright, how much time? I'm almost there. So what are we doing? We're building a cultural corpus, not just a language data set. So we're validating the data right now that was done by online volunteers. Now we want to take national colleagues in mission and actually validate that data. At the same time, we want to use national colleagues to collect more data. to make it more robust and add a cultural corpus into this. So we just an expert in this kind of integrating culture into the AI language models in order to make this language model less Western -centric and more specific to Juba in the context that we're in. Okay, look at that. Okay, so I think I skipped one. Let's see. So it's built with the community. We're doing it three ways. We're always looking for partnerships and a collaborative approach. So we've got local exploring partnerships and collaboration with. We've got international institutions and universities, CSOs, NGOs, and also we want to try, in theory, to get a language cooperative together so that the people that we're tapping into for this information can somehow one day. kind of get it back into them if this data is open source, which it is, but has a paywall to companies and things like that. So we're going to try to figure that out. We're also leaning heavily on national colleagues who have been really supportive of this. In fact, during the online volunteer, which is all that they don't get paid, we actually had language assistants on their own time contributing this because of actually believing in the project and believing in digitizing the language of Juba Arabic. So we're also leaning on national colleagues to make sure that it's kind of linguistically and culturally accepted to be working in, so we're relying heavily on them. UN Volunteers, absolutely fabulous organization to work with. We're working with online volunteers, and we're going to try to get some national online. in Unmiss itself to work there in order to kind of record, pair, and help validate the process as we go along. And then data governance, AI, safety and security, so voices protected data. So we're taking all that comes to that. The human in the loop, once again, we're saying that we're kind of throwing that down maybe too much, but, you know, we really want to do this AI safety by design. We want to integrate it from the beginning and amplify the reach. So basically when we're doing this, we want to make sure we're doing it right. And then safe public text, so protected speech without exposing a voice, because this is voice data, which we found out. It's quite. challenging to work with. I don't know if you're doing it. And then we do it under the do no harm approach, of course, and the biggest thing is community ownership. We want to make sure that we do this in the cooperative if possible. You've mutually defined the guardrails, what's acceptable, what isn't. Trying to build local ownership and work with the government through the national universities and make sure it's culturally and ethically aligned to the community on that. Thanks for the lights. And then all of this under the data protection and oversight guided by the Office of Data Protection and Privacy, a new office at the UN that we're working with who have been absolutely fabulous to work with, but also at the same time it's difficult to move through because they're new as well. And then last but not least, we're talking about governance, assessments, and the other things that we're working on. talking about risk, data impact, human rights, ethical impact. So we're trying to keep that all in mind. That is also over the horizon, but coming up soon. And seven minutes is that? Forty minutes, but Was it really 14 minutes? A little bit more than seven minutes. Oh, I'm sorry. I could have gone faster in the beginning, but I didn't want to lose. I can speak very quickly if you want. Anyway, so we're looking for collaborations, partnerships, discussions. Over to you.
Barbora Bromová
Wonderful. Actually, over to Amy, who will be doing our second presentation on behalf of Clear Global. Over to her now.
Aimee Ansari
Okay. Hi, everybody. I mean, great presentation. Super interesting work you're doing. What can you see? I am going to talk about very, very similar things, but at a slightly larger scale. Thank you. so I'm Amy Ansari I am the chief executive of Clear Global we're going to talk about we've done a lot of work in humanitarian response and why is that not the right screen no I've got it open in two different places and so I just have to find the right one I have this wonderful gentleman over here yeah yeah yeah no no no I'm good I just have to find the right one I've got one open online and one open that should be the right one there we go okay As we've been talking about, language AI is increasingly used in the social impact and humanitarian sector to process information with tools like automatic speech recognition and machine translation. They have the potential to improve efficiency and accessibility, particularly when integrated into your internal workflows, as you were talking about. But sometimes these systems also reflect the inequalities in language representation. When they don't support the language, which is what UNDPKO is trying to get around in South Sudan right now, they risk systematically excluding entire communities. And at worst, the output is wrong or life threatening in nature. So I'm going to talk to you a little bit about the work that we do. Most languages in the world, as you can see on this chart, remain really poorly supported or entirely absent from any frontier model, from any large language model. this from some research done this year it represents how much text data there exists online across languages and that gives you an indication of how well the text -based AI like large language models are likely to perform you can see from this that languages spoken in wealthier countries dominate while others including those with tens or hundreds of millions of intervenants don't have very much data and have limited resources and I was just reading a study today from Cambridge that talked about the Breton language which is spoken by about two hundred thousand people in French did you read this article too it's a fascinating it's a fascinating thing but it has over 50 models two hundred thousand people speak this language there are 50 models in it but a language like Nigerian Pigeon which has about 85 million intervenants has 10 models developed And it's the same for languages like Seraki in northeast Nigeria. And that's talking about text models and data. The picture is even bleaker when you're talking about speech models, the sort of things that develop automatic speech recognition, and only a fraction of languages globally are meaningly supported by ASR. So this map here shows the most common primary languages spoken in Borno, Adamawa, and Yobe states in northeast Nigeria. Of the ten languages that we were able to find, only, to our knowledge, only Hausa has any recognition ASR, including a commercial API. But we don't really know how well that it proves. It performs in the wild in actual... practical practice with real -life people. We only know how it works in a laboratory environment. I'm going to talk a bit more about that. That means when you're implementing ASR -assisted workflows that rely on Hausa or English, you are risking excluding most, the vast, you can see, about 70 % of the population of East Nigeria. You're just not reaching them. No communication is going to them. So why? Well, there are a lot of different reasons why. Developing ASR in low -resource languages comes with a lot of challenges, many of which you've just talked about. But most, some languages, and I'm not sure if it's true of Juba Arabic, actually, but think in New Era like this, they're primarily oral. And so manual transcriptions is a really hard and complex... Transcription in order to develop the ASR model. So terminology and ways of expressing oneself vary widely between the intervenants of those languages. There's a lot of code switching. So mixing two languages when you're speaking. We all do that. Any of us who have worked in a number of countries do that fairly regularly. But that's a really hard challenge when you're building digital language technologies. And as I said, lab metrics don't tell us very much about real world performance. So when you're when you're looking at a speech recognition model, you often hear that it has a low word error rate. I have a really hard time saying word error rates to me. Ours in that, I think. And that generally means that some researcher has tested it in a laboratory. But it doesn't really give you a sense of how well it's going. Of the real world. And at the center of this gap. is a lack of really high -quality voice data and text data, to be honest, in the relevant topics. You need hundreds of hours of recorded data done systematically, collected from across diverse intervenants, including by age, by gender, by educational level, and by dialect, if you really want to make an ASR model that works well and widely in the field. For collaboration, across the language development pipeline, we want to build those kinds of partnerships and have cooperation with communities, governments, social impact organizations, and technologists from the beginning, from initial design through data collection to deployment. Just a little bit. Just a little bit on how we do this. We follow a very community -centered approach to data collection. We collaborate with sectors. and community members to really understand how the models are going to be used, how the data is going to be used, to understand linguistic complexities and the barriers to the models before we do any data collection at all. The quote up here is from, highlights one of the problems with tooling that we have. Just like keyboards don't exist in a lot of the languages that we're talking about. We also train community members, 100 ,000 linguists working in 300 languages. And our background is as linguists and translators, so we bring that linguistic expertise to the work that we do. I wanted to shout out to the interpreters and translators in the booths, but they're all gone now because we don't need them. But they're great people. Thank you. Thank you for your work. Yeah, so the tooling is one of the big challenges. So we train the community members as linguists how to do the recording, including an understanding of how and co -design with them the approaches to how you write down in primarily oral language, because there might not be a standard alphabet or a standard way of writing it. So you have to really think that through. In Canary, we saw one recording have 10 different equally accepted ways to write it down. Like they were all fine, but they could not agree amongst the 10 of them which one was the right way. So we use co -design approaches to manage the complexity to achieve a level of consistency that is required for the speech dialects. I have another nice example. I have another example from Congo, from Congolese, but I think I'll skip that in the interest of time, but I'm happy to talk about it. particularly since there's an Ebola response. So informed consent is a big pillar of the work that we do. It's a guiding principle for our platform, and we do this through workshops, community engagement, to make sure that people really understand how the recordings are going to be used and stored, what the risks are around recording their voices and in contributing, and how they can withdraw consent at any time. Absolutely critical for us. Ultimately, all of our data sets are published openly under non -commercial licenses so they can be used across the sector. Okay, this is two more slides. Is that okay? Do I have time? Okay, okay, okay. Okay. So this... This is just... Last year we... because we used to be Translators Without Borders, TWB. And we are, it's a platform that's integrated within our community and within the platform where we manage that community. And we are currently working on those 12 languages up there, which I'm not going to read out loud, in addition to expanding the work that we've already started collecting, where we've collected about 150 hours of voice data in Hausa Kanuri Shua Arabic, for northeast Nigeria. And these are some examples of prompts designed by community members in Kanuri and Kangali Swahili, which we then translated back into English. The prompts are coded by interaction and and disaster type 2. Yeah, that's it. Thank you very much, everybody. I'm happy to answer any questions. I talk really fast.
Barbora Bromová
All right. Thank you very much, Amy. I will wrap us up for the presentations with a very, hopefully very quick example of what that project might look like in terms of specifically addressing a need that a community has, how that might be put, at least my institution is not always very comfortable with, as a logic of working, and how that might actually work online. I'm going to try to give a very quick demo. Please cross your fingers that my internet connection holds for that. But I'm going to take you to Serbia, one of my favorite projects working on linguistic diversity in collaboration with UNDP, Serbia, and the government of Serbia, specifically the National... library, which is facing a very wicked problem. So as opposed to some of the other languages we've been speaking about, Serbian itself is not necessarily the lowest resource in the world. They have a fair amount of text. They're not doing so hot on speech data either, but they actually have quite a lot of historical sources. Serbian came down, their stories, their experiences, a lot of that material has survived, but it is of course on paper. It's not available for data sets or for data scientists to work on. And so we see that Serbian here is at the very bottom of that distribution here. It is also a language among a very complex South Slavic language family, and so there are closely related but distinct languages like Croatian, Montenegrin, and others. So I'm going to slightly complicate some of these exercises. As I said, this was a project developed in collaboration with the National Library of Serbia. The team at the Digitalna Biblioteka is on a mission to digitize 100% of their resources that they have. For years now, they've gotten to 2%, which is a truly commendable effort. But what we've been trying to do is help them leverage some of the scans that they already have. Because we've been finding that at least standard OCR, so that would be optical character recognition models, were having a very hard time with Serbian at large, partly because the language uses both Cyrillic and Latin alphabets interchangeably. And they were additionally having a particularly hard time with the historical scans. Within the Librarify project, we developed a workflow that prepares the image from the scans that the library has made, segments it to distinguish text from any images or headers that the page might have, enhance these segments, go through optical character recognition, do some post-OCR adjustment as well, evaluate, and then do it all again in hopes to recursively improve that process for that particular data type. We were quite successful with this, and in the initial phase digitized over 400 gigabytes and over 16,000 documents from the library's archives, at least those 2% initially digitized. However, we found that of this experience, this is an issue that many languages around the world are facing, and specifically many public institutions are facing. A lot of documents, a lot of data is trapped on paper legacy formats, and we wanted to develop a way to take this workflow, make it a little bit more implementable across different languages, which is how Loria came about. It has four stages, very much corresponding to the ones I already presented, but it is built for reuse and adaptation. It is an open tool published on GitHub for your download and adjustment in case you're interested in deploying it. And very much building on James' idea or core concept of human in the loop, it has been developed with archivists at the center. One thing that we really found very important throughout this process is that one needs to be really fluent not only in the technology and the data, but also in the process. One needs to understand very well what the technology can and cannot do, but they need to understand also very well the institutions that they're working through, the National Library, their processes, their priorities in digitizing some of these. documents in order for all of this to work together and deployment throughout the library. This is a little bit of an overview of the team behind this. We had two machine learning experts, four full-stack development engineers, a design team and a product management team across the National Library of Serbia, the Mathematical Institute at the Academy of Sciences and UNDP, both country office and the global team. Now, I do definitely thank you for your attention. However, I will try to show you Loria, which is the instance that we have spun out in Serbia that I will hopefully be able to connect into so that you can understand a little bit about how this works. So as I said, this is a little bit of a demo. This is a locally deployable solution. What I'm connecting into now is a server at our office in Serbia. I can log in with my details. And here I can see the interface that the archivist would have access to. my colleagues have kindly uploaded a couple samples from the National Library and specifically this is one of my favorites because it is a woman magazine from 1934 in Serbian and while I myself not very familiar with interesting ideas about fashion in particular to give you a little bit of an idea of what the process used to be back in the day without the automations that we were able to implement first we would start with the image enhancement so we could of course adjust the brightness manually here, adjust the sharpness with the sliders down here until the image is better than it was in the initial scan. What we can also do, thanks to Loria, is have this run automatically. I think this might be my connection problem, which I apologize for. But let me just talk you through it. Essentially, while you would be able to, the same way with photos, you would be used to this. Adjust the sharpness, adjust the brightness. the pages individually. We have an algorithm that does this for you. Then we also have a thresholding algorithm. If it would run, that would be perfect. But essentially that allows the picture to be clear enough for those additional segmenting steps to take place. The next one is layout identification. So that is the one that tries to understand the page, pick the text from among these blocks and essentially make it so that the OCR model only focuses on certain parts of the page. After that, there is the OCR round that's ideally specifically adjusted to Serbian or the language that you are using. And then we also have post -OCR correction that you would have seen. it's not loading for me at the moment but something that we've implemented in the latest update is actually a plug -in for frontier models as well which is something that further allows for better post -ocr correction if that is something that you are after or it allows prompting and batch processing across the entire subset of documents that you are navigating that is particularly interesting if you're trying to analyze for a particular characteristic of the text or directly process the data into a into a format different than plain text I wish I could show it to you but maybe it's better that I don't in the interest of time because we have a short discussion to get to and about the international coordination ongoing about this work so let me turn let me actually take this stop the screen share here and turn to my colleague Daphne from UNESCO who will tell us a little bit more about the initiative that actually unites all of these projects perhaps except for our colleagues at Peacekeeping that we can explore whether they want to join which is the Coalition for Linguistic Diversity in AI led by UNESCO so let me hand it over.
Dafna Feinholz
Thank you and I'll try also to be very brief I don't have any slides but I think it's yeah the idea is to look how we can collaborate together and I want to give you to show to you what we have been doing it's a very young project it's really in the process but I think it's interesting and it's capturing many of the things that are said here and I believe that we are all, I mean, this is kind of obvious, but it is about protecting the languages, but also what it is behind them, which is all the culture that it is behind. And I think this is something very important to remind, to remember that this is not only about the language, because this is, the language is just the way in which people express how they view the world, how they understand, and how do they see themselves, the individuals that become important to keep them. And that's why they are also heritage, and that's why it's so important in the work that UNESCO does, and also in the area of AI. So, as you know, we have a recommendation of ethics of AI, and part of it is the protection of diversity and language diversity in particular. Something that is also moving, something that is behind the project is the idea that linguistic and cultural inclusion in AI, in order to be meaningful, as it was already mentioned, together with the respective linguistic and cultural communities across the entire life cycle, because that's part of the issue, no? That we sometimes tend to include them very late in the process, so I think the example that we just heard two intervenants before, it's really from the very beginning including them. And then throughout and continuously after the development and deployment of the tool. So, it requires cooperation with communities, capacity building and open, which is also, I think, very important. So, what we did is that we established an agreement with the Icelandic Ministry of Culture, Innovation and Higher Education. of the Icelandic Centre for Language and Technology. And this is how we established the Coalition for Linguistic Diversity in Artificial Intelligence. This coalition is a bit of a year ago. It started in June 2025. It was launched every year we have a global forum on ethics of artificial intelligence that we are all very welcome to attend. We bring all the different stakeholders doing areas of AI related to ethics. This year will be in Saudi Arabia from the 14th to the 17th of September. And it was launched in Bangkok last week. Today it comprises more than 30 experts working on community -led digital inclusion of languages from almost all the world regions. And of course it continues to actively grow. And so we definitely invite all of you to join. Another important characteristic, of the work of UNESCO is the all the diversity also in cultures in language, of course language cultures but also different perspectives of everything political but also ethical but also economical so this is really about a multi stakeholder perspective this initiative and that's I think one of the key elements so these coalitions bring together a very broad range of stakeholders it brings together governments it brings together academia communities, technical experts international organizations and the private sector so everybody is sitting together in the same group so basically this coalition is a space for these experts of different areas and different communities sitting together in the same room trying to find solutions of how to preserve these languages and all of them trying to figure it out together what will be the best way of preserving it for their own communities but instead of doing it in silos each of them with their own community This has proven to be very useful because there is a lot of sharing of experiences. There is something that some of them know that can be very useful to others. So this is really something that has been very, very, very useful. And most of the time they have seen that what works for one can work for the other. And also the idea of having this multi -stakeholder partnership is because also you have different scales of partners. So you have big companies, but you have also startups from these big companies without the need of hiring someone to develop their own software. So this is also filling a capacity or skill gap. Now, the idea is not only to have these discussions, but to document them. The idea is to be able to document all these good practices that already are there. So we are building a repository. of these good practices. The idea is to include these that have proven already to be successful. And the idea is also, so for the moment, like a page in which all the information is there, but in order to make sure that this information is systematized and is methodologically also coherent, so there is a questionnaire that is designed in order to identify what is the relevant information to be collected from the project. And this questionnaire is co -created between all of the stakeholders and UNESCO because each of the stakeholders are the only ones that know their project, so they will know what is relevant, who to ask. So that's why it's very important that this will be done by them. I'll give you more documents more examples but this repository can make this knowledge accessible because sometimes that's part of the problem that where do you get access to this knowledge it is a living resource because we want this to really be enriched all the time and again it's a place where others can learn but can also build on and so this is why now we are in the way of creating this what I said the questionnaire so the idea is to once that the repository is established and running we're planning to analyze the data identifying the existing trends in the community led approaches to linguistic and cultural dynamics and diversity in AI but also to identify the areas where the actors working in it need more support and more capacity building so these findings will help us to inform policy guidance on linguistic diversity and AI priority area evoked by many member states that now for example at the AI dialogue because it was also mentioned as one of the things that needs to be taken into account it was in fact one of the findings of the preliminary report of the panel so we want to foster discussions among actors but also working with community led linguistic leaders and also really turn all this into collaboration as the introduction of my presentation was and we are more than happy to work with this table
Barbora Bromová
Thank you very much, Daphne, for your presentation and for the work that you and your team have been doing on coordinating some of these projects and the exciting work ongoing around the globe. I wanted to have a couple of minutes for discussion, which we now won't because we will release you for your well -deserved evenings and afternoons. But perhaps if you do indulge me for just a minute more, I would like to actually ask Amy, because I'm conscious we have a lot of international organizations speaking and we are trying our best, serving both our own objectives when it comes to multilingualism and inclusion and trying to capitalize the ecosystems of builders that we are helping and working off of. But Amy is representing an organization that's a little bit looking outside. So if you, as one of the experts in the Coalition on Linguistic Diversity of AI, or I suppose one of your team's experts, if you look at all these efforts and international organizations in the space, what can we do for you? what would be the most impactful thing for us to help you with to focus on together so that we can make a real difference?
Aimee Ansari
Thanks, Barbara. And I think that a lot of what you are doing already is really great and really helpful. We've really enjoyed the UNDP, but also UNESCO and others. One of the biggest things is it costs a lot, right, because you want to be paying people. The volunteers are great. We work with a lot of volunteers. But you want to be paying people fair wages to collect the data. You want to make sure that you're not exploiting people and that you're not extracting language data, which you are then building a model off of and selling back to people. So I think. That's what. One thing we really struggle with is we know that there's this organization over there and this organization over there and that organization over there who are all trying to develop the technology, but they're not coming together. And if we worked together with all five of them, then we can build the data that is relevant for all of them at reasonable cost for each of them, right? Because this is $50 ,000, $100 ,000. Most grassroots organizations don't have that kind of money. But pulling five or six of them together, they could probably find $10 ,000 each. Trying to do that work, it's a lot of work. It's a lot of coordination work, and we really struggle to do that, and we don't have the kind of overview that you do. So that's one thing on a really grassroots level. I think the other thing is around governments building sovereign models. I think the global dialogue on AI governance is a really good start. I don't see a whole lot of norm setting and standard setting around how all communities and languages and cultures are represented in a sovereign model. And it's a hard thing to talk about in a UN space, but we'll take the government of South Sudan because I know that government. There was a civil war. I was in a civil war in South Sudan. The Dinka and the Newer were fighting against each other. Now, if you have a government that is composed only of Dinka, how likely are they to be able to build a sovereign model or want to build a sovereign model that reflects the Newer culture, the Newer language, that doesn't have any bias in it? like that's a hard ask so I think that one of the roles of the UN has to be around trying to make sure that when we're building global public infrastructure that we're building it in a way that is fair and equitable. that's what I
Barbora Bromová
That's what I those are some great points and wonderful material for my outcome report I'll be compiling from this session which I'm sure you all are eagerly awaiting thank you so much I saw a hand, was that a hand? no? okay, is there a question?
Audience
yeah it's like a combination of comments and questions so first I've got to tell myself I'm a translator the interpreters are the cool kids I'm kind of just a translator okay thanks I have sat in that booth sort of on and off all day taking notes, and I can confidently say that this is the coolest session of the week. Okay? By far. By far. Okay? Make sure we get that on recording. It is. It's love. I helped organize this week with the rest of the team, so I had to come out and listen. That's how cool this session is for me. But I also have to kind of out myself because I work in this space for the enemy, sort of. I work not for, I have to say with because I don't represent them, but I work with Meta on building these LLMs, but my team is I18N. We're languages, right? And what's interesting is what you said about the cost, right? So many of them are rooted in English. The same problem every time. It's that, you know, we. the models are rooted in things that we don't think about every day. Whereas my Vietnamese colleague will go, yeah, there's 80 plus pronouns. And if I use one, I know exactly, you'll know exactly who I'm talking about. I'm talking about my mother's, my paternal aunt's cousin or something like that. Right. And they have that and we don't. But you have to have that cultural knowledge. And to get it, you have to have someone who then, like you said, knows how to read it, knows how to write it, knows how to differentiate it, knows how to do it. But so there's like a level, even at the top of the languages that are in green on your slide there. Like my family's Catalan. So what you were saying about, you know, languages not being preserved or actively fought against, like I get it. Right. So it's just the question is, how do you get the people? And then it's the cost. How do you tackle that? Because it's a two cost problem. Right. It becomes either you're fighting to get a model in that language or you're fighting to get people. Who know the language well enough, which is a privileged question. Right. It's a are you. understanding of the language enough, or then it becomes a do you know the language well enough to be able to communicate with a vast number of people, which is impossible, because then you have fading and Romanized Telugu or Romanized Tamil or something like that, right? There's no canonical version that is totally correct every time. You add an extra A into a Hindi word that's written in Roman scripts, and someone will understand it over text, but how do you figure out what's right? I think I'm just asking a lot of questions, actually, but more to the point that I think how do you tackle the question of the two -cost problem, right? And does the private sector have to be involved? Does the private sector have to be involved in helping you get that data and provide it?
Speaker
I'd love to take that question and answer that firstly I'm speaking on behalf of myself not any organization we are also facing a similar problem where the language that we are targeting is primarily oral and it has no written form so if you try and convert that to a written form because of the way that the language is a lingua franca and it's a creolic language it has body substrate and it has tastes and accents of Arabic it has some stuff in it as well the way to build it is with the community and if you build a standardized orthographic structure with the linguists who are experts in the language as well as who are experts in other languages and have that set as a standard perhaps as IEEE or UNESCO it could really help out in making sure that we don't lose the nuance and the context in when we are transcribing it into ASR because one of the problems that she had mentioned as well as ASR doesn't work for low resource languages because it has to be tuned for that particular language every single time. And translate from speech directly to speech. We are getting there. LLMs might not be the way but we have to start with the data we have to start with the people to build it with the people for the people and democratize access to knowledge and not just data. So it could be a solution. Thank you.
Barbora Bromová
A quick note on the private sector involvement. I am again not necessarily speaking on behalf of UNDP but in our work we are of course open to private sector partners however in approaching them they're relatively low interest in doing this specifically on some of the very low resource languages because of the cost that Amy has previously mentioned. There is also then a some consideration in terms of data ownership and data governance questions that are of course open I must say some of our private sector partners are very open minded about these things and we're grateful for that but not all of them by a long mile so we've been trying to try example or also platforms that allow decentralized data sharing and try to sort of get around it that way but it's an open challenge that we're constantly iterating on that being said if you represent the private sector organization who's interested in doing this please reach out we're very open to those conversations.
Aimee Ansari
I would just say on private sector that there's private sector there's like meta and open AI and anthropic and those guys and then there's Viamo Viamo you know so there's private sector and private sector and we work a lot with the Viamo's and the GSMA's and some of those organizations. Lilapa AI is another really interesting that we work with I think the distinction is a little different we have also worked with Meta to test out models, to test out ideas to see if some of their models are working well enough so that they're safe so it's not the enemy. Meta is not the enemy at all I think the question is more around what Barbara was talking about is data governance and how you build that to communities and they're not people people who speak Dinka are not spending millions of dollars online you know no offense but so they're just not commercially interesting.
Barbora Bromová
thank you very much for spending your last session slot with us we really really appreciate it it's been a pleasure, please do reach out after the session recording stopped thank you

Avertissement : Il ne s'agit pas d'un compte rendu officiel de la session. DiploAI génère ces ressources à partir d'enregistrements audiovisuels ; elles sont présentées telles quelles, y compris d'éventuelles erreurs. En raison de contraintes logistiques (audio/vidéo ou transcriptions), les noms peuvent être mal orthographiés. Nous nous efforçons d'être aussi précis que possible.