WSIS Forum 2026
Rapport généré par l'IA

Better Data for AI – A Possible Task

10 intervenants
Résumé

Résumé

Cette discussion portait sur le défi consistant à garantir que les données utilisées dans les systèmes d'IA soient fiables, bien documentées et accessibles, les statistiques officielles jouant un rôle central à cet égard . Les participants ont examiné pourquoi la qualité des données est importante pour l'IA, comment instaurer la confiance, comment financer les systèmes de données et comment intégrer de nouvelles sources de données de manière responsable . Esperanza Magpantay de l'UIT a soutenu que les systèmes d'IA risquent d'amplifier les biais et de renforcer les inégalités lorsque les données sous-jacentes manquent de qualité, de transparence et de gouvernance éthique . Elle a souligné que les statistiques officielles, produites selon des normes internationalement reconnues et des cadres d'assurance qualité, constituent l'une des rares sources de données probantes bénéficiant d'une confiance mondiale . Elle a également insisté sur le fait que l'avenir ne réside pas dans le remplacement des statistiques officielles par l'IA ou des enquêtes par de nouvelles sources de données, mais dans leur combinaison responsable . Benjamin Rothen de l'Office fédéral de la statistique suisse a présenté le Trusted Data Observatory (TDO), une plateforme mondiale de découverte de métadonnées conçue pour rendre les données fiables trouvables et interopérables, tant pour les humains que pour les machines, sans centraliser ni transférer la propriété des données . Il a noté que les données fiables sont souvent invisibles et non harmonisées, et que l'investissement dans les métadonnées est essentiel pour que les machines puissent localiser et utiliser les données efficacement . Vibeke Østreich-Nielsen de Norad a souligné que les investissements dans les données sont fréquemment sectoriels et déconnectés des systèmes de statistiques officielles, en particulier dans les pays du Sud . Elle a présenté quatre recommandations de la Plateforme d'action SEVIA : traiter les données comme un bien public, améliorer la coordination, renforcer la gouvernance des données et investir dans les capacités statistiques nationales . Daniel Power de Flowminder a démontré la valeur pratique des données de téléphonie mobile lors de l'épidémie d'Ebola en RDC, tout en avertissant que ces données sous-représentent les femmes, les enfants et les populations rurales, et doivent être combinées avec d'autres sources pour corriger les biais . Alexandre Barbosa du Cetic Brésil a présenté un modèle de financement autonome pour les enquêtes TIC, financé par les enregistrements de noms de domaine Internet, en soulignant ses avantages en termes de stabilité, d'indépendance et de capacité d'innovation . La discussion s'est conclue par un large consensus selon lequel aucun acteur unique ne peut fournir des données fiables pour l'IA, et que des partenariats durables entre gouvernements, organisations internationales, secteur privé et société civile sont indispensables pour faire avancer ce programme à grande échelle .

Points clés
  • Points clés

  • Objectif général

  • La discussion visait à explorer comment des données de meilleure qualité et plus fiables peuvent être produites et rendues accessibles pour soutenir un développement responsable de l'IA. Organisée conjointement par l'UIT, la CNUCED et le Comité de coordination des activités statistiques (CCSA), la session a réuni des statisticiens, des décideurs politiques, des donateurs et des représentants du secteur privé pour examiner le rôle des statistiques officielles, des nouvelles sources de données, des mécanismes de financement et des partenariats internationaux dans la construction d'écosystèmes de données prêts pour l'IA.
  • --
  • Principaux points de discussion

  • La qualité et la fiabilité des données constituent le principal facteur limitant pour une IA fiable. Les intervenants ont constamment souligné que les systèmes d'IA ne sont fiables qu'à la mesure des données qui les sous-tendent. Les données actuellement utilisées dans le développement de l'IA sont souvent fragmentées, inégalement documentées et inaccessibles à la vérification, ce qui suscite des préoccupations quant aux biais, à la transparence et à la reproductibilité. Les statistiques officielles, produites selon des normes internationalement reconnues et des méthodologies transparentes, ont été identifiées comme un fondement particulièrement fiable, bien que les intervenants aient insisté sur le fait que les données fiables doivent aller au-delà des statistiques officielles pour inclure toute source gouvernée de manière responsable.
  • Le Trusted Data Observatory (TDO) comme mécanisme pratique pour rendre les données fiables trouvables et interopérables. Benjamin Rothen a décrit comment l'avènement des grands modèles de langage a fondamentalement transformé la manière dont les utilisateurs - y compris les machines - recherchent des données, faisant de la visibilité et de la découvrabilité des ensembles de données fiables un défi crucial. Le TDO, piloté par la Suisse avec la participation d'environ 20 pays et de 30 à 40 organisations internationales, vise à créer une plateforme mondiale de découverte de métadonnées qui ne centralise pas la propriété des données, mais qui permet aux machines et aux humains de localiser des ensembles de données fiables. Un enseignement clé est que l'investissement dans les métadonnées - les informations descriptives sur les ensembles de données - est essentiel, car les systèmes d'IA recherchent du texte plutôt que des données brutes. - Utilisation responsable de nouvelles sources de données, notamment les données de téléphonie mobile, pour combler des lacunes d'information critiques. Daniel Power a illustré comment les données des opérateurs mobiles ont permis à Flowminder d'identifier rapidement les zones à risque sanitaire élevé lors d'une épidémie d'Ebola en RDC, huit des dix zones identifiées ayant été confirmées par la suite. Cependant, les intervenants ont averti que ces données sous-représentent les femmes, les enfants, les populations rurales et les moins aisés, et doivent être combinées avec des données d'enquête pour corriger les biais avant d'être utilisées dans des systèmes d'IA. Les protections de la vie privée - telles que le traitement des données dans les locaux de l'opérateur et l'exportation des seuls résultats agrégés - ont été présentées comme des garanties en amont non négociables. - Un financement durable et diversifié est indispensable pour construire des systèmes de données à long terme, prêts pour l'IA. Vibeke Østreich-Nielsen a noté qu'une grande partie des investissements existants dans les données, notamment les financements des donateurs pour les pays du Sud, est sectorielle et déconnectée des statistiques officielles, produisant rarement la continuité nécessaire à une infrastructure de données robuste. Quatre recommandations de la Plateforme d'action SEVIA ont été présentées : traiter les données comme un bien public, améliorer la coordination nationale, renforcer la gouvernance des données et investir dans les capacités statistiques nationales. Alexandre Barbosa a présenté le modèle Cetic du Brésil - entièrement financé par les revenus du registre de noms de domaine internet .br - comme exemple de mécanisme de financement autonome et indépendant ayant permis 20 ans d'enquêtes TIC continues et représentatives à l'échelle nationale. - Le partenariat entre gouvernements, secteur privé, monde académique et société civile est indispensable. Plusieurs intervenants ont convergé vers l'idée qu'aucun acteur unique ne peut, à lui seul, mettre en place des écosystèmes de données fiables et prêts pour l'IA. L'initiative de l'UIT à l'échelle des Nations Unies sur les données de téléphonie mobile, impliquant des offices nationaux de statistique, des régulateurs et des opérateurs privés dans 25 pays, a été citée comme modèle de collaboration multipartite. La question de savoir s'il convient de rémunérer les données privées - notamment celles des opérateurs mobiles - a été soulevée comme une tension persistante, des approches pragmatiques et propres à chaque pays étant recommandées plutôt qu'une politique universelle unique.
  • --
  • Ton général

  • Le ton de la discussion était constructif, collaboratif et d'une franchise professionnelle. Les intervenants ont abordé l'ampleur des défis : lacunes dans les données, financement fragmenté, biais dans les nouvelles sources de données et difficulté à définir les « données fiables », sans pour autant faire preuve de pessimisme. Un optimisme prudent mais constant a traversé les échanges, les intervenants citant des initiatives concrètes (TDO, Cetic, les travaux de Flowminder en RDC, la plateforme SEVIA) comme preuves que des progrès sont réalisables. La session de questions-réponses a introduit un registre légèrement plus pragmatique et ancré dans la réalité, avec des reconnaissances honnêtes des tensions autour de la monétisation des données et des conditions préalables nécessaires avant que les politiques de données ouvertes puissent être efficaces. Les remarques de clôture ont renforcé un sentiment d'urgence partagé et de responsabilité collective, la modératrice appelant à une coordination internationale continue et soulignant que l'élan existe pour faire avancer ce programme.
Intervenants
EM
Esperanza Magpantay
141 wpm · 6 min
AP
Anu Peltola
140 wpm · 8 min
BR
Benjamin Rothen
181 wpm · 9 min
Vibeke Østreich-Nielsen
134 wpm · 5 min
DP
Daniel Power
167 wpm · 8 min
AB
Alexandre Barbosa
106 wpm · 6 min
GE
GIZ Egypt Representative
119 wpm · 1 min
JP
Jacques Péguet
120 wpm · 24 s
CD
Côte d'Ivoire Representative
146 wpm · 16 s
M
Moderator
142 wpm · 4 min

De meilleures données pour l'IA : une tâche possible ? - Résumé élargi

#

Aperçu général et cadrage

La session, organisée conjointement par l'UIT, la CNUCED et le Comité de coordination des activités statistiques (CCSA) - un organe réunissant 45 organisations internationales et supranationales - était animée par Anu Peltola, Directrice des statistiques, des données et des services numériques de la CNUCED . Les discussions ont porté sur un défi fondamental : à mesure que l'IA façonne de plus en plus la production, l'accès et l'utilisation de l'information, la fiabilité des systèmes d'IA est directement tributaire de la qualité des données qui les sous-tendent . Anu Peltola a posé le problème d'emblée, en soulignant qu'une grande partie des données actuellement utilisées pour développer les systèmes d'IA est fragmentée, inégale en qualité, insuffisamment documentée et souvent inaccessible à des fins de vérification, ce qui soulève des préoccupations en matière de biais, de transparence et de reproductibilité des résultats . Elle a insisté sur le fait qu'il ne s'agit pas uniquement d'une question relevant des statistiques officielles - cela concerne toute source de données utilisée par les outils d'IA - et que le défi n'est pas seulement technique, mais exige des cadres de référence bien pensés sur la manière dont l'IA sélectionne et utilise les données .

La session était structurée autour d'une série de questions interdépendantes : pourquoi de meilleures données pour l'IA sont-elles nécessaires, comment instaurer la confiance, comment financer des données fiables, comment intégrer de nouvelles sources de données de manière responsable, et comment pérenniser ces efforts sur le plan institutionnel . La discussion s'inscrivait dans des processus internationaux plus larges, notamment le Pacte numérique mondial, les travaux sur la gouvernance des données et le programme de financement du développement, qui soulignent tous la nécessité de disposer de données fiables, interopérables et inclusives pour soutenir le développement durable et la transformation numérique .

---

#

Le rôle des statistiques officielles dans la garantie de la qualité des données pour l'IA

Esperanza Magpantay, statisticienne principale à l'UIT forte de plus de trois décennies d'expérience dans les statistiques des TIC, a ouvert la discussion de fond en proposant un nouveau cadrage du débat sur l'IA . Alors que le discours public se concentre largement sur les algorithmes, les modèles et la puissance de calcul, elle a soutenu que le véritable facteur limitant est de plus en plus la qualité des données à partir desquelles ces modèles apprennent . Pour que l'IA serve le bien commun, les données doivent être non seulement abondantes, mais aussi fiables, représentatives, bien documentées, interopérables, obtenues de manière éthique et régies par des principes clairs . Sans ces caractéristiques, les systèmes d'IA risquent d'amplifier les biais, de produire des résultats peu fiables et de renforcer les inégalités .

Esperanza Magpantay a identifié les statistiques officielles comme étant particulièrement bien placées pour relever ce défi. Depuis des décennies, les offices nationaux de statistique et les organisations internationales ont élaboré des normes rigoureuses pour produire des données de haute qualité, en recourant à des méthodologies transparentes, des définitions convenues au niveau international, des cadres d'assurance qualité et de solides protections de la confidentialité . Ces éléments constituent l'une des rares sources mondiales de données probantes fiables sur lesquelles les gouvernements, les entreprises et les citoyens peuvent s'appuyer . À l'UIT, cela se manifeste dans le travail quotidien de mesure du développement numérique - qu'il s'agisse de surveiller la connectivité, de suivre les progrès des ODD ou de mesurer l'inclusion numérique - où les systèmes d'IA ne seront utiles que s'ils reposent sur des statistiques officielles de haute qualité .

Esperanza Magpantay a rejeté une vision binaire de la relation entre les statistiques officielles et les nouvelles sources de données. L'avenir, a-t-elle soutenu, ne consiste pas à remplacer les statistiques officielles par l'IA, ni à substituer des sources de données nouvelles aux enquêtes, mais à combiner des statistiques officielles fiables avec des sources nouvelles gouvernées de manière responsable, afin de créer des données plus riches, plus actuelles et plus pertinentes pour la prise de décision . Ce cadrage - selon lequel la combinaison plutôt que le remplacement est la voie à suivre - a été repris par les intervenants suivants tout au long de la session. Elle a également souligné que le partenariat devient indispensable, car aucun progrès ne peut être accompli sans que les gouvernements, les offices nationaux de statistique, les régulateurs, les organisations internationales, le monde universitaire, le secteur privé et la société civile travaillent ensemble à la construction d'un écosystème de données fiables .

---

#

Construire des écosystèmes de données fiables et prêts pour l'IA : le Trusted Data Observatory

Benjamin Rothen, responsable des affaires internationales et nationales à l'Office fédéral de la statistique suisse, a présenté l'initiative du Trusted Data Observatory (TDO) et proposé un diagnostic original du problème actuel des données . Il a observé que l'émergence des grands modèles de langage - citant l'apparition de ChatGPT fin 2022 comme un tournant - a fondamentalement modifié la manière dont les utilisateurs, y compris les machines, recherchent et interagissent avec les données . Il a formulé l'observation frappante selon laquelle les offices de statistique doivent désormais reconnaître que « nos clients ne sont plus des humains - ce sont souvent des machines », soulignant l'ampleur de ce changement pour les producteurs de données officielles. Le défi n'est plus principalement celui de la rareté des données, mais de leur invisibilité . Les données fiables sont souvent introuvables, non interopérables et non harmonisées - des problèmes que les statisticiens officiels s'efforcent depuis longtemps de résoudre, mais qui ont pris une nouvelle urgence à l'ère de l'IA .

Benjamin Rothen a décrit le TDO comme une réponse à ce défi. L'initiative, pilotée par le gouvernement suisse et impliquant une vingtaine de pays et 30 à 40 organisations internationales du monde entier, vise à créer une plateforme mondiale de découverte de métadonnées basée à Genève . Le principe fondateur est que le TDO n'est pas conçu pour centraliser les données ni pour en transférer la propriété - les données restent là où elles se trouvent - mais pour fournir une couche de découverte partagée permettant aux machines comme aux personnes de localiser des ensembles de données fiables . Il s'est appuyé sur l'expérience de la Suisse, qui a consacré une décennie à la construction d'une plateforme nationale de métadonnées (connue sous le nom d'I14Y, axée sur l'interopérabilité), comme preuve à la fois de la valeur et de la difficulté de ce travail, reconnaissant que même après dix ans, la tâche de rendre les données trouvables et interopérables reste ardue .

Un point technique particulièrement important soulevé par Benjamin Rothen est que l'investissement dans les métadonnées - les informations descriptives sur les ensembles de données - est essentiel, car les grands modèles de langage ne recherchent pas des données brutes ; ils recherchent du texte . Sans métadonnées riches et descriptives, les systèmes d'IA seront incapables de localiser des ensembles de données fiables, quelle que soit leur qualité. Il a également indiqué que la Suisse est en train de refondre le site web de son office de statistique d'ici octobre pour le rendre compatible avec l'IA, permettant aux machines de trouver et d'utiliser les données plus efficacement . Sur la question de ce qui constitue des données fiables, Benjamin Rothen a reconnu que les statistiques officielles - régies par des cadres tels que les Principes fondamentaux des Nations Unies et les principes FAIR - constituent un point de départ naturel, mais a soutenu qu'à plus long terme, le TDO devrait englober toutes les données fiables, et pas seulement les données statistiques, car dans certains cas, les données d'ONG peuvent être plus fiables que celles des offices nationaux de statistique . Il a conclu par une invitation ouverte aux pays, organisations et bailleurs de fonds à rejoindre l'initiative TDO, avec l'ambition de présenter un prototype fonctionnel lors du Sommet sur l'IA à Genève dans environ 11 à 12 mois .

---

#

Financer des systèmes de données durables : défis structurels et solutions innovantes

Vibeke Østreich-Nielsen, conseillère principale à Norad en Norvège et co-responsable de l'initiative Future of Data dans le cadre du financement du développement, a participé à la session à distance et a proposé une critique structurelle de la manière dont les investissements dans les données sont actuellement organisés à l'échelle mondiale . Forte de deux décennies d'expérience dans le renforcement des capacités statistiques, elle a observé que, parce que les statistiques et les données constituent un domaine transversal, les investissements tendent à être sectoriels et fondés sur des projets plutôt que systémiques . Le financement des donateurs pour le Sud global, en particulier, n'est souvent pas lié aux statistiques officielles ni à l'objectif de rendre les données disponibles pour un usage plus large, y compris à des fins d'IA . Cette fragmentation fait que les systèmes de données bénéficient rarement de la continuité d'investissement nécessaire pour renforcer les capacités statistiques à long terme .

En réponse à ce problème structurel, Vibeke Østreich-Nielsen a décrit la Plateforme d'action SEVIA, développée dans le cadre du programme de financement du développement et de la quatrième conférence internationale sur le financement du développement . L'initiative a réuni des pays et des partenaires pour discuter des mesures à prendre, aboutissant à quatre recommandations : traiter les données et les statistiques comme un bien public ; améliorer la coordination nationale des systèmes de données et de statistiques ; renforcer la gouvernance des données et l'innovation ; et investir dans les capacités nationales en matière de données et de statistiques . Elle a souligné que la première recommandation - traiter les données comme un bien public - implique directement que les investissements dans les données et les statistiques doivent avoir pour objectif final de publier les données et de les rendre disponibles dans des formats compatibles avec l'IA, ce qu'elle a identifié comme un défi majeur en cours . Elle a également mis en évidence que de nombreux ministères détiennent des données administratives pertinentes qui ne sont pas rendues disponibles, représentant une ressource inexploitée considérable pour la prise de décision et l'IA, en particulier dans les contextes africains où très peu de données sont accessibles au public . De sa position au sein d'une agence de donateurs, elle a noté que des travaux pratiques sont en cours pour élaborer des orientations à l'intention des représentants des donateurs sur la manière d'évaluer les projets liés aux données et de s'assurer que les travaux sur les données financés sont rendus disponibles dans des formats compatibles avec l'IA .

Alexandre Barbosa, responsable du Centre régional d'études sur le développement de la société de l'information (Cetic) au Brésil, a apporté une réponse concrète et innovante au défi du financement . Il a noté que les enquêtes sur les TIC sont fréquemment financées par des mécanismes ne reposant pas sur des engagements budgétaires solides - programmes gouvernementaux temporaires ou projets financés par des donateurs - qui offrent rarement la continuité nécessaire pour renforcer les capacités statistiques à long terme, entraînant des ruptures dans les séries, des pertes d'expertise et une incapacité à suivre la transformation numérique . La solution développée par Cetic sur 20 ans est un mécanisme de financement auto-durable, entièrement financé par les revenus du registre du domaine de premier niveau national .br, géré par le Comité directeur de l'Internet brésilien - un organe multipartite composé du gouvernement, de la société civile, du monde universitaire et du secteur privé - et NIC.br . Ce modèle présente trois avantages principaux : la stabilité, permettant des programmes d'enquêtes continus et des séries chronologiques statistiques à long terme ; l'indépendance, Cetic ne recevant ni financement gouvernemental ni financement de donateurs et pouvant maintenir une cohérence méthodologique et une planification à long terme ; et l'innovation, un financement stable permettant une mise à jour continue des enquêtes pour aborder les technologies émergentes tout en maintenant la comparabilité internationale . Alexandre Barbosa a toutefois reconnu que ce modèle n'est pas facilement reproductible, car la position du Brésil comme l'un des plus grands registres de noms de domaine parmi les pays du G20 et de l'OCDE lui procure des ressources financières dont la plupart des pays ne disposeraient pas .

---

#

Utilisation responsable de nouvelles sources de données : les données de téléphonie mobile dans l'action humanitaire

Daniel Power, directeur général de la Fondation Flowminder, a illustré la valeur pratique des sources de données non traditionnelles à travers une étude de cas détaillée de la République démocratique du Congo . En mai, lors d'une épidémie d'Ebola dans le nord-est du pays - une région déjà pauvre en données - la question immédiate était de savoir où la maladie se propagerait ensuite . S'appuyant sur un partenariat de longue date de huit ans avec Vodacom Congo, Flowminder a rapidement mené une étude de cohorte, identifiant tous les abonnés qui s'étaient trouvés dans la zone de l'épidémie et suivant les zones de santé qu'ils avaient visitées dans les jours et les semaines suivant le déclenchement de l'épidémie . Les 500 zones sanitaires et plus de la RDC ont été classées en fonction de l'intensité de leur connectivité avec les zones touchées par l'épidémie . Une semaine après la publication du premier rapport, Ebola a été détecté dans dix nouvelles zones sanitaires, dont huit figuraient parmi les régions prioritaires identifiées par Flowminder - démontrant ainsi l'actualité et la richesse des données de téléphonie mobile pour la prise de décision humanitaire .

Daniel Power a toutefois pris soin d'inscrire ce succès dans un ensemble plus large de considérations en amont et en aval pour une utilisation responsable des données . En amont, les données des abonnés sont intrinsèquement sensibles, et Flowminder traite toutes les données dans les locaux de l'opérateur mobile, en déployant un logiciel qui agrège les données avant leur exportation, de sorte que les données individuelles ne quittent jamais les systèmes de l'opérateur . Il a souligné que la protection de la vie privée des personnes contribuant à ces systèmes est essentielle, en particulier compte tenu de la demande croissante de données par l'IA . En aval, il a été franc sur les limites des données de téléphonie mobile : elles ne sont pas pleinement représentatives de la population, sous-représentant systématiquement les femmes, les enfants, les personnes très âgées, les populations rurales et les moins aisés . Il a soutenu que ces biais doivent être corrigés par des données d'enquêtes complémentaires avant que les données de téléphonie mobile ne soient utilisées dans des systèmes d'IA, et que cela s'applique tout autant - sinon plus - lorsque les données alimentent des systèmes analytiques automatisés qui peuvent manquer de discernement pour tenir compte de tels biais .

Esperanza Magpantay a complété l'exposé de Daniel Power en décrivant l'initiative de l'UIT et de la Banque mondiale visant à intégrer les données de téléphonie mobile dans les statistiques officielles de manière durable dans 25 pays . Elle a décrit un modèle dans lequel les offices nationaux de statistique ne paient pas directement les opérateurs mobiles pour leurs données, mais leur offrent plutôt des incitations non monétaires - telles que le partage d'expertise technique en matière d'assurance qualité des données et l'accès à des données statistiques que les opérateurs ne peuvent pas obtenir autrement - en échange de l'accès aux données de téléphonie mobile . Cette approche vise à garantir que les données mobiles sont utilisées de manière responsable et intégrées en tant que source de données officielle plutôt que traitées comme une marchandise commerciale . Elle a confirmé que la Côte d'Ivoire fait partie des 25 pays participants, en réponse à une question posée en français par un représentant de Côte d'Ivoire sur la manière dont le modèle Cetic a été établi et s'il pourrait être reproduit .

---

#

La question du paiement des données : tensions et compromis

Une question du public posée par Jacques Péguet a soulevé la question de savoir si les utilisateurs en aval devraient payer pour de meilleures données . Cela a suscité un échange révélateur qui a mis en lumière de véritables tensions entre différentes perspectives institutionnelles. Benjamin Rothen a adopté une position ferme en faveur du bien public, affirmant qu'en Suisse, les données ne peuvent pas être facturées en vertu de la loi, car elles sont financées par les impôts, bien que des services puissent être facturés . Daniel Power, en revanche, a reconnu une tension réelle et vive : les opérateurs mobiles se voient fréquemment dire que leurs données représentent une source de revenus inexploitée à monétiser, ce qui peut rendre les négociations difficiles . Il a noté pragmatiquement que Flowminder paierait pour des données si cela permettait d'ouvrir une porte et de répondre à une crise, à condition qu'un bailleur de fonds soit prêt à soutenir cette démarche, mais a formulé une mise en garde importante : le montant payé ne correspond pas à la qualité des données, et il existe des cas où des organisations ont payé pour des données qui se sont avérées être de mauvaise qualité . Il a également cité le Ghana Statistical Services comme exemple d'un modèle constructif, notant qu'au Ghana - où Flowminder entretient une relation à long terme avec l'office national de statistique - la consigne est de ne pas facturer les données, reflétant une orientation de bien public. Esperanza Magpantay a proposé une voie médiane, décrivant le modèle d'incitations non monétaires comme préférable au paiement direct, le bénéfice mutuel - plutôt que la transaction commerciale - étant le principe organisateur . Ces positions reflètent une tension réelle et non résolue entre les principes de bien public, les incitations commerciales et le pragmatisme opérationnel, qui nécessitera des analyses politiques et économiques approfondies pour être surmontée.

---

#

Conditions préalables fondamentales et défis pratiques

Une question d'un représentant de GIZ Égypte a introduit une perspective d'ancrage importante, notant que l'Égypte a récemment publié une politique de données ouvertes, mais que la mise en œuvre pratique nécessite de traiter des conditions préalables importantes : numériser les données existantes, permettre le partage de données entre administrations et inciter les acteurs publics et privés . Cette observation a fait écho aux points soulevés par plusieurs panélistes. Benjamin Rothen avait déjà reconnu que même la Suisse, après une décennie de travail sur sa plateforme nationale de métadonnées, peine encore à rendre les données trouvables et interopérables . Vibeke Østreich-Nielsen avait souligné que de nombreux ministères détiennent des données administratives pertinentes qui ne sont pas rendues disponibles . Ensemble, ces contributions ont révélé une tension implicite entre l'ambition d'initiatives mondiales telles que le TDO et les lacunes fondamentales en matière de capacités auxquelles de nombreux pays - en particulier dans le Sud global - sont encore confrontés.

Répondant directement à la question de GIZ Égypte, Anu Peltola a proposé un recadrage constructif : si le TDO détenait des métadonnées sur les données disponibles à l'échelle mondiale, il révélerait également les lacunes en matière de données, ce qui pourrait aider à cibler les investissements là où les lacunes sont les plus critiques pour les besoins politiques . Ce recadrage des plateformes de métadonnées comme instruments d'identification et de comblement des lacunes en matière de données - et pas seulement pour rendre les données existantes découvrables - a ajouté une dimension stratégique à la discussion technique. Elle a également noté que les gouvernements se sont engagés dans le document final sur le financement du développement à investir dans leurs systèmes statistiques nationaux, fournissant une base politique sur laquelle s'appuyer, bien que la mise en œuvre pratique varie considérablement selon les pays .

---

#

Convergences, tensions et questions non résolues

Tout au long de la session, un degré élevé de consensus a émergé sur plusieurs principes fondamentaux : que la qualité des données est la condition préalable centrale à une IA digne de confiance ; que l'avenir réside dans la combinaison des statistiques officielles avec de nouvelles sources de données plutôt que dans le remplacement de l'une par l'autre ; que les métadonnées et la découvrabilité des données constituent une infrastructure critique pour la compatibilité avec l'IA ; que la collaboration multipartite est indispensable ; et que le financement actuel des systèmes de données est structurellement insuffisant, en particulier dans le Sud global . Un domaine notable de consensus a également émergé autour de la décentralisation des données : malgré l'ambition du TDO en tant que plateforme mondiale, tous les intervenants ont accepté sans contestation que les données doivent rester chez leurs propriétaires d'origine, seules les métadonnées étant partagées à l'échelle mondiale .

Sous cette apparente convergence, des tensions significatives subsistaient néanmoins. La définition des « données fiables » a été reconnue comme difficile à résoudre, Benjamin Rothen ayant explicitement déclaré qu'il ne pouvait pas fournir de réponse complète et invitant à un travail collaboratif pour y remédier . La question du paiement des données détenues par des acteurs privés a mis en évidence des positions institutionnelles divergentes qui reflètent des différences structurelles plus profondes entre les institutions statistiques publiques et les organisations opérationnelles travaillant dans des contextes humanitaires . La reproductibilité du modèle de financement brésilien a été reconnue comme limitée par le contexte , et l'écart entre les ambitions du TDO et les défis fondamentaux auxquels de nombreux pays sont confrontés n'a pas été entièrement comblé. La manière dont les outils d'IA sélectionnent et utilisent des données fiables - plutôt que des sources fragmentées ou de mauvaise qualité - lors de la génération de résultats ou d'analyses a été soulevée comme une préoccupation, mais sans qu'une solution concrète soit proposée .

---

#

Conclusions et prochaines étapes

La session s'est conclue par une synthèse des principaux fils de la discussion par la modératrice . Aucun acteur unique ne peut garantir des données fiables pour l'IA ; des partenariats entre gouvernements, organisations internationales, secteur privé et société civile sont indispensables. Les outils d'IA sont puissants et ne doivent pas être sous-estimés, mais ils nécessitent de bonnes données pour produire des résultats robustes, et il est très important d'examiner attentivement les chiffres et les analyses générés par les outils d'IA. La communauté statistique internationale, la communauté géospatiale et les écosystèmes de données privés travaillent activement sur ces défis, et la dynamique existe pour faire avancer le programme. Anu Peltola s'est engagée à porter la discussion au sein du système des Nations Unies et du réseau plus large des statisticiens en chef afin d'examiner comment faire avancer à grande échelle le programme de données fiables pour l'IA et établir des liens avec des partenaires à l'échelle internationale .

Les actions concrètes à court terme identifiées lors de la session comprenaient : l'ambition du TDO de présenter un prototype fonctionnel lors du Sommet sur l'IA à Genève dans environ 11 à 12 mois ; la refonte du site web de l'office de statistique suisse pour le rendre compatible avec l'IA d'ici octobre ; les travaux en cours de l'UIT et de la Banque mondiale pour intégrer les données de téléphonie mobile dans les statistiques officielles de 25 pays ; et l'élaboration par Norad de recommandations pratiques à l'intention des représentants des bailleurs de fonds concernant l'évaluation des projets liés aux données . La session a clairement montré que si les principes sont largement partagés, le travail difficile de leur traduction en cadres opérationnels concrets et adaptés à chaque pays - en particulier pour les contextes pauvres en données du Sud global - reste largement à accomplir.

Anu Peltola
session on Better Data for AI, a Possible Task? This event is organized jointly by ITU and UNCTAD, together with the Committee for the Coordination of Statistical Activities, CCSA, which brings together 45 international and supranational organizations to strengthen and coordinate cooperation in statistics internationally. My name is Anu Peltola. I'm the Director of UNCTAD Statistics, Data, and Digital Service, and it's my pleasure to moderate the discussion today. As AI is increasingly shaping how information is produced, accessed, and used, we are confronted with the challenge that AI is only as trustworthy as the data it uses and shares. Trustworthiness and well -documented and accessible data are not just a technical issue. They are essential foundations for reliable, accountable, and beneficial AI tools. So official statistics, of course, have a role to play in this landscape. Today, much of the data used to develop AI systems is fragmented, uneven in quality, insufficiently documented, and often inaccessible for verification. And as a result, there are concerns about bias, transparency, reproducibility of results when using AI tools. We are now announcing... And we need to ensure the trust. It's not about just official statistics. It's about any data sources that are being used by AI tools where we hope we can find a better way. Official statistics are produced according to internationally agreed standards and methodologies. With continuous public oversight, so that we can have authoritative evidence. and also official statisticians can use any data in society. There's more and more digital data sources. We just need to ensure that the reference of data for AI is well thought through. How will AI tools select what they show as the numbers or what they use in the analysis produced? This discussion is also closely linked to broader international efforts, including the Global Digital Compact, data governance work that is becoming increasingly political and involving many different stakeholders, the financing for development agenda, points to data, the need to invest in data. There are initiatives such as Trusted Data Observatory and the future of data, we will hear about these today. All of these emphasize the need for trusted, interoperable and inclusive data. so that we can support sustainable development, decision -making, digital transformation, and so on. We hope that we can bring together some of the different stakeholders to discuss this issue, statistical geospatial communities, international organizations, the private sector, academia, civil society. We can all work together to see how to share better data for AI. I'll introduce the panelists with some more detail before they each speak. Here I would like to roughly introduce the sequence of the session. So we are going to talk about why we need better data for AI, how to build trust, how to finance trusted data, how to enrich new data sources, and how to sustain all this institutionally and much more. Depending on what you ask. So let me introduce first Esperanza Magpante, Senior Statistician at the International Telecommunication Union. relying on over three decades of experience in statistics. She leads global efforts on ICT statistics and standards and co -chairs major international initiatives on mobile phone data for official statistics. So from a global perspective, why do we need better data for AI? And what is the role of official statistics and partnerships in this effort?
Esperanza Magpantay
Thank you. Thank you so much, Anu, and good afternoon, everyone. So when we speak today about AI, much of the discussion focuses on the use of algorithms, models, and computing power. But the real limiting factor is increasingly the quality of the data that these models learn from. From the official statistics perspective, better data for AI means that the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from. It means data that are not only abundant, but also trusted, representative. Thank you. well -documented, interoperable, ethically sourced, and governed under clear principles. Without these characteristics, AI systems risk amplifying biases, producing unreliable results, and reinforcing inequalities. This is precisely where official statistics have a unique role to play. For decades, national statistical offices and international organizations have developed rigorous standards for producing high -quality data. Official statistics are collected using transparent methodologies, internationally agreed definitions, quality assurance frameworks, and strong confidentiality protections. They represent one of the few global sources of trusted evidence. That governments, businesses, and citizens can rely on. At the ITU, we see this every day in measuring digital development, whether monitoring universal and meaningful connectivity, tracking progress towards sustainable development goals, or measuring digital inclusions, AI systems will only provide useful if they are grounded in high -quality official statistics. Many emerging questions require integrating new data sources, for example, I know mentioned about the work that we are doing on mobile phone data at the UN -wide initiative with traditional statistical systems. These sources have been proven to provide more timely data and more granular data, and of course, they follow governance and transparency measures. And from this experience, we can see that partnership is becoming very, very essential. Because we cannot work without governments, we cannot work without national statistical offices, regulators, international organizations, academia, the private sector, and the civil society in building a trusted data ecosystem. At the ITU, we work together with the experts in the UN Initiative on Mobile Phone Data to bring together partners to brainstorm a methodology and to help countries use this new data source. And these experiences demonstrate an important lesson. The future is not about replacing official statistics with AI, nor is it about replacing surveys with these new data sources. It is about combining. Combining those trusted official statistics with responsibly governed new data sources to create a rich and more timely and more relevant data for decision making. Finally, this is an opportunity for the global statistical community to contribute even more directly to the AI ecosystem. And if we want AI that serves the public good, we need data ecosystems that are built on trust, quality, and collaboration. That is exactly where official statistics and the global partnership have their
Anu Peltola
Thank you, Esperanza. Next, let me turn to Ambassador Benjamin Rotten, Head of International and National Affairs at the Swiss Federal Statistical Office. He leads Switzerland's work on international statistical cooperation, including AI readiness of data and the Trusted Data Observatory initiative. Over to you. The question is, how does the Trusted Data Observatory help build trusted and documented data ecosystems for AI? How do you define what trusted and AI -ready data? What does trusted data look like in practice?
Benjamin Rothen
Fantastic. Thanks, Arno. Thanks for this invitation today. It's a pleasure to be here. And I know you can already see the screen with the QR code, but please don't go yet to the webpage of the TDO. Let me just explain first what's all about. Your question is very relevant. And as we all know, the world changed a bit. So there was in November or December 2022, Chet Chepty come on the planet and everyone is working a bit differently at work. He's searching data and information a bit differently. I guess everyone in the room is doing that, me included, for sure. And the question is, what does that mean for all the official producers of data? And what we have to do, we have to bring out our data to the users. That means also to the machines. Often what we do not lack, there's not a lack of enough data. It's often a lack that our trusted data are not visible. They are not findable. They are often not interoperable. And we have to be very careful about that. they are not harmonized that's all the work that official statisticians know in many years in Switzerland since 165 years we have the task to do that but our work changed completely 2 -3 years ago and that's a huge task for us to change our organization how do we work together in Switzerland but also worldwide together we have since 10 years in Switzerland the task to have like a metadata platform where all the data from the Swiss government are coming together that machines as well as humans can find the right data we're still struggling to do that it's quite a difficult task but that's the idea behind the trusted data observatory it's the idea that we have national metadata platforms that are linked to a global metadata platform and that's the trusted data observatory it's not the idea to centralize data the data stay where they are it's not the idea to take the responsibility or the ownership of the data they stay where they are but we have to show globally a discovery platform that the machines and the people know where to find the trusted data. We all know our customers are not humans anymore. Often they're machines. So what we also do in Switzerland, we're changing our webpage in October completely so it's AI ready. So machines can find the data much better. It's a huge task to do that. It's not easy. It's a huge investment on our side. You asked the question, what are trusted data? We say in the beginning it's often official statistics data because there we have some agreement like the UN fundamental principles. In Europe we have the code of practice and in Switzerland we have a charter, I call it that. And there is of course the FAIR principle. On the long term it's the idea that we have the platform, the TDO, sharing all the trusted data but not only statistical data because we all know... There are also national circumstances, sometimes NGO data much better than national statistical offices data. I mean we all know that. so what we're doing with the TDO at the moment it's an idea led by the Swiss government we have about 20 countries they're working with us from all over the planet we have about 30 to 40 international organizations working with us some are sitting in the room and the idea is this platform has to be Geneva based, that's the idea of our investment, we want to bring all the knowledge that we have already in Geneva together and then enlarging on the global scale we are very convinced that's the way to do, but there's still a long way to go, some of you sitting in the room have seen the first flyer we produced about the TDO and it was about rocket launch going to the stars and sometimes the stars are guiding us the way, but we cannot reach them immediately, so that's a bit where we are at the moment and maybe we have a bit more time later on because there are some people sitting on the panel working very closely with us because we're in a critical situation enlarging the TDO and everyone invited being in the room to be part of that and also working on the question what the trusted data because it's not that easy to answer so I cannot give you the complete answer today but working with all of you hoping answering this question thank you
Anu Peltola
Benjamin I know it was a tricky question a difficult one I think online we are being followed by many people who have registered to listen to our discussions but we are also joined by Vibeke Östreich -Nielsen Senior Advisor at NORAD in Norway she co -leads the Financing for Development Future of Data Initiative and brings two decades of experience in strengthening statistical capacity and data systems you are there ready and listening I'll just share the question with you without sustainable financing better data for AI won't materialize what needs to change in how we fund data systems how can we bring governments private actors and donors together to invest in AI -ready data over to you,
Vibeke Østreich-Nielsen
Thank you very much, Anu. I'm sorry I couldn't be there in person. Are you able to hear and see me? Just a quick check?
Anu Peltola
Yes, yes.
Vibeke Østreich-Nielsen
Great. Thank you. So, yeah, I've been in the statistical system for many years, as you said, and I think one of the realizations that I have had is that because statistics and data are a cross -cutting field, there is a lot of investment in data, but it's often sectoral, particularly when you look at funding and donor funding for the Global South. A lot of the support that goes into data systems is sectoral, and it's not often or always linked to official statistics. The focus for getting data is often your own data needs and not necessarily thinking all the way to making data available for countries in general, but also for... others and for AI purposes. So that's part of the challenge. There's also many other challenges in the data and statistics ecosystem. So in the context of the Financing for Development agenda and the fourth international conference on financing for development last year, we put together this SEVIA Platform for Action initiative that focuses on how we can strengthen national data and statistical systems, also to help governments make a better connection between what we and maybe the statistical world are doing and what their information needs are and how they better can engage and we better can engage with them. And I think that's the first step. There are many other initiatives in the statistics community, but many of them have been maybe more focused within the statistics community and not necessarily... always engaged with decision makers at national level. There are exceptions, of course, but this was at least the effort. And we brought together many different... countries and partners that discussed what needs to be done and came out with four recommendations that we hope can help operationalize the SEVIA commitment for the Financing for Development agenda. And those four are to consider data and statistics as a public good, to improve national data and statistical system coordination, strengthen data governance and innovation, and to invest in national data and statistics capacity. And the first one is very much where I started off by saying investments in data and statistics should also have an end goal of publishing these and make them available in AI -ready formats. From what I have seen, at least, that is still a major challenge. You have many ministries that sit on relevant administrative data that are not being made available. but that could still bring a lot of information both to decision makers, but also in an AI context, maybe particularly in an African context where very little data is publicly available. We also know that there are many challenges in coordinating across different actors. So that's where the second recommendation comes from. The third recommendation is focused on also what this session focuses on, focuses on how we can improve data governance and make governance more organized so that innovation is easier, so that AI finds the information it needs, and also other modern approaches can be more easily moved forward. And then it's the investment in national data and statistics capacity, both from national governments realizing that there is a need to invest in this so that we get more transparent and more... trustworthy information, but also for others to see how we can make this work better. Yeah, so we're hoping that this can be a first step to kind of highlight to a wider community that there is a need to change the approach to work together. But I think also coming now from a donor agency, what we're doing concretely from a donor context is to see how we can maybe further develop practical approaches for representatives in donor organizations on how they look through projects, how they decide on projects, and particularly the data work in this context. But I'll stop here and back to you.
Anu Peltola
At the conference, thank you. Thank you, Rebecca. it's interesting that you mention investment in data in the context of financing for development if you look at the outcome document governments actually committed to investing in their national statistical systems and data so that's something to build on and it depends on country how much we see actually this coming into practice but we are working to support that commitment and if we think of the trusted data observatory the interesting thing is if we would have the metadata of what data are available we would also see the data gaps and this could help us target action into where data gaps are really critical where it's information that policy would really definitely need that would probably help connect the investment and the gaps together next we will hear from Daniel Power, Managing Director from Flowminder Foundation he works on using mobile phone and other novel data sources to generate insights for development and humanitarian action, with a strong focus on responsible data use in addressing global challenges. So, Daniel, let me ask you a question. Humanitarian action often suffers from lack of information that may have critical consequences. Now that we mentioned the word critical and data and gaps, how can mobile phone data help in practice how to ensure responsible data use, avoiding bias and privacy risks? Over to you.
Daniel Power
Great. Thanks very much. Do you hear me okay? Yeah, that's working. Yeah, I'll give an example using this map that I put on the screen. Thanks. And then kind of get into the principles. And, I mean, in any crisis, and we heard this in the previous session as well with Ambassador Jessica Hunter. data is at a premium in a crisis. And that cuts across many different types of data, but a particularly important one, particularly for any crisis which results in displacement or in disease spread, which is what I'm going to talk about here, understanding where people are and how they're moving is critical. And mobile operator data, in fact, several non -traditional data sources can contribute to this. And mobile operator data is a particularly useful one, and that's what we specialize in. Flowminder is a technical operational organization. We don't write policy. And, yeah, so maybe to talk about this example, so in May, probably seen it in the news, there was an Ebola outbreak in eastern DRC in the northeast in an area called the Tauri. And this is a very data -scarce region already. And the immediate question was, you know, where? Where will Ebola spread to next? and there was a lot of speculation about this. And what we were able to do, thanks to our longstanding partnership with Vodacom Congo, we've been working with them for eight years, we very rapidly did a cohort study. We identified all of the subscribers who had been in the area of the outbreak, and we tracked which health zones they visited in the days and weeks following the outbreak. And then we ranked all of the health zones, 500 and something different health zones in DRC, according to the intensity of connectivity with the outbreak areas. And a week after we'd released our first report, so in our first report we said, you know, you need to look very closely at these areas which are highly connected to the outbreak areas where you're not seeing Ebola. Two things, either Ebola's there or it was going to come there very soon. We need to prioritize surveillance in these regions. A week after that report was released, Ebola was detected in ten more health zones, and eight out of those ten were in the top. regions which we had identified using this mobility very rapid, quite straightforward it's not a complicated analysis this one and so that demonstrates the value of this data, it's timeliness and how rich it is, this cohort had hundreds of thousands of subscribers in it now it does come with some points around bias which I'll get onto which are particularly relevant for the question I want to talk actually to split the question to two directions upstream and downstream considerations, so upstream we need to be mindful of where the data comes from this is subscriber data and that's sensitive in and of itself, there's a lot of information in this particular data type about the individual movement of individuals so FlowMinder always processes data at the mobile operators systems so this data has not the individual level data did not leave the Congo, in fact it did not leave Vodacom's premises We deploy software on their premises which aggregates the data before it's exported. And I think it's always important when we're talking about AI and how greedy it is that we protect the privacy of people who are contributing data to these systems. Then I switch to the downstream considerations. Obviously, the data we're producing is used for decisions which affect the population as a whole. And therefore, it should be ideally representative of the population as a whole. Mobile phone data is not. It's pretty good, but it underrepresents women. It underrepresents children, the very old, the more rural, and the less wealthy. So many data types privilege men, wealthy, et cetera. And so if you were to use this, this is a quick analysis where we didn't adjust for any of those biases. When we release our monthly data on population distribution, it is necessary and critical to do so. And so we always... We always endeavor to run surveys. which help us understand how representative the data we are using is of the population of interest. I think that's really important that these data types are not used by themselves. They're combined with other data types to adjust for the bias. And that applies for general use and it applies just as strongly, if not more so, if that data is going to be used for AI, that you take account of those biases before you feed it into a system which may not have the discretion to
Anu Peltola
Thank you, Daniel. I think that's an excellent example of how public good data can also be privately held and can be equally important for our societies and decision makers and help us foresee what will be needed. This leads us to Alexandre Barbosa, Head of the Regional Centre for Studies on the Development of the Information Society in Brazil, leading the production of ICT data for policy and SDG monitoring with extensive experience in survey methodology and digital economy methodology. He is also the head of the International Centre for Studies on the Development of the Information Society in Brazil. He is also the head of the International Centre for Studies on the Development of the Information Society in Brazil. He is also the head of the International Centre for Studies on the Development of the Information Society in Brazil. where he's also chaired the expert group on ICT household indicators. Alex, over to you. For you, I have a specific question. In this AI era, we have huge data gaps on how ICT affects our societies. What is the value of ICT data for national policy in Brazil? And how can we ensure sustainable institutional capacity for production of these statistics?
Alexandre Barbosa
So I would like to highlight how we have found a sustainable financing mechanism in the country to fund a national representative ICT surveys. One of the greatest challenges faced by National Statistical Office, as was already mentioned, is that ICT surveys are frequently financed through mechanisms that are not based on a solid budget commitment. And what we often find is temporarily government programs or donor -funded projects. So while these initiatives are valuable, they rarely provide the continuity needed to build long -term statistical capacity in particular areas of digital economy. And as funding cycles and surveys become irregular, time series are interrupted, expertise is lost, and countries struggle to monitor their own digital transformation. And here is that I would like to give the example of Brazil. In Brazil, the Regional Center for Studies on the Development of the Information Society, CETIC, is running now for 20 years continuous ICT surveys through a self -sustainable financing mechanism. It's funded by the .br country code. It's a top -level domain represented here by two organizations, the Brazilian Internet Steering Committee, which is a multi -stakeholder body composed by the government, civil society, academy, private sector. and NIC .br. NIC .br is the Brazilian Network Information Center, which is responsible for internet domain name registrations, and all the surveys are 100 % made to the government, different ministries, but funded by 100 % for the registry activities for the .br. So this model, I would say, is not very easy to be replicated, because Brazil today ranks sixth largest domain name database among the G20 and OSDD countries, so we are ranked number six right after Russia. The .ru is also very big, and this gives sufficient financial resource to fund the surveys. And we are following all the UNSD and the best practices from the best statistical office and following international frameworks such as the one from the partnership. And this model, I would say, just to conclude my talking, has three major advantages. One is the stability, because it enables continuous survey programs and long -term statistical series that allows policymakers and governments to monitor the digital transformation over time instead of relying on not secure budget to that. The second one is, of course, independence, which is really fundamental. The funding is not limited. We are not tired to involve public budgets. Although we work for the government, we do not receive any funding from the government. or any other individual donor, and we can maintain professional independence, methodological consistence, and long -term planning. And the last and third advantage that I would like to highlight is the dimension of innovation. With a stable finance mechanism, it allows us to continuously update our surveys to address emerging technologies and policy priorities while maintaining international comparability. So besides relying on international agreed standards, we also develop new innovative mechanisms, such as using alternative data source, big data, mobile phone data, to complement the traditional survey data sources. So we have established a laboratory of innovation, innovative data, and new methodologies so that we can really... seek for new data source, new mechanisms for data collection, and the impact of all these three dimensions has been really substantial. Our nationally representative ICT statistics support the design, implementation, and evaluation of public policies on digital inclusion, education, health, several dimensions. We are very much aligned with the dimensions of the SDGs, and this is, I would say, a very innovative model for funding statistical data production. And with that, I would just like to highlight that our data are available in two or three languages, Portuguese, English, and some of them in Spanish. You can access our website and reports
Moderator
Thank you very much. There was an important point about innovation and financing to enable not only production of data as a machinery as you said, what we understand as information in society has changed over the years. It's not that we can just produce statistics and numbers using the same definitions and the same procedures. We have to incorporate innovation into all this. So I promise you can shape the discussion and now we luckily have some time for any questions from the audience. Are there any questions? Yes. We can hear you. The participants may not be hearing you. There's a microphone there.
GIZ Egypt Representative
Thank you. Hello. Thank you. Thank you. I almost forgot my question. Yes, so my question is actually to you. I'm curious because I represent GIZ Egypt, and Egypt has recently released the open data policy, and there's a lot of also conversation on doing data governance while in parallel producing data sets that are open for the industry to develop AI solutions that are really impactful. And GIZ is trying to also support in implementing that. But me as an advisor at GIZ, when I was scoping for such an activity to support, I found some research that says there's a lot of preconditions that need to be in place, and first and foremost digitizing the data that is already there, enabling the governments to speak to each other to begin with, even like with AI. And then G2G, creating similar connected platforms, incentivizing the government, incentivizing the private sector, and so on. So how do we... work around that?
Moderator
I know it's a packed question but we can have a second question. I see there is a queue and then we will try to find some time to answer so please be quick with your question.
Jacques Péguet
Thank you. My name is Jacques Péguet. I'm from Switzerland. Something that has not been touched is downstream money for data. So my question is, is there any consideration on having to pay for better data or is this totally for political reasons whatsoever out of question? Thank you.
Moderator
Any other questions from the audience? Yes, please.
Côte d'Ivoire Representative
Bonjour. Merci pour vos interventions intéressantes. Je suis de la Côte d 'Ivoire et je serais très intéressée par en ce moment de savoir un peu plus sur le modèle de la RTC. Comment vous êtes arrivés à vous mettre en
Moderator
Thank you very much. So we have instant interpretation here in practice, in process. Thanks so much. So I let the panel start answering the questions, I think, at least for Benjamin. There's a question. Yes.
Benjamin Rothen
Thanks for this question. Very difficult. Maybe two or three questions. I don't know. For two, maybe. Very good question and really not easy to answer. Open government data, I mean, that's like the idea behind that you have aggregated data and giving for free to everyone that can work on it. And it's important to say that's aggregated data. So you harmonize the data at the end of the process. what the TDO tries to do we start much earlier to say what kind of data is there, how can we link these different data sets to each other, open government is very important, there's five different steps, how to reach a level and normally I think you have three stories that you are getting published on an open government data platform and that's very important, I think it's always good to invest in all the countries in open government data there's a lot of discussion behind that about the FAIR principles in OGD platforms what we are doing in Switzerland, we are harmonising now we combine the metadata platform we have and it's called I14Y in Switzerland, it calls for interoperability if it's a good name it's another question, but we are changing to metadata .swiss then we could combine open government data .ch I think that's the platform we have together with the metadata platform but the tricky question stays there, I mean how can you bring all this data together what we are trying to do with the TDO is to say all the data sets are visible even though not harmonized data even data set on the micro data level you do not have access to it but you have a description what kind of data sets are out there because that means as a government often we don't know what kind of data sets we have so at least that helps for governments to understand oh there is another ministry they have already data but we do not have access to them that maybe you can start having a conversation I also want to have access maybe we have to change the law that many people can use many organizations can use the data so as Anu said before the TDO helps also to describe a data set in a country or worldwide what's available or at least what exists if the data are available it's often open government data then you have with APIs you can have immediate access but if you are researchers that's not enough you want to have micro data and in Switzerland we have the chance to having a contract and then you're getting access but then it's just a limited access to that so that's a bit what we're doing if I may add here to having all this discussion on OCD platform and more the key word is metadata if we are not investing in metadata you will never be able to find the data sets, so metadata are data about data, so if you have just a figure a number and you don't have the description what this number means and that's the part of the metadata machines will never be able to find the data because the LLMs, they are not looking for data they are looking for text data that's why you have a good description on it I think Anu you said it very well it's not a technical discussion we need to have a technical discussion but at the end it's politics, it's about investing in it and I heard the question a bit the second question was a bit how to pay for it, we as working in national statistical offices and in Switzerland also response for the data system in Switzerland that's a public good, we are getting the money from tax, we cannot charge we can charge for services yes, but not for the data, that's in the Swiss law like that, and it will never change in this, but for the services we should be giving a bit more money to close here to say maybe so the TDO will do a lot of work in the future so if you have money for the TDO just raise your hand and let us know that's fantastic right we hope we can make investment in the next 11 to 12 months that we can show the prototype at the AI summit that's here in Geneva in 11 to 12 months so be part of the TDO initiative, help us to build it up and then we can show to the whole world next year how
Moderator
Thank you very much, I actually had a question if we wanted to see just something concrete as progress in this area, what would it be in the next summit or in the next conference what do we want to see so this is one of the things. We actually have a question to Alexander or Daniel you want to also comment?
Daniel Power
Daniel, yeah. There were a couple of questions there which I wanted to pick up on so maybe just a layer on about the paying for data question first and I don't have enough this is a real tension particularly with mobile operator data because governments maybe have very hard lines around paying for data mobile operators are also being told that their data is a untapped source of revenue that they should be monetizing and coming into these conversations can be quite difficult in terms of expectations. In the Congo we don't pay Vodacom for their data we pay a very small relatively small fee for their server and the management of that server but that's pure cost and you can actually see we know where that invoice goes to so that really isn't paying them for opening up the data to us and in Ghana where we've got a long term relationship with Ghana statistical services tell us don't charge for the data but in other countries we do see operators want to see a return even if it's a modest one for their data and that introduces challenges it is not however the case that I see the amount you pay correlating with data quality it might but it's just not my experience we we work in low and middle income countries and I've definitely seen cases where organisations such as ours have paid for data and it's just been junk really quite poor quality so I'm also not saying you shouldn't pay for data because I think our organisation has to be pragmatic we're not a government, we just want to get the job done and if it helps us open up a door and respond to a particular crisis, well that's something we're prepared to do if our donor, we're always supported by a donor, is also prepared to do it so there's nothing conclusive there just a whole raft of considerations which change and flex for each country you're operating in, which operator you're talking about with respect to DRC thanks for the question I hope it's okay I answer in English and maybe a longer conversation with Esperanza as well because we could follow up on this but we've been very, I have a colleague who's Congolese and is very good at working within the Congo we've got a relationship with a regulator ARP and they've been really critical to have the regulator's support for our work in DRC which is it's a complicated country, you know, this conflict and other considerations that need to be navigated. So having a signal from the regulator that they were comfortable with the work we're doing was really important. But we started with the mobile network operator. For us, as a technical organisation, understand their needs and go from there. And we have had conversations in Cote d 'Ivoire with Orange and the World Bank's global data facility, I think, is operating in Cote d 'Ivoire as well. So those are two avenues where we could continue the
Moderator
Thank you. And Esperanza, you have some final thoughts.
Esperanza Magpantay
Yeah. Thank you so much for the question. I'll address also the question on paying for data. So on the experience of mobile phone data, what we encourage countries, and this is based also on the project that we have with the World Bank, where we are implementing in 25 countries now, where our objective is to integrate mobile phone data. in official statistics and to do it sustainably we have to have all the stakeholders talking to each other and working together. And the idea there is the NSO will not pay for the data but they can provide some incentives with regards to the services that the operators provide because they need to also process the data and they need to pay for those resources. So there are some incentives related to that. But those incentives can also include for example learning from what the National Statistical Office has the technical skills that they have with regards to ensuring data quality or even getting access to some of the data that the National Statistics Office has that operators do not have access to. And so that is the model that we're trying to promote to countries to make sure that they are able to use the data responsibly and at the same time they will have it integrated as one of official data sources. The question from Cote d 'Ivoire, we are happy to share that your country is one of those 25 countries that I mentioned. It's part of the project, and you are benefiting, actually, from the assistance, providing technical assistance with regards to putting all stakeholders together. So I invite you, maybe we can talk after this. There are work that is going on there, and then it can be a fruitful discussion afterwards that you may be able to integrate it in your official statistics. Thank
Moderator
you. Thank you very much. I'm sorry we're out of time. I would love to take more questions, but maybe we can chat after the session. So I won't give the panelists a chance either to react unless you have something really burning. But otherwise, we will continue. Thank you. no single actor can deliver this we need partnerships we need different types of data AI tools are powerful we should not underestimate that it requires good data to ensure that what we get as a result is robust data shapes debates and decisions so it matters what numbers are behind the analysis that we are doing it's so easy to produce outcomes with AI tools but at the same time we must be very critical in here so momentum is here to move this forward it's not easy we don't have easy answers as we heard but the international statistical community the geospatial community private data ecosystem are actively thinking about this moving forward we need to work together with private companies as well and I'm going to take this back to the the UN system and the broader system of the network of chief statisticians so that we can consider internationally how we can best advance this agenda to scale and connect with partners. Thank you very much for your time. I know the doors were locked for a while because of some very important people entering, and some of you came in late, but hopefully you got the chest of the discussion, and I wish you a lovely day and continued engagement.

Avertissement : Il ne s'agit pas d'un compte rendu officiel de la session. DiploAI génère ces ressources à partir d'enregistrements audiovisuels ; elles sont présentées telles quelles, y compris d'éventuelles erreurs. En raison de contraintes logistiques (audio/vidéo ou transcriptions), les noms peuvent être mal orthographiés. Nous nous efforçons d'être aussi précis que possible.