WSIS Forum 2026
AI-generated report

Better Data for AI – A Possible Task

10 speakers
Summary

This discussion, focused on the challenge of ensuring that data used in AI systems is trustworthy, well-documented, and accessible, with official statistics playing a central role . Moderated by Anu Peltola of UNCTAD, participants explored why data quality matters for AI, how trust can be built, how data systems can be financed, and how new data sources can be responsibly integrated .

Esperanza Magpantay of ITU argued that AI systems risk amplifying biases and reinforcing inequalities when the underlying data lacks quality, transparency, and ethical governance . She emphasised that official statistics, produced under internationally agreed standards and quality assurance frameworks, represent one of the few globally trusted sources of evidence . She also stressed that the future lies not in replacing official statistics with AI or surveys with new data sources, but in combining both responsibly .

Benjamin Rothen of the Swiss Federal Statistical Office introduced the Trusted Data Observatory (TDO), a global metadata discovery platform designed to make trusted data findable and interoperable for both humans and machines, without centralising or transferring ownership of the data . He noted that trusted data is often invisible and not harmonised, and that investing in metadata is essential for machines to locate and use data effectively .

Vibeke Østreich-Nielsen of NORAD highlighted that data investment is frequently sectoral and disconnected from official statistics systems, particularly in the Global South . She outlined four recommendations from the SEVIA Platform for Action: treating data as a public good, improving coordination, strengthening data governance, and investing in national statistical capacity . Daniel Power of Flowminder demonstrated the practical value of mobile phone data during the DRC Ebola outbreak, while cautioning that such data underrepresents women, children, and rural populations, and must be combined with other sources to correct for bias . Alexandre Barbosa of CETIC Brazil presented a self-sustaining financing model for ICT surveys funded through internet domain name registrations, highlighting its advantages of stability, independence, and capacity for innovation .

The discussion concluded with broad agreement that no single actor can deliver trusted data for AI, and that sustained partnerships across governments, international organisations, the private sector, and civil society are essential to advance this agenda at scale .

Keypoints
  • Overall Purpose

  • The discussion aimed to explore how better, more trustworthy data can be produced and made accessible to support responsible AI development. Convened jointly by ITU, UNCTAD, and the UN CCSA, Colombia, Norway, and the UK, the session brought together statisticians, policymakers, donors, and private sector representatives to examine the roles of official statistics, novel data sources, financing mechanisms, and international partnerships in building AI-ready data ecosystems.
  • --
  • Major Discussion Points

  • The quality and trustworthiness of data are the central limiting factor for reliable AI. Panellists consistently emphasised that AI systems are only as dependable as the data underpinning them. Current data used in AI development is frequently fragmented, unevenly documented, and inaccessible for verification, leading to concerns about bias, transparency, and reproducibility. Official statistics, produced under internationally agreed standards and transparent methodologies, were identified as a uniquely reliable foundation, though panellists stressed that trusted data must extend beyond official statistics to include any responsibly governed source.
  • The Trusted Data Observatory (TDO) as a practical mechanism for making trusted data findable and interoperable. Benjamin Rothen described how the arrival of large language models has fundamentally changed how users - including machines - search for data, making visibility and discoverability of trusted datasets a critical challenge. The TDO, led by Switzerland with approximately 20 countries and 30-40 international organisations involved, aims to create a global metadata discovery platform that does not centralise data ownership but enables machines and humans to locate trusted datasets. A key insight was that investing in metadata - descriptive information about datasets - is essential, as AI systems search for text rather than raw data. - Responsible use of novel data sources, particularly mobile phone data, to fill critical information gaps. Daniel Power illustrated how mobile operator data enabled Flowminder to rapidly identify high-risk health zones during an Ebola outbreak in the DRC, with eight out of ten subsequently confirmed zones matching their predictions. However, panellists cautioned that such data underrepresents women, children, rural populations, and the less wealthy, and must be combined with survey data to correct for bias before being used in AI systems. Privacy protections - such as processing data on the operator's premises and exporting only aggregated outputs - were highlighted as non-negotiable upstream safeguards. - Sustainable and diversified financing is essential for building long-term, AI-ready data systems. Vibeke Østreich-Nielsen noted that much existing investment in data, particularly donor funding for the Global South, is sectoral and disconnected from official statistics, rarely producing the continuity needed for robust data infrastructure. Four recommendations from the SEVIA Platform for Action were outlined: treating data as a public good, improving national coordination, strengthening data governance, and investing in national statistical capacity. Alexandre Barbosa presented Brazil's CETIC model - funded entirely through revenues from the .br internet domain registry - as an example of a self-sustaining, independent financing mechanism that has enabled 20 years of continuous, nationally representative ICT surveys. - Partnership across governments, the private sector, academia, and civil society is indispensable. Multiple speakers converged on the view that no single actor can deliver trustworthy, AI-ready data ecosystems alone. ITU's UN-wide initiative on mobile phone data, involving national statistical offices, regulators, and private operators across 25 countries, was cited as a model for multistakeholder collaboration. The question of whether to pay for private data - particularly from mobile operators - was raised as a live tension, with pragmatic, country-specific approaches recommended rather than a single universal policy.
  • --
  • Overall Tone

  • The tone throughout the discussion was constructive, collaborative, and professionally candid. Speakers were open about the scale of the challenges - data gaps, fragmented financing, bias in novel data sources, and the difficulty of defining 'trusted data' - without being pessimistic. There was a consistent undercurrent of cautious optimism, with panellists pointing to concrete initiatives (TDO, CETIC, Flowminder's DRC work, the SEVIA platform) as evidence that progress is achievable. The Q&A segment introduced a slightly more pragmatic and grounded register, with honest acknowledgements of tensions around data monetisation and the preconditions required before open data policies can be effective. The closing remarks reinforced a sense of shared urgency and collective responsibility, with the moderator calling for continued international coordination and emphasising that momentum exists to move the agenda forward.
Speakers Overview
EM
Esperanza Magpantay
141 wpm · 6 min
AP
Anu Peltola
140 wpm · 8 min
BR
Benjamin Rothen
181 wpm · 9 min
Vibeke Østreich-Nielsen
134 wpm · 5 min
DP
Daniel Power
167 wpm · 8 min
AB
Alexandre Barbosa
106 wpm · 6 min
GE
GIZ Egypt Representative
119 wpm · 1 min
JP
Jacques Péguet
120 wpm · 24 s
CD
Côte d'Ivoire Representative
146 wpm · 16 s
M
Moderator
142 wpm · 4 min

Better Data for AI: A Possible Task? - Expanded Summary

#

Overview and Framing

The session, convened jointly by ITU, UNCTAD, and the Committee for the Coordination of Statistical Activities (CCSA) - a body bringing together 45 international and supranational organisations - was moderated by Anu Peltola, Director of UNCTAD Statistics, Data, and Digital Service . The discussion centred on a fundamental challenge: as AI increasingly shapes how information is produced, accessed, and used, the trustworthiness of AI systems is only as strong as the data underpinning them . Peltola framed the problem clearly at the outset, noting that much of the data currently used to develop AI systems is fragmented, uneven in quality, insufficiently documented, and often inaccessible for verification, giving rise to concerns about bias, transparency, and the reproducibility of results . Crucially, she emphasised that this is not solely a matter for official statistics - it concerns any data source used by AI tools - and that the challenge is not merely technical but requires well-thought-through reference frameworks for how AI selects and uses data .

The session was structured to address a sequence of interconnected questions: why better data for AI is needed, how trust can be built, how trusted data can be financed, how new data sources can be responsibly integrated, and how these efforts can be sustained institutionally . The discussion was situated within broader international processes, including the Global Digital Compact, data governance work, and the Financing for Development agenda, all of which emphasise the need for trusted, interoperable, and inclusive data to support sustainable development and digital transformation .

---

#

The Role of Official Statistics in Ensuring Data Quality for AI

Esperanza Magpantay, Senior Statistician at ITU with over three decades of experience in ICT statistics, opened the substantive discussion by reframing the AI debate . While much public discourse focuses on algorithms, models, and computing power, she argued that the real limiting factor is increasingly the quality of the data from which these models learn . For AI to serve the public good, data must be not only abundant but also trusted, representative, well-documented, interoperable, ethically sourced, and governed under clear principles . Without these characteristics, AI systems risk amplifying biases, producing unreliable results, and reinforcing inequalities .

Magpantay identified official statistics as uniquely positioned to address this challenge. For decades, national statistical offices and international organisations have developed rigorous standards for producing high-quality data, using transparent methodologies, internationally agreed definitions, quality assurance frameworks, and strong confidentiality protections . These represent one of the few global sources of trusted evidence that governments, businesses, and citizens can rely upon . At ITU, this is visible in the daily work of measuring digital development - whether monitoring connectivity, tracking SDG progress, or measuring digital inclusion - where AI systems will only be useful if grounded in high-quality official statistics .

Critically, Magpantay rejected a binary framing of the relationship between official statistics and new data sources. The future, she argued, is not about replacing official statistics with AI, nor about replacing surveys with novel data sources, but about combining trusted official statistics with responsibly governed new sources to create richer, more timely, and more relevant data for decision-making . This framing - that combination rather than replacement is the path forward - was echoed by subsequent speakers throughout the session. She also stressed that partnership is becoming essential, as no progress can be made without governments, national statistical offices, regulators, international organisations, academia, the private sector, and civil society working together to build a trusted data ecosystem .

---

#

Building Trusted and AI-Ready Data Ecosystems: The Trusted Data Observatory

Benjamin Rothen, Head of International and National Affairs at the Swiss Federal Statistical Office, introduced the Trusted Data Observatory (TDO) initiative and offered a distinctive diagnosis of the current data problem . He observed that the arrival of large language models - referencing the emergence of ChatGPT in late 2022 as a turning point - has fundamentally changed how users, including machines, search for and interact with data . He made the striking observation that statistical offices must now recognise that "our customers are not humans anymore - often they're machines," underscoring the scale of this shift for official data producers. The challenge is no longer primarily one of data scarcity, but of data invisibility . Trusted data is often not findable, not interoperable, and not harmonised - problems that official statisticians have long worked to address, but which have taken on new urgency in the AI era .

Rothen described the TDO as a response to this challenge. The initiative, led by the Swiss government and involving approximately 20 countries and 30 to 40 international organisations from across the globe, aims to create a Geneva-based global metadata discovery platform . The core principle is that the TDO is not designed to centralise data or transfer ownership - data remains where it is - but to provide a shared discovery layer so that both machines and people can locate trusted datasets . He drew on Switzerland's own decade-long experience building a national metadata platform (known as I14Y, focused on interoperability) as evidence of both the value and the difficulty of this work, acknowledging that even after ten years, the task of making data findable and interoperable remains challenging .

A particularly important technical insight Rothen offered was that investing in metadata - descriptive information about datasets - is essential because large language models do not search for raw data; they search for text . Without rich, descriptive metadata, AI systems will be unable to locate trusted datasets regardless of their quality. He also noted that Switzerland is redesigning its statistical office website by October to be AI-ready, enabling machines to find and use data more effectively . On the question of what constitutes trusted data, Rothen acknowledged that official statistics - governed by frameworks such as the UN Fundamental Principles and the FAIR principles - provide a natural starting point, but argued that on a longer-term basis the TDO should encompass all trusted data, not only statistical data, since in some cases NGO data may be more reliable than national statistical office data . He closed with an open invitation for countries, organisations, and funders to join the TDO initiative, with the ambition of presenting a working prototype at the AI Summit in Geneva in approximately 11 to 12 months .

---

#

Financing Sustainable Data Systems: Structural Challenges and Innovative Solutions

Vibeke Østreich-Nielsen, Senior Adviser at NORAD in Norway and co-lead of the Financing for Development Future of Data Initiative, joined the session remotely and offered a structural critique of how data investment is currently organised globally . Drawing on two decades of experience in strengthening statistical capacity, she observed that because statistics and data are a cross-cutting field, investment tends to be sectoral and project-based rather than systemic . Donor funding for the Global South, in particular, is often not linked to official statistics or to the goal of making data available for broader use, including for AI purposes . This fragmentation means that data systems rarely receive the continuity of investment needed to build long-term statistical capacity .

In response to this structural problem, Østreich-Nielsen described the SEVIA Platform for Action, developed in the context of the Financing for Development agenda and the fourth international conference on Financing for Development . The initiative brought together countries and partners to discuss what needs to be done, producing four recommendations: treating data and statistics as a public good; improving national data and statistical system coordination; strengthening data governance and innovation; and investing in national data and statistics capacity . She emphasised that the first recommendation - treating data as a public good - directly implies that investments in data and statistics should have an end goal of publishing data and making it available in AI-ready formats, which she identified as a major ongoing challenge . She also highlighted that many ministries hold relevant administrative data that is not being made available, representing a significant untapped resource for decision-making and AI, particularly in African contexts where very little data is publicly available . From her position in a donor agency, she noted that practical work is under way to develop guidance for donor representatives on how to evaluate data-related projects and ensure that funded data work is made available in AI-ready formats .

Alexandre Barbosa, Head of the Regional Centre for Studies on the Development of the Information Society (CETIC) in Brazil, offered a concrete and innovative response to the financing challenge . He noted that ICT surveys are frequently financed through mechanisms not based on solid budget commitments - temporary government programmes or donor-funded projects - which rarely provide the continuity needed to build long-term statistical capacity, leading to interrupted time series, lost expertise, and an inability to monitor digital transformation . CETIC's solution, developed over 20 years, is a self-sustainable financing mechanism funded entirely by revenues from the .br country code top-level domain registry, managed through the Brazilian Internet Steering Committee - a multi-stakeholder body comprising government, civil society, academia, and the private sector - and NIC.br . This model has three principal advantages: stability, enabling continuous survey programmes and long-term statistical time series; independence, as CETIC does not receive government or donor funding and can maintain methodological consistency and long-term planning; and innovation, as stable financing allows continuous updating of surveys to address emerging technologies while maintaining international comparability . Barbosa acknowledged, however, that this model is not easily replicable, as Brazil's position as one of the largest domain name databases among G20 and OECD countries provides financial resources that most countries would not have .

---

#

Responsible Use of Novel Data Sources: Mobile Phone Data in Humanitarian Action

Daniel Power, Managing Director of Flowminder Foundation, illustrated the practical value of non-traditional data sources through a detailed case study from the Democratic Republic of Congo . In May, during an Ebola outbreak in the northeast of the country - an already data-scarce region - the immediate question was where the disease would spread next . Drawing on a longstanding eight-year partnership with Vodacom Congo, Flowminder rapidly conducted a cohort study, identifying all subscribers who had been in the outbreak area and tracking which health zones they visited in the days and weeks following the outbreak . All 500-plus health zones in the DRC were ranked according to the intensity of connectivity with the outbreak areas . A week after the first report was released, Ebola was detected in ten more health zones, and eight of those ten were among the top regions Flowminder had identified - demonstrating the timeliness and richness of mobile phone data for humanitarian decision-making .

Power was careful, however, to frame this success within a broader set of upstream and downstream considerations for responsible data use . On the upstream side, subscriber data is inherently sensitive, and Flowminder processes all data on the mobile operator's premises, deploying software that aggregates data before it is exported so that individual-level data never leaves the operator's systems . He emphasised that protecting the privacy of individuals contributing data to these systems is essential, particularly given AI's growing demand for data . On the downstream side, he was candid about the limitations of mobile phone data: it is not fully representative of the population, systematically underrepresenting women, children, the very old, the more rural, and the less wealthy . He argued that these biases must be corrected through complementary survey data before mobile phone data is used in AI systems, and that this applies just as strongly - if not more so - when data feeds into automated analytical systems that may lack the discretion to account for such biases .

Esperanza Magpantay complemented Power's account by describing the ITU and World Bank's initiative to integrate mobile phone data into official statistics sustainably across 25 countries . She described a model in which national statistical offices do not pay mobile operators directly for their data, but instead offer non-monetary incentives - such as sharing technical expertise in data quality assurance and providing access to statistical data that operators cannot otherwise obtain - in exchange for access to mobile phone data . This approach aims to ensure that mobile data is used responsibly and integrated as an official data source rather than treated as a commercial commodity . She confirmed that Côte d'Ivoire is among the 25 participating countries, responding to a question raised in French by a representative from Côte d'Ivoire about how the CETIC model was established and whether it could be replicated .

---

#

The Question of Paying for Data: Tensions and Trade-offs

An audience question from Jacques Péguet raised the issue of whether downstream users should pay for better data . This prompted a revealing exchange that exposed genuine tensions between different institutional perspectives. Rothen took a firm public-good position, stating that in Switzerland data cannot be charged for under law, as it is funded by taxes, though services can be charged for . Power, by contrast, acknowledged a real and live tension: mobile operators are frequently told that their data represents an untapped source of revenue to be monetised, which can make negotiations difficult . He noted pragmatically that Flowminder would pay for data if it helped open a door and respond to a crisis, provided a donor was prepared to support this, but offered an important caveat: the amount paid does not correlate with data quality, and he had seen cases where organisations paid for data that turned out to be of poor quality . He also cited Ghana Statistical Services as an example of a constructive model, noting that in Ghana - where Flowminder has a long-term relationship with the national statistical office - the instruction is not to charge for the data, reflecting a public-good orientation. Magpantay offered a middle path, describing the non-monetary incentive model as preferable to direct payment, with mutual benefit - rather than commercial transaction - as the organising principle . These positions reflect a genuine and unresolved tension between public good principles, commercial incentives, and operational pragmatism that will require further policy and economic analysis to navigate.

---

#

Foundational Preconditions and Practical Challenges

A question from a representative of GIZ Egypt introduced an important grounding perspective, noting that Egypt has recently released an open data policy but that practical implementation requires addressing significant preconditions: digitising existing data, enabling government-to-government data sharing, and incentivising both public and private actors . This observation resonated with points made by multiple panellists. Rothen had already acknowledged that even Switzerland, after a decade of work on its national metadata platform, is still struggling to make data findable and interoperable . Østreich-Nielsen had highlighted that many ministries sit on relevant administrative data that is not being made available . Together, these contributions revealed an implicit tension between the ambition of global initiatives such as the TDO and the foundational capacity gaps that many countries - particularly in the Global South - still face.

Responding directly to the GIZ Egypt question, Peltola offered a constructive reframing: if the TDO held metadata on what data are available globally, it would also reveal data gaps, which could help target investment to where gaps are most critical for policy needs . This reframing of metadata platforms as instruments for identifying and addressing data gaps - not merely for making existing data discoverable - added a strategic dimension to the technical discussion. She also noted that governments committed in the Financing for Development outcome document to investing in their national statistical systems, providing a political foundation upon which to build, though practical implementation varies considerably by country .

---

#

Convergences, Tensions, and Unresolved Questions

Across the session, a high degree of consensus emerged on several foundational principles: that data quality is the central prerequisite for trustworthy AI ; that the future lies in combining official statistics with novel data sources rather than replacing one with the other ; that metadata and data discoverability are critical infrastructure for AI-readiness ; that multi-stakeholder collaboration is indispensable ; and that current financing for data systems is structurally inadequate, particularly in the Global South . A notable area of consensus also emerged around data decentralisation: despite the ambition of the TDO as a global platform, all speakers accepted without challenge that data should remain with its original owners, with only metadata shared globally .

Beneath this surface consensus, however, meaningful tensions remained. The definition of 'trusted data' was acknowledged as genuinely difficult to resolve, with Rothen explicitly stating that he could not provide a complete answer and inviting collaborative work to address it . The question of paying for privately held data exposed divergent institutional positions that reflect deeper structural differences between public statistical institutions and operational organisations working in humanitarian contexts . The replicability of Brazil's financing model was acknowledged as limited by context , and the gap between the TDO's ambitions and the foundational challenges facing many countries was not fully bridged. How AI tools select and use trusted data - rather than fragmented or low-quality sources - when generating outputs or analysis was raised as a concern but not addressed with a concrete solution .

---

#

Conclusions and Next Steps

The session closed with the moderator drawing together the key threads of the discussion . No single actor can deliver trusted data for AI; partnerships across governments, international organisations, the private sector, and civil society are essential. AI tools are powerful and should not be underestimated, but they require good data to produce robust results, and it is critically important to scrutinise the numbers and analysis that AI tools generate. The international statistical community, the geospatial community, and private data ecosystems are actively working on these challenges, and momentum exists to move the agenda forward. Peltola committed to bringing the discussion back to the UN system and the broader network of chief statisticians to consider how to advance the trusted data for AI agenda at scale and connect with partners internationally .

Concrete near-term actions identified during the session include: the TDO's ambition to present a working prototype at the AI Summit in Geneva in approximately 11 to 12 months ; Switzerland's redesign of its statistical office website to be AI-ready by October ; the ITU and World Bank's ongoing work to integrate mobile phone data into official statistics across 25 countries ; and NORAD's development of practical guidance for donor representatives on evaluating data-related projects . The session made clear that while the principles are broadly shared, the hard work of translating them into concrete, country-specific operational frameworks - particularly for data-scarce contexts in the Global South - remains very much ahead.

Anu Peltola
session on Better Data for AI, a Possible Task? This event is organized jointly by ITU and UNCTAD, together with the Committee for the Coordination of Statistical Activities, CCSA, which brings together 45 international and supranational organizations to strengthen and coordinate cooperation in statistics internationally. My name is Anu Peltola. I'm the Director of UNCTAD Statistics, Data, and Digital Service, and it's my pleasure to moderate the discussion today. As AI is increasingly shaping how information is produced, accessed, and used, we are confronted with the challenge that AI is only as trustworthy as the data it uses and shares. Trustworthiness and well -documented and accessible data are not just a technical issue. They are essential foundations for reliable, accountable, and beneficial AI tools. So official statistics, of course, have a role to play in this landscape. Today, much of the data used to develop AI systems is fragmented, uneven in quality, insufficiently documented, and often inaccessible for verification. And as a result, there are concerns about bias, transparency, reproducibility of results when using AI tools. We are now announcing... And we need to ensure the trust. It's not about just official statistics. It's about any data sources that are being used by AI tools where we hope we can find a better way. Official statistics are produced according to internationally agreed standards and methodologies. With continuous public oversight, so that we can have authoritative evidence. and also official statisticians can use any data in society. There's more and more digital data sources. We just need to ensure that the reference of data for AI is well thought through. How will AI tools select what they show as the numbers or what they use in the analysis produced? This discussion is also closely linked to broader international efforts, including the Global Digital Compact, data governance work that is becoming increasingly political and involving many different stakeholders, the financing for development agenda, points to data, the need to invest in data. There are initiatives such as Trusted Data Observatory and the future of data, we will hear about these today. All of these emphasize the need for trusted, interoperable and inclusive data. so that we can support sustainable development, decision -making, digital transformation, and so on. We hope that we can bring together some of the different stakeholders to discuss this issue, statistical geospatial communities, international organizations, the private sector, academia, civil society. We can all work together to see how to share better data for AI. I'll introduce the panelists with some more detail before they each speak. Here I would like to roughly introduce the sequence of the session. So we are going to talk about why we need better data for AI, how to build trust, how to finance trusted data, how to enrich new data sources, and how to sustain all this institutionally and much more. Depending on what you ask. So let me introduce first Esperanza Magpante, Senior Statistician at the International Telecommunication Union. relying on over three decades of experience in statistics. She leads global efforts on ICT statistics and standards and co -chairs major international initiatives on mobile phone data for official statistics. So from a global perspective, why do we need better data for AI? And what is the role of official statistics and partnerships in this effort?
Esperanza Magpantay
Thank you. Thank you so much, Anu, and good afternoon, everyone. So when we speak today about AI, much of the discussion focuses on the use of algorithms, models, and computing power. But the real limiting factor is increasingly the quality of the data that these models learn from. From the official statistics perspective, better data for AI means that the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from the data that we use to learn from. It means data that are not only abundant, but also trusted, representative. Thank you. well -documented, interoperable, ethically sourced, and governed under clear principles. Without these characteristics, AI systems risk amplifying biases, producing unreliable results, and reinforcing inequalities. This is precisely where official statistics have a unique role to play. For decades, national statistical offices and international organizations have developed rigorous standards for producing high -quality data. Official statistics are collected using transparent methodologies, internationally agreed definitions, quality assurance frameworks, and strong confidentiality protections. They represent one of the few global sources of trusted evidence. That governments, businesses, and citizens can rely on. At the ITU, we see this every day in measuring digital development, whether monitoring universal and meaningful connectivity, tracking progress towards sustainable development goals, or measuring digital inclusions, AI systems will only provide useful if they are grounded in high -quality official statistics. Many emerging questions require integrating new data sources, for example, I know mentioned about the work that we are doing on mobile phone data at the UN -wide initiative with traditional statistical systems. These sources have been proven to provide more timely data and more granular data, and of course, they follow governance and transparency measures. And from this experience, we can see that partnership is becoming very, very essential. Because we cannot work without governments, we cannot work without national statistical offices, regulators, international organizations, academia, the private sector, and the civil society in building a trusted data ecosystem. At the ITU, we work together with the experts in the UN Initiative on Mobile Phone Data to bring together partners to brainstorm a methodology and to help countries use this new data source. And these experiences demonstrate an important lesson. The future is not about replacing official statistics with AI, nor is it about replacing surveys with these new data sources. It is about combining. Combining those trusted official statistics with responsibly governed new data sources to create a rich and more timely and more relevant data for decision making. Finally, this is an opportunity for the global statistical community to contribute even more directly to the AI ecosystem. And if we want AI that serves the public good, we need data ecosystems that are built on trust, quality, and collaboration. That is exactly where official statistics and the global partnership have their
Anu Peltola
Thank you, Esperanza. Next, let me turn to Ambassador Benjamin Rotten, Head of International and National Affairs at the Swiss Federal Statistical Office. He leads Switzerland's work on international statistical cooperation, including AI readiness of data and the Trusted Data Observatory initiative. Over to you. The question is, how does the Trusted Data Observatory help build trusted and documented data ecosystems for AI? How do you define what trusted and AI -ready data? What does trusted data look like in practice?
Benjamin Rothen
Fantastic. Thanks, Arno. Thanks for this invitation today. It's a pleasure to be here. And I know you can already see the screen with the QR code, but please don't go yet to the webpage of the TDO. Let me just explain first what's all about. Your question is very relevant. And as we all know, the world changed a bit. So there was in November or December 2022, Chet Chepty come on the planet and everyone is working a bit differently at work. He's searching data and information a bit differently. I guess everyone in the room is doing that, me included, for sure. And the question is, what does that mean for all the official producers of data? And what we have to do, we have to bring out our data to the users. That means also to the machines. Often what we do not lack, there's not a lack of enough data. It's often a lack that our trusted data are not visible. They are not findable. They are often not interoperable. And we have to be very careful about that. they are not harmonized that's all the work that official statisticians know in many years in Switzerland since 165 years we have the task to do that but our work changed completely 2 -3 years ago and that's a huge task for us to change our organization how do we work together in Switzerland but also worldwide together we have since 10 years in Switzerland the task to have like a metadata platform where all the data from the Swiss government are coming together that machines as well as humans can find the right data we're still struggling to do that it's quite a difficult task but that's the idea behind the trusted data observatory it's the idea that we have national metadata platforms that are linked to a global metadata platform and that's the trusted data observatory it's not the idea to centralize data the data stay where they are it's not the idea to take the responsibility or the ownership of the data they stay where they are but we have to show globally a discovery platform that the machines and the people know where to find the trusted data. We all know our customers are not humans anymore. Often they're machines. So what we also do in Switzerland, we're changing our webpage in October completely so it's AI ready. So machines can find the data much better. It's a huge task to do that. It's not easy. It's a huge investment on our side. You asked the question, what are trusted data? We say in the beginning it's often official statistics data because there we have some agreement like the UN fundamental principles. In Europe we have the code of practice and in Switzerland we have a charter, I call it that. And there is of course the FAIR principle. On the long term it's the idea that we have the platform, the TDO, sharing all the trusted data but not only statistical data because we all know... There are also national circumstances, sometimes NGO data much better than national statistical offices data. I mean we all know that. so what we're doing with the TDO at the moment it's an idea led by the Swiss government we have about 20 countries they're working with us from all over the planet we have about 30 to 40 international organizations working with us some are sitting in the room and the idea is this platform has to be Geneva based, that's the idea of our investment, we want to bring all the knowledge that we have already in Geneva together and then enlarging on the global scale we are very convinced that's the way to do, but there's still a long way to go, some of you sitting in the room have seen the first flyer we produced about the TDO and it was about rocket launch going to the stars and sometimes the stars are guiding us the way, but we cannot reach them immediately, so that's a bit where we are at the moment and maybe we have a bit more time later on because there are some people sitting on the panel working very closely with us because we're in a critical situation enlarging the TDO and everyone invited being in the room to be part of that and also working on the question what the trusted data because it's not that easy to answer so I cannot give you the complete answer today but working with all of you hoping answering this question thank you
Anu Peltola
Benjamin I know it was a tricky question a difficult one I think online we are being followed by many people who have registered to listen to our discussions but we are also joined by Vibeke Östreich -Nielsen Senior Advisor at NORAD in Norway she co -leads the Financing for Development Future of Data Initiative and brings two decades of experience in strengthening statistical capacity and data systems you are there ready and listening I'll just share the question with you without sustainable financing better data for AI won't materialize what needs to change in how we fund data systems how can we bring governments private actors and donors together to invest in AI -ready data over to you,
Vibeke Østreich-Nielsen
Thank you very much, Anu. I'm sorry I couldn't be there in person. Are you able to hear and see me? Just a quick check?
Anu Peltola
Yes, yes.
Vibeke Østreich-Nielsen
Great. Thank you. So, yeah, I've been in the statistical system for many years, as you said, and I think one of the realizations that I have had is that because statistics and data are a cross -cutting field, there is a lot of investment in data, but it's often sectoral, particularly when you look at funding and donor funding for the Global South. A lot of the support that goes into data systems is sectoral, and it's not often or always linked to official statistics. The focus for getting data is often your own data needs and not necessarily thinking all the way to making data available for countries in general, but also for... others and for AI purposes. So that's part of the challenge. There's also many other challenges in the data and statistics ecosystem. So in the context of the Financing for Development agenda and the fourth international conference on financing for development last year, we put together this SEVIA Platform for Action initiative that focuses on how we can strengthen national data and statistical systems, also to help governments make a better connection between what we and maybe the statistical world are doing and what their information needs are and how they better can engage and we better can engage with them. And I think that's the first step. There are many other initiatives in the statistics community, but many of them have been maybe more focused within the statistics community and not necessarily... always engaged with decision makers at national level. There are exceptions, of course, but this was at least the effort. And we brought together many different... countries and partners that discussed what needs to be done and came out with four recommendations that we hope can help operationalize the SEVIA commitment for the Financing for Development agenda. And those four are to consider data and statistics as a public good, to improve national data and statistical system coordination, strengthen data governance and innovation, and to invest in national data and statistics capacity. And the first one is very much where I started off by saying investments in data and statistics should also have an end goal of publishing these and make them available in AI -ready formats. From what I have seen, at least, that is still a major challenge. You have many ministries that sit on relevant administrative data that are not being made available. but that could still bring a lot of information both to decision makers, but also in an AI context, maybe particularly in an African context where very little data is publicly available. We also know that there are many challenges in coordinating across different actors. So that's where the second recommendation comes from. The third recommendation is focused on also what this session focuses on, focuses on how we can improve data governance and make governance more organized so that innovation is easier, so that AI finds the information it needs, and also other modern approaches can be more easily moved forward. And then it's the investment in national data and statistics capacity, both from national governments realizing that there is a need to invest in this so that we get more transparent and more... trustworthy information, but also for others to see how we can make this work better. Yeah, so we're hoping that this can be a first step to kind of highlight to a wider community that there is a need to change the approach to work together. But I think also coming now from a donor agency, what we're doing concretely from a donor context is to see how we can maybe further develop practical approaches for representatives in donor organizations on how they look through projects, how they decide on projects, and particularly the data work in this context. But I'll stop here and back to you.
Anu Peltola
At the conference, thank you. Thank you, Rebecca. it's interesting that you mention investment in data in the context of financing for development if you look at the outcome document governments actually committed to investing in their national statistical systems and data so that's something to build on and it depends on country how much we see actually this coming into practice but we are working to support that commitment and if we think of the trusted data observatory the interesting thing is if we would have the metadata of what data are available we would also see the data gaps and this could help us target action into where data gaps are really critical where it's information that policy would really definitely need that would probably help connect the investment and the gaps together next we will hear from Daniel Power, Managing Director from Flowminder Foundation he works on using mobile phone and other novel data sources to generate insights for development and humanitarian action, with a strong focus on responsible data use in addressing global challenges. So, Daniel, let me ask you a question. Humanitarian action often suffers from lack of information that may have critical consequences. Now that we mentioned the word critical and data and gaps, how can mobile phone data help in practice how to ensure responsible data use, avoiding bias and privacy risks? Over to you.
Daniel Power
Great. Thanks very much. Do you hear me okay? Yeah, that's working. Yeah, I'll give an example using this map that I put on the screen. Thanks. And then kind of get into the principles. And, I mean, in any crisis, and we heard this in the previous session as well with Ambassador Jessica Hunter. data is at a premium in a crisis. And that cuts across many different types of data, but a particularly important one, particularly for any crisis which results in displacement or in disease spread, which is what I'm going to talk about here, understanding where people are and how they're moving is critical. And mobile operator data, in fact, several non -traditional data sources can contribute to this. And mobile operator data is a particularly useful one, and that's what we specialize in. Flowminder is a technical operational organization. We don't write policy. And, yeah, so maybe to talk about this example, so in May, probably seen it in the news, there was an Ebola outbreak in eastern DRC in the northeast in an area called the Tauri. And this is a very data -scarce region already. And the immediate question was, you know, where? Where will Ebola spread to next? and there was a lot of speculation about this. And what we were able to do, thanks to our longstanding partnership with Vodacom Congo, we've been working with them for eight years, we very rapidly did a cohort study. We identified all of the subscribers who had been in the area of the outbreak, and we tracked which health zones they visited in the days and weeks following the outbreak. And then we ranked all of the health zones, 500 and something different health zones in DRC, according to the intensity of connectivity with the outbreak areas. And a week after we'd released our first report, so in our first report we said, you know, you need to look very closely at these areas which are highly connected to the outbreak areas where you're not seeing Ebola. Two things, either Ebola's there or it was going to come there very soon. We need to prioritize surveillance in these regions. A week after that report was released, Ebola was detected in ten more health zones, and eight out of those ten were in the top. regions which we had identified using this mobility very rapid, quite straightforward it's not a complicated analysis this one and so that demonstrates the value of this data, it's timeliness and how rich it is, this cohort had hundreds of thousands of subscribers in it now it does come with some points around bias which I'll get onto which are particularly relevant for the question I want to talk actually to split the question to two directions upstream and downstream considerations, so upstream we need to be mindful of where the data comes from this is subscriber data and that's sensitive in and of itself, there's a lot of information in this particular data type about the individual movement of individuals so FlowMinder always processes data at the mobile operators systems so this data has not the individual level data did not leave the Congo, in fact it did not leave Vodacom's premises We deploy software on their premises which aggregates the data before it's exported. And I think it's always important when we're talking about AI and how greedy it is that we protect the privacy of people who are contributing data to these systems. Then I switch to the downstream considerations. Obviously, the data we're producing is used for decisions which affect the population as a whole. And therefore, it should be ideally representative of the population as a whole. Mobile phone data is not. It's pretty good, but it underrepresents women. It underrepresents children, the very old, the more rural, and the less wealthy. So many data types privilege men, wealthy, et cetera. And so if you were to use this, this is a quick analysis where we didn't adjust for any of those biases. When we release our monthly data on population distribution, it is necessary and critical to do so. And so we always... We always endeavor to run surveys. which help us understand how representative the data we are using is of the population of interest. I think that's really important that these data types are not used by themselves. They're combined with other data types to adjust for the bias. And that applies for general use and it applies just as strongly, if not more so, if that data is going to be used for AI, that you take account of those biases before you feed it into a system which may not have the discretion to
Anu Peltola
Thank you, Daniel. I think that's an excellent example of how public good data can also be privately held and can be equally important for our societies and decision makers and help us foresee what will be needed. This leads us to Alexandre Barbosa, Head of the Regional Centre for Studies on the Development of the Information Society in Brazil, leading the production of ICT data for policy and SDG monitoring with extensive experience in survey methodology and digital economy methodology. He is also the head of the International Centre for Studies on the Development of the Information Society in Brazil. He is also the head of the International Centre for Studies on the Development of the Information Society in Brazil. He is also the head of the International Centre for Studies on the Development of the Information Society in Brazil. where he's also chaired the expert group on ICT household indicators. Alex, over to you. For you, I have a specific question. In this AI era, we have huge data gaps on how ICT affects our societies. What is the value of ICT data for national policy in Brazil? And how can we ensure sustainable institutional capacity for production of these statistics?
Alexandre Barbosa
So I would like to highlight how we have found a sustainable financing mechanism in the country to fund a national representative ICT surveys. One of the greatest challenges faced by National Statistical Office, as was already mentioned, is that ICT surveys are frequently financed through mechanisms that are not based on a solid budget commitment. And what we often find is temporarily government programs or donor -funded projects. So while these initiatives are valuable, they rarely provide the continuity needed to build long -term statistical capacity in particular areas of digital economy. And as funding cycles and surveys become irregular, time series are interrupted, expertise is lost, and countries struggle to monitor their own digital transformation. And here is that I would like to give the example of Brazil. In Brazil, the Regional Center for Studies on the Development of the Information Society, CETIC, is running now for 20 years continuous ICT surveys through a self -sustainable financing mechanism. It's funded by the .br country code. It's a top -level domain represented here by two organizations, the Brazilian Internet Steering Committee, which is a multi -stakeholder body composed by the government, civil society, academy, private sector. and NIC .br. NIC .br is the Brazilian Network Information Center, which is responsible for internet domain name registrations, and all the surveys are 100 % made to the government, different ministries, but funded by 100 % for the registry activities for the .br. So this model, I would say, is not very easy to be replicated, because Brazil today ranks sixth largest domain name database among the G20 and OSDD countries, so we are ranked number six right after Russia. The .ru is also very big, and this gives sufficient financial resource to fund the surveys. And we are following all the UNSD and the best practices from the best statistical office and following international frameworks such as the one from the partnership. And this model, I would say, just to conclude my talking, has three major advantages. One is the stability, because it enables continuous survey programs and long -term statistical series that allows policymakers and governments to monitor the digital transformation over time instead of relying on not secure budget to that. The second one is, of course, independence, which is really fundamental. The funding is not limited. We are not tired to involve public budgets. Although we work for the government, we do not receive any funding from the government. or any other individual donor, and we can maintain professional independence, methodological consistence, and long -term planning. And the last and third advantage that I would like to highlight is the dimension of innovation. With a stable finance mechanism, it allows us to continuously update our surveys to address emerging technologies and policy priorities while maintaining international comparability. So besides relying on international agreed standards, we also develop new innovative mechanisms, such as using alternative data source, big data, mobile phone data, to complement the traditional survey data sources. So we have established a laboratory of innovation, innovative data, and new methodologies so that we can really... seek for new data source, new mechanisms for data collection, and the impact of all these three dimensions has been really substantial. Our nationally representative ICT statistics support the design, implementation, and evaluation of public policies on digital inclusion, education, health, several dimensions. We are very much aligned with the dimensions of the SDGs, and this is, I would say, a very innovative model for funding statistical data production. And with that, I would just like to highlight that our data are available in two or three languages, Portuguese, English, and some of them in Spanish. You can access our website and reports
Moderator
Thank you very much. There was an important point about innovation and financing to enable not only production of data as a machinery as you said, what we understand as information in society has changed over the years. It's not that we can just produce statistics and numbers using the same definitions and the same procedures. We have to incorporate innovation into all this. So I promise you can shape the discussion and now we luckily have some time for any questions from the audience. Are there any questions? Yes. We can hear you. The participants may not be hearing you. There's a microphone there.
GIZ Egypt Representative
Thank you. Hello. Thank you. Thank you. I almost forgot my question. Yes, so my question is actually to you. I'm curious because I represent GIZ Egypt, and Egypt has recently released the open data policy, and there's a lot of also conversation on doing data governance while in parallel producing data sets that are open for the industry to develop AI solutions that are really impactful. And GIZ is trying to also support in implementing that. But me as an advisor at GIZ, when I was scoping for such an activity to support, I found some research that says there's a lot of preconditions that need to be in place, and first and foremost digitizing the data that is already there, enabling the governments to speak to each other to begin with, even like with AI. And then G2G, creating similar connected platforms, incentivizing the government, incentivizing the private sector, and so on. So how do we... work around that?
Moderator
I know it's a packed question but we can have a second question. I see there is a queue and then we will try to find some time to answer so please be quick with your question.
Jacques Péguet
Thank you. My name is Jacques Péguet. I'm from Switzerland. Something that has not been touched is downstream money for data. So my question is, is there any consideration on having to pay for better data or is this totally for political reasons whatsoever out of question? Thank you.
Moderator
Any other questions from the audience? Yes, please.
Côte d'Ivoire Representative
Bonjour. Merci pour vos interventions intéressantes. Je suis de la Côte d 'Ivoire et je serais très intéressée par en ce moment de savoir un peu plus sur le modèle de la RTC. Comment vous êtes arrivés à vous mettre en
Moderator
Thank you very much. So we have instant interpretation here in practice, in process. Thanks so much. So I let the panel start answering the questions, I think, at least for Benjamin. There's a question. Yes.
Benjamin Rothen
Thanks for this question. Very difficult. Maybe two or three questions. I don't know. For two, maybe. Very good question and really not easy to answer. Open government data, I mean, that's like the idea behind that you have aggregated data and giving for free to everyone that can work on it. And it's important to say that's aggregated data. So you harmonize the data at the end of the process. what the TDO tries to do we start much earlier to say what kind of data is there, how can we link these different data sets to each other, open government is very important, there's five different steps, how to reach a level and normally I think you have three stories that you are getting published on an open government data platform and that's very important, I think it's always good to invest in all the countries in open government data there's a lot of discussion behind that about the FAIR principles in OGD platforms what we are doing in Switzerland, we are harmonising now we combine the metadata platform we have and it's called I14Y in Switzerland, it calls for interoperability if it's a good name it's another question, but we are changing to metadata .swiss then we could combine open government data .ch I think that's the platform we have together with the metadata platform but the tricky question stays there, I mean how can you bring all this data together what we are trying to do with the TDO is to say all the data sets are visible even though not harmonized data even data set on the micro data level you do not have access to it but you have a description what kind of data sets are out there because that means as a government often we don't know what kind of data sets we have so at least that helps for governments to understand oh there is another ministry they have already data but we do not have access to them that maybe you can start having a conversation I also want to have access maybe we have to change the law that many people can use many organizations can use the data so as Anu said before the TDO helps also to describe a data set in a country or worldwide what's available or at least what exists if the data are available it's often open government data then you have with APIs you can have immediate access but if you are researchers that's not enough you want to have micro data and in Switzerland we have the chance to having a contract and then you're getting access but then it's just a limited access to that so that's a bit what we're doing if I may add here to having all this discussion on OCD platform and more the key word is metadata if we are not investing in metadata you will never be able to find the data sets, so metadata are data about data, so if you have just a figure a number and you don't have the description what this number means and that's the part of the metadata machines will never be able to find the data because the LLMs, they are not looking for data they are looking for text data that's why you have a good description on it I think Anu you said it very well it's not a technical discussion we need to have a technical discussion but at the end it's politics, it's about investing in it and I heard the question a bit the second question was a bit how to pay for it, we as working in national statistical offices and in Switzerland also response for the data system in Switzerland that's a public good, we are getting the money from tax, we cannot charge we can charge for services yes, but not for the data, that's in the Swiss law like that, and it will never change in this, but for the services we should be giving a bit more money to close here to say maybe so the TDO will do a lot of work in the future so if you have money for the TDO just raise your hand and let us know that's fantastic right we hope we can make investment in the next 11 to 12 months that we can show the prototype at the AI summit that's here in Geneva in 11 to 12 months so be part of the TDO initiative, help us to build it up and then we can show to the whole world next year how
Moderator
Thank you very much, I actually had a question if we wanted to see just something concrete as progress in this area, what would it be in the next summit or in the next conference what do we want to see so this is one of the things. We actually have a question to Alexander or Daniel you want to also comment?
Daniel Power
Daniel, yeah. There were a couple of questions there which I wanted to pick up on so maybe just a layer on about the paying for data question first and I don't have enough this is a real tension particularly with mobile operator data because governments maybe have very hard lines around paying for data mobile operators are also being told that their data is a untapped source of revenue that they should be monetizing and coming into these conversations can be quite difficult in terms of expectations. In the Congo we don't pay Vodacom for their data we pay a very small relatively small fee for their server and the management of that server but that's pure cost and you can actually see we know where that invoice goes to so that really isn't paying them for opening up the data to us and in Ghana where we've got a long term relationship with Ghana statistical services tell us don't charge for the data but in other countries we do see operators want to see a return even if it's a modest one for their data and that introduces challenges it is not however the case that I see the amount you pay correlating with data quality it might but it's just not my experience we we work in low and middle income countries and I've definitely seen cases where organisations such as ours have paid for data and it's just been junk really quite poor quality so I'm also not saying you shouldn't pay for data because I think our organisation has to be pragmatic we're not a government, we just want to get the job done and if it helps us open up a door and respond to a particular crisis, well that's something we're prepared to do if our donor, we're always supported by a donor, is also prepared to do it so there's nothing conclusive there just a whole raft of considerations which change and flex for each country you're operating in, which operator you're talking about with respect to DRC thanks for the question I hope it's okay I answer in English and maybe a longer conversation with Esperanza as well because we could follow up on this but we've been very, I have a colleague who's Congolese and is very good at working within the Congo we've got a relationship with a regulator ARP and they've been really critical to have the regulator's support for our work in DRC which is it's a complicated country, you know, this conflict and other considerations that need to be navigated. So having a signal from the regulator that they were comfortable with the work we're doing was really important. But we started with the mobile network operator. For us, as a technical organisation, understand their needs and go from there. And we have had conversations in Cote d 'Ivoire with Orange and the World Bank's global data facility, I think, is operating in Cote d 'Ivoire as well. So those are two avenues where we could continue the
Moderator
Thank you. And Esperanza, you have some final thoughts.
Esperanza Magpantay
Yeah. Thank you so much for the question. I'll address also the question on paying for data. So on the experience of mobile phone data, what we encourage countries, and this is based also on the project that we have with the World Bank, where we are implementing in 25 countries now, where our objective is to integrate mobile phone data. in official statistics and to do it sustainably we have to have all the stakeholders talking to each other and working together. And the idea there is the NSO will not pay for the data but they can provide some incentives with regards to the services that the operators provide because they need to also process the data and they need to pay for those resources. So there are some incentives related to that. But those incentives can also include for example learning from what the National Statistical Office has the technical skills that they have with regards to ensuring data quality or even getting access to some of the data that the National Statistics Office has that operators do not have access to. And so that is the model that we're trying to promote to countries to make sure that they are able to use the data responsibly and at the same time they will have it integrated as one of official data sources. The question from Cote d 'Ivoire, we are happy to share that your country is one of those 25 countries that I mentioned. It's part of the project, and you are benefiting, actually, from the assistance, providing technical assistance with regards to putting all stakeholders together. So I invite you, maybe we can talk after this. There are work that is going on there, and then it can be a fruitful discussion afterwards that you may be able to integrate it in your official statistics. Thank
Moderator
you. Thank you very much. I'm sorry we're out of time. I would love to take more questions, but maybe we can chat after the session. So I won't give the panelists a chance either to react unless you have something really burning. But otherwise, we will continue. Thank you. no single actor can deliver this we need partnerships we need different types of data AI tools are powerful we should not underestimate that it requires good data to ensure that what we get as a result is robust data shapes debates and decisions so it matters what numbers are behind the analysis that we are doing it's so easy to produce outcomes with AI tools but at the same time we must be very critical in here so momentum is here to move this forward it's not easy we don't have easy answers as we heard but the international statistical community the geospatial community private data ecosystem are actively thinking about this moving forward we need to work together with private companies as well and I'm going to take this back to the the UN system and the broader system of the network of chief statisticians so that we can consider internationally how we can best advance this agenda to scale and connect with partners. Thank you very much for your time. I know the doors were locked for a while because of some very important people entering, and some of you came in late, but hopefully you got the chest of the discussion, and I wish you a lovely day and continued engagement.

Disclaimer: This is not an official session record. DiploAI generates these resources from audiovisual recordings, and they are presented as-is, including potential errors. Due to logistical challenges, such as discrepancies in audio/video or transcripts, names may be misspelled. We strive for accuracy to the best of our ability.