This discussion, focused on the challenge of ensuring that data used in AI systems is trustworthy, well-documented, and accessible, with official statistics playing a central role . Moderated by Anu Peltola of UNCTAD, participants explored why data quality matters for AI, how trust can be built, how data systems can be financed, and how new data sources can be responsibly integrated .
Esperanza Magpantay of ITU argued that AI systems risk amplifying biases and reinforcing inequalities when the underlying data lacks quality, transparency, and ethical governance . She emphasised that official statistics, produced under internationally agreed standards and quality assurance frameworks, represent one of the few globally trusted sources of evidence . She also stressed that the future lies not in replacing official statistics with AI or surveys with new data sources, but in combining both responsibly .
Benjamin Rothen of the Swiss Federal Statistical Office introduced the Trusted Data Observatory (TDO), a global metadata discovery platform designed to make trusted data findable and interoperable for both humans and machines, without centralising or transferring ownership of the data . He noted that trusted data is often invisible and not harmonised, and that investing in metadata is essential for machines to locate and use data effectively .
Vibeke Østreich-Nielsen of NORAD highlighted that data investment is frequently sectoral and disconnected from official statistics systems, particularly in the Global South . She outlined four recommendations from the SEVIA Platform for Action: treating data as a public good, improving coordination, strengthening data governance, and investing in national statistical capacity . Daniel Power of Flowminder demonstrated the practical value of mobile phone data during the DRC Ebola outbreak, while cautioning that such data underrepresents women, children, and rural populations, and must be combined with other sources to correct for bias . Alexandre Barbosa of CETIC Brazil presented a self-sustaining financing model for ICT surveys funded through internet domain name registrations, highlighting its advantages of stability, independence, and capacity for innovation .
The discussion concluded with broad agreement that no single actor can deliver trusted data for AI, and that sustained partnerships across governments, international organisations, the private sector, and civil society are essential to advance this agenda at scale .
Overall Purpose
- The discussion aimed to explore how better, more trustworthy data can be produced and made accessible to support responsible AI development. Convened jointly by ITU, UNCTAD, and the UN CCSA, Colombia, Norway, and the UK, the session brought together statisticians, policymakers, donors, and private sector representatives to examine the roles of official statistics, novel data sources, financing mechanisms, and international partnerships in building AI-ready data ecosystems.
- --
Major Discussion Points
- The quality and trustworthiness of data are the central limiting factor for reliable AI. Panellists consistently emphasised that AI systems are only as dependable as the data underpinning them. Current data used in AI development is frequently fragmented, unevenly documented, and inaccessible for verification, leading to concerns about bias, transparency, and reproducibility. Official statistics, produced under internationally agreed standards and transparent methodologies, were identified as a uniquely reliable foundation, though panellists stressed that trusted data must extend beyond official statistics to include any responsibly governed source.
- The Trusted Data Observatory (TDO) as a practical mechanism for making trusted data findable and interoperable. Benjamin Rothen described how the arrival of large language models has fundamentally changed how users - including machines - search for data, making visibility and discoverability of trusted datasets a critical challenge. The TDO, led by Switzerland with approximately 20 countries and 30-40 international organisations involved, aims to create a global metadata discovery platform that does not centralise data ownership but enables machines and humans to locate trusted datasets. A key insight was that investing in metadata - descriptive information about datasets - is essential, as AI systems search for text rather than raw data. - Responsible use of novel data sources, particularly mobile phone data, to fill critical information gaps. Daniel Power illustrated how mobile operator data enabled Flowminder to rapidly identify high-risk health zones during an Ebola outbreak in the DRC, with eight out of ten subsequently confirmed zones matching their predictions. However, panellists cautioned that such data underrepresents women, children, rural populations, and the less wealthy, and must be combined with survey data to correct for bias before being used in AI systems. Privacy protections - such as processing data on the operator's premises and exporting only aggregated outputs - were highlighted as non-negotiable upstream safeguards. - Sustainable and diversified financing is essential for building long-term, AI-ready data systems. Vibeke Østreich-Nielsen noted that much existing investment in data, particularly donor funding for the Global South, is sectoral and disconnected from official statistics, rarely producing the continuity needed for robust data infrastructure. Four recommendations from the SEVIA Platform for Action were outlined: treating data as a public good, improving national coordination, strengthening data governance, and investing in national statistical capacity. Alexandre Barbosa presented Brazil's CETIC model - funded entirely through revenues from the .br internet domain registry - as an example of a self-sustaining, independent financing mechanism that has enabled 20 years of continuous, nationally representative ICT surveys. - Partnership across governments, the private sector, academia, and civil society is indispensable. Multiple speakers converged on the view that no single actor can deliver trustworthy, AI-ready data ecosystems alone. ITU's UN-wide initiative on mobile phone data, involving national statistical offices, regulators, and private operators across 25 countries, was cited as a model for multistakeholder collaboration. The question of whether to pay for private data - particularly from mobile operators - was raised as a live tension, with pragmatic, country-specific approaches recommended rather than a single universal policy.
- --
Overall Tone
- The tone throughout the discussion was constructive, collaborative, and professionally candid. Speakers were open about the scale of the challenges - data gaps, fragmented financing, bias in novel data sources, and the difficulty of defining 'trusted data' - without being pessimistic. There was a consistent undercurrent of cautious optimism, with panellists pointing to concrete initiatives (TDO, CETIC, Flowminder's DRC work, the SEVIA platform) as evidence that progress is achievable. The Q&A segment introduced a slightly more pragmatic and grounded register, with honest acknowledgements of tensions around data monetisation and the preconditions required before open data policies can be effective. The closing remarks reinforced a sense of shared urgency and collective responsibility, with the moderator calling for continued international coordination and emphasising that momentum exists to move the agenda forward.
Better Data for AI: A Possible Task? - Expanded Summary
#
Overview and Framing
The session, convened jointly by ITU, UNCTAD, and the Committee for the Coordination of Statistical Activities (CCSA) - a body bringing together 45 international and supranational organisations - was moderated by Anu Peltola, Director of UNCTAD Statistics, Data, and Digital Service . The discussion centred on a fundamental challenge: as AI increasingly shapes how information is produced, accessed, and used, the trustworthiness of AI systems is only as strong as the data underpinning them . Peltola framed the problem clearly at the outset, noting that much of the data currently used to develop AI systems is fragmented, uneven in quality, insufficiently documented, and often inaccessible for verification, giving rise to concerns about bias, transparency, and the reproducibility of results . Crucially, she emphasised that this is not solely a matter for official statistics - it concerns any data source used by AI tools - and that the challenge is not merely technical but requires well-thought-through reference frameworks for how AI selects and uses data .
The session was structured to address a sequence of interconnected questions: why better data for AI is needed, how trust can be built, how trusted data can be financed, how new data sources can be responsibly integrated, and how these efforts can be sustained institutionally . The discussion was situated within broader international processes, including the Global Digital Compact, data governance work, and the Financing for Development agenda, all of which emphasise the need for trusted, interoperable, and inclusive data to support sustainable development and digital transformation .
---
#
The Role of Official Statistics in Ensuring Data Quality for AI
Esperanza Magpantay, Senior Statistician at ITU with over three decades of experience in ICT statistics, opened the substantive discussion by reframing the AI debate . While much public discourse focuses on algorithms, models, and computing power, she argued that the real limiting factor is increasingly the quality of the data from which these models learn . For AI to serve the public good, data must be not only abundant but also trusted, representative, well-documented, interoperable, ethically sourced, and governed under clear principles . Without these characteristics, AI systems risk amplifying biases, producing unreliable results, and reinforcing inequalities .
Magpantay identified official statistics as uniquely positioned to address this challenge. For decades, national statistical offices and international organisations have developed rigorous standards for producing high-quality data, using transparent methodologies, internationally agreed definitions, quality assurance frameworks, and strong confidentiality protections . These represent one of the few global sources of trusted evidence that governments, businesses, and citizens can rely upon . At ITU, this is visible in the daily work of measuring digital development - whether monitoring connectivity, tracking SDG progress, or measuring digital inclusion - where AI systems will only be useful if grounded in high-quality official statistics .
Critically, Magpantay rejected a binary framing of the relationship between official statistics and new data sources. The future, she argued, is not about replacing official statistics with AI, nor about replacing surveys with novel data sources, but about combining trusted official statistics with responsibly governed new sources to create richer, more timely, and more relevant data for decision-making . This framing - that combination rather than replacement is the path forward - was echoed by subsequent speakers throughout the session. She also stressed that partnership is becoming essential, as no progress can be made without governments, national statistical offices, regulators, international organisations, academia, the private sector, and civil society working together to build a trusted data ecosystem .
---
#
Building Trusted and AI-Ready Data Ecosystems: The Trusted Data Observatory
Benjamin Rothen, Head of International and National Affairs at the Swiss Federal Statistical Office, introduced the Trusted Data Observatory (TDO) initiative and offered a distinctive diagnosis of the current data problem . He observed that the arrival of large language models - referencing the emergence of ChatGPT in late 2022 as a turning point - has fundamentally changed how users, including machines, search for and interact with data . He made the striking observation that statistical offices must now recognise that "our customers are not humans anymore - often they're machines," underscoring the scale of this shift for official data producers. The challenge is no longer primarily one of data scarcity, but of data invisibility . Trusted data is often not findable, not interoperable, and not harmonised - problems that official statisticians have long worked to address, but which have taken on new urgency in the AI era .
Rothen described the TDO as a response to this challenge. The initiative, led by the Swiss government and involving approximately 20 countries and 30 to 40 international organisations from across the globe, aims to create a Geneva-based global metadata discovery platform . The core principle is that the TDO is not designed to centralise data or transfer ownership - data remains where it is - but to provide a shared discovery layer so that both machines and people can locate trusted datasets . He drew on Switzerland's own decade-long experience building a national metadata platform (known as I14Y, focused on interoperability) as evidence of both the value and the difficulty of this work, acknowledging that even after ten years, the task of making data findable and interoperable remains challenging .
A particularly important technical insight Rothen offered was that investing in metadata - descriptive information about datasets - is essential because large language models do not search for raw data; they search for text . Without rich, descriptive metadata, AI systems will be unable to locate trusted datasets regardless of their quality. He also noted that Switzerland is redesigning its statistical office website by October to be AI-ready, enabling machines to find and use data more effectively . On the question of what constitutes trusted data, Rothen acknowledged that official statistics - governed by frameworks such as the UN Fundamental Principles and the FAIR principles - provide a natural starting point, but argued that on a longer-term basis the TDO should encompass all trusted data, not only statistical data, since in some cases NGO data may be more reliable than national statistical office data . He closed with an open invitation for countries, organisations, and funders to join the TDO initiative, with the ambition of presenting a working prototype at the AI Summit in Geneva in approximately 11 to 12 months .
---
#
Financing Sustainable Data Systems: Structural Challenges and Innovative Solutions
Vibeke Østreich-Nielsen, Senior Adviser at NORAD in Norway and co-lead of the Financing for Development Future of Data Initiative, joined the session remotely and offered a structural critique of how data investment is currently organised globally . Drawing on two decades of experience in strengthening statistical capacity, she observed that because statistics and data are a cross-cutting field, investment tends to be sectoral and project-based rather than systemic . Donor funding for the Global South, in particular, is often not linked to official statistics or to the goal of making data available for broader use, including for AI purposes . This fragmentation means that data systems rarely receive the continuity of investment needed to build long-term statistical capacity .
In response to this structural problem, Østreich-Nielsen described the SEVIA Platform for Action, developed in the context of the Financing for Development agenda and the fourth international conference on Financing for Development . The initiative brought together countries and partners to discuss what needs to be done, producing four recommendations: treating data and statistics as a public good; improving national data and statistical system coordination; strengthening data governance and innovation; and investing in national data and statistics capacity . She emphasised that the first recommendation - treating data as a public good - directly implies that investments in data and statistics should have an end goal of publishing data and making it available in AI-ready formats, which she identified as a major ongoing challenge . She also highlighted that many ministries hold relevant administrative data that is not being made available, representing a significant untapped resource for decision-making and AI, particularly in African contexts where very little data is publicly available . From her position in a donor agency, she noted that practical work is under way to develop guidance for donor representatives on how to evaluate data-related projects and ensure that funded data work is made available in AI-ready formats .
Alexandre Barbosa, Head of the Regional Centre for Studies on the Development of the Information Society (CETIC) in Brazil, offered a concrete and innovative response to the financing challenge . He noted that ICT surveys are frequently financed through mechanisms not based on solid budget commitments - temporary government programmes or donor-funded projects - which rarely provide the continuity needed to build long-term statistical capacity, leading to interrupted time series, lost expertise, and an inability to monitor digital transformation . CETIC's solution, developed over 20 years, is a self-sustainable financing mechanism funded entirely by revenues from the .br country code top-level domain registry, managed through the Brazilian Internet Steering Committee - a multi-stakeholder body comprising government, civil society, academia, and the private sector - and NIC.br . This model has three principal advantages: stability, enabling continuous survey programmes and long-term statistical time series; independence, as CETIC does not receive government or donor funding and can maintain methodological consistency and long-term planning; and innovation, as stable financing allows continuous updating of surveys to address emerging technologies while maintaining international comparability . Barbosa acknowledged, however, that this model is not easily replicable, as Brazil's position as one of the largest domain name databases among G20 and OECD countries provides financial resources that most countries would not have .
---
#
Responsible Use of Novel Data Sources: Mobile Phone Data in Humanitarian Action
Daniel Power, Managing Director of Flowminder Foundation, illustrated the practical value of non-traditional data sources through a detailed case study from the Democratic Republic of Congo . In May, during an Ebola outbreak in the northeast of the country - an already data-scarce region - the immediate question was where the disease would spread next . Drawing on a longstanding eight-year partnership with Vodacom Congo, Flowminder rapidly conducted a cohort study, identifying all subscribers who had been in the outbreak area and tracking which health zones they visited in the days and weeks following the outbreak . All 500-plus health zones in the DRC were ranked according to the intensity of connectivity with the outbreak areas . A week after the first report was released, Ebola was detected in ten more health zones, and eight of those ten were among the top regions Flowminder had identified - demonstrating the timeliness and richness of mobile phone data for humanitarian decision-making .
Power was careful, however, to frame this success within a broader set of upstream and downstream considerations for responsible data use . On the upstream side, subscriber data is inherently sensitive, and Flowminder processes all data on the mobile operator's premises, deploying software that aggregates data before it is exported so that individual-level data never leaves the operator's systems . He emphasised that protecting the privacy of individuals contributing data to these systems is essential, particularly given AI's growing demand for data . On the downstream side, he was candid about the limitations of mobile phone data: it is not fully representative of the population, systematically underrepresenting women, children, the very old, the more rural, and the less wealthy . He argued that these biases must be corrected through complementary survey data before mobile phone data is used in AI systems, and that this applies just as strongly - if not more so - when data feeds into automated analytical systems that may lack the discretion to account for such biases .
Esperanza Magpantay complemented Power's account by describing the ITU and World Bank's initiative to integrate mobile phone data into official statistics sustainably across 25 countries . She described a model in which national statistical offices do not pay mobile operators directly for their data, but instead offer non-monetary incentives - such as sharing technical expertise in data quality assurance and providing access to statistical data that operators cannot otherwise obtain - in exchange for access to mobile phone data . This approach aims to ensure that mobile data is used responsibly and integrated as an official data source rather than treated as a commercial commodity . She confirmed that Côte d'Ivoire is among the 25 participating countries, responding to a question raised in French by a representative from Côte d'Ivoire about how the CETIC model was established and whether it could be replicated .
---
#
The Question of Paying for Data: Tensions and Trade-offs
An audience question from Jacques Péguet raised the issue of whether downstream users should pay for better data . This prompted a revealing exchange that exposed genuine tensions between different institutional perspectives. Rothen took a firm public-good position, stating that in Switzerland data cannot be charged for under law, as it is funded by taxes, though services can be charged for . Power, by contrast, acknowledged a real and live tension: mobile operators are frequently told that their data represents an untapped source of revenue to be monetised, which can make negotiations difficult . He noted pragmatically that Flowminder would pay for data if it helped open a door and respond to a crisis, provided a donor was prepared to support this, but offered an important caveat: the amount paid does not correlate with data quality, and he had seen cases where organisations paid for data that turned out to be of poor quality . He also cited Ghana Statistical Services as an example of a constructive model, noting that in Ghana - where Flowminder has a long-term relationship with the national statistical office - the instruction is not to charge for the data, reflecting a public-good orientation. Magpantay offered a middle path, describing the non-monetary incentive model as preferable to direct payment, with mutual benefit - rather than commercial transaction - as the organising principle . These positions reflect a genuine and unresolved tension between public good principles, commercial incentives, and operational pragmatism that will require further policy and economic analysis to navigate.
---
#
Foundational Preconditions and Practical Challenges
A question from a representative of GIZ Egypt introduced an important grounding perspective, noting that Egypt has recently released an open data policy but that practical implementation requires addressing significant preconditions: digitising existing data, enabling government-to-government data sharing, and incentivising both public and private actors . This observation resonated with points made by multiple panellists. Rothen had already acknowledged that even Switzerland, after a decade of work on its national metadata platform, is still struggling to make data findable and interoperable . Østreich-Nielsen had highlighted that many ministries sit on relevant administrative data that is not being made available . Together, these contributions revealed an implicit tension between the ambition of global initiatives such as the TDO and the foundational capacity gaps that many countries - particularly in the Global South - still face.
Responding directly to the GIZ Egypt question, Peltola offered a constructive reframing: if the TDO held metadata on what data are available globally, it would also reveal data gaps, which could help target investment to where gaps are most critical for policy needs . This reframing of metadata platforms as instruments for identifying and addressing data gaps - not merely for making existing data discoverable - added a strategic dimension to the technical discussion. She also noted that governments committed in the Financing for Development outcome document to investing in their national statistical systems, providing a political foundation upon which to build, though practical implementation varies considerably by country .
---
#
Convergences, Tensions, and Unresolved Questions
Across the session, a high degree of consensus emerged on several foundational principles: that data quality is the central prerequisite for trustworthy AI ; that the future lies in combining official statistics with novel data sources rather than replacing one with the other ; that metadata and data discoverability are critical infrastructure for AI-readiness ; that multi-stakeholder collaboration is indispensable ; and that current financing for data systems is structurally inadequate, particularly in the Global South . A notable area of consensus also emerged around data decentralisation: despite the ambition of the TDO as a global platform, all speakers accepted without challenge that data should remain with its original owners, with only metadata shared globally .
Beneath this surface consensus, however, meaningful tensions remained. The definition of 'trusted data' was acknowledged as genuinely difficult to resolve, with Rothen explicitly stating that he could not provide a complete answer and inviting collaborative work to address it . The question of paying for privately held data exposed divergent institutional positions that reflect deeper structural differences between public statistical institutions and operational organisations working in humanitarian contexts . The replicability of Brazil's financing model was acknowledged as limited by context , and the gap between the TDO's ambitions and the foundational challenges facing many countries was not fully bridged. How AI tools select and use trusted data - rather than fragmented or low-quality sources - when generating outputs or analysis was raised as a concern but not addressed with a concrete solution .
---
#
Conclusions and Next Steps
The session closed with the moderator drawing together the key threads of the discussion . No single actor can deliver trusted data for AI; partnerships across governments, international organisations, the private sector, and civil society are essential. AI tools are powerful and should not be underestimated, but they require good data to produce robust results, and it is critically important to scrutinise the numbers and analysis that AI tools generate. The international statistical community, the geospatial community, and private data ecosystems are actively working on these challenges, and momentum exists to move the agenda forward. Peltola committed to bringing the discussion back to the UN system and the broader network of chief statisticians to consider how to advance the trusted data for AI agenda at scale and connect with partners internationally .
Concrete near-term actions identified during the session include: the TDO's ambition to present a working prototype at the AI Summit in Geneva in approximately 11 to 12 months ; Switzerland's redesign of its statistical office website to be AI-ready by October ; the ITU and World Bank's ongoing work to integrate mobile phone data into official statistics across 25 countries ; and NORAD's development of practical guidance for donor representatives on evaluating data-related projects . The session made clear that while the principles are broadly shared, the hard work of translating them into concrete, country-specific operational frameworks - particularly for data-scarce contexts in the Global South - remains very much ahead.
Official statistics provide trusted, representative, and well-documented data that AI systems require to avoid bias and unreliable results - Official statistics as a foundation for trustworthy AI
Arg. 1Esperanza Magpantay argues that the real limiting factor for AI is not algorithms or computing power, but the quality of the data these models learn from. Official statistics, produced with transparent methodologies and internationally agreed definitions, represent one of the few truly trusted global sources of evidence. Without such characteristics, AI systems risk amplifying biases and reinforcing inequalities.
She noted that for decades, national statistical offices and international organisations have developed rigorous standards, including transparent methodologies, internationally agreed definitions, quality assurance frameworks, and strong confidentiality protections . She further illustrated this with ITU's work on measuring digital development, connectivity, and digital inclusion, arguing that AI systems will only be useful if grounded in high-quality official statistics .
on: Data quality is the fundamental prerequisite for trustworthy and beneficial AI systems
on: The definition and scope of 'trusted data' — whether it should be limited to official statistics or extended to other sources
The future lies in combining trusted official statistics with responsibly governed new data sources, not replacing one with the other - Combining traditional and new data sources
Arg. 2Magpantay emphasises that the goal is not to replace official statistics with AI or to substitute surveys with new data sources, but rather to combine them. By integrating trusted official statistics with responsibly governed novel data sources, the result is richer, more timely, and more relevant data for decision-making. This combination approach is presented as the key lesson from existing experience.
She referenced the UN-wide initiative on mobile phone data as an example of integrating new data sources with traditional statistical systems, noting these sources provide more timely and granular data while following governance and transparency measures . She explicitly stated that the future is about combining trusted official statistics with responsibly governed new data sources .
on: The future lies in combining trusted official statistics with responsibly governed new data sources, not replacing one with the other
No single actor can deliver trusted data for AI; governments, national statistical offices, regulators, international organisations, academia, the private sector, and civil society must all collaborate - Multi-stakeholder collaboration as essential
Arg. 3Magpantay argues that partnership is becoming essential because no single entity can build a trusted data ecosystem alone. Collaboration across governments, national statistical offices, regulators, international organisations, academia, the private sector, and civil society is required. This multi-stakeholder approach is presented as a prerequisite for AI that serves the public good.
She stated that partnership is becoming very essential because work cannot proceed without governments, national statistical offices, regulators, international organisations, academia, the private sector, and civil society . She also described ITU's work with experts in the UN Initiative on Mobile Phone Data as an example of bringing partners together to develop methodology and help countries use new data sources .
on: Multi-stakeholder collaboration across governments, statistical offices, regulators, international organisations, academia, private sector, and civil society is essential for building trusted data ecosystems for AI
Integrating mobile phone data into official statistics requires bringing all stakeholders together, with national statistical offices offering technical skills and data access as incentives rather than direct payment - Multi-stakeholder model for mobile data integration
Arg. 4Magpantay describes a model where national statistical offices do not pay mobile operators directly for their data, but instead offer incentives such as technical skills in data quality assurance and access to statistical data that operators do not otherwise have. This approach is being implemented across 25 countries through a World Bank project. The aim is to integrate mobile phone data sustainably into official statistics.
She described a project implemented in 25 countries with the World Bank, where the objective is to integrate mobile phone data into official statistics sustainably by having all stakeholders work together . She explained that national statistical offices can provide incentives such as technical skills and access to their own data, rather than paying operators directly for data . She also confirmed that Côte d'Ivoire is one of the 25 countries benefiting from this technical assistance .
on: Privacy protection and responsible data governance must accompany the use of novel data sources such as mobile phone data
on: Whether to pay for privately held data such as mobile operator data
AI tools are only as trustworthy as the data they use, and official statistics produced under internationally agreed standards offer authoritative evidence for AI systems - Data trustworthiness as a prerequisite for AI
Arg. 1Anu Peltola frames the central challenge of the session: AI is only as trustworthy as the data it uses and shares. Trustworthiness and well-documented, accessible data are not merely technical issues but essential foundations for reliable, accountable, and beneficial AI tools. Official statistics, produced according to internationally agreed standards with continuous public oversight, are positioned as a key source of authoritative evidence.
She noted that much of the data used to develop AI systems is fragmented, uneven in quality, insufficiently documented, and often inaccessible for verification, leading to concerns about bias, transparency, and reproducibility . She highlighted that official statistics are produced according to internationally agreed standards and methodologies with continuous public oversight .
on: Data quality is the fundamental prerequisite for trustworthy and beneficial AI systems
on: The definition and scope of 'trusted data' — whether it should be limited to official statistics or extended to other sources
Governments committed in the Financing for Development outcome document to investing in national statistical systems, providing a political foundation to build upon - Political commitment to data investment
Arg. 2Peltola points out that governments have already made a formal commitment to investing in their national statistical systems and data through the Financing for Development outcome document. This political commitment provides a foundation upon which further action can be built, though its practical implementation varies by country. The international community is working to support the realisation of this commitment.
She referenced the outcome document of the Financing for Development conference, noting that governments committed to investing in their national statistical systems and data, and acknowledged that the extent to which this is put into practice depends on the country .
Understanding what data exists through metadata platforms can help identify data gaps and target investment where it is most critically needed for policy - Using metadata to identify and address data gaps
Arg. 3Peltola argues that a metadata platform such as the Trusted Data Observatory would not only make data findable but would also reveal where data gaps exist. This visibility into gaps could help direct investment and action towards areas where information is most critically needed for policy. The connection between metadata, gap identification, and targeted investment is presented as a practical benefit of such platforms.
She observed that if the TDO held metadata on what data are available, it would also reveal data gaps, which could help target action to where gaps are most critical for policy needs, thereby helping to connect investment with identified gaps .
The Trusted Data Observatory (TDO) aims to create a global metadata discovery platform so that machines and people can find trusted data, without centralising or transferring ownership of the data - TDO as a global metadata platform
Arg. 1Benjamin Rothen explains that the TDO is a global discovery platform designed to make trusted data findable by both machines and humans, without centralising the data or transferring ownership. National metadata platforms are linked to a global metadata platform, allowing users to discover where trusted data resides. The data itself remains with its original custodians.
He described the TDO as a national metadata platform linked to a global metadata platform, emphasising that data stays where it is and ownership is not transferred, but a global discovery platform is created so machines and people know where to find trusted data . He noted that Switzerland has had a metadata platform for ten years where all government data comes together so machines and humans can find the right data .
on: Metadata platforms and data discoverability are foundational prerequisites for making data findable and usable by AI systems
on: The definition and scope of 'trusted data' — whether it should be limited to official statistics or extended to other sources
Trusted data must be findable, interoperable, and harmonised; investing in metadata is essential because machines search for text descriptions, not raw figures - Metadata as the key to data discoverability
Arg. 2Rothen argues that a critical problem is not a lack of data but a lack of visibility: trusted data is often not findable, not interoperable, and not harmonised. Investing in metadata — descriptions of what data exists and what it means — is essential because large language models and machines search for text, not raw numbers. Without good metadata, machines cannot locate or use the underlying data.
He stated that the problem is often not a lack of data but that trusted data is not visible, not findable, and not interoperable . He emphasised that if organisations do not invest in metadata, data sets will never be found, because LLMs search for text data and not raw figures, making good textual descriptions essential .
on: Metadata platforms and data discoverability are foundational prerequisites for making data findable and usable by AI systems
National statistical offices must adapt their digital presence to be AI-ready, enabling machines to find and use data more effectively - Making statistical offices AI-ready
Arg. 3Rothen highlights that the work of national statistical offices has changed fundamentally in the past two to three years, requiring organisations to restructure how they operate and present their data. Switzerland is redesigning its entire website to be AI-ready so that machines can find data more effectively. This is described as a significant investment and a major organisational challenge.
He noted that Switzerland is changing its webpage completely in October so it is AI-ready, enabling machines to find data much better, describing this as a huge task and a significant investment . He also acknowledged that Switzerland has been working for ten years on a metadata platform where all government data comes together for both machines and humans, and that this remains a difficult task .
Open government data and FAIR principles are important steps, but metadata platforms that describe what data exists are needed even before data is made openly accessible - Open data and FAIR principles as building blocks
Arg. 4Rothen acknowledges the importance of open government data and FAIR principles as foundational steps, but argues that metadata platforms are needed at an even earlier stage to describe what data exists. Even data that is not openly accessible should be described in metadata so that governments and researchers know it exists and can begin conversations about access. This approach helps governments understand their own data landscape.
He described five steps to reaching a level of open government data publication and noted that the TDO tries to start much earlier by identifying what data exists and how different datasets can be linked . He explained that even micro-level data that is not openly accessible should have a description in the metadata platform, so governments can discover what data other ministries hold and begin conversations about access .
on: Whether to pay for privately held data such as mobile operator data
The TDO initiative involves approximately 20 countries and 30 to 40 international organisations working together to build a Geneva-based global trusted data platform - TDO as a collaborative international initiative
Arg. 5Rothen describes the TDO as a collaborative initiative led by the Swiss government, involving around 20 countries from across the world and 30 to 40 international organisations. The platform is intended to be Geneva-based, drawing on the existing concentration of international knowledge in Geneva and expanding globally. He invites further participation and investment to help build the platform.
He stated that approximately 20 countries and 30 to 40 international organisations are working with Switzerland on the TDO, with the platform intended to be Geneva-based to bring together existing knowledge and expand globally . He expressed the ambition to show a prototype at the AI summit in Geneva in 11 to 12 months and invited all present to be part of the initiative .
on: Multi-stakeholder collaboration across governments, statistical offices, regulators, international organisations, academia, private sector, and civil society is essential for building trusted data ecosystems for AI
Investment in data and statistics is often sectoral and donor-driven, lacking the continuity needed to build long-term statistical capacity, particularly in the Global South - Fragmented and unsustainable data financing
Arg. 1Vibeke Østreich-Nielsen identifies a structural problem in how data and statistics are funded: investment tends to be sectoral and project-based rather than systemic, and donor funding for the Global South is often not linked to official statistics or long-term capacity building. This fragmentation means that data systems lack the continuity needed to develop sustained statistical capacity. The focus is often on meeting specific data needs rather than making data broadly available.
She observed that because statistics and data are a cross-cutting field, investment is often sectoral and not always linked to official statistics, with the focus being on specific data needs rather than making data available for countries generally or for AI purposes . She noted that this is part of the broader challenge facing data and statistics ecosystems .
on: Sustainable and long-term financing for data and statistical systems is critical but currently fragmented and insufficient, particularly in the Global South
on: The replicability of Brazil's self-sustainable financing model for ICT statistics
The SEVIA Platform for Action recommends treating data and statistics as a public good, improving coordination, strengthening data governance, and investing in national statistical capacity - Four recommendations for sustainable data financing
Arg. 2Østreich-Nielsen describes the SEVIA Platform for Action, developed in the context of the Financing for Development agenda, which produced four recommendations to operationalise commitments to data and statistics. These recommendations cover treating data as a public good, improving national coordination, strengthening data governance and innovation, and investing in national statistical capacity. The first recommendation explicitly includes making data available in AI-ready formats.
She outlined the four recommendations: treating data and statistics as a public good, improving national data and statistical system coordination, strengthening data governance and innovation, and investing in national data and statistics capacity . She noted that the first recommendation includes an end goal of publishing data and making it available in AI-ready formats, which remains a major challenge .
Donor organisations need to develop practical approaches for evaluating data-related projects and ensuring that funded data work is made available in AI-ready formats - Donor community's role in data collaboration
Arg. 3Østreich-Nielsen, speaking from the perspective of a donor agency, argues that donor organisations need to develop more practical approaches for how they evaluate and decide on data-related projects. This includes ensuring that data work funded through donor projects is ultimately made available in AI-ready formats. She presents this as a concrete step the donor community can take to improve the data ecosystem.
She described work from a donor agency context to develop practical approaches for representatives in donor organisations on how they review and decide on projects, particularly regarding data work . She noted that this is intended to help ensure funded data work contributes to making data available in AI-ready formats .
on: Multi-stakeholder collaboration across governments, statistical offices, regulators, international organisations, academia, private sector, and civil society is essential for building trusted data ecosystems for AI
Many ministries hold relevant administrative data that are not publicly available, representing a major untapped resource for decision-making and AI, particularly in African contexts - Unlocking administrative data as a priority
Arg. 4Østreich-Nielsen highlights that a significant untapped resource exists in the form of administrative data held by government ministries that is not being made publicly available. This data could provide valuable information for both decision-makers and AI systems, particularly in African contexts where very little data is publicly accessible. Making this data available is identified as a major challenge and priority.
She noted that many ministries sit on relevant administrative data that are not being made available, but that could bring a lot of information to decision-makers and in an AI context, particularly in an African context where very little data is publicly available .
Mobile phone data enabled rapid identification of health zones at risk during the DRC Ebola outbreak, demonstrating the timeliness and richness of such data for humanitarian crisis response - Mobile phone data for humanitarian crisis response
Arg. 1Daniel Power uses the example of the May 2023 Ebola outbreak in eastern DRC to demonstrate the practical value of mobile phone data in humanitarian crises. By tracking the movements of subscribers who had been in the outbreak area, Flowminder was able to rank health zones by their connectivity to the outbreak and predict where Ebola was likely to spread. This analysis proved highly accurate within a week of publication.
He described how Flowminder, using data from Vodacom Congo through an eight-year partnership, conducted a rapid cohort study identifying all subscribers who had been in the outbreak area and tracking which health zones they visited in subsequent days and weeks . He noted that a week after releasing their first report, Ebola was detected in ten more health zones, eight of which were among the top regions Flowminder had identified using mobility data .
Mobile phone data underrepresents women, children, the elderly, the rural, and the less wealthy, so it must be combined with survey data to adjust for bias before being used in AI systems - Addressing bias in mobile phone data
Arg. 2Power acknowledges that mobile phone data, while valuable, is not representative of the whole population and systematically underrepresents certain groups. This bias must be addressed by combining mobile data with survey data to adjust for these imbalances before the data is used in AI systems. He argues this is critical both for general use and even more so when data feeds into AI systems that may lack the discretion to account for such biases.
He stated that mobile phone data underrepresents women, children, the very old, the more rural, and the less wealthy, and that many data types privilege men and the wealthy . He noted that when releasing monthly data on population distribution, it is necessary and critical to run surveys to understand how representative the data is and to adjust for bias, and that this applies equally or more strongly when data is used for AI .
on: The future lies in combining trusted official statistics with responsibly governed new data sources, not replacing one with the other
Privacy must be protected by processing data at the mobile operator's premises and aggregating it before export, ensuring individual-level data does not leave the source - Privacy protection in mobile data processing
Arg. 3Power argues that protecting the privacy of individuals who contribute data to mobile systems is essential, particularly given how much information subscriber data contains about individual movements. Flowminder's approach is to deploy software on the mobile operator's premises that aggregates data before it is exported, meaning individual-level data never leaves the operator's systems. This is presented as a model for responsible data use in AI contexts.
He explained that Flowminder always processes data at the mobile operator's systems, so individual-level data did not leave the Congo and did not leave Vodacom's premises; instead, software deployed on their premises aggregates the data before export . He emphasised the importance of protecting the privacy of people contributing data to AI systems .
on: Privacy protection and responsible data governance must accompany the use of novel data sources such as mobile phone data
Data governance frameworks, including regulator support and clear principles, are essential for responsible use of novel data sources such as mobile phone data - Data governance for novel data sources
Arg. 4Power highlights that having the support of regulators is critical for legitimising and enabling the responsible use of mobile phone data, particularly in complex country contexts. Long-term relationships with both mobile network operators and regulators provide the foundation for accessing and using such data responsibly. Clear governance principles and frameworks are presented as prerequisites for this work.
He noted that having a signal from the regulator ARP in DRC that they were comfortable with Flowminder's work was really important for navigating the complex environment of that country . He also described the importance of understanding operator needs and building long-term relationships, as well as the role of the World Bank's global data facility in Côte d'Ivoire .
on: Privacy protection and responsible data governance must accompany the use of novel data sources such as mobile phone data
The question of whether to pay for privately held data, such as mobile operator data, involves real tensions between public good principles, operator monetisation expectations, and data quality outcomes - Paying for private data: tensions and trade-offs
Arg. 5Power describes the genuine tension that arises when negotiating access to mobile operator data, as operators are often told their data is a source of revenue to be monetised, while governments may have hard lines against paying for data. He notes that paying for data does not necessarily correlate with better data quality, and that his organisation takes a pragmatic approach depending on the country and operator. The considerations are complex and vary significantly by context.
He described the tension between governments' reluctance to pay for data and operators' expectations of monetisation, noting that in the Congo, Flowminder pays only a small fee for server management rather than for the data itself . He observed that paying for data does not necessarily correlate with data quality, citing cases where organisations paid for data that turned out to be of poor quality .
Long-term relationships with mobile network operators and regulators are critical for accessing and responsibly using mobile data in complex country contexts - Long-term operator and regulator partnerships
Arg. 6Power emphasises that Flowminder's ability to respond rapidly to crises such as the DRC Ebola outbreak was made possible by an eight-year partnership with Vodacom Congo. Similarly, having regulator support is critical for navigating complex country environments. These long-term relationships are presented as essential infrastructure for responsible mobile data use.
He noted that Flowminder's rapid response to the DRC Ebola outbreak was enabled by a longstanding eight-year partnership with Vodacom Congo . He also described how support from the regulator ARP in DRC was really important for legitimising their work in a complicated country context involving conflict and other considerations .
ICT statistics produced through rigorous, internationally aligned methodologies support the design and evaluation of public policies on digital inclusion and SDG monitoring - ICT data for national policy
Arg. 1Alexandre Barbosa argues that nationally representative ICT statistics are essential for designing, implementing, and evaluating public policies across multiple domains including digital inclusion, education, and health. Brazil's CETIC model produces data aligned with international standards and SDG frameworks. The availability of this data in multiple languages further enhances its accessibility and utility.
He stated that nationally representative ICT statistics from CETIC support the design, implementation, and evaluation of public policies on digital inclusion, education, health, and several other dimensions, and are aligned with SDG dimensions . He noted that data are available in Portuguese, English, and some in Spanish, accessible via their website and reports .
Brazil's CETIC model, funded by revenues from the .br country code domain registry, provides a self-sustainable, independent, and innovative financing mechanism for continuous ICT surveys - Self-sustainable financing through domain registry revenues
Arg. 2Barbosa presents Brazil's CETIC as an example of a self-sustainable financing model for ICT statistics, funded entirely by revenues from the .br country code domain registry rather than government budgets or donor funding. This model has enabled 20 years of continuous ICT surveys. The Brazilian Internet Steering Committee, a multi-stakeholder body, oversees this arrangement.
He explained that CETIC has been running continuous ICT surveys for 20 years through a self-sustainable financing mechanism funded by the .br country code domain registry, managed by the Brazilian Internet Steering Committee and NIC.br, with surveys 100% funded by registry activities . He noted that Brazil ranks sixth largest domain name database among G20 and OECD countries, providing sufficient financial resources to fund the surveys .
Stable financing enables long-term statistical time series, methodological independence, and continuous innovation in data collection, including the use of alternative and big data sources - Benefits of stable financing for statistical innovation
Arg. 3Barbosa identifies three major advantages of the CETIC financing model: stability enabling continuous survey programmes and long-term time series; independence from government budgets and individual donors, ensuring methodological consistency; and innovation, allowing continuous updating of surveys to address emerging technologies. Stable financing also enables the development of innovative data collection methods such as big data and mobile phone data.
He outlined three advantages of the model: stability enabling continuous survey programmes and long-term statistical series for monitoring digital transformation ; independence from public budgets and individual donors, maintaining professional independence and methodological consistency ; and innovation, allowing continuous updating of surveys and development of alternative data sources including big data and mobile phone data .
on: The future lies in combining trusted official statistics with responsibly governed new data sources, not replacing one with the other
Many countries face foundational challenges such as digitising existing data, enabling government-to-government data sharing, and incentivising both public and private actors before AI-ready data can be produced - Foundational preconditions for AI-ready data
Arg. 1The GIZ Egypt Representative raises the practical challenge that many countries, including Egypt, face foundational preconditions that must be met before AI-ready data can be produced. These include digitising existing data, enabling government-to-government data sharing, and creating incentives for both public and private sector actors. Research suggests these preconditions must be in place before more advanced data governance and AI data initiatives can succeed.
She described Egypt's recent open data policy and GIZ's efforts to support its implementation, noting that research she encountered identified many preconditions that need to be in place first, including digitising existing data, enabling government-to-government communication, creating connected platforms, and incentivising government and private sector actors .
The question of whether to pay for better data or treat it as a free public good raises important political and practical considerations - Paying for private data: tensions and trade-offs
Arg. 1Jacques Péguet raises the question of downstream payment for data, asking whether there is any consideration of charging for better data or whether this is ruled out for political reasons. This question touches on the tension between treating data as a public good and the practical need to fund its production and maintenance. The question invites reflection on sustainable financing models for high-quality data.
on: Whether to pay for privately held data such as mobile operator data
Interest from countries such as Côte d'Ivoire in replicating mobile data integration models highlights the need for practical technical assistance and knowledge sharing across countries - Practical technical assistance for data integration
Arg. 1The Côte d'Ivoire Representative expresses interest in learning more about the CETIC model and how it was established, reflecting a broader demand from countries in the Global South for practical guidance on replicating successful data integration and financing models. This highlights the need for knowledge sharing and technical assistance across countries seeking to build similar capacities.
The representative from Côte d'Ivoire expressed interest in learning more about the CETIC model and how it was set up , prompting Esperanza Magpantay to confirm that Côte d'Ivoire is one of the 25 countries already benefiting from the World Bank mobile phone data integration project and to invite further discussion .
Statistics production must incorporate innovation because what counts as information in society has fundamentally changed, and the same definitions and procedures are no longer sufficient - Innovation as a necessity in statistics
Arg. 1The Moderator argues that it is not enough to produce statistics and numbers using the same definitions and procedures as before, because the nature of information in society has changed over time. Innovation must be built into the entire process of data production, not treated as an optional add-on. This point is made in response to Alexandre Barbosa's discussion of financing and innovation.
The Moderator noted that what we understand as information in society has changed over the years, and that statistics cannot simply be produced using the same definitions and the same procedures, emphasising that innovation must be incorporated into all aspects of data production .
Concrete and measurable progress milestones should be identified for advancing trusted data ecosystems, such as demonstrating a prototype at the next AI summit - Setting concrete targets for progress
Arg. 2The Moderator raises the question of what concrete progress would look like at the next summit or conference, pushing the discussion beyond general principles towards specific, demonstrable outcomes. This framing encourages panellists to identify tangible deliverables rather than aspirational goals. The TDO prototype at the Geneva AI summit is cited as one such concrete milestone.
The Moderator asked what a single concrete piece of progress in this area would look like at the next summit or conference, prompting Benjamin Rothen to identify the TDO prototype at the AI summit in Geneva in 11 to 12 months as a key deliverable .
No single actor can deliver trusted data for AI; partnerships across different types of actors and data sources are essential, and AI outputs must be critically evaluated - Multi-stakeholder collaboration and critical use of AI
Arg. 3The Moderator synthesises the session's key messages by emphasising that no single actor can deliver trusted data for AI and that partnerships are indispensable. She also stresses that while AI tools are powerful and should not be underestimated, the results they produce must be critically evaluated, because data shapes debates and decisions and it matters what numbers underlie any analysis.
The Moderator concluded the session by stating that no single actor can deliver this, that partnerships and different types of data are needed, and that AI tools are powerful but require good data to ensure robust results, adding that it is easy to produce outcomes with AI tools but that one must be very critical about them .
on: Multi-stakeholder collaboration across governments, statistical offices, regulators, international organisations, academia, private sector, and civil society is essential for building trusted data ecosystems for AI
The outcomes of the session should be taken back to the UN system and the network of chief statisticians to consider how to advance the data-for-AI agenda at scale and connect with partners internationally - Scaling up through the UN statistical system
Arg. 4The Moderator commits to carrying the discussion forward within the UN system and the broader network of chief statisticians, framing this as the appropriate institutional channel for scaling up and connecting with partners. This reflects a recognition that the issues raised require coordinated international action beyond the session itself.
The Moderator stated that she would take the discussion back to the UN system and the broader network of chief statisticians so that the international community can consider how best to advance this agenda to scale and connect with partners .
Publicly held data can also be privately held and equally important for societies and decision-makers, including for anticipating future needs - Private data as public good
Arg. 5The Moderator draws attention to the fact that data serving the public good is not exclusively held by public institutions; privately held data, such as mobile operator data, can be equally important for societies and decision-makers. She also highlights the value of such data for anticipating future needs, as demonstrated by the Ebola outbreak example.
Following Daniel Power's presentation on the DRC Ebola outbreak, the Moderator observed that public good data can also be privately held and can be equally important for societies and decision-makers, and can help foresee what will be needed .
Session Knowledge Graph
Speakers · Topics · Arguments · Relationships
All speakers converged on the view that AI systems are only as trustworthy as the data they learn from . Peltola noted that much of the data used to develop AI systems is fragmented, uneven in quality, insufficiently documented, and often inaccessible for verification, leading to concerns about bias, transparency, and reproducibility . Magpantay reinforced this, stating that without trusted, representative, well-documented, interoperable, and ethically sourced data, AI systems risk amplifying biases, producing unreliable results, and reinforcing inequalities . Rothen highlighted that trusted data is often not visible, findable, or interoperable, making investment in metadata essential . Power demonstrated this concretely by noting that mobile phone data systematically underrepresents certain groups and must be adjusted before feeding into AI systems .
AI tools are only as trustworthy as the data they use, and official statistics produced under internationally agreed standards offer authoritative evidence for AI systems - Data trustworthiness as a prerequisite for AI
Official statistics provide trusted, representative, and well-documented data that AI systems require to avoid bias and unreliable results - Official statistics as a foundation for trustworthy AI
Trusted data must be findable, interoperable, and harmonised; investing in metadata is essential because machines search for text descriptions, not raw figures - Metadata as the key to data discoverability
Mobile phone data underrepresents women, children, the elderly, the rural, and the less wealthy, so it must be combined with survey data to adjust for bias before being used in AI systems - Addressing bias in mobile phone data
Magpantay explicitly stated that the future is not about replacing official statistics with AI, nor replacing surveys with new data sources, but about combining trusted official statistics with responsibly governed new data sources to create richer, more timely, and more relevant data for decision-making . Power reinforced this by arguing that mobile phone data types should not be used by themselves but combined with other data types to adjust for bias, applying equally or more strongly when data feeds into AI systems . Barbosa similarly described how CETIC has established a laboratory of innovative data and new methodologies to seek new data sources to complement traditional survey data .
The future lies in combining trusted official statistics with responsibly governed new data sources, not replacing one with the other - Combining traditional and new data sources
Mobile phone data underrepresents women, children, the elderly, the rural, and the less wealthy, so it must be combined with survey data to adjust for bias before being used in AI systems - Addressing bias in mobile phone data
Stable financing enables long-term statistical time series, methodological independence, and continuous innovation in data collection, including the use of alternative and big data sources - Benefits of stable financing for statistical innovation
Magpantay stated that partnership is becoming very essential because work cannot proceed without governments, national statistical offices, regulators, international organisations, academia, the private sector, and civil society in building a trusted data ecosystem . Rothen described the TDO as involving approximately 20 countries and 30 to 40 international organisations working together, with the platform intended to be Geneva-based to bring together existing knowledge and expand globally . Østreich-Nielsen described the SEVIA Platform for Action as bringing together many different countries and partners to discuss what needs to be done . The Moderator concluded the session by emphasising that no single actor can deliver this and that partnerships and different types of data are needed .
No single actor can deliver trusted data for AI; governments, national statistical offices, regulators, international organisations, academia, the private sector, and civil society must all collaborate - Multi-stakeholder collaboration as essential
The TDO initiative involves approximately 20 countries and 30 to 40 international organisations working together to build a Geneva-based global trusted data platform - TDO as a collaborative international initiative
Donor organisations need to develop practical approaches for evaluating data-related projects and ensuring that funded data work is made available in AI-ready formats - Donor community's role in data collaboration
No single actor can deliver trusted data for AI; partnerships across different types of actors and data sources are essential, and AI outputs must be critically evaluated - Multi-stakeholder collaboration and critical use of AI
Østreich-Nielsen identified a structural problem whereby investment in data and statistics tends to be sectoral and project-based rather than systemic, with donor funding for the Global South often not linked to official statistics or long-term capacity building . Barbosa illustrated this challenge by noting that ICT surveys are frequently financed through mechanisms not based on solid budget commitments, such as temporary government programmes or donor-funded projects, which rarely provide the continuity needed to build long-term statistical capacity . Peltola noted that governments committed in the Financing for Development outcome document to investing in their national statistical systems, providing a political foundation to build upon, though practical implementation varies by country .
Investment in data and statistics is often sectoral and donor-driven, lacking the continuity needed to build long-term statistical capacity, particularly in the Global South - Fragmented and unsustainable data financing
Self-sustainable financing through domain registry revenues - Self-sustainable financing through domain registry revenues
Political commitment to data investment - Political commitment to data investment
Rothen argued that the problem is often not a lack of data but that trusted data is not visible, findable, or interoperable, and that if organisations do not invest in metadata, data sets will never be found because large language models search for text data and not raw figures . He described the TDO as a global discovery platform designed to make trusted data findable by both machines and humans without centralising the data or transferring ownership . Peltola reinforced this by observing that if the TDO held metadata on what data are available, it would also reveal data gaps, which could help target investment to where gaps are most critical for policy needs . Magpantay similarly emphasised that data must be trusted, representative, well-documented, and interoperable .
The Trusted Data Observatory (TDO) aims to create a global metadata discovery platform so that machines and people can find trusted data, without centralising or transferring ownership of the data - TDO as a global metadata platform
Trusted data must be findable, interoperable, and harmonised; investing in metadata is essential because machines search for text descriptions, not raw figures - Metadata as the key to data discoverability
Using metadata to identify and address data gaps - Using metadata to identify and address data gaps
Power argued that protecting the privacy of individuals who contribute data to mobile systems is essential, describing Flowminder's approach of deploying software on the mobile operator's premises that aggregates data before export so that individual-level data never leaves the operator's systems . He also emphasised that having regulator support was really important for legitimising their work in complex country contexts . Magpantay described a model where national statistical offices do not pay mobile operators directly for their data but instead offer incentives such as technical skills and access to statistical data, ensuring data is used responsibly and integrated as an official data source .
Privacy must be protected by processing data at the mobile operator's premises and aggregating it before export, ensuring individual-level data does not leave the source - Privacy protection in mobile data processing
Data governance frameworks, including regulator support and clear principles, are essential for responsible use of novel data sources such as mobile phone data - Data governance for novel data sources
Integrating mobile phone data into official statistics requires bringing all stakeholders together, with national statistical offices offering technical skills and data access as incentives rather than direct payment - Multi-stakeholder model for mobile data integration
Both Magpantay and Power share the view that mobile phone data and other novel data sources are valuable but must be integrated with official statistics rather than used in isolation. Magpantay stated that the future is about combining trusted official statistics with responsibly governed new data sources , while Power argued that mobile data types should not be used by themselves but combined with other data types to adjust for bias, applying equally or more strongly when data feeds into AI systems . Both also emphasised the importance of governance frameworks and stakeholder collaboration in making this integration work responsibly . Both Rothen and Peltola share the view that metadata platforms are not only tools for data discoverability but also instruments for identifying data gaps and targeting investment. Rothen described the TDO as a global discovery platform that makes trusted data findable without centralising it , while Peltola observed that if the TDO held metadata on what data are available, it would also reveal data gaps, helping to connect investment with identified gaps . Both see metadata infrastructure as a practical bridge between data availability and policy action. Both Østreich-Nielsen and Barbosa share the diagnosis that current financing models for data and statistical systems are structurally inadequate, though they approach solutions from different angles. Østreich-Nielsen identified that investment is often sectoral and not linked to official statistics or long-term capacity building , while Barbosa illustrated this by noting that ICT surveys are frequently financed through mechanisms not based on solid budget commitments, which rarely provide the continuity needed . Barbosa's CETIC model — funded by domain registry revenues — is presented as a practical response to the very structural problem Østreich-Nielsen identifies, with both agreeing that stability, independence, and innovation are the desired outcomes . Magpantay, Rothen, Peltola, and the Moderator all share the view that multi-stakeholder collaboration is not optional but essential for building trusted data ecosystems for AI. Magpantay stated that partnership is becoming very essential because work cannot proceed without governments, NSOs, regulators, international organisations, academia, the private sector, and civil society . Rothen described the TDO as involving approximately 20 countries and 30 to 40 international organisations . The Moderator concluded by stating that no single actor can deliver this and committed to taking the discussion back to the UN system and the network of chief statisticians . Both Power and Magpantay share a nuanced view on the question of paying for mobile operator data, agreeing that direct payment is not the preferred model but that operators need some form of incentive or cost recovery. Power described the tension between governments' reluctance to pay for data and operators' monetisation expectations, noting that in the Congo, Flowminder pays only a small fee for server management rather than for the data itself, and that paying for data does not necessarily correlate with data quality . Magpantay described a model where NSOs offer incentives such as technical skills and access to their own data rather than paying operators directly . Both converge on the view that non-monetary incentives and mutual benefit are preferable to straightforward data purchase.
It might have been expected that a global platform initiative like the TDO would involve some degree of data centralisation for efficiency or quality control. However, Rothen explicitly and emphatically stated that the TDO is not the idea to centralise data - the data stays where it is - and is not the idea to take the responsibility or the ownership of the data . This decentralised approach was not challenged by any speaker, including those from international organisations that might have institutional interests in centralisation. This consensus around data sovereignty and distributed ownership, even within a global discovery framework, represents an unexpected convergence between national statistical offices, international organisations, and technical practitioners .
The question raised by Jacques Péguet about downstream payment for data might have been expected to generate debate between those favouring market-based approaches and those favouring public good principles. Instead, both Power and Magpantay converged on a nuanced position that neither simply endorses nor rejects payment. Power noted that paying for data does not necessarily correlate with data quality, citing cases where organisations paid for data that turned out to be of poor quality , while Magpantay described a model where NSOs offer non-monetary incentives such as technical skills and data access rather than direct payment . This unexpected consensus across a technical NGO and an international statistical organisation suggests a shared pragmatic view that transcends the public good versus market debate.
It might have been expected that a technical operational organisation like Flowminder, which works with novel data sources, would emphasise the sufficiency of those sources, while official statisticians would emphasise the primacy of traditional methods. Instead, Power - speaking from the perspective of a practitioner who uses mobile data - explicitly acknowledged its limitations and argued that it must be combined with surveys to correct for bias, stating that this applies just as strongly, if not more so, if that data is going to be used for AI . This methodological humility from a novel data practitioner aligned closely with the position of official statisticians like Magpantay and Barbosa, representing an unexpected bridging of the traditional versus novel data divide .
An unexpected area of consensus emerged around the observation that governments frequently do not know what data they themselves hold, making internal data discovery a necessary first step before any broader sharing or AI-readiness agenda can proceed. Rothen noted that the TDO helps governments understand what data other ministries hold, enabling conversations about access, and that governments often do not know what data sets they have . Østreich-Nielsen similarly observed that many ministries sit on relevant administrative data that are not being made available . The GIZ Egypt Representative reinforced this from a practical implementation perspective, noting that digitising existing data and enabling government-to-government communication are foundational preconditions . This convergence across a national statistical office, a donor agency, and a development cooperation practitioner was not anticipated given their different institutional roles.
The session demonstrated a remarkably high level of consensus across speakers from very different institutional backgrounds - international organisations (ITU, UNCTAD), national statistical offices (Switzerland, Brazil), donor agencies (NORAD), technical NGOs (Flowminder), and development cooperation practitioners (GIZ Egypt). The core areas of agreement centred on: (1) data quality as the fundamental prerequisite for trustworthy AI ; (2) the need to combine official statistics with novel data sources rather than replace one with the other ; (3) the indispensability of multi-stakeholder collaboration ; (4) the structural inadequacy of current data financing, particularly in the Global South ; (5) the centrality of metadata and data discoverability for AI-readiness ; and (6) the necessity of responsible data governance, including privacy protection and bias correction, when using novel data sources . Unexpected consensus emerged around data decentralisation, the limited value of paying for data, the need for bias correction from novel data practitioners themselves, and the foundational challenge of governments not knowing what data they hold.
Jacques Péguet raised the question of whether downstream payment for data should be considered . Benjamin Rothen took a firm public-good position, stating that in Switzerland data cannot be charged for by law and that official statistics data is funded by taxes . By contrast, Daniel Power acknowledged a real tension: mobile operators are told their data is a source of revenue to be monetised, while governments may have hard lines against paying . He noted pragmatically that his organisation would pay if necessary to open a door and respond to a crisis, and observed that paying does not necessarily correlate with data quality . Esperanza Magpantay offered a middle path, suggesting that national statistical offices should not pay directly for data but can offer non-monetary incentives such as technical skills and access to their own datasets . These positions reflect a genuine disagreement about whether and how privately held data should be compensated.
Open government data and FAIR principles are important steps, but metadata platforms that describe what data exists are needed even before data is made openly accessible - Open data and FAIR principles as building blocks
The question of whether to pay for better data or treat it as a free public good raises important political and practical considerations - Paying for private data: tensions and trade-offs
Paying for private data: tensions and trade-offs
Integrating mobile phone data into official statistics requires bringing all stakeholders together, with national statistical offices offering technical skills and data access as incentives rather than direct payment - Multi-stakeholder model for mobile data integration
Rothen acknowledged that in the beginning trusted data is often equated with official statistics because of agreed frameworks such as the UN fundamental principles and FAIR principles, but argued that on the long term the TDO platform should include all trusted data, not only statistical data, because sometimes NGO data is better than national statistical office data . Peltola and Magpantay, while acknowledging new data sources, placed greater emphasis on official statistics as the primary foundation of trustworthiness . This reflects a tension between a broader, more inclusive definition of trusted data and a narrower, standards-based definition anchored in official statistics.
The Trusted Data Observatory (TDO) aims to create a global metadata discovery platform so that machines and people can find trusted data, without centralising or transferring ownership of the data - TDO as a global metadata platform
AI tools are only as trustworthy as the data they use, and official statistics produced under internationally agreed standards offer authoritative evidence for AI systems - Data trustworthiness as a prerequisite for AI
Official statistics provide trusted, representative, and well-documented data that AI systems require to avoid bias and unreliable results - Official statistics as a foundation for trustworthy AI
Barbosa presented Brazil's CETIC model - funded by revenues from the .br country code domain registry - as an innovative and stable financing mechanism that has enabled 20 years of continuous ICT surveys . However, he himself acknowledged that this model is not very easy to replicate, because Brazil ranks sixth largest domain name database among G20 and OECD countries, providing sufficient financial resources that most countries would not have . Østreich-Nielsen, speaking from a donor agency perspective, described the broader reality for the Global South as one of fragmented, sectoral, and donor-driven investment that lacks continuity . These positions implicitly disagree on whether sustainable self-financing is a broadly applicable solution or a context-specific exception.
Self-sustainable financing through domain registry revenues
Investment in data and statistics is often sectoral and donor-driven, lacking the continuity needed to build long-term statistical capacity, particularly in the Global South - Fragmented and unsustainable data financing
It was unexpected that a session focused on trusted data for the public good would surface a pragmatic acceptance of paying for private data. Rothen, representing a national statistical office, took a firm principled position that data is a public good funded by taxes and cannot be charged for under Swiss law . Power, however, stated that his organisation would pay for data if it helped open a door and respond to a crisis, provided a donor was also prepared to do so, and noted that the amount paid does not correlate with data quality . This tension was unexpected because both speakers are ostensibly working towards the same public-interest goal, yet their operational realities lead them to fundamentally different stances on commodification of data. The disagreement reveals a structural fault line between public statistical institutions and technical operational organisations working in humanitarian contexts.
Rothen presented the TDO as an ambitious global metadata discovery platform involving 20 countries and 30 to 40 international organisations, with a prototype planned for the Geneva AI summit in 11 to 12 months . Yet he also acknowledged that Switzerland, after ten years of work on its own metadata platform, is still struggling to make data findable and interoperable . The GIZ Egypt Representative highlighted that many countries face much more basic preconditions - digitising existing data, enabling government-to-government communication - before any metadata platform could be useful . This created an unexpected implicit disagreement about the readiness of the international community to benefit from a global metadata platform, with the TDO's ambition sitting in tension with the foundational challenges many countries still face.
Both Power and Magpantay agreed on the value of mobile phone data, but Power's detailed acknowledgement of its systematic biases - underrepresenting women, children, the elderly, the rural, and the less wealthy - raised an unexpected implicit tension with the broader enthusiasm for mobile data as a complement to official statistics. Power noted that for rapid crisis response, such as the DRC Ebola outbreak, bias adjustments were not applied , yet the data was still used to inform critical public health decisions. This sits in some tension with his own principle that bias must be corrected before data is fed into AI systems . Magpantay's presentation of mobile data as a valuable complement to official statistics did not address these bias limitations with the same degree of candour, creating an unexpected divergence in how honestly the limitations of novel data sources were presented.
The session was characterised by a high degree of surface consensus around the shared goal of better data for AI, with all speakers agreeing on the importance of trusted, interoperable, and well-documented data, the need for multi-stakeholder collaboration, and the value of combining official statistics with novel data sources. However, meaningful disagreements emerged beneath this consensus on three main issues: (1) whether and how to pay for privately held data such as mobile operator data, with positions ranging from a principled public-good stance to pragmatic acceptance of payment to a non-monetary incentive model ; (2) the definition and scope of trusted data, with some speakers anchoring it firmly in official statistics and others arguing for a broader, more inclusive definition ; and (3) the replicability and sequencing of solutions, with ambitious global platform proposals sitting in tension with acknowledgements of very basic foundational challenges in many countries . Additional unexpected tensions arose around the honest acknowledgement of mobile data bias and the gap between the TDO's ambitions and the operational realities of data infrastructure in the Global South.
All three speakers agreed that the future lies in combining official statistics with new data sources rather than replacing one with the other . However, they differed on emphasis and method. Magpantay stressed the combination of trusted official statistics with responsibly governed new sources as the key lesson from the mobile phone data initiative . Power agreed that mobile phone data must be combined with survey data to adjust for bias before being used in AI systems, and that these data types should not be used by themselves . Rothen agreed that metadata platforms should eventually include all trusted data, not just official statistics . The shared goal is integration, but the speakers diverged on how much weight to give official statistics versus novel sources, and on the governance mechanisms needed to ensure responsible combination.
The future lies in combining trusted official statistics with responsibly governed new data sources, not replacing one with the other - Combining traditional and new data sources Addressing bias in mobile phone data Open data and FAIR principles are important steps, but metadata platforms that describe what data exists are needed even before data is made openly accessible - Open data and FAIR principles as building blocks
All three speakers agreed on the need for stable, long-term financing for data systems and on treating data as a public good . However, they proposed different mechanisms to achieve this. Østreich-Nielsen focused on reforming donor funding approaches and improving national coordination through the SEVIA recommendations . Barbosa advocated for self-sustainable financing models tied to specific revenue streams such as domain registries . Rothen called for investment in the TDO as a shared international infrastructure, inviting participants to contribute funding . The shared goal is sustainable data financing, but the pathways proposed — donor reform, self-financing, and international platform investment — reflect different assumptions about where responsibility and resources should lie.
The SEVIA Platform for Action recommends treating data and statistics as a public good, improving coordination, strengthening data governance, and investing in national statistical capacity - Four recommendations for sustainable data financing Stable financing enables long-term statistical time series, methodological independence, and continuous innovation in data collection, including the use of alternative and big data sources - Benefits of stable financing for statistical innovation The TDO initiative involves approximately 20 countries and 30 to 40 international organisations working together to build a Geneva-based global trusted data platform - TDO as a collaborative international initiative
Both Power and Magpantay agreed that long-term partnerships with mobile operators and regulators are essential for responsible use of mobile phone data . However, they differed on the institutional model. Power, as a technical operational organisation, described a bilateral relationship with a single operator (Vodacom Congo) built over eight years, with regulator support as a secondary but important factor . Magpantay described a broader multi-stakeholder model involving national statistical offices, operators, regulators, and international organisations across 25 countries, with the explicit goal of integrating mobile data into official statistics sustainably . The shared goal is responsible mobile data use, but Power's model is more operationally pragmatic and bilateral, while Magpantay's is more institutionally structured and multilateral.
Long-term operator and regulator partnerships are critical for accessing and responsibly using mobile data in complex country contexts - Long-term operator and regulator partnerships Integrating mobile phone data into official statistics requires bringing all stakeholders together, with national statistical offices offering technical skills and data access as incentives rather than direct payment - Multi-stakeholder model for mobile data integration
All three speakers agreed that foundational preconditions must be addressed before AI-ready data can be produced at scale . The GIZ Egypt Representative identified digitising existing data and enabling government-to-government data sharing as immediate priorities . Østreich-Nielsen highlighted the problem of ministries sitting on administrative data that is not being made available . Rothen acknowledged that even in Switzerland, after ten years of work on a metadata platform, the task of making data findable and interoperable remains difficult . However, they differed on the sequencing and urgency of these steps, with the GIZ representative emphasising the very basic starting point of digitisation, while Rothen and Østreich-Nielsen were already discussing more advanced interoperability and governance challenges.
Foundational preconditions for AI-ready data Many ministries hold relevant administrative data that are not publicly available, representing a major untapped resource for decision-making and AI, particularly in African contexts - Unlocking administrative data as a priority Trusted data must be findable, interoperable, and harmonised; investing in metadata is essential because machines search for text descriptions, not raw figures - Metadata as the key to data discoverability
- AI systems are only as trustworthy as the data they are built upon; official statistics, produced under internationally agreed standards with continuous public oversight, provide a uniquely authoritative and trusted foundation for AI.
- The future of data for AI lies not in replacing official statistics with new data sources, but in combining trusted official statistics with responsibly governed novel sources such as mobile phone data, administrative data, and big data.
- Trusted data must be findable, interoperable, harmonised, and well-documented through metadata; without investment in metadata, machines cannot locate or use data effectively, regardless of its quality.
- The Trusted Data Observatory (TDO) initiative, led by Switzerland and involving approximately 20 countries and 30–40 international organisations, aims to create a Geneva-based global metadata discovery platform that makes trusted data visible without centralising ownership.
- Mobile phone data offers significant value for humanitarian and development purposes, as demonstrated by the rapid identification of Ebola-risk health zones in the DRC, but it carries inherent biases (underrepresenting women, children, the elderly, rural, and less wealthy populations) that must be corrected through complementary survey data before use in AI systems.
- Privacy protection in the use of mobile operator data requires processing and aggregating data at the operator's premises so that individual-level data never leaves the source.
- Financing for data and statistics is often fragmented, sectoral, and donor-driven, lacking the continuity needed to build long-term statistical capacity, particularly in the Global South.
- Brazil's CETIC model, funded by revenues from the .br country code domain registry, demonstrates that self-sustainable, independent, and innovative financing mechanisms for continuous ICT surveys are achievable, though not easily replicable in all contexts.
- The SEVIA Platform for Action, developed in the context of the Financing for Development agenda, proposes four recommendations: treating data and statistics as a public good, improving national coordination, strengthening data governance and innovation, and investing in national statistical capacity.
- No single actor can deliver trusted data for AI; effective collaboration among governments, national statistical offices, regulators, international organisations, academia, the private sector, and civil society is essential.
- Many countries face foundational preconditions before AI-ready data can be produced, including digitising existing data, enabling government-to-government data sharing, and incentivising both public and private actors.
- Many ministries hold relevant administrative data that are not publicly available, representing a major untapped resource for decision-making and AI, particularly in African contexts.
- Metadata platforms that describe what data exists can help identify data gaps and target investment where it is most critically needed for policy purposes.
- Governments committed in the Financing for Development outcome document to investing in national statistical systems, providing a political foundation upon which to build further action.
“The real limiting factor is increasingly the quality of the data that these models learn from... Without these characteristics, AI systems risk amplifying biases, producing unreliable results, and reinforcing inequalities. The future is not about replacing official statistics with AI, nor is it about replacing surveys with new data sources. It is about combining.”
“Often what we do not lack, there's not a lack of enough data. It's often a lack that our trusted data are not visible. They are not findable. They are often not interoperable... Our customers are not humans anymore. Often they're machines.”
“Because statistics and data are a cross-cutting field, there is a lot of investment in data, but it's often sectoral... A lot of the support that goes into data systems is sectoral, and it's not often or always linked to official statistics. The focus for getting data is often your own data needs and not necessarily thinking all the way to making data available for countries in general, but also for others and for AI purposes.”
“A week after we'd released our first report... Ebola was detected in ten more health zones, and eight out of those ten were in the top regions which we had identified using this mobility data... Mobile phone data is not [representative of the population]. It underrepresents women. It underrepresents children, the very old, the more rural, and the less wealthy.”
“In Brazil, the Regional Center for Studies on the Development of the Information Society, CETIC, is running now for 20 years continuous ICT surveys through a self-sustainable financing mechanism. It's funded by the .br country code top-level domain... all the surveys are 100% funded by the registry activities for the .br.”
“If we are not investing in metadata you will never be able to find the data sets... machines will never be able to find the data because the LLMs, they are not looking for data, they are looking for text data, that's why you have a good description on it.”
“We don't pay Vodacom for their data... In other countries we do see operators want to see a return even if it's a modest one for their data and that introduces challenges. It is not however the case that I see the amount you pay correlating with data quality.”
What does 'trusted data' look like in practice, and how do we define AI-ready data?
Benjamin Rothen acknowledged this is not an easy question to answer fully, and invited collaboration to define it. This is foundational to the entire Trusted Data Observatory initiative and to ensuring AI systems are built on reliable foundations.
How can the Trusted Data Observatory be expanded and sustained institutionally, and what investment is needed to build a working prototype by the AI Summit in Geneva in 11-12 months?
Rothen explicitly called for partners and funding to help build the TDO prototype, indicating this is an open and urgent area requiring further action and research into governance, financing, and technical architecture.
How can countries with significant preconditions (e.g., undigitised data, lack of inter-governmental data sharing) practically move towards open data and AI-ready data ecosystems?
The question highlighted that many countries, including Egypt, face foundational barriers before AI-ready data is achievable. Understanding how to sequence and prioritise these preconditions is critical for practical implementation, especially in the Global South.
Should users or downstream actors pay for better data, and if so, how should pricing models be structured without compromising data quality or access?
This question was raised by an audience member and addressed partially by multiple panellists. There is no consensus, and the tension between data as a public good and commercial incentives for private data holders (e.g., mobile operators) remains unresolved and requires further policy and economic research.
How can the model used in Côte d'Ivoire for integrating mobile phone data into official statistics be replicated or adapted in other countries?
The Côte d'Ivoire representative asked about replicating the mobile data integration model. Esperanza confirmed Côte d'Ivoire is part of a 25-country project, but a fuller explanation of the model and its transferability was deferred to a post-session conversation, indicating a need for further documentation and dissemination.
How can bias in non-traditional data sources such as mobile phone data (which underrepresents women, children, rural populations, and the less wealthy) be systematically corrected before such data is used in AI systems?
Power highlighted that mobile phone data is not fully representative of populations and that bias correction is essential, especially when data feeds into AI systems. Further research is needed into standardised methodologies for bias adjustment across different data types and contexts.
How can sectoral donor funding for data in the Global South be better coordinated and redirected towards building sustainable, AI-ready national statistical systems rather than fragmented, project-based data collection?
Vibeke identified a structural problem where donor funding is siloed and not linked to official statistics or long-term data availability. Developing practical frameworks for donor organisations to evaluate and align their data investments is an identified area for further work.
How can the Brazil CETIC financing model (funded through country-code domain name registrations) be adapted or inspire sustainable financing mechanisms for ICT and digital statistics in other countries?
Barbosa acknowledged the model is difficult to replicate directly due to Brazil's unique scale of domain registrations, but the underlying principle of linking digital economy revenues to statistical capacity is worth exploring for other national contexts.
How can metadata standards and platforms be developed and invested in globally so that machines (including AI systems) can discover and use trusted data more effectively?
Rothen stressed that without investment in metadata, data sets remain invisible to AI systems. This points to a significant gap in current infrastructure and a need for further research into metadata standards, interoperability frameworks, and the resources required to implement them at scale.
How can official statistics and new data sources (such as mobile phone data, administrative data, and big data) be systematically combined to produce richer, more timely, and more representative data for AI and policy use?
Multiple panellists emphasised that the future lies in combining traditional official statistics with responsibly governed new data sources, but the methodological, governance, and institutional frameworks for doing so at scale remain underdeveloped and require further research.
How can governments be better incentivised to publish administrative data held by ministries in AI-ready formats, particularly in contexts such as Africa where publicly available data is scarce?
Vibeke noted that many ministries sit on relevant administrative data that is not made publicly available, which limits both policy use and AI applications. Understanding the political, legal, and technical barriers to releasing such data is an important area for further investigation.
What concrete progress or deliverables should the international statistical and AI community aim to demonstrate at the next major international summit or conference?
The moderator explicitly posed this as a forward-looking question to the panel, and while the TDO prototype was mentioned as one goal, a broader and more comprehensive answer was not fully developed, suggesting this requires further strategic planning across the community.
How can the privacy of individuals contributing data to mobile operator datasets be protected as AI systems become increasingly data-hungry, particularly in conflict-affected or fragile contexts such as the DRC?
Power raised the importance of processing data on-premises at mobile operators to avoid individual-level data leaving the country, but broader questions about privacy-preserving techniques, legal frameworks, and governance in fragile states remain open for further research.
How can the international statistical community scale up and connect its work on trusted data with private sector partners and broader digital governance initiatives such as the Global Digital Compact?
The moderator concluded by noting the need to bring this agenda to the UN system and the network of chief statisticians, indicating that the institutional and political pathways for scaling trusted data initiatives internationally remain to be fully mapped out.
