WSIS Forum 2026
AI-generated report

Teaching AI to Speak our Language: A Showcase of Global Efforts to Bridge the ‘Last Mile’ of AI Inclusion for Local Impact

6 speakers
Summary

This discussion centred on the challenge of developing AI-powered language tools for low-resource and underrepresented languages, with a particular focus on humanitarian and peacekeeping contexts. Three main projects were presented, alongside a broader international coordination initiative.

James D'Ercole described a UN Peacekeeping project based in Juba, South Sudan, aimed at building a real-time translation tool to address communication gaps among approximately 50,000 uniformed personnel from around 120 countries . The tool is designed to operate entirely offline in low-connectivity environments , with a human-in-the-loop approach that ensures AI assists rather than replaces human judgement . A key ambition is to build not just a language dataset but a cultural corpus that incorporates local meaning and context , and to preserve Juba Arabic as a living language that could eventually be accessible online .

Aimee Ansari of CLEAR Global highlighted the broader global disparity in language AI support, noting that most languages spoken in lower-income countries remain poorly represented in frontier models . She illustrated this with northeast Nigeria, where only Hausa has any meaningful automatic speech recognition support, potentially excluding around 70% of the population from AI-assisted communication . She emphasised the need for high-quality, diverse voice data collected systematically across speakers of different ages, genders, and dialects , and stressed the importance of informed consent and open, non-commercial data licensing .

Barbora Bromová presented the Loria tool, developed with the National Library of Serbia and UNDP, which automates the digitisation of historical documents through optical character recognition and post-processing . Dafna Feinholz introduced UNESCO's Coalition for Linguistic Diversity in AI, launched in June 2025, which brings together over 30 experts from governments, academia, communities, and the private sector to share good practices and build a repository of community-led approaches .

A closing discussion highlighted two persistent challenges: the high cost of ethical, community-centred data collection , and the difficulty of coordinating fragmented efforts across organisations . Ansari noted that while private sector involvement is valuable, commercially uninteresting language communities risk being left behind, underscoring the need for international organisations to help ensure equitable representation in global AI infrastructure .

Keypoints
  • Overall Purpose

  • The discussion centres on the challenge of developing AI-powered language and translation tools for low-resource and underrepresented languages, particularly in humanitarian and peacekeeping contexts. Speakers from UN Peacekeeping (UNDPKO), CLEAR Global, UNDP, and UNESCO share their respective projects and explore how collaboration, community involvement, and ethical data governance can help bridge the global language technology gap.
  • --
  • Major Discussion Points

  • The critical communication gap in peacekeeping and humanitarian operations: UN peacekeeping missions involve approximately 50,000 uniformed personnel from around 120 troop- and police-contributing countries, yet communication with local communities and between peacekeepers themselves remains a persistent, long-standing problem. A real-time, offline translation tool is being piloted in Juba, South Sudan, specifically designed to function in low-connectivity, austere environments, with a human-in-the-loop approach to ensure AI assists rather than replaces human judgement. - The severe underrepresentation of low-resource languages in AI models: Most languages spoken by large populations in the Global South remain poorly supported or entirely absent from frontier AI models, with text and speech data heavily skewed towards wealthier, English-dominant countries. For example, in northeast Nigeria, of ten commonly spoken languages, only Hausa has any meaningful automatic speech recognition (ASR) support, risking the exclusion of approximately 70% of the population from AI-assisted communication workflows. Challenges include the primarily oral nature of many languages, code-switching, lack of standardised orthographies, and the fact that laboratory performance metrics do not reflect real-world accuracy. - The necessity of community-centred, culturally sensitive data collection: Multiple speakers emphasised that building effective language AI requires deep community involvement from the very beginning of the design process, not as an afterthought. This includes co-designing recording approaches for primarily oral languages, obtaining informed consent, protecting voice data, and ensuring that cultural context - not just linguistic data - is embedded in the models. The UNDPKO project in South Sudan, for instance, is working to build a cultural corpus with national colleagues to make the language model less Western-centric. - Data governance, ownership, and the risk of exploitation: A recurring concern across presentations was ensuring that communities whose language data is collected retain meaningful ownership and are not exploited. Amy Ansari highlighted the tension between the high cost of ethical data collection - requiring fair wages and proper consent - and the limited resources of grassroots organisations, suggesting that pooling resources across multiple organisations could make data collection more financially viable. The question of private sector involvement was also raised, with acknowledgement that while some partners are open-minded about data governance, commercial interest in very low-resource languages remains limited because speakers are not large consumers in digital markets. - International coordination and the Coalition for Linguistic Diversity in AI: UNESCO, in partnership with Iceland, has established the Coalition for Linguistic Diversity in Artificial Intelligence, launched in June 2025, which brings together over 30 experts from governments, academia, communities, technical experts, international organisations, and the private sector. The coalition aims to document good practices in a shared repository, identify gaps in capacity building, and inform policy guidance on linguistic diversity in AI - moving away from siloed efforts towards a collaborative, multi-stakeholder model. Dafna Feinholz stressed that linguistic preservation is inseparable from cultural preservation and that communities must be included across the entire AI development lifecycle. ---
  • Overall Tone

  • The overall tone of the discussion is collaborative, earnest, and solutions-oriented, with an undercurrent of urgency. Presenters speak with genuine passion for their work and a shared commitment to equity and inclusion in AI development. The tone is largely optimistic - particularly when speakers describe community engagement successes and the growing coalition of partners - but is tempered by frank acknowledgements of significant structural challenges, including funding constraints, data governance complexities, and the limited commercial incentive for private sector actors to invest in low-resource languages. Towards the end, during the audience Q&A, the tone becomes slightly more candid and informal, with speakers adding a grounded, practical dimension to what had been a more formal set of presentations. The closing remarks maintain a warm, collegial tone and offer a clear invitation to continued collaboration.
Speakers Overview
JD
James D'Ercole
135 wpm · 12 min
AA
Aimee Ansari
133 wpm · 16 min
BB
Barbora Bromová
137 wpm · 13 min
DF
Dafna Feinholz
138 wpm · 8 min
A
Audience
167 wpm · 3 min
S
Speaker
172 wpm · 1 min

Expanded Summary: AI-Powered Language Tools for Low-Resource Languages in Humanitarian and Peacekeeping Contexts

#

Overview and Framing

This discussion brought together practitioners from UN Peacekeeping (UNDPKO), Clear Global, UNDP, and UNESCO to examine the challenge of developing AI-powered language and translation tools for low-resource and underrepresented languages, particularly in humanitarian and peacekeeping contexts. The session was structured around three project presentations followed by an introduction to an international coordination initiative and a candid audience discussion. Bromová served a dual role throughout as both session moderator and presenter of the Serbian digitisation project. A unifying conceptual thread was introduced at the outset by James D'Ercole, drawing on his experience as a regional administrative officer in East Timor: the principle of focusing on the "last mile" . His argument was that if a system functions correctly at the furthest, most resource-constrained point from headquarters - whether in a mission, in New York, or in Geneva - then everything along the chain must necessarily be working . This framing set the philosophical tone for the entire session, grounding the technical work in operational reality rather than institutional convenience.

#

The Communication Gap in UN Peacekeeping

D'Ercole opened by describing the scale and persistence of the communication problem facing UN peacekeeping operations. Having spent a little over 20 years in peacekeeping, he characterised this as a systemic and enduring challenge. Approximately 50,000 or more uniformed personnel are deployed daily across 11 missions, drawn from approximately 120 troop- and police-contributing countries . Communication - both between peacekeepers and the communities they serve, and among peacekeepers themselves - is foundational to the peacekeeping mission, which depends on listening and understanding . Yet this has remained a persistent, systemic problem throughout D'Ercole's career, documented repeatedly in field reports, C-34 submissions from governments, military and police advisory committee visits, and human rights reports . Interpreters and translators have never been available in sufficient numbers, even in better-resourced periods , and the diversity of peacekeepers' own national languages compounds the challenge further .

The project being piloted in Juba, South Sudan, is specifically designed to address this gap through real-time translation . Crucially, it is built around a human-centred oversight approach: AI is conceived as assisting the process, with a person always present to validate outputs, never replacing human judgement or eliminating the roles of language assistants and interpreters . D'Ercole argued that this approach would, if anything, amplify the reach and influence of existing language professionals rather than diminish them . A central technical requirement is that the tool must function completely offline, given the low or absent connectivity in the environments where peacekeeping missions operate . An edge device - described as a mini-server capable of running on a battery pack - is currently on loan and being tested as the hardware platform for this capability .

The translation model at the heart of the project was developed with contributions from a Harvard-linked effort and an NYU capstone project involving graduate students, reflecting the collaborative academic partnerships that have shaped the initiative from its early stages. D'Ercole described a specific iterative improvement cycle: offline testing feeds into dataset building and enhancement, which involves UN Volunteers working online, whose contributions flow into the translation model, which is then deployed on the edge device. The proof-of-concept phase operates in what D'Ercole called "shadow mode," in which the device sits and listens to an interpreter working, and the interpreter then reviews the device's outputs - validating them, correcting them, and in doing so helping to build the guardrails for the system. This validation loop then feeds back into the next improvement cycle, creating a continuously refined model grounded in real operational experience.

#

Building a Cultural Corpus, Not Just a Language Dataset

A distinctive ambition of the Juba project is to build what D'Ercole described as a cultural corpus rather than merely a language dataset . The project has engaged an expert in integrating culture into AI language models specifically to make the model less Western-centric and more contextually appropriate for the Juba Arabic-speaking community . This involves validating data collected by online UN Volunteers, then engaging national colleagues within the mission to further validate and enrich that data with cultural context . Language assistants have already contributed voluntarily on their own time, motivated by genuine belief in the project and in the importance of digitising Juba Arabic .

D'Ercole identified the preservation dimension of this work as the aspect he values most , describing the goal as creating a "living language" - preserving not only the language itself but the culture embedded within it . He further noted that once a language is digitised and made available as open-source data, commercial partners could eventually develop further tools enabling communities to access the internet in their own language . D'Ercole articulated three core benefits of the project: better communications, deeper trust, and a living language. The project is being built through collaborative partnerships with local institutions, international universities, civil society organisations, and NGOs, with a proposed language cooperative intended to ensure that the communities contributing data can eventually benefit from it . Data governance, AI safety, and voice data protection are being addressed under a "do no harm" approach, with community ownership and mutually defined guardrails treated as non-negotiable . This work is being conducted under the oversight of the UN Office of Data Protection and Privacy, a new office that D'Ercole described as both valuable and, given its novelty, sometimes challenging to navigate .

#

The Global Disparity in Language AI Coverage

Aimee Ansari of Clear Global situated the Juba project within a much broader global pattern of inequality in AI language coverage. Drawing on research conducted in 2025, she described how the amount of text data available online across languages - a key determinant of how well large language models perform - is heavily skewed towards languages spoken in wealthier countries, while languages with tens or hundreds of millions of speakers in the Global South remain poorly supported or entirely absent from frontier models . She illustrated this disparity with a striking comparison: Breton, spoken by approximately 200,000 people in France, has over 50 AI models, while Nigerian Pidgin, with approximately 85 million speakers, has only 10 . She also cited Seraki in northeast Nigeria as a further example of a language with negligible AI coverage. This gap is not driven by communicative need or speaker population, but by economic and geopolitical power.

The situation is even more acute for speech models. Ansari noted that only a fraction of languages globally are meaningfully supported by automatic speech recognition (ASR) technology . In northeast Nigeria, of the ten most commonly spoken primary languages in Borno, Adamawa, and Yobe states, only Hausa has any ASR support, including a commercial API . Even for Hausa, performance has only been assessed in laboratory conditions, and real-world field performance remains largely unknown . The practical consequence is that ASR-assisted workflows relying on Hausa or English risk excluding approximately 70% of the population of northeast Nigeria from AI-assisted communication entirely .

#

Challenges Specific to Low-Resource and Oral Languages

Ansari identified several interconnected challenges that make developing ASR for low-resource languages particularly difficult. Many such languages are primarily oral, meaning that manual transcription - a prerequisite for building ASR models - is an extremely complex undertaking . Terminology and ways of expressing oneself vary widely among speakers, and code-switching between two languages is common, creating further complexity for digital language technologies . A critical methodological point she raised is that laboratory-based word error rate metrics do not reliably predict real-world performance: a low word error rate typically means a researcher has tested the model in a controlled setting, which tells us little about how it will function with real speakers in real environments . At the centre of this gap is a shortage of high-quality, diverse voice data: building a reliable ASR model requires hundreds of hours of recorded data collected systematically from speakers of different ages, genders, educational levels, and dialects .

Clear Global's response to these challenges is a community-centred approach to data collection. The organisation collaborates with community members from the very beginning of the design process - before any data collection begins - to understand how models will be used, to grasp linguistic complexities, and to identify barriers . This includes training community members as linguists and co-designing approaches to writing down primarily oral languages, since there may be no standard alphabet or agreed written form . Ansari gave the example of a recording in Canary where ten different, equally accepted ways of writing the same word existed and contributors could not agree on a single correct version . Co-design approaches are used to manage this complexity and achieve sufficient consistency for model development . Informed consent is a foundational principle: community workshops are held to ensure speakers understand how their recordings will be used and stored, what the risks are, and how they can withdraw consent at any time . All datasets are published openly under non-commercial licences so they can be used across the sector .

#

Digitising Historical Documents: The Loria Project in Serbia

Barbora Bromová presented a complementary project addressing a different dimension of the language data gap: the digitisation of historical documents trapped in paper formats. Working with the National Library of Serbia and UNDP, the team developed a workflow - within the Librarify project - to process scans from the library's archives using optical character recognition, layout identification, image enhancement, and post-OCR correction . Serbian presents particular challenges: despite having a significant body of historical written sources, it sits at the bottom of language data distributions , partly because the language uses both Cyrillic and Latin alphabets interchangeably, which confounds standard OCR models . The initial phase of the project digitised over 400 gigabytes and over 16,000 documents .

From this experience, the team developed Loria, a more generalisable, open-source tool published on GitHub for reuse and adaptation across different languages and institutions . Loria follows the same four-stage workflow - image enhancement, layout identification, OCR, and post-OCR correction - but is built for reuse and includes a plug-in for frontier models to enable further post-OCR correction and batch processing . Critically, it was developed with archivists at the centre of the human oversight process , reflecting the broader session consensus that domain expertise and institutional knowledge are as important as technical capability . The project involved a multidisciplinary team spanning machine learning experts, full-stack developers, designers, product managers, the National Library of Serbia, the Mathematical Institute at the Academy of Sciences, and UNDP .

#

UNESCO's Coalition for Linguistic Diversity in AI

Dafna Feinholz of UNESCO introduced the international coordination initiative that seeks to connect these various efforts: the Coalition for Linguistic Diversity in Artificial Intelligence, established in partnership with Iceland's Ministry of Culture, Innovation and Higher Education and the Icelandic Centre for Language and Technology . Feinholz described the coalition as having started approximately a year prior and as having been formally launched at UNESCO's Global Forum on Ethics of Artificial Intelligence in Bangkok "last week," with the next such forum scheduled in Saudi Arabia from 14 to 17 September . The coalition already comprises more than 30 experts working on community-led digital inclusion of languages from almost all world regions . Feinholz grounded the coalition's work in a philosophical argument: language is not merely a communication tool but the medium through which people express how they view the world, understand themselves, and transmit their heritage . Preserving language is therefore inseparable from preserving culture, and this principle is embedded in UNESCO's Recommendation on the Ethics of AI, which includes the protection of linguistic diversity .

The coalition takes a deliberately multi-stakeholder approach, bringing together governments, academia, communities, technical experts, international organisations, and the private sector in the same space . Feinholz emphasised that this collaborative model has proven valuable because knowledge that works for one community often works for others, and because having large companies and small startups in the same group allows smaller actors to benefit from capabilities they could not develop independently . A key output of the coalition is a repository of good practices, built around a co-created questionnaire designed in collaboration between all stakeholders and UNESCO - on the basis that each stakeholder knows their own project best - to systematise and make accessible knowledge about community-led approaches to linguistic and cultural diversity in AI . This repository is conceived as a living resource that will inform policy guidance on linguistic diversity in AI - an area identified as a priority by many member states . Feinholz stressed that meaningful linguistic and cultural inclusion requires communities to be involved across the entire life cycle of AI development, from initial design through deployment, rather than being brought in at late stages .

#

Closing Discussion: Funding, Coordination, and the Private Sector

The session's closing discussion, prompted by Bromová's question about what international organisations could most usefully do to support organisations like Clear Global, produced some of the most candid and practically significant exchanges of the session . Ansari identified two principal areas where international organisations could make a difference. The first is financial: ethical, community-centred data collection is expensive, requiring fair wages for contributors rather than reliance on volunteers, and most grassroots organisations typically cannot afford the $50,000-$100,000 cost of building a language model . However, if five or six organisations working on related languages could be brought together, each contributing approximately $10,000, the collective cost becomes manageable . International organisations, with their systemic overview of who is working on what, are uniquely positioned to perform this convening function - a capacity that organisations like Clear Global lack .

The second area Ansari identified was norm-setting around sovereign AI models. She argued that while the global dialogue on AI governance is a good start, there is insufficient norm-setting around how all communities and languages are represented in nationally built AI models . She used South Sudan as a pointed illustration: in a government composed predominantly of one ethnic group following a civil war, the likelihood of building a sovereign model that fairly represents the language and culture of other groups is low . She argued that one of the UN's roles must be to ensure that global public infrastructure is built in a way that is fair and equitable .

Bromová, noting that she was not speaking on behalf of UNDP, observed that while some private sector partners are genuinely open-minded about data governance, many are not, and the challenge of data ownership remains unresolved . She also mentioned that her team has been exploring decentralised data-sharing platforms as one way of navigating these ownership challenges.

An audience member who works with Meta on building large language models introduced themselves with self-deprecating humour as working "for the enemy, sort of" , before reframing the challenge as a shared technical and cost problem. They noted that even within large technology companies, teams focused on internationalisation grapple with the same issues of cultural nuance, code-switching, and non-canonical written forms . They posed what they called the "two-cost problem": the dual difficulty of either fighting to get a model built in a given language or fighting to find people who know the language well enough to contribute meaningfully . Ansari responded by explicitly stating that "Meta is not the enemy at all" and describing Clear Global's collaboration with Meta to test model safety , while simultaneously maintaining that speakers of languages like Dinka are "not commercially interesting" to large companies because they are not significant consumers in digital markets . She also noted that Clear Global works extensively with organisations such as Viamo, the GSMA, and Lilapa AI - smaller and mid-sized private sector actors who occupy a different position in the ecosystem from the large frontier model companies.

A further contribution from the floor, from a separate unnamed participant, proposed that building standardised orthographic structures for primarily oral languages - potentially endorsed by bodies such as IEEE or UNESCO - could help preserve nuance and context in transcription and ASR development . This proposal sat in productive tension with Ansari's earlier empirical observation that even community members themselves sometimes cannot agree on a single written standard, suggesting that top-down standardisation may be neither straightforwardly achievable nor always appropriate .

#

Unresolved Challenges and Forward Agenda

The session concluded with a shared recognition that, despite strong conceptual consensus across all speakers on the principles of community-centred design, human oversight of AI outputs, cultural preservation, and ethical data governance, significant structural challenges remain unresolved. These include the sustainable funding of language model development for commercially unattractive languages , the governance of voice data and the risk of extractivism , the political dimensions of sovereign AI model building in divided societies , and the practical difficulty of coordinating fragmented efforts across the ecosystem . The UNESCO Coalition for Linguistic Diversity in AI, the Loria open-source tool, and the proposed language cooperative in Juba all represent concrete steps towards addressing these challenges, but all speakers acknowledged that the work is at an early stage and that the barriers ahead are as much political and economic as they are technical . The session closed with an open invitation for collaboration, with Bromová encouraging any interested private sector organisations to make contact and all speakers expressing enthusiasm for continued partnership across the sector .

James D'Ercole
I think it's okay. I think we're okay. Yeah, we're fine. Okay. All right. So Pilot is in Juba, South Sudan, and it's specifically built with the community and focusing on the last mile. Now, for me, the last mile kind of means a lot because in my past, I was a regional administrative officer for the mission in East Timor. So it was the last mile. I was out there, and I knew that if you get it right at the furthest distance away from headquarters, whether that's headquarters in a mission or headquarters in New York or Geneva, then everything along the way has to be working. Focus on the last mile, and this is what we're focusing on specifically on this project. So the project. Problem. What we're dealing with is we've got about 50 ,000 or more uniformed personnel out there every day in 11 missions. And this represents approximately 120 or so troop and police contributing countries. Now, it might not be as much as maybe UNDP collectively, but it's quite significant for us. And all of this peacekeeping, it depends on communication. I mean, this is a of peacekeeping is listening and understanding to the communities that we're working in. Now, since I've been in peacekeeping a little over 20 years now, it's always been a problem. And every year you get countless field reports, whether it's C -34 coming from governments, whether it's a military police advisors committee is going out to mission, whether it's human rights reports coming back. And it's always been a problem. we're just, we're not communicating as well as we could be. Interpreters, translators, and even if we did have the money in the better days of yesteryear, you still can't have enough, right? So this project is trying to look at that. But also, not only that, we've got 120 different countries of actual peacekeepers. So the communication between peacekeepers is also a major challenge. And the people that we serve, the peacekeepers themselves, are coming from low -resource languages. So the solution. So what we're doing is we're focusing on this real -time translation. It's a tool to help close this gap that I was telling you about. What we're doing is specifically with the human in the loop. we're looking at it as AI as assisting the process person will always validate should always be there it never replaces human judgment or it doesn't remove the jobs if anything it kind of amplifies and elevates by language assistance or interpreters because the focus now will turn into the individual having a larger influence on several but I'll get back to that later on we're also focusing on cultural sensitivity we're the United Nations so of course we need to build the device or the tool making sure that local meaning and register and context are all built into the whole process and specifically in the environments that we work there's little to no connectivity. So no signal. The device, the tool, has to work completely offline and in kind of austere environments. I don't know if I'm going Yeah, you may pick it up. Okay. So, all right. So I'll pick it up a little bit. Sorry about that. So the benefits of what we're trying to do are better communications, deeper trust, trust, and a living language. So improving communications, it's obvious, right? The more we have the better communications, the better delivery, we can deliver the mandate better. With more communications and understanding with the people that we're serving, it builds trust, listening, inclusion, and hopefully a durable peace. And the part I like the most is that we're actually, by digitizing these languages, we're preserving the language. When we preserve the language, we preserve the culture. All right. And that to me is really important. And when we do that, it has the language in theory, then commercial partners can one day pick up that language and you may be able to actually surf the web in that in that language itself. OK, so. Real time translation, the whole cycle of this is first we try it offline, again, human in the loop, so we're testing it offline. We're kind of building enhancing data sets. We've had UNVs do this online UNVs. They worked very well. A translation model that we've built in basically Gartner. Harvard was originally a part of that. We had a UN cap NYU capstone project. So graduate students working on that. Then we have the edge device. So we have a sample edge device on loan right now. That is like a mini. server that can go out there and as long as we have the battery pack and can plug it in, it can actually do this remotely without any connectivity. And then the proof of concept. So what we want to do is we want to take that, we want to get it out there and put it into shadow mode where it sits and it listens to an interpreter do what it does and then the interpreter itself will now go through it, say yes, no, help build the guardrails and we'll go through it like that and then with that improvement cycle. Alright, how much time? I'm almost there. So what are we doing? We're building a cultural corpus, not just a language data set. So we're validating the data right now that was done by online volunteers. Now we want to take national colleagues in mission and actually validate that data. At the same time, we want to use national colleagues to collect more data. to make it more robust and add a cultural corpus into this. So we just an expert in this kind of integrating culture into the AI language models in order to make this language model less Western -centric and more specific to Juba in the context that we're in. Okay, look at that. Okay, so I think I skipped one. Let's see. So it's built with the community. We're doing it three ways. We're always looking for partnerships and a collaborative approach. So we've got local exploring partnerships and collaboration with. We've got international institutions and universities, CSOs, NGOs, and also we want to try, in theory, to get a language cooperative together so that the people that we're tapping into for this information can somehow one day. kind of get it back into them if this data is open source, which it is, but has a paywall to companies and things like that. So we're going to try to figure that out. We're also leaning heavily on national colleagues who have been really supportive of this. In fact, during the online volunteer, which is all that they don't get paid, we actually had language assistants on their own time contributing this because of actually believing in the project and believing in digitizing the language of Juba Arabic. So we're also leaning on national colleagues to make sure that it's kind of linguistically and culturally accepted to be working in, so we're relying heavily on them. UN Volunteers, absolutely fabulous organization to work with. We're working with online volunteers, and we're going to try to get some national online. in Unmiss itself to work there in order to kind of record, pair, and help validate the process as we go along. And then data governance, AI, safety and security, so voices protected data. So we're taking all that comes to that. The human in the loop, once again, we're saying that we're kind of throwing that down maybe too much, but, you know, we really want to do this AI safety by design. We want to integrate it from the beginning and amplify the reach. So basically when we're doing this, we want to make sure we're doing it right. And then safe public text, so protected speech without exposing a voice, because this is voice data, which we found out. It's quite. challenging to work with. I don't know if you're doing it. And then we do it under the do no harm approach, of course, and the biggest thing is community ownership. We want to make sure that we do this in the cooperative if possible. You've mutually defined the guardrails, what's acceptable, what isn't. Trying to build local ownership and work with the government through the national universities and make sure it's culturally and ethically aligned to the community on that. Thanks for the lights. And then all of this under the data protection and oversight guided by the Office of Data Protection and Privacy, a new office at the UN that we're working with who have been absolutely fabulous to work with, but also at the same time it's difficult to move through because they're new as well. And then last but not least, we're talking about governance, assessments, and the other things that we're working on. talking about risk, data impact, human rights, ethical impact. So we're trying to keep that all in mind. That is also over the horizon, but coming up soon. And seven minutes is that? Forty minutes, but Was it really 14 minutes? A little bit more than seven minutes. Oh, I'm sorry. I could have gone faster in the beginning, but I didn't want to lose. I can speak very quickly if you want. Anyway, so we're looking for collaborations, partnerships, discussions. Over to you.
Barbora Bromová
Wonderful. Actually, over to Amy, who will be doing our second presentation on behalf of Clear Global. Over to her now.
Aimee Ansari
Okay. Hi, everybody. I mean, great presentation. Super interesting work you're doing. What can you see? I am going to talk about very, very similar things, but at a slightly larger scale. Thank you. so I'm Amy Ansari I am the chief executive of Clear Global we're going to talk about we've done a lot of work in humanitarian response and why is that not the right screen no I've got it open in two different places and so I just have to find the right one I have this wonderful gentleman over here yeah yeah yeah no no no I'm good I just have to find the right one I've got one open online and one open that should be the right one there we go okay As we've been talking about, language AI is increasingly used in the social impact and humanitarian sector to process information with tools like automatic speech recognition and machine translation. They have the potential to improve efficiency and accessibility, particularly when integrated into your internal workflows, as you were talking about. But sometimes these systems also reflect the inequalities in language representation. When they don't support the language, which is what UNDPKO is trying to get around in South Sudan right now, they risk systematically excluding entire communities. And at worst, the output is wrong or life threatening in nature. So I'm going to talk to you a little bit about the work that we do. Most languages in the world, as you can see on this chart, remain really poorly supported or entirely absent from any frontier model, from any large language model. this from some research done this year it represents how much text data there exists online across languages and that gives you an indication of how well the text -based AI like large language models are likely to perform you can see from this that languages spoken in wealthier countries dominate while others including those with tens or hundreds of millions of speakers don't have very much data and have limited resources and I was just reading a study today from Cambridge that talked about the Breton language which is spoken by about two hundred thousand people in French did you read this article too it's a fascinating it's a fascinating thing but it has over 50 models two hundred thousand people speak this language there are 50 models in it but a language like Nigerian Pigeon which has about 85 million speakers has 10 models developed And it's the same for languages like Seraki in northeast Nigeria. And that's talking about text models and data. The picture is even bleaker when you're talking about speech models, the sort of things that develop automatic speech recognition, and only a fraction of languages globally are meaningly supported by ASR. So this map here shows the most common primary languages spoken in Borno, Adamawa, and Yobe states in northeast Nigeria. Of the ten languages that we were able to find, only, to our knowledge, only Hausa has any recognition ASR, including a commercial API. But we don't really know how well that it proves. It performs in the wild in actual... practical practice with real -life people. We only know how it works in a laboratory environment. I'm going to talk a bit more about that. That means when you're implementing ASR -assisted workflows that rely on Hausa or English, you are risking excluding most, the vast, you can see, about 70 % of the population of East Nigeria. You're just not reaching them. No communication is going to them. So why? Well, there are a lot of different reasons why. Developing ASR in low -resource languages comes with a lot of challenges, many of which you've just talked about. But most, some languages, and I'm not sure if it's true of Juba Arabic, actually, but think in New Era like this, they're primarily oral. And so manual transcriptions is a really hard and complex... Transcription in order to develop the ASR model. So terminology and ways of expressing oneself vary widely between the speakers of those languages. There's a lot of code switching. So mixing two languages when you're speaking. We all do that. Any of us who have worked in a number of countries do that fairly regularly. But that's a really hard challenge when you're building digital language technologies. And as I said, lab metrics don't tell us very much about real world performance. So when you're when you're looking at a speech recognition model, you often hear that it has a low word error rate. I have a really hard time saying word error rates to me. Ours in that, I think. And that generally means that some researcher has tested it in a laboratory. But it doesn't really give you a sense of how well it's going. Of the real world. And at the center of this gap. is a lack of really high -quality voice data and text data, to be honest, in the relevant topics. You need hundreds of hours of recorded data done systematically, collected from across diverse speakers, including by age, by gender, by educational level, and by dialect, if you really want to make an ASR model that works well and widely in the field. For collaboration, across the language development pipeline, we want to build those kinds of partnerships and have cooperation with communities, governments, social impact organizations, and technologists from the beginning, from initial design through data collection to deployment. Just a little bit. Just a little bit on how we do this. We follow a very community -centered approach to data collection. We collaborate with sectors. and community members to really understand how the models are going to be used, how the data is going to be used, to understand linguistic complexities and the barriers to the models before we do any data collection at all. The quote up here is from, highlights one of the problems with tooling that we have. Just like keyboards don't exist in a lot of the languages that we're talking about. We also train community members, 100 ,000 linguists working in 300 languages. And our background is as linguists and translators, so we bring that linguistic expertise to the work that we do. I wanted to shout out to the interpreters and translators in the booths, but they're all gone now because we don't need them. But they're great people. Thank you. Thank you for your work. Yeah, so the tooling is one of the big challenges. So we train the community members as linguists how to do the recording, including an understanding of how and co -design with them the approaches to how you write down in primarily oral language, because there might not be a standard alphabet or a standard way of writing it. So you have to really think that through. In Canary, we saw one recording have 10 different equally accepted ways to write it down. Like they were all fine, but they could not agree amongst the 10 of them which one was the right way. So we use co -design approaches to manage the complexity to achieve a level of consistency that is required for the speech dialects. I have another nice example. I have another example from Congo, from Congolese, but I think I'll skip that in the interest of time, but I'm happy to talk about it. particularly since there's an Ebola response. So informed consent is a big pillar of the work that we do. It's a guiding principle for our platform, and we do this through workshops, community engagement, to make sure that people really understand how the recordings are going to be used and stored, what the risks are around recording their voices and in contributing, and how they can withdraw consent at any time. Absolutely critical for us. Ultimately, all of our data sets are published openly under non -commercial licenses so they can be used across the sector. Okay, this is two more slides. Is that okay? Do I have time? Okay, okay, okay. Okay. So this... This is just... Last year we... because we used to be Translators Without Borders, TWB. And we are, it's a platform that's integrated within our community and within the platform where we manage that community. And we are currently working on those 12 languages up there, which I'm not going to read out loud, in addition to expanding the work that we've already started collecting, where we've collected about 150 hours of voice data in Hausa Kanuri Shua Arabic, for northeast Nigeria. And these are some examples of prompts designed by community members in Kanuri and Kangali Swahili, which we then translated back into English. The prompts are coded by interaction and and disaster type 2. Yeah, that's it. Thank you very much, everybody. I'm happy to answer any questions. I talk really fast.
Barbora Bromová
All right. Thank you very much, Amy. I will wrap us up for the presentations with a very, hopefully very quick example of what that project might look like in terms of specifically addressing a need that a community has, how that might be put, at least my institution is not always very comfortable with, as a logic of working, and how that might actually work online. I'm going to try to give a very quick demo. Please cross your fingers that my internet connection holds for that. But I'm going to take you to Serbia, one of my favorite projects working on linguistic diversity in collaboration with UNDP, Serbia, and the government of Serbia, specifically the National... library, which is facing a very wicked problem. So as opposed to some of the other languages we've been speaking about, Serbian itself is not necessarily the lowest resource in the world. They have a fair amount of text. They're not doing so hot on speech data either, but they actually have quite a lot of historical sources. Serbian came down, their stories, their experiences, a lot of that material has survived, but it is of course on paper. It's not available for data sets or for data scientists to work on. And so we see that Serbian here is at the very bottom of that distribution here. It is also a language among a very complex South Slavic language family, and so there are closely related but distinct languages like Croatian, Montenegrin, and others. So I'm going to slightly complicate some of these exercises. As I said, this was a project developed in collaboration with the National Library of Serbia. The team at the Digitalna Biblioteka is on a mission to digitize 100% of their resources that they have. For years now, they've gotten to 2%, which is a truly commendable effort. But what we've been trying to do is help them leverage some of the scans that they already have. Because we've been finding that at least standard OCR, so that would be optical character recognition models, were having a very hard time with Serbian at large, partly because the language uses both Cyrillic and Latin alphabets interchangeably. And they were additionally having a particularly hard time with the historical scans. Within the Librarify project, we developed a workflow that prepares the image from the scans that the library has made, segments it to distinguish text from any images or headers that the page might have, enhance these segments, go through optical character recognition, do some post-OCR adjustment as well, evaluate, and then do it all again in hopes to recursively improve that process for that particular data type. We were quite successful with this, and in the initial phase digitized over 400 gigabytes and over 16,000 documents from the library's archives, at least those 2% initially digitized. However, we found that of this experience, this is an issue that many languages around the world are facing, and specifically many public institutions are facing. A lot of documents, a lot of data is trapped on paper legacy formats, and we wanted to develop a way to take this workflow, make it a little bit more implementable across different languages, which is how Loria came about. It has four stages, very much corresponding to the ones I already presented, but it is built for reuse and adaptation. It is an open tool published on GitHub for your download and adjustment in case you're interested in deploying it. And very much building on James' idea or core concept of human in the loop, it has been developed with archivists at the center. One thing that we really found very important throughout this process is that one needs to be really fluent not only in the technology and the data, but also in the process. One needs to understand very well what the technology can and cannot do, but they need to understand also very well the institutions that they're working through, the National Library, their processes, their priorities in digitizing some of these. documents in order for all of this to work together and deployment throughout the library. This is a little bit of an overview of the team behind this. We had two machine learning experts, four full-stack development engineers, a design team and a product management team across the National Library of Serbia, the Mathematical Institute at the Academy of Sciences and UNDP, both country office and the global team. Now, I do definitely thank you for your attention. However, I will try to show you Loria, which is the instance that we have spun out in Serbia that I will hopefully be able to connect into so that you can understand a little bit about how this works. So as I said, this is a little bit of a demo. This is a locally deployable solution. What I'm connecting into now is a server at our office in Serbia. I can log in with my details. And here I can see the interface that the archivist would have access to. my colleagues have kindly uploaded a couple samples from the National Library and specifically this is one of my favorites because it is a woman magazine from 1934 in Serbian and while I myself not very familiar with interesting ideas about fashion in particular to give you a little bit of an idea of what the process used to be back in the day without the automations that we were able to implement first we would start with the image enhancement so we could of course adjust the brightness manually here, adjust the sharpness with the sliders down here until the image is better than it was in the initial scan. What we can also do, thanks to Loria, is have this run automatically. I think this might be my connection problem, which I apologize for. But let me just talk you through it. Essentially, while you would be able to, the same way with photos, you would be used to this. Adjust the sharpness, adjust the brightness. the pages individually. We have an algorithm that does this for you. Then we also have a thresholding algorithm. If it would run, that would be perfect. But essentially that allows the picture to be clear enough for those additional segmenting steps to take place. The next one is layout identification. So that is the one that tries to understand the page, pick the text from among these blocks and essentially make it so that the OCR model only focuses on certain parts of the page. After that, there is the OCR round that's ideally specifically adjusted to Serbian or the language that you are using. And then we also have post -OCR correction that you would have seen. it's not loading for me at the moment but something that we've implemented in the latest update is actually a plug -in for frontier models as well which is something that further allows for better post -ocr correction if that is something that you are after or it allows prompting and batch processing across the entire subset of documents that you are navigating that is particularly interesting if you're trying to analyze for a particular characteristic of the text or directly process the data into a into a format different than plain text I wish I could show it to you but maybe it's better that I don't in the interest of time because we have a short discussion to get to and about the international coordination ongoing about this work so let me turn let me actually take this stop the screen share here and turn to my colleague Daphne from UNESCO who will tell us a little bit more about the initiative that actually unites all of these projects perhaps except for our colleagues at Peacekeeping that we can explore whether they want to join which is the Coalition for Linguistic Diversity in AI led by UNESCO so let me hand it over.
Dafna Feinholz
Thank you and I'll try also to be very brief I don't have any slides but I think it's yeah the idea is to look how we can collaborate together and I want to give you to show to you what we have been doing it's a very young project it's really in the process but I think it's interesting and it's capturing many of the things that are said here and I believe that we are all, I mean, this is kind of obvious, but it is about protecting the languages, but also what it is behind them, which is all the culture that it is behind. And I think this is something very important to remind, to remember that this is not only about the language, because this is, the language is just the way in which people express how they view the world, how they understand, and how do they see themselves, the individuals that become important to keep them. And that's why they are also heritage, and that's why it's so important in the work that UNESCO does, and also in the area of AI. So, as you know, we have a recommendation of ethics of AI, and part of it is the protection of diversity and language diversity in particular. Something that is also moving, something that is behind the project is the idea that linguistic and cultural inclusion in AI, in order to be meaningful, as it was already mentioned, together with the respective linguistic and cultural communities across the entire life cycle, because that's part of the issue, no? That we sometimes tend to include them very late in the process, so I think the example that we just heard two speakers before, it's really from the very beginning including them. And then throughout and continuously after the development and deployment of the tool. So, it requires cooperation with communities, capacity building and open, which is also, I think, very important. So, what we did is that we established an agreement with the Icelandic Ministry of Culture, Innovation and Higher Education. of the Icelandic Centre for Language and Technology. And this is how we established the Coalition for Linguistic Diversity in Artificial Intelligence. This coalition is a bit of a year ago. It started in June 2025. It was launched every year we have a global forum on ethics of artificial intelligence that we are all very welcome to attend. We bring all the different stakeholders doing areas of AI related to ethics. This year will be in Saudi Arabia from the 14th to the 17th of September. And it was launched in Bangkok last week. Today it comprises more than 30 experts working on community -led digital inclusion of languages from almost all the world regions. And of course it continues to actively grow. And so we definitely invite all of you to join. Another important characteristic, of the work of UNESCO is the all the diversity also in cultures in language, of course language cultures but also different perspectives of everything political but also ethical but also economical so this is really about a multi stakeholder perspective this initiative and that's I think one of the key elements so these coalitions bring together a very broad range of stakeholders it brings together governments it brings together academia communities, technical experts international organizations and the private sector so everybody is sitting together in the same group so basically this coalition is a space for these experts of different areas and different communities sitting together in the same room trying to find solutions of how to preserve these languages and all of them trying to figure it out together what will be the best way of preserving it for their own communities but instead of doing it in silos each of them with their own community This has proven to be very useful because there is a lot of sharing of experiences. There is something that some of them know that can be very useful to others. So this is really something that has been very, very, very useful. And most of the time they have seen that what works for one can work for the other. And also the idea of having this multi -stakeholder partnership is because also you have different scales of partners. So you have big companies, but you have also startups from these big companies without the need of hiring someone to develop their own software. So this is also filling a capacity or skill gap. Now, the idea is not only to have these discussions, but to document them. The idea is to be able to document all these good practices that already are there. So we are building a repository. of these good practices. The idea is to include these that have proven already to be successful. And the idea is also, so for the moment, like a page in which all the information is there, but in order to make sure that this information is systematized and is methodologically also coherent, so there is a questionnaire that is designed in order to identify what is the relevant information to be collected from the project. And this questionnaire is co -created between all of the stakeholders and UNESCO because each of the stakeholders are the only ones that know their project, so they will know what is relevant, who to ask. So that's why it's very important that this will be done by them. I'll give you more documents more examples but this repository can make this knowledge accessible because sometimes that's part of the problem that where do you get access to this knowledge it is a living resource because we want this to really be enriched all the time and again it's a place where others can learn but can also build on and so this is why now we are in the way of creating this what I said the questionnaire so the idea is to once that the repository is established and running we're planning to analyze the data identifying the existing trends in the community led approaches to linguistic and cultural dynamics and diversity in AI but also to identify the areas where the actors working in it need more support and more capacity building so these findings will help us to inform policy guidance on linguistic diversity and AI priority area evoked by many member states that now for example at the AI dialogue because it was also mentioned as one of the things that needs to be taken into account it was in fact one of the findings of the preliminary report of the panel so we want to foster discussions among actors but also working with community led linguistic leaders and also really turn all this into collaboration as the introduction of my presentation was and we are more than happy to work with this table
Barbora Bromová
Thank you very much, Daphne, for your presentation and for the work that you and your team have been doing on coordinating some of these projects and the exciting work ongoing around the globe. I wanted to have a couple of minutes for discussion, which we now won't because we will release you for your well -deserved evenings and afternoons. But perhaps if you do indulge me for just a minute more, I would like to actually ask Amy, because I'm conscious we have a lot of international organizations speaking and we are trying our best, serving both our own objectives when it comes to multilingualism and inclusion and trying to capitalize the ecosystems of builders that we are helping and working off of. But Amy is representing an organization that's a little bit looking outside. So if you, as one of the experts in the Coalition on Linguistic Diversity of AI, or I suppose one of your team's experts, if you look at all these efforts and international organizations in the space, what can we do for you? what would be the most impactful thing for us to help you with to focus on together so that we can make a real difference?
Aimee Ansari
Thanks, Barbara. And I think that a lot of what you are doing already is really great and really helpful. We've really enjoyed the UNDP, but also UNESCO and others. One of the biggest things is it costs a lot, right, because you want to be paying people. The volunteers are great. We work with a lot of volunteers. But you want to be paying people fair wages to collect the data. You want to make sure that you're not exploiting people and that you're not extracting language data, which you are then building a model off of and selling back to people. So I think. That's what. One thing we really struggle with is we know that there's this organization over there and this organization over there and that organization over there who are all trying to develop the technology, but they're not coming together. And if we worked together with all five of them, then we can build the data that is relevant for all of them at reasonable cost for each of them, right? Because this is $50 ,000, $100 ,000. Most grassroots organizations don't have that kind of money. But pulling five or six of them together, they could probably find $10 ,000 each. Trying to do that work, it's a lot of work. It's a lot of coordination work, and we really struggle to do that, and we don't have the kind of overview that you do. So that's one thing on a really grassroots level. I think the other thing is around governments building sovereign models. I think the global dialogue on AI governance is a really good start. I don't see a whole lot of norm setting and standard setting around how all communities and languages and cultures are represented in a sovereign model. And it's a hard thing to talk about in a UN space, but we'll take the government of South Sudan because I know that government. There was a civil war. I was in a civil war in South Sudan. The Dinka and the Newer were fighting against each other. Now, if you have a government that is composed only of Dinka, how likely are they to be able to build a sovereign model or want to build a sovereign model that reflects the Newer culture, the Newer language, that doesn't have any bias in it? like that's a hard ask so I think that one of the roles of the UN has to be around trying to make sure that when we're building global public infrastructure that we're building it in a way that is fair and equitable. that's what I
Barbora Bromová
That's what I those are some great points and wonderful material for my outcome report I'll be compiling from this session which I'm sure you all are eagerly awaiting thank you so much I saw a hand, was that a hand? no? okay, is there a question?
Audience
yeah it's like a combination of comments and questions so first I've got to tell myself I'm a translator the interpreters are the cool kids I'm kind of just a translator okay thanks I have sat in that booth sort of on and off all day taking notes, and I can confidently say that this is the coolest session of the week. Okay? By far. By far. Okay? Make sure we get that on recording. It is. It's love. I helped organize this week with the rest of the team, so I had to come out and listen. That's how cool this session is for me. But I also have to kind of out myself because I work in this space for the enemy, sort of. I work not for, I have to say with because I don't represent them, but I work with Meta on building these LLMs, but my team is I18N. We're languages, right? And what's interesting is what you said about the cost, right? So many of them are rooted in English. The same problem every time. It's that, you know, we. the models are rooted in things that we don't think about every day. Whereas my Vietnamese colleague will go, yeah, there's 80 plus pronouns. And if I use one, I know exactly, you'll know exactly who I'm talking about. I'm talking about my mother's, my paternal aunt's cousin or something like that. Right. And they have that and we don't. But you have to have that cultural knowledge. And to get it, you have to have someone who then, like you said, knows how to read it, knows how to write it, knows how to differentiate it, knows how to do it. But so there's like a level, even at the top of the languages that are in green on your slide there. Like my family's Catalan. So what you were saying about, you know, languages not being preserved or actively fought against, like I get it. Right. So it's just the question is, how do you get the people? And then it's the cost. How do you tackle that? Because it's a two cost problem. Right. It becomes either you're fighting to get a model in that language or you're fighting to get people. Who know the language well enough, which is a privileged question. Right. It's a are you. understanding of the language enough, or then it becomes a do you know the language well enough to be able to communicate with a vast number of people, which is impossible, because then you have fading and Romanized Telugu or Romanized Tamil or something like that, right? There's no canonical version that is totally correct every time. You add an extra A into a Hindi word that's written in Roman scripts, and someone will understand it over text, but how do you figure out what's right? I think I'm just asking a lot of questions, actually, but more to the point that I think how do you tackle the question of the two -cost problem, right? And does the private sector have to be involved? Does the private sector have to be involved in helping you get that data and provide it?
Speaker
I'd love to take that question and answer that firstly I'm speaking on behalf of myself not any organization we are also facing a similar problem where the language that we are targeting is primarily oral and it has no written form so if you try and convert that to a written form because of the way that the language is a lingua franca and it's a creolic language it has body substrate and it has tastes and accents of Arabic it has some stuff in it as well the way to build it is with the community and if you build a standardized orthographic structure with the linguists who are experts in the language as well as who are experts in other languages and have that set as a standard perhaps as IEEE or UNESCO it could really help out in making sure that we don't lose the nuance and the context in when we are transcribing it into ASR because one of the problems that she had mentioned as well as ASR doesn't work for low resource languages because it has to be tuned for that particular language every single time. And translate from speech directly to speech. We are getting there. LLMs might not be the way but we have to start with the data we have to start with the people to build it with the people for the people and democratize access to knowledge and not just data. So it could be a solution. Thank you.
Barbora Bromová
A quick note on the private sector involvement. I am again not necessarily speaking on behalf of UNDP but in our work we are of course open to private sector partners however in approaching them they're relatively low interest in doing this specifically on some of the very low resource languages because of the cost that Amy has previously mentioned. There is also then a some consideration in terms of data ownership and data governance questions that are of course open I must say some of our private sector partners are very open minded about these things and we're grateful for that but not all of them by a long mile so we've been trying to try example or also platforms that allow decentralized data sharing and try to sort of get around it that way but it's an open challenge that we're constantly iterating on that being said if you represent the private sector organization who's interested in doing this please reach out we're very open to those conversations.
Aimee Ansari
I would just say on private sector that there's private sector there's like meta and open AI and anthropic and those guys and then there's Viamo Viamo you know so there's private sector and private sector and we work a lot with the Viamo's and the GSMA's and some of those organizations. Lilapa AI is another really interesting that we work with I think the distinction is a little different we have also worked with Meta to test out models, to test out ideas to see if some of their models are working well enough so that they're safe so it's not the enemy. Meta is not the enemy at all I think the question is more around what Barbara was talking about is data governance and how you build that to communities and they're not people people who speak Dinka are not spending millions of dollars online you know no offense but so they're just not commercially interesting.
Barbora Bromová
thank you very much for spending your last session slot with us we really really appreciate it it's been a pleasure, please do reach out after the session recording stopped thank you

Disclaimer: This is not an official session record. DiploAI generates these resources from audiovisual recordings, and they are presented as-is, including potential errors. Due to logistical challenges, such as discrepancies in audio/video or transcripts, names may be misspelled. We strive for accuracy to the best of our ability.