Teaching AI to Speak our Language: A Showcase of Global Efforts to Bridge the ‘Last Mile’ of AI Inclusion for Local Impact
This discussion centred on the challenge of developing AI-powered language tools for low-resource and underrepresented languages, with a particular focus on humanitarian and peacekeeping contexts. Three main projects were presented, alongside a broader international coordination initiative.
James D'Ercole described a UN Peacekeeping project based in Juba, South Sudan, aimed at building a real-time translation tool to address communication gaps among approximately 50,000 uniformed personnel from around 120 countries . The tool is designed to operate entirely offline in low-connectivity environments , with a human-in-the-loop approach that ensures AI assists rather than replaces human judgement . A key ambition is to build not just a language dataset but a cultural corpus that incorporates local meaning and context , and to preserve Juba Arabic as a living language that could eventually be accessible online .
Aimee Ansari of CLEAR Global highlighted the broader global disparity in language AI support, noting that most languages spoken in lower-income countries remain poorly represented in frontier models . She illustrated this with northeast Nigeria, where only Hausa has any meaningful automatic speech recognition support, potentially excluding around 70% of the population from AI-assisted communication . She emphasised the need for high-quality, diverse voice data collected systematically across speakers of different ages, genders, and dialects , and stressed the importance of informed consent and open, non-commercial data licensing .
Barbora Bromová presented the Loria tool, developed with the National Library of Serbia and UNDP, which automates the digitisation of historical documents through optical character recognition and post-processing . Dafna Feinholz introduced UNESCO's Coalition for Linguistic Diversity in AI, launched in June 2025, which brings together over 30 experts from governments, academia, communities, and the private sector to share good practices and build a repository of community-led approaches .
A closing discussion highlighted two persistent challenges: the high cost of ethical, community-centred data collection , and the difficulty of coordinating fragmented efforts across organisations . Ansari noted that while private sector involvement is valuable, commercially uninteresting language communities risk being left behind, underscoring the need for international organisations to help ensure equitable representation in global AI infrastructure .
Overall Purpose
- The discussion centres on the challenge of developing AI-powered language and translation tools for low-resource and underrepresented languages, particularly in humanitarian and peacekeeping contexts. Speakers from UN Peacekeeping (UNDPKO), CLEAR Global, UNDP, and UNESCO share their respective projects and explore how collaboration, community involvement, and ethical data governance can help bridge the global language technology gap.
- --
Major Discussion Points
- The critical communication gap in peacekeeping and humanitarian operations: UN peacekeeping missions involve approximately 50,000 uniformed personnel from around 120 troop- and police-contributing countries, yet communication with local communities and between peacekeepers themselves remains a persistent, long-standing problem. A real-time, offline translation tool is being piloted in Juba, South Sudan, specifically designed to function in low-connectivity, austere environments, with a human-in-the-loop approach to ensure AI assists rather than replaces human judgement. - The severe underrepresentation of low-resource languages in AI models: Most languages spoken by large populations in the Global South remain poorly supported or entirely absent from frontier AI models, with text and speech data heavily skewed towards wealthier, English-dominant countries. For example, in northeast Nigeria, of ten commonly spoken languages, only Hausa has any meaningful automatic speech recognition (ASR) support, risking the exclusion of approximately 70% of the population from AI-assisted communication workflows. Challenges include the primarily oral nature of many languages, code-switching, lack of standardised orthographies, and the fact that laboratory performance metrics do not reflect real-world accuracy. - The necessity of community-centred, culturally sensitive data collection: Multiple speakers emphasised that building effective language AI requires deep community involvement from the very beginning of the design process, not as an afterthought. This includes co-designing recording approaches for primarily oral languages, obtaining informed consent, protecting voice data, and ensuring that cultural context - not just linguistic data - is embedded in the models. The UNDPKO project in South Sudan, for instance, is working to build a cultural corpus with national colleagues to make the language model less Western-centric. - Data governance, ownership, and the risk of exploitation: A recurring concern across presentations was ensuring that communities whose language data is collected retain meaningful ownership and are not exploited. Amy Ansari highlighted the tension between the high cost of ethical data collection - requiring fair wages and proper consent - and the limited resources of grassroots organisations, suggesting that pooling resources across multiple organisations could make data collection more financially viable. The question of private sector involvement was also raised, with acknowledgement that while some partners are open-minded about data governance, commercial interest in very low-resource languages remains limited because speakers are not large consumers in digital markets. - International coordination and the Coalition for Linguistic Diversity in AI: UNESCO, in partnership with Iceland, has established the Coalition for Linguistic Diversity in Artificial Intelligence, launched in June 2025, which brings together over 30 experts from governments, academia, communities, technical experts, international organisations, and the private sector. The coalition aims to document good practices in a shared repository, identify gaps in capacity building, and inform policy guidance on linguistic diversity in AI - moving away from siloed efforts towards a collaborative, multi-stakeholder model. Dafna Feinholz stressed that linguistic preservation is inseparable from cultural preservation and that communities must be included across the entire AI development lifecycle. ---
Overall Tone
- The overall tone of the discussion is collaborative, earnest, and solutions-oriented, with an undercurrent of urgency. Presenters speak with genuine passion for their work and a shared commitment to equity and inclusion in AI development. The tone is largely optimistic - particularly when speakers describe community engagement successes and the growing coalition of partners - but is tempered by frank acknowledgements of significant structural challenges, including funding constraints, data governance complexities, and the limited commercial incentive for private sector actors to invest in low-resource languages. Towards the end, during the audience Q&A, the tone becomes slightly more candid and informal, with speakers adding a grounded, practical dimension to what had been a more formal set of presentations. The closing remarks maintain a warm, collegial tone and offer a clear invitation to continued collaboration.
Expanded Summary: AI-Powered Language Tools for Low-Resource Languages in Humanitarian and Peacekeeping Contexts
#
Overview and Framing
This discussion brought together practitioners from UN Peacekeeping (UNDPKO), Clear Global, UNDP, and UNESCO to examine the challenge of developing AI-powered language and translation tools for low-resource and underrepresented languages, particularly in humanitarian and peacekeeping contexts. The session was structured around three project presentations followed by an introduction to an international coordination initiative and a candid audience discussion. Bromová served a dual role throughout as both session moderator and presenter of the Serbian digitisation project. A unifying conceptual thread was introduced at the outset by James D'Ercole, drawing on his experience as a regional administrative officer in East Timor: the principle of focusing on the "last mile" . His argument was that if a system functions correctly at the furthest, most resource-constrained point from headquarters - whether in a mission, in New York, or in Geneva - then everything along the chain must necessarily be working . This framing set the philosophical tone for the entire session, grounding the technical work in operational reality rather than institutional convenience.
#
The Communication Gap in UN Peacekeeping
D'Ercole opened by describing the scale and persistence of the communication problem facing UN peacekeeping operations. Having spent a little over 20 years in peacekeeping, he characterised this as a systemic and enduring challenge. Approximately 50,000 or more uniformed personnel are deployed daily across 11 missions, drawn from approximately 120 troop- and police-contributing countries . Communication - both between peacekeepers and the communities they serve, and among peacekeepers themselves - is foundational to the peacekeeping mission, which depends on listening and understanding . Yet this has remained a persistent, systemic problem throughout D'Ercole's career, documented repeatedly in field reports, C-34 submissions from governments, military and police advisory committee visits, and human rights reports . Interpreters and translators have never been available in sufficient numbers, even in better-resourced periods , and the diversity of peacekeepers' own national languages compounds the challenge further .
The project being piloted in Juba, South Sudan, is specifically designed to address this gap through real-time translation . Crucially, it is built around a human-centred oversight approach: AI is conceived as assisting the process, with a person always present to validate outputs, never replacing human judgement or eliminating the roles of language assistants and interpreters . D'Ercole argued that this approach would, if anything, amplify the reach and influence of existing language professionals rather than diminish them . A central technical requirement is that the tool must function completely offline, given the low or absent connectivity in the environments where peacekeeping missions operate . An edge device - described as a mini-server capable of running on a battery pack - is currently on loan and being tested as the hardware platform for this capability .
The translation model at the heart of the project was developed with contributions from a Harvard-linked effort and an NYU capstone project involving graduate students, reflecting the collaborative academic partnerships that have shaped the initiative from its early stages. D'Ercole described a specific iterative improvement cycle: offline testing feeds into dataset building and enhancement, which involves UN Volunteers working online, whose contributions flow into the translation model, which is then deployed on the edge device. The proof-of-concept phase operates in what D'Ercole called "shadow mode," in which the device sits and listens to an interpreter working, and the interpreter then reviews the device's outputs - validating them, correcting them, and in doing so helping to build the guardrails for the system. This validation loop then feeds back into the next improvement cycle, creating a continuously refined model grounded in real operational experience.
#
Building a Cultural Corpus, Not Just a Language Dataset
A distinctive ambition of the Juba project is to build what D'Ercole described as a cultural corpus rather than merely a language dataset . The project has engaged an expert in integrating culture into AI language models specifically to make the model less Western-centric and more contextually appropriate for the Juba Arabic-speaking community . This involves validating data collected by online UN Volunteers, then engaging national colleagues within the mission to further validate and enrich that data with cultural context . Language assistants have already contributed voluntarily on their own time, motivated by genuine belief in the project and in the importance of digitising Juba Arabic .
D'Ercole identified the preservation dimension of this work as the aspect he values most , describing the goal as creating a "living language" - preserving not only the language itself but the culture embedded within it . He further noted that once a language is digitised and made available as open-source data, commercial partners could eventually develop further tools enabling communities to access the internet in their own language . D'Ercole articulated three core benefits of the project: better communications, deeper trust, and a living language. The project is being built through collaborative partnerships with local institutions, international universities, civil society organisations, and NGOs, with a proposed language cooperative intended to ensure that the communities contributing data can eventually benefit from it . Data governance, AI safety, and voice data protection are being addressed under a "do no harm" approach, with community ownership and mutually defined guardrails treated as non-negotiable . This work is being conducted under the oversight of the UN Office of Data Protection and Privacy, a new office that D'Ercole described as both valuable and, given its novelty, sometimes challenging to navigate .
#
The Global Disparity in Language AI Coverage
Aimee Ansari of Clear Global situated the Juba project within a much broader global pattern of inequality in AI language coverage. Drawing on research conducted in 2025, she described how the amount of text data available online across languages - a key determinant of how well large language models perform - is heavily skewed towards languages spoken in wealthier countries, while languages with tens or hundreds of millions of speakers in the Global South remain poorly supported or entirely absent from frontier models . She illustrated this disparity with a striking comparison: Breton, spoken by approximately 200,000 people in France, has over 50 AI models, while Nigerian Pidgin, with approximately 85 million speakers, has only 10 . She also cited Seraki in northeast Nigeria as a further example of a language with negligible AI coverage. This gap is not driven by communicative need or speaker population, but by economic and geopolitical power.
The situation is even more acute for speech models. Ansari noted that only a fraction of languages globally are meaningfully supported by automatic speech recognition (ASR) technology . In northeast Nigeria, of the ten most commonly spoken primary languages in Borno, Adamawa, and Yobe states, only Hausa has any ASR support, including a commercial API . Even for Hausa, performance has only been assessed in laboratory conditions, and real-world field performance remains largely unknown . The practical consequence is that ASR-assisted workflows relying on Hausa or English risk excluding approximately 70% of the population of northeast Nigeria from AI-assisted communication entirely .
#
Challenges Specific to Low-Resource and Oral Languages
Ansari identified several interconnected challenges that make developing ASR for low-resource languages particularly difficult. Many such languages are primarily oral, meaning that manual transcription - a prerequisite for building ASR models - is an extremely complex undertaking . Terminology and ways of expressing oneself vary widely among speakers, and code-switching between two languages is common, creating further complexity for digital language technologies . A critical methodological point she raised is that laboratory-based word error rate metrics do not reliably predict real-world performance: a low word error rate typically means a researcher has tested the model in a controlled setting, which tells us little about how it will function with real speakers in real environments . At the centre of this gap is a shortage of high-quality, diverse voice data: building a reliable ASR model requires hundreds of hours of recorded data collected systematically from speakers of different ages, genders, educational levels, and dialects .
Clear Global's response to these challenges is a community-centred approach to data collection. The organisation collaborates with community members from the very beginning of the design process - before any data collection begins - to understand how models will be used, to grasp linguistic complexities, and to identify barriers . This includes training community members as linguists and co-designing approaches to writing down primarily oral languages, since there may be no standard alphabet or agreed written form . Ansari gave the example of a recording in Canary where ten different, equally accepted ways of writing the same word existed and contributors could not agree on a single correct version . Co-design approaches are used to manage this complexity and achieve sufficient consistency for model development . Informed consent is a foundational principle: community workshops are held to ensure speakers understand how their recordings will be used and stored, what the risks are, and how they can withdraw consent at any time . All datasets are published openly under non-commercial licences so they can be used across the sector .
#
Digitising Historical Documents: The Loria Project in Serbia
Barbora Bromová presented a complementary project addressing a different dimension of the language data gap: the digitisation of historical documents trapped in paper formats. Working with the National Library of Serbia and UNDP, the team developed a workflow - within the Librarify project - to process scans from the library's archives using optical character recognition, layout identification, image enhancement, and post-OCR correction . Serbian presents particular challenges: despite having a significant body of historical written sources, it sits at the bottom of language data distributions , partly because the language uses both Cyrillic and Latin alphabets interchangeably, which confounds standard OCR models . The initial phase of the project digitised over 400 gigabytes and over 16,000 documents .
From this experience, the team developed Loria, a more generalisable, open-source tool published on GitHub for reuse and adaptation across different languages and institutions . Loria follows the same four-stage workflow - image enhancement, layout identification, OCR, and post-OCR correction - but is built for reuse and includes a plug-in for frontier models to enable further post-OCR correction and batch processing . Critically, it was developed with archivists at the centre of the human oversight process , reflecting the broader session consensus that domain expertise and institutional knowledge are as important as technical capability . The project involved a multidisciplinary team spanning machine learning experts, full-stack developers, designers, product managers, the National Library of Serbia, the Mathematical Institute at the Academy of Sciences, and UNDP .
#
UNESCO's Coalition for Linguistic Diversity in AI
Dafna Feinholz of UNESCO introduced the international coordination initiative that seeks to connect these various efforts: the Coalition for Linguistic Diversity in Artificial Intelligence, established in partnership with Iceland's Ministry of Culture, Innovation and Higher Education and the Icelandic Centre for Language and Technology . Feinholz described the coalition as having started approximately a year prior and as having been formally launched at UNESCO's Global Forum on Ethics of Artificial Intelligence in Bangkok "last week," with the next such forum scheduled in Saudi Arabia from 14 to 17 September . The coalition already comprises more than 30 experts working on community-led digital inclusion of languages from almost all world regions . Feinholz grounded the coalition's work in a philosophical argument: language is not merely a communication tool but the medium through which people express how they view the world, understand themselves, and transmit their heritage . Preserving language is therefore inseparable from preserving culture, and this principle is embedded in UNESCO's Recommendation on the Ethics of AI, which includes the protection of linguistic diversity .
The coalition takes a deliberately multi-stakeholder approach, bringing together governments, academia, communities, technical experts, international organisations, and the private sector in the same space . Feinholz emphasised that this collaborative model has proven valuable because knowledge that works for one community often works for others, and because having large companies and small startups in the same group allows smaller actors to benefit from capabilities they could not develop independently . A key output of the coalition is a repository of good practices, built around a co-created questionnaire designed in collaboration between all stakeholders and UNESCO - on the basis that each stakeholder knows their own project best - to systematise and make accessible knowledge about community-led approaches to linguistic and cultural diversity in AI . This repository is conceived as a living resource that will inform policy guidance on linguistic diversity in AI - an area identified as a priority by many member states . Feinholz stressed that meaningful linguistic and cultural inclusion requires communities to be involved across the entire life cycle of AI development, from initial design through deployment, rather than being brought in at late stages .
#
Closing Discussion: Funding, Coordination, and the Private Sector
The session's closing discussion, prompted by Bromová's question about what international organisations could most usefully do to support organisations like Clear Global, produced some of the most candid and practically significant exchanges of the session . Ansari identified two principal areas where international organisations could make a difference. The first is financial: ethical, community-centred data collection is expensive, requiring fair wages for contributors rather than reliance on volunteers, and most grassroots organisations typically cannot afford the $50,000-$100,000 cost of building a language model . However, if five or six organisations working on related languages could be brought together, each contributing approximately $10,000, the collective cost becomes manageable . International organisations, with their systemic overview of who is working on what, are uniquely positioned to perform this convening function - a capacity that organisations like Clear Global lack .
The second area Ansari identified was norm-setting around sovereign AI models. She argued that while the global dialogue on AI governance is a good start, there is insufficient norm-setting around how all communities and languages are represented in nationally built AI models . She used South Sudan as a pointed illustration: in a government composed predominantly of one ethnic group following a civil war, the likelihood of building a sovereign model that fairly represents the language and culture of other groups is low . She argued that one of the UN's roles must be to ensure that global public infrastructure is built in a way that is fair and equitable .
Bromová, noting that she was not speaking on behalf of UNDP, observed that while some private sector partners are genuinely open-minded about data governance, many are not, and the challenge of data ownership remains unresolved . She also mentioned that her team has been exploring decentralised data-sharing platforms as one way of navigating these ownership challenges.
An audience member who works with Meta on building large language models introduced themselves with self-deprecating humour as working "for the enemy, sort of" , before reframing the challenge as a shared technical and cost problem. They noted that even within large technology companies, teams focused on internationalisation grapple with the same issues of cultural nuance, code-switching, and non-canonical written forms . They posed what they called the "two-cost problem": the dual difficulty of either fighting to get a model built in a given language or fighting to find people who know the language well enough to contribute meaningfully . Ansari responded by explicitly stating that "Meta is not the enemy at all" and describing Clear Global's collaboration with Meta to test model safety , while simultaneously maintaining that speakers of languages like Dinka are "not commercially interesting" to large companies because they are not significant consumers in digital markets . She also noted that Clear Global works extensively with organisations such as Viamo, the GSMA, and Lilapa AI - smaller and mid-sized private sector actors who occupy a different position in the ecosystem from the large frontier model companies.
A further contribution from the floor, from a separate unnamed participant, proposed that building standardised orthographic structures for primarily oral languages - potentially endorsed by bodies such as IEEE or UNESCO - could help preserve nuance and context in transcription and ASR development . This proposal sat in productive tension with Ansari's earlier empirical observation that even community members themselves sometimes cannot agree on a single written standard, suggesting that top-down standardisation may be neither straightforwardly achievable nor always appropriate .
#
Unresolved Challenges and Forward Agenda
The session concluded with a shared recognition that, despite strong conceptual consensus across all speakers on the principles of community-centred design, human oversight of AI outputs, cultural preservation, and ethical data governance, significant structural challenges remain unresolved. These include the sustainable funding of language model development for commercially unattractive languages , the governance of voice data and the risk of extractivism , the political dimensions of sovereign AI model building in divided societies , and the practical difficulty of coordinating fragmented efforts across the ecosystem . The UNESCO Coalition for Linguistic Diversity in AI, the Loria open-source tool, and the proposed language cooperative in Juba all represent concrete steps towards addressing these challenges, but all speakers acknowledged that the work is at an early stage and that the barriers ahead are as much political and economic as they are technical . The session closed with an open invitation for collaboration, with Bromová encouraging any interested private sector organisations to make contact and all speakers expressing enthusiasm for continued partnership across the sector .
Communication barriers between peacekeepers and local communities are a persistent, systemic problem across 11 UN missions involving 50,000+ uniformed personnel from 120 countries - Last mile communication problem (James D'Ercole)
Arg. 1James D'Ercole highlights that communication has been a persistent challenge in UN peacekeeping for over 20 years, with field reports consistently identifying it as a problem. With approximately 50,000 uniformed personnel from around 120 troop and police contributing countries deployed across 11 missions, the scale of the communication gap is significant. The diversity of peacekeepers' home countries means that inter-peacekeeper communication is also a major challenge, compounded by the fact that many come from low-resource language backgrounds.
D'Ercole noted that there are about 50,000 or more uniformed personnel in 11 missions, representing approximately 120 troop and police contributing countries . He stated that since being in peacekeeping for over 20 years, communication has always been a problem, with countless field reports from governments, military and police advisors, and human rights bodies consistently identifying it . He also pointed out that peacekeepers themselves come from low-resource language backgrounds, making inter-peacekeeper communication a major challenge .
on: Structural inequality in AI language coverage systematically excludes low-resource language communities
A real-time translation tool is being developed specifically for UN peacekeeping in Juba, South Sudan, designed to work completely offline in austere, low-connectivity environments - Offline translation tool for peacekeeping (James D'Ercole)
Arg. 2The project focuses on developing a real-time translation tool to close the communication gap between peacekeepers and local communities in Juba, South Sudan. A critical design requirement is that the tool must function completely offline, as the environments where peacekeepers operate have little to no connectivity. An edge device functioning as a mini-server with battery pack capability is being used to enable this offline functionality.
D'Ercole explained that the solution focuses on real-time translation as a tool to help close the communication gap , and that the device must work completely offline in austere environments with no signal . He described an edge device on loan that functions like a mini-server, capable of operating remotely without any connectivity as long as it has a battery pack .
The tool is built with a human-in-the-loop approach, where AI assists rather than replaces human judgment, and language assistants validate outputs to build guardrails - Human-in-the-loop design principle (James D'Ercole)
Arg. 3D'Ercole emphasises that the AI translation tool is designed to assist human interpreters and language assistants rather than replace them, with a person always present to validate outputs. The human-in-the-loop principle is central to the project's AI safety approach, ensuring that human judgment is never removed from the process. Rather than eliminating jobs, the tool is intended to amplify the reach and influence of language assistants.
D'Ercole stated that the human is always in the loop, AI assists the process, a person will always validate, and it never replaces human judgment or removes jobs, but instead amplifies and elevates language assistants . He described the proof of concept as a shadow mode where the tool listens to an interpreter and the interpreter then validates outputs to help build guardrails .
An edge device functioning as a mini-server with battery pack capability is being tested in shadow mode, listening to interpreters and learning from their corrections to improve iteratively - Edge device proof of concept (James D'Ercole)
Arg. 4The project is using a physical edge device on loan that acts as a portable mini-server, enabling offline translation capabilities in remote field environments. The proof of concept involves deploying this device in shadow mode, where it observes and listens to human interpreters working in real situations. The interpreter then reviews the device's outputs, providing yes/no feedback to help build guardrails and iteratively improve the model.
D'Ercole described a sample edge device on loan that functions like a mini-server, operable remotely without connectivity as long as it has a battery pack . He outlined the proof of concept plan to deploy it in shadow mode, where it listens to an interpreter and the interpreter validates outputs to build guardrails through an improvement cycle .
Cultural sensitivity must be embedded into the tool, ensuring local meaning, register, and context are incorporated to make the model less Western-centric and more specific to the Juba context - Cultural corpus over language dataset (James D'Ercole)
Arg. 5D'Ercole argues that the project is building a cultural corpus rather than merely a language dataset, recognising that language cannot be separated from its cultural context. The tool must incorporate local meaning, register, and context specific to Juba Arabic to avoid producing outputs that are culturally inappropriate or misleading. National colleagues in the mission are being engaged to validate existing data and collect additional data to make the corpus more robust and culturally grounded.
D'Ercole stated that the project is building a cultural corpus, not just a language dataset, and that national colleagues in mission are being used to validate data and collect more to add cultural depth . He noted that an expert has been engaged to integrate culture into the AI language models in order to make the model less Western-centric and more specific to Juba .
Digitising low-resource languages through AI tools preserves not only the language but also the culture behind it, and may eventually enable communities to access the internet in their own language - Language digitisation as cultural preservation (James D'Ercole)
Arg. 6D'Ercole argues that digitising languages such as Juba Arabic has a preservation benefit that extends beyond immediate communication needs, as preserving the language also preserves the culture it carries. Once a language is digitised and its data made open source, it could eventually attract commercial partners who might develop further tools, potentially enabling speakers to access the internet in their own language. This cultural preservation dimension is described as one of the most important aspects of the project for D'Ercole personally.
D'Ercole stated that by digitising languages, the project is preserving the language and, in doing so, preserving the culture, which he described as really important . He added that once digitised and open source, commercial partners could one day pick up the language, potentially enabling communities to surf the web in that language .
AI safety must be integrated by design from the beginning, with community ownership, mutually defined guardrails, and alignment with cultural and ethical norms of the community being served - AI safety by design and community ownership (James D'Ercole)
Arg. 7D'Ercole argues that AI safety cannot be an afterthought but must be embedded from the very beginning of the project's design. Community ownership is central to this approach, with guardrails and acceptable use boundaries being mutually defined with the community rather than imposed externally. The project also seeks to ensure cultural and ethical alignment with the community through engagement with national universities and government.
D'Ercole stated that the project wants to do AI safety by design, integrating it from the beginning and amplifying reach . He described the goal of mutually defining guardrails with the community, building local ownership, and working with the government through national universities to ensure cultural and ethical alignment .
Voice data presents particular challenges for data protection; protected speech without exposing individual voices is a critical concern, guided by the UN Office of Data Protection and Privacy - Voice data protection challenges (James D'Ercole)
Arg. 8D'Ercole highlights that working with voice data introduces specific and complex data protection challenges, particularly around protecting individual speakers' identities while still making the data useful. The project is guided by the UN's Office of Data Protection and Privacy, a new office that has been both helpful and challenging to work with given its own nascent status. The do-no-harm approach and community ownership principles are applied throughout to manage these risks.
D'Ercole noted that voice data is quite challenging to work with and that the project aims for safe public text and protected speech without exposing individual voices . He stated that all of this is done under the guidance of the Office of Data Protection and Privacy, a new UN office that has been fabulous but also difficult to move through because of its newness .
The project in Juba is built through collaborative partnerships with local institutions, international universities, CSOs, NGOs, and a proposed language cooperative to ensure data benefits flow back to contributing communities - Collaborative and cooperative partnership model (James D'Ercole)
Arg. 9D'Ercole describes a multi-layered partnership approach for the Juba project, involving local institutions, international universities, civil society organisations, and NGOs. A proposed language cooperative would allow the communities contributing language data to eventually benefit from that data, particularly if it is made open source with a paywall for commercial entities. National colleagues have been especially supportive, with language assistants voluntarily contributing their own time because they believe in the project.
D'Ercole described exploring partnerships with local institutions, international universities, CSOs, and NGOs, and the idea of forming a language cooperative so that contributing communities can benefit from the open-source data . He noted that language assistants contributed on their own time during the online volunteer phase because they believed in digitising Juba Arabic .
Most languages in the world remain poorly supported or entirely absent from frontier AI models, with wealthier-country languages dominating available text data - Language representation inequality (Aimee Ansari)
Arg. 1Ansari argues that the distribution of text data available online, which underpins the performance of large language models, is deeply skewed towards languages spoken in wealthier countries. Languages with tens or hundreds of millions of speakers in lower-income countries have very limited data and resources, meaning they are poorly served or entirely absent from frontier AI models. This structural inequality in data availability directly translates into inequality in AI performance and accessibility.
Ansari presented research showing how much text data exists online across languages, indicating that languages spoken in wealthier countries dominate while others, including those with hundreds of millions of speakers, have limited resources . She cited the example of Breton, spoken by about 200,000 people in France, which has over 50 AI models, compared to Nigerian Pidgin with about 85 million speakers, which has only 10 models .
Speech recognition models cover only a fraction of global languages; in northeast Nigeria, only Hausa has any ASR support, meaning roughly 70% of the population risks being excluded from AI-assisted communication - ASR coverage gap (Aimee Ansari)
Arg. 2Ansari argues that the gap in language coverage is even more severe for speech models than for text models, with only a fraction of global languages meaningfully supported by automatic speech recognition. In northeast Nigeria, of the ten most common primary languages spoken in Borno, Adamawa, and Yobe states, only Hausa has any ASR support, and even that is only tested in laboratory conditions. This means that ASR-assisted workflows relying on Hausa or English risk excluding approximately 70% of the population of northeast Nigeria.
Ansari presented a map showing the most common primary languages in northeast Nigeria, noting that of the ten languages identified, only Hausa has any ASR recognition including a commercial API . She stated that this means roughly 70% of the population of northeast Nigeria risks being excluded from AI-assisted communication workflows .
Many low-resource languages are primarily oral, making transcription for model development extremely complex, with wide variation in terminology, code-switching, and lack of standardised writing systems - Oral language complexity (Aimee Ansari)
Arg. 3Ansari highlights that many low-resource languages present unique challenges for AI model development because they are primarily oral, with no standardised written form. This makes the transcription process required to build ASR models extremely difficult, as terminology and expression vary widely between speakers, and code-switching between languages is common. The absence of a standard alphabet or writing system means that even the act of writing down a language requires co-design with communities to achieve sufficient consistency.
Ansari noted that some languages are primarily oral, making manual transcription for ASR development a really hard and complex process . She described wide variation in terminology, code-switching between languages, and the challenge this poses for building digital language technologies . She gave the example from Canary where one recording had ten different equally accepted ways to write it down, with contributors unable to agree on the correct version .
Even languages with significant speaker populations, such as Nigerian Pidgin with 85 million speakers, have far fewer AI models than smaller European languages like Breton - Disparity between speaker population and model availability (Aimee Ansari)
Arg. 4Ansari draws attention to the stark disparity between the number of speakers a language has and the number of AI models developed for it, arguing that this reflects structural inequalities rather than linguistic need. Breton, spoken by only about 200,000 people in France, has over 50 AI models, while Nigerian Pidgin, with approximately 85 million speakers, has only 10. This disparity illustrates how commercial and geopolitical factors, rather than speaker population size, drive investment in language AI.
Ansari cited a study she had read that day from Cambridge, noting that Breton, spoken by about 200,000 people, has over 50 AI models, while Nigerian Pidgin with about 85 million speakers has only 10 models . She noted the same pattern applies to languages like Seraki in northeast Nigeria .
Developing high-quality ASR models requires hundreds of hours of recorded data collected systematically from diverse speakers across age, gender, educational level, and dialect - Requirements for quality voice data (Aimee Ansari)
Arg. 5Ansari argues that the quality of ASR models depends fundamentally on the quality and diversity of the voice data used to train them, and that achieving this requires a systematic and resource-intensive data collection process. Hundreds of hours of recorded data must be gathered from speakers who are diverse across multiple dimensions, including age, gender, educational level, and dialect. Without this breadth of data, ASR models will not perform reliably in real-world field conditions.
Ansari stated that you need hundreds of hours of recorded data done systematically, collected from across diverse speakers including by age, gender, educational level, and dialect, if you really want to make an ASR model that works well and widely in the field .
Clear Global follows a community-centred approach, collaborating with community members from initial design through data collection to deployment, training over 100,000 linguists working in 300 languages - Community-centred data collection methodology (Aimee Ansari)
Arg. 6Ansari describes Clear Global's methodology as deeply community-centred, involving community members at every stage of the language technology development pipeline rather than treating them merely as data sources. The organisation trains community members as linguists, equipping them with the skills to conduct recordings and navigate the complexities of writing down primarily oral languages. With 100,000 linguists working across 300 languages, Clear Global brings significant linguistic expertise to this work.
Ansari described a community-centred approach to data collection, collaborating with community members to understand how models will be used and to understand linguistic complexities before any data collection begins . She noted that Clear Global trains over 100,000 linguists working in 300 languages, and that their background as linguists and translators informs the work . She also described co-design approaches to manage the complexity of writing down primarily oral languages .
Informed consent is a guiding principle for data collection, requiring community workshops to ensure speakers understand how their voice data will be used, stored, and how they can withdraw consent - Informed consent in voice data collection (Aimee Ansari)
Arg. 7Ansari argues that informed consent is a non-negotiable guiding principle for all of Clear Global's data collection work, particularly given the sensitive nature of voice data. Community workshops and engagement processes are used to ensure that contributors genuinely understand how their recordings will be used and stored, what the risks are, and how they can withdraw consent at any time. All datasets are ultimately published openly under non-commercial licences to ensure broad sector access.
Ansari described informed consent as a big pillar and guiding principle for their platform, implemented through workshops and community engagement to ensure people understand how recordings will be used and stored, what the risks are, and how they can withdraw consent at any time . She noted that all datasets are published openly under non-commercial licences so they can be used across the sector .
All data sets should be published openly under non-commercial licences so they can be used across the sector, while ensuring communities are not exploited through extraction of their language data - Open data under non-commercial licences (Aimee Ansari)
Arg. 8Ansari argues that language data collected from communities should be made openly available under non-commercial licences to maximise its benefit across the humanitarian and social impact sector. At the same time, she emphasises the ethical imperative of ensuring that communities are not exploited through the extraction of their language data, which is then used to build commercial models sold back to them. Fair wages for data contributors, rather than reliance solely on volunteers, are part of this ethical framework.
Ansari stated that all of Clear Global's datasets are published openly under non-commercial licences so they can be used across the sector . She also emphasised the importance of paying people fair wages to collect data and not exploiting people or extracting language data to build a model that is then sold back to communities .
Governments building sovereign AI models risk embedding political and ethnic biases, as illustrated by the example of South Sudan, where a government dominated by one ethnic group may not equitably represent other languages and cultures - Political bias risk in sovereign models (Aimee Ansari)
Arg. 9Ansari raises the concern that when governments build sovereign AI models, they risk embedding the political and ethnic biases of those in power, potentially excluding minority languages and cultures from the national digital infrastructure. She uses South Sudan as a concrete example, noting that a government composed primarily of Dinka people may have little incentive to build a sovereign model that fairly represents the Nuer language and culture. She argues that the UN has a role in ensuring that global public infrastructure is built in a fair and equitable manner.
Ansari used the example of South Sudan, where she noted that if a government is composed only of Dinka, it is unlikely to want to build a sovereign model that reflects Nuer culture and language without bias . She argued that one of the roles of the UN must be to ensure that when building global public infrastructure, it is built in a way that is fair and equitable .
International organisations can play a convening role by bringing together fragmented grassroots organisations to pool resources for data collection, since individual organisations cannot afford the $50,000–$100,000 cost alone but could collectively contribute $10,000 each - Convening role to pool resources (Aimee Ansari)
Arg. 10Ansari identifies a critical coordination gap in the language AI ecosystem, where multiple organisations are independently trying to develop similar technologies without collaborating, resulting in duplicated effort and prohibitive costs for each. She argues that international organisations are well-placed to play a convening role, bringing these fragmented actors together so they can pool resources and share the cost of data collection. Individual grassroots organisations typically cannot afford the $50,000–$100,000 required, but five or six of them contributing $10,000 each could make it feasible.
Ansari described the challenge of knowing that multiple organisations are all trying to develop similar technology but not coming together, and that if they worked together they could build relevant data at reasonable cost for each . She noted that data collection costs $50,000-$100,000, which most grassroots organisations cannot afford alone, but that five or six organisations contributing $10,000 each could manage it . She stated that this coordination work is very difficult and that Clear Global does not have the overview that international organisations do .
A distinction must be drawn between large technology companies and smaller private sector actors such as Viamo and Lilapa AI; collaboration with large companies like Meta can be appropriate for testing model safety without treating them as adversaries - Distinguishing private sector actors (Aimee Ansari)
Arg. 11Ansari argues that the private sector should not be treated as a monolithic entity, and that a meaningful distinction exists between large technology companies such as Meta, OpenAI, and Anthropic, and smaller mission-aligned private sector actors such as Viamo and Lilapa AI. She notes that Clear Global has worked with Meta to test models and assess their safety, and that this kind of collaboration is appropriate and valuable. The core concern is not about the private sector per se, but about data governance and ensuring that communities are not commercially exploited.
Ansari distinguished between large companies like Meta, OpenAI, and Anthropic, and smaller private sector actors like Viamo, GSMA, and Lilapa AI, noting that Clear Global works a lot with the latter . She noted that Clear Global has also worked with Meta to test models and assess whether they are working well enough to be safe, and that Meta is not the enemy . She clarified that the key concern is data governance and how it is built to communities, noting that speakers of languages like Dinka are not commercially interesting to large companies .
Serbian, despite having historical written sources, sits at the bottom of language data distributions and faces challenges with dual Cyrillic and Latin alphabets in OCR processing - Serbian digitisation challenge (Barbora Bromová)
Arg. 1Bromová argues that Serbian presents a distinctive digitisation challenge because, unlike many low-resource languages, it has a significant body of historical written material, yet this material is trapped on paper and unavailable for digital use. Serbian sits at the bottom of language data distributions despite this historical richness, partly because standard OCR models struggle with the language's use of both Cyrillic and Latin alphabets interchangeably. Historical scans present additional difficulties for OCR processing beyond those encountered with contemporary text.
Bromová noted that Serbian is not necessarily the lowest-resource language in the world and has quite a lot of historical sources, but that this material is on paper and not available for datasets or data scientists . She stated that Serbian sits at the very bottom of the data distribution and that standard OCR models had a very hard time with Serbian partly because the language uses both Cyrillic and Latin alphabets interchangeably, with historical scans presenting additional challenges .
The Librarify and Loria projects in Serbia demonstrate a replicable workflow for digitising historical documents trapped in paper formats, using OCR, layout identification, and post-OCR correction tailored to specific languages - Document digitisation workflow for linguistic heritage (Barbora Bromová)
Arg. 2Bromová describes the Librarify project, developed in collaboration with the National Library of Serbia, which created a workflow to digitise historical documents using image enhancement, segmentation, OCR, and post-OCR correction. The success of this project led to the development of Loria, a more generalisable and reusable open tool published on GitHub, designed to be adapted across different languages and institutions. Both projects place archivists at the centre of the human-in-the-loop process, recognising that institutional knowledge is as important as technical expertise.
Bromová described the Librarify project workflow, which prepares images from scans, segments them to distinguish text from images or headers, enhances segments, applies OCR, and performs post-OCR adjustment in a recursive improvement cycle . She noted that in the initial phase, over 400 gigabytes and over 16,000 documents from the library's archives were digitised . She described Loria as an open tool published on GitHub for reuse and adaptation, built with archivists at the centre of the human-in-the-loop process .
There is limited private sector interest in very low-resource languages because speakers of those languages are not commercially significant online audiences, making cost recovery difficult - Low commercial interest in low-resource languages (Barbora Bromová)
Arg. 3Bromová acknowledges that while UNDP is open to private sector partnerships in language digitisation work, there is relatively low interest from private sector actors in very low-resource languages because the speaker communities are not commercially significant online audiences. This makes it difficult for private sector partners to justify the investment required, creating a structural gap in funding for these languages. Some private sector partners are open-minded about data ownership and governance questions, but this is not universal.
Bromová noted that in approaching private sector partners, there is relatively low interest in doing work specifically on very low-resource languages because of the cost involved . She also noted that data ownership and data governance questions are open considerations, with some private sector partners being very open-minded but not all of them .
The Loria tool is published openly on GitHub for reuse and adaptation across different languages and institutions, with archivists placed at the centre of the human-in-the-loop process - Open reusable digitisation tool (Barbora Bromová)
Arg. 4Bromová argues that the Loria tool represents a scalable and reusable solution to the problem of documents and data trapped in paper legacy formats, which is a challenge faced by many languages and public institutions around the world. By publishing the tool openly on GitHub, UNDP enables other institutions to download, adapt, and deploy it for their own languages and document types. The design places archivists at the centre of the process, recognising that understanding institutional processes and priorities is as important as technical expertise.
Bromová described Loria as an open tool published on GitHub for download and adjustment, built for reuse and adaptation across different languages . She emphasised that it was developed with archivists at the centre, and that being fluent not only in the technology and data but also in the processes and priorities of the institutions is critical for successful deployment .
Linguistic and cultural inclusion in AI must involve communities across the entire life cycle of a tool, from initial design through development and deployment, not only at late stages - Full life cycle community inclusion (Dafna Feinholz)
Arg. 1Feinholz argues that meaningful linguistic and cultural inclusion in AI requires communities to be involved throughout the entire life cycle of a tool, from initial design through data collection, development, and deployment. A common failure mode is to include communities only at late stages of the process, which undermines genuine inclusion and ownership. This principle is embedded in UNESCO's approach to AI ethics and is reflected in the work of the Coalition for Linguistic Diversity in AI.
Feinholz stated that linguistic and cultural inclusion in AI, in order to be meaningful, must be done together with the respective linguistic and cultural communities across the entire life cycle, noting that communities are sometimes included very late in the process . She cited the example of speakers two presentations before as a model of including communities from the very beginning and throughout .
UNESCO established the Coalition for Linguistic Diversity in Artificial Intelligence with Iceland, bringing together over 30 experts from almost all world regions to share good practices and avoid working in silos - Coalition for Linguistic Diversity in AI (Dafna Feinholz)
Arg. 2Feinholz describes the Coalition for Linguistic Diversity in Artificial Intelligence, established by UNESCO in partnership with the Icelandic Ministry of Culture, Innovation and Higher Education and the Icelandic Centre for Language and Technology. The coalition, launched in June 2025, brings together over 30 experts from almost all world regions who are working on community-led digital inclusion of languages. A key motivation is to enable sharing of experiences and good practices across communities rather than each working in isolation.
Feinholz described the establishment of an agreement with the Icelandic Ministry of Culture, Innovation and Higher Education and the Icelandic Centre for Language and Technology to create the Coalition for Linguistic Diversity in Artificial Intelligence . She noted that the coalition was launched in Bangkok and currently comprises more than 30 experts working on community-led digital inclusion of languages from almost all world regions .
The coalition takes a multi-stakeholder approach, bringing together governments, academia, communities, technical experts, international organisations, and the private sector to find collaborative solutions for language preservation - Multi-stakeholder coalition model (Dafna Feinholz)
Arg. 3Feinholz argues that the Coalition for Linguistic Diversity in AI is distinctive in its genuinely multi-stakeholder composition, bringing together governments, academia, communities, technical experts, international organisations, and the private sector in the same space. This diversity of perspectives is seen as essential for finding solutions that are technically sound, culturally appropriate, and politically viable. The coalition has found that what works for one community can often work for another, making cross-community sharing particularly valuable.
Feinholz described the coalition as bringing together governments, academia, communities, technical experts, international organisations, and the private sector, all sitting together in the same group . She noted that this has proven very useful because there is a lot of sharing of experiences, and that what works for one community can often work for another .
A repository of good practices is being built with a co-created questionnaire to systematise and make accessible knowledge about community-led approaches to linguistic and cultural diversity in AI - Repository of good practices (Dafna Feinholz)
Arg. 4Feinholz describes the coalition's plan to build a repository of good practices in community-led linguistic and cultural inclusion in AI, making this knowledge accessible to practitioners who may not know where to find it. A co-created questionnaire, developed with all stakeholders, is being designed to systematise the collection of relevant information from projects in a methodologically coherent way. The repository is conceived as a living resource that will be continuously enriched and will eventually inform policy guidance on linguistic diversity in AI.
Feinholz described the plan to build a repository of good practices, noting that the idea is to include practices that have proven successful and to make this knowledge accessible . She explained that a questionnaire is being co-created between all stakeholders and UNESCO to identify what relevant information should be collected from each project, with stakeholders involved because they know their own projects best . She noted that findings will help inform policy guidance on linguistic diversity and AI, which has been identified as a priority area by many member states .
The two-cost problem involves both the expense of building models in low-resource languages and the difficulty of finding people with sufficient linguistic expertise to contribute, with no canonical written form in many cases - Two-cost problem in language AI development (Audience)
Arg. 1The audience member, who works with Meta on building large language models with a focus on internationalisation, identifies a dual cost problem in language AI development for low-resource languages. The first cost is the financial expense of building models in these languages, and the second is the challenge of finding people with sufficient linguistic expertise to contribute meaningfully, particularly when there is no canonical written form. This is compounded by the fact that even well-resourced languages face challenges around cultural nuance, such as the 80-plus pronouns in Vietnamese that carry precise relational meaning.
The audience member described the two-cost problem as either fighting to get a model in a language or fighting to get people who know the language well enough, noting that this is a privileged question . They gave the example of Vietnamese having 80-plus pronouns that carry precise relational meaning, illustrating the depth of cultural knowledge required . They also noted the challenge of Romanised scripts for languages like Telugu or Tamil, where there is no canonical version that is totally correct every time .
Building language models with and for communities, starting from the data and democratising access to knowledge rather than just data, is the foundational approach needed regardless of private sector involvement - Community-first model building approach (Speaker)
Arg. 1The speaker argues that the foundational approach to building language models for low-resource languages must start with the community and with the data, rather than with technology or commercial considerations. The goal should be to democratise access to knowledge, not merely to data, ensuring that the benefits of language AI flow back to the communities that contributed to it. Building a standardised orthographic structure with linguists and community members, potentially endorsed by bodies such as IEEE or UNESCO, is proposed as a way to preserve nuance and context in transcription.
The speaker described the approach of building a standardised orthographic structure with linguists who are experts in the language, potentially standardised through IEEE or UNESCO, to ensure nuance and context are not lost in transcription . They stated that the way to build it is with the community, and that the goal is to start with the data, start with the people, build it with the people for the people, and democratise access to knowledge and not just data .
Session Knowledge Graph
Speakers · Topics · Arguments · Relationships
All three speakers converge on the principle that AI must assist rather than replace human judgment. D'Ercole explicitly states that 'a person will always validate, should always be there, it never replaces human judgment' , and describes the proof of concept as a shadow mode where interpreters validate outputs to build guardrails . Bromová echoes this in the Loria project, noting it was 'developed with archivists at the centre' and that 'human in the loop' is a core concept . The Speaker similarly argues that the approach must be 'with the community' and 'for the people' , emphasising that democratising access to knowledge requires human agency throughout.
Human-in-the-loop design principle (James D'Ercole)
Document digitisation workflow for linguistic heritage (Barbora Bromová)
Community-first model building approach (Speaker)
All three speakers agree that communities must be involved from the very beginning of language AI development, not brought in at late stages. D'Ercole describes leaning heavily on national colleagues to ensure the tool is 'linguistically and culturally accepted' and that community ownership is central . Ansari describes collaborating with community members 'to really understand how the models are going to be used' before any data collection begins . Feinholz explicitly states that 'linguistic and cultural inclusion in AI, in order to be meaningful' must be done 'together with the respective linguistic and cultural communities across the entire life cycle', criticising the tendency to include communities 'very late in the process' .
Collaborative and cooperative partnership model (James D'Ercole)
Community-centred data collection methodology (Aimee Ansari)
Full life cycle community inclusion (Dafna Feinholz)
Both D'Ercole and Feinholz explicitly connect language preservation to cultural preservation, treating them as inseparable. D'Ercole states that 'by digitising these languages, we're preserving the language' and 'when we preserve the language, we preserve the culture', describing this as 'really important' . Feinholz reinforces this, arguing that 'this is not only about the language, because the language is just the way in which people express how they view the world' and that languages are 'also heritage' . Both frame this cultural dimension as a core motivation for their work.
Language digitisation as cultural preservation (James D'Ercole)
Full life cycle community inclusion (Dafna Feinholz)
All three speakers identify a structural gap in AI language coverage that systematically disadvantages low-resource language communities. D'Ercole notes that peacekeepers come from 'low-resource languages' and that communication has 'always been a problem' for over 20 years . Ansari presents research showing that 'languages spoken in wealthier countries dominate while others including those with tens or hundreds of millions of speakers don't have very much data' , and that in northeast Nigeria only Hausa has any ASR support, risking exclusion of roughly 70% of the population . Bromová notes that Serbian 'sits at the very bottom of that distribution' despite having historical written sources , illustrating that even languages with some resources face significant gaps.
Communication barriers between peacekeepers and local communities are a persistent, systemic problem across 11 UN missions involving 50,000+ uniformed personnel from 120 countries - Last mile communication problem (James D'Ercole)
Language representation inequality (Aimee Ansari)
Serbian digitisation challenge (Barbora Bromová)
Both D'Ercole and Ansari place data governance and community ownership at the centre of their ethical frameworks. D'Ercole describes wanting to 'mutually define the guardrails' with the community and ensure 'cultural and ethical alignment' , and notes that voice data is 'quite challenging to work with' from a protection standpoint . Ansari describes informed consent as 'a big pillar' and 'guiding principle', implemented through workshops ensuring people understand how recordings will be used and how they can withdraw consent . Both also emphasise that data should not be extracted from communities for commercial benefit without their knowledge or fair compensation .
AI safety by design and community ownership (James D'Ercole)
Voice data protection challenges (James D'Ercole)
Informed consent in voice data collection (Aimee Ansari)
Open data under non-commercial licences (Aimee Ansari)
All four speakers agree that no single organisation can address language AI gaps alone and that multi-stakeholder collaboration is essential. D'Ercole describes exploring partnerships with 'local institutions, international universities, CSOs, NGOs' and a proposed language cooperative . Ansari argues that international organisations can play a convening role to bring fragmented actors together to pool resources, since individual grassroots organisations cannot afford the $50,000-$100,000 cost alone . Bromová publishes Loria openly on GitHub for reuse and adaptation and notes openness to private sector partners . Feinholz describes the UNESCO coalition as bringing together 'governments, academia, communities, technical experts, international organisations, and the private sector' in the same space .
Collaborative and cooperative partnership model (James D'Ercole)
Convening role to pool resources (Aimee Ansari)
Open reusable digitisation tool (Barbora Bromová)
Multi-stakeholder coalition model (Dafna Feinholz)
Both D'Ercole and Ansari emphasise that building language AI tools requires going beyond mere language datasets to incorporate cultural context and address the complexity of primarily oral languages. D'Ercole states that the project is 'building a cultural corpus, not just a language dataset' and has engaged an expert to integrate culture into the AI language models to make them 'less Western-centric and more specific to Juba' . Ansari similarly notes that many low-resource languages are 'primarily oral' and that 'terminology and ways of expressing oneself vary widely between the speakers', making transcription 'a really hard and complex' process . Both recognise that standard approaches rooted in text-heavy, Western-centric models are inadequate for these contexts. All three practitioners share the view that successful language AI projects require deep engagement with local and national colleagues who possess both linguistic and cultural knowledge. D'Ercole notes that language assistants contributed on their own time 'because of actually believing in the project and believing in digitising the language of Juba Arabic' . Ansari describes training over 100,000 community members as linguists and using co-design approaches to manage the complexity of writing down primarily oral languages . Bromová emphasises that 'one needs to be really fluent not only in the technology and the data, but also in the process' and in understanding the institutions being worked through . All three treat local expertise as indispensable rather than supplementary. Both Ansari and Bromová acknowledge the structural financial challenge that low-resource languages face in attracting private sector investment, and both see international organisations as having a role in bridging this gap. Bromová notes that there is 'relatively low interest' from private sector actors in very low-resource languages 'because of the cost involved' and because speaker communities are not commercially significant online audiences . Ansari similarly notes that speakers of languages like Dinka 'are not spending millions of dollars online' and are therefore 'just not commercially interesting' to large companies . Both see the convening and coordination capacity of international organisations as a potential solution to this market failure. Both D'Ercole and Feinholz share the view that AI safety and ethical alignment must be embedded from the very beginning of a project rather than added retrospectively. D'Ercole states that the project wants to 'do AI safety by design, integrating it from the beginning' and that guardrails should be 'mutually defined' with the community . Feinholz similarly argues that inclusion must happen 'across the entire life cycle' and that communities are 'sometimes included very late in the process', which undermines genuine inclusion . Both frame this as a matter of principle rather than merely procedural compliance. Both Ansari and the audience Speaker converge on the view that building language AI for low-resource languages must start with the community and the data rather than with technology or commercial considerations. Ansari describes a community-centred approach where collaboration with community members precedes any data collection and emphasises the importance of paying people fair wages rather than exploiting them . The Speaker argues that 'the way to build it is with the community' and that the goal should be to 'start with the data, start with the people, build it with the people for the people, and democratise access to knowledge and not just data' . Both treat community agency as foundational. Both Ansari and Feinholz share a commitment to making knowledge and data openly accessible to the sector rather than keeping it siloed. Ansari states that 'all of our data sets are published openly under non-commercial licences so they can be used across the sector' . Feinholz describes building a repository of good practices that is 'a living resource' intended to make knowledge 'accessible because sometimes that's part of the problem, where do you get access to this knowledge' . Both see open sharing as essential to avoiding duplication of effort and enabling smaller organisations to benefit from collective learning.
Given the broader context of concerns about data extraction and commercial exploitation of community language data, it might be expected that speakers would take a uniformly critical stance towards large technology companies. Instead, both Ansari and Bromová express nuanced and pragmatic positions. Ansari explicitly states that 'Meta is not the enemy at all' and describes working with Meta to test models and assess whether they are 'working well enough so that they're safe' . Bromová notes that 'some of our private sector partners are very open-minded about these things and we're grateful for that' and invites private sector organisations to reach out . The consensus is that the concern is not about private sector involvement per se, but about data governance and ensuring communities are not exploited .
Both Ansari and the audience Speaker, who works with Meta on building large language models, converge on the inadequacy of laboratory metrics for assessing real-world performance of speech models. This is somewhat unexpected given that the audience member represents a large technology company that produces such models. Ansari states that 'lab metrics don't tell us very much about real world performance' and that a low word error rate 'generally means that some researcher has tested it in a laboratory' but 'doesn't really give you a sense of how well it's going in the real world' . The Speaker implicitly reinforces this by noting the complexity of cultural nuance, code-switching, and non-canonical written forms that laboratory conditions cannot capture . Both agree that real-world community testing is essential.
Both Ansari and the audience Speaker, coming from quite different institutional backgrounds, converge on the need for standardised approaches to writing down primarily oral languages. Ansari describes the challenge of one recording having 'ten different equally accepted ways to write it down' and the use of 'co-design approaches to manage the complexity to achieve a level of consistency' . The Speaker goes further, proposing that 'building a standardised orthographic structure with the linguists who are experts in the language' and having it 'set as a standard perhaps as IEEE or UNESCO' could help ensure nuance and context are not lost in transcription . This convergence between a humanitarian linguist and a technology industry professional on the value of formal standardisation is notable.
Both D'Ercole and Ansari, approaching from different angles, converge on the idea that digitising low-resource languages has transformative long-term potential for digital inclusion. D'Ercole, coming from a peacekeeping operational perspective, notes that once a language is digitised and open source, 'commercial partners can one day pick up that language and you may be able to actually surf the web in that language itself' . Ansari, coming from a humanitarian linguistics perspective, frames the entire problem around the fact that most languages 'remain really poorly supported or entirely absent from any frontier model' , implying that digitisation is the pathway to inclusion. The unexpected element is that a UN peacekeeping official and a humanitarian linguist both independently arrive at the same vision of internet access in one's own language as the ultimate goal.
The discussion reveals a remarkably high level of consensus across all speakers on the core principles and challenges of language AI development for low-resource languages. All speakers agree that: (1) structural inequality in AI language coverage systematically excludes low-resource language communities ; (2) community involvement must span the entire life cycle of AI tools, from design through deployment ; (3) the human-in-the-loop principle is non-negotiable ; (4) language preservation is inseparable from cultural preservation ; (5) data governance, informed consent, and community ownership are ethical imperatives ; and (6) multi-stakeholder collaboration is essential because no single organisation can address these challenges alone . There is also strong consensus on the inadequacy of purely technical or commercial approaches, with all speakers emphasising that the communities whose languages are being digitised must be treated as partners and owners rather than data sources. The unexpected areas of consensus include a pragmatic rather than adversarial stance towards the private sector, agreement on the insufficiency of laboratory metrics for real-world deployment, and convergence on the long-term vision of internet access in low-resource languages.
Ansari drew a careful distinction between large technology companies such as Meta, OpenAI, and Anthropic, and smaller mission-aligned private sector actors such as Viamo and Lilapa AI, noting that Clear Global has worked with Meta to test model safety and that 'Meta is not the enemy' . However, she also emphasised that speakers of languages like Dinka are 'not commercially interesting' to large companies , implying structural limits to their engagement. Bromová acknowledged that some private sector partners are open-minded about data governance but 'not all of them by a long mile' . The audience member, who works with Meta on building large language models, implicitly pushed back on any adversarial framing by identifying themselves as working 'for the enemy, sort of' , and framed the challenge as a shared technical and cost problem rather than a commercial exploitation issue . This created a subtle but real tension about whether large tech companies are partners, bystanders, or obstacles in this space.
Distinguishing private sector actors (Aimee Ansari)
Low commercial interest in low-resource languages (Barbora Bromová)
Two-cost problem in language AI development (Audience)
D'Ercole described language assistants contributing data 'on their own time' as volunteers during the online volunteer phase, framing this positively as evidence of community belief in the project . He also described working with UN Volunteers (UNVs) as an online volunteer model . Ansari, by contrast, explicitly argued that 'you want to be paying people fair wages to collect the data' and that relying on volunteers risks exploiting people and extracting language data that is then used to build commercial models . While both speakers value community involvement, they diverge on whether unpaid volunteer contributions are ethically acceptable or whether fair wages are a non-negotiable requirement.
Collaborative and cooperative partnership model (James D'Ercole)
Open data under non-commercial licences (Aimee Ansari)
Ansari raised a pointed concern about the risk of governments building sovereign AI models that embed political and ethnic biases, using South Sudan as a concrete example: if a government is composed only of Dinka, it is unlikely to build a model that fairly represents Nuer culture and language . She argued that the UN must ensure global public infrastructure is built fairly and equitably . Feinholz, while agreeing on the importance of community inclusion across the full life cycle , did not directly address the political bias risk in sovereign models and instead focused on the multi-stakeholder coalition model as the solution . This represents a divergence in how deeply each speaker is willing to engage with the political dimensions of AI governance, with Ansari naming the problem explicitly and Feinholz offering a more diplomatically cautious framing.
Political bias risk in sovereign models (Aimee Ansari)
Full life cycle community inclusion (Dafna Feinholz)
D'Ercole described the project's data as open source but with 'a paywall to companies and things like that' , suggesting a model where commercial actors pay to access data while the broader sector benefits freely. He also expressed hope that commercial partners might one day develop further tools enabling communities to access the internet in their own language . Ansari, by contrast, stated that all of Clear Global's datasets are published 'openly under non-commercial licences' , and explicitly warned against 'extracting language data, which you are then building a model off of and selling back to people' . These represent meaningfully different positions on how data should be governed and monetised, with D'Ercole more open to commercial engagement and Ansari more cautious about it.
Language digitisation as cultural preservation (James D'Ercole)
Open data under non-commercial licences (Aimee Ansari)
This disagreement was unexpected because the session was broadly collaborative and focused on shared challenges. The audience member, who works with Meta on building large language models, introduced themselves as working 'for the enemy, sort of' , signalling awareness of a potential adversarial framing. They then reframed the challenge as a shared technical and cost problem, noting that even within large tech companies, teams focused on internationalisation are grappling with the same cultural complexity issues . Ansari responded by explicitly stating 'Meta is not the enemy at all' and noting that Clear Global has worked with Meta to test models for safety . However, she simultaneously maintained that speakers of languages like Dinka are 'not commercially interesting' to large companies , implying a structural limit to their engagement that the audience member's framing did not fully address. The unexpected element is that what appeared to be a potential confrontation between civil society and big tech actually revealed a more nuanced internal tension within the large tech sector itself, with the audience member implicitly arguing that the problem is systemic rather than a matter of corporate bad faith.
This disagreement was unexpected because both speakers are working on similar problems and might have been expected to align fully on methodology. Ansari explicitly argued that laboratory metrics such as word error rates do not tell us much about real-world performance , and that even Hausa, the only language with ASR support in northeast Nigeria, has only been tested in laboratory conditions . D'Ercole, by contrast, described a proof-of-concept approach using shadow mode - where the device listens to a human interpreter in real situations and the interpreter validates outputs - which is closer to real-world testing. However, D'Ercole's model still relies on a controlled proof-of-concept phase before full deployment, and the edge device is described as 'on loan' and in early testing . The unexpected tension is that Ansari's critique of laboratory-only testing implicitly applies to the early-stage proof-of-concept approach that D'Ercole is still working through, even though D'Ercole's shadow mode approach is more field-oriented than pure laboratory testing.
This disagreement was unexpected because it emerged from the floor discussion rather than the prepared presentations, and it touched on a genuinely unresolved technical and governance question. The speaker from the floor argued that building a standardised orthographic structure with linguists, potentially endorsed by IEEE or UNESCO, could help preserve nuance and context in transcription for primarily oral languages . Ansari, however, had earlier described the difficulty of achieving such standardisation in practice, giving the example from Canary where one recording had ten different equally accepted ways to write it down and contributors could not agree on the correct version . She described using co-design approaches to manage this complexity , but did not suggest that external standardisation bodies like IEEE or UNESCO could resolve it. The unexpected element is that the speaker's proposed solution - external standardisation - appears to conflict with Ansari's empirical experience that even community members themselves cannot agree on a single standard, suggesting that top-down standardisation may be neither achievable nor appropriate.
The discussion was characterised by a high degree of surface-level consensus around shared goals — community-centred design, human-in-the-loop AI, cultural preservation, and the need to address language gaps in AI systems — but revealed meaningful divergences in approach, governance philosophy, and political framing when examined closely. Key areas of genuine disagreement included: the role of large private sector technology companies (adversary, partner, or structurally limited bystander); whether volunteer or paid data collection is ethically appropriate; how open-source data should be governed and whether commercial paywalls are acceptable; and how deeply international organisations should engage with the political dimensions of sovereign AI model building, including ethnic and political bias risks. The most unexpected tensions arose around the adequacy of laboratory testing versus real-world deployment, the feasibility of external orthographic standardisation for oral languages, and the implicit challenge posed by the audience member working with Meta to any adversarial framing of large tech companies.
All four speakers agreed that communities must be placed at the centre of language AI development and that human oversight is essential. D'Ercole emphasised that 'the human in the loop' means AI assists rather than replaces human judgment , and that community ownership and mutually defined guardrails are critical . Ansari described a community-centred approach involving community members from initial design through deployment . Bromová noted that Loria was built with archivists at the centre . Feinholz argued that meaningful inclusion requires communities across the entire life cycle . However, they differed on implementation: D'Ercole's model relies partly on volunteers , Ansari insists on paid contributors , Bromová focuses on institutional archivists , and Feinholz emphasises multi-stakeholder coalitions . The shared goal of community-centred design thus masks significant differences in how that principle is operationalised.
Human-in-the-loop design principle (James D'Ercole) Community-centred data collection methodology (Aimee Ansari) Document digitisation workflow for linguistic heritage (Barbora Bromová) Full life cycle community inclusion (Dafna Feinholz)
All three speakers agreed that preserving low-resource languages through digitisation is important not only for communication but for cultural heritage. D'Ercole stated that 'when we preserve the language, we preserve the culture' and described this as the aspect of the project he values most . Ansari's work with Clear Global similarly treats language preservation as central to its mission, training over 100,000 linguists in 300 languages . Feinholz argued that language is 'just the way in which people express how they view the world' and that preserving it is about preserving heritage . However, they diverged on the mechanism: D'Ercole envisioned commercial partners eventually enabling internet access in digitised languages , Ansari focused on non-commercial open licences , and Feinholz emphasised policy guidance and multi-stakeholder coalitions . The shared goal of cultural preservation thus coexists with different visions of how digitised language data should be governed and used.
Language digitisation as cultural preservation (James D'Ercole) Community-centred data collection methodology (Aimee Ansari) Coalition for Linguistic Diversity in AI (Dafna Feinholz)
Both Ansari and Bromová agreed that international organisations have an important role to play in supporting language AI development for low-resource languages, and that private sector interest is structurally limited. Ansari argued that international organisations have a convening role to bring fragmented grassroots organisations together to pool resources, since individual organisations cannot afford the $50,000–$100,000 cost alone . Bromová acknowledged that private sector interest in very low-resource languages is 'relatively low' because of cost and that data governance questions remain open . However, they differed in emphasis: Ansari focused on the coordination and convening function , while Bromová focused on the data governance and ownership challenges that complicate private sector partnerships . Both agreed that the private sector alone cannot solve the problem, but neither offered a fully resolved model for how international organisations should fill the gap.
Convening role to pool resources (Aimee Ansari) Low commercial interest in low-resource languages (Barbora Bromová)
Both D'Ercole and Ansari agreed that building language AI for low-resource languages requires going beyond simple language datasets to incorporate cultural context, and that the complexity of primarily oral languages presents particular challenges. D'Ercole described building a 'cultural corpus, not just a language data set' and engaging an expert to make the model 'less Western-centric and more specific to Juba' . Ansari described the challenges of primarily oral languages, including wide variation in terminology, code-switching, and the absence of standardised writing systems . However, they approached the solution differently: D'Ercole focused on national colleagues in mission validating and collecting data , while Ansari described co-design approaches with community members and training 100,000 linguists . Both recognised the problem but operated at different scales and with different methodological emphases.
Cultural corpus over language dataset (James D'Ercole) Oral language complexity (Aimee Ansari)
- Communication barriers between peacekeepers and local communities represent a persistent, systemic problem across 11 UN missions involving over 50,000 uniformed personnel from approximately 120 countries, with language gaps undermining the core peacekeeping mission of listening and understanding communities.
- The vast majority of the world's languages remain poorly supported or entirely absent from frontier AI models, with wealthier-country languages dominating available text and speech data; even languages with tens or hundreds of millions of speakers, such as Nigerian Pidgin, have far fewer AI models than smaller European languages like Breton.
- Speech recognition technology covers only a fraction of global languages; in northeast Nigeria, only Hausa has any meaningful ASR support, meaning roughly 70% of the population risks being excluded from AI-assisted communication workflows.
- A real-time translation tool is being developed for UN peacekeeping in Juba, South Sudan, designed to function completely offline in austere, low-connectivity environments, with a human-in-the-loop approach ensuring AI assists rather than replaces human judgement.
- Digitising low-resource languages through AI tools preserves not only the language itself but also the culture embedded within it, and may eventually enable communities to access the internet in their own language, with potential for commercial partners to build upon open datasets.
- High-quality ASR model development requires hundreds of hours of systematically recorded voice data collected from diverse speakers across age, gender, educational level, and dialect, and lab-based word error rate metrics do not reliably predict real-world performance.
- Community-centred approaches to data collection, involving communities from initial design through development and deployment, are essential for linguistic and cultural inclusion in AI; late-stage community involvement is insufficient.
- Informed consent is a foundational principle for voice data collection, requiring community workshops to ensure speakers understand how their data will be used, stored, and how they can withdraw consent at any time.
- AI safety must be integrated by design from the outset, with community ownership, mutually defined guardrails, and alignment with the cultural and ethical norms of the communities being served, rather than retrofitted after development.
- Governments building sovereign AI models risk embedding political and ethnic biases, as illustrated by the South Sudan example, where a government dominated by one ethnic group may not equitably represent other languages and cultures in a national model.
- UNESCO's Coalition for Linguistic Diversity in Artificial Intelligence, established with Iceland, brings together over 30 experts from almost all world regions to share good practices, avoid siloed working, and inform policy guidance on linguistic diversity in AI.
- International organisations can play a critical convening role by bringing together fragmented grassroots organisations to pool resources for data collection, since individual organisations cannot afford the full cost alone but could collectively contribute smaller amounts to achieve shared goals.
- A distinction must be drawn between large technology companies and smaller private sector actors; collaboration with large companies can be appropriate for testing model safety, but data governance and community ownership concerns must be carefully managed.
- The Loria tool, developed in collaboration with the National Library of Serbia, demonstrates a replicable, open-source workflow for digitising historical documents trapped in paper formats, with archivists placed at the centre of the human-in-the-loop process.
- All datasets should be published openly under non-commercial licences to enable broad sectoral use, while ensuring that communities are not exploited through extraction of their language data for commercial gain without reciprocal benefit.
“Focus on the last mile. If you get it right at the furthest distance away from headquarters, whether that's headquarters in a mission or headquarters in New York or Geneva, then everything along the way has to be working.”
“By digitizing these languages, we're preserving the language. When we preserve the language, we preserve the culture. And when we do that, the language in theory can one day allow commercial partners to pick it up and you may be able to actually surf the web in that language itself.”
“Most languages in the world remain really poorly supported or entirely absent from any frontier model. A language like Breton, spoken by about 200,000 people, has over 50 models, while Nigerian Pidgin, with about 85 million speakers, has 10 models developed.”
“Lab metrics don't tell us very much about real world performance. When you're looking at a speech recognition model, you often hear that it has a low word error rate, but it doesn't really give you a sense of how well it works in the real world.”
“Take the government of South Sudan. There was a civil war. The Dinka and the Nuer were fighting against each other. If you have a government composed only of Dinka, how likely are they to want to build a sovereign model that reflects the Nuer culture, the Nuer language, that doesn't have any bias in it? One of the roles of the UN has to be around trying to make sure that when we're building global public infrastructure, we're building it in a way that is fair and equitable.”
“One of the biggest things is it costs a lot, because you want to be paying people fair wages to collect the data. You want to make sure that you're not exploiting people and that you're not extracting language data, which you are then building a model off of and selling back to people.”
“If we worked together with all five of them, then we can build the data that is relevant for all of them at reasonable cost for each of them. Most grassroots organisations don't have $50,000–$100,000. But pulling five or six of them together, they could probably find $10,000 each. We don't have the kind of overview that you do.”
“People who speak Dinka are not spending millions of dollars online, so they're just not commercially interesting.”
“The language is just the way in which people express how they view the world, how they understand, and how they see themselves. That's why they are also heritage, and that's why it's so important in the work that UNESCO does, and also in the area of AI.”
“How do you tackle the two-cost problem? It's either you're fighting to get a model in that language, or you're fighting to get people who know the language well enough. And does the private sector have to be involved in helping you get that data?”
How can the real-time translation tool be effectively tested and validated in shadow mode with human interpreters in the field, and what improvement cycles are needed?
James described a proof-of-concept phase where the tool would sit alongside interpreters and learn from their corrections, but the methodology for iterating on this process and measuring success in austere, offline environments remains an open area requiring further development and research.
How can cultural corpora be meaningfully integrated into AI language models to make them less Western-centric and more contextually appropriate for specific communities like Juba Arabic speakers?
James highlighted the need to go beyond a language dataset to build a cultural corpus, noting they had just brought in an expert on integrating culture into AI language models. This is an emerging and under-researched area, particularly for low-resource and oral languages.
How can a language cooperative or community ownership model be structured so that communities whose language data is collected can benefit from and have governance over that data, even when it is open source?
James raised the idea of a language cooperative where contributing communities could receive something back from the data they provide, but acknowledged this structure has not yet been worked out. This is a significant governance and equity question for the field.
How well do existing automatic speech recognition (ASR) models actually perform in real-world, field conditions for low-resource languages, as opposed to laboratory environments with controlled metrics?
Aimee pointed out that word error rates measured in laboratories do not reflect real-world performance, and that even for languages like Hausa that have some ASR support, actual field performance is largely unknown. This gap between lab metrics and practical utility is a critical area for further research.
What are the minimum requirements for high-quality voice data collection — in terms of speaker diversity by age, gender, educational level, and dialect — needed to build ASR models that work reliably across a language community?
Aimee noted that hundreds of hours of systematically collected, diverse voice data are needed, but the specific standards and methodologies for achieving this across different low-resource language contexts remain an area of research and practice.
How can the two-cost problem — the cost of building models in low-resource languages and the cost of finding sufficiently knowledgeable speakers to contribute data — be practically addressed, and what role should the private sector play?
The audience member articulated a dual challenge: it is expensive both to develop models for low-resource languages and to find people with sufficient linguistic expertise to contribute meaningfully. Whether and how private sector involvement can help resolve this without exploiting communities is a pressing open question.
How can multiple organisations working on language data for the same or related communities be brought together to share costs and avoid duplication of effort?
Aimee identified that fragmentation among organisations working on similar language data challenges is a major inefficiency. Coordinating five or six organisations to jointly fund data collection could make projects viable, but the coordination work itself is burdensome and requires an entity with sufficient overview to facilitate it.
How can international norms and standards be developed to ensure that sovereign AI models built by governments are linguistically and culturally representative of all communities within a country, including minority or historically marginalised groups?
Aimee used the example of South Sudan to illustrate how a government dominated by one ethnic group may not build a sovereign model that fairly represents other groups. She argued this is a role for the UN, but acknowledged it is a politically sensitive area with little current norm-setting activity.
How can standardised orthographic structures for primarily oral languages be established and recognised — potentially through bodies like IEEE or UNESCO — to enable consistent transcription and ASR development without losing linguistic nuance?
The speaker noted that for oral lingua franca languages like Juba Arabic, the absence of a written form creates fundamental challenges for ASR development. Establishing a community-validated, internationally recognised orthographic standard could be a prerequisite for scalable language technology development.
What data governance frameworks and ownership models are appropriate when working with private sector partners on low-resource language data, and how can communities retain meaningful control?
Both Barbora and Aimee acknowledged that while some private sector partners are open-minded about data governance, many are not, and communities speaking low-resource languages are often not commercially interesting to large technology companies. Developing robust, community-centred data governance models that can function across different types of private sector partnerships is an unresolved challenge.
How can the Loria OCR workflow be adapted and deployed across different languages and institutional contexts beyond Serbian, and what are the barriers to reuse?
Barbora presented Loria as an open, reusable tool published on GitHub, but the practical challenges of adapting it to other languages, alphabets, and institutional processes — particularly for languages using non-standard scripts or with limited digitised resources — remain to be explored.
How can UNESCO's Coalition for Linguistic Diversity in AI repository of good practices be systematically built, validated, and made accessible so that it genuinely informs policy guidance and capacity building across member states?
Dafna described the repository as a living resource still in development, with a co-created questionnaire to systematise data collection. The methodological and governance challenges of building a coherent, policy-relevant knowledge base from highly diverse community-led projects across world regions is an ongoing area of work.
How can the UN Peacekeeping real-time translation initiative (UNDPKO) be connected to or integrated with broader coalitions such as UNESCO's Coalition for Linguistic Diversity in AI?
Barbora explicitly raised the question of whether the UNDPKO peacekeeping project could join the Coalition for Linguistic Diversity in AI, suggesting this is an unexplored avenue for collaboration and coordination that could benefit both the peacekeeping mission and the wider coalition.
How can speech-to-speech translation be developed for low-resource oral languages, bypassing the need for intermediate text transcription, and what technical pathways exist for achieving this?
The speaker suggested that translating directly from speech to speech, rather than going through a written intermediate form, could be a viable approach for primarily oral languages. This is an emerging research direction that could circumvent many of the orthographic and transcription challenges discussed throughout the session.
How can informed consent processes for voice data collection be designed and implemented in communities where literacy is low, oral traditions dominate, and people may not fully understand the long-term implications of contributing their voice data?
Aimee identified informed consent as a guiding principle for Clear Global's work, including ensuring people understand how recordings will be used, stored, and the risks involved, and how they can withdraw consent. Designing consent processes that are genuinely meaningful in low-literacy, oral-tradition contexts is a significant practical and ethical research challenge.
