This panel discussion, moderated by Sonja Schmer-Galunder of the University of Florida , examines the risks of 'monoculture' in AI, described as the tendency towards homogenisation in AI models, training data, and outputs and its consequences for global resilience and cultural diversity . To illustrate the concept, Schmer-Galunder draws on the example of German forestry monoculture, where optimising timber production by removing ecological diversity ultimately led to the collapse of entire forests , and argues that AI faces analogous risks of correlated failure and brittleness .
Dr Supheakmungkol Sarin highlighted that current AI models are predominantly developed in the US or China using dominant datasets, making them ill-suited to respond appropriately to under-resourced cultures and languages . He noted that of approximately 7,000 languages worldwide, only about 100 are represented in AI models, meaning the vast majority of the world's ways of thinking, values, and cultures are entirely absent . He argued that the solution lies not in post-hoc localisation but in incorporating diverse cultural data at the foundational design stage .
Ambassador Muhammadou Kah emphasised that much of the Global South's knowledge is oral rather than written, making it especially vulnerable to exclusion from AI systems . He strongly rejected the model of data extraction, arguing that countries must retain ownership of their data and negotiate equitable benefit-sharing arrangements, drawing parallels with the historical exploitation of natural resources . He called for multilateralism and inclusive norm-setting to ensure the Global South can co-shape AI governance rather than merely accept rules set by others .
Wallace Cheng identified three layers of the monoculture problem: the dominance of a handful of English-language AI tools , the risk of collective intellectual homogenisation and reduced human judgement , and the danger that small biases embedded in widely deployed foundational models could become global defaults with irreversible consequences, particularly in healthcare and military applications .
Schmer-Galunder concluded by warning that the erosion of epistemic diversity poses a systemic resilience risk, citing the 2007-2009 financial crisis - in which reliance on a single flawed mathematical model caused USD11 trillion in household wealth losses - as a cautionary analogy . Ambassador Kah reinforced this, arguing that truly resilient AI systems must be comprehensive, inclusive, and broadly representative, and that non-representative models risk producing dysfunctional outputs with serious consequences when applied across different geographies . The panel collectively underscored the urgency of addressing cultural and linguistic diversity in AI before homogenisation becomes irreversible .
Overall Purpose
- The discussion aims to define and examine the concept of "monoculture" in artificial intelligence - the risks arising from a lack of diversity in AI models, training data, languages, and cultural representation. The panel seeks to explore the consequences of this homogenisation for global societies, particularly for the Global South, and to consider potential governance and diplomatic solutions.
- --
Major Discussion Points
- The analogy of ecological and systemic monoculture as a framework for understanding AI risks. The moderator introduces the concept of monoculture through historical and biological examples - including the collapse of geometrically planted German forests and the 2007-2009 financial crisis - to argue that AI systems optimised for efficiency at the expense of diversity risk becoming brittle and non-resilient.
- Cultural and linguistic underrepresentation in AI models poses serious safety and appropriateness risks. Panellists highlight that the vast majority of the world's approximately 7,000 languages are absent from current AI models, with only around 100 represented. This means that models trained predominantly on Western data fail to reflect the values, ways of thinking, and contextual needs of underrepresented communities, leading to potentially harmful or inappropriate outputs.
- Data sovereignty and the risk of extractive practices replicating colonial dynamics. Ambassador Kah argues forcefully that the Global South must reject data extraction without equitable benefit-sharing, drawing parallels to the historical extraction of natural resources. He advocates for data ownership, fair negotiation of terms, and win-win models of collaboration rather than one-sided extraction.
- Governance, multilateralism, and tech diplomacy as necessary responses to AI monoculture. Panellists argue that multilateral frameworks, norm-setting, and codes of conduct are essential tools for ensuring inclusivity and equitable representation in AI development. The Global South must move beyond being "rule takers" to actively co-shaping the rules governing AI.
- The irreversibility of epistemic diversity loss and its systemic consequences. The moderator and panellists warn that the homogenisation of knowledge and culture through AI monoculture could lead to irreversible damage - unlike financial crises, which can recover, the loss of cultural and epistemic diversity may be permanent. This creates systemic risks including reduced innovation, entrenched bias, and non-resilient societies and models.
- --
Overall Tone
- The discussion is earnest, intellectually engaged, and at times urgent. The moderator sets a thoughtful, academic tone from the outset, grounding the conversation in analogy and theory. As the panel progresses, the tone becomes increasingly passionate, particularly when Ambassador Kah speaks about data sovereignty and the parallels to colonial resource extraction. Wallace Cheng adds a measured, analytical perspective, while Supheakmungkol Sarin speaks with practical concern for underrepresented communities. Overall, the tone remains collaborative and solution-oriented, though tinged with concerns at the scale and urgency of the challenges described.
Expanded Summary: AI Monoculture - Risks, Representation, and Resilience
#
Introduction and Panel Overview
The afternoon panel session, moderated by Sonja Schmer-Galunder - Glenn and Deborah Renwick Leadership Professor of AI Ethics at the University of Florida - brought together experts from across the globe to examine the concept of "monoculture" in artificial intelligence and its consequences for global resilience, cultural diversity, and equitable development . The panel included Dr Supheakmungkol Sarin, Executive Director of AI Safety Asia and appointed AI expert to the United Nations Secretary General's High-Level Advisory Board on Artificial Intelligence , and Ambassador Muhammadou Kah, Gambian diplomat and Permanent Representative to the United Nations Office at Geneva . Wallace Cheng, Professor and Programme Director for Frontier Technologies and Governance at the Geneva School of Diplomacy, also contributed, bringing expertise in governance, trade, and sustainable development; he has additionally worked with the UN World Food Programme, Globe Ethics, and the International Centre for Trade and Sustainable Development, and has been involved with the World Economic Forum . Adam Russell and Gwyneth Sutherland - who had backgrounds in the national defence sector and were described as having particular expertise in the topic - were unable to attend, and Ambassador Kah arrived late, though the discussion proceeded substantively nonetheless .
#
Defining AI Monoculture Through Analogy
Schmer-Galunder opened by grounding the abstract concept of AI monoculture in a series of concrete historical and biological analogies, arguing that the risks facing AI systems are best understood through parallel examples from ecology, finance, and developmental biology . Her primary illustration drew on the political scientist James Scott's work, describing how German forests in the early twentieth century were planted in perfect geometric rows to maximise timber production and facilitate taxation . Everything considered "noise" - undergrowth, bushes, insects, and the broader ecosystem - was systematically removed in pursuit of optimisation . While this approach generated significant revenue for approximately one hundred years, it ultimately led to the complete collapse and death of the forest . Schmer-Galunder drew a direct parallel to AI: a lack of diversity, both in the models themselves and in their outputs, risks making human society and the models less resilient, more brittle, and potentially prone to collapse .
She reinforced this argument with two further analogies. The first concerned financial markets, where she later elaborated that between 2007 and 2009 the United States lost eleven trillion dollars in household wealth because every financial institution was operating under the assumption that housing prices were correlated with local economic conditions, when in fact they had become correlated with something the models were not measuring at all - the norms that lenders used to decide mortgage credibility . Drawing on Nassim Taleb's work, she argued that the markets became non-resilient because the system had optimised away the very variability that would have allowed it to absorb shocks . An audience member offered a contrasting interpretation of the same crisis, suggesting it was not merely a modelling failure but involved a deliberate misalignment of incentives - with sales incentives focused on volume rather than repayment ability - and structured financial products that were not independent of one another, resulting in what they characterised as a planned transfer of wealth from pension funds to hedge funds. The second analogy was biological: in embryonic development, early growth involves proliferation of cells, but genuine maturation means differentiation - cells specialising, forming relationships, and becoming interdependent systems . Healthy development, she argued, is not about size but about the integration of difference . These three analogies collectively established the intellectual framework for the entire discussion: that optimising for a single metric at the expense of diversity is a systemic pattern with potentially catastrophic consequences.
#
Cultural and Linguistic Underrepresentation in AI Models and the Need for Design-Stage Inclusion
Dr Sarin opened the substantive discussion by addressing the safety risks arising from the cultural and linguistic homogeneity of current AI models . He explained that models are predominantly developed in the United States or China using dominant datasets, meaning they are not designed to respond appropriately to the contexts of under-resourced cultures and languages . Speaking from the perspective of communities with under-resourced languages, he noted that outputs from such models would not be appropriate and would instead be catered towards dominant cultural contexts - creating real security and safety risks for communities whose needs the models were never designed to serve .
Sarin provided a striking empirical illustration of the scale of this problem, citing a figure he had encountered at a recent Tech Diplomacy Conference in Paris: of approximately 7,000 languages in the world, only around 100 are represented in current AI models, meaning the vast majority of the world's linguistic heritage is entirely absent . Crucially, he emphasised that these absent languages are not merely communication tools but carriers of distinct ways of thinking, values, and cultures . This point was reinforced by an audience member who noted that the challenge extends beyond language to fundamentally different cultural value systems - for example, differing moral weights assigned to elderly people versus infants across cultures - and recalled how Microsoft's attempt to create a universal encyclopaedia in the 1990s immediately encountered cultural divergences over questions such as who invented the telephone .
Ambassador Kah deepened this analysis by highlighting a dimension of the problem that goes beyond written language altogether . He observed that in many parts of the Global South, a significant proportion of valuable knowledge assets - spanning health, agriculture, wisdom, and values - is not written but oral, making its capture and inclusion in AI training pipelines especially difficult yet critically important . He illustrated the scale of linguistic diversity within single nations by citing Cameroon, which has over 200 languages, each representing a distinct culture and heritage with embedded knowledge assets . He further noted that in many Global South countries, the majority of the population is not formally educated, meaning that the informal population holding this knowledge represents the demographic majority rather than a marginal group .
Schmer-Galunder complemented these observations with concrete examples of how Western bias produces culturally inappropriate AI outputs - recommending picnics in countries where temperatures reach fifty degrees, or suggesting poetry based on birdsong in cultures where birds are considered a nuisance . While acknowledging these may not constitute catastrophic risks, she used them to illustrate how Western assumptions seep through into model outputs, and described her own work developing a cultural evaluation dataset called AntroBench in collaboration with Google to assess and address such biases .
A central argument advanced by Sarin - and broadly endorsed by other panellists - was that the solution to cultural underrepresentation cannot be found in post-hoc fine-tuning or localisation . He argued explicitly that making a model speak a language does not equate to representing the ideology or cultural context of the community that speaks it; localisation addresses surface form but not deep cultural meaning . The genuine solution, he contended, lies in ensuring that diverse cultural data, thinking, and values are incorporated at the very design stage of model development - a stage at which such inclusion is currently absent . He called on communities to engage in design thinking first: identifying what data, language, and cultural values they wish to see represented before contributing to model development .
Kah supported this position, arguing that the centrality of building capacity and competence in the Global South needs to be revisited in fundamentally new ways, and that the current approach of attempting to incorporate diverse data after the fact is fundamentally inadequate . Schmer-Galunder acknowledged a structural obstacle to this aspiration: foundational models require enormous computational resources, and there are currently no strong commercial incentives for providers of foundational models to incorporate diverse cultural data at the design stage . This tension between the structural logic of the AI industry and the imperative of inclusive design remained one of the discussion's central unresolved challenges.
#
Data Sovereignty and the Risk of Extractive Practices
The discussion took on a distinctly geopolitical character when Ambassador Kah addressed the question of data extraction and sovereignty . He drew an explicit and forceful parallel between the historical extraction of natural resources from the Global South - where value was taken and sold back at higher cost - and the emerging risk of data extraction without equitable benefit-sharing . He argued that the Global South's inability to convert its data into value should not result in giving that data away for free, and that the consequences of doing so would be far more severe than the historical resource extraction that had already impoverished many nations .
Kah reframed the Global South's position from one of passive vulnerability to one of latent leverage, arguing that sophisticated AI models still need Global South data and natural resources to function and improve, providing non-financial intangible assets that can be mobilised in negotiations . He proposed a win-win model in which originating communities retain ownership of their data while technology providers contribute the capital, infrastructure, and know-how to convert it into value, with benefits shared equitably between both parties . He cited the concrete example of Congo's natural resource data, preserved in Belgium, as an illustration of how ownership disputes over data can mirror those over physical resources .
Schmer-Galunder identified a fundamental tension - a "catch-22" - in this discussion: on one hand, countries risk data colonialisation by sharing their data; on the other, withholding data risks the permanent loss of endangered languages and cultures, particularly oral traditions that are especially vulnerable to disappearance . She also noted that the culture of data extraction is embedded at the very foundation of AI development, citing Stuart Russell's observation, made at an earlier panel, that every book ever written has been scanned into those models - yet this still does not resolve the diversity issue, as the corpus remains largely shaped by dominant languages and cultures . An audience member extended this point, noting that across Europe, monuments and cultural heritage had been digitised without fees, resulting in lost commercial opportunities - and proposed that rather than providing raw cultural material, communities should package their data together with its full cultural context to increase its value and strengthen their negotiating position .
#
Governance, Multilateralism, and Tech Diplomacy
Given the scale of the structural challenges identified, the panel converged on multilateralism and tech diplomacy as the primary governance mechanisms available to address AI monoculture and data extraction . Kah argued that many Global South countries lack the domestic regulatory capacity to deter extractive data practices unilaterally, making multilateral norm-setting and codes of conduct essential . He was emphatic that the Global South is not seeking exclusion from AI development but rather fairness, equity, transparency, and shared benefits - so that models become more adaptable to local health and agricultural realities, and so that communities receive tangible returns from their contributions . He stressed that the Global South must move beyond being "rule takers" to actively co-shaping the rules governing AI development .
Schmer-Galunder referenced the second global summit for tech diplomacy as a mechanism for equipping Global South countries not merely with a voice but with negotiation skills and diplomatic capacity to engage with technology companies that act like state actors . However, she also expressed scepticism about whether multilateral norm-setting alone can overcome the deeply embedded culture of extraction, questioning how this foundational dynamic can be changed for the Global South when it has not even been resolved within Western societies . This tension between Kah's relative optimism about multilateral mechanisms and Schmer-Galunder's more cautious assessment represented one of the discussion's more productive points of divergence.
#
Wallace Cheng's Three-Layer Framework
Wallace Cheng offered a structured analytical framework for understanding the problem of AI market concentration, identifying three distinct but interconnected layers . The first layer concerns tool concentration: despite the diversity of the people using AI, a handful of tools - fewer than ten, predominantly English-language - dominate globally, meaning that the outputs of these tools reflect a narrow cultural and linguistic perspective . The second layer concerns collective cognitive degradation: while AI may make individuals more efficient, Cheng raised the provocative question of whether it might make humanity collectively less creative, and lead to less human judgement, less human interaction, less local knowledge, and reduced communications about regional experience and mutual learning - a form of intellectual homogenisation that reduces the richness of regional experience . The third and most immediately consequential layer concerns the amplification of bias: when foundational models are deployed as infrastructure across critical sectors such as healthcare, education, military, and defence, even a small bias becomes a global default .
Cheng illustrated the stakes of this third layer with two concrete examples. In healthcare, he noted that while medicine may aspire to universality, healthcare systems are not universal - they reflect different infrastructures, cultural preferences, and economic situations, meaning that a biased model can produce clinically inappropriate recommendations across different health systems . In military and defence contexts, he warned that AI models with biases against certain groups, or that cannot distinguish civilians from soldiers, could create irreversible damages . These examples grounded the abstract concept of AI monoculture in sectors where the consequences of error are not merely inconvenient but potentially fatal and permanent .
#
Epistemic Diversity, Resilience, and Irreversibility
The final substantive phase of the discussion addressed what Schmer-Galunder framed as the deepest risk of AI monoculture: the erosion of epistemic diversity and its consequences for systemic resilience . She argued that if humanity loses the diversity of knowledge, culture, and ways of thinking that currently exists, the result will be not merely cultural impoverishment but the creation of non-resilient societies and non-resilient models - systems that are optimised for a specific task but incapable of adapting to unexpected contexts or shocks . She drew a qualitative distinction between recoverable crises - such as the financial markets, which recovered from the 2008 crash - and irrecoverable ones, such as the monoculture forest, which never recovered . This distinction, she argued, makes the loss of epistemic diversity potentially irreversible in a way that financial or technical failures are not .
Kah reinforced this analysis, arguing that lack of representation introduces biases that produce systems which are seemingly functional but generate non-optimal outputs, creating systemic risks that can spread globally with serious safety consequences . He argued that genuine resilience can only be achieved if AI systems are comprehensive, systematically inclusive, and broadly representative - and that a model fed flawed or geographically narrow data will produce harmful outcomes when applied in different contexts, such as a medical device designed for one geography being deployed in a rural village in Gambia, Tanzania, or Nigeria . An audience member added a further dimension to this concern, arguing that if only one dominant AI organism exists, diverse cultural data fed into it will ultimately serve a single political or ideological purpose, and proposing the development of multiple AI organisms as a structural safeguard .
#
Conclusion
The panel concluded with a shared sense of urgency about the scale and complexity of the challenges posed by AI monoculture, and a recognition that the discussion had only scratched the surface of what is a multi-dimensional problem spanning technical, cultural, economic, geopolitical, and civilisational dimensions . The collective message was clear: the homogenisation of AI models, training data, and outputs poses risks that go far beyond technical inefficiency or cultural insensitivity, rising to the level of systemic fragility and potentially irreversible loss. The discussion pointed to several interconnected imperatives: embedding diversity at the design stage of model development, establishing equitable data sovereignty frameworks, building negotiation capacity in the Global South, and developing inclusive multilateral governance mechanisms - all before the window for reversible intervention closes .
The Forest Monoculture Collapse Analogy
Arg. 1Sonja Schmer-Galunder uses the historical example of German forestry monoculture to illustrate how optimising a system for a single metric—timber production—by removing all diversity eventually leads to systemic collapse. She draws a direct parallel to AI, arguing that a lack of diversity in models and their outputs makes both human society and the models themselves less resilient and more brittle.
She describes how, at the beginning of the last century in Germany, trees were planted in perfect geometrical positions to maximise timber production and facilitate taxation, while all undergrowth, bushes, and insects were removed . This worked well for about 100 years, maximising revenue, but ultimately led to the complete collapse and dying of the forest . She explicitly draws the analogy to AI, warning of correlated failure risks when models share similar training data, languages, and weights .
on: AI monoculture creates systemic fragility and reduces resilience, with potentially irreversible consequences
The Financial Monoculture Risk
Arg. 2Sonja Schmer-Galunder argues that the 2007–2009 financial crisis is a powerful analogy for the dangers of AI monoculture, where reliance on a single flawed mathematical model across all financial institutions wiped out enormous wealth. The model optimised away the very variability that would have allowed the system to absorb shocks, making it non-resilient.
She states that between 2007 and 2009, the United States lost 11 trillion dollars in household wealth due to a financial crisis built on a single mathematical model . Every financial institution assumed housing prices were correlated with local economic conditions, when in fact they had become correlated with mortgage lending norms that the models were not measuring . She references Nassim Taleb's work on how markets became non-resilient by optimising away variability .
on: AI monoculture creates systemic fragility and reduces resilience, with potentially irreversible consequences
on: Whether the 2008 financial crisis was an accidental consequence of monoculture or a deliberately engineered transfer of wealth
Biological Differentiation as a Model for Diversity
Arg. 3Sonja Schmer-Galunder uses the biological process of embryonic development to argue that healthy growth is not about proliferation or size but about differentiation and integration of difference. She suggests this principle—that maturation means specialisation and interdependence—should inform how we think about AI development and diversity.
She explains that in early embryonic growth, the initial phase involves proliferation of cells, but true maturation means differentiation, where cells specialise, form relationships, and become interdependent systems . She concludes that from the earliest beginning of life, healthy development is integration of difference rather than mere size .
Cultural Inappropriateness of Biased Model Outputs
Arg. 4Sonja Schmer-Galunder argues that Western bias in AI models produces outputs that are culturally inappropriate for other contexts, even if not catastrophically harmful. She uses concrete examples to illustrate how models trained predominantly on Western data fail to account for the lived realities of people in other parts of the world.
She gives examples of culturally inappropriate AI outputs, such as recommending picnic activities in countries where temperatures reach 50 degrees and nobody goes for picnics, or suggesting poetry based on bird songs in countries where birds are considered annoying . She also mentions her work with Google on an evaluation dataset called AntroBench, designed to provide cultural evaluations for AI models .
on: Diversity and cultural representation must be built into AI models at the design stage, not added as an afterthought
Tension Between Preservation and Extractive Risk
Arg. 5Sonja Schmer-Galunder identifies a catch-22 in the debate over data inclusion: on one hand, countries risk having their data extracted without reciprocal benefit; on the other hand, withholding data risks the permanent loss of endangered languages and cultures. This tension is particularly acute for oral data, which is harder to collect and more easily lost.
She describes the dilemma where countries giving up their data for free may have a product sold back to them at a cost, raising concerns about data colonialisation . Conversely, some countries may want their data included precisely to prevent its loss, especially for oral traditions that are more difficult to capture and more easily disappear .
Extraction Culture Is Foundational to AI Development
Arg. 6Sonja Schmer-Galunder argues that the culture of data extraction is not a peripheral problem but is embedded at the very foundation of AI development, as demonstrated by the mass scanning of books and collection of personal data without explicit consent, even within Western societies. She questions how this deeply ingrained culture can be changed, particularly for the Global South.
She notes that Stuart Russell, speaking at an earlier panel, acknowledged that every book ever written has been scanned into AI models, which, while providing a large volume of data, does not solve the diversity issue as it remains largely English-language content . She also observes that even within Western cultures, individuals have not been asked for consent before their data was incorporated into AI models .
on: Data extraction from the Global South mirrors historical natural resource extraction and must be resisted
on: Whether multilateralism is a sufficient or realistic mechanism to counter data extraction
Tech Diplomacy to Build Negotiation Capacity
Arg. 7Sonja Schmer-Galunder argues that tech diplomacy summits are a necessary mechanism to equip Global South countries with the negotiation skills needed to engage with technology companies that now wield state-like power. She frames this as going beyond simply having a voice at the table to developing substantive diplomatic capacity.
She references a global summit on tech diplomacy-described as the second global summit for tech diplomacy-aimed at equipping countries of the Global South not just with a voice but with real negotiation skills and diplomatic tools to negotiate terms with technology companies that act like state actors .
on: Multilateralism and inclusive global dialogue are essential governance mechanisms for addressing AI monoculture
Some Diversity Losses Are Irreversible Unlike Financial Crashes
Arg. 8Sonja Schmer-Galunder draws a qualitative distinction between recoverable and irrecoverable forms of systemic collapse, arguing that the loss of epistemic and cultural diversity through AI monoculture may be irreversible in a way that financial crises are not. She uses the contrast between the financial market recovery and the permanent death of monoculture forests to make this point.
She notes that while the financial markets recovered from the 2008 crash, the monoculture forest never recovered, illustrating a special qualitative distinction between these two types of systemic failure . She warns that if current conditions do not allow for sufficient diversity, future societies may lose the adaptability needed to respond to outside shocks, constituting potentially irreversible damage .
on: AI monoculture creates systemic fragility and reduces resilience, with potentially irreversible consequences
on: Whether a single dominant AI system or multiple AI organisms is the preferable architecture
Western-Centric Training Data Creates Contextual Safety Risks
Arg. 1Supheakmungkol Sarin argues that AI models developed in the US or China using predominantly Western data are fundamentally ill-suited to respond appropriately to under-resourced cultural and linguistic contexts, creating significant safety and security risks. Because the models are designed to answer questions for a Western cultural context, their outputs in other contexts will be inappropriate or even harmful.
He explains that models are currently developed in the US or China with dominant data, and will therefore not be designed to respond or give appropriate answers for other contexts . He states that in most cases, outputs will be catered towards dominant probabilities suitable for other contexts, and that this creates security and safety issues because the model is not designed to fit the purpose of under-resourced communities .
Vast Majority of World's Languages Absent from AI Models
Arg. 2Supheakmungkol Sarin highlights that of approximately 7,000 languages in the world, only around 100 are represented in current AI models, meaning the vast majority of the world's linguistic and cultural heritage is entirely absent. He emphasises that languages are not merely communication tools but carriers of distinct ways of thinking, values, and cultures.
He references a discussion at the Tech Diplomacy Conference in Paris where the question of language representation in AI models arose, noting that of approximately 7,000 languages, only about 100 are represented in models . He stresses that the missing languages carry not just linguistic content but distinct ways of thinking, values, and cultures that are entirely unrepresented .
on: Languages are carriers of values, ways of thinking, and cultural heritage, not merely communication tools
Diversity Must Be Built In at the Design Stage
Arg. 3Supheakmungkol Sarin argues that the solution to cultural under-representation in AI is not post-hoc fine-tuning or localisation but ensuring that diverse cultural data, thinking, and values are incorporated at the very design stage of model development. He contends that localisation—making a model speak a language—does not equate to genuine cultural representation.
He argues that the way to solve the problem is not fine-tuning at a later stage or localising the model, but ensuring that the data of diverse cultures is represented when the model itself is being developed, at the very design stage . He illustrates the inadequacy of localisation by noting that a model can speak a language while still representing the ideology or context of another culture .
on: Diversity and cultural representation must be built into AI models at the design stage, not added as an afterthought
on: Whether cultural under-representation should be addressed at the design stage or through post-hoc localisation
Design Thinking Must Precede Data Contribution
Arg. 4Supheakmungkol Sarin argues that communities must first engage in design thinking—identifying what data, language, and cultural values they want represented—before contributing to model development. Simply adding more data without this prior reflection does not address the underlying problem of cultural misrepresentation.
He states that communities must start from design thinking: identifying what data represents their community and culture, and how to bring that knowledge into a model . He contrasts this with the current approach of simply trying to get more data in and make it work, arguing that speaking a language alone does not solve the problem of cultural representation .
on: Diversity and cultural representation must be built into AI models at the design stage, not added as an afterthought
Oral Knowledge Assets Are Excluded from AI Models
Arg. 1Muhammadou Kah argues that a significant portion of valuable knowledge assets in many parts of the world, particularly in Africa, is oral rather than written, making their inclusion in AI models especially challenging yet critically important. The exclusion of this oral knowledge means that entire bodies of cultural, sectoral, and heritage knowledge are absent from AI systems.
He notes that quite a number of useful knowledge assets in his part of the world are not even written, raising the question of how verbal and visible data can find its way into learning models . He emphasises that these oral knowledge assets span sectoral areas including health, agriculture, wisdom, and values, and that they need to be captured in AI models .
on: Languages are carriers of values, ways of thinking, and cultural heritage, not merely communication tools
Linguistic Diversity Within Single Nations Is Ignored
Arg. 2Muhammadou Kah uses the example of Cameroon, which has over 200 languages, to illustrate that the scale of linguistic and cultural diversity within a single nation far exceeds what current AI models capture. Each of these languages carries distinct cultural heritage and sectoral knowledge that must be incorporated into AI systems.
He uses Cameroon as an example, noting that it has over 200 languages, each representing a distinct culture and heritage with embedded knowledge assets . He explains that these knowledge assets extend beyond mere exchanges into sectoral data covering health, agriculture, wisdom, and values .
on: Languages are carriers of values, ways of thinking, and cultural heritage, not merely communication tools
Capacity Building Is Prerequisite for Inclusive Model Development
Arg. 3Muhammadou Kah argues that building capacity and competence in the Global South is essential to enable informal and non-literate populations—who represent the majority in many countries—to contribute their data to AI models. Without this investment in human capacity, the data of the majority will remain excluded.
He states that the centrality of building capacity and competence needs to be revisited in ways not previously considered, and that there is an informal population with data that needs to get into models . He points out that in many Global South countries, the majority of the population is not formally educated, meaning literacy rates skew towards those without formal education, making capacity building all the more critical .
on: Diversity and cultural representation must be built into AI models at the design stage, not added as an afterthought
on: Whether cultural under-representation should be addressed at the design stage or through post-hoc localisation
Reject Extraction; Negotiate Equitable Data Ownership
Arg. 4Muhammadou Kah argues that the Global South must firmly reject data extraction and instead negotiate rules of engagement that retain data ownership while establishing equitable, win-win economic models for data utilisation and benefit-sharing. He frames this as a matter of sovereignty and economic justice.
He states that data extraction is an evolving reality that must be rejected, at least from the Global South's perspective . He argues that the trade-off between extraction and preservation depends on negotiating rules of engagement that retain data ownership and establish an equitable, fair model for sharing the value extracted from data .
on: Data extraction from the Global South mirrors historical natural resource extraction and must be resisted
on: Whether providing raw cultural data or contextualised data packages is the better strategy for the Global South
Data Sovereignty as Counter to Historical Extraction Patterns
Arg. 5Muhammadou Kah draws a direct parallel between historical natural resource extraction and the emerging risk of data extraction, arguing that the Global South must not allow the same pattern to repeat itself with data. He contends that the Global South's data and natural resources give it non-financial intangible assets that can be leveraged as negotiating power.
He describes how historically, natural resources were extracted from Global South countries only for the value to be sold back to them at a higher cost, and insists this must not happen with data . He argues that the Global South's inability to convert data into value should not result in giving it away for free, and that sophisticated AI models still need Global South data and natural resources to function, providing negotiating leverage .
on: Data extraction from the Global South mirrors historical natural resource extraction and must be resisted
Multilateralism as the Primary Governance Mechanism
Arg. 6Muhammadou Kah argues that multilateralism must play a central role in norm-setting and codes of conduct to deter extractive data practices, since many Global South countries lack the regulatory capacity to do so unilaterally. He sees multilateral forums as the key arena for establishing fair rules around AI and data governance.
He acknowledges that some argue the Global South has no choice because data can be taken without their knowledge or capacity to prevent it, which is precisely why the ethics, transparency, and responsibility of AI matter . He argues that multilateralism can encourage states and private sector actors to think differently about these issues, and that the Global Dialogue on AI exists to ensure inclusivity and that all voices and perspectives are heard .
on: Multilateralism and inclusive global dialogue are essential governance mechanisms for addressing AI monoculture
on: Whether multilateralism is a sufficient or realistic mechanism to counter data extraction
Global South Seeks Co-Shaped Rules, Not Isolation
Arg. 7Muhammadou Kah clarifies that the Global South is not seeking to exclude itself from AI development or to prevent data sharing, but rather to co-shape the rules governing AI on fair, equitable, and transparent terms. He argues that inclusive models will ultimately produce better outcomes for all, including more adaptable AI for local health and agricultural contexts.
He states that the Global South is not saying the data is none of others' business, but rather calling for fairness, equity, transparency, and shared benefits in a win-win situation . He argues that models with more representative data will be more adaptable to local health and agricultural realities, rather than relying on synthetic data that creates problems in specific communities .
on: Multilateralism and inclusive global dialogue are essential governance mechanisms for addressing AI monoculture
Epistemic Monoculture Produces Systemic Dysfunction
Arg. 8Muhammadou Kah argues that the loss of epistemic diversity in AI leads to a lack of innovation, the introduction of bias, and systems that appear functional but produce non-optimal outputs, creating systemic risks that can spread globally with serious safety consequences. He frames this as a fundamental threat to the integrity and utility of AI systems.
He argues that lack of representation introduces biases, and biases produce dysfunctional systems that are seemingly functional, with outputs that become non-optimal and create systemic risks that can spread globally with safety concerns and dire consequences . He also notes that lack of diversity triggers a lack of innovation in the system .
on: AI monoculture creates systemic fragility and reduces resilience, with potentially irreversible consequences
Resilience Requires Comprehensive Representation
Arg. 9Muhammadou Kah argues that true resilience in AI systems can only be achieved if models are comprehensive, systematically inclusive, and broadly representative. He warns that a model fed flawed or geographically narrow data will produce harmful outcomes when applied in different contexts, using the example of medical devices applied in rural African communities.
He states that resilience can only be achieved if it is comprehensive, systematically inclusive, and broadly representative, enabling better innovation, optimisation, and adaptation . He gives the example of a medical device using non-representative data that may work well in one geography but create havoc when applied to a rural village in Gambia, Tanzania, or Nigeria .
on: AI monoculture creates systemic fragility and reduces resilience, with potentially irreversible consequences
Three-Layer Problem of AI Market Concentration
Arg. 1Wallace Cheng identifies three interconnected problems arising from AI market concentration: a handful of predominantly English-language AI tools dominate globally; collective human judgment, creativity, and local knowledge may diminish; and small biases in foundational models become global defaults when deployed across critical sectors. He frames these as compounding risks that must be addressed collectively.
He describes the first problem as the dominance of fewer than ten AI tools globally, the majority of which are in English . The second problem is that while individuals may become more efficient, collectively humanity may become less creative and rely less on local knowledge and regional experience . The third problem is that foundational models set up as infrastructure are widely deployed in healthcare, education, military, and defence, meaning a small bias becomes a global default .
on: Multilateralism and inclusive global dialogue are essential governance mechanisms for addressing AI monoculture
Irreversible Harm from Biased AI in Critical Sectors
Arg. 2Wallace Cheng argues that biased AI models deployed in military or healthcare contexts can cause irreversible harm, such as misidentifying civilians as soldiers or providing clinically inappropriate recommendations across different health systems. He distinguishes between the universality of medicine as a science and the non-universality of healthcare systems, which require local cultural and infrastructural knowledge.
He notes that in healthcare, while medicine itself may be universal, healthcare systems are not, requiring knowledge of different infrastructures, cultural preferences, and economic situations . In military and security contexts, he warns that AI models with biases against certain groups, or that cannot distinguish civilians from soldiers, will create irreversible damages . He states he is focusing specifically on these last two sectors in his work .
on: AI monoculture creates systemic fragility and reduces resilience, with potentially irreversible consequences
Cultural Value Systems Must Be Embedded in AI Responses
Arg. 1Speaker 1 argues that the challenge of AI monoculture goes beyond language to encompass fundamentally different cultural value systems, and that AI must be able to respond according to the culture of the user. Different cultures assign different moral and social weights to demographic groups, and AI systems must reflect these differences rather than imposing a single cultural framework.
He gives the example that in some distant cultures, older people are considered more relevant than babies, while in Western cultures the reverse is true, and this makes a significant difference in AI responses . He also references the historical example of Microsoft's attempt to create a universal encyclopaedia in the 1990s, which immediately encountered the problem that the telephone was attributed to different inventors-Graham Bell in the US, Meucci in Italy, and Popov in Russia-illustrating how cultural framing shapes knowledge .
on: Languages are carriers of values, ways of thinking, and cultural heritage, not merely communication tools
Package Cultural Context with Raw Data to Strengthen Bargaining Power
Arg. 2Speaker 1 argues that rather than providing raw oral or cultural material to AI developers, communities should package their data together with its full cultural context, thereby increasing its value and improving their negotiating position. He warns that providing raw material alone is a win-lose approach that cedes control once the data is extracted.
He argues that transferring raw material is a win-loser approach because once data is extracted, it is very hard to maintain a balanced position . He suggests that communities should provide a package that includes not only raw oral traditions but all the cultural context that improves the value of those traditions, making it easier to achieve a balanced relationship between the two sides . He draws on the European experience of monuments being digitised without fees, resulting in lost commercial opportunities .
on: Data extraction from the Global South mirrors historical natural resource extraction and must be resisted
on: Whether providing raw cultural data or contextualised data packages is the better strategy for the Global South
Deliberate Structural Flaw in Financial Monoculture
Arg. 1Speaker 2 argues that the 2008 financial crash was not merely an accident of model error but was structurally engineered through misaligned incentives and deliberately non-orthogonal financial products, resulting in a planned transfer of wealth from pension funds to hedge funds. This challenges the framing of the crisis as an unintended consequence of monoculture and suggests intentional design.
He explains that the crash occurred because mortgage sales incentives were aligned with volume rather than ability to repay , and that financial products were structured in a deliberately non-orthogonal way, which was clearly known at the time and resulted in a transfer of money from pension funds to hedge funds .
on: Whether the 2008 financial crisis was an accidental consequence of monoculture or a deliberately engineered transfer of wealth
Single AI Organism Poses Ideological Capture Risk
Arg. 2Speaker 2 argues that if only one dominant AI system exists, the diverse cultural data fed into it will ultimately serve a single political or ideological purpose, endangering those who provide that information. He proposes that having multiple AI organisms—rather than a single dominant one—is a preferable and safer architecture.
He suggests that if only one organism exists, the data provided by various cultures will serve one political or ideological purpose, and the more information there is about different ideologies, the greater the danger for those who provide that information . He therefore proposes multiple AI organisms as a solution .
on: AI monoculture creates systemic fragility and reduces resilience, with potentially irreversible consequences
on: Whether a single dominant AI system or multiple AI organisms is the preferable architecture
Session Knowledge Graph
Speakers · Topics · Arguments · Relationships
