Beyond State vs. Corporate: Community Data Governance Models from the Global Majority
The session discussed community-centric data governance, looking into what it means for communities to govern data amid concerns about both state surveillance and private-sector extraction. The discussion focused on Global Majority contexts where debates over digital public infrastructure and data localisation are intensifying . The speakers discussed the development of a policy brief and structured the session around three pillars: data ownership and stewardship, participatory decision-making, and benefit sharing, which they would test against a hypothetical case study .
Bridgette Ndlovu presented the first pillar by arguing that data should be treated as a collective asset rather than an individual commodity, with communities rather than corporations or states holding authority over it . She described existing approaches such as sectoral stewardship codes, community data commons, participatory models, and data trusts, citing examples from the African Union framework, Mozambique, and South Africa . The policies, according to Ndlovu, should include recognition of self-determination, prohibitions on harmful extractive practices, free, prior and informed consent, sector-specific codes of conduct, and stronger transparency and accountability for platform and AI systems .
Fernanda Campagnucci explained the second pillar as enforceable community power throughout a project’s lifecycle, distinguishing genuine participation from token consultation and emphasising monitoring, review, and even suspension when harms emerge . She pointed to examples and inspirations from Brazil’s CGI, Chile’s public algorithm repositories, Kenya’s Huduma Namba and DPIA experience, Amsterdam’s AI pause, and FPIC procedures in the Philippines . Her proposals included FPIC for projects affecting identifiable communities, permanent oversight boards for high-impact systems, funding for independent community intermediaries, and limits on trade-secret claims in public systems .
Shumaila Shahani presented benefit sharing, arguing that communities in the global majority bear the risks of data extraction while value flows to companies headquartered in the global north, so communities should share in the value as a matter of right rather than charity . She noted that benefit-sharing principles already exist in mining and oil arrangements, and argued that data-extractive enterprises should meet similar standards, while recognising that value includes money, infrastructure, capacity, access, and improved services . She added that fairness requires auditable valuation over time and may justify lifecycle-long, renegotiable agreements, public registries, community-controlled funds, and repeated audited value-impact reports .
The hypothetical case concerned Windfall Valley, where a foreign health technology company proposed an AI diagnostic and disease-surveillance system using local symptom, location, and biometric data, while retaining rights to models that would later be licensed abroad . The case highlighted tensions between urgent public-health needs and weak consultation, noting that the indigenous group had not been consulted separately, migrant workers lacked representation, and the valley was offered services rather than direct financial returns, despite likely substantial commercial value abroad . Audience participants raised questions about who validates knowledge across medical traditions, whether data rights and template agreements are needed, how taxation or token-based systems might redistribute value, and how governance can represent both organised communities and unorganised individuals . The session ended without firm conclusions, but it established a shared need for frameworks that combine community authority, state-backed enforcement, meaningful participation, and fair, long-term value-sharing in data governance .
- The overall purpose of the discussion was to explore what 'community-centric data governance' should mean in practice, especially in Global South and Global Majority contexts where both state control and private-sector extraction create risks. The organisers were also testing principles for a draft policy brief against a hypothetical health-data case study to see whether their ideas on ownership, participation and benefit-sharing hold up in real-world scenarios.
- The discussion began by framing data governance as a power issue rather than merely a privacy issue, with concern about both government surveillance/data localisation and extractive corporate practices. The speakers argued that communities in the Global Majority are caught between these two models and need an alternative approach.
- A major theme was data stewardship and community agency: data should be treated as a collective asset, with communities-not only states, corporations, or isolated individuals-holding meaningful authority over how it is used and how benefits are distributed. The speakers pointed to models such as data trusts, community data commons and sector-specific codes, and called for self-determination, free, prior and informed consent, and stronger transparency and accountability.
- A second major point concerned participatory decision-making. Fernanda Campagnucci stressed that genuine participation must go beyond symbolic consultation and give communities enforceable power from project design through operation and review, including the ability to monitor systems and trigger suspension if harms emerge. Suggested mechanisms included FPIC procedures, permanent community oversight boards, technical intermediaries, and limits on trade-secret claims in high-impact public systems.
- Benefit-sharing was presented as a core justice issue. Shumaila Shahani argued that communities whose data generates value and who bear the risks should receive a fair share of that value as a right, not as charity or CSR. She noted that value may include money, infrastructure, capacity-building, access to insights, and improved services, and proposed binding benefit-sharing agreements, public disclosure of agreements, community-controlled funds, and recurring audited data-value impact reports.
- The hypothetical 'Windfall Valley' case study focused the discussion on practical tensions: urgent public-health needs versus weak consultation, cross-border data control, lack of representation for Indigenous people and migrant workers, and uncertain adequacy of non-monetary benefits. Audience interventions raised related issues, including plural understandings of valid knowledge, the need for fundamental data rights and agreement templates, taxation of AI firms, collective rather than purely individual data governance, data cooperatives, and the need for flexible legal frameworks adapted to the distinctive nature of data.
- The overall tone was thoughtful, collaborative and exploratory throughout. At the start, it was analytical and agenda-setting, as the organisers introduced the problem and their draft framework. In the middle, it became more practical and deliberative as the speakers tested policy ideas against the case study and invited audience responses. Toward the end, the tone became slightly more urgent and compressed because of time pressure, though it remained constructive and open to further engagement beyond the session.
Shumaila Shahani welcomed in-person and online participants and introduced the session as a discussion on what “community-centric data governance” should mean in practice, especially in Global Majority and Global South contexts where people face risks from both state surveillance and private-sector extraction . She stressed that digital public infrastructure is expanding through systems people often cannot opt out of and may not meaningfully access . She explained that she, Fernanda Campagnucci and Bridgette Ndlovu were developing a policy brief that was still in progress and wanted to test its principles against practical realities through expert discussion and a hypothetical case study . Bridgette introduced herself as being from Paradigm Initiative in Zimbabwe, and Fernanda introduced herself as being from InternetLab in São Paulo, Brazil .
Shumaila then framed the discussion as a question of power rather than privacy alone. With AI systems and digital public infrastructures spreading into everyday life, she said the key questions are who owns data, who benefits from it, and who bears the risks . She said the draft brief was organised around three pillars: data ownership and stewardship, participatory decision-making, and benefit sharing .
Bridgette Ndlovu presented the first pillar, data stewardship and community agency. She argued that data should be treated as a collective asset rather than an individual commodity, and that communities and not corporations or states must hold legal authority over it . She described the core idea as keeping the community at the centre of all decisions concerning its data . She defined stewardship as collective authority over data generated within a community’s boundaries, with community decision-making on use, sharing and the distribution of benefits .
Bridgette pointed to existing and emerging governance models including sectoral stewardship codes, community data commons, participatory data scout models and data trust models . In the African context, she highlighted the African Union Data Policy Framework, which recognises data trusts as one model for countries to adopt, while noting implementation gaps because many governments have not followed that framework consistently . She added that Mozambique has adopted national data governance policies alongside data protection laws, and that South Africa’s POPIA framework allows representative bodies to submit codes dealing with data quality, consent and redress, which she linked to sectoral stewardship codes .
Her policy asks made this pillar more concrete. She called for recognition of a right to self-determination grounded in broader international standards . She proposed prohibitions on harmful and unfair market practices in extractive contexts where communities face unequal bargaining power . She also called for free, prior and informed consent so that communities can decide whether they want to participate in data processes at all . She repeated the need for sector-specific codes of conduct to guide responsible sharing, reuse and repurposing of personal data . Finally, she called for mandatory transparency and accountability for key platform services, including recommendation systems, especially as governments increasingly adopt AI policies and AI-driven systems .
Fernanda Campagnucci presented the second pillar, participatory decision-making. She argued that communities must have enforceable power to shape projects from beginning to end, rather than being offered only symbolic or discretionary consultation . She distinguished between performative participation and genuine participation, saying that consultation is often episodic, discretionary and non-binding, whereas real participation requires institutional arrangements that let communities shape the terms of collection, use and sharing before implementation begins . She added that participation must continue during implementation through monitoring, and after deployment through the ability to trigger review or even suspension when harms emerge .
Fernanda said there was no perfect one-size-fits-all example, but cited Brazil’s CGI and a preceding discussion on stakeholder guidelines as inspirations . She also referred to a repository of public algorithms in Chile, Kenya’s Huduma Namba litigation and later DPIA guidance, Amsterdam’s pause in an AI programme, and the Philippines’ free, prior and informed consent procedures .
Her policy proposals followed directly from these examples. She argued that projects affecting identifiable communities should be subject to free, prior and informed consent procedures . She also called for permanent community oversight boards for high-impact data systems . Another practical measure was public funding, including potentially from states and large technology firms, for independent community-based technical intermediaries so communities have the expertise to participate effectively . Finally, she argued that trade-secret and commercial-confidentiality protections should not apply in public systems or systems affecting communities where visibility into how the system operates is necessary .
Shumaila Shahani presented the third pillar, benefit sharing, as a matter of justice in the global political economy of data. She argued that data from billions of people in the Global Majority powers digital systems, but much of the resulting value flows to companies headquartered elsewhere rather than back to the communities that supplied the data and carried the risks . She said that if communities bear the risks of sharing data, they should also share in the value created from it . She was explicit that this should not be treated as charity or corporate social responsibility, but as a duty on the part of companies and a right of communities .
To illustrate the principle, Shumaila pointed to extractive industries. She noted that mining codes across Africa already require benefit-sharing agreements and that oil revenue funds in places such as Alaska and elsewhere distribute direct dividends to residents . At the same time, she acknowledged that data differs from traditional resources because the value derived from it is harder to define, measure and time .
She then unpacked what “value” should mean in data governance. In her view, value is not just money; it also includes infrastructure, capacity building, access to data, access to insights generated from data, and improved services . She warned that these forms of value are not currently being audited, making it difficult to judge whether what communities receive is fair or adequate . She also stressed that fairness cannot mean only a token share of very large profits, especially when data collected today may continue generating value many years later . For that reason, she suggested lifecycle-long agreements and renegotiable rights .
Her specific policy asks followed from this view of value creation. She proposed binding benefit-sharing agreements as a condition for data licensing, while stressing that licensing was only one part of the broader governance picture . She also called for public disclosure and registries of all such agreements so communities can compare terms and benchmark what others have received . A further principle was that any resulting funds should be community-controlled . Finally, she proposed repeated, audited data value impact reports to track when and how much value has been generated from community data over time .
The hypothetical case study, Windfall Valley, was then introduced to test the three pillars . Shumaila described the valley as a rural region of about 400,000 people in a lower-middle-income country, with an ethnically mixed population, a majority farming community, a smaller indigenous group in the uplands, and seasonal migrant agricultural workers . Mobile phone penetration was high, but formal healthcare was weak, consisting of one understaffed district hospital and a network of community health workers . A foreign company, VitalReach, proposed deploying an AI-driven diagnostic and disease-surveillance system in partnership with the national Ministry of Health, which was enthusiastic and had signalled it would sign . Community health workers would collect symptom, location and biometric data through a mobile app, while VitalReach’s models would detect outbreak patterns and suggest diagnoses .
The proposed arrangement raised several governance issues. VitalReach would train the system on the valley’s data, retain all rights to the resulting models, and license them abroad to other countries . In return, the valley would receive five years of free use of the app, training for fifty community health workers, and a new cold-storage facility for vaccines at the district hospital; no money would change hands directly . The data, including indigenous health information and migrant workers’ biometric records, would be transferred to VitalReach servers abroad and governed under the law of the country where those servers were located . Shumaila highlighted three tensions in the scenario: the valley genuinely needed better diagnostics and refusal had real human costs; the indigenous group had not been separately consulted and the migrant workers had no representation; and the models would likely generate significant commercial value abroad while the valley was offered only unassessed services in return .
Bridgette opened discussion of the case by asking who the legitimate steward of the valley’s data would be, whether there could be more than one steward, and how representation should work in such a mixed community . Shumaila also made clear that the session aimed to hear practical responses rather than simply present a finished framework .
The first audience intervention focused on differing understandings of valid knowledge. The speaker argued that different communities do not necessarily share the same understanding of valid knowledge, particularly in medicine, citing Ayurveda, Chinese medicine, Western medicine and Indigenous healing traditions . The comment raised an additional question about who defines valid knowledge and evidence across medical traditions .
Another audience participant proposed a more institutional and economic response. They suggested a fundamental right to data at both individual and community levels, so governments could not negotiate away data as if it were ownerless . They also recommended model templates for data-sharing agreements to help governments and communities negotiate fairer terms . Finally, they argued that because single data points become highly valuable through aggregation, AI firms may need to be taxed through an internationally coordinated regime that can compensate communities for the global value generated from their data . Bridgette connected this to competition concerns and cited a recent Nigerian competition decision against Meta, Facebook and WhatsApp as an example of data-sharing practices being treated as anti-competitive .
An online participant, Angelica Saldana, asked where to learn more about international standards on data self-determination, and referred to related issues such as digital identity, social safety nets and mobile network governance . Bridgette said the organisers would share their thoughts later and continue the conversation beyond the event .
When the conversation moved to the second pillar and the question of meaningful participation in the valley case, Fernanda asked what genuine participation would look like before the Ministry signed the deal and whether existing multi-stakeholder or oversight bodies in other sectors might be adapted to data governance . An audience member responded that current governance remains trapped in a post-Westphalian nation-state model that does not fit the global nature of the internet . They criticised both centralised systems and anarchic decentralisation, and proposed governance structures that connect individuals, families and communities directly into systems while still maintaining legal compliance and interoperability . They also floated tokenomic and programmable-money mechanisms as possible ways to route benefits back to data contributors when models built from their data are monetised .
Fernanda responded that even if the goal is to move beyond a simple government-versus-corporate model, states would still be needed to legitimise and enforce any new governance arrangement, because companies are unlikely to respect community validation processes without legal authority behind them . She then raised another challenge: community governance discussions often presume an organised community, such as an indigenous or traditional group, but data governance also affects many individuals who are not organised at community level, especially in cities . She said community data should be treated respectfully and governed with strong collective rights, but that there were unresolved questions about how to represent dispersed individuals who do not belong to clear community structures .
Abhinav, a researcher from IT for Change, then contributed a technically grounded intervention. He argued that individual informed consent, while central in liberal democracies, is inadequate for health AI because such systems do not function on isolated individual data points . Diagnostic AI requires high-quality datasets about a whole community, especially if the model is meant to make decisions about that community . For that reason, he argued there is real value in collective governance and that representative structures such as data cooperatives or councils could be involved at every stage of the process, from collection through model training and use . He referred to examples from India, where indigenous and community decision-making traditions already exist in climate and forest-rights contexts, and linked this to the idea of collective data as an asset .
In the final minutes, Shumaila returned to the benefit-sharing problem in the case and asked how anyone could judge whether the offered services were enough, not enough, and against what baseline such adequacy should be measured . Angel Akunle from the Internet Society argued that the answer begins with a data governance framework establishing common rules and a baseline for how data is defined and handled . He suggested that such a framework would help different agencies align their data practices and combine data more seamlessly, which in turn would allow a clearer baseline for evaluating value and adequacy .
A final audience speaker cautioned against too-simple analogies with mining or forestry. They said that data is far more amorphous than timber or mineral extraction, and that communities around data may also be more fluid and difficult to define . They argued that data can later generate implications in other domains, such as the environment, and that legal frameworks therefore need to be flexible and creative rather than copied directly from older extractive-sector models .
The session ended without formal conclusions because time ran short . Still, several themes recurred throughout the discussion. Speakers framed data governance as a question of power, ownership, benefit and risk rather than privacy alone . Across the three pillars, they argued for collective or community authority over data, meaningful participation throughout a project’s lifecycle, and binding benefit-sharing arrangements backed by transparency and auditability . The discussion also surfaced unresolved questions about who can legitimately represent mixed or unorganised populations, how to include indigenous groups and migrant workers, what counts as valid knowledge in health systems, and how to judge whether non-monetary benefits are fair . Because time was running short, Shumaila closed by inviting participants to continue sharing resources and ideas after the forum or bilaterally .
Practical Implementation Challenges #The Convenience-Privacy Trade-off An audience member from Brazil raised the practical challenge of how people "trade convenience for data" without fully understanding risks, c...
The knowledge base supports this broader framing by noting concerns about both state misuse of data and corporate extraction or opacity. [S106] warns that more data can become a source of government control and misuse, while [S101] highlights insufficient corporate oversight and the need for transparency and accountability in how private actors use personal data.
This is consistent with wider concerns in the knowledge base. [S96] notes that digital public infrastructure is being actively developed and can harness huge amounts of data, making governance critical. [S106] adds that governments often move services online without adequate alternatives for those who do not want to participate or who are excluded by lack of access or skills.
The knowledge base confirms Fernanda's affiliation with InternetLab and identifies her as a Brazil-based Global South participant. [S9] explicitly refers to Fernanda in connection with Internet Lab and Brazil.
This is corroborated by multiple sources emphasising collective approaches to data governance. [S23] argues that data value should be rethought collectively through models such as data cooperatives, and [S40] states that current regulation often over-focuses on individual data rights while neglecting group dynamics and collective rights.
The knowledge base confirms that data trusts, data commons and related collective governance models are recognised approaches in current debates. [S23] explicitly lists data cooperatives, data commons and data trusts, while [S40] discusses data trusts and cooperatives as institutional models for collective data handling.
The knowledge base confirms the relevance of the African Union Data Policy Framework to data governance on the continent. [S40] presents the framework as extending data governance beyond first-generation rights and as a potential response to asymmetrical power relations in data sharing. [S62] also refers to the framework as an example of a people-centred approach adopted by African Union member states.
While the knowledge base does not verify this exact implementation assessment, it adds supporting context on enforcement and participation weaknesses in African data governance. [S40] highlights limited civil-society participation in many African public processes and stresses the need for stronger enforcement and access to information, which helps explain implementation gaps.
The knowledge base does not directly confirm the POPIA codes point, but it does support the broader claim that South Africa has data governance enforcement structures. [S40] notes that a data protection and information regulator in South Africa took strong action against WhatsApp and other groups, indicating an institutional framework capable of oversight and code-based governance.
The knowledge base adds relevant normative context by showing parallel calls for stronger rights-based and sovereignty-oriented approaches to data governance. [S104] argues that people have often ceded sovereignty over their data without meaningful understanding, and [S23] argues for greater voice, choice and stake online through collective governance mechanisms.
This is consistent with broader concerns in the knowledge base about dominant business models built on data extraction and asymmetrical power. [S23] states that the most lucrative model remains one that extracts users' data for advertising, and [S40] discusses asymmetrical power relations in data sharing and the need for broader economic regulation.
Practical Implementation Challenges #The Convenience-Privacy Trade-off An audience member from Brazil raised the practical challenge of how people "trade convenience for data" without fully understanding risks, c...
Data governance is about power, not only privacy - short name (Shumaila Shahani)
Arg. 1Shumaila argues that the spread of AI systems and digital public infrastructure has made data governance fundamentally a question of power. Her point is that the central issue is who owns data, who benefits from it, and who carries the risks, rather than privacy alone.
She explicitly says that with AI systems and DPIs proliferating, data governance is "not about privacy anymore" but instead concerns power, ownership, benefit and risk allocation . She also frames the wider debate as one between state surveillance and private extraction, especially in global majority countries that worry about both government localisation practices and corporate abuse .
on: Data governance should be community-centred rather than left primarily to states or corporations.
In the case study, stewardship is complicated because Indigenous groups and migrant workers are not equally represented, raising the question of who can legitimately speak for absent or unorganised groups - short name (Shumaila Shahani)
Arg. 2Shumaila highlights that the hypothetical valley is not a single uniform community, so deciding who can govern the data is politically complex. She stresses that the Indigenous council was not separately consulted and migrant workers have no representation, making legitimate stewardship and consent difficult.
In presenting the case study, she notes that the Indigenous upland group has its own council but was not separately consulted because the ministry treats the valley as one undifferentiated population . She also points out that migrant workers have no representation and may not even be present when decisions are made, which directly raises the problem of who can speak for absent groups .
on: Representation is a major challenge, especially for Indigenous peoples, migrant workers, and other absent or unorganised groups.
on: Whether there can be a single legitimate steward or uniform community voice in the case study
Communities that bear the risks of data sharing should receive a fair share of the value created, and this is a right rather than charity or corporate social responsibility - short name (Shumaila Shahani)
Arg. 3Shumaila argues that benefit sharing should be understood as justice, not goodwill. If communities take on the risks of supplying data, they should have a rightful claim to the value produced from that data.
She states that systems are powered by data from billions of people in the global majority, while value flows back to companies headquartered in the global north . She then says that those who bear the risks of data sharing should also share in the value, and emphasises that this is neither charity nor corporate social responsibility but a company duty and a community right .
on: Benefit sharing from data should be fair, structured, and not treated as charity.
on: Whether fair benefit sharing should rely mainly on direct community negotiation and agreements or also on internationally coordinated taxation and compensation systems
Value from data is not only financial; it also includes infrastructure, training, access to insights and improved services, so fairness must be assessed across multiple forms of benefit - short name (Shumaila Shahani)
Arg. 4She argues that value generated from data cannot be reduced to cash payments alone. Fairness has to consider non-monetary returns such as infrastructure, capacity building, access to data-derived insights, and better services.
She says explicitly that value is "not just money" and includes infrastructure, capacity building, access to data, access to insights, and improved services . She adds that these forms of value are not currently being audited, which makes it difficult to determine whether what is offered is fair or sufficient .
on: Benefit sharing from data should be fair, structured, and not treated as charity.
Because data gains value over time and through aggregation, communities need lifecycle agreements, renegotiation rights and transparency over realised value - short name (Shumaila Shahani)
Arg. 5Shumaila contends that data can continue generating value long after it is collected, so one-off agreements are inadequate. Communities therefore need long-term arrangements, the right to renegotiate when value compounds, and transparent reporting on what value has actually been realised.
She explains that data collected today can produce value ten years later, so a small one-time share cannot be considered fair . On that basis, she proposes lifecycle-long agreements, renegotiable rights when value compounds, public disclosure of agreements, and repeated audited data value impact reports so communities can understand when and how much value has been generated .
on: Benefit sharing from data should be fair, structured, and not treated as charity.
on: Whether fair benefit sharing should rely mainly on direct community negotiation and agreements or also on internationally coordinated taxation and compensation systems
The proposed deal offers needed health improvements, but refusal also carries human costs, making community governance decisions ethically difficult - short name (Shumaila Shahani)
Arg. 6In the case study, Shumaila presents the decision as morally difficult rather than straightforwardly extractive. The valley needs better diagnostics, and refusing the proposal could mean preventable deaths, so governance must balance immediate welfare with long-term rights and risks.
She describes the valley as having weak formal healthcare, with one understaffed district hospital and community health workers, while VitalReach offers an AI diagnostic and disease surveillance system . She then stresses that the valley genuinely needs better diagnostics, that people are dying from conditions that earlier detection could address, and that refusing the deal has a real human cost .
on: The health AI case involves real development benefits but also serious governance risks, so decisions are ethically and politically difficult rather than straightforward.
The ministry’s treatment of the valley as a single population ignores differentiated rights and needs, especially of the Indigenous council and migrant workers - short name (Shumaila Shahani)
Arg. 7Shumaila argues that the case study reveals a governance failure in treating a diverse population as if it were homogeneous. This erases the particular rights, representation needs and vulnerabilities of Indigenous people and migrant labourers.
She states that the Indigenous upland group has its own council but has not been separately consulted, because the ministry treats the valley as one undifferentiated population . She also notes that migrant workers have no representation at all and may be absent when decisions are made, showing that the ministry's approach overlooks their specific position .
on: Representation is a major challenge, especially for Indigenous peoples, migrant workers, and other absent or unorganised groups.
The adequacy of non-cash benefits such as free services and training cannot be assumed; it must be independently assessed against an agreed baseline - short name (Shumaila Shahani)
Arg. 8She questions whether offers like free app access, staff training and cold storage are actually fair compensation for the commercial value extracted from community data. Her point is that such benefits need independent assessment against a defined baseline rather than being accepted at face value.
In the case study, she notes that VitalReach offers five years of free use of the diagnostic app, training for 50 community health workers and a cold storage facility, with no direct cash payment . She adds that the model will generate substantial commercial value abroad while the valley receives only services "whose adequacy no one has independently assessed", and later asks how to know whether the services offered are enough, against what baseline, and how that baseline should be created .
on: Practical governance tools are needed, including frameworks, codes, templates, transparency rules, and oversight mechanisms.
on: Whether baseline fairness and governance should be defined through fixed frameworks and templates or through more flexible, evolving legal approaches
Communities should hold collective authority over data rather than corporations or states - short name (Bridgette Ndlovu)
Arg. 1Bridgette argues that data should not be controlled primarily by governments or private firms. Instead, community-centred governance should place legal authority over data in the hands of the communities from which that data is generated.
She states that the core idea is to see data as a collective asset rather than an individual commodity, and says communities, not corporations or states, must hold legal authority over it . She reinforces this by saying the community should be at the core of all decisions concerning its own data .
on: Data governance should be community-centred rather than left primarily to states or corporations.
on: Whether community-centric data governance should minimise the role of the state or still depend on state enforcement
Data should be treated as a collective asset, with stewardship defined through community decision-making over use, sharing and benefits - short name (Bridgette Ndlovu)
Arg. 2Bridgette defines data stewardship in collective rather than individual terms. For her, stewardship means that communities exercise authority over how their data is used, shared, and how any resulting benefits are distributed.
She explains that data should be understood as a collective asset and that data stewardship and community agency concern collective authority over data generated within community boundaries . She further says that stewardship is about collective decision-making regarding the use and sharing of data and the distribution of benefits within the community .
on: Collective or community-based governance is necessary because individual consent alone is inadequate for high-impact data and AI systems.
on: Whether data governance should be built primarily around collective/community authority or around self-sovereign individual control
Free, prior and informed consent and permanent community oversight mechanisms are needed for high-impact data systems - short name (Bridgette Ndlovu)
Arg. 3Bridgette argues that communities should be able to decide in advance whether they participate in data processes, especially in extractive contexts. She also supports durable oversight arrangements for powerful systems rather than one-off consultation exercises.
She identifies recognition of free, prior and informed consent as a key ask so that communities can determine whether they are willing to participate in data processes or governance arrangements . She also calls for mandatory transparency and accountability for major platform services and AI-related systems, indicating a need for ongoing scrutiny of high-impact systems .
on: Collective or community-based governance is necessary because individual consent alone is inadequate for high-impact data and AI systems.
Policy should recognise self-determination, prohibit extractive and unfair practices, require sector-specific codes of conduct, and mandate transparency and accountability for platforms and AI systems - short name (Bridgette Ndlovu)
Arg. 4Bridgette lays out a policy package for community-centred governance that combines rights, prohibitions, sectoral rules and transparency duties. The aim is to align data governance with self-determination and to curb extractive market behaviour while establishing operational standards.
She calls for recognising the right to self-determination in line with international standards so it is not defined arbitrarily by each country . She also asks for prohibitions on harmful and unfair market practices in extractive contexts, recognition of free, prior and informed consent, sector-specific codes of conduct to guide responsible data reuse, and mandatory transparency and accountability for key platform services including recommendation systems as governments adopt AI .
on: Practical governance tools are needed, including frameworks, codes, templates, transparency rules, and oversight mechanisms.
Existing governance remains trapped between centralised state control and unaccountable private power, so systems must connect individual and community sovereignty to enforceable rules - short name (Audience)
Arg. 1The audience speaker argues that current governance models are stuck between over-centralised state authority and overly decentralised or unaccountable private systems. They propose a model that protects individual and community sovereignty while still embedding compliance with legal and regulatory rules.
One audience participant argues that governance is still rooted in a Westphalian nation-state model that creates a disconnect between community-level realities and international governance . The same speaker says decentralised systems cannot simply become lawless spaces, so governance must connect individuals, families and communities to a compliance layer that respects existing laws and regulations rather than choosing between centralised systems and anarchic decentralisation .
on: Data governance should be community-centred rather than left primarily to states or corporations.
on: Whether public-interest algorithmic systems should be exempt from trade secret protections
Validity of data depends on who defines legitimate knowledge, so governance must account for plural knowledge systems rather than only dominant Western standards - short name (Audience)
Arg. 2The audience argues that data governance is inseparable from questions of epistemic authority. What counts as valid data will vary across medical and cultural traditions, so governance frameworks should not assume a single universal standard based only on dominant Western knowledge systems.
An audience participant says that communities do not agree on what counts as valid or proper knowledge, giving examples from Ayurveda, Chinese medicine and Western medicine . The speaker adds that traditional medicine systems were long dismissed by Western scientists, and warns that whoever validates data will shape whether certain forms of knowledge are recognised at all, including whether a Mayan healer and a New York doctor would assess health in the same way .
on: Representation is a major challenge, especially for Indigenous peoples, migrant workers, and other absent or unorganised groups.
on: Whether there can be a single legitimate steward or uniform community voice in the case study
AI in health relies on collective datasets, so community-level representation such as data cooperatives is more appropriate than relying only on individual consent - short name (Audience)
Arg. 3The audience speaker argues that healthcare AI systems need rich collective datasets and therefore cannot be governed adequately through individual consent alone. They suggest community representation structures, such as data cooperatives and representative councils, as more suitable mechanisms for collective decision-making.
Abhinav states that artificial intelligence does not operate meaningfully on isolated individual data points and that diagnostic tools in healthcare require high-quality datasets about whole communities . He therefore argues for the value of collectives and points to Indian experiences with data cooperatives and community-driven decision-making in indigenous and climate governance contexts, as well as collective understandings of data in Brazil, as possible models for representation across the full data lifecycle .
on: Representation is a major challenge, especially for Indigenous peoples, migrant workers, and other absent or unorganised groups.
on: Whether data governance should be built primarily around collective/community authority or around self-sovereign individual control
Internationally coordinated taxation and compensation mechanisms may be needed because aggregated global datasets create disproportionate profits for AI companies - short name (Audience)
Arg. 4The audience argues that the value of data comes largely from aggregation at scale, which allows major AI firms to extract outsized profits. Because this value creation is global, compensation and taxation measures should also be coordinated internationally rather than left only to local bargaining.
An audience member says that a single data point is worth little, but aggregated datasets become exponentially more valuable . On that basis, the speaker argues there is a need to tax AI companies for multiple purposes and to build an internationally coordinated taxation regime that could eventually compensate communities .
on: Benefit sharing from data should be fair, structured, and not treated as charity.
on: Whether fair benefit sharing should rely mainly on direct community negotiation and agreements or also on internationally coordinated taxation and compensation systems
Governments and regulators would benefit from template agreements and baseline rights frameworks to structure data-sharing deals more fairly - short name (Audience)
Arg. 5The audience argues that fairer data-sharing requires both substantive rights and practical tools. A baseline right to data at both individual and community levels, combined with standard agreement templates, would help governments negotiate complex deals more consistently and more protectively.
One audience participant suggests a fundamental right to data at both individual and community level so governments negotiating agreements must begin from those principles rather than treating data as ownerless . The same speaker also proposes two or three template forms for data-sharing agreements because such negotiations are complex and templates would make it easier for governments and others to structure them fairly .
on: Practical governance tools are needed, including frameworks, codes, templates, transparency rules, and oversight mechanisms.
on: Whether baseline fairness and governance should be defined through fixed frameworks and templates or through more flexible, evolving legal approaches
Effective baselines require governance frameworks that standardise how data is defined, gathered and combined across institutions - short name (Audience)
Arg. 6The audience speaker argues that a meaningful baseline depends on shared governance rules for data collection and integration. Standardised frameworks make it possible for different agencies and institutions to align and combine data consistently.
Angel Akunle says that to establish a baseline, a country first needs a data governance framework that defines the rules for producing baseline data . He explains that such a framework would let multiple agencies collect and later merge their data in the same way, making it possible to synchronise datasets seamlessly across institutions .
on: Practical governance tools are needed, including frameworks, codes, templates, transparency rules, and oversight mechanisms.
on: Whether baseline fairness and governance should be defined through fixed frameworks and templates or through more flexible, evolving legal approaches
Data-related legal frameworks must remain flexible because data use evolves in ways unlike older extractive-resource models - short name (Audience)
Arg. 7The audience argues that data cannot be regulated exactly like older extractive resources such as minerals, forests or timber. Because data uses are diffuse and can generate new applications over time, legal frameworks must be more adaptable and imaginative.
An audience participant observes that data is different from mineral extraction or forests, where compensation and legal structures were more clearly defined . The speaker adds that in the health-data example, uses could later expand into environmental or other domains, so legal frameworks must allow for change because data is not like "logs of wood" and requires more creative treatment .
on: The health AI case involves real development benefits but also serious governance risks, so decisions are ethically and politically difficult rather than straightforward.
on: Whether baseline fairness and governance should be defined through fixed frameworks and templates or through more flexible, evolving legal approaches
Genuine participation requires enforceable co-determination from design through operation and review, not token consultation - short name (Fernanda Campagnucci)
Arg. 1Fernanda argues that participation must give communities real decision-making power throughout the life of a project. Consultation alone is insufficient unless communities can shape the terms before implementation, monitor the system during operation, and trigger review or suspension if harms emerge.
She says communities must have enforceable power from inception to completion, and that participation is about co-determination rather than merely having a voice or a token seat at the table . She distinguishes genuine participation from consultation by explaining that consultation is often episodic, discretionary and non-binding, whereas real participation requires institutional arrangements enabling communities to shape data collection, use and sharing before implementation, monitor systems during operation, and seek review or suspension afterwards if harms appear .
on: Collective or community-based governance is necessary because individual consent alone is inadequate for high-impact data and AI systems.
States still have to play an enforcement role, because companies are unlikely to respect community processes without legal authority behind them - short name (Fernanda Campagnucci)
Arg. 2Fernanda accepts the need to move beyond simple government-versus-corporate models, but insists that states remain necessary for enforcement. Community participation and representation will not be effective unless companies are legally required to comply with the outcomes.
In response to discussion from the room, she says that even if governance moves away from state-versus-corporate models, some role for the state remains necessary to establish and legitimate new arrangements . She specifically argues that companies will comply with meaningful participation, broader consultation and validation processes only if some kind of enforcement exists behind them .
on: The health AI case involves real development benefits but also serious governance risks, so decisions are ethically and politically difficult rather than straightforward.
on: Whether there can be a single legitimate steward or uniform community voice in the case study
Public-interest systems should not be shielded by trade secrets when communities need visibility into how algorithmic systems operate - short name (Fernanda Campagnucci)
Arg. 3Fernanda argues that transparency obligations should override trade secret claims when algorithmic systems are used in ways that affect communities or public services. Her point is that public-interest systems require visibility and scrutiny, especially where governments invoke proprietary restrictions to avoid disclosure.
She cites Chile's repository of public algorithms as an example that advances transparency and open-source approaches to algorithmic systems . She then says governments often excuse non-disclosure by claiming algorithms are protected by trade secrets and commercial rules, but argues that when systems are public or affect communities, those trade secret protections should not apply .
on: Practical governance tools are needed, including frameworks, codes, templates, transparency rules, and oversight mechanisms.
on: Whether public-interest algorithmic systems should be exempt from trade secret protections
Session Knowledge Graph
Speakers · Topics · Arguments · Relationships
There was broad agreement that the core problem is not only privacy, but who holds power over data and how communities can exercise authority over it. Shumaila framed data governance as a question of ownership, benefit and risk rather than privacy alone . Bridgette argued that data should be understood as a collective asset and that communities, not corporations or states, should hold legal authority over it . Fernanda added that communities need enforceable co-determination throughout a project's lifecycle rather than symbolic consultation . An audience speaker similarly criticised the gap between centralised state models and unaccountable decentralised systems, arguing for arrangements that connect individual and community sovereignty to enforceable rules .
Data governance is about power, not only privacy - short name (Shumaila Shahani)
Communities should hold collective authority over data rather than corporations or states - short name (Bridgette Ndlovu)
Data should be treated as a collective asset, with stewardship defined through community decision-making over use, sharing and benefits - short name (Bridgette Ndlovu)
Genuine participation requires enforceable co-determination from design through operation and review, not token consultation - short name (Fernanda Campagnucci)
Existing governance remains trapped between centralised state control and unaccountable private power, so systems must connect individual and community sovereignty to enforceable rules - short name (Audience)
This aligns with people-centred and community-oriented governance approaches discussed in digital policy forums, including calls to bring governance closer to affected communities [S59], recognition of community data as a collective resource or data commons [S68], and examples of community-owned data collection and organising in humanitarian settings [S64].
Speakers converged on the view that governance of data-intensive and AI systems must be collective rather than reducible to isolated individual choices. Bridgette defined stewardship as collective authority over data, including decisions on use, sharing and benefits , and called for free, prior and informed consent alongside ongoing accountability for powerful systems . Fernanda argued that real participation requires institutional arrangements allowing communities to shape, monitor and review systems across their lifecycle . Abhinav from the audience reinforced this by arguing that health AI depends on high-quality community-level datasets and therefore cannot sensibly rely only on individual consent, pointing instead to data cooperatives and representative councils .
Data should be treated as a collective asset, with stewardship defined through community decision-making over use, sharing and benefits - short name (Bridgette Ndlovu)
Free, prior and informed consent and permanent community oversight mechanisms are needed for high-impact data systems - short name (Bridgette Ndlovu)
Genuine participation requires enforceable co-determination from design through operation and review, not token consultation - short name (Fernanda Campagnucci)
AI in health relies on collective datasets, so community-level representation such as data cooperatives is more appropriate than relying only on individual consent - short name (Audience)
This is supported by policy discussions noting that regulation focused mainly on individual rights is insufficient because important data decisions are also group decisions, creating a need for collective or group-rights approaches and stewardship institutions [S61]. Community data has likewise been framed as a collective resource governed in the public interest [S68].
There was clear agreement that community governance becomes difficult when the affected population is internally diverse or not equally organised. Shumaila stressed that the Indigenous group had its own council but was not separately consulted, while migrant workers had no representation at all and might be absent when decisions are made . Fernanda expanded this concern by noting that even where community governance is desirable, representation is hard when people are not organised at community level, especially in cities or among groups such as migrants . Audience interventions echoed this by asking who validates knowledge across different traditions and by proposing representative councils or data cooperatives to involve communities across the data lifecycle .
In the case study, stewardship is complicated because Indigenous groups and migrant workers are not equally represented, raising the question of who can legitimately speak for absent or unorganised groups - short name (Shumaila Shahani)
The ministry’s treatment of the valley as a single population ignores differentiated rights and needs, especially of the Indigenous council and migrant workers - short name (Shumaila Shahani)
States still have to play an enforcement role, because companies are unlikely to respect community processes without legal authority behind them - short name (Fernanda Campagnucci)
Validity of data depends on who defines legitimate knowledge, so governance must account for plural knowledge systems rather than only dominant Western standards - short name (Audience)
AI in health relies on collective datasets, so community-level representation such as data cooperatives is more appropriate than relying only on individual consent - short name (Audience)
Authoritative and policy discussions stress that meaningful participation and representation of marginalised groups are essential in data governance [S62], while recent IGF discussion highlights the particular sensitivity of Indigenous data and the need for stronger protections for small communities and minorities [S69].
Speakers agreed that when communities bear risks and supply data, they should receive a fair share of the value generated, and that such arrangements need clearer structures. Shumaila argued that communities' claims to value are rights rather than charity or corporate social responsibility , and said value includes not only money but infrastructure, training, insights and services . She also argued that because value can emerge years later, agreements should be long-term, transparent and renegotiable . Audience speakers complemented this by stressing that aggregated datasets create outsized profits for AI firms and may require taxation and compensation mechanisms , and by suggesting baseline rights and model agreement templates to make negotiations fairer .
Communities that bear the risks of data sharing should receive a fair share of the value created, and this is a right rather than charity or corporate social responsibility - short name (Shumaila Shahani)
Value from data is not only financial; it also includes infrastructure, training, access to insights and improved services, so fairness must be assessed across multiple forms of benefit - short name (Shumaila Shahani)
Because data gains value over time and through aggregation, communities need lifecycle agreements, renegotiation rights and transparency over realised value - short name (Shumaila Shahani)
Internationally coordinated taxation and compensation mechanisms may be needed because aggregated global datasets create disproportionate profits for AI companies - short name (Audience)
Governments and regulators would benefit from template agreements and baseline rights frameworks to structure data-sharing deals more fairly - short name (Audience)
This reflects development-oriented data governance frameworks that call for equitable distribution of benefits from the digital economy [S67] and UN-level commitments to consider sharing the benefits of data within interoperable governance arrangements [S63].
A strong area of agreement concerned the need to translate principles into concrete governance instruments. Bridgette called for recognition of self-determination, prohibitions on unfair extractive practices, sector-specific codes of conduct, and mandatory transparency and accountability for platform and AI systems . Fernanda argued that trade secret claims should not block visibility into public-interest or community-affecting algorithmic systems . Audience speakers proposed rights-based templates for agreements and national frameworks to standardise data collection and integration . Shumaila likewise insisted that offered services must be independently assessed against a defined baseline rather than accepted at face value .
Policy should recognise self-determination, prohibit extractive and unfair practices, require sector-specific codes of conduct, and mandate transparency and accountability for platforms and AI systems - short name (Bridgette Ndlovu)
Public-interest systems should not be shielded by trade secrets when communities need visibility into how algorithmic systems operate - short name (Fernanda Campagnucci)
Governments and regulators would benefit from template agreements and baseline rights frameworks to structure data-sharing deals more fairly - short name (Audience)
Effective baselines require governance frameworks that standardise how data is defined, gathered and combined across institutions - short name (Audience)
The adequacy of non-cash benefits such as free services and training cannot be assumed; it must be independently assessed against an agreed baseline - short name (Shumaila Shahani)
This is consistent with current multilateral policy direction calling for ethical frameworks, transparency, accountability, and inclusivity [S48], as well as concrete standards, classifications, auditing, and interoperable governance mechanisms in the Pact for the Future [S63]. Operational sandboxes and multi-stakeholder oversight have also been proposed as practical tools [S62].
There was shared recognition that the hypothetical health-data deal is not a simple case of rejecting extraction, because it also promises urgently needed health benefits. Shumaila emphasised that the valley has weak healthcare infrastructure, that better diagnostics are genuinely needed, and that refusing the deal has a real human cost . Audience speakers added that data is unlike older extractive resources because its uses can expand over time into new domains, requiring flexible legal treatment . Fernanda stressed that even if communities participate meaningfully, enforceable state-backed arrangements are still needed to make companies comply .
The proposed deal offers needed health improvements, but refusal also carries human costs, making community governance decisions ethically difficult - short name (Shumaila Shahani)
Data-related legal frameworks must remain flexible because data use evolves in ways unlike older extractive-resource models - short name (Audience)
States still have to play an enforcement role, because companies are unlikely to respect community processes without legal authority behind them - short name (Fernanda Campagnucci)
This matches established policy framing that health and AI governance requires balancing innovation and public benefit against privacy, ethical, and governance risks [S61]. Wider policy analysis also warns against overdependence on algorithms without critical human oversight in complex contexts [S48], and recent discussions emphasise that health-data use raises difficult questions of trust, selection, and benefit allocation [S62].
Both speakers framed data governance as a structural question about who controls data and who benefits from it, rather than a narrow privacy issue. Shumaila said the debate is about power, ownership, benefit and risk , while Bridgette argued that legal authority over data should rest with communities instead of corporations or states . These speakers all supported sustained community involvement in data governance rather than one-off consultation or purely individual consent. Bridgette called for free, prior and informed consent and ongoing accountability . Fernanda said communities must shape projects before implementation, monitor them during operation, and trigger review or suspension when harms emerge . Abhinav argued that because health AI depends on collective datasets, representative structures such as data cooperatives are more suitable than isolated individual consent . Shumaila and audience participants agreed that value generated from data is concentrated through aggregation and should be redistributed more fairly. Shumaila argued that communities who bear the risks should share in the value as a matter of right, and that agreements must remain open to renegotiation as value compounds over time . An audience speaker similarly said single data points have little value but aggregation produces exponential value, justifying taxation or compensation mechanisms for communities . These speakers shared a practical orientation towards enforceable governance tools. Fernanda argued that public-interest systems affecting communities must not be behind trade secret claims . Audience contributors proposed rights frameworks, templates, and standardised baseline frameworks for data collection and comparison . Shumaila similarly pressed for independent assessment of service offers against an agreed baseline . Both speakers highlighted that representation is politically complex and cannot be assumed. Shumaila pointed to the exclusion of the Indigenous council and the lack of any representation for migrant workers in the case study . Fernanda similarly noted that governance debates often presume organised communities, but many individuals are not represented through such structures, making it difficult to identify who speaks for them .
An unexpected convergence emerged between a speaker emphasising community empowerment and an audience speaker critical of centralised governance: both still accepted the need for enforceable legal structures. Fernanda explicitly said that companies will only respect meaningful participation if some enforcement exists and that states therefore still have a role . The audience speaker likewise rejected both centralised systems and anarchic decentralisation, calling for a compliance layer that adheres to existing laws and regulations .
Although benefit-sharing debates often focus on money, speakers also converged on the importance of assessing non-cash benefits through clear baselines. Shumaila stressed that value includes infrastructure, training, insights and improved services, and warned that these are not being adequately audited . An audience speaker then responded by arguing that baseline assessment requires governance frameworks that standardise how data is gathered and combined, showing agreement on the need to measure service adequacy rather than assume it .
A notable consensus emerged that, despite strong concern for rights and harms, health AI cannot be addressed through individualism alone because it depends on community-scale data. Shumaila presented the case as ethically difficult because better diagnostics could save lives . Bridgette had already framed data as a collective asset governed through community authority . Abhinav then made the technical argument that health AI requires community-level datasets, reinforcing the appropriateness of collective governance models .
The strongest areas of agreement were that data governance is fundamentally about power and community authority rather than privacy alone; that high-impact data and AI systems require meaningful collective participation and oversight; that representation of Indigenous peoples, migrants and other absent groups is a central governance challenge; that benefit sharing should be fair, transparent and structured; and that practical governance tools such as frameworks, templates, baselines, transparency duties and oversight bodies are necessary .
Bridgette frames the ideal as one in which communities, rather than corporations or states, hold legal authority over data and remain at the centre of decisions about it . By contrast, one audience speaker argues that governance must avoid both centralised systems and lawless decentralisation, and therefore needs a compliance layer that connects individuals, families and communities to existing legal and regulatory structures . Fernanda sharpens this by explicitly saying that, even if governance moves beyond a simple state-versus-corporate model, companies will only respect meaningful participation if there is some state-backed enforcement and legitimation behind it .
Communities should hold collective authority over data rather than corporations or states - short name (Bridgette Ndlovu)
Existing governance remains trapped between centralised state control and unaccountable private power, so systems must connect individual and community sovereignty to enforceable rules - short name (Audience)
States still have to play an enforcement role, because companies are unlikely to respect community processes without legal authority behind them - short name (Fernanda Campagnucci)
This tension reflects a long-running policy divide between socially anchored governance beyond the state and the need for public authority, oversight, and enforcement. Multi-stakeholder models emphasise collaboration rather than single-actor control [S51], while critiques of digital sovereignty warn against overconcentration of power in the state and call for socially driven governance arrangements [S60].
Bridgette argues that data is a collective asset and that stewardship should be defined through collective community authority over data use, sharing and benefits . Abhinav from the audience supports this in the health AI context, saying AI diagnostics require high-quality community-level datasets and therefore collective forms such as data cooperatives are more suitable than relying only on individual consent . However, another audience speaker proposes a model based on self-sovereign identity and self-sovereign data owned by each person, with rules applied to that individual ownership structure . Fernanda also introduces a practical challenge to purely community-based thinking by noting that many people, especially in cities, are not organised as communities, so representation of individuals remains unresolved .
Data should be treated as a collective asset, with stewardship defined through community decision-making over use, sharing and benefits - short name (Bridgette Ndlovu)
AI in health relies on collective datasets, so community-level representation such as data cooperatives is more appropriate than relying only on individual consent - short name (Audience)
States still have to play an enforcement role, because companies are unlikely to respect community processes without legal authority behind them - short name (Fernanda Campagnucci)
This disagreement maps onto an established policy contrast between collective governance of community data as a commons [S68] and individual-centred models such as digital self-determination [S57] and self-sovereign identity, which emphasise personal control, consent, portability, and discretion [S58].
Shumaila presents the valley as internally diverse and politically uneven, stressing that the Indigenous council was not separately consulted and migrant workers have no representation, which makes legitimate stewardship difficult . An audience participant deepens that challenge by arguing that communities do not even share a common view of what counts as valid knowledge, especially in health, and that whoever validates the data determines which knowledge systems are recognised . Fernanda adds that even when community rights are recognised, representation is still difficult because not all affected people are organised as communities, particularly urban individuals or absent groups, making a single voice or steward hard to identify .
In the case study, stewardship is complicated because Indigenous groups and migrant workers are not equally represented, raising the question of who can legitimately speak for absent or unorganised groups - short name (Shumaila Shahani)
Validity of data depends on who defines legitimate knowledge, so governance must account for plural knowledge systems rather than only dominant Western standards - short name (Audience)
States still have to play an enforcement role, because companies are unlikely to respect community processes without legal authority behind them - short name (Fernanda Campagnucci)
Relevant policy discussions caution against assuming one actor can manage data governance alone, stressing that there is no single road or silver bullet and that stakeholders have different roles [S51]. Representation challenges and the need for meaningful participation from contextually informed actors also suggest that a single uniform community voice may be difficult to establish [S62].
Shumaila argues that communities bearing the risks of data sharing have a right to value sharing, and proposes binding benefit-sharing agreements, public disclosure of agreements, community-controlled funds, lifecycle arrangements and renegotiation rights as value compounds over time . An audience member agrees on the extraction problem but argues that because single data points become valuable only through global aggregation, compensation cannot depend only on local bargaining; it may require internationally coordinated taxation of AI firms and redistribution mechanisms . The disagreement is not over fairness as a goal, but over the scale and mechanism through which it should be secured.
Communities that bear the risks of data sharing should receive a fair share of the value created, and this is a right rather than charity or corporate social responsibility - short name (Shumaila Shahani)
Because data gains value over time and through aggregation, communities need lifecycle agreements, renegotiation rights and transparency over realised value - short name (Shumaila Shahani)
Internationally coordinated taxation and compensation mechanisms may be needed because aggregated global datasets create disproportionate profits for AI companies - short name (Audience)
This is enriched by global policy debates on taxing the digital economy, including arguments for coordinated international tax rules, stakeholder engagement, and fair allocation of revenues [S52]. UN policy proposals also call for global tax transparency and information-sharing frameworks that benefit developing countries [S53].
Shumaila asks how the offered services in the hypothetical deal can be judged as enough or not enough, against what baseline, and how such a baseline should be created . Some audience interventions respond by calling for a fundamental right to data, template agreements for governments, and standardised frameworks to define and align baseline data across institutions . However, another audience speaker cautions that data is unlike older extractive resources, that future uses may shift across domains, and that legal frameworks must therefore remain adaptable rather than too fixed or static .
Governments and regulators would benefit from template agreements and baseline rights frameworks to structure data-sharing deals more fairly - short name (Audience)
Effective baselines require governance frameworks that standardise how data is defined, gathered and combined across institutions - short name (Audience)
The adequacy of non-cash benefits such as free services and training cannot be assumed; it must be independently assessed against an agreed baseline - short name (Shumaila Shahani)
Data-related legal frameworks must remain flexible because data use evolves in ways unlike older extractive-resource models - short name (Audience)
This reflects a recognised governance tension. On one hand, international frameworks call for standards, definitions, classifications, and interoperability [S63]. On the other, policy discussions emphasise adaptive tools such as sandboxes, incubators, and evolving responses to digital unknowns rather than rigid one-size-fits-all rules [S55].
Fernanda argues that in public systems or systems affecting communities, governments should not be allowed to hide algorithms behind trade secrets or commercial confidentiality, and she presents transparency and openness as necessary for scrutiny . By contrast, an audience speaker discussing tokenomic and ecosystem models still envisions benefit flows to model builders and firms within a broader architecture, suggesting a governance model that accommodates proprietary actors rather than simply removing such protections . The difference is limited but real: Fernanda prioritises public-interest transparency over commercial secrecy, while the audience framing is more accommodating of commercial participation provided benefit sharing is built in .
Public-interest systems should not be shielded by trade secrets when communities need visibility into how algorithmic systems operate - short name (Fernanda Campagnucci)
Existing governance remains trapped between centralised state control and unaccountable private power, so systems must connect individual and community sovereignty to enforceable rules - short name (Audience)
This sits within a well-developed trade and AI governance debate. Trade secret and source code protections are firmly embedded in trade law frameworks [S72], yet recent policy analysis argues these protections can obstruct accountability in sectors such as health and public services and suggests limiting or removing such barriers where public-interest oversight is needed [S70]. IGF 2025 discussions likewise raised possible exceptions for essential public systems [S69].
An unexpected tension emerged because the session was framed around ownership, stewardship, participation and benefits, but one audience intervention shifted the disagreement to epistemology: communities may not agree on what counts as valid medical knowledge in the first place, whether Western, Ayurvedic, Chinese or other systems . This complicates Shumaila's and Bridgette's focus on who should steward or decide over data, because the issue is not only who governs the data but also which knowledge systems the data governance framework validates .
A somewhat unexpected disagreement appeared within the audience contributions themselves. Some participants proposed templates, rights frameworks and standard baselines to make data-sharing negotiations fairer and more manageable . Another audience intervention warned that data is unlike timber or minerals, can generate new uses over time, and therefore requires more creative and adaptable legal treatment . This sits in tension with Shumaila's search for a baseline to judge adequacy, because it suggests that any baseline may need to remain revisable rather than fixed .
The discussion showed broad normative agreement that existing data governance is inadequate, that communities should have more power, that extractive models are problematic, and that the hypothetical health AI case raises serious issues of representation, consent and fair value sharing . The main disagreements concerned institutional design: how much authority communities versus states should have, whether governance should prioritise collective stewardship or self-sovereign individuals, how to represent absent or unorganised groups, and whether compensation should come mainly through contracts, oversight and auditing or through broader international taxation and redistribution .
All sides broadly agree that current data governance arrangements are inadequate and too concentrated in state or corporate hands. Shumaila frames the issue as one of power, ownership, benefits and risks rather than privacy alone . Bridgette argues communities should hold authority over data . Fernanda insists participation must be real and enforceable across the lifecycle of projects . Audience speakers similarly criticise the gap between community realities and current governance arrangements . They disagree, however, on the institutional route forward: stronger community authority, state-backed enforcement, or hybrid sovereignty models .
Data governance is about power, not only privacy - short name (Shumaila Shahani) Communities should hold collective authority over data rather than corporations or states - short name (Bridgette Ndlovu) Genuine participation requires enforceable co-determination from design through operation and review, not token consultation - short name (Fernanda Campagnucci) Existing governance remains trapped between centralised state control and unaccountable private power, so systems must connect individual and community sovereignty to enforceable rules - short name (Audience)
There is agreement that communities supplying data should not be left without a fair return. Shumaila argues this is a right and that value includes money, infrastructure, training, insights and improved services . An audience participant agrees that AI firms derive disproportionate value from aggregated data and therefore communities should be compensated . The disagreement concerns mechanism: Shumaila centres negotiated benefit-sharing agreements and auditing , while the audience stresses taxation and internationally coordinated redistribution .
Communities that bear the risks of data sharing should receive a fair share of the value created, and this is a right rather than charity or corporate social responsibility - short name (Shumaila Shahani) Value from data is not only financial; it also includes infrastructure, training, access to insights and improved services, so fairness must be assessed across multiple forms of benefit - short name (Shumaila Shahani) Internationally coordinated taxation and compensation mechanisms may be needed because aggregated global datasets create disproportionate profits for AI companies - short name (Audience)
These speakers agree that meaningful representation of affected groups is essential in the health AI case. Shumaila points to the failure to separately consult the Indigenous council and to represent migrant workers . Fernanda argues communities need enforceable power from design to review, not token consultation . Abhinav agrees that collective representation is needed throughout the data lifecycle and suggests representative councils or data cooperatives . They differ mainly on the precise representational form and on how to handle people who are absent or not organised .
In the case study, stewardship is complicated because Indigenous groups and migrant workers are not equally represented, raising the question of who can legitimately speak for absent or unorganised groups - short name (Shumaila Shahani) Genuine participation requires enforceable co-determination from design through operation and review, not token consultation - short name (Fernanda Campagnucci) AI in health relies on collective datasets, so community-level representation such as data cooperatives is more appropriate than relying only on individual consent - short name (Audience)
There is a shared goal of making data-sharing deals assessable and less arbitrary. Shumaila questions how to know whether promised services are sufficient and asks for a baseline to judge adequacy . Audience speakers respond by calling for baseline rights, template agreements and governance frameworks that standardise data collection and integration across agencies . The remaining difference is that audience speakers emphasise pre-defined frameworks and standardisation, whereas Shumaila poses the baseline question more open-endedly in relation to fairness and independent assessment .
The adequacy of non-cash benefits such as free services and training cannot be assumed; it must be independently assessed against an agreed baseline - short name (Shumaila Shahani) Governments and regulators would benefit from template agreements and baseline rights frameworks to structure data-sharing deals more fairly - short name (Audience) Effective baselines require governance frameworks that standardise how data is defined, gathered and combined across institutions - short name (Audience)
- The discussion framed data governance as a question of power, not only privacy, focusing on who controls data, who benefits from it, and who bears the risks.
- A central conclusion was that community-centric data governance should offer an alternative to both state-centric control and extractive corporate ownership, with communities treated as core decision-makers over data generated within their boundaries.
- Participants emphasised that data should often be understood as a collective asset rather than solely an individual commodity, especially in contexts such as health AI where models rely on aggregated community-level datasets.
- Legitimate data stewardship remains complex: the discussion highlighted that Indigenous groups, migrant workers, and other absent or unorganised populations may not be adequately represented by state institutions or by a single community body.
- The discussion stressed that governance must account for plural knowledge systems, since what counts as valid data or legitimate expertise may differ across medical, cultural, and community traditions.
- There was broad agreement that meaningful participation must go beyond consultation and include enforceable co-determination throughout the lifecycle of a project, from design to deployment, monitoring, review, and possible suspension.
- Free, prior and informed consent was highlighted as a key safeguard, particularly for identifiable communities affected by high-impact data systems.
- Participants noted that states still have an essential enforcement role, because companies are unlikely to honour community processes without legal authority and regulatory backing.
- Benefit sharing was presented as a right rather than charity: communities that provide data and bear associated risks should receive a fair share of the value created from that data.
- The discussion concluded that value from data is multi-dimensional, including money, infrastructure, training, access to insights, and improved services, and that the adequacy of these benefits must be independently assessed.
- Because data can generate value long after collection and through aggregation across populations, communities should have lifecycle-based agreements, transparency over realised value, and the ability to renegotiate terms over time.
- Existing and emerging policy tools were identified as useful building blocks, including data trusts, community data commons, sector-specific codes of conduct, representative bodies, transparency requirements, and benefit-sharing agreements.
- Participants argued that public-interest AI and data systems should not be shielded by trade secret claims where such secrecy prevents accountability to affected communities.
- The hypothetical health AI case illustrated the ethical difficulty of governance decisions: accepting the deal could improve healthcare access, but the proposed arrangement concentrated rights and profits abroad while offering uncertain and non-cash local benefits.
- A key lesson from the case study was that communities should not be treated as homogeneous; governance arrangements must reflect differentiated rights, needs, and representation for Indigenous councils, migrant workers, and other distinct groups.
- The discussion also highlighted that data governance frameworks need flexibility, because data extraction and reuse evolve differently from older extractive-resource models such as mining or forestry.
“An audience member argued that what counts as ‘valid knowledge’ is not universally agreed, especially in health: a Mayan healer, Ayurvedic practitioner, Chinese medicine practitioner and a doctor in New York may define health and evidence differently, so whoever validates data also determines what is treated as legitimate.”
“Another audience member proposed three linked ideas: a fundamental right to data at both individual and community level, practical templates for data-sharing agreements, and an internationally coordinated taxation regime for AI companies because aggregated data becomes exponentially more valuable than single data points.”
“A participant reflected that current governance remains trapped in a post-Westphalian nation-state model, while digital systems are global; they argued for governance that maximises individual sovereignty, links individuals, families and communities directly into the system, preserves interoperability, and could use tokenomic mechanisms to route benefits back to data contributors.”
“Fernanda Campagnucci responded that even if the goal is to move beyond government-versus-corporate models, states are still needed to legitimate and enforce community decisions against companies; she also raised a second challenge: unlike organised indigenous or traditional groups, many data subjects are ordinary individuals in cities who are not organised in community form, so who represents them remains unresolved.”
“Abhinav argued that AI, especially in digital health, does not function meaningfully on isolated individual data points but requires high-quality collective datasets; therefore, community-level governance is not only normatively desirable but technically necessary. He added that data cooperatives and representative councils could play a role at each step of collection, training and use.”
“Angel Akunle suggested that to decide whether the services offered in exchange for data are ‘enough’, there must first be a data governance framework establishing common rules and a baseline, so that agencies and actors can align data practices and compare value consistently.”
“A final audience intervention stressed that data is unlike timber, minerals or forests: communities around data are more amorphous, data can generate unforeseen downstream uses such as environmental implications, and legal frameworks therefore need to be flexible and more creative than past extractive-compensation models.”
Who is the legitimate steward of the Valley’s data, and can there be more than one legitimate steward?
This is central to community-centric data governance because the case involves multiple affected groups, including the general Valley population, an Indigenous council and migrant workers. Determining stewardship affects authority over consent, access, control and benefit-sharing.
How should different knowledge systems be recognised when validating health data and deciding what counts as ‘proper’ or legitimate knowledge?
The participant highlighted that Western medicine, Ayurveda, Chinese medicine and Indigenous healing traditions may define valid knowledge differently. This matters because data governance standards and AI systems may exclude or devalue non-Western knowledge systems, shaping whose data is recognised and how health interventions are designed.
Should there be a fundamental right to data at both individual and community level?
A rights-based foundation could constrain how governments and companies negotiate data-sharing arrangements. It is important because without recognised rights, communities may have little legal basis to challenge extractive agreements or demand meaningful control.
What model templates or standard forms should exist for data-sharing agreements so that governments and communities can negotiate fairer terms?
The participant noted that negotiating such agreements is complex. Research into model clauses or templates could help less-resourced governments and communities benchmark fair terms, including consent, benefit-sharing, redress and governance obligations.
What internationally coordinated taxation regime for AI companies could compensate communities whose aggregated data creates commercial value?
The participant argued that individual data points gain value through large-scale aggregation by AI firms. Research is needed because taxation could become a mechanism for redistributing value, addressing safety, knowledge and energy externalities, and compensating source communities.
Where can more information be found about international standards for self-determination of data?
This is a direct request for further material on relevant international norms. It is important because self-determination standards could guide policy development, especially for communities engaging with digital identity, mobile networks and social safety net infrastructures.
How would genuine participation look in the Valley case, and what concrete mechanisms could ensure communities shape the project before, during and after deployment?
This question goes to the heart of participatory decision-making. It is important because token consultation is insufficient; effective mechanisms are needed for co-determination, monitoring, review and suspension when harms emerge.
Can existing multi-stakeholder committees or oversight bodies in areas such as education, environment or public policy be adapted to govern data and AI systems?
This suggests a practical research avenue into institutional design. It matters because building entirely new structures may be slow or politically difficult, while adapting existing bodies could offer a feasible route to oversight.
How can governance systems connect individual, family and community-level sovereignty to broader legal and technical frameworks without falling into either centralised control or decentralised anarchy?
The participant raised a structural challenge about designing governance that preserves individual sovereignty while remaining lawful and interoperable. This is important for future models of digital commons, self-sovereign identity and community-centred infrastructures.
What architecture would enable benefit-sharing to flow back to those who provide data when AI models built from that data are monetised?
The participant noted that such architecture does not yet exist. This is important because benefit-sharing is a core principle of the discussion, and practical financial and technical mechanisms are needed to make it operational rather than rhetorical.
What role must the state play in enforcing community decisions and ensuring companies comply with community-centred governance models?
Fernanda suggested that even alternative governance models still require state legitimacy and enforcement. This is important because without enforceability, community participation may have no effect on corporate behaviour.
Who represents individuals who are not organised in identifiable communities, especially in cities or dispersed populations?
Fernanda pointed to a major representational gap in data governance. This is important because many people affected by data systems are not part of formal communities, yet their interests still need protection and representation.
How should absent or marginalised groups, such as Indigenous councils not separately consulted and migrant workers with no representation, be represented in data governance decisions?
This issue was posed directly by Bridgette and expanded by Abhinav. It is important because exclusions in representation can make consent invalid and deepen structural injustice, especially in high-stakes health data systems.
How can community involvement be built into every stage of the AI data lifecycle, from collection to model training to deployment and review?
Abhinav proposed examining the role of communities across the full process rather than at a single consultation point. This is important because harms and value extraction can occur at multiple stages, and governance must cover the entire lifecycle.
How should collective data be governed differently from individual data, especially where AI systems rely on high-quality community-level datasets rather than isolated individual data points?
Abhinav argued that AI diagnostics depend on collective datasets, making purely individual consent frameworks inadequate. This is important because it implies the need for legal and policy models tailored to collective data rights and governance.
What can be learned from data cooperatives and other community decision-making models, including examples from India and Brazil, for governing collective data?
This is a concrete comparative research avenue. It is important because existing cooperative and community-rights models may offer institutional lessons for representation, consent and governance in data-intensive sectors.
How do we know whether the non-monetary services offered in exchange for data are adequate, and against what baseline should adequacy be assessed?
This question addresses fairness in benefit-sharing. It is important because the Valley is offered services rather than cash, and without a baseline or independent assessment, there is no way to judge whether the exchange is just.
What kind of data governance framework is needed to establish a baseline for evaluating and harmonising data across agencies or institutions?
Angel argued that a common framework is necessary to define how baseline data is established and synchronised. This is important because interoperable governance standards affect data quality, comparability and the ability to oversee complex systems.
How should legal frameworks for data governance remain flexible enough to address the changing, cross-sectoral and amorphous nature of data compared with extractive resources such as minerals or forests?
The participant stressed that data differs fundamentally from physical resources and can generate unexpected downstream uses. This is important because rigid legal analogies may fail to capture the fluidity of data, community boundaries and evolving harms or benefits.
