Microsoft gives communities greater control over AI image training

Community-built image libraries aim to fix gaps in AI training data for Microsoft’s tool.

Microsoft's new tool lets communities help shape how AI systems represent them.

Microsoft has developed a research tool that allows communities to help shape how they are represented in AI-generated images, addressing long-standing concerns that AI training data often reflects incomplete or biased online content rather than the perspectives of the people depicted.

Instead of relying solely on existing online datasets, communities, including people with disabilities and others with shared lived experiences, can work through advocacy organisations to create their own image and video libraries. Each image is accompanied by descriptions explaining what is important about the scene, giving AI systems contextual information based on the community’s own perspective rather than visual appearance alone.

Developed by Microsoft’s Accessibility Team, the process guides participants from selecting meaningful images to curating a library of around 400 real-world examples organised around shared themes such as family life and everyday routines. These annotated collections are then used to generate AI images, which community members review and evaluate against their own standards of accurate representation, creating a feedback loop intended to improve future outputs.

A central feature of the project is that advocacy organisations retain ownership of the resulting datasets and decide whether, how and with whom they are shared, including publication on platforms such as Hugging Face. Communities can also withdraw their data if consent changes over time.

Why does it matter?

The project addresses one of the central challenges in AI governance: who decides how people and communities are represented in training data. Many AI systems are built using large-scale datasets collected from the internet with limited transparency, consent or community involvement, increasing the risk of inaccurate, stereotypical or exclusionary outputs.

By giving communities ownership of their data and a direct role in defining what constitutes fair representation, Microsoft’s approach shifts part of the AI development process towards participatory governance. If adopted more widely, it could provide a model for building datasets that are not only more representative but also more transparent, accountable and consent-based.

Would you like to learn more about AI, tech and digital diplomacy? If so, ask our Diplo chatbot