OECD warns of growing shadow AI in government

Shadow AI exposes gaps in public-sector safeguards as civil servants adopt GenAI without consistent approval or monitoring.

OECD logo over a blurred government technology centre, illustrating concerns about Shadow AI spreading across public administration

The OECD has warned that ‘shadow AI’ is spreading across public administrations as civil servants adopt generative AI without formal approval, consistent oversight or clear operational rules.

A review of official guidance from 14 countries found that while governments have widely adopted AI ethics principles, practical guidance for designing, monitoring, and evaluating generative AI experiments remains fragmented. The result is inconsistent practices, duplicated pilot projects and excessive risk aversion.

The OECD argues that structured experimentation should bridge the gap between national AI strategies and day-to-day public administration. Governments should begin with controlled, low-risk applications before extending AI to sensitive services or decisions affecting citizens.

Evaluation remains a major weakness. Many pilot projects rely on basic indicators such as usage rates and user satisfaction rather than systematically assessing output quality, costs, institutional impact or regulatory compliance.

The OECD therefore recommends evaluating experiments across five dimensions: performance, public value, feasibility, usability and risk management. Because generative AI can produce different outputs from identical prompts, continuous monitoring may be more appropriate than one-off testing.

The framework presented in the report shows that trustworthy adoption depends on three connected elements: enablers such as skills, data and infrastructure; guardrails including legislation, oversight and ethical principles; and engagement with public servants, citizens and external partners.

To contain shadow AI, the OECD recommends shared tools, secure testing environments, stronger AI literacy, clearer accountability and better coordination across projects. Experiments and evaluation methods should be designed together from the outset. Without these measures, governments risk creating isolated pilot projects that cannot be safely scaled, compared or discontinued when they fail.

Why does it matter?

The rapid spread of shadow AI suggests that public servants are adopting generative AI faster than governments can establish formal governance processes. Without common standards for testing, evaluation and oversight, informal AI use could undermine accountability, expose sensitive information and create inconsistent public services.

The OECD’s framework reflects a broader shift in AI governance from developing high-level principles to establishing practical methods for responsible deployment. As governments move from experimentation to implementation, structured evaluation may become as important as regulation itself.

Would you like to learn more about AI, tech, and digital diplomacy? If so, ask our Diplo chatbot!