Atelier 4: La numérisation des contenus inexploités: stratégies et outils
The discussion focused on transforming “dormant” content into useful, reliable, and well-governed digital resources for artificial intelligence, while preserving authors’ rights, data quality, and the public interest . The moderator also framed the exchange around the experience of heritage digitization and the governance and copyright issues related to content used by AI .
Speaker 1 argued that immense volumes of heritage content still remain inaccessible online, even though AI models depend not only on the quantity but also on the quality of data, with risks of cultural and linguistic bias if only data already dominant on the web are used . He presented this as a dual issue of sovereignty and representativeness for the Francophonie, noting that the Royal Library of Belgium has digitized only about 10% of its collections and that, within the digital Francophone network, rates often range from just a few percent to 25% at most . He stressed that heritage digitization is not just about producing images, but about creating contextualized, reliable heritage data that can be used by machines, notably through high-quality full text and collections transformed into structured data . In his view, this justifies public investment, harmonization of practices, consideration of copyright, and a collective increase in digital maturity within the Francophone network .
Speaker 2 complemented this approach with an intellectual property perspective, recalling that AI systems need both older heritage data and contemporary content that is often protected by copyright . He explained that WIPO promotes a balanced framework between public access and respect for rights, notably through guides on the preservation of and access to works in heritage institutions . He emphasized that copyright exceptions are possible, but only within the limits of the “three-step test,” and that a distinction must be made between public access, text and data mining, and use as training data for AI . He also stressed transparency of uses, value-sharing, licensing, public-private partnerships, and the need for a robust semantic infrastructure with metadata and identifiers to attribute works and manage rights .
Questions from the audience broadened the debate to oral knowledge, particularly in African contexts, highlighting the need to think about archiving, costs, sovereignty criteria, and the risk that some societies might become mere providers of knowledge exploited by others . In response, the speakers emphasized the trust of communities, the need to frame future uses of local knowledge, and the fact that even seemingly marginal content can have high value for AI . Asked about priorities, Speaker 1 proposed starting with a better inventory of collections, then giving priority to fragile or threatened heritage-notably magnetic tapes, the press, audio sources, and unique documents-while multiplying access channels and interoperability . In conclusion, the discussion converged on the idea that a digitization strategy for AI must combine preservation, reliability, cultural diversity, trust, funding, common standards, and rights governance so that heritage becomes a resource for the future rather than a blind spot of the digital age .
- The debate first concerns the transformation of “dormant” content into useful resources for AI, with a balance between usefulness, reliability, governance, copyright, and the public interest. The moderator explicitly raises this central question at the outset, and the speakers then return to it from the perspective of heritage digitization and machine use.
- A major point is the urgency of digitizing heritage that is still absent from the web, because a large share of collections remains outside the digital sphere, creating a problem of sovereignty, cultural and linguistic representativeness, and bias in AI systems. Speaker 1 stresses that the data available online are limited, that heritage collections constitute a reservoir still largely untapped, and that their absence makes certain cultures and languages invisible in future models.
- The discussion underscores that digitization is not enough in itself: content must be reliable, contextualized, structured, interoperable, and converted into data that machines can genuinely use. Speaker 1 explains that heritage digitization involves choices about quality, control, metadata, and the production of full text or structured data (“collections as data”), without which content remains difficult to use for AI.
- Another central focus concerns the legal framework for works protected by copyright, particularly twentieth-century content and uses by AI. Speaker 2 recalls the need to reconcile access, preservation, and respect for rights, relying on copyright limitations and exceptions, the “three-step test,” licenses, public-private partnerships, and remuneration mechanisms.
- Lastly, participants stress the need for a collective strategy based on trust, consideration of oral traditions, clarification of digitization priorities, infrastructure, funding, and standardization. The discussion highlights African and Indigenous oral knowledge, the risks of value capture, the fragility of certain media, as well as the need for robust metadata, precise inventories, common standards, and coordinated public policies.
- Overall objective of the discussion:
- This discussion aims to define how to digitize and govern heritage and cultural content-sometimes still analog or oral-so that it becomes a reliable, visible resource that can be used by artificial intelligence, while respecting copyright, the communities concerned, cultural diversity, and the requirements of information sovereignty.
- General tone:
- The overall tone is serious, expert, and constructive. It is initially pedagogical and strategic, with rich presentations by the two main speakers on the technical, cultural, and legal issues. It then becomes more interactive and critical during questions from the audience, particularly around oral traditions, risks of exclusion, costs, sovereignty, and the economic capture of content. Despite these tensions and questions, the discussion returns to a pragmatic and action-oriented tone, focused on priorities, trust, standards, and collective strategies.
The workshop opens with a central question: how can still “dormant” content be transformed into useful, reliable, and well-governed resources for artificial intelligence, without sacrificing authors’ rights, data quality, or the public interest . The moderator immediately clarifies that the discussion is intended to be more open than a traditional panel, and announces two complementary entry points: the experience of heritage digitization presented by Speaker 1, and the issues of content governance and copyright related to AI presented by Speaker 2 . He also reformulates a key point very early on: in this context, digitization only makes sense if it produces usable and interoperable content, not merely digital copies .
In his opening remarks, Speaker 1 returns to an idea raised earlier in the day: the harvesting of data available on the web is approaching a form of saturation, making the vast amounts of content still absent from the internet all the more visible . He emphasizes that behind every AI system, the decisive issue remains data, both in quantity and quality . In his view, the imbalances already present online - between languages, cultures, viewpoints, or historical periods - risk being reproduced in the outputs of AI systems . From this perspective, the Francophone world faces a challenge of both sovereignty and representativeness, but it also has particular potential because of its cultural and linguistic diversity .
Speaker 1 supports this observation with several figures. At the Royal Library of Belgium, he estimates that at most 10% of content has now been digitized and made available online, without this yet meaning that it is directly usable by AI . In the Réseau francophone numérique, which brings together around one hundred heritage institutions from across the Francophone world, digitization rates often range from a few percent to a maximum of 25% . In other words, very large swathes of knowledge and culture remain outside the digital sphere and are therefore much less likely to feed AI systems . For Speaker 1, digitization is thus also a matter of public policy choice, since these reserves of knowledge are held in heritage institutions that are largely publicly funded .
His central point, however, is that digitizing is not enough. He distinguishes between the simple production of a digital facsimile and the creation of genuine heritage data . The latter must be contextualized, reliable, and structured if the goal is to avoid producing content that is of little use or itself carries bias . He describes heritage digitization as a complete professional chain: harmonization of inventories, revision of old bibliographic data, consideration of the legal status of works, technical capture, quality control, structuring, and dissemination . He also stresses the place of human intervention in this process, comparing digitization staff to “knowledge brokers” or “modern-day copyists” .
Speaker 1 adds that having beautiful images or basic metadata is not enough either . For content to be truly useful to machines, high-quality full text must be produced and collections must be made readable by automated systems, including for printed, manuscript, iconographic, sound, and audiovisual corpora . It is in this context that he refers to the idea of “collections as data” . In his view, this shift leads heritage data to be seen no longer merely as memory of the past, but also as a strategic asset and a raw material for the future . He therefore calls for public investment in digitization to continue and expand so that this content can be made genuinely usable in the age of AI .
He also presents the Réseau francophone numérique, created in 2006 at the time Google Books emerged, as a tool for solidarity and for building digital maturity . He sums up this orientation with a clear formula: “there is no AI without digitization” . To reach a level at which heritage data can support advanced uses, institutions must master the entire chain: inventory, cataloguing, legal status, capture, structuring, dissemination, and interoperability . The network thus seeks to harmonize practices and develop training, in a context where members’ capacities remain highly unequal .
Speaker 2 then addresses the issue from the standpoint of copyright and governance of uses. He recalls that AI systems need high-quality and culturally relevant data, but also contemporary content that is often still protected, not only older works that have entered the public domain . His remarks therefore focus in particular on twentieth-century works still existing in analog form - books, phonograms, films on magnetic tape - and on the conditions for their digitization within a balanced framework between public access and respect for rights . In this regard, he presents the role of WIPO, which supports Member States through guides to good practice on the preservation of and access to works in heritage collections .
At the heart of his intervention is the issue of limitations and exceptions to copyright. Speaker 2 recalls that international law allows States to provide exceptions to the exclusive right of reproduction, but only in compliance with the three-step test: certain special cases, no conflict with the normal exploitation of the work, and no undue prejudice to the legitimate interests of authors . He then stresses three distinctions. Providing access to researchers is not equivalent to publishing on the internet; enabling text search with excerpts is not equivalent to providing full access; and digitizing a work does not automatically authorize its use for training an AI . He cites the example of Google Books, where the digitization of protected works was accepted in the United States for text search and previews, but not for full access .
Speaker 2 adds that “making accessible” does not necessarily mean “free of charge” . He highlights the growing number of licensing agreements between content holders and AI companies, showing that a market exists and that free access cannot be presumed . He also mentions that certain public policy choices may lead to the opening of public data, while some rights holders may prefer open licences . He also draws attention to the potential value of content that is not widely commercialized, including in rare, Indigenous, or minority languages, for which companies are already seeking to build specialized datasets, especially audio and multilingual ones .
From there, Speaker 2 sets out several guiding messages. First, any digitization strategy must integrate copyright issues from the outset, as well as folklore and intangible heritage, particularly when it involves knowledge held by local communities or Indigenous peoples . Next, he stresses transparency regarding the objectives being pursued: preservation, access, text search, or the training of AI systems . He also mentions, in the European context, the existence of an “opt-out clause” allowing opposition to certain uses by AI systems to be signaled, with reference to Article 4 of the Directive on copyright in the digital single market . He also calls for explicit reflection on value-sharing among creators, heritage institutions, and AI companies, through licences, collective management, or public-private partnerships . On this point, he does not advocate a single model: several arrangements are possible depending on whether an institution wishes to transfer certain rights, retain copies, or pursue other dissemination and valorization objectives . Finally, he emphasizes the need for a robust rights management infrastructure, based on precise metadata and standardized identifiers such as ISBN, ISSN, ISAN, ISRC, or ISWC .
The discussion with the audience then shifts the debate toward four concrete questions: how to integrate orality, how to archive even before digitization, how to avoid cultural extractivism, and how to define operational and fundable priorities . Speaker 3 refocuses the discussion on oral knowledge, particularly in several African societies where information is not primarily preserved in structured documentary form . He emphasizes that a strategy based solely on libraries and already organized archives would risk excluding these forms of knowledge . He also stresses a step that is often overlooked: before digitization itself, thought must be given to the primary archiving of oral knowledge collected in the field, its preservation, its classification, and the cost of this transition phase . He also recalls the weight of material constraints: the cost of digitization, the need for servers and storage, the fragility of energy infrastructure, and the need to design sustainable strategies . In the same intervention, he warns against a risk of cultural extractivism: without a strategy for local valorization, some regions could provide AI with cultural raw material while allowing the capture of the economic value produced to take place elsewhere .
Speaker 4 intervenes in a more conceptual and skeptical register. Through the imaginary figure of a “Martian” coming in a thousand years to write the history of Earth, he recalls the potential scale of unexploited content: telephone conversations, receipts, letters, private images, unsold stocks of books, or various videos . His point is above all to highlight the gap between the ideal of exhaustiveness and the human, legal, moral, and institutional constraints involved . Speaker 1 responds on a pragmatic level: generative AI has already entered everyday use, and the issue is no longer whether or not to take part in this movement, but to ensure that it relies on data that are as reliable, complete, and diverse as possible .
The question of trust and consent then becomes an important thread in the discussion. The moderator introduces this point with the example of a Kanak religious song recorded in a non-profit context in the 1980s and then commercially reused much later, without the originating community being credited . Speaker 2 uses this example to underscore the need to obtain communities’ consent and to define the conditions of use in advance . He also recalls that AI systems are highly diverse and that even apparently marginal content can acquire strong economic value in specialized uses . His example of “helicopter theft” illustrates the fact that an AI producer may need very specific images, collected and annotated from television news, documentaries, or films . He also cites the French ReLIRE database to show that out-of-print twentieth-century works, sometimes niche ones, may retain real value for specialized audiences . Speaker 1, for his part, responds mainly on the strategic level: since AI use is already here, it is necessary to ensure that systems rely on better-quality data .
Several interventions from the audience then anchor these issues in African and national concerns. Speaker 5, Senegal’s Director of Digital Affairs, explains that the absence of digitization of oral transmission could lead to distorted reproduction, or even falsification, of national historical or heroic references in AI systems . In her view, these forms of heritage must therefore be digitized, structured, and made reliable, even if that comes at a cost, so that each society retains control over what culturally belongs to it . She also links this strengthening of reliability to intellectual property, seen as a way to secure content and its attribution at an earlier stage .
The moderator then refocuses the discussion on the Francophone space by asking which prioritization criteria should guide digitization efforts, and whether the aim should be a common Francophone corpus or rather a federation of commons despite the diversity of national and institutional rules . In response, Speaker 1 proposes a fairly clear hierarchy of action. The first priority is to improve or redo inventories so as to know precisely what is being preserved and what is at risk of disappearing . The second is to deal urgently with fragile or endangered heritage, in particular magnetic tapes, sound sources, newspapers printed on acidic paper, manuscripts, and unique documents . He also argues for multiplying access channels in order to avoid dependence on a single intermediary and increase the discoverability of content . He illustrates this with Belgian examples, explaining that certain documents held in Brussels have little chance of being found if they are visible only on a national portal and therefore need to circulate through several specialized or thematic portals .
Speaker 1 also returns to the institutional conditions for this policy. He stresses the need for libraries, archives, and museums to know one another better, understand their respective realities, and collectively build trust, because there can be no AI without that trust . He calls for an open digital ecosystem that is useful both to researchers, students, the general public, and machines . This requires cross-cutting training, better command of public procurement, shared standards, suitable equipment, digital libraries, and aggregator portals .
Asked whether priorities have already been identified and whether concrete action plans exist, Speaker 1 replies that AI has at least had the merit of bringing digitization back to the forefront of policy, whereas some thought this work was already complete . He reaffirms that priority must go to endangered heritage: magnetic tapes, newspapers, sound sources, recorded intangible heritage, manuscripts, and unique documents . He notes, however, that in several Francophone contexts, existing public policies are still aimed primarily at regulating the use of AI in the public sector, rather than refinancing the entire documentary chain needed to build genuinely usable corpora .
In his closing remarks, Speaker 2 summarizes three priorities: building trust with rights holders and the communities carrying oral traditions, ensuring transparency of AI systems regarding the content they use, and putting in place a robust semantic infrastructure for attribution and rights management . He had also specified that conditions of use, digitization objectives, and any restrictions could be recorded in the semantic layer describing digitized objects, so that they are visible to machines as well .
Overall, the workshop reveals broad agreement on several points. The speakers consider that the digitization of content still offline has become an important condition for AI that is more useful, more reliable, and more representative . They also agree that simple digital conversion is not enough: content must be structured, contextualized, interoperable, accompanied by full text, rich metadata, and explicit governance of rights and uses . Finally, several interventions converge on the importance of integrating oral knowledge, fragile heritage, and underrepresented languages, particularly in the Francophone and African space .
The tensions concern less the objectives than the trade-offs. A first sets the ambition to broaden corpora against recognition of the practical, legal, and material limits of any digitization policy . A second concerns the risk of cultural extractivism and the issue of value-sharing . A third relates to method: whether to start from institutional heritage corpora and their standards, or to rethink strategies on the basis of oral knowledge, prior archiving, and community consent .
The conclusion that emerges from the workshop is therefore pragmatic. The point is not to digitize everything indiscriminately, but to build corpora that are more reliable, more representative, and better governed, based on precise inventories, clearly assumed priorities, and strengthened institutional capacities . The clearest avenues for action concern prioritizing fragile media, harmonizing standards and practices, public investment, early integration of copyright and the communities concerned, as well as putting in place robust semantic layers to make content genuinely usable by machines . Several questions nevertheless remain open: sustainable financing, common criteria for prioritization in the Francophone space, governance of a possible common corpus, fair value-sharing, and how to reconcile access, cultural sovereignty, community consent, and the commercial exploitation of data .
La base confirme explicitement que « l'IA est construite sur des données », et que la gouvernance des données et celle de l’IA vont de pair [S19].
Cette idée est corroborée par [S29], qui rappelle que les données utilisées par l’IA sont majoritairement anglophones et que les plateformes recommandent souvent des contenus anglophones, ce qui pose un défi direct pour la diversité linguistique francophone en ligne [S29].
La base apporte un contexte utile en montrant que la concentration du pouvoir numérique et l’inégalité numérique sont identifiées comme des risques majeurs, en particulier pour les régions moins favorisées, ce qui renforce l’idée d’un enjeu de souveraineté informationnelle et culturelle [S22].
La base n’aborde pas directement les institutions patrimoniales, mais elle confirme plus largement que l’usage accru des données par les organisations et pouvoirs publics doit servir des politiques fondées sur des preuves, tandis que la gouvernance des données devient plus politique [S19]. Cela donne du poids à la dimension de politique publique évoquée dans le rapport.
La base confirme que l’essor d’outils comme ChatGPT a fait émerger de nouvelles questions de gouvernance pratique et de protection des droits de propriété intellectuelle [S19].
Dormant content to be made usable for AI, with requirements for reliability, governance, and respect for rights (Moderator)
Arg. 1The moderator immediately sets the framework for the discussion: the issue is not merely digitizing content, but turning dormant content into resources that are genuinely useful for AI. He stresses the need to maintain a balance among usefulness, reliability, content governance, data quality, the public interest, and respect for the rights of authors and creators.
The moderator explicitly frames the workshop’s central question as the transformation of dormant content into resources that are « useful, reliable, and governed » for artificial intelligence, without losing sight of authors’ rights, data quality, and the public interest .
on: Digitization must be aligned with respect for rights, the trust of rights holders and communities, and clear governance of uses.
In the French-speaking world, prioritization criteria need to be defined and, where appropriate, shared corpora or a federation of commons should be developed (Moderator)
Arg. 2The moderator then refocuses the discussion on the French-speaking world by asking which content should be prioritized and according to what criteria. He also opens the prospect of more structured Francophone cooperation, in the form of shared corpora or a federation of commons, while recognizing the diversity of national and institutional frameworks.
The moderator explicitly asks what the criteria for prioritizing digitization should be in the French-speaking world and how to develop sovereignty through a common Francophone corpus or a « federation of commons », distinguishing, where necessary, between national and institutional corpora according to the applicable rules .
on: Digitization priorities must be defined, funded, and supported by institutional capacities and appropriate infrastructure.
The accessible web is reaching saturation, while a vast amount of heritage content remains offline; digitizing it is necessary to feed AI (Speaker 1)
Arg. 1Speaker 1 explains that AI has already extensively exploited content accessible on the web, leading to a form of saturation. At the same time, a significant mass of heritage content remains offline and could enrich AI models if it were digitized and made accessible.
Speaker 1 raises the idea of a future "crash" in the AI market linked to the gradual exhaustion of data harvested from the web, then responds that, on the contrary, there is still an enormous amount of content that remains completely inaccessible on the Internet, sometimes even unknown, which justifies making it available in the digital society . He also notes that, in his own library, only 10% of the content is digitized and accessible, leaving 90% outside the digital sphere .
on: Digitizing dormant content is a necessary condition for having resources useful to AI, beyond what is already accessible on the web.
The quality of data matters just as much as its quantity; otherwise AI reproduces cultural, linguistic, and historical biases (Speaker 1)
Arg. 2For Speaker 1, the challenge is not only to increase the volume of data available for AI, but above all to improve its quality and representativeness. If training data mainly reflect dominant languages and cultures, AI will mechanically reproduce biases in diversity, perspective, and historical depth.
Speaker 1 states that behind every AI model lies the question of data, both in quantity and quality, and connects this to the issue of which cultures, languages, and worldviews will be represented in future systems . He stresses that when AI responses are based on data representing only part of human knowledge, this necessarily introduces biases in diversity, perspective, and diachrony .
on: Cultural and linguistic diversity, particularly Francophone and African, must be better represented in corpora in order to limit AI bias.
Heritage digitization is a matter of sovereignty, because cultural data are becoming a strategic raw material for AI (Speaker 1)
Arg. 3Speaker 1 presents heritage and cultural data as a strategic resource comparable to a raw material for AI. Accordingly, their production, availability, and control are matters of public sovereignty, especially because their digitization depends largely on public investment.
He explains that heritage and cultural data have become the raw material of artificial intelligence and that, when they are not born digital, they must be retrieved from libraries, archives, and museums, which requires investments often borne by the public sector . He concludes that making these contents available and discoverable is part of public policies also aimed at avoiding monopolistic situations and preserving alternative means of access .
on: Should strategic priority go first to maximizing the use of content for AI, or first to protecting against the external capture of value?
La Francophonie has a particular role to play in ensuring the representation of plural cultures, languages, and imaginaries in AI (Speaker 1)
Arg. 4Speaker 1 believes that the Francophonie can play a specific role because of its cultural and linguistic plurality. It can help prevent the erasure of knowledge, imaginaries, and languages in AI systems dominated by content that is already in the majority on the web.
He argues that issues of bias place the Francophonie before a dual challenge, while also giving it an important role precisely because of its plural dimension . He adds that, without corrective action, Francophone knowledge, cultures, and imaginaries risk becoming invisible in the AI ecosystem, even though most of the relevant collections remain inaccessible .
on: Cultural and linguistic diversity, particularly Francophone and African diversity, must be better represented in corpora in order to limit AI bias.
Digitizing is not enough: we need to produce contextualized, reliable, and structured heritage data, not just digital images (Speaker 1)
Arg. 5Speaker 1 argues that heritage digitization cannot be limited to creating visual digital reproductions. Content must become genuine heritage data, meaning contextualized, reliable, and structured so as to be truly useful to machines as well as humans.
He clearly distinguishes the simple provision of a digital facsimile from the more demanding work involved in producing “heritage data,” which must be contextualized or capable of being contextualized, and reliable . He also specifies that such data should be regarded as a strategic asset for society, which justifies investing in their quality and structuring .
on: Interoperability, common standards, and the multiplication of access channels are essential for the discoverability and use of content.
on: Should the main response to gaps in corpora be heritage-based and institutional, or should it first be adapted to the oral and unstructured forms of knowledge?
The quality of capture, control, OCR, full text, and structuring determines the real usability of content by machines (Speaker 1)
Arg. 6For Speaker 1, effective use by AI depends on the entire technical digitization chain, not just the existence of a digital file. The quality of capture, quality control, full-text transcription, and collection structuring directly determines the reliability and usefulness of content for training models.
He describes the crucial role of digitization operators, whom he compares to “knowledge brokers,” explaining that the quality of their work and the calibration of digitization workflows determine the reliability of the final digital copy and therefore its value for training AI models . He adds that it is not enough to have attractive images and structured metadata: high-quality full text must also be produced, printed works and manuscripts alike must be processed, and collections must be transformed into structured data that machines can use directly .
on: The quality, reliability, structuring, and metadata of digitized content are just as important as digitization itself.
Interoperability, common standards, and the multiplication of dissemination channels strengthen access, visibility, and the autonomy of institutions (Speaker 1)
Arg. 7Speaker 1 maintains that an effective digitization strategy requires shared standards, harmonized practices, and interoperable systems. The multiplication of portals and dissemination channels helps content circulate more widely, increases its visibility, and reduces dependence on single access points.
He explains that the French-speaking digital network is working to harmonize inventory practices, catalog data processing, and the consideration of copyright in order to raise the average level of digital maturity among its members . Later, he emphasizes the need to share norms and standards, promote interoperability, and multiply dissemination channels, giving the example of Belgian documents, which are more likely to be consulted if they are made available through specialized or international portals rather than a single institutional website .
on: Interoperability, common standards, and the multiplication of access channels are essential for the discoverability and use of content.
Public funding must support digitization, because it serves the public interest and policies for access to knowledge (Speaker 1)
Arg. 8Speaker 1 believes that the digitization of heritage cannot be left solely to private interests, because it serves a mission of public interest. Since these investments support access to knowledge, cultural diversity, and information sovereignty, they fall under public policy and public funding.
He recalls that digitizing heritage collections held in libraries, archives, and museums is costly and that this investment is most often made by the public sector . He adds that states have economic, legal, and moral responsibilities to protect, preserve, and enhance this shared heritage, while institutions must do more to advance these operations .
on: Digitization priorities must be defined, funded, and supported by institutional capacities and appropriate infrastructure.
Priority should go to fragile or endangered heritage: magnetic tapes, newspapers, unique documents, manuscripts, and recorded oral sources (Speaker 1)
Arg. 9For Speaker 1, prioritization must first respond to the urgency of preservation. The most fragile media and unique content must be dealt with first, because their disappearance would be irreversible and some institutions are their sole custodians.
He explains that the first step is to redo or refine the inventory in order to know what is being preserved and what is at risk of disappearing . He cites magnetic tapes, which are highly fragile, as the top priority, followed by the press, audio sources, unique documents, and manuscript heritage, stressing that if libraries do not handle them, no one will .
on: Digitization priorities must be defined, funded, and supported by institutional capacities and appropriate infrastructure.
The digital maturity of institutions must be strengthened so that they can choose their equipment, partners, and infrastructure without excessive dependence (Speaker 1)
Arg. 10Speaker 1 insists that digitization useful for AI requires institutions capable of mastering the entire technical, legal, and organizational chain. Strengthening this digital maturity helps avoid becoming dependent on service providers or more powerful economic actors.
He explains that, to get from heritage data to AI, it is necessary to master the entire chain of library and archival professions, which requires a certain degree of digital maturity; this is why the réseau francophone numérique is developing training programs and a maturity pyramid whose apex is AI . He adds that this capacity-building must make it possible to better manage public procurement, equipment, infrastructure, and partnerships, so as not to have terms dictated by economic actors due to a lack of expertise .
on: Digitization priorities must be defined, funded, and supported by institutional capacities and appropriate infrastructure.
AI can serve as a political argument for reviving funding for heritage digitization, but this requires clear public priorities (Speaker 1)
Arg. 11Speaker 1 sees the current rise of AI as an opportunity to put heritage digitization back on the political agenda. But he emphasizes that this window of opportunity will have an effect only if it is translated into explicit public priorities and funding programs suited to the entire digitization chain.
He states that AI is bringing digitization issues back to the forefront and offers an opportunity to once again raise awareness among public authorities of the need to fund the preservation and digitization of heritage . He notes, however, that for the moment, the policies observed are aimed mainly at regulating the use of AI in the public sector, rather than refinancing all the upstream chains that make this AI possible, which he describes as "the next battle" .
on: Digitization priorities must be defined, funded, and supported by institutional capacities and appropriate infrastructure.
In the face of information overload and the growing use of generative AI, it is better to improve the quality and completeness of available data than to stand aside (Speaker 1)
Arg. 12Speaker 1 acknowledges that one could choose to keep a distance from AI, but believes that this option has already been largely overtaken by actual use. Since generative AI is increasingly being used instead of search engines, action must be taken on the quality and completeness of the data that feeds it.
In response to Speaker 4's intervention, he explains that many people already use generative AI applications to get quick answers, which makes the question of whether such AI should exist partly outdated . He also invokes the idea of “information overload” to point out that the Internet already contains so much information that people rely on opaque ranking mechanisms; in this context, his goal is to make training data as complete and reliable as possible .
on: Should the goal be to maximize the expansion of corpora for AI, or to accept that a large share of content will remain unused?
Twentieth-century works and contemporary content are also necessary for AI, not just older public-domain works (Speaker 2)
Arg. 1Speaker 2 complements the heritage-based approach by recalling that AI also needs recent and culturally relevant content. He therefore emphasizes that twentieth-century works and contemporary content, often still protected by copyright, are also essential to high-quality AI.
He states at the outset that AI systems need high-quality, culturally relevant data drawn not only from the public domain and older works, but also from contemporary and recent data generally covered by copyright .
on: The digitization of dormant content is a necessary condition for having resources useful to AI, beyond what is already accessible on the web.
Oral content and community knowledge must be digitized with the communities’ consent and with safeguards regarding future uses (Speaker 2)
Arg. 2Speaker 2 insists that the digitization of oral traditions and community knowledge cannot be conceived without consent and safeguards. To build trust, the intended uses must be clarified and the communities concerned must be allowed to retain some form of control over the future exploitation of their content.
He warns that questions of copyright and the protection of folklore must be addressed from the outset when dealing with unpublished heritage of local communities or Indigenous peoples, in order to avoid mistakes already made in certain experiences . He adds that, without control over the conditions of use and subsequent uses by AI systems, these communities will have very little desire to share their knowledge .
on: Digitization must be aligned with respect for rights, the trust of rights holders and communities, and clear governance of uses.
Non-consensual commercial uses of recorded cultural content show the risk of losing control and the need for a framework of trust (Speaker 2)
Arg. 3Speaker 2 shows that the risks are not theoretical: content collected for non-profit purposes can later be reused in commercial channels without recognition of the source communities. This type of example justifies putting a framework of trust in place from the digitization strategy stage.
In the discussion on oral content, he cites the example of a Kanak religious song recorded in New Caledonia by musicologists in the 1980s for non-profit purposes, then commercially reused twenty years later by a successful musician without crediting the source community . He concludes that future uses must be considered from the very design stage of the strategy in order to regulate such reuses .
on: Digitization must be aligned with respect for rights, the trust of rights holders and communities, and clear governance of uses.
It is necessary to reconcile public access, heritage preservation, and respect for authors’ rights through balanced limitations and exceptions (Speaker 2)
Arg. 4Speaker 2 presents copyright not as an absolute obstacle, but as a framework to be balanced with preservation and access. He highlights limitations and exceptions as a tool enabling States to support the public interest while respecting authors' rights.
He explains that his remarks concern how to digitize 20th-century analogue works while respecting authors' rights, particularly through the issue of limitations and exceptions to copyright for preservation and public access . He sets out the “three-step test”: exceptions must apply to certain special cases, must not conflict with the normal exploitation of the work, and must not unreasonably prejudice the legitimate interests of authors .
on: Digitization must be aligned with respect for rights, the trust of rights holders and communities, and clear governance of uses.
Access to digitized content, putting it online, and using it as training data for AI are three legally distinct situations (Speaker 2)
Arg. 5Speaker 2 calls for avoiding confusion between several operations that are often conflated in public debate. Authorizing access to content, making it visible online, and using it to train AI fall under different legal regimes and have different legal implications.
He explicitly distinguishes several scenarios: giving researchers access to collections is not the same as making a document accessible on the Internet . He adds that digitization intended for text research with limited previews, such as Google Books, differs from using digitized works as training data for AI .
on: Digitization must be aligned with respect for rights, the trust of rights holders and communities, and clear governance of uses.
Free access does not automatically follow from digitization; licensing arrangements and content markets for AI already exist (Speaker 2)
Arg. 6Speaker 2 emphasizes that making digital content accessible does not necessarily mean it is free. He points to the growing existence of licensing contracts and structured content markets for AI, which requires thinking through the economics of these uses.
He states clearly that “digitized” and “accessible” do not mean free of charge, then mentions the proliferation of licensing contracts between content holders and AI systems . He adds that companies are presenting to WIPO specialized marketplaces for structuring multilingual datasets, demonstrating the existence of real economic demand for this content .
Consideration must be given to value-sharing among creators, heritage institutions, and AI companies, particularly in public-private partnerships (Speaker 2)
Arg. 7For Speaker 2, the issue is not only access, but also the distribution of the value created from digitized content. This means considering licensing, remuneration mechanisms, collective management, and the effects of public-private partnerships on rights and benefits.
He identifies as a third message the need to reflect on value sharing between content creators and AI companies, taking into account the possibilities of licensing agreements . He also calls for examining the consequences of public-private partnerships for the digitization of content, particularly regarding the transfer or retention of rights, emphasizing that everything depends on the policy objectives being pursued .
Precise metadata, standardized identifiers, and a robust semantic layer are essential for attributing works and managing rights (Speaker 2)
Arg. 8Speaker 2 stresses that a strong information infrastructure is essential for making digitization work in the age of AI. Metadata, identifiers and semantic layers make it possible to identify authors, ensure attribution and effectively manage the associated rights.
He states that a robust infrastructure for attribution and rights management must be put in place, based on a semantic layer describing works and making it possible to identify their authors . He illustrates this requirement by citing several standardized identifiers - ISBN, ISSN, DOI, ISAN, ISRC, ISWC - and mentions standardization work carried out notably by the national libraries of Finland and Latvia .
on: Interoperability, common standards, and the multiplication of access channels are essential for the discoverability and use of content.
Metadata can also include the conditions for use, restriction, or openness of digitized content (Speaker 2)
Arg. 9Speaker 2 explains that metadata are not used only to describe works, but also to govern their circulation. They can incorporate the purposes of digitization, restrictions on use, and levels of accessibility, so that these rules are readable and enforceable by machines.
He answers a technical question by saying that it is entirely possible to include, in the semantic layer describing digitized objects, the purpose of digitization and the conditions of use, whether the access is broad, restricted, or sensitive . He adds that extremely precise metadata make these objects visible to machines and usable in accordance with the policy set by the State .
on: Interoperability, common standards, and the multiplication of access channels are essential for the discoverability and use of content.
Even content considered trivial or of little value can have significant use value for specialized AI systems and targeted data markets (Speaker 2)
Arg. 10Speaker 2 responds to the idea that some unused content would remain worthless by showing that its value depends heavily on its uses. Elements that may appear banal or commercially insignificant can become very valuable for training specialized AI systems or meeting niche documentation needs.
He gives the example of a generative AI company needing images of helicopters in flight: to correct a weakness in its system, it may request the creation of a dataset from television news broadcasts, reports, documentaries, or films, and then have those images annotated, which is "worth a great deal" . He adds that a work with no commercial success is not necessarily without value, citing the French ReLIRE database, which contains niche works and self-references that are useful for specialists .
on: Is the value of unused content mainly limited by its marginal nature, or is it instead revealed by specialized AI uses?
Societies with a strong oral tradition, particularly in Africa, must not be excluded by an approach centered on already structured data (Speaker 3)
Arg. 1Speaker 3 recalls that not all societies preserve their memory in the form of documents already organized as they are in libraries. A digitization strategy focused solely on structured content would therefore risk excluding oral knowledge, particularly African knowledge, even though it is essential and relevant.
He points out that libraries work with structured data, whereas in some African societies information is primarily oral, which means digitization strategies must be adapted to avoid exclusion . He also stresses that oral information can be extremely important and relevant even if it is not based on any already formalized medium .
on: Cultural and linguistic diversity, particularly French-speaking and African diversity, must be better represented in corpora in order to reduce AI bias.
Without a strategy for local value creation, some regions risk supplying the cultural raw material while leaving others to capture the economic value (Speaker 3)
Arg. 2Speaker 3 warns of a risk of reproducing the economic imbalances already seen in other extractive sectors. The regions that produce knowledge could become mere suppliers of cultural raw material, while others would capture the value through commercialization and AI applications.
He reframes the issue by saying that we must avoid having producers of cultural content become merely consumers of products derived from their own heritage . He explicitly draws a parallel with raw materials in Africa, stating that AI could reproduce the situation in which one ‘produces, but remains poor,’ this time with knowledge .
on: Should strategic priority be given first to maximizing the use of content for AI, or first to protecting against the external capture of value?
Digitization strategies must incorporate orality, transitional archiving methods, and the associated costs even before full digitization takes place (Speaker 3)
Arg. 3Speaker 3 argues that a realistic strategy must begin upstream of digitization itself. It is necessary to consider the collection and temporary archiving of oral material, preservation methods, classification trade-offs, as well as the material and energy costs this entails.
He asks how orally transmitted knowledge should be archived, for example when an elder is recorded with a voice recorder, and stresses the need to think through a transitional phase between collection, archiving, and digitization . He adds that the strategy must also take into account preservation methods, trade-offs over what is sovereign or shareable, as well as the costs linked to servers, storage, data centers, and even the lack of electricity .
on: Digitization priorities must be defined, financed, and supported by institutional capacity and appropriate infrastructure.
on: Should the main response to gaps in corpora be heritage-based and institutional, or should it first be adapted to the oral and unstructured forms of knowledge?
Energy costs, storage, servers, and the lack of infrastructure such as electricity complicate large-scale digitization strategies (Speaker 3)
Arg. 4Speaker 3 recalls that digitization is not just a cultural or technical project: it depends on substantial physical infrastructure. In contexts where energy and basic equipment are lacking, a large-scale shift to digital poses major financial and environmental constraints.
He states that the major problem is funding, because digitization is expensive and requires servers, data centers, and storage capacity . He adds that, in some contexts, the very absence of electricity makes these strategies difficult and calls for reflection on what is sustainable for the planet if everyone digitizes everything .
on: Digitization priorities must be defined, financed, and supported by appropriate institutional capacities and infrastructure.
Without the digitization of local heritage, history and culture risk being distorted or rendered invisible in AI tools (Speaker 5)
Arg. 1Speaker 5 warns of future falsification or distortion of historical narratives if local heritage, particularly that transmitted orally, is not digitized. Without this work, AI tools may reproduce incomplete, translated, or distorted versions of important cultural references.
She gives the example of Senegal and suggests that a national hero could one day be reproduced by AI in a degraded or Frenchified form simply because oral transmission had not been digitized and made usable . She adds that if systems are queried about information that has not been digitized, there is a risk of ending up with a falsification of history .
on: Cultural and linguistic diversity, particularly Francophone and African diversity, must be better represented in corpora in order to reduce AI bias.
Ensuring the reliability of oral knowledge and structuring it are essential if it is to be transmitted properly in the digital environment and in AI (Speaker 5)
Arg. 2Speaker 5 believes that simply collecting oral knowledge is not enough: it must be structured and made reliable in order to be transmitted properly in the digital age. In her view, this structuring is a prerequisite for preserving cultural belonging and ensuring reliable uses in AI.
She stresses that heritage must be digitized and structured, particularly in Africa where much information remains oral, so that everyone can know what culturally belongs to them within diversity . She explicitly raises the question of when information is sufficiently reliable to be made available to everyone, making reliability a central element of the strategy .
on: The quality, reliability, structure, and metadata of digitized content are just as important as digitization itself.
Intellectual property can help ensure the reliability of content and regulate its circulation from the strategy design stage onward (Speaker 5)
Arg. 3Speaker 5 presents intellectual property as a tool for ensuring the reliability of digitized content and framing its use. Integrating it early into the strategy would help better secure the circulation of knowledge and strengthen its credibility in digital environments.
She emphasizes the importance of also being able to rely on intellectual property, which she associates with a form of information reliability . She states that if this aspect is integrated very early into the strategy, it could help better regulate and secure content .
on: Digitization must be aligned with respect for rights, the trust of rights holders and communities, and clear governance of uses.
Concrete digitization priorities must be identified to reduce bias in online content and move to action (Speaker 6)
Arg. 1Speaker 6 asks how concerns about information bias can be translated into operational digitization priorities. The intervention emphasizes the need for an action plan, recommendations, standards, and political and financial commitment to make the objectives effective.
Speaker 6 asks whether priorities have already been identified in libraries and in data classification, along with recommendations for digitizing certain categories of information . She also emphasizes that without an action plan, funding, national political priorities, and strong mobilization, it will be difficult to achieve the goal of reducing bias in online content .
on: Digitization priorities must be defined, funded, and supported by appropriate institutional capacities and infrastructure.
In any case, some content will remain unusable or unused, because not everything can be processed or made accessible, given human, legal, and material constraints (Speaker 4)
Arg. 1Speaker 4 takes a more skeptical position and points out that there is an irreducible gap between the encyclopedic ideal of total access to content and the constraints of the real world. In his view, a great deal of content will remain unused, either for legitimate reasons or because no institution can materially process or make everything accessible.
He offers the fiction of a “Martian” arriving in a thousand years to write the history of Earth and wanting to consult a virtually infinite mass of archives, conversations, tickets, private correspondence, unsold collections, photos, and videos, in order to illustrate the scale of potentially unused content . He concludes that such content remains so in part for legitimate reasons today, and that libraries and legislators cannot be expected to work according to such a totalizing ideal .
on: Is the value of unused content mainly limited by its marginal nature, or is it instead revealed through specialized uses of AI?
Session Knowledge Graph
Speakers · Topics · Arguments · Relationships
The moderator frames as the central question the transformation of dormant content into useful, reliable, and well-governed resources for AI . Speaker 1 argues that accessible web data is reaching saturation while a mass of heritage content remains offline, with only 10% of content digitized in his institution . Speaker 2 adds that AI also needs recent and contemporary works, not just older public-domain material . Speaker 5 warns that, without digitizing local heritage, AI risks distorting history and culture . Finally, Speaker 6 calls for operational digitization priorities to correct these biases and move to action .
Dormant content must be made usable for AI, with requirements for reliability, governance, and respect for rights (Moderator)
The accessible web is reaching saturation, while a vast amount of heritage content remains offline; digitizing it is necessary to feed AI (Speaker 1)
Twentieth-century works and contemporary content are also necessary for AI, not just old public-domain works (Speaker 2)
Without digitizing local heritage, history and culture risk being distorted or rendered invisible in AI tools (Speaker 5)
Concrete digitization priorities must be identified to reduce bias in online content and move to action (Speaker 6)
From the outset, the moderator stresses usefulness, reliability, governance, and data quality . Speaker 1 explains that heritage digitization must not be limited to a facsimile, but must produce contextualized, reliable, and structured data, as well as full text that machines can use . Speaker 2 echoes this point by underscoring the need for a robust semantic layer, standardized identifiers, and precise metadata for attribution, rights management, and defining conditions of use . Speaker 5 also emphasizes the structuring and validation of knowledge, especially oral knowledge, before it is made available .
Dormant content must be made usable for AI, with requirements for reliability, governance, and respect for rights (Moderator)
Digitizing is not enough: contextualized, reliable, and structured heritage data must be produced, not just digital images (Speaker 1)
The quality of capture, quality control, OCR, full text, and structuring determines whether content can actually be used by machines (Speaker 1)
Precise metadata, standardized identifiers, and a robust semantic layer are essential to attribute works and manage rights (Speaker 2)
Metadata can also carry the conditions of use, restriction, or openness for digitized content (Speaker 2)
Validating oral knowledge and structuring it are essential so that it can be transmitted properly in the digital and AI environment (Speaker 5)
The moderator refocuses the discussion on the Francophone space by asking which content should be prioritized and whether a common Francophone corpus or a federation of commons should be built . Speaker 1 states that the underrepresentation of certain languages and cultures introduces bias, and that La Francophonie has a particular role to play in preventing its knowledge systems and imaginaries from being made invisible . Speaker 3 recalls that societies with strong oral traditions, especially in Africa, risk being excluded if the focus is only on already structured data . Speaker 5 warns against the falsification or degradation of cultural heritage that has not been digitized . Speaker 6 directly links the question of digitization priorities to reducing bias in online content .
In the Francophone space, prioritization criteria need to be defined and possibly common corpora or a federation of commons should be built (Moderator)
The quality of data matters as much as its quantity; otherwise AI reproduces cultural, linguistic, and historical biases (Speaker 1)
La Francophonie has a particular role to play in ensuring the representation of plural cultures, languages, and imaginaries in AI (Speaker 1)
Societies with strong oral traditions, especially in Africa, must not be excluded by an approach centered on already structured data (Speaker 3)
Without digitizing local heritage, history and culture risk being distorted or rendered invisible in AI tools (Speaker 5)
Concrete digitization priorities must be identified to reduce bias in online content and move to action (Speaker 6)
This point is strongly supported by earlier work on linguistic plurality online, which emphasizes the need for rigorous methods to measure the actual place of languages and improve the discoverability of Francophone content [S29]. It also fits within a broader perspective of African digital diplomacy calling for greater African participation in global digital governance so that its interests and representations are better taken into account [S26].
The moderator explicitly includes authors' rights and content governance in the central issue . Speaker 2 develops at length the idea that preservation, access, and copyright must be reconciled through balanced exceptions, while also clearly distinguishing among access, putting content online, and AI training . He also stresses the need to build trust with communities that hold oral knowledge and to frame future uses . Speaker 5 aligns with this view by presenting intellectual property as a tool for making content more reliable and securing it at an early stage .
Dormant content must be made usable for AI, with requirements for reliability, governance, and respect for rights (Moderator)
Public access, heritage preservation, and respect for authors' rights must be reconciled through balanced limitations and exceptions (Speaker 2)
Access to digitized content, putting it online, and using it as training data for AI are three legally distinct situations (Speaker 2)
Oral content and community knowledge must be digitized with the communities' consent and with safeguards regarding future uses (Speaker 2)
Non-consensual commercial uses of recorded cultural content show the risk of loss of control and the need for a framework of trust (Speaker 2)
Intellectual property can help make content more reliable and govern its circulation from the strategy-design stage onward (Speaker 5)
The moderator explicitly asks for prioritization criteria in the Francophone space . Speaker 1 argues that digitization serves the public interest, should be publicly funded, and should prioritize fragile or endangered heritage, such as magnetic tapes, newspapers, unique documents, and manuscripts . He adds that institutions must strengthen their digital maturity to better choose equipment, infrastructure, and partners . Speaker 3, for his part, highlights material, energy, and archiving costs, as well as electricity and storage constraints . Speaker 6 emphasizes the need for an action plan, funding, and national political priorities to make these goals effective .
In the Francophone space, prioritization criteria need to be defined and possibly common corpora or a federation of commons should be built (Moderator)
Public funding must support digitization, because it serves the public interest and policies of access to knowledge (Speaker 1)
Priority should go to fragile or endangered heritage: magnetic tapes, newspapers, unique documents, manuscripts, and recorded oral sources (Speaker 1)
The digital maturity of institutions must be strengthened so they can choose their equipment, partners, and infrastructure without excessive dependence (Speaker 1)
AI can serve as a political argument to revive funding for heritage digitization, but this requires clear public priorities (Speaker 1)
Digitization strategies must incorporate orality, transitional archiving methods, and associated costs even before full digitization (Speaker 3)
Energy costs, storage, servers, and the lack of infrastructure such as electricity complicate large-scale digitization strategies (Speaker 3)
Concrete digitization priorities must be identified to reduce bias in online content and move to action (Speaker 6)
This finding aligns with the framing of African digital diplomacy policies, which stress the need to mobilize human and institutional resources to enable effective engagement in the digital sphere [S26]. It is also consistent with analyses of digital governance according to which digital transformation requires structural integration and sustainable organizational capacities, not just one-off initiatives [S27].
The moderator, reformulating Speaker 1's intervention, emphasizes that it is not enough to digitize: content must be usable and interoperable . Speaker 1 stresses the harmonization of practices, shared standards, interoperability, and the multiplication of dissemination channels to improve access to and visibility of content . Speaker 2 supports this approach by emphasizing precise metadata, identifiers, and a robust semantic layer, including to carry conditions of use .
Digitizing is not enough: contextualized, reliable, and structured heritage data must be produced, not just digital images (Speaker 1)
Interoperability, common standards, and multiple dissemination channels strengthen access, visibility, and the autonomy of institutions (Speaker 1)
Precise metadata, standardized identifiers, and a robust semantic layer are essential to attribute works and manage rights (Speaker 2)
Metadata can also carry the conditions of use, restriction, or openness for digitized content (Speaker 2)
This point echoes earlier debates on discoverability, where methodological frameworks, indicators, and access to platform data were already identified as conditions for the effective visibility of content [S29]. It also fits into an international context in which rules on data flows, digital regulation, and common frameworks are becoming central governance issues [S28].
Both speakers directly link the absence of digitization, or the poor representation of cultural content, to biases or distortions in AI systems. Speaker 1 refers to bias in diversity, perspective, and diachrony when only certain cultures dominate the data . Speaker 5 extends this idea by warning that undigitized heritage can lead to a falsification of history or to a degraded version of national references . Speaker 2 and Speaker 5 agree that oral and cultural knowledge must not be digitized without a framework of rights and control. Speaker 2 stresses community consent, the protection of folklore, and control over future uses . Speaker 5, for his part, believes that intellectual property must be integrated at a very early stage in order to make the circulation of content more reliable and properly regulated . The two experts share a very similar vision of "high-quality" digitization based on data structuring. Speaker 1 emphasizes contextualization, reliability, full text, and the transformation of collections into data that machines can use . Speaker 2 adds the need for rich metadata, standardized identifiers, and a semantic layer enabling attribution and rights management . These contributions converge on the need to move from a general debate to concrete strategic choices. Speaker 3 calls for thinking through types of information, archiving, preservation, and costs before digitization . Speaker 6 calls for identified priorities, recommendations, funding, and an action plan . The moderator expresses the same requirement in terms of prioritization criteria and French-speaking cooperation . Speaker 1 and Speaker 3 both link digitization to cultural sovereignty. Speaker 1 highlights the challenge of ensuring the representativeness of French-speaking languages, cultures, and imaginaries in AI . Speaker 3 takes the argument further by warning of the risk that societies, especially African ones, may provide the cultural raw material without capturing the value generated .
This consensus is unexpected because it includes even the most skeptical participant. Speaker 4 argues that it will never be possible to process everything, but makes clear that this does not mean the content is uninteresting-quite the opposite . Speaker 2 shows concretely that apparently ordinary content can have great use value for a specialized AI system, such as images of helicopters in flight . Speaker 1 and Speaker 3 stress the heritage value of fragile documents and oral knowledge . Speaker 5 finally links this value to the preservation of history and cultural identity .
Even though their perspectives differ, the speakers converge on the need to create trust. The moderator sets out a framework based on reliability, governance, and respect for rights . Speaker 2 makes trust an explicit objective of any strategy toward rights holders and communities . Speaker 1 also states that AI is not possible without trust and that heritage institutions must help build an open and reliable ecosystem .
The main area of agreement is that digitizing heritage and content still offline has become strategic for AI, provided that it produces reliable, structured, interoperable data within a proper legal framework . There is also strong agreement on the need to better represent cultural, linguistic, and oral diversity, especially within the French-speaking world and Africa, in order to limit AI bias . Lastly, several speakers converge on the importance of priorities, public funding, institutional capacity, and trust to make these objectives achievable .
Speaker 4 argues that there is an irreducible gap between the ideal of total access to content and real-world constraints: even for a fictional future observer, a vast mass of archives and traces will remain unexploited, and it is neither realistic nor desirable to expect libraries and lawmakers to aim for total encyclopedism . Speaker 1 responds that, given the already established uses of generative AI, the question is no longer whether or not to engage with this logic, but rather to ensure that training data are as complete and reliable as possible, including by opening up more libraries and other corpora that are underused today .
Given information overload and the growing use of generative AI, it is better to improve the quality and completeness of the available data than to stand aside (Speaker 1)
Some of the content will in any case remain unusable or unused, because not everything can be processed or made accessible given human, legal, and material constraints (Speaker 4)
This disagreement is illuminated by earlier discussions about the tension between putting content online, making it discoverable, and preserving its meaning. In the workshop on cultural and linguistic diversity in AI, it was recalled that cultures exceed classification logics and that not all content is necessarily reducible to standardized AI exploitation [S30]. At the same time, debates on global digital governance show that expanding data uses must be considered within broader political frameworks and not as an automatic end in itself [S27].
Speaker 4 emphasizes the sheer scale of content that may remain unexploited and the fact that much of it will remain so for legitimate and practical reasons . Speaker 2 implicitly challenges the idea that such content is of little interest by showing that apparently ordinary or low-valued content can acquire significant value in specialized data markets, for example to train visual systems to recognize helicopters in flight, or for niche scholarly uses .
Some of the content will in any case remain unusable or unused, because not everything can be processed or made accessible given human, legal, and material constraints (Speaker 4)
Even content considered trivial or of little value can have high use value for specialized AI systems and for targeted data markets (Speaker 2)
Speaker 1 stresses the importance of digitizing heritage content on a large scale so that it can feed AI, making such digitization an issue of sovereignty and access to knowledge . Speaker 3 shifts the focus to the economic and geopolitical risk: if regions such as Africa digitize without a strategy for local valorization, they could simply supply the cultural raw material while others capture the economic value, reproducing an extractive logic that is already well known .
Heritage digitization is an issue of sovereignty, because cultural data are becoming a strategic raw material for AI (Speaker 1)
Without a strategy for local valorization, some regions risk supplying the cultural raw material while allowing others to capture the economic value (Speaker 3)
This contrast echoes explicit concerns about digital sovereignty, the concentration of digital power, and the unequal distribution of value created by the digital economy [S28]. It is also consistent with the call to strengthen African digital diplomacy so that it can defend national and continental interests in the face of geopolitical realities and power asymmetries in global digital governance [S26].
Speaker 1 mainly develops an approach centered on heritage institutions, collections, metadata, capture quality, OCR, and the transformation of collections into structured data that machines can use . Speaker 3 emphasizes that this approach risks remaining too closely tied to already structured documentary logics and points out that, in some societies, information is primarily oral; collection, transitional archiving, classification, and the specific costs associated with orality must therefore be considered upstream .
Digitization is not enough: we must produce contextualized, reliable, and structured heritage data, not just digital images (Speaker 1)
Digitization strategies must incorporate orality, transitional archiving methods, and the associated costs even before full digitization (Speaker 3)
This disagreement is directly enriched by reflections on museums, living libraries, and oral knowledge, which challenge a purely institutional or classificatory vision of digitized heritage [S30]. The source emphasizes that cultural objects and knowledge exist within linguistic, ritual, and social contexts that cannot easily be converted into standard corpora for AI [S30].
The disagreement is unexpected because Speaker 4 himself makes clear that he never said this content was of no interest . His point is rather that it will never be possible to process everything or make everything open . Speaker 2 responds, however, by showing that very specific or seemingly trivial content can become economically and technically highly valuable for AI . The gap therefore concerns less the intrinsic value of the content than the feasibility and the logics of its exploitation.
The exchange begins from an apparent agreement on the need to digitize heritage, but Speaker 3 introduces an unexpected angle: the reproduction, in AI, of a pattern comparable to that of raw materials, in which those who hold cultural resources would remain poor while others industrialize and commercialize the value . Speaker 1 did indeed speak of sovereignty and strategic assets , but mainly from the perspective of cultural representation and access. Speaker 3 thus reveals an unexpected divide between cultural sovereignty and economic sovereignty.
The disagreements mainly concern four fault lines: the realistic ambition of digitization versus the encyclopedic ideal; the value and exploitability of marginal content; the respective place of heritage urgency, oral knowledge, and community rights; and, finally, the sharing of value between the public interest, cultural sovereignty, and data markets .
Both speakers agree that AI needs broader documentary bases. Speaker 1 emphasizes the enormous mass of heritage content still absent from the web and the need to digitize it . Speaker 2 shares the goal of enriching data that are useful for AI, but believes that recent and contemporary works covered by copyright should also be included, not only older heritage corpora or public-domain materials . They therefore converge on the need to expand corpora, but differ on the priority scope and on the weight of legal constraints .
The accessible web is reaching saturation, while a large mass of heritage content remains offline; digitizing it is necessary to feed AI (Speaker 1) Twentieth-century works and contemporary content are also necessary for AI, not just old public-domain works (Speaker 2)
Everyone recognizes the importance of preserving and making visible oral and fragile knowledge. Speaker 1 ranks sound sources, magnetic tapes, and unique documents among the priorities . Speaker 3 stresses that societies with oral traditions must not be excluded from a digitization strategy designed around libraries and structured data . Speaker 5 emphasizes the structuring and validation of this knowledge to avoid falsification or cultural invisibility . Speaker 2 adds a strong condition: this digitization must not take place without the agreement of communities or without guarantees regarding future uses . They therefore share the same goal of preservation and visibility, but differ on the priority method: heritage urgency, structuring, or governance through consent and control over uses .
Priorities should focus on fragile or threatened heritage: magnetic tapes, newspapers, unique documents, manuscripts, and recorded oral sources (Speaker 1) Societies with strong oral traditions, particularly in Africa, must not be excluded by a vision centered on already structured data (Speaker 3) Ensuring the reliability of oral knowledge and structuring it are essential so that it can be transmitted properly in the digital environment and in AI (Speaker 5) Oral content and community knowledge must be digitized with the consent of the communities and with safeguards regarding future uses (Speaker 2)
These three contributions converge on the need for standards, structuring, and an operational framework. Speaker 1 emphasizes the harmonization of practices, interoperability, and the multiplication of portals to improve visibility and access . Speaker 2 stresses a robust semantic layer, identifiers, and metadata that also enable rights management . Speaker 6 pushes the discussion toward action by asking whether priorities, recommendations, and a concrete digitization plan to reduce biases already exist . The agreement concerns the need for structuring and standards; the divergence concerns the level of progress and what should be put in place first: technical standards, rights management, or a political and financial plan .
Interoperability, common standards, and the multiplication of dissemination channels strengthen access, visibility, and institutional autonomy (Speaker 1) Precise metadata, standardized identifiers, and a robust semantic layer are essential for attributing works and managing rights (Speaker 2) Concrete digitization priorities must be identified to reduce bias in online content and move to action (Speaker 6)
Speaker 1 and Speaker 2 agree that digitization has a strategic dimension and cannot be considered apart from its economic models. Speaker 1 highlights the responsibility of states and public funding in protecting, preserving, and enhancing shared heritage . Speaker 2 also acknowledges the importance of public policy choices, but places greater emphasis on licensing, remuneration, collective management, and public-private partnerships as possible levers for value creation and value sharing . They therefore share the same goal of economic sustainability, but differ on the balance between public-interest public funding and market-based or contractual mechanisms .
Public funding must support digitization, because it serves the public interest and policies for access to knowledge (Speaker 1) The sharing of value among creators, heritage institutions, and AI companies must be considered, particularly in public-private partnerships (Speaker 2)
- The discussion concluded that heritage digitization is a prerequisite for having AI that is useful, reliable, governed, and representative of cultural, linguistic, and historical diversity.
- The speakers emphasized that the accessible web is approaching a form of saturation, while a very large body of heritage, documentary, and oral content remains offline and therefore absent from AI systems.
- The quantity of data is not enough: the quality, reliability, contextualization, structuring, and interoperability of digitized content are essential to avoid bias and enable real use by machines.
- La Francophonie was presented as a strategic space for defending cultural sovereignty and the representation of diverse languages, imaginaries, and bodies of knowledge in AI.
- Societies with oral traditions, particularly in Africa, must not be excluded from digitization strategies; orality, transitional archives, and community knowledge must be integrated from the very design stage of policies.
- Digitization must not be limited to producing images: high-quality full text, rich metadata, standardized identifiers, and a robust semantic layer are also needed for attribution, rights management, and use by AI.
- Copyright, exceptions and limitations, licenses, and conditions of access must be considered upstream, with a clear distinction between preservation, public access, online publication, and the use of content as training data for AI.
- Contemporary content and twentieth-century works are also necessary for AI; the public domain alone is not enough to build relevant corpora.
- The trust of rights holders and communities emerged as a central condition: transparency regarding objectives, uses, and restrictions is essential, particularly for sensitive or oral heritage.
- The discussion highlighted that even content that appears marginal or undervalued may have strong use value for specialized AI, which reinforces the importance of value-sharing and partnership governance.
- Digitization priorities should focus first on fragile or endangered heritage, such as magnetic tapes, newspapers, unique documents, manuscripts, and oral sources that have already been recorded.
- Public funding, the growing digital maturity of heritage institutions, and control over infrastructure, standards, and public procurement are seen as essential levers for avoiding excessive dependence on private actors.
““There is no AI without digitization,” and heritage data are no longer just the memory of the past, but “the raw material of our future.””
“Digitizing heritage is not a simple technical operation: it is not enough to have “beautiful images”; you need contextualized, reliable data, high-quality full text, and collections transformed into structured data that machines can use.”
“Several things must be distinguished: access for researchers, putting content online on the Internet, text search without full access, and use as training data for AI; “making something accessible does not mean it is free.””
“For oral content and local knowledge, the strategy must be conceived differently: “we do not all follow the same path toward digitization”; much knowledge, especially in Africa, is oral, unstructured, and at risk of being excluded.”
“The strategy must also define what is “sovereign” and what may or may not be shared, so as to anticipate regulation rather than scramble to catch up with it.”
“The scenario of the “Martian in a thousand years” who would want to write the history of Earth from all unused content, including telephone conversations, receipts, videos, private correspondence, and unsold collections.”
““Many people have already started today to stop using search engines” in favor of generative AI tools; therefore, it is better for these systems to be trained on data that are “as complete as possible and as reliable as possible.””
“The example of a Kanak religious song recorded for non-profit purposes and then commercially reused without crediting the source community.”
“Even seemingly harmless content has value for AI: to train a system to correctly generate “helicopter flight,” images from television news, reports, films, and documentaries must be collected and annotated.”
““Knowledge in AI also runs the risk of repeating the same pattern we saw with raw materials. We produce, but we remain poor.””
“The Senegalese example: without digitizing oral transmission, there is a risk that tomorrow history or national heroes will be reconstructed in a falsified way, for example “in French,” because local sources will not have been structured and integrated.”
“According to Speaker 1, the priority is first to “take stock again,” to know what exists and what is at risk of disappearing, and then to prioritize fragile, unique, or threatened heritage, particularly magnetic tapes, the press, manuscripts, and documents for which institutions are the sole holders.”
“It is necessary to “build trust,” ensure transparency of objectives, and put in place a robust “semantic layer” with identifiers and metadata to manage rights, authorized uses, and the machine visibility of content.”
How can oral knowledge and unstructured content be integrated into a digitization and AI strategy without excluding societies where transmission is primarily oral?
This question is important to prevent entire bodies of heritage, particularly African and Indigenous, from remaining invisible in digital corpora and AI systems. It touches on cultural diversity, inclusion, and data representativeness.
What interim archiving methods should be put in place before digitizing oral or fragile content?
Before digitization, it is necessary to know how to collect, preserve, and document content that is sometimes unique and vulnerable. This is an essential prerequisite to avoid the loss of knowledge even before its digital processing.
How should a digitization strategy define what falls under informational sovereignty and what can be shared publicly?
The distinction between sovereign, public, and restricted content is central to anticipating regulation, protecting cultural interests, and organizing the conditions of access and reuse, particularly within the Francophone space.
How can digitization, storage, infrastructure, and the energy required be funded sustainably, while taking environmental constraints into account?
The cost of digitization and its infrastructure is a major obstacle, especially in contexts where material and energy resources are limited. This issue determines the real feasibility of the proposed policies.
Where should the line be drawn between the encyclopedic ambition of exploiting all possible content and the practical, legal, and social limits of that ambition?
This question probes the very purposes of digitization and AI: preserving everything and exploiting everything is neither simple nor necessarily desirable. It compels reflection on selection criteria, uses, and trade-offs.
How can countries or communities that produce content be prevented from becoming merely suppliers of raw material while others capture the economic value of AI?
This question concerns value-sharing, economic justice, and the asymmetry between knowledge producers and technology companies. It is crucial to avoid reproducing the inequalities already observed with raw materials.
How can the trust of local communities, Indigenous peoples, linguistic minorities, and rights holders be built in digitization systems and AI uses?
Without trust, holders of knowledge or rights may refuse the digitization or opening up of their content. The issue is decisive for enabling sustainable, ethical, and accepted policies.
What specific uses should be authorized after digitization: preservation, scientific access, public dissemination, text-based search, AI training, or others?
Clarifying uses is necessary in order to define an appropriate legal and technical framework. Not all uses raise the same issues in terms of copyright, consent, access, or remuneration.
How can the conditions of use, access restrictions, and digitization objectives be technically indicated in metadata and the semantic layer?
This line of research is essential to make rights and access policies machine-readable, facilitate content governance, and ensure uses comply with the choices of institutions and states.
What prioritization criteria should be adopted to decide what should be digitized first, particularly in the French-speaking world?
Since resources are limited, clear prioritization criteria are needed. The debate brings out several possible avenues: endangered heritage, magnetic tapes, fragile press archives, unique documents, national heritage, and underrepresented content.
How can a shared corpus or a federation of digital commons be built within the French-speaking world, despite differing national or institutional rules?
This question is important for strengthening Francophone digital sovereignty, improving interoperability, and increasing the presence of Francophone languages and cultures in AI systems.
Are there already national priorities or concrete action plans in member countries to fund and guide AI-related digitization?
The issue is whether the discourse has already been translated into public policy, budgets, and operational programs. Without an action plan, the goals of representativeness and sovereignty risk remaining theoretical.
To what extent can norms and standards facilitate the classification, interoperability, and prioritization of data to be digitized?
Standards are presented as a lever for sharing practices, improving data quality, and enabling machine exploitation of data. This is a key issue for moving from isolated initiatives to compatible ecosystems.
How can information drawn from oral traditions be made reliable before being made accessible and reusable by all?
The reliability of content is essential to prevent historical falsification and the spread of inaccurate versions in digital tools and AI. The issue is particularly sensitive for heritage transmitted orally.
How can intellectual property be reconciled with ensuring the reliability of heritage content, particularly oral content?
The link between intellectual property, attribution, and content credibility is central to protecting authors and communities while ensuring a trusted framework for the circulation of knowledge.
What lessons can be learned from past experiences where overly broad copyright exceptions failed or triggered negative reactions from rights holders?
This line of research is important for designing balanced policies and avoiding the reproduction of arrangements perceived as unfair or ineffective. It helps identify the conditions under which reforms will be acceptable.
Which public-private partnership models for digitization are the most relevant, and what are the implications for rights, copies, access, and the value created?
Partnerships with private-sector actors can accelerate digitization, but they raise questions about control, rights transfers, and benefit-sharing. This is an area that requires careful analysis of the possible models.
What is the economic and strategic value of niche, non-commercial, or seemingly harmless content for training specialized AI systems?
The discussion shows that content with low visibility can still be highly useful for certain AI applications. Better understanding this value would help guide digitization and licensing strategies.
How can access channels and dissemination portals be multiplied without losing interoperability, so as to avoid dependence on a single actor or a single interface?
Diversifying access is linked to sovereignty, discoverability, and combating monopolistic situations. It is an important condition for ensuring that digitized content is genuinely visible and useful.
How can the digital maturity of heritage institutions be collectively raised, particularly in countries with limited resources?
Digitization that can be effectively used by AI requires skills, methods, and a well-controlled technical workflow. The issues of training, harmonizing practices, and institutional solidarity are therefore strategic.
How can the specific urgency of endangered media, such as magnetic tapes, fragile newspapers, or unique documents, be incorporated into national and international strategies?
These media are at risk of disappearing rapidly. Giving them priority treatment is a matter of irreversible preservation, but also of future availability for research, public access, and the training of AI systems.
How can the falsification of history or the misrepresentation of cultural figures be avoided when local digitized and structured content is lacking?
Without reliable sources drawn from local heritage, AI systems risk spreading inaccurate or external narratives. This issue directly affects collective memory, education, and cultural self-representation.
How can heritage data be made not only accessible, but truly usable by machines, particularly through high-quality full text and structured data?
The transition from digitized images to data usable by AI is a major technical challenge. Without structuring, high-quality OCR/HTR, metadata, and contextualization, digitization remains incomplete from the standpoint of advanced uses.