WSIS Forum 2026
Rapport généré par l'IA

Science in the Age of AI: Knowledge, Data, and Trust

8 intervenants
Résumé

Résumé

Cette discussion, animée par le Pr David Castle, a réuni un panel international pour explorer l'impact de l'intelligence artificielle sur la connaissance scientifique, la qualité des données et l'intégrité de la recherche. Le Dr Vanessa McBride a ouvert les débats en soulignant que l'IA n'est pas seulement un produit de la science, mais qu'elle remodèle désormais fondamentalement la manière dont la science est pratiquée , et a mis en évidence que la plupart des stratégies nationales en matière d'IA ne traitent pas du secteur scientifique en tant que tel . Sur la question de la qualité des données, le Dr Kamil Dziubek a utilisé l'exemple d'AlphaFold pour illustrer comment les modèles d'IA dépendent de données d'entraînement qui sont elles-mêmes basées sur des modèles et sujettes à des biais et des erreurs , en soutenant que des mesures robustes de qualité des données, incluant l'exactitude, la provenance et la traçabilité, sont essentielles pour une science pilotée par l'IA digne de confiance . Le Dr Moses Thiga a soulevé des préoccupations relatives aux pays du Sud, avertissant que l'IA favorise l'émergence d'une génération de chercheurs dépourvus de compétences empiriques fondamentales , et que l'insuffisance des infrastructures de données et des capacités de calcul risque d'aggraver les inégalités existantes . Le Dr Marion Mercier a mis en lumière la façon dont l'IA transforme des disciplines entières, citant la découverte de médicaments comme un domaine où l'IA pourrait réduire les délais de développement de plusieurs années à quelques jours , tout en soulevant la question plus profonde de savoir si la science peut rester porteuse de sens si l'IA génère des connaissances que les humains ne peuvent pas interpréter . Le Pr Vukosi Marivate a noté que le volume considérable de soumissions assistées par l'IA aux conférences académiques crée de sérieux défis en matière d'intégrité , et a souligné que l'IA amplifie à la fois le meilleur et le pire des systèmes de recherche existants . Sur la souveraineté des données et la science ouverte, Vukosi Marivate a soutenu que des modèles de licences équitables sont nécessaires pour empêcher les grandes entreprises technologiques de bénéficier de manière disproportionnée des données partagées librement , en citant des initiatives telles que la licence Esethu comme alternatives concrètes . Alistair Nolan a ajouté que les grandes entreprises technologiques orientent le programme de recherche vers une IA à haute intensité de calcul et de données, potentiellement au détriment de l'intérêt public plus large . Le panel s'est globalement accordé sur le fait que les institutions, les gouvernements et la communauté scientifique doivent repenser la formation à la recherche, la réglementation et la gouvernance des données afin de garantir que l'IA serve la science de manière équitable et responsable .

Points clés

Objectif général

La discussion visait à explorer l'impact multidimensionnel de l'intelligence artificielle sur la pratique scientifique, la production de connaissances, la qualité et l'accès aux données, ainsi que l'intégrité de la recherche. Organisée par le Conseil international des sciences et CODATA, la session a réuni des panélistes issus de contextes mondiaux variés pour examiner à la fois les opportunités et les risques que l'IA présente pour les systèmes scientifiques, avec une attention particulière portée à l'équité et aux pays du Sud. ---

Principaux points de discussion

- L'IA transforme la pratique scientifique à chaque étape du processus de recherche, apportant à la fois des opportunités considérables et des risques sérieux. Les panélistes ont noté que l'IA est utilisée à chaque étape du processus scientifique et dans tous les domaines , notamment pour la gestion des flux de travail scientifiques et l'identification des revues prédatrices . Des mises en garde ont toutefois été formulées contre l'adoption de flux de travail pilotés par l'IA sans garde-fous adéquats . Vukosi Marivate a illustré l'ampleur de cette perturbation en décrivant comment l'IA a contribué à une explosion des soumissions d'articles - passant de quelques milliers à 13 000 en un seul cycle - soulevant des questions urgentes sur l'intégrité scientifique et la prévalence de contenus générés par l'IA sans véritable contribution scientifique . - La qualité des données est fondamentale pour une IA digne de confiance en science, et la « vérité terrain » utilisée pour entraîner les modèles est intrinsèquement dynamique et imparfaite. En prenant AlphaFold comme étude de cas, Kamil Dziubek a expliqué que les modèles d'IA en science sont entraînés sur des données qui sont elles-mêmes des modèles - des reconstructions issues de techniques physiques - et sont donc sujettes à des biais, des inexactitudes et une obsolescence . Il a insisté sur le fait que la vérité terrain d'hier n'est pas celle d'aujourd'hui, rendant une validation continue indispensable . Le principe a été résumé ainsi : « L'IA est bénéfique si elle repose sur de bonnes données », ce qui exige des mesures convenues de qualité, d'incertitude, d'exactitude et de provenance . - Les pays du Sud font face à un désavantage cumulé : l'IA offre la promesse d'un saut technologique, mais risque d'approfondir l'incompétence empirique, l'exclusion des données et la dépendance technologique. Moses Thiga a souligné que si l'IA permet la simulation, la revue de littérature et l'analyse dans des environnements aux ressources limitées , elle produit simultanément une génération de scientifiques incapables de mener de véritables expériences, de lire des articles de manière critique ou d'analyser des données de façon autonome . Cette situation est aggravée par des directions académiques peu familières avec l'IA et par le fait que la plupart des modèles sont principalement entraînés sur des données occidentales, tandis que l'infrastructure de données dans les pays du Sud reste embryonnaire . Vukosi Marivate a ajouté que l'IA amplifie les inégalités existantes . - La souveraineté des données, les licences équitables et la concentration du pouvoir de recherche en IA constituent de sérieuses menaces pour la science ouverte et l'intérêt public. Vukosi Marivate a décrit comment les mandats de données ouvertes ont conduit à un « open washing », où les grandes entreprises technologiques bénéficient de manière disproportionnée des données sous licence ouverte sans réinvestir dans les communautés qui les ont créées . En réponse, de nouveaux cadres de licences tels que la licence Esethu et la licence NOODL ont émergé pour garantir le partage des bénéfices en fonction du statut géographique ou économique . Alistair Nolan a renforcé cette préoccupation en notant que les grandes entreprises technologiques surpassent les universités en matière de dépenses en recherche et développement sur l'IA, collaborent principalement avec des institutions d'élite américaines et orientent la recherche vers des modèles à haute intensité de calcul et de données qui servent des intérêts corporatifs plutôt que publics . - Les institutions, les gouvernements et la communauté scientifique doivent repenser d'urgence la réglementation, l'éthique et la définition même de la connaissance scientifique à l'ère de l'IA. Moses Thiga a soutenu que les universités doivent reconsidérer leur mission fondamentale en recentrant leur attention sur le rôle de gardien éthique, l'enseignement d'une bonne méthode scientifique et l'investissement dans les capacités de calcul plutôt que dans les campus . Un membre du public a soulevé le risque que les stratégies nationales en matière d'IA, qui ignorent largement le secteur scientifique , puissent par inadvertance mal réguler la science ou ne pas la réguler du tout, et a appelé à l'élaboration de codes de pratique disciplinaires pour démontrer que la communauté scientifique est capable de s'autoréguler . Moses Thiga a répondu que la nature de contrôle et d'équilibre propre à la science doit être préservée, et que les gouvernements - citant l'Inde comme modèle - pourraient avoir besoin d'égaler les investissements de l'industrie pour protéger la souveraineté scientifique . ---

Ton général

Les intervenants ont eu une discussion honnête, engagée et légèrement optimiste, tout en exprimant une légère inquiétude. Les remarques d'ouverture étaient largement contextuelles et mesurées, le Dr Vanessa McBride et les panélistes cadrant l'impact de l'IA sur la science comme vaste et lourd de conséquences . Au fil de la conversation, le ton est devenu plus direct et parfois franc, notamment avec les observations de Vukosi Marivate sur l'état des communautés de recherche en IA et les avertissements de Moses Thiga concernant « une génération de scientifiques empiriquement incompétents » . Cependant, la discussion n'est jamais devenue pessimiste ; les panélistes ont constamment équilibré la critique avec les possibilités. Vers la fin de la discussion, le ton a évolué vers une résolution constructive des problèmes, notamment autour des licences de données et de la réglementation, pour se conclure sur une note collaborative et tournée vers l'avenir.

Intervenants

- Dr Vanessa McBride - Rôle : Représentante du Conseil international des sciences et de CODATA - Domaine d'expertise : Politique scientifique, impact de l'IA sur les systèmes scientifiques - Pr David Castle - Rôle : Président du projet Science Systems Futures ; modérateur de la session - Domaine d'expertise : Systèmes scientifiques, politique de la recherche - M. Alistair Nolan - Affiliation : OCDE (participation en ligne depuis Paris) - Domaine d'expertise : Politique de l'IA, économie de la science et de l'innovation, productivité de la recherche - Dr Kamil Dziubek - Affiliation : Université de Vienne ; CODATA (co-président du groupe de travail sur la gestion de la qualité des données de recherche tout au long du cycle de vie des données) - Domaine d'expertise : Gestion de la qualité des données de recherche, biologie structurale, IA en science (notamment la prédiction de la structure des protéines et la validation des données) - Dr Moses Thiga - Affiliation : Université d'Egerton, Kenya - Rôle : Responsable chargé de promouvoir l'utilisation des technologies au sein de son université - Domaine d'expertise : IA dans l'enseignement supérieur, adoption des technologies dans les pays du Sud, intégrité de la recherche - Dr Marion Mercier - Affiliation : Geneva Science and Diplomacy Anticipator (GESDA), fondation indépendante à but non lucratif - Domaine d'expertise : Anticipation en science et technologie, diplomatie scientifique, dialogue prospectif sur les technologies émergentes - Pr Vukosi Marivate - Affiliation : Université de Pretoria (Directeur de l'Institut africain pour la science des données et l'IA ; titulaire de la chaire de science des données) ; co-fondateur de Lelapa AI ; membre du Panel scientifique indépendant des Nations Unies sur l'IA - Domaine d'expertise : Traitement du langage naturel, IA pour les langues à faibles ressources, évaluation de l'IA, équité des données et licences Intervenants supplémentaires : - Public - Affiliation : Afrique du Sud ; CODATA - Question portant sur la viabilité financière des grands modèles d'IA et ses implications pour les pays du Sud - Public (représentant de l'IFLA) - Affiliation : Fédération internationale des associations de bibliothécaires et des bibliothèques - Question portant sur les codes de pratique disciplinaires et l'autorégulation en science dans le contexte de l'IA - Public (journaliste indépendant local) - Question portant sur la motivation académique au Kenya et la nature de l'enquête scientifique (si la valeur réside dans la question ou dans la réponse) - Public (représentante de Women in Technology, Nigeria) - Affiliation : Women in Technology in Nigeria - Question portant sur la question de savoir si la souveraineté des données constitue une menace pour la science ouverte

Intervenants
MA
Mr. Alistair Nolan
165 wpm · 5 min
DK
Dr. Kamil Dziubek
127 wpm · 7 min
DM
Dr. Moses Thiga
140 wpm · 7 min
DV
Dr. Vanessa McBride
134 wpm · 4 min
PV
Prof. Vukosi Marivate
168 wpm · 15 min
DM
Dr. Marion Mercier
194 wpm · 6 min
A
Audience
136 wpm · 5 min
PD
Prof. David Castle
152 wpm · 6 min

La science à l'ère de l'IA : connaissance, données et confiance

#

Ouverture et mise en contexte

La session a été convoquée par le Conseil international des sciences et le Comité de données pour la science et la technologie (CODATA) afin d'examiner l'impact multidimensionnel de l'intelligence artificielle sur la pratique scientifique, la production de connaissances, la qualité des données et l'intégrité de la recherche. Le Dr Vanessa McBride a ouvert les travaux en présentant l'IA non pas simplement comme un produit de la science, mais comme une force qui remodèle en profondeur la manière dont la science elle-même est pratiquée . Elle a attiré l'attention sur l'ampleur de l'impact de l'IA sur la littérature scientifique, en évoquant des développements allant de l'essor des agents IA pour gérer les flux de travail scientifiques et identifier les revues prédatrices, jusqu'aux mises en garde sérieuses contre l'adoption de flux de travail pilotés par l'IA sans garde-fous adéquats . L'une des observations centrales de Vanessa McBride était que, malgré la prolifération des stratégies nationales en matière d'IA, presque aucune d'entre elles ne contient de volet significatif consacré au secteur scientifique lui-même, se concentrant plutôt sur des domaines d'application en aval tels que la santé et l'agriculture . Ce vide en matière de gouvernance, a-t-elle soutenu, risque de négliger les fondements scientifiques mêmes dont émergent en définitive les nouvelles technologies et applications.

Vanessa McBride a également mis en lumière un rapport publié plus tôt dans l'année par le Conseil international des sciences, intitulé Preparing National Research Ecosystems for AI, qui synthétisait des études de cas menées dans 26 pays . Plusieurs des auteurs de ce rapport étaient présents dans le panel, notamment des contributeurs d'Afrique du Sud et du Kenya. La conclusion centrale du rapport, qui disait que les stratégies nationales en matière d'IA restent largement muettes sur le secteur scientifique, a constitué la toile de fond thématique de la discussion qui a suivi . Vanessa McBride a esquissé trois thèmes que le panel allait aborder : l'IA et la production de connaissances, les données scientifiques comme fondement de l'IA, et les enjeux en aval liés à la fiabilité, à la confiance et à l'intégrité de la recherche .

#

L'IA comme amplificateur : opportunités et hyperboles

Alistair Nolan, participant au panel en ligne depuis l'OCDE, a offert la perspective la plus optimiste de la session. Il a décrit l'IA comme « un merveilleux auxiliaire et amplificateur de l'intelligence humaine » et a déclaré être « très confiant quant aux implications à long terme de l'IA pour la science », notant que l'IA est désormais utilisée à chaque étape du processus scientifique et dans tous les domaines . Il a toutefois reconnu qu'il existe « beaucoup d'hyperboles » autour de l'IA et qu'elle engendrera « une série de tensions institutionnelles » qui devront être gérées avec soin .

La contribution la plus substantielle d' Alistair Nolan a porté sur l'évolution de la nature des goulots d'étranglement scientifiques. S'appuyant sur des recherches de l'OCDE dans le domaine des sciences des matériaux, il a soutenu que la découverte elle-même pourrait cesser d'être l'étape la plus complexe des sciences et de la technologie. Le défi principal réside de plus en plus dans le passage à l'échelle industrielle des découvertes réalisées en laboratoire . Il a également cité les propos du mathématicien lauréat de la médaille Fields, Terence Tao, pour étayer l'idée que l'IA permettra aux jeunes scientifiques prometteurs de s'attaquer plus tôt dans leur carrière à des problèmes de pointe, en réduisant la charge cognitive consacrée à des tâches de bas niveau telles que la mémorisation de vastes corpus de littérature et l'exécution de longs calculs . Dans la perspective d'Alistair Nolan, l'IA ne remplace pas le talent scientifique, mais en accélère le développement et élargit le cercle de ceux qui peuvent participer à la recherche de pointe .

#

La qualité des données et la cible mouvante de la vérité terrain

Le Dr Kamil Dziubek, co-président du groupe de travail de CODATA sur la gestion de la qualité des données de recherche tout au long du cycle de vie des données, a ancré la discussion dans une étude de cas concrète et instructive en parlant d'AlphaFold, la famille de programmes IA capables de prédire la structure tridimensionnelle des protéines, qui a valu à ses auteurs le prix Nobel de chimie en 2024 . AlphaFold est entraîné sur des données issues de la Protein Data Bank, qui contient plus de deux cent cinquante mille structures déterminées expérimentalement à partir de techniques telles que la diffraction des rayons X, la cryo-microscopie électronique et la résonance magnétique nucléaire . Le point critique soulevé par Kamil Dziubek était que ces structures sont elles-mêmes des modèles - des reconstructions à partir de données physiques brutes - et portent donc toutes les limites inhérentes aux modèles . Invoquant l'aphorisme du statisticien George Box selon lequel « tous les modèles sont faux, mais certains sont utiles », il a soutenu que la communauté scientifique doit activement s'employer à identifier et à éliminer les modèles biaisés, de mauvaise qualité ou contextuellement inappropriés .

Kamil Dziubek a également souligné que la « vérité terrain » utilisée pour valider les systèmes d'IA n'est pas statique mais constitue une cible mouvante : de nouveaux ensembles de données émergent chaque jour, et ce qui était considéré comme vérité terrain hier ne l'est peut-être plus aujourd'hui . Cette nature dynamique de la connaissance scientifique rend la validation continue indispensable et exige des mesures de qualité des données convenues entre disciplines, incluant l'exactitude, l'incertitude, la traçabilité et la provenance . Il a résumé le principe de manière concise : « l'IA est bénéfique si elle repose sur de bonnes données », et a averti que si ceux qui déploient des méthodes d'IA ne peuvent pas faire preuve de transparence concernant leurs données d'entraînement, « c'est un signal d'alarme majeur » . CODATA est actuellement engagé dans un travail de cartographie des différentes mesures de qualité des données entre disciplines, dans le but de définir et de convenir de normes pouvant être appliquées concrètement .

#

Incompétence empirique et désavantage cumulé du Sud global

Le Dr Moses Thiga, s'exprimant à partir de son expérience dans la conduite de l'adoption technologique à l'Université Igaton au Kenya, a offert une perspective nettement plus prudente, en particulier concernant le Sud global. Il a reconnu que l'IA présente de véritables opportunités de saut technologique en permettant nottament la simulation de laboratoires, la revue de littérature, le brainstorming et l'analyse de données dans des environnements aux ressources limitées . Cependant, il a identifié une contre-tendance profondément préoccupante : l'émergence d'une génération de scientifiques « empiriquement incompétents » . Les chercheurs dans son contexte, a-t-il averti, produisent des publications sans être capables de mener de véritables expériences, de lire des articles de manière critique, de collecter des données ou d'évaluer de manière indépendante les résultats de leurs recherches . Ce problème est aggravé par des directions académiques et de recherche qui ne maîtrisent pas l'IA, conduisant à une situation où l'IA est soit diabolisée et reléguée dans l'ombre sous forme d'« IA fantôme », soit adoptée sans esprit critique, sans les compétences nécessaires pour évaluer ses résultats .

Moses Thiga a également mis en évidence des désavantages structurels qui rendent le Sud global particulièrement vulnérable. La plupart des modèles d'IA sont principalement entraînés sur des données occidentales, tandis que l'infrastructure de données dans le Sud global reste embryonnaire et loin d'être mature . Sans la capacité de développer des modèles localement pertinents, et sans infrastructure de calcul propre, l'opportunité de saut technologique risque de devenir un mécanisme supplémentaire par lequel le Sud global prend encore plus de retard . Cette préoccupation a été renforcée par le Pr Vukosi Marivate, qui a noté que l'IA amplifie les inégalités existantes, aggravant les problèmes préexistants pour le Sud global .

#

L'IA transformant la pratique scientifique dans toutes les disciplines

Le Dr Marion Mercier, de la Geneva Science and Diplomacy Anticipator - une fondation indépendante à but non lucratif travaillant avec des scientifiques pour anticiper les avancées sur des horizons de cinq, dix et vingt-cinq ans - a décrit l'impact de l'IA comme « omniprésent » et « catalytique » dans toutes les disciplines scientifiques que son organisation examine . Elle a noté que son organisation avait récemment tenu son premier comité d'anticipation entièrement consacré à l'IA pour la science, présidé par Hiroaki Kitano, qui dirige le Nobel Turing Challenge, et que les enseignements de ce comité étaient à venir et pertinents pour la discussion de la session.

Marion Mercier s'est appuyée sur un atelier d'anticipation sur l'amélioration cognitive pour illustrer une implication particulièrement frappante : alors que les neuroscientifiques ne comprennent toujours pas pleinement le cerveau , la combinaison des interfaces cerveau-ordinateur et de l'IA pourrait atteindre un point où l'IA comprend le cerveau même si les humains ne le comprennent pas . Cela a soulevé ce qu'elle a décrit comme une question profonde sur l'interprétabilité : si l'IA peut atteindre les objectifs des neurosciences - moduler la fonction cérébrale, traiter des maladies - mais que les humains ne peuvent pas comprendre comment elle le fait, est-ce scientifiquement et éthiquement acceptable ? La question de savoir si des résultats obtenus sans compréhension humaine constituent une connaissance scientifique valide a été délibérément laissée ouverte, mais elle a introduit une dimension épistémologique qui a résonné tout au long du reste de la session.

Marion Mercier a également noté que la découverte de médicaments est l'un des domaines où l'automatisation pilotée par l'IA est la plus bienvenue et la plus avancée. Les comités d'anticipation ont suggéré que dans un horizon de vingt-cinq ans, les délais de découverte de médicaments pourraient pourraient passer de plusieurs années à quelques jours, grâce à l'exploration par l'IA des données cliniques et à la synthèse de composés chimiques . Elle a noté que ce niveau d'automatisation, bien que potentiellement transformateur dans la découverte de médicaments, pourrait ne pas être bienvenu dans tous les aspects de la science .

#

La crise d'intégrité dans la publication de recherches en IA

Le Pr Vukosi Marivate, Directeur de l'African Institute for Data Science and AI et titulaire de la chaire de science des données à l'Université de Pretoria, co-fondateur de Lilapa AI et membre du Groupe scientifique indépendant des Nations Unies sur l'IA, a offert une analyse et un témoignage sur les effets perturbateurs de l'IA sur la publication scientifique. Il se concentre actuellement sur les questions d'évaluation des modèles d'IA et sur la manière dont celles-ci peuvent être améliorées. Il a décrit la situation dans les conférences de traitement du langage naturel comme « un chaos », notant que des cycles de soumission qui recevaient auparavant deux à trois mille articles sont passés à treize mille soumissions en un seul cycle . Il a expliqué que l'Association for Computational Linguistics (ACL) était passée à un système d'évaluation en continu - fonctionnant initialement toutes les six semaines, puis étendu à huit ou dix semaines - dans lequel les articles entrent dans un pool commun et les auteurs choisissent à quelle conférence les présenter après acceptation. Cela constitue un changement structurel qui a contribué de manière significative à l'explosion des volumes de soumissions. Les évaluateurs sont désormais chargés de vérifier si les références dans les articles soumis sont réelles, un développement qu'il a décrit comme symptomatique d'une crise d'intégrité plus large . Il a également noté que NeurIPS, l'une des conférences phares du domaine, attire désormais entre 20 000 et 30 000 participants, illustrant l'ampleur considérable du défi. Son évaluation était sans détour : « ça nous dévore en tant que chercheurs en IA » .

Selon Vukosi Marivate il n'y a pas raison de désespérer, mais il s'agit d'un défi que la communauté scientifique doit relever, nécessitant de nouvelles façons de penser l'intégrité scientifique, la nature de la découverte et ce qui constitue une véritable contribution scientifique . Il a également reconnu l'enthousiasme sincère actuel, notant que l'IA a ouvert des pistes de recherche qu'il avait laissées inexplorées lors de son doctorat il y a onze ans . Son message global était que l'IA est un amplificateur : elle peut amplifier le bien, et la tâche consiste à renforcer cela tout en réduisant autant que possible les effets négatifs .

#

Repenser la finalité des institutions scientifiques

En réponse à la question sur les implications à long terme de l'IA pour la production de connaissances scientifiques Moses Thiga a soutenu que la communauté scientifique doit fondamentalement redéfinir ce que signifient la connaissance, la science et la recherche à l'ère de l'IA . Il a affirmé que les universités doivent reconsidérer leur mission fondamentale : plutôt que de se concentrer sur des campus physiques, les institutions devraient investir dans une meilleure infrastructure de calcul et se positionner comme gardiennes éthiques, enseignant la bonne méthode scientifique et les valeurs qui sous-tendent une recherche responsable . La connaissance, a-t-il observé, est déjà « disponible partout », ce qui soulève des questions urgentes sur ce que les universités enseignent réellement et sur la finalité de la recherche .

Alistair Nolan a complété cela en soutenant que l'IA changera non seulement la manière dont la science est pratiquée, mais aussi qui la pratique, permettant à un plus grand nombre de personnes de participer à de grands projets scientifiques grâce à des initiatives de science citoyenne et permettant aux jeunes chercheurs de s'engager plus tôt à la frontière de la recherche . Vukosi Marivate a ajouté une réflexion personnelle : bien qu'il travaille dans le domaine de l'IA, il n'a jamais possédé autant de carnets de notes qu'aujourd'hui, car s'asseoir avec ses propres pensées - plutôt que de se tourner immédiatement vers la boîte de dialogue de l'IA - est essentiel pour manier l'IA comme un outil plutôt que de la laisser faire la réflexion à sa place . Cette observation a souligné une préoccupation partagée par l'ensemble du panel : la méthode scientifique et le processus d'enquête authentique doivent être activement préservés, et non supposés survivre passivement à la transition vers l'IA.

#

Souveraineté des données, science ouverte et licences équitables

Le troisième grand thème de la session a porté sur la relation entre les données scientifiques en tant que bien public et les risques d'exploitation déloyale ou de souveraineté des données trop restrictive. Vukosi Marivate a soutenu que les licences ouvertes standard telles que Creative Commons Zero (CC0) et CCBY ont permis une forme d'« écoblanchiment de l'ouverture », dans laquelle des acteurs bien dotés en ressources - en particulier les grandes entreprises technologiques - bénéficient de manière disproportionnée des données sous licence ouverte sans réinvestir dans les communautés qui les ont créées . Il a noté que Creative Commons est actuellement en cours de révision précisément en raison de ces préoccupations, et a établi un parallèle avec le mouvement copyleft dans les logiciels libres .

En réponse, Vukosi Marivate a décrit deux nouveaux cadres de licences conçus pour remédier à ce déséquilibre. La licence Esethu, développée par sa startup Lelapa AI, distingue entre les utilisateurs qui s'identifient comme africains et ceux qui ne le font pas : les utilisateurs africains peuvent utiliser les données à des fins commerciales ou non commerciales librement, tandis que les utilisateurs non africains sont limités à un usage non commercial et doivent négocier des arrangements de partage des bénéfices pour les applications commerciales . De même, la licence NOODL, développée à la faculté de droit de l'Université de Pretoria, distingue selon que l'utilisateur provient d'un pays développé ou en développement, exigeant un partage des bénéfices de la part des utilisateurs bien dotés en ressources . Ces cadres représentent une tentative de rendre les données localement ouvertes aux communautés qu'elles représentent, tout en empêchant leur exploitation par des acteurs externes disposant de ressources plus importantes.

Vukosi Marivate a également révélé une réalité contre-intuitive : malgré les mandats de science ouverte émanant des gouvernements et des bailleurs de fonds, les données sont déjà dissimulées, et les communautés « trouvent des moyens de les cacher encore davantage » par crainte d'exploitation . Il a cité la pratique courante des articles indiquant que les « données sont disponibles sur demande » tout en tenant rarement cette promesse . Sa conclusion n'était pas d'abandonner l'ouverture, mais de corriger le cadre des licences : « c'est ce qu'est la science - nous continuons à nous améliorer » .

Marion Mercier a proposé un recadrage complémentaire, suggérant que la souveraineté des données n'a pas à s'opposer à la science ouverte, mais pourrait au contraire faire « partie de la solution pour rendre les données ouvertes aux réseaux locaux auxquels elles devraient être ouvertes » , une vision que Vukosi Marivate a confirmée . Kamil Dziubek a ajouté une dimension technique à ce débat, avertissant que lorsque des données sont exclues ou restreintes des ensembles d'entraînement de l'IA, la question cruciale est de savoir si l'ensemble de données restant est représentatif et non biaisé . Dans des domaines à enjeux élevés tels que la découverte de médicaments, la recherche en STIM et la modélisation du langage, des données d'entraînement non représentatives constituent un risque critique pour la fiabilité des résultats de l'IA . Il a également réitéré que dans les sciences expérimentales, le test ultime de toute réponse générée par l'IA reste l'expérience physique ou l'étude clinique, qui fournit la validation finale et ne peut être contournée .

#

Pouvoir des entreprises et des agendas de recherche

Alistair Nolan a introduit une critique structurelle de l'écosystème de recherche en IA qui reliait la discussion sur la souveraineté des données à des questions plus larges de pouvoir et d'intérêt public. S'appuyant sur des recherches de l'OCDE publiées en 2023, il a noté que les grandes entreprises technologiques dépensent dans la recherche et le développement en IA des montants bien supérieurs à ceux des universités publiques, et que le taux de croissance de leurs investissements en IA est significativement plus élevé que celui des universités . Ces entreprises tendent à collaborer principalement avec des institutions de recherche américaines d'élite, dont les profils de recherche sont considérablement plus étroits que ceux du système universitaire dans son ensemble . La recherche sur laquelle elles se concentrent implique des types d'IA qui reposent sur une puissance de calcul élevée et de grands volumes de données - précisément les actifs détenus par les entreprises elles-mêmes . Alistair Nolan a soutenu que cela crée une boucle de rétroaction qui oriente l'agenda de recherche d'une manière potentiellement « préjudiciable à l'intérêt public à long terme », et a suggéré qu'un investissement accru dans des modèles plus petits, moins gourmands en énergie et en données, pourrait nécessiter une forme d'intervention structurelle .

Une membre du public, venant d'Afrique du Sud et étant membre de CODATA, a soulevé la question de la viabilité financière du modèle d'investissement dans l'IA, notant que les entreprises d'IA dépenseraient environ 1 400 milliards USD pour des revenus d'environ 613 milliards USD . Vukosi Marivate a répondu en soutenant que les nations, en particulier dans le Sud global, ne devraient pas se sentir obligées de reproduire ce modèle . Des modèles plus petits et spécifiques à des tâches peuvent rivaliser efficacement pour de nombreuses applications scientifiques, et la justification de dépenses massives est portée par la fausse promesse d'une « machine universelle » qui résoudra tous les problèmes de l'humanité . Il a averti que la bulle de l'IA pourrait éclater et que le développement d'une véritable capacité scientifique et technique est plus important que la dépendance à l'égard de grands systèmes propriétaires .

#

Réglementation, gouvernance et rôle des gouvernements

Une membre du public a soulevé le risque que les stratégies nationales en matière d'IA, en ignorant largement le secteur scientifique, puissent par inadvertance mal réguler la science. Cela pourrait se produire, selon elle, soit en appliquant des réglementations générales sur l'IA mal calibrées pour la pratique scientifique, soit en ne réglementant pas du tout . Elle a appelé à l'élaboration de codes de pratique et de protocoles au niveau des disciplines comme moyen de démontrer que la communauté scientifique est capable d'autorégulation d'une manière qui reflète ses propres valeurs, et a demandé quelles approches semblent efficaces pour accélérer ce processus de manière inclusive .

Moses Thiga a répondu en ancrant la question de la gouvernance dans la nature fondamentale des mécanismes de contrôle et d'équilibre scientifiques. La science, a-t-il soutenu, n'a jamais reposé sur des scientifiques infaillibles ; elle a toujours dépendu de systèmes de vérification et de responsabilité . Le défi consiste maintenant à repenser la manière dont ces mécanismes de contrôle et d'équilibre s'appliquent à l'IA, couvrant l'éthique, les données, la puissance de calcul et le déséquilibre de pouvoir entre l'industrie et le monde académique . Tout en approuvant le principe d'autorégulation, il a également soutenu que « la responsabilité incombe en premier lieu aux gouvernements », car seuls les gouvernements peuvent en définitive égaler les niveaux d'investissement des grandes entreprises technologiques . Il a cité l'investissement national de l'Inde dans l'IA comme modèle illustrant comment les considérations de souveraineté peuvent stimuler l'investissement public nécessaire pour contrebalancer l'influence des entreprises .

#

La méthode scientifique comme valeur fondamentale en jeu

Uun journaliste indépendant a alors demandé si la science, à une époque de réponses générées par l'IA en abondance, réside davantage dans la question ou dans la réponse. La réponse de Moses Thiga a recadré l'ensemble du débat : la valeur de la science ne réside ni dans la question ni dans la réponse, mais dans la méthode d'enquête elle-même, qui implique les étapes de l'enquête, de l'observation, de l'hypothèse, de l'expérience, de la collecte de données, de la découverte » . C'est ce processus, a-t-il soutenu, que l'IA est « sur le point de voler à la science » .

Cette intervention un peu plus philosophique était en lien directe avec l'intervention de Kamil Dziubek, qui avait affirmé que dans les sciences expérimentales, le test final est toujours l'expérience . Elle était également en lien avec la question de Marion Mercier sur l'interprétation par les humains d'une connaissance générée par l'IA qu'ils ne comprennent pas entièrement . Ensemble, ces contributions ont suggéré que la préoccupation la plus profonde partagée par les intervenants n'était pas simplement la qualité des données, l'intégrité de la publication ou la gouvernance institutionnelle, mais la préservation de la science en tant que processus humain et particulièrement rigoureux.

#

Conclusions et questions non résolues

La session s'est conclue sur un large accord selon lequel les défis posés par l'IA à la science sont systémiques, urgents et nécessitent des réponses coordonnées à plusieurs niveaux - des chercheurs individuels et des disciplines scientifiques, aux institutions, aux gouvernements et aux organismes internationaux. Le rapport du Conseil international pour la science a été présenté comme une ressource et une invitation à des contributions supplémentaires . Les travaux en cours de CODATA sur les mesures de qualité des données entre disciplines, la publication à venir de la Geneva Science and Diplomacy Anticipator issue de son premier comité IA pour la science, et le paysage émergent des cadres de licences de données équitables ont tous été identifiés comme des étapes pratiques dans la bonne direction.

Néanmoins, des questions importantes sont restées sans réponse : Comment réformer les stratégies nationales en matière d'IA pour prendre en compte le secteur scientifique ? Comment standardiser les mesures de qualité des données entre disciplines ? Comment gérer la crise d'intégrité dans la publication scientifique ? Comment prévenir l'érosion de la compétence empirique chez la prochaine génération de chercheurs ? Et, comment garantir que les bénéfices de l'IA dans la science soient équitablement distribués à l'échelle mondiale ? La diversité du panel a garanti que ces questions ont été examinées sous de multiples perspectives géographiques, disciplinaires et institutionnelles, faisant de cette session une contribution véritablement multidimensionnelle à une conversation mondiale de plus en plus urgente.

Dr. Vanessa McBride
Good. Thanks very much for joining us today in the session on science in the age of AI. We're going to talk a bit about knowledge, data, and trust, and specifically on the impact on how we practice science. So I'm just going to start with a kind of scene setting on behalf of the International Science Council and CODATA, and then I'll hand over to my colleague David Castle, who will moderate our esteemed panel, and we'll do some introductions of our panel members shortly. Thank you. So I think part of the conversation we've been having this week is really about how artificial intelligence is impacting many aspects of our society. But this is kind of a feedback loop because not only has science been fundamental to developing this kind of technology, but the impact of the technology is similarly feeding back into science systems and really changing the way that we do science. And if you just take a look at some of the articles that we've seen published in the scientific literature over the, this is just over the first six months of this year. You can see that the impact is incredibly broad on science. We're seeing the rise of AI agents to manage scientific workflows. We're seeing benefits, for example, the fact that AI tools are available to identify some predatory journals. But we're also seeing warning bills around adopting these kinds of workflows without the guardrails in place needed to do that. We're seeing the rise of AI agents to ensure scientific integrity and trust. and so that's really the setting in which we're discussing things today i wanted to highlight this this report that the international science council published earlier this year and it was on preparing national research ecosystems for ai it is a synthesis of case studies across 26 countries some of the case study authors are with you on the stage today of of course from south africa and moses from kenya and it was really looking at how national research ecosystems are preparing for ai and i think the thing we wanted to highlight is just that very first um learning from the report in that there are lots of national ai strategies under development and i think that's a very important part of the report development but almost none of them have any focus whatsoever on the science sector itself. We see a lot of health sector, we see a lot of agriculture, and we see how AI can be applied, but we don't see much consideration given to how it's changing the science that will result in new technologies and the applications itself. We welcome your input and feedback on this report. You can download it with the QR code or at the link. And then just to say about our panel today, we sort of thought that there were these three interconnected dimensions that we wanted to talk about today. We wanted to talk about AI and knowledge generation. We also wanted to talk about scientific data and how foundational it is for AI. And then we wanted to talk about AI and knowledge generation. We also wanted to talk about the downstream impact. The issues of reliability, trust, and research integrity. And so with that, I'm happy to hand over to David Castle, who's the chair for the Science Systems Futures project that we're working on, and to introduce our esteemed panel who are going to talk about
Prof. David Castle
Thank you, Vanessa, for the introduction. So we are a little bit behind, and we asked each one of our panelists, including Alistair, who's from the OECD and is joining us online, to say just a few remarks by way of their thoughts about the topic. We prepared a short briefing note about the panel and some of the questions that we wanted to raise. And so I think maybe out of courtesy to Alistair, we'd like to start with you, since you're away from us. And in person, I want to leave you to last. So did you have a few introductory remarks that you would like to share with us?
Mr. Alistair Nolan
Yeah, thank you very much. And it's an honor to be invited to this.
Prof. David Castle
ust a second. I'll just say we're just making sure we can hear you.
Mr. Alistair Nolan
Okay.
Prof. David Castle
Your pieces you see.
Mr. Alistair Nolan
Yes can you hear me now.
Prof. David Castle
okay good.
Mr. Alistair Nolan
Oh shall I go ahead okay very good so again thank you for inviting me I'm honored to be here um I'll just make a couple of comments very briefly one is about the structural importance for our economies and societies of advancing the productivity of science and as our economies age we will need science to feed into the development of more technologies to keep our economies productive enough to pay for the older age cohorts we will also need more discovery in order to address obviously the kind of global challenges that we face now say around the climate around disease and so forth um I am I saw you put up a question there at the beginning which will probably come to but overall I'm very bullish about the long -term implications of AI and science I think that they're very positive we see AI being used across every step in the scientific process and in every domain of science I think it's a wonderful adjunct to an amplifier of human intelligence I don't want to sound starry -eyed. There's a lot of hyperbole here as well. And I think that AI will create a series of institutional stresses that we'll
Prof. David Castle
Great. Thanks very much, Alasdair. Why don't we go with Camille?
Dr. Kamil Dziubek
So I'm Camille Tubeck. I work at the University of Vienna, and I'm also in CODATE. I'm co -chairing the task group on research data quality management across the data lifecycle. I choose one working example to start with, and just a question to everyone. Please raise your hand if you have heard about AlphaFold. So is most of people in this room. So AlphaFold is a family of programs that actually can predict based on the AI. The AI models, the three -dimensional structure of the protein. and it was a great invention. It earned the authors of this program the Nobel Prize in Chemistry in 2024. It's based on the deep learning algorithms and it uses the training sets from the experimentally determined protein structures. Those data sets are actually in the Protein Data Bank, which is a large database, over a quarter of a million of the structures, determined with experimental techniques like X -ray diffraction, cryo -EM, or nuclear magnetic resonance. But there is a catch about that, because all those data are models. We've never seen the molecules with our naked eye, we've never seen the molecules under the microscope, because they are too small, so we have to reconstruct them from the raw data, from different physical techniques. and they have all the problems that models have, actually. There was an aphorism attributed to the British statistician George Box that says that all models are wrong, but some are useful. So we are looking for the useful models, actually. We try to eliminate the models that are biased. We try to eliminate the models that are of the low quality. We try to eliminate the models that don't fit into the context of chemistry and physics that we know. And now there is another story, because it's not only about the quality of the models. You can verify and validate this quality, but there is also a question of these models being updated. So we will speak about the ground truth in the AI, this ground truth is a moving target, actually. every day you have the new data sets and every day you're checking against something else because the structures the ground truth of yesterday is not the ground truth of today now how we can actually know that our structures and the structures that are predicted with the AI methods are good we need to validate them we need to have the checks to validate them we need to define them and agree on them and actually this is the big question about the data quality so we need to define the measures of data quality across the disciplines this is what we are aiming for in Codata at the moment landscaping the different measures for data quality and asking people Because if you ask about the quality of the training data, you have to ask what are the measures of the quality, which you have to ask about the uncertainties, about the accuracy, about the traceability, about the provenance. And if the people who actually use the AI methods, they don't provide you with this data, it's a big red flag. So, in principle, just to wrap up, if you say AI for good, there is this big slogan, AI is for good if it's based on good data. And we need to have some measures to check if the data are good or not. So, I'll finish.
Prof. David Castle
Great. Thank you very much. Moses, why don't you go next?
Dr. Moses Thiga
Thank you. My name is Dr. Moses. Thank you. I'm Dr. Moses Iga from Igaton University in Kenya. And taking you on a different tangent, AI. is impacting the science system. And I come from an environment where we are still trying things out, we are still thinking things through. We still have a few people who say no way to AI, and we have shadow AI, you know. So in essence, AI in the global south, which is what I see it, in an academic context as an individual who is charged with driving the use of technology in my university. We see AI as a double -edged sword. And in the global south, it's given us the ability to leapfrog a lot, you know, do things that we possibly wouldn't be able to do, because AI gives us the ability to simulate laboratories and data and, you know, do lots of analysis, you know, get your literature done, your brainstorming. It's absolutely wonderful when it comes to that. But the real danger that we are seeing, at least internationally, my institution in my country is... a generation of scientists who are empirically incompetent. There's a lot of research going on, a lot of publication going on by some researchers who actually cannot do that experiment in a real lab, in a real world setting. So we are finding researchers who cannot actually read a paper. You know, they can't read, they can't do literature review, they cannot collect data, they cannot analyze data, they cannot critically evaluate or analyze the outputs of their said research. It's such a big problem and that's what is happening in the global South. Now, that is compounded by academic leadership and research leadership that is not conversant with AI. They don't use AI. You have this scenario where AI is demonized. so then it goes into the shadows and we're sort of in a very big confusion if I may say there is beginnings of policy and strategies coming up across the global south but I say not much of that comes from an experiential perspective that said again we are still at a disadvantage a lot of the models again predominantly fed by western data our data systems are not yet robust still not even maturing are very nascent stages so while AI is presenting opportunities to leapfrog it's also creating another scenario where again we might be left really far behind without the capacity to develop our models, we don't have the data infrastructure, the computes we still have to buy that we don't have native infrastructure in the global south so AI is doing good in science but presenting lots of more dangers.
Dr. Marion Mercier
Thank you. Hi. Thank you. So I'm from, I should correct, I'm from the Geneva Science and Diplomacy Anticipator. It's fine. The only reason I'm highlighting it is because anticipation is so key to what we do. So just a little bit of background to explain. We're an independent non -profit foundation that was founded by the Swiss government and the canton of Geneva with the mission of working with scientists to anticipate the science and technology advances that may be coming over the next 5, 10, 25 years and to then work with various stakeholder communities to have forward -looking dialogues and ensure that these advances can be developed in a way that will be beneficial to society. So it's a pretty big mission and what that looks like in practice is that we work with scientists from a very very broad range of fields we look at lots of different topics ranging from the social sciences all the way through to very fundamental disciplines like maths and we ask them to give us their vision of what they think might be possible they're not predictions they're possibilities over the next 5, 10 and 25 years and that puts us in a position to really see the pervasive catalytic impact that AI is having across every single one of the disciplines that we look at and it's very interesting to see the kind of the way that AI is being applied to very specific problems in different disciplines but I think what's more interesting is the broader question that's coming out which is how AI is changing the very practice of science and the pursuit of knowledge you and I think if I can use an example from one committee from one anticipation workshop that we did that really struck me and that illustrates this So we did a workshop on cognitive enhancement, so neuroscience -based, and all of our workshops are structured across some sub -themes, and they often range from the very fundamental to the more applied. And so in this case, we started off with discussing fundamentals of cognition. How well do we currently understand the brain, and do we think, you know, how is that going to advance over the next 5, 10, 25 years? And the conclusion there was that, you know, we still don't really understand the brain, and there's lots happening, but really we're not very close to kind of crafting it. And then we got to the end of the workshop where we were discussing brain -computer interfaces and how that technology is developing, how AI is developing, and how by bringing those two together, you know, these implants and even actually the non -invasive technology, you bring that together with AI, and you get to the point where actually AI is going to understand the brain. Even if we don't. And I thought that was so striking because... it raises so many implications, not least the question of interpretability. You know, if AI understands how the brain works and it can do all the things that we're trying to achieve with neuroscience, like, you know, modulation, treating diseases. But if we don't understand how it's doing that, you know, are we OK with that? Are we OK with the outcome without it advancing human understanding? So, you know, there's a lot of other implications, but I just thought that was a really good illustration of some of the things that we're faced with in terms of bringing AI into science and into discovery. And just to close to say that because we're seeing this kind of wide impact this year for the first time, we had our first committee, anticipation committee devoted entirely to AI for science, which was chaired by Hiroaki Kitano, who recently or not recently, I should remember the date, but he is leading on the Nobel Turing Challenge, which is where he's launched the challenge. To develop AI that will be able to develop Nobel worthy insights. so that was a very interesting discussion which I think lots of insights from that which will be published soon and lots of insights from that will be relevant to today's discussion. Thank you.
Prof. Vukosi Marivate
Thank you, I'm Vukosi I am one of the members on the UN independent scientific panel on AI so yeah, it's also being scientists looking at the science to try and give perspectives. I'm at the University of Pretoria as the director of the African Institute for Data Science and AI and I'm also the chair of data science there and then on the other side I have a startup company I'm a co -founder of Lilapa AI where we work on building language systems using AI and I'm on sabbatical this year trying to think about assessment or evaluation for AI models and how we can improve that, because we need it for being able to scale, especially in areas where you don't have enough representation, whether it's the way that evaluation is built, models are built, or the data. How does it represent the majority of the world? Yeah, but with that, maybe I'll start off and say I come to you from the future. AI is boring now, and it is not like, you know, it's something that we've accepted, it's pervasive, and we've now learned to live with it. Unfortunately, today, we have to go through it. We have to go through all of the emotions, all of the hurt, all of the opportunity of what AI is doing to much of the way that we think about research and science. So, like, just briefly. building on everything that is already being said, both online and by the fellow panelists who are in the room. There's much that is going on. I come from natural language processing as a researcher and AI, so at the moment I am area chair and senior area chair for two different conferences, one being in natural language processing and the other one being NeurIPS. It's a mess. I'm telling you from the AI researchers, it is a mess. We've gone maybe from, hey, a few years ago you were getting during submission cycles maybe 2 ,000 to 3 ,000 papers in submission cycle. Oh, just to give you context, in ACL, which is Associated Computational Linguistics, we have something called rolling review. So our major areas is not journals, but it's conferences, so there's a number of conferences a year. Let's say there's eight. So initially you used to just try to get your paper into one of those conferences. And you would get a deadline. And then there was a decision to say, let's just have rolling review. I think at the beginning was every six weeks you could submit a paper. and it went into a common pool. There would be reviews, and then after your paper gets accepted, you could choose which conference you would then present that your paper would then be published in those proceedings. Already before AI, it was overwhelming, and then I think they moved it to like an eight -week cycle or ten -week cycle. That's what we're living on now. This cycle that just finished, which we are now doing reviews for, had 13 ,000 papers submitted. And now there's a question that comes in from Integrity. How many of those are like, you know, a, what is it? It's not a good effort, but it was actually true. Somebody doing some experimentation or trying to show, like doing the science and then going to whether AI assisted or not. And what other ones are just, it's just completely AI generated, has no real. Kind of goal, it was just, can you please get me like, you know, write a paper similar to AlphaFold and get it out there. Right? So that's what's happening to the people who are building the tools. That's what's happened to our practice. Right? And then in NeurIPS, which is the other one, NeurIPS typically now has like 20 ,000 to 30 ,000 participants who come to that conference. It's crazy. So the amount of papers they are is now we're going through and checking are the references actually real or not. And this is, again, like I know there's this tension that comes from other parts of science who are saying like, but we're being affected. I'm like, oh, it's eating us as AI researchers. So we're going to have to go through this. We're going to have to figure out like new ways of thinking about what is scientific integrity, show adjustments to where discovery comes in, and then what it actually is. Then it's going to change in our practice. It's just that it's going to affect us all differently. Moses was saying, if you're coming from the Jehovah majority, it's going to be already, there was like a lot of like vibrations in the system that we were trying to get out. Now this is just making it worse. Right? So you can amplify good. And you can, that's the things with technology, right? It's an amplifier. It can amplify the good. And we can try to reinforce that. And we have to reduce the downside as much as possible. And that's the big thing that we're going through. And that's what I just wanted as like an opening statement of saying like, yes, from the future, it's boring. Right now, it hurts. And then at the same time, it's amazing. It's amazing the kind of experimentation you can try out that you couldn't. There's even thoughts I've been having, I actually want to go back to my PhD 11 years ago. And there's all these things that I'd left on the table that I think I can explore now. And that's where we are.
Prof. David Castle
Thank you very much for all your interesting opening remarks we also had some questions prepared three of them and so we uh we're about half past now so what we'd like to do pardon me yeah do all the questions up at once oh we had four sorry um so uh what we do is maybe spend a few minutes on each one of the questions and then take questions from uh the audience and uh perhaps there's some things that have come in in the chat on online that Felix will be able to uh tell us about so um anybody who wants on the panel to including of course you Alistair um as AI continues to permeate various aspects of the research process what are the long -term implications for the production of scientific knowledge so who would like to comment on this particular question.
Dr. Moses Thiga
Um I could in a sentence um We probably need to redefine knowledge And what is science What is research, what is learning That's probably changing a lot How do we even teach research in the first place What kind of skills do we need today And if anybody is not Rethinking this It's probably not futuristic at the moment AI is creating illusions of competence You know Of capabilities, lots of illusions Institutions must rethink What are we here for What's a university for What do you do in the lecture hall You know, what's the purpose of that Because the knowledge is all out there Everything is out there So what are you teaching What is research, the literature can get done Hallucinations notwithstanding So what must institutions be In this day and age I think institutions probably need to focus On having better compute Than better campuses Because that's the future of science They need to probably think of Being ethical gatekeepers teach ethics, teach good science, you know, like there is a process to it. There's a human at the end of your discovery. You know, that's probably.
Prof. David Castle
Thank you very much. I see Alistair, you have your hand up as well.
Mr. Alistair Nolan
Yeah. If I just make a comment on the first question, I think that a bottom line for me is perhaps that discovery itself, the generation of new knowledge may become less of a rate limiting factor in the way that science has an impact on our technology, on our economies and societies. So, for example, in the era of material science, many of our fundamental challenges around the climate, around battery technology, et cetera, buildings that can cool themselves. So material science is a fundamental, will be a fundamental contributor to that. And there's a lot of sort of a lot of hyperbole, actually, and a lot of enthusiasm about the role that can play material science. But it turns out we've just been doing work on this subject that it's not actually discovery, which is the primary bottleneck. It's translating discovery of what you find in laboratory to an industrial scale when you're producing this new material in terms of thousands of millions of tons. So we need the new technologies. And I think discovery will be less of the challenge in getting in getting those discoveries. Another very briefly, another long term implication is I think about who will be producing new knowledge. It won't be that anyone could do science. Through the instrument of AI, almost certainly not. But what is already happening is that AI can help bring more people into large science projects, say through citizen science initiatives. But more importantly than this. It will help students to engage in higher level thinking about novel problems at a younger age. So you may have seen the recent remarks. You can find them on YouTube and just Google them by Terence Tao, who's a Fields Medal winner in mathematics and a professor of maths at the University of California, Los Angeles. And he recently pointed out that his main argument is the promising young mathematicians will now be able to reach the point where they can seriously experiment at the frontier sooner because less of their cognitive bandwidth is going to be spent on lower level tasks like memorizing large bodies of literature, developing an aptitude for long calculations and long proofs and so on. So I think those will be two points where I think we'll see changes. The discovery will become less fundamental and as on the path to new technologies. And we'll see a difference. In who's generating new knowledge as well.
Prof. David Castle
Okay thanks thanks Alice um so let's go on to the second question here and see what other of our panel wants to say what what other um uh areas do you think that uh yeah it might contribute to generating scientific knowledge we provocatively use the word independently here but maybe in say some sort of supervised or semi -supervised environment any any thoughts about where this will happen?
Dr. Marion Mercier
um so i'm absolutely not an ai expert but just again um just insights that we get from these anticipation committees that we have and i think you know drug discovery is definitely one of the areas where we're seeing you know huge impact and scope for automation along every part of the pipeline um and one of the anticipations that we had in one of our in a 25 -year time frame we might start to see drug discovery time cut from years, which is where it currently stands, to a matter of days due to AI mining of the clinical data to synthesis of chemical compounds. And so I think that's definitely an area where there is scope for automation and where it would be very welcome as well. I think it might not be so welcome in other aspects of science, but I think in drug discovery, there's a lot of room for it. Great. Thanks. Any other quick... So I'm absolutely not an AI expert, but just again, just insights that we get from these anticipation committees that we have. And I think, you know, drug discovery is definitely one of the areas where we're seeing, you know, huge impact and scope for automation along every part of the pipeline. And one of the anticipations that we had in one of our committees was that within a 25 year time frame, we might start to see, you know, drug discovery time cut from years, which is where it currently stands to a matter of days due to AI mining of the clinical data to synthesis of chemical compounds. And so I think that's definitely an area where there is, you know, scope. For automation and where it would be very welcome as well. I think it might not be so welcome in other aspects of science, but I think in drug discovery, there's a lot of room for it.
Prof. David Castle
reat, thanks. Any other quick thoughts from the panel on this?
Prof. Vukosi Marivate
Yeah, so I'm working in natural language processing the things that people look at even in building language models it doesn't need to be an LLM, it's still very much based around thinking about English or very high resource languages there's so much opportunity here and you might say hey there's a bit of engineering here of now discovering more and more about our human knowledge by being able to model more and more of our lower resource languages and that's a place that yes some of the recipes are repeatable and being able to explore that so in terms of being independent it's there and then being able to then hit that kind of the edge and then go over and see and say like oh for these languages that have these different types of scripts here's actually new rules that we kind of didn't anticipate and new knowledge so i'm excited for that uh but yes again uh being a tool how you wield it is going to be interesting and that's the thing we're going to have to as mozo say uh teach the new scientists what to what to do better and also online um it's it's going to be really important and and maybe one thought to leave for everybody is it's i've never i'm not a person who likes writing notes down on paper uh but i've never had as many notebooks as i have right now uh because in now on thinking and being able to sit with your thoughts and actually then yes it's easy to to kind of be in in the prompt box but then you don't think but then while you're sitting and you can sit with a pen like you know with a pen and and and paper now you're learning to wield the AI more as a tool and not do the thinking for you in a way. So that's another.
Prof. David Castle
Great, thanks. So as Vanessa said, we've structured this forum as an opportunity to talk about AI and science systems, which we've done a little bit. Now we'd like to go to the other end of the triangle and talk about the data issue. And I wonder, Camilla, if you would like to comment on number three, because this question is really actually about what's at stake in terms of thinking about science as a public good, a generator of knowledge that all people can benefit from and potentially use. When in fact we do know that at the core of AI is data, and as you said at the beginning, you were focused on data quality, but then this is a question about access and use, and how do we protect against unfair or unwarranted access versus over -restrictive data sovereignty and protection of data on the other hand. So what would you say about these issues briefly, Agnes?
Dr. Kamil Dziubek
Thank you. So in principle it also goes back to the previous question because you mentioned the scientific knowledge and there is a difference between data and knowledge. There is data, information, knowledge and wisdom if you remember the pyramid. In principle I know that CTI will massively create data. The question is about the quality of the data and the usefulness of the data. So in principle the risk in data harvesting I can say that if we massively produce the data, there is a bottleneck at some time. And we will not be able to actually comprehend the data. Then of course we'll have to say if all the data should be open and fair because fairness is not exactly the openness we need to safeguard the fairness, we need to safeguard quality of the data but of course be of course look also at risks with data sovereignty and some issues with data with the personal data for example so just need to take the balance but I guess the data deluge is a big problem at the moment and it will be in the near future
Prof. David Castle
Okay, thank you, you have a comment? Go ahead.
Prof. Vukosi Marivate
yeah, so so So here, in terms of thinking about the universally accessible, yes. But thinking about balance and equity, it is likely that we have to think about things like equitable licenses. So this year, if people don't know, Creative Commons is going through a review because they've seen that there's a challenge of open washing, that the people who do have the compute, the engineers, the scientists, then really benefit off, they accrue most of these benefits of data as available, and it's been made open, like let's say just on a CC0 or a CCBY, while they don't necessarily invest back into the open source software communities that have gone through this before, and that's why you have copy left. And now we're going to do that in data as well, right? If you want to see this, put our data on CCBY SA, and sometimes you get hate mail. right because people don't want to actually continue really like you know building on the data that you've released and also releasing those um those things at the same time you need we need to understand that there's parts of the world where then there is a benefit that can accrue um uh to to smaller players if they did get access so for example um uh for people there's at the university of pretoria in our health uh scholars i'm sorry law um uh faculty there's a new license called the no letter or bordeaux open data license so noodl which does kind of discriminate by where you are accessing uh the data from where they are coming from a developed country or a developing country and then it then says hey you can either use the data or whatever freely commercial non -commercial reasons but if you come from a developed country uh you must then get back to the communities that have created that data and then talk about benefit sharing At Lilapa AI Our startup company We have the ESETU license E -S -E -T -H -U And that one says do you identify as African or non -African If you are African You can use commercial or non -commercial Doesn't matter If you don't identify as African You can only use that data for non -commercial reasons And for commercial reasons Again get in touch with The creators or the communities That that data represents And then negotiate benefit sharing These are things that are going to Because we are not on a level playing field So saying And this is some of the challenges that The communities have seen Of just saying we have open science We have open data What then happens is actually you have people hiding Their data And it doesn't show up And the reason I know about this is because I work on language And language data is there It's just that people are trying to make sure it stays in their garage And tapes and under beds Because they are just Just
Prof. David Castle
Great. Thank you very much. I see that, Alistair, you have a comment about this. And then if we could maybe hear from you, Alistair, and then take a couple of questions from the audience or online. So, Alistair?
Mr. Alistair Nolan
Just an observation that links the issue of equity and sort of power in this AI ecosystem and data. And that's the nature of the research which is being done by the institutions that have the deepest pockets. That's the large tech companies that are outspending by multiples, the kind of AI R &D which is done in public universities, and the rate of growth of their investments in AI R &D much higher than in the universities, and they're drawing off talent, as we know, et cetera, et cetera. So, we put out a publication in 2023 that amongst others things showed that the large tech companies... They tend to collaborate with the elite in the United States, the elite research institutions. And you look at the breadth of the research profile of these elite research bodies, and it's much narrower than is the case for the rest of the university system. And what are they concentrating on in their research? They're concentrating on types of AI that rely on high compute and large volumes of data. Now, these obviously are the assets which are held by the companies that they're working with. So there's a sort of steering of the research agenda in ways which I think would be in the long term prejudicial to the public interest. So it would be good, for example, if there was more investment in less energy, less data, compute hungry at smaller models, for example. And I think that may require some kind of
Prof. David Castle
Great. Thank you very much for that. That's very insightful and interesting. Okay. We started a bit late, but we'd like to take any questions that we could. from the audience. So anybody want to ask? Okay. I think you were first. You, your man. Can you use the microphone? Thank you.
Audience
I'm Daisy. I'm from South Africa and part of Core Data. And I just want to check with the panelists their thoughts around the involvement of AI and what's going to happen, especially what my counterpart from South Africa, Fugosi, highlighted. We are aware that AI companies are spending 1 .4 trillion. That's the loss. 1 .4 trillion. And the revenue is just 613 billion US dollars. So is this model sustainable? How are we, especially in the global south, we're leapfrogging and this is happening? Thank you.
Prof. David Castle
Yes, yes, yes. Go ahead.
Prof. Vukosi Marivate
Yeah, yeah, quick one. I think I will agree with my fellow panelists online. You don't need to follow this model at all. It's just, it's like, you know, sometimes we get nations coming and saying, we should build our own insert AI model that requires a lot of data, a lot of compute, because we need to show that we can do it. Or we need it because that's the thing that gives you something. Like, no, you don't. You can build small language models, if you're in language, that can compete depending on the task and what you actually need. You do not need everything machines. You can say, I'm working, like, if you're saying, I'm working on protein folding in this specific way, and I can have a model that is literally a couple of megabytes, and it actually does well. Why? Because in science, is that you go in and you say, a model is an abstraction of the real world, and I was able to figure it out, and it actually works. So there is, yes, the data -driven way, and it is one way, and we keep on seeing the models as they get bigger. But at the same time, the reason you're looking at this spending that is ridiculous is because of the promise of this is an everything machine. They will solve everything, and as such, we just have to keep on throwing cash in, and then we're going to get to a point of human nirvana, where we all apparently just get money for living. This is the thing that then becomes very enticing to just say, hey, as countries, don't think about your sovereignty, because we just need to keep on shoveling money and data into these systems, and we will solve everything in humanity, and that's not true. right and that's the part that we have to that is not true and it doesn't mean you shouldn't be doing basic high impact research work do it because we still have to learn because the bubble is going to burst or it's going to be like you know somewhere that it actually has burst and we need to then survive past that as humanity so let's yeah that is just one of the we yeah there's more to do here.
Prof. David Castle
There was one other question I think it was yours oh sorry I'm really sorry about the position here anybody else so you and then you and we'll see if we have time for an online. Okay, go ahead.
Audience
Thank you very much so see my from the international federation of library associations I think Vanessa you mentioned at the moment that a lot of national AI strategies don't actually mention science and there's still a risk that either they regulate science or they regulate it and they regulate it and they regulate it by accident because they haven't bothered thinking about it and they're just bring up science in the rest of things or indeed they get worried and say science, not sure about that and they actually regulate without thinking I think this all points to the value of developing some of these codes of practice these protocols on a disciplinary basis in order to demonstrate that the community is still capable of regulating itself in a way that actually supports the values and I suppose the question is what seems to work, what opportunities are there to accelerate this process of developing the ability to update protocols, to update ethics in a way that is truly inclusive, that doesn't hold everyone to one particular standard and so have that answer, have that way of saying no, science can actually regulate itself we can find solutions within the community rather than needing to depend on government regulation coming in which may not.
Prof. David Castle
That's a really interesting comment Moses, do you have something to say in response to that? I feel like that's up your alley
Dr. Moses Thiga
You know my thoughts on what you said is it borders around regulation. And science, the science system has never been about infallible scientists who cannot make mistakes. It's about checks and balances. And the way forward really is, how do we continue with this regulation? There is now a new player on board. So how do we now regulate how we do this? How do we do that? I think that's really the point around regulation. We need to rethink how we regulate. And we regulate on ethics, on data, on compute, and back to the issue of the power imbalance, the industry, academia, high end. I think the responsibility falls at the floor of governments. Only governments really, at some point, can match what some industry players are doing. and please read the case of India I'm always fascinated at the kind of investment they are doing to match what industry is doing it's a matter of national pride it's a matter of sovereignty it's probably how things will work going forward
Prof. David Castle
Thank you one more question from here and I think we have one online and then that will be our session thank you.
Audience
I'm just an unbearable local journalist freelance so two questions I heard that in Kenya the quality or at least motivation of academics and researchers is poor but the motivation and quality of students is outstanding since I watched a documentary saying that half the PhDs in Oxford and even research and lecturing is indeed outstanding written in in coffee shops in Nairobi by penniless students. So should I conclude that the more you rise, you step up in an academic career, the less quality and motivation you have for science? So that's my first question. The second one, I cannot find a good example. I will take a lousy one. One, is science, especially when total easy information is available, more in the answer or in the question? For instance, I have a question as a non -scientist, is there a common origin to the word mother and water because in many languages ma, Mayim, may can be encountered? Or another question, why has Latin left so little, trace in Arabic -speaking countries? So. probably with AI I can find a lot of answers I tried, I haven't found but I think with good tools I could find a lot of things making possible advances in research what about the question, in that case is the real science in the question and why do people suddenly raise a question which more learned people didn't raise or is it in the answer we also know that deciphering the Maya scripture came from the experts being on holiday and the child being brought there just for holiday started deciphering and he didn't have the prejudice of learned people so he went through, so I don't know if you.
Dr. Moses Thiga
get your very well framed question. I'll do it quickly. The paper paper mills did very well Before AI So the business is no longer viable A lot of the students That did the papers That are rumoured to be published At Oxford etc That was the pre -AI era That business is no longer Viable, it's really not working Are the students more motivated? Yes for money Are the faculty and researchers less motivated? In a sense because Not much research funding is available And support from the government But that was pre -AI Is the value in the question or the answer When you think about the scientific Method The value is in the method And this is where we are losing it as scientists This is where we are losing The whole thing about The inquiry, the observation The hypothesis, the experiment The data collection The discovery And this is what AI Is about to steal from science The beauty Of the AI Of the inquiry so the value is neither in the question nor in the answer but the beauty is
Audience
this is an excellent title for a book write the book, I'll be your first reader.
Prof. David Castle
There you go Moses alright let's take one question online I think it's from Nipu can you use a microphone yes because it's online as well.
Audience
So I'm wondering if I'm wondering if the sovereignty of data in AI would be a threat to open up open science ok.
Prof. David Castle
Good question we'll take that and we'll also take the one online as you suggest especially in the global south great question ok and a question online Nipu unmute not. Not there. Okay. Then we're just going to close with this one question and let's move on. If Parley can... Parley, do you want to unmute?
Audience
My name is from women in technology in Nigeria so because I work with a lot of scientists who we are pushing for citizens data so I'm just wondering if that push for data sovereignty will threaten the open science we are all fighting about.
Prof. David Castle
Alright good nothing from online we've asked twice panel in response to this last question sure quick one.
Prof. Vukosi Marivate
Already at the moment even as I said like with the pushes for open science and open source open data I'm not making aspersions on code data or other other things I work in language. Finding language data has been very interesting, understanding traditions that different fields have about how to treat. So even with open science requirements by governments and all those things, data is hiding. Data is hiding. And people are figuring out ways to hide it even more. Right? We, in machine learning, used to go like, oh, show me, like, you know, you say here's the experiment that you did. Show me the data. Oh, and then they put it in the paper. Paper, data available on request. Lots of papers have come out showing that that never happens. Very rarely when you ask do you get the data. Right? And there's all these things. But, yes, machine learning and AI has also progressed very rapidly because it's become easier to just literally you see the paper and they have a link and you click it and it brings up a notebook and you run it and it pulls the data. There's been huge things. But we do understand the challenge of people then saying we do not want to be exploited. This has happened in. in machine learning where people work on health data, agricultural data. I sat on the Lacuna Fund Steering Committee where we were funding people to create data for lots, and we were requiring either a CCBY or more liberal license. And then what people said that came back as feedback to us as an organization was, oh, it's great that you want us to have a kind of open data at the end of this, but then the people who end up getting the recognition are the big tech companies who take that data and then release an update. And that's the thing that brought these huge conversations about, oh, what should we be doing about the licensing in such a way that people, and that's why, whether it's a CETU or it's Noodle or the new reviews that are coming in, because Creative Commons used to have a developing country license. So it's not that we just abandon it. We fix it. That is what science is about, right? We keep on. We're improving.
Prof. David Castle
I wonder, Marion, do you happen to have a comment about this as well from your perspective? And then we'll wrap up the session.
Dr. Marion Mercier
Not so much. More, I had a follow -up question, which was rather than, you know, is data sovereignty a threat to open science? Is it actually, you know, part of the solution to make data open, you know, to stop people hiding their data, like what you're describing, and to make data locally open, you know, to the networks where it should be open to, rather than, so, yeah, rather than being opposed. That was how I kind of understood your answer.
Prof. Vukosi Marivate
That is the part that it's trying to do.
Dr. Marion Mercier
eah.
Prof. Vukosi Marivate
The local network you're trying to impact.
Dr. Marion Mercier
Yeah. Yeah, sorry.
Prof. David Castle
Okay, great. Thanks very much. Did you have one short? Okay. Very short comment on this.
Dr. Kamil Dziubek
So the main threat, if you don't have fully open data, is you have to ask yourself a question. Is the data set representative? Is it not biased if it's not fully open? And if you're sure that the data that you're excluding from the data set doesn't cause that it's biased, it's all right. But if it causes some kind of bias, it can be really crucial. And I'm not talking about the data in languages, the data in STEM, for example, the data for the drug discovery. If you don't have the input from different groups. And one very short comment also to the gentleman who was asking the question about the question or the answer. In experimental science, the final test is always the experiment. So if you have the answer from the AI, what type of drug you're designing or what type of the material you're designing, there is a clinical study or there is an experiment. And then you know.
Prof. David Castle
Great, thank you Thank you panel, it was a wonderful session Thank you Alistair for joining us from Paris and audience please join me in thanking our panelists for this interesting conversation It was a good session

Avertissement : Il ne s'agit pas d'un compte rendu officiel de la session. DiploAI génère ces ressources à partir d'enregistrements audiovisuels ; elles sont présentées telles quelles, y compris d'éventuelles erreurs. En raison de contraintes logistiques (audio/vidéo ou transcriptions), les noms peuvent être mal orthographiés. Nous nous efforçons d'être aussi précis que possible.