WSIS Forum 2026
AI-generated report

Science in the Age of AI: Knowledge, Data, and Trust

8 speakers
Summary

This discussion, moderated by Prof. David Castle, brought together an international panel to explore the impact of AI on scientific knowledge, data quality, and research integrity. Dr. Vanessa McBride opened by noting that AI is not only a product of science but is now fundamentally reshaping how science is practised , and highlighted that most national AI strategies fail to address the science sector itself .

On the question of data quality, Dr. Kamil Dziubek used the example of AlphaFold to illustrate how AI models depend on training data that is itself model-based and subject to bias and error , arguing that robust measures of data quality, including accuracy, provenance, and traceability, are essential for trustworthy AI-driven science . Dr. Moses Thiga raised concerns from the Global South, warning that AI is enabling a generation of researchers who lack fundamental empirical skills , and that inadequate data infrastructure and compute capacity risk widening existing inequalities .

Dr. Marion Mercier highlighted how AI is transforming entire disciplines, citing drug discovery as an area where AI could reduce development timelines from years to days , while also raising the deeper question of whether science can remain meaningful if AI generates knowledge that humans cannot interpret . Prof. Vukosi Marivate noted that the sheer volume of AI-assisted submissions to academic conferences is creating serious integrity challenges , and emphasised that AI amplifies both the good and the bad in existing research systems .

On data sovereignty and open science, Marivate argued that equitable licensing models are needed to prevent large tech companies from disproportionately benefiting from openly shared data , pointing to initiatives such as the ESETU licence as practical alternatives . Alistair Nolan added that large tech companies are steering the research agenda towards high-compute, data-intensive AI, potentially at the expense of broader public interest .

The panel broadly agreed that institutions, governments, and the scientific community must rethink research training, regulation, and data governance to ensure that AI serves science equitably and responsibly .

Keypoints
  • Overall Purpose

  • The discussion aimed to explore the multifaceted impact of AI on scientific practice, knowledge generation, data quality and access, and research integrity. Convened by the International Science Council and Committee on Data of the ISC, the session brought together panellists from diverse global contexts to examine both the opportunities and risks AI presents to science systems, with particular attention to equity and the Global South.
  • --
  • Major Discussion Points

  • AI is transforming the practice of science across every stage of the research process, bringing both significant opportunities and serious risks. Panellists noted that AI is being used across every step of the scientific process and in every domain , including managing scientific workflows and identifying predatory journals . However, warnings were raised about adopting AI-driven workflows without adequate guardrails . Vukosi Marivate illustrated the scale of disruption by describing how AI has contributed to a surge in paper submissions - from a few thousand to 13,000 in a single cycle - raising urgent questions about scientific integrity and the prevalence of AI-generated content with no genuine scientific contribution .
  • Data quality is foundational to trustworthy AI in science, and the "ground truth" used to train models is inherently dynamic and imperfect. Using AlphaFold as a case study, Kamil Dziubek explained that AI models in science are trained on data that are themselves models - reconstructions from physical techniques - and are therefore subject to bias, inaccuracy, and obsolescence . He stressed that yesterday's ground truth is not today's, making ongoing validation essential . The principle was summarised as: 'AI is for good if it's based on good data,' requiring agreed measures of quality, uncertainty, accuracy, and provenance .
  • The Global South faces a compounded disadvantage: AI offers the promise of leapfrogging, but risks deepening empirical incompetence, data exclusion, and technological dependency. Moses Thiga highlighted that while AI enables simulation, literature review, and analysis in resource-constrained environments , it is simultaneously producing a generation of scientists who cannot conduct real experiments, read papers critically, or analyse data independently . This is worsened by academic leadership that is not conversant with AI and by the fact that most models are predominantly trained on Western data, while data infrastructure in the Global South remains nascent . Marivate added that AI amplifies existing inequalities, worsening problems for the Global Majority .
  • Data sovereignty, equitable licensing, and the concentration of AI research power pose serious threats to open science and the public interest. Marivate described how open data mandates have led to 'open washing,' where large tech companies disproportionately benefit from openly licensed data without reinvesting in the communities that created it . In response, new licensing frameworks such as the ESETU licence and the Noodle licence have emerged to ensure benefit-sharing based on geographic or economic status . Alistair Nolan reinforced this concern by noting that large tech companies are outspending universities on AI R&D, collaborating primarily with elite US institutions, and steering research towards high-compute, data-hungry models that serve corporate rather than public interests .
  • Institutions, governments, and the scientific community must urgently rethink regulation, ethics, and the very definition of scientific knowledge in the AI era. Moses Thiga argued that universities need to reconsider their fundamental purpose - shifting focus towards ethical gatekeeping, teaching good scientific method, and investing in computing rather than campuses . An audience member raised the risk that national AI strategies, which largely ignore the science sector , may inadvertently regulate science poorly or not at all, and called for discipline-level codes of practice to demonstrate that the scientific community can self-regulate . Thiga responded that the checks-and-balances nature of science must be preserved, and that governments - citing India as a model - may need to match industry investment to protect scientific sovereignty .
  • --
  • Overall Tone

  • The overall tone of the discussion was earnest, intellectually engaged, and cautiously optimistic, with recurring undercurrents of urgency and concern. The opening remarks were largely scene-setting and measured, with Dr McBride and the panellists framing AI's impact on science as broad and consequential . As the conversation progressed, the tone became more candid and, at times, frank , introducing a more sobering register. However, the discussion never became pessimistic; panellists consistently balanced critique with possibility, as seen in Marivate's metaphor of coming 'from the future' where AI is normalised , and in Nolan's bullish long-term outlook . Towards the close, the tone shifted to constructive problem-solving, particularly around data licensing and regulation, ending on a collaborative and forward-looking note.
Speakers Overview
MA
Mr. Alistair Nolan
165 wpm · 5 min
DK
Dr. Kamil Dziubek
127 wpm · 7 min
DM
Dr. Moses Thiga
140 wpm · 7 min
DV
Dr. Vanessa McBride
134 wpm · 4 min
PV
Prof. Vukosi Marivate
168 wpm · 15 min
DM
Dr. Marion Mercier
194 wpm · 6 min
A
Audience
136 wpm · 5 min
PD
Prof. David Castle
152 wpm · 6 min

Science in the Age of AI: Knowledge, Data, and Trust - An Expanded Summary

#

Opening and Scene-Setting

The session was convened by the International Science Council and CODATA to examine the multifaceted impact of artificial intelligence on scientific practice, knowledge generation, data quality, and research integrity. Dr Vanessa McBride opened proceedings by framing AI not merely as a product of science but as a force that is fundamentally reshaping how science itself is practised . She drew attention to the breadth of AI's impact on the scientific literature, noting developments ranging from the rise of AI agents to manage scientific workflows and identify predatory journals, to serious warnings about adopting AI-driven workflows without adequate guardrails . A central observation from McBride was that, despite the proliferation of national AI strategies, almost none of them contain any meaningful focus on the science sector itself, concentrating instead on downstream application domains such as health and agriculture . This governance gap, she argued, risks neglecting the very scientific foundations from which new technologies and applications ultimately emerge.

McBride also highlighted a report published earlier in the year by the International Science Council, titled Preparing National Research Ecosystems for AI, which synthesised case studies across 26 countries . Several of the report's authors were present on the panel, including contributors from South Africa and Kenya. The report's central finding - that national AI strategies are largely silent on the science sector - set the thematic backdrop for the discussion that followed . McBride outlined three interconnected dimensions the panel would address: AI and knowledge generation, scientific data as the foundation of AI, and the downstream issues of reliability, trust, and research integrity .

#

AI as Amplifier: Opportunities and Hyperbole

Alistair Nolan, joining the panel online from the OECD, offered the most broadly optimistic perspective of the session. He described AI as "a wonderful adjunct and amplifier of human intelligence" and stated that he was "very bullish about the long-term implications of AI and science," noting that AI is now being used across every step of the scientific process and in every domain . He acknowledged, however, that there is "a lot of hyperbole" surrounding AI and that it will create "a series of institutional stresses" that must be carefully managed .

Nolan's most substantive contribution concerned the changing nature of scientific bottlenecks. Drawing on OECD research into materials science, he argued that discovery itself may become less of a rate-limiting factor in translating science into technology; the primary challenge is increasingly one of scaling laboratory discoveries to industrial application . He also cited the remarks of Fields Medal-winning mathematician Terence Tao to support the view that AI will enable promising young scientists to engage with frontier problems earlier in their careers, by reducing the cognitive bandwidth spent on lower-level tasks such as memorising large bodies of literature and performing lengthy calculations . In Nolan's framing, AI does not replace scientific talent but accelerates its development and broadens who can participate in frontier research .

#

Data Quality and the Moving Target of Ground Truth

Dr Kamil Dziubek (the name as rendered here is the likely correct form of the phonetically transcribed "Camille Tubeck" in the transcript), co-chairing CODATA's task group on research data quality management across the data lifecycle, grounded the discussion in a concrete and instructive case study: AlphaFold, the family of AI programmes capable of predicting the three-dimensional structure of proteins, which earned its authors the Nobel Prize in Chemistry in 2024 . AlphaFold is trained on data from the Protein Data Bank, which contains over a quarter of a million experimentally determined structures derived from techniques such as X-ray diffraction, cryo-electron microscopy, and nuclear magnetic resonance . Dziubek's critical point was that these structures are themselves models - reconstructions from raw physical data - and therefore carry all the limitations that models inherently possess . Invoking the statistician George Box's aphorism that "all models are wrong, but some are useful," he argued that the scientific community must actively work to identify and eliminate biased, low-quality, or contextually inappropriate models .

Dziubek further emphasised that the "ground truth" used to validate AI systems is not static but a moving target: new data sets emerge daily, and what was considered ground truth yesterday may not be so today . This dynamic quality of scientific knowledge makes continuous validation essential and demands agreed measures of data quality across disciplines, including accuracy, uncertainty, traceability, and provenance . He summarised the principle concisely: "AI is for good if it's based on good data," and warned that if those deploying AI methods cannot provide transparency about their training data, "it's a big red flag" . CODATA is currently engaged in landscaping different measures of data quality across disciplines, with the aim of defining and agreeing on standards that can be practically applied .

#

Empirical Incompetence and the Global South's Compounded Disadvantage

Dr Moses Thiga, speaking from his experience driving technology adoption at Igaton University in Kenya, offered a markedly more cautionary perspective, particularly regarding the Global South. He acknowledged that AI presents genuine opportunities for leapfrogging - enabling simulation of laboratories, literature review, brainstorming, and data analysis in resource-constrained environments . However, he identified a deeply troubling countertrend: the emergence of a generation of scientists who are "empirically incompetent" . Researchers in his context, he warned, are producing publications without being able to conduct real experiments, read papers critically, collect data, or independently evaluate the outputs of their research . This problem is compounded by academic and research leadership that is not conversant with AI, leading to a situation where AI is either demonised and driven underground as "shadow AI," or adopted uncritically without the necessary skills to evaluate its outputs .

Thiga also highlighted structural disadvantages that make the Global South particularly vulnerable. Most AI models are predominantly trained on Western data, while data infrastructure in the Global South remains nascent and far from mature . Without the capacity to develop locally relevant models, and without native compute infrastructure, the leapfrogging opportunity risks becoming another mechanism through which the Global South falls further behind . This concern was reinforced by Prof Vukosi Marivate, who noted that AI amplifies existing inequalities, making pre-existing problems worse for the Global South .

#

AI Changing the Practice of Science Across Disciplines

Dr Marion Mercier, from the Geneva Science and Diplomacy Anticipator - an independent non-profit foundation working with scientists to anticipate advances over five, ten, and twenty-five year horizons - described AI's impact as "pervasive" and "catalytic" across every scientific discipline her organisation examines . She noted that her organisation had recently held its first anticipation committee devoted entirely to AI for science, chaired by Hiroaki Kitano, who is leading the Nobel Turing Challenge - an initiative to develop AI capable of producing Nobel-worthy scientific insights - and that insights from this committee were forthcoming and directly relevant to the session's discussion.

Mercier drew on an anticipation workshop on cognitive enhancement to illustrate a particularly striking implication: while neuroscientists still do not fully understand the brain , the combination of brain-computer interfaces and AI may reach a point where AI understands the brain even if humans do not . This raised what she described as a profound question about interpretability: if AI can achieve the goals of neuroscience - modulating brain function, treating diseases - but humans cannot understand how it is doing so, is that scientifically and ethically acceptable ? The question of whether outcomes without human understanding constitute valid scientific knowledge was left deliberately open, but it introduced an epistemological dimension that resonated throughout the remainder of the session.

Mercier also noted that drug discovery is one of the areas where AI-driven automation is most welcome and most advanced. Anticipation committees have suggested that within a twenty-five year timeframe, drug discovery timelines could be compressed from years to a matter of days, through AI mining of clinical data and synthesis of chemical compounds . She noted that this level of automation, while potentially transformative in drug discovery, may not be equally welcome across all aspects of science .

#

The Integrity Crisis in AI Research Publishing

Prof Vukosi Marivate, Director of the African Institute for Data Science and AI and Chair of Data Science at the University of Pretoria, a co-founder of Lilapa AI, and a member of the UN Independent Scientific Panel on AI, offered perhaps the most candid account of AI's disruptive effects on scientific publishing - from the inside of the AI research community itself. He is currently on sabbatical, focusing on questions of assessment and evaluation for AI models and how these can be improved. He described the situation in natural language processing conferences as "a mess," noting that submission cycles which previously received two to three thousand papers have now grown to thirteen thousand submissions in a single cycle . He explained that the Association for Computational Linguistics (ACL) had moved to a rolling review system - initially operating every six weeks, later extended to eight or ten weeks - in which papers enter a common pool and authors choose which conference to present at after acceptance, a structural change that contributed significantly to the explosion in submission volumes. Reviewers are now tasked with checking whether references in submitted papers are actually real, a development he described as symptomatic of a broader integrity crisis . He also noted that NeurIPS, one of the field's flagship conferences, now attracts between 20,000 and 30,000 participants, illustrating the sheer scale of the challenge. His assessment was blunt: "it's eating us as AI researchers" .

Marivate framed this not as a reason for despair but as a challenge that the scientific community must work through, requiring new ways of thinking about scientific integrity, the nature of discovery, and what constitutes a genuine scientific contribution . He used the metaphor of coming "from the future" - a future in which AI has been normalised and is "boring" - to encourage a longer-term perspective on the current turbulence . At the same time, he acknowledged the genuine excitement of the present moment, reflecting that AI has opened up lines of inquiry he had left unexplored during his PhD eleven years earlier . His overall message was that AI is an amplifier: it can amplify the good, and the task is to reinforce that while reducing the downside as much as possible .

#

Rethinking the Purpose of Scientific Institutions

In response to the panel's first formal question - on the long-term implications of AI for the production of scientific knowledge - Thiga argued that the scientific community must fundamentally redefine what knowledge, science, and research mean in the AI era . He contended that universities need to reconsider their core purpose: rather than focusing on physical campuses, institutions should invest in better compute infrastructure and position themselves as ethical gatekeepers, teaching good scientific method and the values that underpin responsible research . The knowledge, he observed, is already "all out there," which raises urgent questions about what universities are actually teaching and what research is for .

Nolan complemented this by arguing that AI will change not only how science is done but who does it, enabling a broader range of people to participate in large science projects through citizen science initiatives and allowing younger researchers to engage at the frontier sooner . Marivate added a personal reflection: despite working in AI, he has never owned as many notebooks as he does now, because sitting with one's own thoughts - rather than immediately turning to the AI prompt box - is essential to wielding AI as a tool rather than allowing it to do the thinking . This observation underscored a shared concern across the panel that the scientific method and the process of genuine inquiry must be actively preserved, not passively assumed to survive the AI transition.

#

Data Sovereignty, Open Science, and Equitable Licensing

The session's third major theme concerned the relationship between scientific data as a public good and the risks of unfair exploitation or over-restrictive data sovereignty. Marivate argued that standard open licences such as Creative Commons CC0 and CCBY have enabled a form of "open washing," in which well-resourced actors - particularly large technology companies - disproportionately benefit from openly licensed data without reinvesting in the communities that created it . He noted that Creative Commons is currently undergoing a review precisely because of these concerns, and drew a parallel with the copy-left movement in open-source software .

In response, Marivate described two novel licensing frameworks designed to address this imbalance. The ESETHU licence, developed by his startup Lilapa AI, distinguishes between users who identify as African and those who do not: African users may use the data for commercial or non-commercial purposes freely, while non-African users are restricted to non-commercial use and must negotiate benefit-sharing arrangements for commercial applications . Similarly, the NOODL (No Letter or Bordeaux Open Data Licence), developed at the University of Pretoria's law faculty, discriminates by whether the user comes from a developed or developing country, requiring benefit-sharing from well-resourced users . These frameworks represent an attempt to make data locally open to the communities it represents, while preventing exploitation by external actors with greater resources.

Marivate also revealed a counterintuitive reality: despite open science mandates from governments and funders, data is already being hidden, and communities are "figuring out ways to hide it even more" because they fear exploitation . He cited the common practice of papers stating "data available on request" while rarely delivering on that promise . His conclusion was not to abandon openness but to fix the licensing framework: "that is what science is about - we keep on improving" .

Mercier offered a complementary reframing, suggesting that data sovereignty need not be opposed to open science but could instead be "part of the solution to make data open... to the local networks where it should be open to" - a view Marivate confirmed . Dziubek added a technical dimension to this debate, warning that when data is excluded or restricted from AI training sets, the critical question is whether the remaining dataset is representative and unbiased . In high-stakes domains such as drug discovery, STEM research, and language modelling, non-representative training data poses a critical risk to the reliability of AI outputs . He also reiterated that in experimental science, the ultimate test of any AI-generated answer remains the physical experiment or clinical study, which provides the final validation and cannot be bypassed .

#

Corporate Power, Research Agendas, and the Public Interest

Nolan introduced a structural critique of the AI research ecosystem that connected the data sovereignty discussion to broader questions of power and public interest. Drawing on OECD research published in 2023, he noted that large technology companies are outspending public universities in AI research and development by large multiples, and that the rate of growth of their AI investment is significantly higher than in universities . These companies tend to collaborate predominantly with elite US research institutions, whose research profiles are considerably narrower than the broader university system . Crucially, the research they concentrate on involves types of AI that rely on high compute and large volumes of data - precisely the assets held by the companies themselves . Nolan argued that this creates a feedback loop that steers the research agenda in ways that may be "prejudicial to the public interest in the long term," and suggested that greater investment in smaller, less energy- and data-intensive models may require some form of structural intervention .

Daisy, an audience member from South Africa and CODATA, raised the question of the financial sustainability of the AI investment model, noting that AI companies are reportedly spending approximately 1.4 trillion USD against revenues of around 613 billion USD . Marivate responded by arguing that nations, particularly in the Global South, should not feel compelled to replicate this model . Smaller, task-specific models can compete effectively for many scientific purposes, and the rationale for massive spending is driven by the false promise of an "everything machine" that will solve all of humanity's problems . He warned that the AI bubble may burst - or may already be deflating - and that building genuine scientific and technical capacity is more important than dependency on large proprietary systems .

#

Regulation, Governance, and the Role of Governments

An audience member from the International Federation of Library Associations raised the risk that national AI strategies, by largely ignoring the science sector, may inadvertently regulate science poorly - either by applying general AI regulations that are poorly calibrated to scientific practice, or by failing to regulate at all . She called for the development of discipline-level codes of practice and protocols as a means of demonstrating that the scientific community is capable of self-regulation in a way that reflects its own values, and asked what approaches seem to work in accelerating this process inclusively .

Thiga responded by grounding the governance question in the foundational nature of scientific checks and balances. Science, he argued, has never been about infallible scientists; it has always depended on systems of verification and accountability . The challenge now is to rethink how those checks and balances apply to AI, covering ethics, data, compute, and the power imbalance between industry and academia . While he endorsed the principle of self-regulation, he also argued that "the responsibility falls at the floor of governments," as only governments can ultimately match the investment levels of large technology companies . He cited India's national AI investment as a model of how sovereignty considerations can drive the public investment needed to counterbalance corporate influence .

#

The Scientific Method as the Ultimate Value at Stake

The session's most philosophically resonant moment came in response to a question from a freelance journalist, who asked whether science - in an age of abundant AI-generated answers - lies more in the question or in the answer. Thiga's response reframed the entire debate: the value of science lies in neither the question nor the answer, but in the method of inquiry itself - "the inquiry, the observation, the hypothesis, the experiment, the data collection, the discovery" . It is this process, he argued, that AI is "about to steal from science" . The journalist's immediate response - "this is an excellent title for a book" - reflected the resonance of the formulation with the audience .

This philosophical point connected directly to Dziubek's earlier insistence that in experimental science, the final test is always the experiment , and to Mercier's question about whether AI-generated knowledge that humans cannot interpret constitutes genuine scientific understanding . Together, these contributions suggested that the panel's deepest shared concern was not merely about data quality, publishing integrity, or institutional governance, but about the preservation of science as a distinctively human and epistemically rigorous endeavour.

#

Conclusions and Unresolved Questions

The session closed with broad agreement that the challenges posed by AI to science are systemic, urgent, and require coordinated responses across multiple levels - from individual researchers and scientific disciplines, to institutions, governments, and international bodies. The International Science Council's report on national research ecosystems was offered as a resource and an invitation for further input . CODATA's ongoing work on data quality measures across disciplines, the Geneva Science and Diplomacy Anticipator's forthcoming publication from its first AI-for-science committee, and the emerging landscape of equitable data licensing frameworks were all identified as practical steps in the right direction.

Nevertheless, significant questions remained unresolved: how to reform national AI strategies to address the science sector; how to standardise data quality measures across disciplines; how to manage the integrity crisis in scientific publishing; how to prevent the erosion of empirical competence in the next generation of researchers; and how to ensure that the benefits of AI in science are equitably distributed globally. The panel's diversity - spanning an OECD policy analyst, a Kenyan university technologist, a data scientist based in Vienna, a Geneva-based science anticipator, and a South African AI researcher and entrepreneur - ensured that these questions were examined from multiple geographic, disciplinary, and institutional perspectives, making the session a genuinely multidimensional contribution to an increasingly urgent global conversation.

Dr. Vanessa McBride
Good. Thanks very much for joining us today in the session on science in the age of AI. We're going to talk a bit about knowledge, data, and trust, and specifically on the impact on how we practice science. So I'm just going to start with a kind of scene setting on behalf of the International Science Council and CODATA, and then I'll hand over to my colleague David Castle, who will moderate our esteemed panel, and we'll do some introductions of our panel members shortly. Thank you. So I think part of the conversation we've been having this week is really about how artificial intelligence is impacting many aspects of our society. But this is kind of a feedback loop because not only has science been fundamental to developing this kind of technology, but the impact of the technology is similarly feeding back into science systems and really changing the way that we do science. And if you just take a look at some of the articles that we've seen published in the scientific literature over the, this is just over the first six months of this year. You can see that the impact is incredibly broad on science. We're seeing the rise of AI agents to manage scientific workflows. We're seeing benefits, for example, the fact that AI tools are available to identify some predatory journals. But we're also seeing warning bills around adopting these kinds of workflows without the guardrails in place needed to do that. We're seeing the rise of AI agents to ensure scientific integrity and trust. and so that's really the setting in which we're discussing things today i wanted to highlight this this report that the international science council published earlier this year and it was on preparing national research ecosystems for ai it is a synthesis of case studies across 26 countries some of the case study authors are with you on the stage today of of course from south africa and moses from kenya and it was really looking at how national research ecosystems are preparing for ai and i think the thing we wanted to highlight is just that very first um learning from the report in that there are lots of national ai strategies under development and i think that's a very important part of the report development but almost none of them have any focus whatsoever on the science sector itself. We see a lot of health sector, we see a lot of agriculture, and we see how AI can be applied, but we don't see much consideration given to how it's changing the science that will result in new technologies and the applications itself. We welcome your input and feedback on this report. You can download it with the QR code or at the link. And then just to say about our panel today, we sort of thought that there were these three interconnected dimensions that we wanted to talk about today. We wanted to talk about AI and knowledge generation. We also wanted to talk about scientific data and how foundational it is for AI. And then we wanted to talk about AI and knowledge generation. We also wanted to talk about the downstream impact. The issues of reliability, trust, and research integrity. And so with that, I'm happy to hand over to David Castle, who's the chair for the Science Systems Futures project that we're working on, and to introduce our esteemed panel who are going to talk about
Prof. David Castle
Thank you, Vanessa, for the introduction. So we are a little bit behind, and we asked each one of our panelists, including Alistair, who's from the OECD and is joining us online, to say just a few remarks by way of their thoughts about the topic. We prepared a short briefing note about the panel and some of the questions that we wanted to raise. And so I think maybe out of courtesy to Alistair, we'd like to start with you, since you're away from us. And in person, I want to leave you to last. So did you have a few introductory remarks that you would like to share with us?
Mr. Alistair Nolan
Yeah, thank you very much. And it's an honor to be invited to this.
Prof. David Castle
ust a second. I'll just say we're just making sure we can hear you.
Mr. Alistair Nolan
Okay.
Prof. David Castle
Your pieces you see.
Mr. Alistair Nolan
Yes can you hear me now.
Prof. David Castle
okay good.
Mr. Alistair Nolan
Oh shall I go ahead okay very good so again thank you for inviting me I'm honored to be here um I'll just make a couple of comments very briefly one is about the structural importance for our economies and societies of advancing the productivity of science and as our economies age we will need science to feed into the development of more technologies to keep our economies productive enough to pay for the older age cohorts we will also need more discovery in order to address obviously the kind of global challenges that we face now say around the climate around disease and so forth um I am I saw you put up a question there at the beginning which will probably come to but overall I'm very bullish about the long -term implications of AI and science I think that they're very positive we see AI being used across every step in the scientific process and in every domain of science I think it's a wonderful adjunct to an amplifier of human intelligence I don't want to sound starry -eyed. There's a lot of hyperbole here as well. And I think that AI will create a series of institutional stresses that we'll
Prof. David Castle
Great. Thanks very much, Alasdair. Why don't we go with Camille?
Dr. Kamil Dziubek
So I'm Camille Tubeck. I work at the University of Vienna, and I'm also in CODATE. I'm co -chairing the task group on research data quality management across the data lifecycle. I choose one working example to start with, and just a question to everyone. Please raise your hand if you have heard about AlphaFold. So is most of people in this room. So AlphaFold is a family of programs that actually can predict based on the AI. The AI models, the three -dimensional structure of the protein. and it was a great invention. It earned the authors of this program the Nobel Prize in Chemistry in 2024. It's based on the deep learning algorithms and it uses the training sets from the experimentally determined protein structures. Those data sets are actually in the Protein Data Bank, which is a large database, over a quarter of a million of the structures, determined with experimental techniques like X -ray diffraction, cryo -EM, or nuclear magnetic resonance. But there is a catch about that, because all those data are models. We've never seen the molecules with our naked eye, we've never seen the molecules under the microscope, because they are too small, so we have to reconstruct them from the raw data, from different physical techniques. and they have all the problems that models have, actually. There was an aphorism attributed to the British statistician George Box that says that all models are wrong, but some are useful. So we are looking for the useful models, actually. We try to eliminate the models that are biased. We try to eliminate the models that are of the low quality. We try to eliminate the models that don't fit into the context of chemistry and physics that we know. And now there is another story, because it's not only about the quality of the models. You can verify and validate this quality, but there is also a question of these models being updated. So we will speak about the ground truth in the AI, this ground truth is a moving target, actually. every day you have the new data sets and every day you're checking against something else because the structures the ground truth of yesterday is not the ground truth of today now how we can actually know that our structures and the structures that are predicted with the AI methods are good we need to validate them we need to have the checks to validate them we need to define them and agree on them and actually this is the big question about the data quality so we need to define the measures of data quality across the disciplines this is what we are aiming for in Codata at the moment landscaping the different measures for data quality and asking people Because if you ask about the quality of the training data, you have to ask what are the measures of the quality, which you have to ask about the uncertainties, about the accuracy, about the traceability, about the provenance. And if the people who actually use the AI methods, they don't provide you with this data, it's a big red flag. So, in principle, just to wrap up, if you say AI for good, there is this big slogan, AI is for good if it's based on good data. And we need to have some measures to check if the data are good or not. So, I'll finish.
Prof. David Castle
Great. Thank you very much. Moses, why don't you go next?
Dr. Moses Thiga
Thank you. My name is Dr. Moses. Thank you. I'm Dr. Moses Iga from Igaton University in Kenya. And taking you on a different tangent, AI. is impacting the science system. And I come from an environment where we are still trying things out, we are still thinking things through. We still have a few people who say no way to AI, and we have shadow AI, you know. So in essence, AI in the global south, which is what I see it, in an academic context as an individual who is charged with driving the use of technology in my university. We see AI as a double -edged sword. And in the global south, it's given us the ability to leapfrog a lot, you know, do things that we possibly wouldn't be able to do, because AI gives us the ability to simulate laboratories and data and, you know, do lots of analysis, you know, get your literature done, your brainstorming. It's absolutely wonderful when it comes to that. But the real danger that we are seeing, at least internationally, my institution in my country is... a generation of scientists who are empirically incompetent. There's a lot of research going on, a lot of publication going on by some researchers who actually cannot do that experiment in a real lab, in a real world setting. So we are finding researchers who cannot actually read a paper. You know, they can't read, they can't do literature review, they cannot collect data, they cannot analyze data, they cannot critically evaluate or analyze the outputs of their said research. It's such a big problem and that's what is happening in the global South. Now, that is compounded by academic leadership and research leadership that is not conversant with AI. They don't use AI. You have this scenario where AI is demonized. so then it goes into the shadows and we're sort of in a very big confusion if I may say there is beginnings of policy and strategies coming up across the global south but I say not much of that comes from an experiential perspective that said again we are still at a disadvantage a lot of the models again predominantly fed by western data our data systems are not yet robust still not even maturing are very nascent stages so while AI is presenting opportunities to leapfrog it's also creating another scenario where again we might be left really far behind without the capacity to develop our models, we don't have the data infrastructure, the computes we still have to buy that we don't have native infrastructure in the global south so AI is doing good in science but presenting lots of more dangers.
Dr. Marion Mercier
Thank you. Hi. Thank you. So I'm from, I should correct, I'm from the Geneva Science and Diplomacy Anticipator. It's fine. The only reason I'm highlighting it is because anticipation is so key to what we do. So just a little bit of background to explain. We're an independent non -profit foundation that was founded by the Swiss government and the canton of Geneva with the mission of working with scientists to anticipate the science and technology advances that may be coming over the next 5, 10, 25 years and to then work with various stakeholder communities to have forward -looking dialogues and ensure that these advances can be developed in a way that will be beneficial to society. So it's a pretty big mission and what that looks like in practice is that we work with scientists from a very very broad range of fields we look at lots of different topics ranging from the social sciences all the way through to very fundamental disciplines like maths and we ask them to give us their vision of what they think might be possible they're not predictions they're possibilities over the next 5, 10 and 25 years and that puts us in a position to really see the pervasive catalytic impact that AI is having across every single one of the disciplines that we look at and it's very interesting to see the kind of the way that AI is being applied to very specific problems in different disciplines but I think what's more interesting is the broader question that's coming out which is how AI is changing the very practice of science and the pursuit of knowledge you and I think if I can use an example from one committee from one anticipation workshop that we did that really struck me and that illustrates this So we did a workshop on cognitive enhancement, so neuroscience -based, and all of our workshops are structured across some sub -themes, and they often range from the very fundamental to the more applied. And so in this case, we started off with discussing fundamentals of cognition. How well do we currently understand the brain, and do we think, you know, how is that going to advance over the next 5, 10, 25 years? And the conclusion there was that, you know, we still don't really understand the brain, and there's lots happening, but really we're not very close to kind of crafting it. And then we got to the end of the workshop where we were discussing brain -computer interfaces and how that technology is developing, how AI is developing, and how by bringing those two together, you know, these implants and even actually the non -invasive technology, you bring that together with AI, and you get to the point where actually AI is going to understand the brain. Even if we don't. And I thought that was so striking because... it raises so many implications, not least the question of interpretability. You know, if AI understands how the brain works and it can do all the things that we're trying to achieve with neuroscience, like, you know, modulation, treating diseases. But if we don't understand how it's doing that, you know, are we OK with that? Are we OK with the outcome without it advancing human understanding? So, you know, there's a lot of other implications, but I just thought that was a really good illustration of some of the things that we're faced with in terms of bringing AI into science and into discovery. And just to close to say that because we're seeing this kind of wide impact this year for the first time, we had our first committee, anticipation committee devoted entirely to AI for science, which was chaired by Hiroaki Kitano, who recently or not recently, I should remember the date, but he is leading on the Nobel Turing Challenge, which is where he's launched the challenge. To develop AI that will be able to develop Nobel worthy insights. so that was a very interesting discussion which I think lots of insights from that which will be published soon and lots of insights from that will be relevant to today's discussion. Thank you.
Prof. Vukosi Marivate
Thank you, I'm Vukosi I am one of the members on the UN independent scientific panel on AI so yeah, it's also being scientists looking at the science to try and give perspectives. I'm at the University of Pretoria as the director of the African Institute for Data Science and AI and I'm also the chair of data science there and then on the other side I have a startup company I'm a co -founder of Lilapa AI where we work on building language systems using AI and I'm on sabbatical this year trying to think about assessment or evaluation for AI models and how we can improve that, because we need it for being able to scale, especially in areas where you don't have enough representation, whether it's the way that evaluation is built, models are built, or the data. How does it represent the majority of the world? Yeah, but with that, maybe I'll start off and say I come to you from the future. AI is boring now, and it is not like, you know, it's something that we've accepted, it's pervasive, and we've now learned to live with it. Unfortunately, today, we have to go through it. We have to go through all of the emotions, all of the hurt, all of the opportunity of what AI is doing to much of the way that we think about research and science. So, like, just briefly. building on everything that is already being said, both online and by the fellow panelists who are in the room. There's much that is going on. I come from natural language processing as a researcher and AI, so at the moment I am area chair and senior area chair for two different conferences, one being in natural language processing and the other one being NeurIPS. It's a mess. I'm telling you from the AI researchers, it is a mess. We've gone maybe from, hey, a few years ago you were getting during submission cycles maybe 2 ,000 to 3 ,000 papers in submission cycle. Oh, just to give you context, in ACL, which is Associated Computational Linguistics, we have something called rolling review. So our major areas is not journals, but it's conferences, so there's a number of conferences a year. Let's say there's eight. So initially you used to just try to get your paper into one of those conferences. And you would get a deadline. And then there was a decision to say, let's just have rolling review. I think at the beginning was every six weeks you could submit a paper. and it went into a common pool. There would be reviews, and then after your paper gets accepted, you could choose which conference you would then present that your paper would then be published in those proceedings. Already before AI, it was overwhelming, and then I think they moved it to like an eight -week cycle or ten -week cycle. That's what we're living on now. This cycle that just finished, which we are now doing reviews for, had 13 ,000 papers submitted. And now there's a question that comes in from Integrity. How many of those are like, you know, a, what is it? It's not a good effort, but it was actually true. Somebody doing some experimentation or trying to show, like doing the science and then going to whether AI assisted or not. And what other ones are just, it's just completely AI generated, has no real. Kind of goal, it was just, can you please get me like, you know, write a paper similar to AlphaFold and get it out there. Right? So that's what's happening to the people who are building the tools. That's what's happened to our practice. Right? And then in NeurIPS, which is the other one, NeurIPS typically now has like 20 ,000 to 30 ,000 participants who come to that conference. It's crazy. So the amount of papers they are is now we're going through and checking are the references actually real or not. And this is, again, like I know there's this tension that comes from other parts of science who are saying like, but we're being affected. I'm like, oh, it's eating us as AI researchers. So we're going to have to go through this. We're going to have to figure out like new ways of thinking about what is scientific integrity, show adjustments to where discovery comes in, and then what it actually is. Then it's going to change in our practice. It's just that it's going to affect us all differently. Moses was saying, if you're coming from the Jehovah majority, it's going to be already, there was like a lot of like vibrations in the system that we were trying to get out. Now this is just making it worse. Right? So you can amplify good. And you can, that's the things with technology, right? It's an amplifier. It can amplify the good. And we can try to reinforce that. And we have to reduce the downside as much as possible. And that's the big thing that we're going through. And that's what I just wanted as like an opening statement of saying like, yes, from the future, it's boring. Right now, it hurts. And then at the same time, it's amazing. It's amazing the kind of experimentation you can try out that you couldn't. There's even thoughts I've been having, I actually want to go back to my PhD 11 years ago. And there's all these things that I'd left on the table that I think I can explore now. And that's where we are.
Prof. David Castle
Thank you very much for all your interesting opening remarks we also had some questions prepared three of them and so we uh we're about half past now so what we'd like to do pardon me yeah do all the questions up at once oh we had four sorry um so uh what we do is maybe spend a few minutes on each one of the questions and then take questions from uh the audience and uh perhaps there's some things that have come in in the chat on online that Felix will be able to uh tell us about so um anybody who wants on the panel to including of course you Alistair um as AI continues to permeate various aspects of the research process what are the long -term implications for the production of scientific knowledge so who would like to comment on this particular question.
Dr. Moses Thiga
Um I could in a sentence um We probably need to redefine knowledge And what is science What is research, what is learning That's probably changing a lot How do we even teach research in the first place What kind of skills do we need today And if anybody is not Rethinking this It's probably not futuristic at the moment AI is creating illusions of competence You know Of capabilities, lots of illusions Institutions must rethink What are we here for What's a university for What do you do in the lecture hall You know, what's the purpose of that Because the knowledge is all out there Everything is out there So what are you teaching What is research, the literature can get done Hallucinations notwithstanding So what must institutions be In this day and age I think institutions probably need to focus On having better compute Than better campuses Because that's the future of science They need to probably think of Being ethical gatekeepers teach ethics, teach good science, you know, like there is a process to it. There's a human at the end of your discovery. You know, that's probably.
Prof. David Castle
Thank you very much. I see Alistair, you have your hand up as well.
Mr. Alistair Nolan
Yeah. If I just make a comment on the first question, I think that a bottom line for me is perhaps that discovery itself, the generation of new knowledge may become less of a rate limiting factor in the way that science has an impact on our technology, on our economies and societies. So, for example, in the era of material science, many of our fundamental challenges around the climate, around battery technology, et cetera, buildings that can cool themselves. So material science is a fundamental, will be a fundamental contributor to that. And there's a lot of sort of a lot of hyperbole, actually, and a lot of enthusiasm about the role that can play material science. But it turns out we've just been doing work on this subject that it's not actually discovery, which is the primary bottleneck. It's translating discovery of what you find in laboratory to an industrial scale when you're producing this new material in terms of thousands of millions of tons. So we need the new technologies. And I think discovery will be less of the challenge in getting in getting those discoveries. Another very briefly, another long term implication is I think about who will be producing new knowledge. It won't be that anyone could do science. Through the instrument of AI, almost certainly not. But what is already happening is that AI can help bring more people into large science projects, say through citizen science initiatives. But more importantly than this. It will help students to engage in higher level thinking about novel problems at a younger age. So you may have seen the recent remarks. You can find them on YouTube and just Google them by Terence Tao, who's a Fields Medal winner in mathematics and a professor of maths at the University of California, Los Angeles. And he recently pointed out that his main argument is the promising young mathematicians will now be able to reach the point where they can seriously experiment at the frontier sooner because less of their cognitive bandwidth is going to be spent on lower level tasks like memorizing large bodies of literature, developing an aptitude for long calculations and long proofs and so on. So I think those will be two points where I think we'll see changes. The discovery will become less fundamental and as on the path to new technologies. And we'll see a difference. In who's generating new knowledge as well.
Prof. David Castle
Okay thanks thanks Alice um so let's go on to the second question here and see what other of our panel wants to say what what other um uh areas do you think that uh yeah it might contribute to generating scientific knowledge we provocatively use the word independently here but maybe in say some sort of supervised or semi -supervised environment any any thoughts about where this will happen?
Dr. Marion Mercier
um so i'm absolutely not an ai expert but just again um just insights that we get from these anticipation committees that we have and i think you know drug discovery is definitely one of the areas where we're seeing you know huge impact and scope for automation along every part of the pipeline um and one of the anticipations that we had in one of our in a 25 -year time frame we might start to see drug discovery time cut from years, which is where it currently stands, to a matter of days due to AI mining of the clinical data to synthesis of chemical compounds. And so I think that's definitely an area where there is scope for automation and where it would be very welcome as well. I think it might not be so welcome in other aspects of science, but I think in drug discovery, there's a lot of room for it. Great. Thanks. Any other quick... So I'm absolutely not an AI expert, but just again, just insights that we get from these anticipation committees that we have. And I think, you know, drug discovery is definitely one of the areas where we're seeing, you know, huge impact and scope for automation along every part of the pipeline. And one of the anticipations that we had in one of our committees was that within a 25 year time frame, we might start to see, you know, drug discovery time cut from years, which is where it currently stands to a matter of days due to AI mining of the clinical data to synthesis of chemical compounds. And so I think that's definitely an area where there is, you know, scope. For automation and where it would be very welcome as well. I think it might not be so welcome in other aspects of science, but I think in drug discovery, there's a lot of room for it.
Prof. David Castle
reat, thanks. Any other quick thoughts from the panel on this?
Prof. Vukosi Marivate
Yeah, so I'm working in natural language processing the things that people look at even in building language models it doesn't need to be an LLM, it's still very much based around thinking about English or very high resource languages there's so much opportunity here and you might say hey there's a bit of engineering here of now discovering more and more about our human knowledge by being able to model more and more of our lower resource languages and that's a place that yes some of the recipes are repeatable and being able to explore that so in terms of being independent it's there and then being able to then hit that kind of the edge and then go over and see and say like oh for these languages that have these different types of scripts here's actually new rules that we kind of didn't anticipate and new knowledge so i'm excited for that uh but yes again uh being a tool how you wield it is going to be interesting and that's the thing we're going to have to as mozo say uh teach the new scientists what to what to do better and also online um it's it's going to be really important and and maybe one thought to leave for everybody is it's i've never i'm not a person who likes writing notes down on paper uh but i've never had as many notebooks as i have right now uh because in now on thinking and being able to sit with your thoughts and actually then yes it's easy to to kind of be in in the prompt box but then you don't think but then while you're sitting and you can sit with a pen like you know with a pen and and and paper now you're learning to wield the AI more as a tool and not do the thinking for you in a way. So that's another.
Prof. David Castle
Great, thanks. So as Vanessa said, we've structured this forum as an opportunity to talk about AI and science systems, which we've done a little bit. Now we'd like to go to the other end of the triangle and talk about the data issue. And I wonder, Camilla, if you would like to comment on number three, because this question is really actually about what's at stake in terms of thinking about science as a public good, a generator of knowledge that all people can benefit from and potentially use. When in fact we do know that at the core of AI is data, and as you said at the beginning, you were focused on data quality, but then this is a question about access and use, and how do we protect against unfair or unwarranted access versus over -restrictive data sovereignty and protection of data on the other hand. So what would you say about these issues briefly, Agnes?
Dr. Kamil Dziubek
Thank you. So in principle it also goes back to the previous question because you mentioned the scientific knowledge and there is a difference between data and knowledge. There is data, information, knowledge and wisdom if you remember the pyramid. In principle I know that CTI will massively create data. The question is about the quality of the data and the usefulness of the data. So in principle the risk in data harvesting I can say that if we massively produce the data, there is a bottleneck at some time. And we will not be able to actually comprehend the data. Then of course we'll have to say if all the data should be open and fair because fairness is not exactly the openness we need to safeguard the fairness, we need to safeguard quality of the data but of course be of course look also at risks with data sovereignty and some issues with data with the personal data for example so just need to take the balance but I guess the data deluge is a big problem at the moment and it will be in the near future
Prof. David Castle
Okay, thank you, you have a comment? Go ahead.
Prof. Vukosi Marivate
yeah, so so So here, in terms of thinking about the universally accessible, yes. But thinking about balance and equity, it is likely that we have to think about things like equitable licenses. So this year, if people don't know, Creative Commons is going through a review because they've seen that there's a challenge of open washing, that the people who do have the compute, the engineers, the scientists, then really benefit off, they accrue most of these benefits of data as available, and it's been made open, like let's say just on a CC0 or a CCBY, while they don't necessarily invest back into the open source software communities that have gone through this before, and that's why you have copy left. And now we're going to do that in data as well, right? If you want to see this, put our data on CCBY SA, and sometimes you get hate mail. right because people don't want to actually continue really like you know building on the data that you've released and also releasing those um those things at the same time you need we need to understand that there's parts of the world where then there is a benefit that can accrue um uh to to smaller players if they did get access so for example um uh for people there's at the university of pretoria in our health uh scholars i'm sorry law um uh faculty there's a new license called the no letter or bordeaux open data license so noodl which does kind of discriminate by where you are accessing uh the data from where they are coming from a developed country or a developing country and then it then says hey you can either use the data or whatever freely commercial non -commercial reasons but if you come from a developed country uh you must then get back to the communities that have created that data and then talk about benefit sharing At Lilapa AI Our startup company We have the ESETU license E -S -E -T -H -U And that one says do you identify as African or non -African If you are African You can use commercial or non -commercial Doesn't matter If you don't identify as African You can only use that data for non -commercial reasons And for commercial reasons Again get in touch with The creators or the communities That that data represents And then negotiate benefit sharing These are things that are going to Because we are not on a level playing field So saying And this is some of the challenges that The communities have seen Of just saying we have open science We have open data What then happens is actually you have people hiding Their data And it doesn't show up And the reason I know about this is because I work on language And language data is there It's just that people are trying to make sure it stays in their garage And tapes and under beds Because they are just Just
Prof. David Castle
Great. Thank you very much. I see that, Alistair, you have a comment about this. And then if we could maybe hear from you, Alistair, and then take a couple of questions from the audience or online. So, Alistair?
Mr. Alistair Nolan
Just an observation that links the issue of equity and sort of power in this AI ecosystem and data. And that's the nature of the research which is being done by the institutions that have the deepest pockets. That's the large tech companies that are outspending by multiples, the kind of AI R &D which is done in public universities, and the rate of growth of their investments in AI R &D much higher than in the universities, and they're drawing off talent, as we know, et cetera, et cetera. So, we put out a publication in 2023 that amongst others things showed that the large tech companies... They tend to collaborate with the elite in the United States, the elite research institutions. And you look at the breadth of the research profile of these elite research bodies, and it's much narrower than is the case for the rest of the university system. And what are they concentrating on in their research? They're concentrating on types of AI that rely on high compute and large volumes of data. Now, these obviously are the assets which are held by the companies that they're working with. So there's a sort of steering of the research agenda in ways which I think would be in the long term prejudicial to the public interest. So it would be good, for example, if there was more investment in less energy, less data, compute hungry at smaller models, for example. And I think that may require some kind of
Prof. David Castle
Great. Thank you very much for that. That's very insightful and interesting. Okay. We started a bit late, but we'd like to take any questions that we could. from the audience. So anybody want to ask? Okay. I think you were first. You, your man. Can you use the microphone? Thank you.
Audience
I'm Daisy. I'm from South Africa and part of Core Data. And I just want to check with the panelists their thoughts around the involvement of AI and what's going to happen, especially what my counterpart from South Africa, Fugosi, highlighted. We are aware that AI companies are spending 1 .4 trillion. That's the loss. 1 .4 trillion. And the revenue is just 613 billion US dollars. So is this model sustainable? How are we, especially in the global south, we're leapfrogging and this is happening? Thank you.
Prof. David Castle
Yes, yes, yes. Go ahead.
Prof. Vukosi Marivate
Yeah, yeah, quick one. I think I will agree with my fellow panelists online. You don't need to follow this model at all. It's just, it's like, you know, sometimes we get nations coming and saying, we should build our own insert AI model that requires a lot of data, a lot of compute, because we need to show that we can do it. Or we need it because that's the thing that gives you something. Like, no, you don't. You can build small language models, if you're in language, that can compete depending on the task and what you actually need. You do not need everything machines. You can say, I'm working, like, if you're saying, I'm working on protein folding in this specific way, and I can have a model that is literally a couple of megabytes, and it actually does well. Why? Because in science, is that you go in and you say, a model is an abstraction of the real world, and I was able to figure it out, and it actually works. So there is, yes, the data -driven way, and it is one way, and we keep on seeing the models as they get bigger. But at the same time, the reason you're looking at this spending that is ridiculous is because of the promise of this is an everything machine. They will solve everything, and as such, we just have to keep on throwing cash in, and then we're going to get to a point of human nirvana, where we all apparently just get money for living. This is the thing that then becomes very enticing to just say, hey, as countries, don't think about your sovereignty, because we just need to keep on shoveling money and data into these systems, and we will solve everything in humanity, and that's not true. right and that's the part that we have to that is not true and it doesn't mean you shouldn't be doing basic high impact research work do it because we still have to learn because the bubble is going to burst or it's going to be like you know somewhere that it actually has burst and we need to then survive past that as humanity so let's yeah that is just one of the we yeah there's more to do here.
Prof. David Castle
There was one other question I think it was yours oh sorry I'm really sorry about the position here anybody else so you and then you and we'll see if we have time for an online. Okay, go ahead.
Audience
Thank you very much so see my from the international federation of library associations I think Vanessa you mentioned at the moment that a lot of national AI strategies don't actually mention science and there's still a risk that either they regulate science or they regulate it and they regulate it and they regulate it by accident because they haven't bothered thinking about it and they're just bring up science in the rest of things or indeed they get worried and say science, not sure about that and they actually regulate without thinking I think this all points to the value of developing some of these codes of practice these protocols on a disciplinary basis in order to demonstrate that the community is still capable of regulating itself in a way that actually supports the values and I suppose the question is what seems to work, what opportunities are there to accelerate this process of developing the ability to update protocols, to update ethics in a way that is truly inclusive, that doesn't hold everyone to one particular standard and so have that answer, have that way of saying no, science can actually regulate itself we can find solutions within the community rather than needing to depend on government regulation coming in which may not.
Prof. David Castle
That's a really interesting comment Moses, do you have something to say in response to that? I feel like that's up your alley
Dr. Moses Thiga
You know my thoughts on what you said is it borders around regulation. And science, the science system has never been about infallible scientists who cannot make mistakes. It's about checks and balances. And the way forward really is, how do we continue with this regulation? There is now a new player on board. So how do we now regulate how we do this? How do we do that? I think that's really the point around regulation. We need to rethink how we regulate. And we regulate on ethics, on data, on compute, and back to the issue of the power imbalance, the industry, academia, high end. I think the responsibility falls at the floor of governments. Only governments really, at some point, can match what some industry players are doing. and please read the case of India I'm always fascinated at the kind of investment they are doing to match what industry is doing it's a matter of national pride it's a matter of sovereignty it's probably how things will work going forward
Prof. David Castle
Thank you one more question from here and I think we have one online and then that will be our session thank you.
Audience
I'm just an unbearable local journalist freelance so two questions I heard that in Kenya the quality or at least motivation of academics and researchers is poor but the motivation and quality of students is outstanding since I watched a documentary saying that half the PhDs in Oxford and even research and lecturing is indeed outstanding written in in coffee shops in Nairobi by penniless students. So should I conclude that the more you rise, you step up in an academic career, the less quality and motivation you have for science? So that's my first question. The second one, I cannot find a good example. I will take a lousy one. One, is science, especially when total easy information is available, more in the answer or in the question? For instance, I have a question as a non -scientist, is there a common origin to the word mother and water because in many languages ma, Mayim, may can be encountered? Or another question, why has Latin left so little, trace in Arabic -speaking countries? So. probably with AI I can find a lot of answers I tried, I haven't found but I think with good tools I could find a lot of things making possible advances in research what about the question, in that case is the real science in the question and why do people suddenly raise a question which more learned people didn't raise or is it in the answer we also know that deciphering the Maya scripture came from the experts being on holiday and the child being brought there just for holiday started deciphering and he didn't have the prejudice of learned people so he went through, so I don't know if you.
Dr. Moses Thiga
get your very well framed question. I'll do it quickly. The paper paper mills did very well Before AI So the business is no longer viable A lot of the students That did the papers That are rumoured to be published At Oxford etc That was the pre -AI era That business is no longer Viable, it's really not working Are the students more motivated? Yes for money Are the faculty and researchers less motivated? In a sense because Not much research funding is available And support from the government But that was pre -AI Is the value in the question or the answer When you think about the scientific Method The value is in the method And this is where we are losing it as scientists This is where we are losing The whole thing about The inquiry, the observation The hypothesis, the experiment The data collection The discovery And this is what AI Is about to steal from science The beauty Of the AI Of the inquiry so the value is neither in the question nor in the answer but the beauty is
Audience
this is an excellent title for a book write the book, I'll be your first reader.
Prof. David Castle
There you go Moses alright let's take one question online I think it's from Nipu can you use a microphone yes because it's online as well.
Audience
So I'm wondering if I'm wondering if the sovereignty of data in AI would be a threat to open up open science ok.
Prof. David Castle
Good question we'll take that and we'll also take the one online as you suggest especially in the global south great question ok and a question online Nipu unmute not. Not there. Okay. Then we're just going to close with this one question and let's move on. If Parley can... Parley, do you want to unmute?
Audience
My name is from women in technology in Nigeria so because I work with a lot of scientists who we are pushing for citizens data so I'm just wondering if that push for data sovereignty will threaten the open science we are all fighting about.
Prof. David Castle
Alright good nothing from online we've asked twice panel in response to this last question sure quick one.
Prof. Vukosi Marivate
Already at the moment even as I said like with the pushes for open science and open source open data I'm not making aspersions on code data or other other things I work in language. Finding language data has been very interesting, understanding traditions that different fields have about how to treat. So even with open science requirements by governments and all those things, data is hiding. Data is hiding. And people are figuring out ways to hide it even more. Right? We, in machine learning, used to go like, oh, show me, like, you know, you say here's the experiment that you did. Show me the data. Oh, and then they put it in the paper. Paper, data available on request. Lots of papers have come out showing that that never happens. Very rarely when you ask do you get the data. Right? And there's all these things. But, yes, machine learning and AI has also progressed very rapidly because it's become easier to just literally you see the paper and they have a link and you click it and it brings up a notebook and you run it and it pulls the data. There's been huge things. But we do understand the challenge of people then saying we do not want to be exploited. This has happened in. in machine learning where people work on health data, agricultural data. I sat on the Lacuna Fund Steering Committee where we were funding people to create data for lots, and we were requiring either a CCBY or more liberal license. And then what people said that came back as feedback to us as an organization was, oh, it's great that you want us to have a kind of open data at the end of this, but then the people who end up getting the recognition are the big tech companies who take that data and then release an update. And that's the thing that brought these huge conversations about, oh, what should we be doing about the licensing in such a way that people, and that's why, whether it's a CETU or it's Noodle or the new reviews that are coming in, because Creative Commons used to have a developing country license. So it's not that we just abandon it. We fix it. That is what science is about, right? We keep on. We're improving.
Prof. David Castle
I wonder, Marion, do you happen to have a comment about this as well from your perspective? And then we'll wrap up the session.
Dr. Marion Mercier
Not so much. More, I had a follow -up question, which was rather than, you know, is data sovereignty a threat to open science? Is it actually, you know, part of the solution to make data open, you know, to stop people hiding their data, like what you're describing, and to make data locally open, you know, to the networks where it should be open to, rather than, so, yeah, rather than being opposed. That was how I kind of understood your answer.
Prof. Vukosi Marivate
That is the part that it's trying to do.
Dr. Marion Mercier
eah.
Prof. Vukosi Marivate
The local network you're trying to impact.
Dr. Marion Mercier
Yeah. Yeah, sorry.
Prof. David Castle
Okay, great. Thanks very much. Did you have one short? Okay. Very short comment on this.
Dr. Kamil Dziubek
So the main threat, if you don't have fully open data, is you have to ask yourself a question. Is the data set representative? Is it not biased if it's not fully open? And if you're sure that the data that you're excluding from the data set doesn't cause that it's biased, it's all right. But if it causes some kind of bias, it can be really crucial. And I'm not talking about the data in languages, the data in STEM, for example, the data for the drug discovery. If you don't have the input from different groups. And one very short comment also to the gentleman who was asking the question about the question or the answer. In experimental science, the final test is always the experiment. So if you have the answer from the AI, what type of drug you're designing or what type of the material you're designing, there is a clinical study or there is an experiment. And then you know.
Prof. David Castle
Great, thank you Thank you panel, it was a wonderful session Thank you Alistair for joining us from Paris and audience please join me in thanking our panelists for this interesting conversation It was a good session

Disclaimer: This is not an official session record. DiploAI generates these resources from audiovisual recordings, and they are presented as-is, including potential errors. Due to logistical challenges, such as discrepancies in audio/video or transcripts, names may be misspelled. We strive for accuracy to the best of our ability.