WSIS Forum 2026
AI-generated report

Open Source and AI for Sustainable Development: Digital Infrastructure, Open Data, and Multistakeholder Cooperation

5 speakers
Summary

This panel discussion, moderated by Yue Gao at WSIS 2026, focused on the role of open source software and artificial intelligence in advancing sustainable development goals, bringing together perspectives from academia, open source governance, and industry practice .

Wei Wang opened by framing open source AI as a critical tool for transparency, arguing that just as Richard Stallman once asked "who controls your computer?", the rise of large language models raises the equally pressing question of who controls users' knowledge and minds . He also emphasised that open source projects provide students with real-world engineering tasks, arguing that the most important skills in the AI age are judgement, taste, and quality control rather than mere code generation . Christofer Dutz highlighted a practical challenge facing open source foundations: the surge of AI-generated pull requests is placing enormous strain on project maintainers, with some Apache projects receiving up to 170 pull requests per day . In response, Apache launched a project called Apache Magpie, an AI-based tool designed to review, sort, and consolidate pull requests to reduce this burden .

A significant portion of the discussion addressed the tension between open AI models and closed training data. Jianmin Wang noted that industrial enterprises regard their data as private assets, and suggested that open source tools could allow companies to train models locally, while artificially generated datasets might be shared publicly as a workaround . Dutz added that even willing companies may be legally prevented from releasing training data due to copyright constraints , and Wei Wang called for a standardised transparency spectrum to clarify what is genuinely open in any given AI model .

On the topic of multimodal AI, the panellists distinguished between language models, vision models, world models, and time series models, noting that each presents distinct opportunities and challenges within the open source ecosystem . Dutz expressed concern about recursive data poisoning, whereby AI-generated code enters open source repositories and is subsequently used to retrain AI models, potentially degrading quality over time . Jianmin Wang concluded by suggesting that software engineering has evolved through three paradigms - high-level languages, foundation models, and natural language prompting - and that future software will likely combine all three .

Overall, the panel converged on the view that open source remains the most viable path towards inclusive and trustworthy AI, provided that the community can navigate challenges around data rights, governance standards, and the sustainability of open source maintainership in the AI era .

Keypoints
  • Overall Purpose

  • The panel, moderated by Yue Gao at WSIS 2026, aimed to explore the intersection of open source software and artificial intelligence in the context of sustainable development. Bringing together academics and open source governance practitioners, the discussion sought to examine how open source and AI can serve global development goals, address tensions between open models and closed data, and consider the challenges and opportunities posed by multi-modal large language models. ---
  • Major Discussion Points

  • The philosophical and governance importance of open source in the AI era: Panellists framed open source not merely as a technical practice but as a governance philosophy critical to transparency and user empowerment. Wei Wang drew on Richard Stallman's question "who controls your computer?" to pose a new challenge in the AI age: "who controls your mind and your knowledge?" Open source AI was seen as a tool to make AI mechanisms more transparent to end users, covering models, datasets, source code, and publication reports. - The tension between open models and closed data, including copyright complications: A central theme was the conflict between openly available AI models and proprietary or legally restricted training data. Christofer Dutz distinguished between "open source," "open weights," and "open data" models, noting that even willing companies may be unable to release training data due to copyright constraints - for instance, training on entire library collections. Jianmin Wang suggested that enterprises view their data as private assets, and proposed solutions such as local training on private infrastructure and the use of AI-generated synthetic datasets for sharing. Wei Wang called for a standardised transparency spectrum to help users understand what is truly "open" in any given model. - The impact of AI-generated contributions on open source communities and the risk of recursive data poisoning: Dutz described a "tsunami" of AI-generated pull requests flooding open source projects - some Apache projects receiving up to 170 per day - placing enormous strain on maintainers. In response, Apache launched the Magpie project, an AI-based tool to help review and triage contributions. Looking further ahead, Dutz warned of a recursive poisoning risk: AI-generated code entering open source repositories could feed back into future AI training data, potentially degrading quality over time. - Multi-modal AI models and the distinct challenges and opportunities they present within open source: The panel discussed how AI extends well beyond language models to include vision, time series, and world models. Vision models raise acute copyright and deep-fake concerns. World models - described by Dutz as the "holy grail" - would require training data and compute capacity that currently exceeds available infrastructure. Time series models, such as those supported by IoTDB, were highlighted as a practical near-term tool for industrial optimisation. Jianmin Wang noted a lack of standardised data formats for physical and embodied AI datasets, presenting an opportunity for the open source community to contribute. - The future of software engineering and open source in an AI-driven world: Panellists reflected on how AI is transforming software development itself - from high-level programming languages towards natural language and specification-driven development. Jianmin Wang described a three-stage evolution of software engineering: high-level languages, foundation models as software components, and natural language prompts. Wei Wang argued that open source remains the best vehicle for making AI a genuine public good - akin to air and water - particularly in education, where access to AI tools should be free or low-cost. ---
  • Overall Tone

  • The overall tone of the discussion was collegial, intellectually engaged, and cautiously optimistic. The moderator and panellists maintained a respectful, collaborative atmosphere throughout, consistent with an academic and policy-oriented conference setting.
  • Early in the discussion, the tone was largely exploratory and philosophical, as panellists introduced their perspectives on open source and AI. It became more technically specific and candid in the middle sections, particularly when Dutz raised concerns about maintainer burnout, the pull request tsunami, and recursive data poisoning - moments that introduced a note of genuine concern.
  • However, the tone never became pessimistic. Panellists consistently reframed challenges as opportunities, with phrases such as "we just need to survive this tsunami" and "the whole world is broken, maybe we can build a new one." By the closing exchanges, the tone had shifted towards forward-looking optimism, with discussion of specification-driven development, AI as digital public infrastructure, and the potential for open source to emerge stronger from current pressures.
Speakers Overview
WW
Wei Wang
104 wpm · 9 min
YG
Yue Gao
108 wpm · 15 min
CD
Christofer Dutz
143 wpm · 13 min
JW
Jianmin Wang
107 wpm · 7 min
A
Audience
99 wpm · 1 min

Expanded Summary: Open Source and AI for Sustainable Development - WSIS 2026 Panel Discussion

#

Introduction and Panel Overview

The panel, moderated by Yue Gao at WSIS 2026, was convened on the final day of the conference to explore the intersection of open source software and artificial intelligence in the context of sustainable development . The session's full theme - "Open Source and AI for Sustainable Development: Digital Infrastructure, Open Data, and Multi-Stakeholder Cooperation" - reflected the breadth of issues the panellists were invited to address . In her opening remarks, Gao framed the discussion by observing that artificial intelligence is reshaping society and the economy "at an unprecedented pace," and that open source software and open data form the foundational layer of this transformation . She cited examples ranging from Linux and Apache to Hugging Face and PyTorch to illustrate that open source is "not merely a technical collaboration model" but rather "a governance philosophy that promotes global knowledge sharing and narrowing the digital divide" . Gao further noted that achieving the Sustainable Development Goals depends on open, trustworthy, and collaborative digital infrastructure, with applications in climate monitoring, public health early warning, and smart city development .

Three panellists were introduced: Professor Jianmin Wang, Dean of the School of Software at Tsinghua University, whose work spans software engineering, data management, and open source education in China ; Christofer Dutz, a board member of the Apache Software Foundation (ASF), one of the world's most influential open source foundations, stewarding more than 350 open source projects ; and Professor Wei Wang from East China Normal University, whose research focuses on open source software supply chains, open source governance, and digital education . The panel was structured around a series of questions posed by the moderator, supplemented by audience contributions, covering the philosophical importance of open source in the AI era, the tension between open models and closed data, the challenges of multimodal AI, and the future of software engineering .

#

The Philosophical and Governance Importance of Open Source in the AI Era

Wei Wang opened the substantive discussion by invoking the foundational question posed by Richard Stallman - rendered in the transcript as "Richard Stoneman," an apparent transcription error - namely, "who controls your computer?", and extending it to the present moment . In the age of large language models, Wang argued, the more pressing question has become: "who controls your mind and your knowledge?" . This reframing elevated the open source debate from a technical matter of software licensing to a deeper question about who controls access to knowledge and information. Wang argued that open source AI is "very, very, very critical for every user and developer" precisely because it offers a mechanism for making AI more transparent to end users, giving them the right to understand how models retrieve and process information from large datasets .

Wei Wang also described his personal journey with open source, beginning with Apache and Tomcat during the early internet era, and later attending open source summits, which shifted his focus from building projects to studying the communities and developers behind them . He drew on this academic experience to argue that open source projects and communities serve as irreplaceable real-world training environments for engineering talent . He described how his research group mines data from GitHub and Hugging Face using their own open source project called OpenDigger to gain insight into developer behaviour . In the AI age, he contended, the ability to write code is no longer a differentiating skill - what matters is "your judgement, your taste, and the quantity you can control" . Open source projects, he argued, provide students with real industry-relevant tasks that cannot be replicated in a conventional classroom setting, making open source communities "the real course" for engineering talent cultivation .

Christofer Dutz offered a more operationally grounded perspective, drawing on his experience with the Apache PLC4X project - a niche industrial automation project within the Apache ecosystem . He described how the emergence of AI had initially manifested in open source communities through a surge of trivial pull requests, such as suggestions about punctuation in documentation, which he and his colleagues initially found puzzling . It soon became apparent that many contributors were attempting to build reputations by accumulating merged pull requests across numerous projects . On larger Apache projects such as Apache Airflow, this phenomenon had escalated to as many as 170 pull requests per day, each requiring review, triage, and feedback . Dutz described this as placing "a huge stress on maintainers of larger or more famous open source projects," warning of "quite a threat of burning out core assets of our major open source projects" .

In response to this challenge, the Apache Software Foundation launched a project called Apache Magpie - an AI-based tool designed to review, sort, and consolidate pull requests, thereby reducing the burden on human maintainers . Dutz expressed cautious optimism, arguing that open source communities simply need to "survive this tsunami," after which the code that endures will be in a stronger position for having been reviewed by so many "digital eyes" . This dynamic - using AI to manage the problems created by AI - would recur at several points in the discussion.

#

The Tension Between Open Models and Closed Data

A central and sustained theme of the panel was the conflict between openly available AI models and proprietary or legally restricted training data. Dutz introduced an important technical distinction, noting that most AI models are distributed as "open weights" rather than being truly open source in the traditional sense . He referenced the Linux Foundation's taxonomy, which distinguishes between open models, open weights, open data, and potentially a fourth level of openness - though he acknowledged uncertainty about what that fourth level entailed . Crucially, he argued that the Apache Software Foundation's foundational principle - that two people building from the same source package should obtain identical results - is fundamentally incompatible with AI model training, because "even if I run the same training on the same data on the same infrastructure at different times, I'll get different results" . This non-determinism, he suggested, means that traditional open source reproducibility principles simply cannot apply to AI.

Jianmin Wang approached the same tension from an industrial perspective, observing that "there is a conflict today about the open data closed data set and the open source training code" . Drawing on his experience building IoTDB, a database for industrial applications, he noted that enterprises regard their operational data as private assets and properties . His proposed solution was twofold: first, open source training code could be made available for enterprises to download and use to train models within their own private environments, allowing models to absorb the characteristics of proprietary datasets without exposing the data itself ; second, AI-generated synthetic or simulated datasets could be shared publicly as a workaround, enabling domain knowledge to be disseminated without compromising confidential information .

Wei Wang added a further layer of complexity by challenging the very concept of "open source AI" as it is currently used in industry. He argued that open source AI is inherently more complex than open source software, encompassing at minimum the model, the data, the publication report, and the source code . There is currently "no agreement" on what level of openness is required across these elements, and many corporations claim their models are open source as a "marketing story" without genuine transparency . In response, his laboratory is working to develop a standardised transparency spectrum - a catalogue that classifies the actual level of openness of AI models - so that users can understand precisely what they are permitted to do with a given model, whether for commercial use, research, or other purposes .

Dutz introduced an additional and often overlooked barrier to open data: copyright law. Drawing on the example of AI companies training models on entire library collections, he argued that while the act of learning from such data is analogous to a person reading books - which is legally permissible - requiring companies to release that training data would in many cases constitute copyright infringement . This means that even organisations genuinely wishing to publish their training data may be legally prevented from doing so, a constraint that no technical solution can easily circumvent .

#

Multimodal AI: Distinct Challenges and Opportunities Within the Open Source Ecosystem

The moderator broadened the discussion to consider how large language models of different modalities - language, vision, time series, and world models - present distinct challenges and opportunities within the open source ecosystem . Jianmin Wang observed that industrial and enterprise environments generate not only text but also large volumes of relational data from ERP and supply chain management systems, as well as time series data from machines . He argued that there is a significant opportunity to train domain-specific multi-modal foundation models combining these data types, and that this represents "several tracks to develop the open source and for the special domains large models" .

Dutz offered a more differentiated assessment of the different modalities. He described language models as relatively mainstream - "my dad knows how to use them" - in contrast to vision models, which he characterised as raising more pressing legal and ethical challenges . Vision models, he argued, create acute copyright conflicts around the line between a generated image and a copy of a copyrighted work, as well as serious concerns about deep fakes and the misuse of other people's images . World models - which he described as the "holy grail" - would theoretically enable the prediction and simulation of physical system behaviour based on device specifications alone, but the training data and computational capacity required "just exceeds our human data centre capacity" . Time series models, by contrast, he characterised as a practical "world model light," citing IoTDB - Jianmin Wang's project, introduced earlier in the discussion - as a ready-to-use tool that allows existing industrial systems to be observed and optimised over time without requiring full physics modelling .

Jianmin Wang reinforced this point by noting that while language models benefit from the vast quantities of text available on the internet, world models and physical AI face a significant data scarcity problem . He described his team's contribution of a time series file format - the TS file - to Hugging Face as a practical step towards standardising data formats for domain-specific AI , and called for equivalent standards for physical data collections, framing this as an "AI for good" initiative . Wei Wang added a broader observation about the economic dimension of AI-generated content, arguing that the AI age raises urgent questions about how to return commercial benefit to the contributors whose work underpins open source and AI systems . He suggested that building economic incentive mechanisms for contributors could make open source AI ecosystems more sustainable, concluding with the provocative observation that "the whole world is broken - maybe we can build a new one" .

#

Computing Resources as a Barrier to Open Source AI Participation

An audience member raised the question of whether open source foundations such as the Apache Software Foundation could address the unequal access to computing resources - particularly GPUs - that is required for AI training . Dutz was candid in his response, explaining that the ASF "doesn't own or run very much hardware on its own" and relies on contractors for its infrastructure needs . He acknowledged that building data centres full of GPUs would require donations in the "two to three digit million" range from large AI companies before such infrastructure could be contemplated . Jianmin Wang framed this as "another dimension for open source - the open source for the computing power resource," implying that equitable access to computing infrastructure is a distinct and important aspect of openness that the community must address . This exchange highlighted the significant gap in computational resources between open source communities and large technology companies.

#

The Future Outlook for Open Source Software in the Age of AI

A second audience question, from a representative of the Chinese Institute for AI Development Strategy and the World Federation of Engineering Organisations, invited the panellists to reflect on the long-term outlook for open source after surviving the current wave of AI disruption . Dutz returned to the PLC4X project as an illustration, arguing that open source's transparency - once seen as a vulnerability because it exposes code to scrutiny - becomes a long-term competitive advantage in the AI era . Having been reviewed by "so many human eyes and so many artificial eyes," surviving open source projects will be able to demonstrate a level of robustness and trustworthiness that proprietary software cannot match .

Gao offered a historical perspective, tracing the evolution of software development from assembly language through high-level languages to natural language programming, and suggesting that the tools and skills required for software development are undergoing a fundamental transformation . Jianmin Wang formalised this observation into a three-stage evolution of software engineering: the first stage centred on high-level programming languages, the second on foundation models as software components, and the third on natural language prompting . He argued that future software production will combine all three formats, representing a structural shift in how software is conceived, built, and maintained.

Wei Wang articulated a vision of AI as a digital public good, comparable to air and water, particularly in the domain of education . He argued that students and teachers should not be priced out of using AI tools, and that open source represents the best mechanism for delivering AI as a genuinely accessible public resource . Dutz elaborated on the practical implications of this shift through the concept of specification-driven development - writing rules and architecture in plain text such as Markdown, which AI then translates into working code . He described a personal experiment in which he wrote the specification based on an existing Java driver, then asked Claude to implement it in Rust - a language he does not know - receiving a perfectly functioning result . This, he argued, means that "being able to formulate what you actually want to do without sort of like dealing with the syntax" will open many more doors to participation in open source and software development more broadly .

#

Recursive Data Poisoning and the Sustainability of Open Source AI Ecosystems

In the panel's closing exchanges, Dutz introduced one of its most consequential insights: the risk of recursive data poisoning. He acknowledged that AI is already removing barriers to open source participation by helping contributors with strong domain knowledge but limited software engineering skills to contribute more effectively . However, he warned that scaling this up would result in large volumes of AI-generated code entering open source repositories, which would subsequently feed back into AI training datasets - a feedback loop he described as "recursive poisoning of AI datasets, which everybody's afraid of" . He suggested that the donations open source foundations are receiving from large AI companies may be partly motivated by these companies' fear of "starting to eat their own dog food all the time," as the quality of their training data degrades through this recursive process .

Jianmin Wang responded by reaffirming the importance of data as "fundamental materials for software development" and by situating the current moment within his three-stage model of software engineering evolution . He suggested that the combination of high-level languages, foundation models, and natural language prompting represents a new production style for software that the community is only beginning to understand .

#

Conclusion

The panel concluded with Gao summarising the key themes addressed: the tension between open models and closed data, the challenge of balancing openness with legal and commercial constraints, and the integration of AI into open source practice and governance . The discussion suggested broad agreement that open source represents an important path towards inclusive and trustworthy AI, though significant challenges around data rights, governance standards, contributor sustainability, and the maintenance of code quality in an era of AI-generated contributions remain . The session closed with an invitation for all participants to take a group photograph, reflecting the collegial and collaborative spirit that characterised the discussion throughout .

Yue Gao
Hello everyone. I know it's the last day of the conference, but still we've got exciting topics to discuss. So, dear distinguished guests, ladies and gentlemen, good morning. It's my great honor to moderate this panel at WSIS 2026. The theme of our panel today is Open Source and AI for Sustainable Development. Digital Infrastructure, Open Data, and Multi -Stakeholder Cooperation. As we are witnessing today, artificial intelligence is reshaping our society and economy at an unprecedented pace. At the foundation of this transformation, open source software and open data continues the backlog of digital infrastructure from digital infrastructure to open data. From Linux to Apachefrom HyRub to PyTorch. Open source is not merely a technical collaboration model. It is a governance philosophy that promotes global knowledge, sharing, and narrowing the digital divide. Meanwhile, the achievement of the Sustainable Development Goals, SDGs, relies on open, trustworthy, and collaborative digital infrastructure. Whether in climate monitoring, public health early warning, or smart city development, we have witnessed that open source means the AI when the fluidity and open data emerge. With the synergy of multi -stakeholder collaborations, truly inclusive, and, and resilience and digital solutions will emerge. Today, we're gathering here to explore From the perspectives of digital infrastructure builders, open source governor practitioners, and academic frontier explorers, how can open source and AI better serve global sustainable development goals? What tension exists between the open models and closed data? And how does the large language model of different modalities present distinct characteristics within the open source ecosystem? Now, please join me in welcoming three distinguished panelists to the stage. So our first panelist is... Professor Jianmin Wang, Dean of the School of Software at Tsinghua University. Professor Wang has long been dedicated to research in software engineering and data management and making outstanding contributions to advancing open source software education, industry software innovation, and data infrastructure development in China. Welcome, Professor Wang. Please be seated here. Our second guest is Christopher Duce. Board of the Apache Software Foundation, as one of the world's most influential open source foundations. ASF stammers over 350 open source projects, including such as community overcode. That's the philosophy of the organization, has profoundly shaped global open source governance. So welcome, Chris. Our third panelist is Professor Wei Wang from East China Normal University. Professor Wang has conducted extensive research in open source software, supply chains, open source governance, and digital education, and is a key advocate for open source academic research and talent cultivation in China. Welcome, Professor Wang. Okay, let's begin our first question to our panelists. Could each of you briefly introduce yourself? What is your journey with open source and AI? and sharing your understanding of what open source and AI means to you. Who would you like to first? Yes. Okay, Professor Wong, please.
Wei Wang
with open source and AI? and sharing your understanding of what open source and AI means to you. Who would you like to first? Yes. Okay, Professor Wong, please. Okay. Distinguished Professor Gong and distinguished guests, and I welcome you to our session about the open source. I think yesterday we have also a panel on the open source, and I think the open source has become very important in the AI area. I think as the leader of the open source, the pioneer of open source, Richard Stoneman, had asked a question, who controls your computer? Yes. And... closed -source software area. But I think today the large -language model, everything we can ask the large -language model. So who controls your mind and your knowledge? I think it's a very important thing. So the open -source in the AI area is very, very, very critical for every user and developer. So I just give my view the open -source of the AI would be a great tool to make the AI more transparent to the end users. And it gives the right if the user wanted to know how the mechanism in the big data is reach and get the information from open source code models and the data set is this is my opinion thank you.
Yue Gao
thank you for sharing very insightful insights about the open source and the AI so who would like to follow in on maybe Chris yeah
Christofer Dutz
yeah maybe a little explanation this is my favorite open source project at Apache and I'll just take that project as a little example to how I observed the emergence of AI so at first we started getting pull requests pull requests about very important stuff like should there be a comma at this at this point in our documentation or should there be a full stop at the end of this code comment At first, we were a bit worried. So why is somebody paying attention on this level? But we noticed pretty soon that a lot of people seem to be trying to build up a reputation by bragging with having pull requests merged in many, many different projects. Well, for our project, we're a pretty small project. The Apache Pulse for X project is, let's say, a niche project. But Apache, we have some projects like Apache Airflow, for example. There, the projects are getting up to 170 pull requests per day. And imagine that every pull request has to be reviewed, triaged, and potentially merged or given feedback. So I think you can all imagine that it's put a huge stress on maintainers of larger or more famous open source projects. So I currently see... quite a threat of burning out core assets of our major open source projects. But I'm also very confident that we'll be able to manage this, let's call it a tsunami of pull requests coming in. Because at Apache, we just recently started a project called Apache Magpie. It's a project mainly aimed at helping us ourselves. So it's an AI -based tool for reviewing pull requests to sort them, merge them together, because many people are reporting the same things. So it's helping reduce the load. And so after quite a while of thinking AI will break open source, I think we just need to survive this tsunami. And as soon as that is done, all open source that survived this time will be in a lot better place because so many APIs are now available. And so I think we just need to survive this tsunami. And so many, well, let's say digital eyes. have reviewed the code, and I think open source will be in a much, much better state after this time. We just need to survive it.
Yue Gao
Okay, thank you, Chris, for sharing the very, very interesting product, as well as bringing your pet to our audience. Thank you. Professor Wong, do you have any comments?
Wei Wang
in university, the first open source project was Apache, and then Tomcat. Because in the age of the Internet, I used them to build websites. And when I was a staff in university, I had reached the offline open source event, for example, the open source summit, et cetera. And I realized not only... I was building the open source project, but the community... the developer and the maintainer behind the open source project. This is very interesting. So I focus not only my teaching course. I have put some open source project in my course, but also I have doing some research on the open source because my school is the data science. So I study and mine the data of the GitHub and the HackingFace so that I can have an insight of what the developers do and when they do. The report is based on the data analysis of our own open source project named OpenDigger. So in that way, in the age of the artificial intelligence, I think open source... open source is a very useful way to training our engineers. because anyone can code and can write code using AI. But the most important ability is your judgment, your taste, and the quantity you can control. So open source is the real case related to the industry. They are the real task you can sell when you're in the campus. So the open source project and the community is the real course I think in the university can train our talent of the engineering. So I think the open source is very wonderful.
Yue Gao
Thank you all for sharing your insights regarding the AI and open source. From your remarks, we have heard diverse perspectives from academia, foundation governance, and industry practice. Exactly what we're hoping through the panel providing value. Thank you to our audience. But following the first question, actually, there is a following on question to be followed. That, you know, you mentioned quite a lot on the AI and open source. And normally the models are some, at least some of the large language models are open models, open source. And then normally the data are closed. So how do you think the closed data work with the open models, especially from the open source perspective? As the example, you guys has been given early that you have open source product, meaning that those products are open data. Will that continuously to creating more open source products? How the AI will help you to creating the products? all the AI will answer your questions as Chris mentioned earlier. Although you have reached a large number of questions from the society, how do you find, are they from the AI or from the real users or human users? So that's the quite intricate question. I believe the audience would like to hear your response or your thoughts about that.
Christofer Dutz
Yeah. Well, let's say you mentioned that most AI models or that there are some open source ones. However, I would say most of them are, let's say, open weights. And I think the Linux Foundation even distinguishes, I think, three or four different levels, just open model where you sort of get the, let's call it the AI binary, open weights, which includes, let's say, the settings and the build infrastructure and open data, which also contains all of the information things were trained on. Gee. And I think the last one, I don't even know what the fourth one was, but please forgive me for that. But I just wanted to say most are distributed as open weight models. but I think we can't at Apache for example we never actually release binaries we always call them convenience binaries what the Apache Software Foundation releases is source code so the idea is that two people here in this room who download the same source package run the build should get the same result out and with AI that's out of the question because even if I run the same training on the same data on the same infrastructure at different times I'll get different results they will never be identical the open data will definitely help that I get a model out that has let's say the same biases the same yeah vibe to it will produce similar results but talking about a deterministic build of something that's out of the question in the whole AI world anyway
Yue Gao
Thank you Chris Yes
Jianmin Wang
just like said by Chris yes there is a conflict today about the open data closed data set and the open source training code I think and as my experience to build the IoTDB a database for the industrial applications I think the industrial enterprise they considered considered their The data they own is their private assets, properties. So I think the open source can make the training set, the training code, downloaded by the enterprises, and they can train the models in their private circumstances, in their private environment. And the models can take some characteristics of their data set. And for the open source, they can do some simulation and generating the artificial data set. Yes, the artificial data set, I think, may be public. And we can share the data set generated by a domain. Large foundation models, yes, then we can share the simulation data together. I think maybe there's a way to solve that in the special domains, especially in the industrial domains, how to share their data sets. This is my opinion.
Wei Wang
Okay. Thank you very much. So when we talk about open source AI, it's a big part because the open source artificial intelligence is more complex than software. The open source AI not only includes at least the models, the data, the publication report, as well as the source code. So when and what level the openness of these elements, we have no agreement on this. All of the corporates say the model they have is the open source model. Which is the model? Marketing story. But essentially, the user wants to know what we can do with the model. So in our university, in my lab, we want to build a standard. That is the spectrum. We want to catalog what the level of the transparency of the model, what really they open, and what the openness level of this. So in this way, the user can know what we can do with the model. You use it in the commercial or using with the data and the publication for the research way, which is the most important way, I think, in the following years. Okay.
Christofer Dutz
Thank you. Yeah, of course. Because it just came to me that one really important aspect that we haven't talked about is trademarks, for example. I think I read some articles on how Google trained their models. So like. like feeding all of the books of a library into training. And this is okay because it's like if I had the time and would enjoy going to a library and just read all of the books, I could run through the world and quote all sorts of books, and that would be fine. But if Google now would need to release all of the training data, they would actually be violating a lot of copyrights. So sometimes even if the companies would really like to publish the training data, it's just not possible from a copyright perspective.
Yue Gao
Absolutely. Thanks for sharing. And this is indeed a very intriguing problem and is a big challenge ahead between the open model and copyrights. Close data and how they're influencing the future AI. different type of models. So as we mentioned earlier, that large language model now spans multiple modalities, we already touched a bit, language model, vision models, word models, time series models, as well as there are more coming up. So in your view, what distinguished opportunities and challenges do large language models of different modalities facing within the open source ecosystem?
Jianmin Wang
Yes, I think, yes, nowadays the large language model is the popular one, the problem. And I think this in the industries and in the enterprise, we also have a lot of big data, such as relational data in our ERP systems, SCM systems. and we also collect a lot of time series that are generated by the machines. There is also a time series model, right? So I think we're not only paying attention to the large language model. There is some other multi -model, yes, multi -type data set. Yes, we can combine and train the synthesis, the multi -enterprise large language models, yes. So now in the enterprises and the industries, we find there is a chance to train their traditional large language model and the large time series and the large time series models. for their own business. So I think it's several tracks to develop the open source and for the special domains large models, the foundation models. Yes, it's my opinion.
Yue Gao
Okay, thank you very much.
Christofer Dutz
So for me, let's say the language models, they're sort of that, yeah, let's say my dad knows how to use them, so I think it's pretty much more mainstream. It's what we all know with ChatGPT and CloudCode, but let's say the vision models, that's a bit more tricky because first of all, the training is a lot more yeah, let's say, going to be in conflict with ... with copyrights and stuff like that. So is this image, is this a real copy of a copyrighted image, or is it a very good? So where to cut the line? And with all of these deep fakes in all sorts, so not only there's a much more pressing legal problem here, but you also, let's say, the safeguards that need to be put up. So it's totally fine if I upload a picture of myself and say, well, put me in a Star Wars costume, that's fine. But uploading the image of someone else, that's a problem. So how do we detect this? Then I'd say the world models, that would be absolute my, yeah, the holy grail of models. Because for me, well, with PLC4X and my focus on industrial automation, with a world model, I would be able, let's say, I'd call it a physics model. It would be able to predict. and calculate up front the behavior of a system just based on, let's say, the specification of the device. But I think we can all imagine the training data and the amount of training that we would need for such a model just exceeds our human data center capacity, I'd say. But if we want to improve production, for example, the time series models are, for me, sort of like a world model light. There are loads of tools that are already finished and ready to use, like IoTDB, for example, where we can observe an existing system. We don't need to model its physics. We just need to observe it. And based on that, we can do predictions on optimizations of that. We can't use it to plan a system up front, but we can use it to take an existing system and optimize that over time. yeah so for me yeah that's sort of like classification.
Yue Gao
thank you for sharing yeah
Wei Wang
Yeah I think the AI have changed a lot but only in the copyright and the law side and in the AI age I think the digital public good is very important and because the AI can make the content very easily they so how can you come to come back this benefit to the contributors for example in the open source software age many contributors have to really their time to contribute to the software but they have not benefited from the open source project of the commercial side and and also in the age of AI if we can construct a system that adding contributors to the open source system can also get feedback or get commercial benefit from this system I think this system
Jianmin Wang
can be run even more better and I think this is an opportunity we can because the whole world is broken maybe we can build a new one ok I just give another comment I think just said Chris and Professor Wong and there is in the traditional area such as large language model there is a lot of data in the internet and in the web but as to the world model for the physical model I think it's lack of the data such as for the space and time so I think So we must collect more data for the embodied AI and for the physical AI. So nowadays there is a nice standard to build these data sets. They can be used and exchanged and synthesized together for the data set. So I think maybe this is a chance for us. Such as for the time series data, we just provide a time series file format. We call it as a TS file. And we provide it to the Hugging Face as a data format for the time series. But I think there is a lack of data format, a data set format, for the... physical data collections. Yes. I think it's maybe the AI for good. We can make the general data set format for the physical AI or for the world models. Thank you.
Yue Gao
Okay. Thank you, our panelists, for sharing their thoughts. I think we have throughout enough kicking off questions. I believe our audience might have their own view and seeking opportunities to ask questions to our panelists. So in our audience, if you have any questions, please, we have a microphone. And tell your name and your affiliation and then shoot your questions. Questions or and closed data and open source models and your opinions as well. Anyone? Any questions or comments?
Audience
Yeah, sorry. Yeah. I want to ask a question to Chris. Actually, I think the conflict does not only exist between the open model and closed data set, but also the computing resource and the model. So will the open source software foundation like ASF will solve this problem or, for example, provide computing resource to the contributors for later?
Christofer Dutz
Yes. Yes. I want to solve the world foundation. to AI or yeah well the Apache Software Foundation actually doesn't own or run very much hardware on its own so so that's things we usually have have contractors for so right now I don't really see it's like the ASF building up data centers with full of GPUs that we can start training things on well maybe some of the big AI companies will drop a few well two to three digit million number on us and donations and then we can maybe think about that but I think right now the ASF doesn't have the the funds to set up data centers for
Jianmin Wang
I think professor Huang asked a question about another dimension for open source the open source for the computing power resource
Yue Gao
thank you for your question anyone else?
Audience
I just want to I'm from the Chinese Institute for AI Development Strategy and the World Federation of Engineering Organizations I just want to put a general question to all of you what is your view to the technical trends of open source software Chris talked about the tsunami of AI to the open source what is your vision to that if you survive from the tsunami what is the new So, Outlook. What is Outlook? After Survive.
Christofer Dutz
Well, let me take Pilted Rex as an example. Because when I'm trying to advocate use of open source in the automation industry, I usually run into walls. People say, well, Siemens, Schneider, Rockwell, they all have this one million years of experience, so you can never compete with that. And, well, if it's open source, anybody can spot the issues, the zero day exploits and stuff like that. But I think after we've survived that, open source, being open source actually becomes an advantage because we can say, well, so many human eyes and so many artificial eyes have reviewed all of this code and we've addressed all of the major issues. So we're very confident that we've closed. almost all of them and this is something that doesn't count for proprietary software because the companies themselves will probably run similar models on them but not the whole world. So I think
Yue Gao
Thank you Chris. Yes. This is really I think it's really challenging for the software engineering and for the workers in this domain. I think look back to the history of the software development we have evolved from the assembly language to the high level language made up to the natural language programming. Yes, this is one phenomenon we can observe today. And I think there is another side. and what's the truth we call it as a matter programming and how to make the language the natural language to be more generally to more sufficient efficient program I think it's also needed a tourist however I say today with world AI for AI I think the code for code yes but I think at the matter of philosophy there is the truth have changed notice maybe from the GPU to GPU and to the AI infrastructures I think there is a lot of work to do for the software and in years and also there she's also phenomena as the development of the software production in mood the software application applied widely, yes, and there is a large room for the software application. But I agree with you. The language or the tools or the skills for the future software may be changed, not just upon the high -level languages today. Thank you.
Wei Wang
Thank you. I think the artificial intelligence is a new digital infra of the public, just like the air and the water, especially in some public domain, for example, education. Because every student and the teachers want to learn and teach with AI. They must not pay much for this token. I think in the age of the education, they should free to get the AI and get the token in order to they want to play to be more efficient. So in this way, I think the open source is the best way to implement implementation list with the economic system of the open source. But with a lot of challenges at this time. But I think maybe we have enough smart way to build this system together to make a really AI for good and AI for our next
Yue Gao
Okay, thank you. Thank you for the question. Chris, if you want anything to add on.
Christofer Dutz
You brought up a very interesting thing. And I think it's worth emphasizing on that. So till now, participating in open source was sort of like, yeah, this project is written in Java and you need to select this girl. in Java, but actually the problem you're trying to solve isn't a Java problem. It's an algorithmic problem or a specification problem. And one thing that I started playing with years ago was so -called specification -driven development. So you actually just write text. You define the rules and the architecture of a project in Markdown or ASCII .org and it's checked in with the project. And with the emergence of AI, this suddenly started becoming just like developing on steroids. So I do everything with AI. I do specification -driven. And it even goes so far that I did a little experiment. I wrote one of my industrial drivers in Java, but I did it by specifying each detail in a specification. And then I said, Claude, please implement this driver. And I personally, I like Rust, but I have no idea how to do it. to use it. But it was actually able to write a driver that was perfectly communicating in a programming language I had no idea about. And I think this, so like being able to formulate what you actually want to do without sort of like dealing with the syntax, that's going to be opening a lot more doors in the future.
Yue Gao
Absolutely. Yeah. Thank you for sharing the insights. Let me ask our audience first. I have four or five questions actually. But I think I'll give the opportunity to our audience. Any questions? No? Okay. I'll ask further questions regarding the questions we discussed earlier. It's a lot more to think, right? For example, Chris, you just mentioned a very, very typical example. Most of software engineers experience. that he or her, she is good at one type of language. I think we transfer from the natural language to programming language. They are all language models. That's why you've got clunky code or related codecs or related things. And then you mentioned that you are experiencing one particularly good at one language that you can transfer to the others. And then there will be a lot of cases like this. Will this impact the open source? You know, the open source from the AI -generated codes and from all different type of inputs. And those will get in... Whether we'll have flooding of more available, you know, codes in your open source domain, or especially, for example, your Apache foundation. Have you seen the trend? And... Or if... You haven't, all four professors... do you feel this is the trend? This will be a wave for open source. And that's one way we're thinking AI helping the software or open source community. And the other way is that the open source data, that these are all programming data, right? Those have to be, I have to say they are quality data, even Github or more. That's why you're making this domain even better than the region model has been existing for years. But the program is only like an autonomous car or driving. Those that have confined quality data, they can use and then they build good region models. But for programming, that will have a large community and they are able to build such data sets available already. And on the other hand, we'll, those data sets help the AI. So this is a two -way exchange. I wanted to see your opinion as the final question of this panel.
Christofer Dutz
Okay, thank you. Well, short -term, I see a lot of benefit as it's removing barriers. For example, in PLC -FREX, we have several people that have a strong automation background, but they don't have a strong software engineering background. In the past, some of the contributions were, well, let's say, sub -ideal quality, or they ruined architecture and stuff like that. But AI is helping them to participate. But if we scale this up, we will have a lot of code generated by AI, which is going into open source, which will go again into training. And this is sort of like the recursive poisoning. of AI data sets, which everybody's afraid of. And I guess a lot of the donations... the open source foundations are getting are because the companies behind these big models are really scared about starting to eat their own dog food all the time. So, yes, it's got its advantages, but we also need to be careful about how it's affecting the quality in the future.
Yue Gao
Thank you.
Jianmin Wang
Yes, just as you said, I think in the future, the generation and the production efficiency may be accelerated by the AI. I think it's another side is we need the data. Yes, the data set is fundamental materials for the software development. Yes, I think last year. Someone has said that the software engineering has up to great to the three point year old. Yes, the one is just for the high level language. the second one is foundation models yes the foundation model is a part of software and software system and the third is a prompt we just write with natural language I think maybe the the software is a combination of the three of the format for the new production styles, thank you.
Yue Gao
Okay, thank you very much for our three panelists and due to the time constraints and what we have today we discussed a number of issues regarding the open models closed data and in between and how we're balancing these two as well as we have humanity said how we implemented it and into the open source as well as the large language model or AI elements. So thank you very much for coming, and that's all for our today. Let's give a round of applause to our panelists, and thank you very much for your interest. I want to invite all the participants to take a photo together. Okay? Okay. Would you be able to take a photo, all of us, together? Thank you. Yes, you have a question?

Disclaimer: This is not an official session record. DiploAI generates these resources from audiovisual recordings, and they are presented as-is, including potential errors. Due to logistical challenges, such as discrepancies in audio/video or transcripts, names may be misspelled. We strive for accuracy to the best of our ability.