Open Source and AI for Sustainable Development: Digital Infrastructure, Open Data, and Multistakeholder Cooperation
This panel discussion, moderated by Yue Gao at WSIS 2026, focused on the role of open source software and artificial intelligence in advancing sustainable development goals, bringing together perspectives from academia, open source governance, and industry practice .
Wei Wang opened by framing open source AI as a critical tool for transparency, arguing that just as Richard Stallman once asked "who controls your computer?", the rise of large language models raises the equally pressing question of who controls users' knowledge and minds . He also emphasised that open source projects provide students with real-world engineering tasks, arguing that the most important skills in the AI age are judgement, taste, and quality control rather than mere code generation . Christofer Dutz highlighted a practical challenge facing open source foundations: the surge of AI-generated pull requests is placing enormous strain on project maintainers, with some Apache projects receiving up to 170 pull requests per day . In response, Apache launched a project called Apache Magpie, an AI-based tool designed to review, sort, and consolidate pull requests to reduce this burden .
A significant portion of the discussion addressed the tension between open AI models and closed training data. Jianmin Wang noted that industrial enterprises regard their data as private assets, and suggested that open source tools could allow companies to train models locally, while artificially generated datasets might be shared publicly as a workaround . Dutz added that even willing companies may be legally prevented from releasing training data due to copyright constraints , and Wei Wang called for a standardised transparency spectrum to clarify what is genuinely open in any given AI model .
On the topic of multimodal AI, the panellists distinguished between language models, vision models, world models, and time series models, noting that each presents distinct opportunities and challenges within the open source ecosystem . Dutz expressed concern about recursive data poisoning, whereby AI-generated code enters open source repositories and is subsequently used to retrain AI models, potentially degrading quality over time . Jianmin Wang concluded by suggesting that software engineering has evolved through three paradigms - high-level languages, foundation models, and natural language prompting - and that future software will likely combine all three .
Overall, the panel converged on the view that open source remains the most viable path towards inclusive and trustworthy AI, provided that the community can navigate challenges around data rights, governance standards, and the sustainability of open source maintainership in the AI era .
Overall Purpose
- The panel, moderated by Yue Gao at WSIS 2026, aimed to explore the intersection of open source software and artificial intelligence in the context of sustainable development. Bringing together academics and open source governance practitioners, the discussion sought to examine how open source and AI can serve global development goals, address tensions between open models and closed data, and consider the challenges and opportunities posed by multi-modal large language models. ---
Major Discussion Points
- The philosophical and governance importance of open source in the AI era: Panellists framed open source not merely as a technical practice but as a governance philosophy critical to transparency and user empowerment. Wei Wang drew on Richard Stallman's question "who controls your computer?" to pose a new challenge in the AI age: "who controls your mind and your knowledge?" Open source AI was seen as a tool to make AI mechanisms more transparent to end users, covering models, datasets, source code, and publication reports. - The tension between open models and closed data, including copyright complications: A central theme was the conflict between openly available AI models and proprietary or legally restricted training data. Christofer Dutz distinguished between "open source," "open weights," and "open data" models, noting that even willing companies may be unable to release training data due to copyright constraints - for instance, training on entire library collections. Jianmin Wang suggested that enterprises view their data as private assets, and proposed solutions such as local training on private infrastructure and the use of AI-generated synthetic datasets for sharing. Wei Wang called for a standardised transparency spectrum to help users understand what is truly "open" in any given model. - The impact of AI-generated contributions on open source communities and the risk of recursive data poisoning: Dutz described a "tsunami" of AI-generated pull requests flooding open source projects - some Apache projects receiving up to 170 per day - placing enormous strain on maintainers. In response, Apache launched the Magpie project, an AI-based tool to help review and triage contributions. Looking further ahead, Dutz warned of a recursive poisoning risk: AI-generated code entering open source repositories could feed back into future AI training data, potentially degrading quality over time. - Multi-modal AI models and the distinct challenges and opportunities they present within open source: The panel discussed how AI extends well beyond language models to include vision, time series, and world models. Vision models raise acute copyright and deep-fake concerns. World models - described by Dutz as the "holy grail" - would require training data and compute capacity that currently exceeds available infrastructure. Time series models, such as those supported by IoTDB, were highlighted as a practical near-term tool for industrial optimisation. Jianmin Wang noted a lack of standardised data formats for physical and embodied AI datasets, presenting an opportunity for the open source community to contribute. - The future of software engineering and open source in an AI-driven world: Panellists reflected on how AI is transforming software development itself - from high-level programming languages towards natural language and specification-driven development. Jianmin Wang described a three-stage evolution of software engineering: high-level languages, foundation models as software components, and natural language prompts. Wei Wang argued that open source remains the best vehicle for making AI a genuine public good - akin to air and water - particularly in education, where access to AI tools should be free or low-cost. ---
Overall Tone
- The overall tone of the discussion was collegial, intellectually engaged, and cautiously optimistic. The moderator and panellists maintained a respectful, collaborative atmosphere throughout, consistent with an academic and policy-oriented conference setting.
- Early in the discussion, the tone was largely exploratory and philosophical, as panellists introduced their perspectives on open source and AI. It became more technically specific and candid in the middle sections, particularly when Dutz raised concerns about maintainer burnout, the pull request tsunami, and recursive data poisoning - moments that introduced a note of genuine concern.
- However, the tone never became pessimistic. Panellists consistently reframed challenges as opportunities, with phrases such as "we just need to survive this tsunami" and "the whole world is broken, maybe we can build a new one." By the closing exchanges, the tone had shifted towards forward-looking optimism, with discussion of specification-driven development, AI as digital public infrastructure, and the potential for open source to emerge stronger from current pressures.
Expanded Summary: Open Source and AI for Sustainable Development - WSIS 2026 Panel Discussion
#
Introduction and Panel Overview
The panel, moderated by Yue Gao at WSIS 2026, was convened on the final day of the conference to explore the intersection of open source software and artificial intelligence in the context of sustainable development . The session's full theme - "Open Source and AI for Sustainable Development: Digital Infrastructure, Open Data, and Multi-Stakeholder Cooperation" - reflected the breadth of issues the panellists were invited to address . In her opening remarks, Gao framed the discussion by observing that artificial intelligence is reshaping society and the economy "at an unprecedented pace," and that open source software and open data form the foundational layer of this transformation . She cited examples ranging from Linux and Apache to Hugging Face and PyTorch to illustrate that open source is "not merely a technical collaboration model" but rather "a governance philosophy that promotes global knowledge sharing and narrowing the digital divide" . Gao further noted that achieving the Sustainable Development Goals depends on open, trustworthy, and collaborative digital infrastructure, with applications in climate monitoring, public health early warning, and smart city development .
Three panellists were introduced: Professor Jianmin Wang, Dean of the School of Software at Tsinghua University, whose work spans software engineering, data management, and open source education in China ; Christofer Dutz, a board member of the Apache Software Foundation (ASF), one of the world's most influential open source foundations, stewarding more than 350 open source projects ; and Professor Wei Wang from East China Normal University, whose research focuses on open source software supply chains, open source governance, and digital education . The panel was structured around a series of questions posed by the moderator, supplemented by audience contributions, covering the philosophical importance of open source in the AI era, the tension between open models and closed data, the challenges of multimodal AI, and the future of software engineering .
#
The Philosophical and Governance Importance of Open Source in the AI Era
Wei Wang opened the substantive discussion by invoking the foundational question posed by Richard Stallman - rendered in the transcript as "Richard Stoneman," an apparent transcription error - namely, "who controls your computer?", and extending it to the present moment . In the age of large language models, Wang argued, the more pressing question has become: "who controls your mind and your knowledge?" . This reframing elevated the open source debate from a technical matter of software licensing to a deeper question about who controls access to knowledge and information. Wang argued that open source AI is "very, very, very critical for every user and developer" precisely because it offers a mechanism for making AI more transparent to end users, giving them the right to understand how models retrieve and process information from large datasets .
Wei Wang also described his personal journey with open source, beginning with Apache and Tomcat during the early internet era, and later attending open source summits, which shifted his focus from building projects to studying the communities and developers behind them . He drew on this academic experience to argue that open source projects and communities serve as irreplaceable real-world training environments for engineering talent . He described how his research group mines data from GitHub and Hugging Face using their own open source project called OpenDigger to gain insight into developer behaviour . In the AI age, he contended, the ability to write code is no longer a differentiating skill - what matters is "your judgement, your taste, and the quantity you can control" . Open source projects, he argued, provide students with real industry-relevant tasks that cannot be replicated in a conventional classroom setting, making open source communities "the real course" for engineering talent cultivation .
Christofer Dutz offered a more operationally grounded perspective, drawing on his experience with the Apache PLC4X project - a niche industrial automation project within the Apache ecosystem . He described how the emergence of AI had initially manifested in open source communities through a surge of trivial pull requests, such as suggestions about punctuation in documentation, which he and his colleagues initially found puzzling . It soon became apparent that many contributors were attempting to build reputations by accumulating merged pull requests across numerous projects . On larger Apache projects such as Apache Airflow, this phenomenon had escalated to as many as 170 pull requests per day, each requiring review, triage, and feedback . Dutz described this as placing "a huge stress on maintainers of larger or more famous open source projects," warning of "quite a threat of burning out core assets of our major open source projects" .
In response to this challenge, the Apache Software Foundation launched a project called Apache Magpie - an AI-based tool designed to review, sort, and consolidate pull requests, thereby reducing the burden on human maintainers . Dutz expressed cautious optimism, arguing that open source communities simply need to "survive this tsunami," after which the code that endures will be in a stronger position for having been reviewed by so many "digital eyes" . This dynamic - using AI to manage the problems created by AI - would recur at several points in the discussion.
#
The Tension Between Open Models and Closed Data
A central and sustained theme of the panel was the conflict between openly available AI models and proprietary or legally restricted training data. Dutz introduced an important technical distinction, noting that most AI models are distributed as "open weights" rather than being truly open source in the traditional sense . He referenced the Linux Foundation's taxonomy, which distinguishes between open models, open weights, open data, and potentially a fourth level of openness - though he acknowledged uncertainty about what that fourth level entailed . Crucially, he argued that the Apache Software Foundation's foundational principle - that two people building from the same source package should obtain identical results - is fundamentally incompatible with AI model training, because "even if I run the same training on the same data on the same infrastructure at different times, I'll get different results" . This non-determinism, he suggested, means that traditional open source reproducibility principles simply cannot apply to AI.
Jianmin Wang approached the same tension from an industrial perspective, observing that "there is a conflict today about the open data closed data set and the open source training code" . Drawing on his experience building IoTDB, a database for industrial applications, he noted that enterprises regard their operational data as private assets and properties . His proposed solution was twofold: first, open source training code could be made available for enterprises to download and use to train models within their own private environments, allowing models to absorb the characteristics of proprietary datasets without exposing the data itself ; second, AI-generated synthetic or simulated datasets could be shared publicly as a workaround, enabling domain knowledge to be disseminated without compromising confidential information .
Wei Wang added a further layer of complexity by challenging the very concept of "open source AI" as it is currently used in industry. He argued that open source AI is inherently more complex than open source software, encompassing at minimum the model, the data, the publication report, and the source code . There is currently "no agreement" on what level of openness is required across these elements, and many corporations claim their models are open source as a "marketing story" without genuine transparency . In response, his laboratory is working to develop a standardised transparency spectrum - a catalogue that classifies the actual level of openness of AI models - so that users can understand precisely what they are permitted to do with a given model, whether for commercial use, research, or other purposes .
Dutz introduced an additional and often overlooked barrier to open data: copyright law. Drawing on the example of AI companies training models on entire library collections, he argued that while the act of learning from such data is analogous to a person reading books - which is legally permissible - requiring companies to release that training data would in many cases constitute copyright infringement . This means that even organisations genuinely wishing to publish their training data may be legally prevented from doing so, a constraint that no technical solution can easily circumvent .
#
Multimodal AI: Distinct Challenges and Opportunities Within the Open Source Ecosystem
The moderator broadened the discussion to consider how large language models of different modalities - language, vision, time series, and world models - present distinct challenges and opportunities within the open source ecosystem . Jianmin Wang observed that industrial and enterprise environments generate not only text but also large volumes of relational data from ERP and supply chain management systems, as well as time series data from machines . He argued that there is a significant opportunity to train domain-specific multi-modal foundation models combining these data types, and that this represents "several tracks to develop the open source and for the special domains large models" .
Dutz offered a more differentiated assessment of the different modalities. He described language models as relatively mainstream - "my dad knows how to use them" - in contrast to vision models, which he characterised as raising more pressing legal and ethical challenges . Vision models, he argued, create acute copyright conflicts around the line between a generated image and a copy of a copyrighted work, as well as serious concerns about deep fakes and the misuse of other people's images . World models - which he described as the "holy grail" - would theoretically enable the prediction and simulation of physical system behaviour based on device specifications alone, but the training data and computational capacity required "just exceeds our human data centre capacity" . Time series models, by contrast, he characterised as a practical "world model light," citing IoTDB - Jianmin Wang's project, introduced earlier in the discussion - as a ready-to-use tool that allows existing industrial systems to be observed and optimised over time without requiring full physics modelling .
Jianmin Wang reinforced this point by noting that while language models benefit from the vast quantities of text available on the internet, world models and physical AI face a significant data scarcity problem . He described his team's contribution of a time series file format - the TS file - to Hugging Face as a practical step towards standardising data formats for domain-specific AI , and called for equivalent standards for physical data collections, framing this as an "AI for good" initiative . Wei Wang added a broader observation about the economic dimension of AI-generated content, arguing that the AI age raises urgent questions about how to return commercial benefit to the contributors whose work underpins open source and AI systems . He suggested that building economic incentive mechanisms for contributors could make open source AI ecosystems more sustainable, concluding with the provocative observation that "the whole world is broken - maybe we can build a new one" .
#
Computing Resources as a Barrier to Open Source AI Participation
An audience member raised the question of whether open source foundations such as the Apache Software Foundation could address the unequal access to computing resources - particularly GPUs - that is required for AI training . Dutz was candid in his response, explaining that the ASF "doesn't own or run very much hardware on its own" and relies on contractors for its infrastructure needs . He acknowledged that building data centres full of GPUs would require donations in the "two to three digit million" range from large AI companies before such infrastructure could be contemplated . Jianmin Wang framed this as "another dimension for open source - the open source for the computing power resource," implying that equitable access to computing infrastructure is a distinct and important aspect of openness that the community must address . This exchange highlighted the significant gap in computational resources between open source communities and large technology companies.
#
The Future Outlook for Open Source Software in the Age of AI
A second audience question, from a representative of the Chinese Institute for AI Development Strategy and the World Federation of Engineering Organisations, invited the panellists to reflect on the long-term outlook for open source after surviving the current wave of AI disruption . Dutz returned to the PLC4X project as an illustration, arguing that open source's transparency - once seen as a vulnerability because it exposes code to scrutiny - becomes a long-term competitive advantage in the AI era . Having been reviewed by "so many human eyes and so many artificial eyes," surviving open source projects will be able to demonstrate a level of robustness and trustworthiness that proprietary software cannot match .
Gao offered a historical perspective, tracing the evolution of software development from assembly language through high-level languages to natural language programming, and suggesting that the tools and skills required for software development are undergoing a fundamental transformation . Jianmin Wang formalised this observation into a three-stage evolution of software engineering: the first stage centred on high-level programming languages, the second on foundation models as software components, and the third on natural language prompting . He argued that future software production will combine all three formats, representing a structural shift in how software is conceived, built, and maintained.
Wei Wang articulated a vision of AI as a digital public good, comparable to air and water, particularly in the domain of education . He argued that students and teachers should not be priced out of using AI tools, and that open source represents the best mechanism for delivering AI as a genuinely accessible public resource . Dutz elaborated on the practical implications of this shift through the concept of specification-driven development - writing rules and architecture in plain text such as Markdown, which AI then translates into working code . He described a personal experiment in which he wrote the specification based on an existing Java driver, then asked Claude to implement it in Rust - a language he does not know - receiving a perfectly functioning result . This, he argued, means that "being able to formulate what you actually want to do without sort of like dealing with the syntax" will open many more doors to participation in open source and software development more broadly .
#
Recursive Data Poisoning and the Sustainability of Open Source AI Ecosystems
In the panel's closing exchanges, Dutz introduced one of its most consequential insights: the risk of recursive data poisoning. He acknowledged that AI is already removing barriers to open source participation by helping contributors with strong domain knowledge but limited software engineering skills to contribute more effectively . However, he warned that scaling this up would result in large volumes of AI-generated code entering open source repositories, which would subsequently feed back into AI training datasets - a feedback loop he described as "recursive poisoning of AI datasets, which everybody's afraid of" . He suggested that the donations open source foundations are receiving from large AI companies may be partly motivated by these companies' fear of "starting to eat their own dog food all the time," as the quality of their training data degrades through this recursive process .
Jianmin Wang responded by reaffirming the importance of data as "fundamental materials for software development" and by situating the current moment within his three-stage model of software engineering evolution . He suggested that the combination of high-level languages, foundation models, and natural language prompting represents a new production style for software that the community is only beginning to understand .
#
Conclusion
The panel concluded with Gao summarising the key themes addressed: the tension between open models and closed data, the challenge of balancing openness with legal and commercial constraints, and the integration of AI into open source practice and governance . The discussion suggested broad agreement that open source represents an important path towards inclusive and trustworthy AI, though significant challenges around data rights, governance standards, contributor sustainability, and the maintenance of code quality in an era of AI-generated contributions remain . The session closed with an invitation for all participants to take a group photograph, reflecting the collegial and collaborative spirit that characterised the discussion throughout .
Open source AI as a tool for transparency and user empowerment, raising the question of who controls knowledge in the age of large language models
Arg. 1Wei Wang draws a parallel between Richard Stallman's foundational question about who controls your computer in the closed-source software era and a new, more pressing question in the AI age: who controls your mind and knowledge. He argues that open source AI is critical for making AI more transparent to end users, giving them the right to understand how large models retrieve and process information. This transparency is essential for both users and developers.
Wei Wang invoked Richard Stallman's question 'who controls your computer?' as a historical reference point for the open source movement , then reframed it for the AI era by asking 'who controls your mind and your knowledge?' . He argued that open source AI would be a great tool to make AI more transparent to end users and give them the right to understand the mechanisms by which big data models retrieve information from open source code, models, and datasets .
on: Open source is essential for AI transparency, democratisation, and serving the public good
Open source projects and communities serve as real-world training environments for engineering talent, providing practical industry-relevant experience that AI alone cannot replicate
Arg. 2Wei Wang contends that while AI has made it easy for anyone to write code, the most important engineering skills — judgement, taste, and quality control — cannot be developed through AI alone. Open source projects provide real, industry-relevant tasks that students can engage with on campus, making open source communities the ideal practical course for training future engineers. He integrates open source projects directly into his university teaching.
Wei Wang described how he incorporates open source projects into his university courses and conducts research by mining data from GitHub and HuggingFace using his own open source project, OpenDigger . He argued that in the age of AI, where anyone can write code, the most important abilities are judgement, taste, and quality control, and that open source projects provide real industry-relevant tasks students can work on while still on campus .
on: Whether the primary challenge of AI for open source is a threat to be survived or an opportunity to be embraced
There is a need for a standardised spectrum or catalogue to measure the transparency and openness levels of AI models, so users understand what they can actually do with a given model
Arg. 3Wei Wang highlights that open source AI is far more complex than open source software, encompassing models, data, publication reports, and source code, yet there is no agreed standard for what constitutes openness across these elements. He notes that many companies market their models as open source without genuine transparency, which misleads users. His lab is therefore working to build a spectrum or catalogue that classifies the actual level of openness of AI models.
Wei Wang noted that open source AI includes at minimum models, data, publication reports, and source code, and that there is currently no agreement on what level of openness is required across these elements . He observed that many corporations claim their models are open source as a marketing story, without genuine transparency . He described his lab's effort to build a standard spectrum that catalogues the transparency level of models so users know what they can actually do with them, whether for commercial use, research, or other purposes .
on: Standardisation of data formats and openness levels is necessary to advance open source AI, particularly for non-language modalities
on: What constitutes 'open source' in the context of AI models
Contributors to open source and AI systems are not adequately compensated commercially; building economic incentive mechanisms for contributors could make open source AI ecosystems more sustainable
Arg. 4Wei Wang points out that in the open source software era, contributors invested significant time without receiving commercial benefit, and this problem is amplified in the AI age where AI can generate content easily. He argues that constructing a system that provides commercial feedback or benefit to contributors would make open source AI ecosystems more sustainable and vibrant. He frames this as an opportunity to build a new, better system.
Wei Wang noted that in the open source software age, many contributors dedicated their time to projects without benefiting commercially from them . He argued that in the AI age, if a system could be constructed whereby contributors to open source receive commercial feedback or benefit, the ecosystem could run even better, and framed this as an opportunity to build a new system since 'the whole world is broken' .
AI should be treated as a digital public good, akin to air and water, particularly in education, where free or low-cost access to AI tools is essential for equitable learning
Arg. 5Wei Wang argues that AI has become a new digital infrastructure for the public, comparable in importance to air and water, and that this is especially true in education. Students and teachers should not have to pay heavily for AI tokens in order to learn and teach effectively. Open source is presented as the best mechanism to implement an economic system that ensures equitable, affordable access to AI for educational purposes.
Wei Wang described AI as 'a new digital infra of the public, just like the air and the water', particularly in public domains such as education . He argued that students and teachers must not pay much for AI tokens and should be free to access AI in order to be more efficient in learning and teaching . He concluded that open source is the best way to implement this vision, despite the many challenges involved .
Open source as a governance philosophy that promotes global knowledge sharing and narrows the digital divide, foundational to sustainable development
Arg. 1Yue Gao frames open source not merely as a technical collaboration model but as a governance philosophy with broad societal implications. She argues that open source software and open data form the backbone of digital infrastructure and that their synergy with multi-stakeholder collaboration is essential for achieving the SDGs. She highlights concrete application domains such as climate monitoring, public health, and smart city development.
Yue Gao described open source as 'not merely a technical collaboration model' but 'a governance philosophy that promotes global knowledge sharing and narrowing the digital divide' . She cited examples including Linux, Apache, HyRub, and PyTorch as illustrations of open source's foundational role . She also noted that the achievement of the SDGs relies on open, trustworthy, and collaborative digital infrastructure, with applications in climate monitoring, public health early warning, and smart city development .
on: Open source is essential for AI transparency, democratisation, and serving the public good
AI is creating a "tsunami" of pull requests in open source projects, placing enormous stress on maintainers of major projects, but open source communities are developing AI-based tools (e.g., Apache Magpie) to manage this load
Arg. 1Christofer Dutz describes how AI-generated contributions, initially appearing as trivial pull requests about punctuation, have grown into a massive volume of submissions that overwhelm maintainers of large open source projects. He warns of a risk of burning out core contributors to major projects. However, he expresses confidence that AI-based tools developed within the open source community itself, such as Apache Magpie, can help manage this load.
Dutz described observing an influx of trivial pull requests - such as whether a comma or full stop should appear in documentation - which he attributed to people trying to build reputations by having pull requests merged across many projects . He noted that large projects like Apache Airflow receive up to 170 pull requests per day, each requiring review, triage, and feedback, creating enormous stress on maintainers . He referenced the Apache Magpie project as an AI-based tool designed to review, sort, and merge pull requests to reduce this load .
on: Whether the primary challenge of AI for open source is a threat to be survived or an opportunity to be embraced
Most AI models are distributed as "open weights" rather than truly open source, and unlike traditional software, AI training is non-deterministic, making reproducible builds impossible
Arg. 2Dutz draws a distinction between open weights, open models, and fully open data, referencing the Linux Foundation's multi-level classification framework. He argues that the Apache Software Foundation's principle of releasing source code — so that any two people building from the same package get identical results — cannot apply to AI, because training the same model on the same data at different times yields different results. This non-determinism fundamentally challenges the open source ethos as applied to AI.
Dutz referenced the Linux Foundation's classification of AI openness into approximately three or four levels: open model (the AI binary), open weights (including settings and build infrastructure), and open data (including training information) . He contrasted this with Apache's principle of releasing source code rather than binaries, so that two people building from the same source package get the same result . He argued that this deterministic reproducibility is impossible in AI, because even running the same training on the same data on the same infrastructure at different times yields different results .
on: There is a fundamental tension between open AI models and closed training data that requires new frameworks and compromises to resolve
on: What constitutes 'open source' in the context of AI models
Copyright presents a significant barrier to open data release, as training data often incorporates copyrighted material that companies cannot legally publish even if they wished to
Arg. 3Dutz argues that even companies that would genuinely like to publish their training data are often legally prevented from doing so because that data includes copyrighted material such as books from libraries. He uses the analogy of a person reading all the books in a library and then quoting from them, which is acceptable for an individual but becomes a copyright violation if the content is republished at scale. This creates a structural tension between the desire for open data and intellectual property law.
Dutz referenced reports about Google training its models by feeding entire library collections into the training process . He drew an analogy: just as an individual could read all the books in a library and quote from them without legal issue, a company doing the equivalent is acceptable - but if that company were required to release all the training data, it would be violating many copyrights . He concluded that sometimes, even if companies would like to publish training data, it is simply not possible from a copyright perspective .
on: There is a fundamental tension between open AI models and closed training data that requires new frameworks and compromises to resolve
on: How to resolve the tension between open models and closed data in industrial/enterprise contexts
Vision models raise more pressing legal and ethical challenges than language models, including copyright conflicts, deep fakes, and the need for robust safeguards
Arg. 4Dutz argues that while language models have become relatively mainstream, vision models present a more complex set of legal and ethical challenges. These include determining whether a generated image constitutes a copy of a copyrighted work, and the risks posed by deep fakes, including the non-consensual use of other people's images. He emphasises that the safeguards required for vision models are significantly more demanding than those for language models.
Dutz noted that vision model training is more likely to conflict with copyright law, raising the question of where to draw the line between a generated image and a copy of a copyrighted image . He also highlighted the problem of deep fakes, noting that while uploading a picture of oneself for creative purposes is acceptable, uploading someone else's image is a problem, and detecting this distinction is a significant challenge .
World models (physics-based models) represent the "holy grail" but require training data and computational capacity that currently exceeds available infrastructure; time series models offer a practical intermediate step for industrial optimisation
Arg. 5Dutz describes world models — which he also calls physics models — as the ultimate goal for industrial AI, capable of predicting system behaviour from specifications alone. However, he acknowledges that the training data and computational resources required for such models currently exceed available data centre capacity. He presents time series models as a practical near-term alternative, capable of observing existing systems and optimising them over time without needing to model their physics from scratch.
Dutz described world models as his 'holy grail', explaining that for his work in industrial automation with PLC4X, a physics model could predict and calculate system behaviour based solely on device specifications . He noted that the training data and computational capacity required for such a model 'just exceeds our human data center capacity' . He then described time series models as a 'world model light', citing tools like IoTDB as ready-to-use examples that allow observation of existing systems and prediction-based optimisation without requiring full physics modelling .
The conflict between open models and closed data is compounded by unequal access to computing resources; open source foundations such as ASF currently lack the funds or infrastructure to provide GPU resources for AI training
Arg. 6In response to an audience question, Dutz acknowledges that computing resources represent a significant additional barrier to open source AI participation, beyond the open model versus closed data tension. He explains that the Apache Software Foundation does not own or operate significant hardware infrastructure and relies on contractors, making it currently unable to provide GPU resources for AI training. He suggests that substantial donations from large AI companies would be needed to change this.
Dutz explained that the Apache Software Foundation does not own or run much hardware on its own and relies on contractors for such needs . He stated that he does not foresee the ASF building data centres full of GPUs for AI training, and suggested that only very large donations - in the two-to-three digit million range - from big AI companies might make this feasible .
on: Computing resources represent an additional and distinct barrier to open source AI participation, beyond the open model versus closed data divide
on: The feasibility and role of open source foundations in providing computing resources for AI
Open source code that survives the current AI-driven tsunami will be in a stronger position, having been reviewed by both human and artificial intelligence, making it more robust than proprietary alternatives
Arg. 7Dutz argues that while the current wave of AI-generated contributions poses a serious threat to open source maintainers, the projects that survive this period will emerge stronger. Having been reviewed by both human and artificial eyes, surviving open source code will have had its issues identified and addressed at a scale that proprietary software cannot match. This makes openness itself a long-term competitive advantage.
Dutz used the example of advocating for open source in the industrial automation sector, where he often faces the objection that proprietary vendors like Siemens, Schneider, and Rockwell have decades of experience and that open source exposes zero-day exploits . He argued that after surviving the AI tsunami, open source can claim that 'so many human eyes and so many artificial eyes have reviewed all of this code and we've addressed all of the major issues', which is something proprietary software cannot claim at the same scale .
on: Open source is essential for AI transparency, democratisation, and serving the public good
Specification-driven development, enabled by AI, lowers barriers to participation in open source by allowing contributors to define intent in natural language rather than requiring mastery of specific programming languages
Arg. 8Dutz describes specification-driven development as an approach where project rules and architecture are written in natural language (e.g., Markdown) and checked in with the project, rather than requiring contributors to master a specific programming language. He argues that AI has supercharged this approach, enabling contributors to implement complex functionality in languages they do not know by simply specifying what they want. This dramatically lowers the barrier to participation in open source projects.
Dutz described experimenting with specification-driven development years ago, writing project rules and architecture in Markdown or ASCII and checking them in with the project . He described a concrete experiment in which he wrote a specification for an industrial driver in Java, then asked Claude (an AI assistant) to implement it in Rust - a language he does not know - and the AI produced a driver that communicated perfectly . He argued that being able to formulate intent without dealing with syntax will open many more doors in the future .
AI-generated code feeding back into open source training data risks recursive quality degradation ("poisoning"), which is a concern driving donations from large AI companies to open source foundations
Arg. 9Dutz warns that as AI-generated code enters open source repositories and is subsequently used to train future AI models, there is a risk of recursive quality degradation — a feedback loop in which AI trains on its own outputs, progressively degrading quality. He notes that this concern is a significant motivation behind large AI companies donating to open source foundations, as they fear their models will begin training on AI-generated rather than human-generated code.
Dutz acknowledged that while AI helps contributors with weaker software engineering backgrounds to participate in open source projects like PLC4X , scaling this up means large volumes of AI-generated code will enter open source repositories and then feed back into AI training data, which he described as 'recursive poisoning of AI datasets, which everybody's afraid of' . He suggested that donations from large AI companies to open source foundations are partly motivated by these companies' fear of their models 'starting to eat their own dog food all the time' .
Industrial enterprises treat their data as private assets; open source training code can allow enterprises to train models locally on private data, while synthetic/simulated datasets could be shared publicly as a compromise
Arg. 1Jianmin Wang acknowledges the fundamental tension between open models and closed data, drawing on his experience building IoTDB for industrial applications. He argues that enterprises view their data as private property and will not share it openly, but that open source training code can enable them to train models in their own private environments. As a compromise, he proposes that synthetic or simulated datasets generated by domain foundation models could be shared publicly.
Jianmin Wang described his experience building IoTDB, a database for industrial applications, noting that industrial enterprises consider their data to be private assets and properties . He proposed that open source training code could be downloaded by enterprises to train models in their private environments, allowing models to take on characteristics of their private datasets . He further suggested that open source tools could generate artificial/simulated datasets that could be made public and shared, potentially using domain foundation models to generate simulation data .
on: There is a fundamental tension between open AI models and closed training data that requires new frameworks and compromises to resolve
on: How to resolve the tension between open models and closed data in industrial/enterprise contexts
Beyond language models, industrial and enterprise environments generate large volumes of relational and time series data, creating opportunities for domain-specific multi-modal foundation models
Arg. 2Jianmin Wang argues that the focus on large language models overlooks the rich variety of data types generated in industrial and enterprise settings, including relational data from ERP and SCM systems and time series data from machines. He sees an opportunity to combine these data types to train multi-modal, domain-specific foundation models tailored to enterprise needs. He identifies this as a distinct track for open source development alongside general-purpose language models.
Jianmin Wang noted that industries and enterprises generate large volumes of relational data from ERP and SCM systems, as well as time series data from machines . He argued that attention should not be limited to large language models, as there are opportunities to combine multiple data types to train multi-enterprise large language models and large time series models for specific business domains .
There is a lack of standardised data formats for physical/embodied AI data collection; providing open file formats (e.g., TS file for time series on Hugging Face) could help address this gap
Arg. 3Jianmin Wang argues that while world models and physical AI hold great promise, progress is hampered by a lack of standardised, interoperable data formats for physical and embodied AI data. He draws on his own work providing a time series file format (TS file) to Hugging Face as a positive example, and calls for similar standardisation efforts for physical data collections. He frames this as an opportunity for the AI for good community to contribute.
Jianmin Wang noted that unlike language models, which benefit from abundant internet data, world models and physical AI suffer from a lack of data, particularly for space and time . He described his team's contribution of a time series file format (TS file) to Hugging Face as a data format for time series data . He argued that there is currently a lack of equivalent data format standards for physical data collections and called for the development of a general data set format for physical AI and world models as an 'AI for good' initiative .
on: Standardisation of data formats and openness levels is necessary to advance open source AI, particularly for non-language modalities
Access to computing resources for AI training is an additional dimension of openness that the open source community must address
Arg. 4Jianmin Wang, responding to an audience question, frames computing resource access as a distinct and important dimension of openness in the AI era, separate from the open model versus closed data debate. He suggests that the open source community needs to consider how to address unequal access to the computational power required for AI training, not just access to code and data.
Jianmin Wang characterised the audience question as raising 'another dimension for open source - the open source for the computing power resource' , framing computing access as a distinct and important aspect of openness that the community must grapple with.
on: Computing resources represent an additional and distinct barrier to open source AI participation, beyond the open model versus closed data divide
on: The feasibility and role of open source foundations in providing computing resources for AI
Software engineering is evolving through three paradigms: high-level programming languages, foundation models as software components, and natural language prompting; future software will combine all three formats
Arg. 5Jianmin Wang argues that software engineering has undergone a fundamental upgrade to what he calls 'version 3.0', characterised by three distinct paradigms: traditional high-level programming languages, foundation models as integral components of software systems, and natural language prompting. He contends that future software production will be a combination of all three formats rather than a replacement of one by another.
Jianmin Wang referenced a claim made the previous year that software engineering had been upgraded to 'three point year old' (version 3.0), comprising: high-level programming languages as the first paradigm, foundation models as a part of software and software systems as the second, and natural language prompting as the third . He concluded that future software will be a combination of all three formats as new production styles emerge .
on: Whether the primary challenge of AI for open source is a threat to be survived or an opportunity to be embraced
The conflict between open source and AI extends beyond open models versus closed data to include unequal access to computing resources, and open source foundations should consider providing GPU resources to contributors
Arg. 1An audience member raises the point that the tension in open source AI is not solely about the divide between open models and closed datasets, but also encompasses the significant barrier posed by access to computing resources. They question whether open source software foundations like the Apache Software Foundation could step in to provide computing infrastructure, such as GPUs, to contributors who lack access.
The audience member explicitly stated that 'the conflict does not only exist between the open model and closed data set, but also the computing resource and the model' . They then directed a question to Chris asking whether the Apache Software Foundation would solve this problem or provide computing resources to contributors .
on: Computing resources represent an additional and distinct barrier to open source AI participation, beyond the open model versus closed data divide
on: The feasibility and role of open source foundations in providing computing resources for AI
After surviving the current AI-driven disruption to open source, the community needs a clear vision or outlook for what the future of open source software will look like
Arg. 2A second audience member, representing the Chinese Institute for AI Development Strategy and the World Federation of Engineering Organizations, poses a broader question about the long-term technical trends and future vision for open source software. They acknowledge Chris's framing of the current AI wave as a 'tsunami' and ask what the outlook is for open source once that challenge has been navigated.
The audience member asked all panellists for their view on the technical trends of open source software, referencing Chris's tsunami metaphor, and specifically asked 'what is your vision to that if you survive from the tsunami, what is the new outlook?' .
Session Knowledge Graph
Speakers · Topics · Arguments · Relationships
All panelists converge on the view that open source is not merely a technical model but a foundational philosophy for making AI transparent, trustworthy, and broadly beneficial. Wei Wang framed this in terms of user empowerment, asking 'who controls your mind and your knowledge?' and arguing that open source AI is 'very, very, very critical for every user and developer' . Yue Gao described open source as 'a governance philosophy that promotes global knowledge sharing and narrowing the digital divide' , with applications in climate monitoring, public health, and smart city development . Dutz argued that open source projects that survive the current AI wave will be stronger for having been reviewed by 'so many human eyes and so many artificial eyes' , making openness itself a long-term competitive advantage. Jianmin Wang supported open source training code as a mechanism for enterprises to benefit from AI while retaining control of their private data .
Open source AI as a tool for transparency and user empowerment, raising the question of who controls knowledge in the age of large language models
Open source as a governance philosophy that promotes global knowledge sharing and narrows the digital divide, foundational to sustainable development
Open source code that survives the current AI-driven tsunami will be in a stronger position, having been reviewed by both human and artificial intelligence, making it more robust than proprietary alternatives
Industrial enterprises treat their data as private assets; open source training code can allow enterprises to train models locally on private data, while synthetic/simulated datasets could be shared publicly as a compromise
All three panelists explicitly acknowledged the tension between open models and closed data as a central challenge. Dutz distinguished between open model, open weights, and open data as distinct levels of openness , and argued that AI's non-deterministic training makes traditional open source reproducibility impossible . Wei Wang noted that there is currently 'no agreement' on what level of openness is required across models, data, publication reports, and source code , and criticised companies that claim open source as a 'marketing story' . Jianmin Wang observed that 'there is a conflict today about the open data closed data set and the open source training code' , and proposed synthetic datasets as a compromise . Dutz further noted that copyright law often prevents companies from releasing training data even when they would like to .
There is a need for a standardised spectrum or catalogue to measure the transparency and openness levels of AI models, so users understand what they can actually do with a given model
Most AI models are distributed as "open weights" rather than truly open source, and unlike traditional software, AI training is non-deterministic, making reproducible builds impossible
Copyright presents a significant barrier to open data release, as training data often incorporates copyrighted material that companies cannot legally publish even if they wished to
Industrial enterprises treat their data as private assets; open source training code can allow enterprises to train models locally on private data, while synthetic/simulated datasets could be shared publicly as a compromise
When an audience member raised the issue of computing resources , both Jianmin Wang and Dutz acknowledged this as a distinct and important dimension of openness. Jianmin Wang characterised it as 'another dimension for open source - the open source for the computing power resource' . Dutz confirmed that the Apache Software Foundation 'doesn't own or run very much hardware on its own' and relies on contractors , making it currently unable to provide GPU resources for AI training, and suggested that only very large donations from big AI companies might change this .
Access to computing resources for AI training is an additional dimension of openness that the open source community must address
The conflict between open models and closed data is compounded by unequal access to computing resources; open source foundations such as ASF currently lack the funds or infrastructure to provide GPU resources for AI training
The conflict between open source and AI extends beyond open models versus closed data to include unequal access to computing resources, and open source foundations should consider providing GPU resources to contributors
Both Wei Wang and Jianmin Wang emphasised the need for standardisation, though from different angles. Wei Wang described his lab's effort to build a 'spectrum' or catalogue that classifies the actual level of openness of AI models , so users can understand what they can do with a given model . Jianmin Wang focused on data format standardisation, describing his team's contribution of a time series file format (TS file) to Hugging Face and calling for equivalent standards for physical data collections , framing this as an 'AI for good' initiative .
There is a need for a standardised spectrum or catalogue to measure the transparency and openness levels of AI models, so users understand what they can actually do with a given model
There is a lack of standardised data formats for physical/embodied AI data collection; providing open file formats (e.g., TS file for time series on Hugging Face) could help address this gap
Both Dutz and Wei Wang recognise that AI is fundamentally changing the nature of participation and contribution in open source communities, and that human judgement remains irreplaceable. Dutz described the 'tsunami' of AI-generated pull requests overwhelming maintainers , while Wei Wang argued that although AI has made it easy for anyone to write code, 'the most important ability is your judgement, your taste, and the quantity you can control' . Both see open source communities as environments where these human qualities are developed and tested, with Wei Wang explicitly stating that 'open source project and the community is the real course I think in the university can train our talent of the engineering' . Both Dutz and Jianmin Wang share a strong interest in industrial and time series data as an underexplored frontier for open source AI, and both reference IoTDB as a concrete example. Dutz described time series models as a 'world model light' and cited IoTDB as a ready-to-use tool for observing and optimising existing industrial systems . Jianmin Wang, as the builder of IoTDB, similarly argued that industrial environments generate rich time series and relational data that could be used to train domain-specific multi-modal foundation models . Both see this as a distinct and important track for open source AI development alongside general-purpose language models. Dutz, Jianmin Wang, and Yue Gao all share the view that software engineering is undergoing a fundamental paradigm shift driven by AI, moving towards natural language as a primary interface. Dutz described specification-driven development as 'developing on steroids' with AI , and demonstrated how AI enabled him to implement a driver in Rust — a language he does not know — simply by specifying intent . Jianmin Wang framed this as software engineering reaching 'version 3.0', combining high-level languages, foundation models, and natural language prompting . Yue Gao observed the evolution 'from the assembly language to the high level language made up to the natural language programming' , agreeing that 'the language or the tools or the skills for the future software may be changed' . Both Wei Wang and Dutz are concerned about the long-term sustainability and economic dynamics of open source AI ecosystems, though from complementary angles. Wei Wang argued that contributors to open source have historically not benefited commercially and that building a system to provide commercial feedback to contributors would make the ecosystem 'run even more better' . Dutz similarly noted that large AI companies are donating to open source foundations partly out of fear of 'recursive poisoning' of their training data , suggesting that economic incentives and sustainability concerns are increasingly intertwined in the open source AI ecosystem. Both Wei Wang and Dutz recognise that copyright and legal frameworks present significant challenges for open source AI, particularly as AI models become more capable of generating and reproducing content. Dutz highlighted that vision models raise pressing copyright questions about where to draw the line between a generated image and a copy , and that deep fakes create additional ethical and legal challenges . Wei Wang noted that 'the AI have changed a lot but only in the copyright and the law side' and argued that the lack of agreed standards for what constitutes openness in AI models compounds these legal ambiguities .
Given that the panel was framed around the positive potential of open source and AI for sustainable development , it is notable that all three panelists converged on a candid acknowledgement of the risks AI poses to open source communities. Dutz warned of 'a threat of burning out core assets of our major open source projects' and described the risk of 'recursive poisoning of AI datasets' . Wei Wang raised the concern that contributors are not adequately compensated . Jianmin Wang implicitly acknowledged these sustainability challenges by proposing synthetic datasets as a workaround for the closed data problem . This consensus on the risks and vulnerabilities of open source in the AI era was unexpected in a panel primarily focused on opportunities and positive applications.
It was unexpected that a panel representing major open source institutions - including the Apache Software Foundation - would so candidly acknowledge the structural resource disadvantage of open source communities relative to large AI companies. Dutz openly stated that the ASF 'doesn't own or run very much hardware on its own' and that building data centres full of GPUs would require 'two to three digit million' donations from big AI companies . Jianmin Wang framed this as a distinct dimension of openness that the community must address . This frank acknowledgement of resource asymmetry, rather than a more optimistic framing, represents an unexpected area of consensus.
There is an unexpected meta-level consensus that the open source community's response to the challenges posed by AI must itself be open source and AI-driven. Dutz described Apache Magpie - an AI-based tool for reviewing and sorting pull requests - as the community's own solution to the AI-generated contribution tsunami , essentially using AI to manage the problems created by AI. Wei Wang similarly argued that open source communities provide the real-world environment for developing the human judgement skills that AI cannot replicate , suggesting that the solution to AI's limitations lies within open source practice itself. This recursive, self-referential consensus was not an obvious outcome of the panel's framing.
The panel exhibited a high degree of consensus on several foundational issues: the critical importance of open source for AI transparency and democratisation ; the fundamental tension between open models and closed data ; the need for new standards and frameworks to define and measure openness in AI ; the role of open source communities as training grounds for human engineering talent ; and the emerging paradigm shift in software engineering towards natural language interfaces . There was also notable consensus on the risks AI poses to open source sustainability, including maintainer burnout , recursive data poisoning , and inadequate contributor compensation . The panelists agreed that computing resources represent a distinct and underaddressed barrier to open source AI participation , and that open source communities must develop their own AI-based tools to manage AI-driven challenges .
Wei Wang emphasised that open source AI is more complex than open source software, encompassing models, data, publication reports, and source code, and that companies often use 'open source' as a marketing story without genuine transparency . He proposed building a standardised spectrum to catalogue actual openness levels . Dutz approached the same problem from a technical angle, arguing that most models are merely 'open weights' and that the Apache principle of deterministic, reproducible builds from source code simply cannot apply to AI, because training the same model on the same data at different times yields different results . While both agree the current labelling is inadequate, Wang focuses on a user-facing transparency standard, whereas Dutz focuses on the fundamental technical incompatibility between AI and traditional open source reproducibility principles.
There is a need for a standardised spectrum or catalogue to measure the transparency and openness levels of AI models, so users understand what they can actually do with a given model
Most AI models are distributed as "open weights" rather than truly open source, and unlike traditional software, AI training is non-deterministic, making reproducible builds impossible
Dutz framed AI's impact on open source primarily as a 'tsunami' of AI-generated pull requests that threatens to burn out core maintainers of major projects , requiring survival strategies such as Apache Magpie . Wei Wang, by contrast, focused on the positive role of open source as a training ground for engineering talent in the AI age, arguing that judgement, taste, and quality control remain irreplaceable human skills . Jianmin Wang took a more evolutionary view, describing software engineering as having upgraded to a three-paradigm model combining high-level languages, foundation models, and natural language prompting , framing AI as a structural transformation rather than a threat. These perspectives reflect different professional vantage points: foundation governance (threat management), academia (talent development), and industry/research (paradigm evolution).
AI is creating a "tsunami" of pull requests in open source projects, placing enormous stress on maintainers of major projects, but open source communities are developing AI-based tools (e.g., Apache Magpie) to manage this load
Open source projects and communities serve as real-world training environments for engineering talent, providing practical industry-relevant experience that AI alone cannot replicate
Software engineering is evolving through three paradigms: high-level programming languages, foundation models as software components, and natural language prompting; future software will combine all three formats
Jianmin Wang approached the closed data problem from an industrial perspective, proposing that open source training code could be downloaded by enterprises to train models privately, and that synthetic or simulated datasets generated by domain foundation models could be shared publicly as a workaround . Dutz addressed the same tension but from a legal and copyright angle, arguing that even companies genuinely wishing to publish training data are often legally prevented from doing so because the data incorporates copyrighted material such as library books . Wang's solution is technical and structural (local training plus synthetic data sharing), while Dutz's analysis highlights a legal barrier that no technical solution can easily circumvent.
Industrial enterprises treat their data as private assets; open source training code can allow enterprises to train models locally on private data, while synthetic/simulated datasets could be shared publicly as a compromise
Copyright presents a significant barrier to open data release, as training data often incorporates copyrighted material that companies cannot legally publish even if they wished to
An audience member raised the question of whether foundations like the Apache Software Foundation could provide GPU computing resources to contributors . Dutz was direct in his scepticism, explaining that the ASF does not own or operate significant hardware, relies on contractors, and would need two-to-three digit million donations from large AI companies before such infrastructure could be contemplated . Jianmin Wang, however, framed computing resource access as a distinct and important new dimension of openness that the open source community must address , implying a broader obligation without dismissing the possibility. The audience member's question assumed foundations could or should take on this role, which Dutz effectively refuted on financial grounds.
The conflict between open models and closed data is compounded by unequal access to computing resources; open source foundations such as ASF currently lack the funds or infrastructure to provide GPU resources for AI training
Access to computing resources for AI training is an additional dimension of openness that the open source community must address
The conflict between open source and AI extends beyond open models versus closed data to include unequal access to computing resources, and open source foundations should consider providing GPU resources to contributors
It was unexpected that Dutz simultaneously expressed both optimism and alarm about AI's impact on open source. On one hand, he warned of a serious threat of burning out core maintainers and described recursive poisoning of training data . On the other hand, he argued that surviving open source will be stronger for having been reviewed by both human and artificial eyes . Wei Wang, by contrast, largely bypassed the threat narrative and focused on AI as a digital public good comparable to air and water , with open source as the best delivery mechanism . This divergence was unexpected given that both speakers are open source advocates: Dutz's practitioner experience led him to a more ambivalent, threat-aware position, while Wang's academic perspective led to a more idealistic framing. The tension between these two positions has significant implications for how the open source community should prioritise its responses.
It was unexpected that copyright emerged as a point of implicit disagreement about its nature and solvability. Dutz treated copyright as a hard legal constraint that prevents data openness even when companies are willing , framing it as an external barrier. Wei Wang, however, reframed the issue as one of contributor compensation and digital public goods , implicitly suggesting that new economic systems could realign incentives rather than simply accepting copyright as a fixed constraint. Jianmin Wang sidestepped copyright entirely, focusing on synthetic data generation as a technical workaround . The disagreement was unexpected because copyright was not a stated focus of the panel, yet it surfaced as a fundamental structural tension with different speakers implicitly proposing incompatible framings: legal constraint (Dutz), economic redesign (Wei Wang), and technical circumvention (Jianmin Wang).
The panel exhibited a broadly collaborative tone with moderate but substantive underlying disagreements. The main areas of disagreement centred on: (1) the definition and measurability of 'open source' as applied to AI models, with Dutz emphasising technical non-determinism and Wei Wang emphasising the need for a user-facing transparency spectrum ; (2) the primary nature of AI's impact on open source, ranging from existential threat (Dutz's tsunami metaphor ) to structural evolution (Jianmin Wang's three-paradigm model ) to public good opportunity (Wei Wang's air-and-water framing ); (3) how to resolve the open model/closed data tension, with Jianmin Wang proposing synthetic data sharing , Dutz highlighting copyright barriers , and Wei Wang calling for contributor compensation systems ; and (4) the feasibility of open source foundations providing computing infrastructure, where Dutz was sceptical and Jianmin Wang framed it as an open obligation .
All three panellists agreed that the current state of 'open source AI' is inadequate and that greater transparency is needed. Wei Wang called for a standardised spectrum to measure actual openness , Dutz distinguished between open weights and genuinely open source models , and Jianmin Wang acknowledged the conflict between open data and closed datasets from an industrial perspective . However, they diverged on the solution: Wang favoured a user-facing transparency catalogue, Dutz emphasised the technical impossibility of full reproducibility , and Jianmin Wang proposed synthetic data sharing as a practical compromise .
Open source AI as a tool for transparency and user empowerment, raising the question of who controls knowledge in the age of large language models Most AI models are distributed as "open weights" rather than truly open source, and unlike traditional software, AI training is non-deterministic, making reproducible builds impossible Industrial enterprises treat their data as private assets; open source training code can allow enterprises to train models locally on private data, while synthetic/simulated datasets could be shared publicly as a compromise
Both Dutz and Wei Wang agreed that the current economic and incentive structures of open source are under strain in the AI era. Dutz warned of recursive poisoning of AI datasets as AI-generated code re-enters training pipelines , while Wei Wang highlighted that contributors have historically not received commercial benefit from their open source work and argued for building systems that provide economic feedback to contributors . Both see the sustainability of open source ecosystems as at risk, but Dutz focuses on data quality degradation as the mechanism of harm, while Wang focuses on the lack of contributor compensation as the structural weakness.
AI-generated code feeding back into open source training data risks recursive quality degradation ("poisoning"), which is a concern driving donations from large AI companies to open source foundations Contributors to open source and AI systems are not adequately compensated commercially; building economic incentive mechanisms for contributors could make open source AI ecosystems more sustainable
Both Dutz and Jianmin Wang agreed that time series models represent an important and practical near-term opportunity for industrial AI, and that world/physics models are a longer-term aspiration. Dutz described time series models as a 'world model light' and cited IoTDB as a ready-to-use tool for observing and optimising existing systems . Jianmin Wang similarly identified time series data from machines as a key data type in industrial settings and noted opportunities to train large time series models for specific business domains . However, Dutz framed world models as currently infeasible due to data centre capacity constraints , while Jianmin Wang focused more on the opportunity to develop standardised data formats to enable future progress .
World models (physics-based models) represent the "holy grail" but require training data and computational capacity that currently exceeds available infrastructure; time series models offer a practical intermediate step for industrial optimisation Beyond language models, industrial and enterprise environments generate large volumes of relational and time series data, creating opportunities for domain-specific multi-modal foundation models
All three panellists agreed that AI is changing the skills and methods required for software development and open source participation. Wei Wang argued that judgement, taste, and quality control remain the most important skills even as AI makes coding accessible to all . Dutz described specification-driven development as a way AI enables contributors to work in languages they do not know . Jianmin Wang described a three-paradigm evolution of software engineering incorporating natural language prompting . All agree AI lowers barriers, but they differ in emphasis: Wang stresses irreplaceable human judgement, Dutz stresses the practical empowerment of non-expert contributors, and Jianmin Wang frames it as a structural evolution of the discipline.
Open source projects and communities serve as real-world training environments for engineering talent, providing practical industry-relevant experience that AI alone cannot replicate Specification-driven development, enabled by AI, lowers barriers to participation in open source by allowing contributors to define intent in natural language rather than requiring mastery of specific programming languages Software engineering is evolving through three paradigms: high-level programming languages, foundation models as software components, and natural language prompting; future software will combine all three formats
- Open source in the AI era is fundamentally a question of transparency and control: just as Richard Stallman asked 'who controls your computer?', the emergence of large language models raises the question of who controls knowledge and information access, making open source AI critical for user empowerment.
- Most AI models are distributed as 'open weights' rather than being truly open source; unlike traditional software, AI training is non-deterministic, meaning reproducible builds are impossible, which fundamentally challenges conventional open source principles of verifiable, identical outputs from the same source.
- AI is generating a 'tsunami' of pull requests in open source projects, placing severe stress on maintainers; the Apache Software Foundation is responding with AI-based tools such as Apache Magpie to help triage, review, and manage this influx, suggesting that AI must be used to manage the problems AI itself creates.
- Industrial enterprises treat their operational data as private assets, creating a structural tension between open source training code and closed proprietary datasets; this tension is a central challenge for the deployment of AI in enterprise and industrial settings.
- Copyright law presents a significant and often overlooked barrier to open data release, as training datasets frequently incorporate copyrighted material that organisations cannot legally publish even if they wished to do so.
- There is currently no agreed standard for measuring the openness of AI models; researchers are working to develop a transparency spectrum or catalogue so that users can understand precisely what they are permitted to do with any given model.
- Multi-modal AI presents distinct challenges and opportunities: language models are relatively mainstream, vision models raise acute legal and ethical issues around copyright and deep fakes, world models represent a theoretical ideal but exceed current infrastructure capacity, and time series models offer a practical intermediate solution for industrial optimisation.
- There is a significant gap in standardised data formats for physical and embodied AI data collection, which limits the development of world models and domain-specific foundation models.
- Access to computing resources for AI training is an additional and underappreciated dimension of openness; open source foundations such as the Apache Software Foundation currently lack the funds or infrastructure to provide GPU resources for AI training, creating a barrier to equitable participation.
- Software engineering is evolving through three paradigms: high-level programming languages, foundation models as software components, and natural language prompting; future software production will likely combine all three formats.
- AI-generated code feeding back into open source training datasets risks recursive quality degradation, a phenomenon sometimes described as 'poisoning', which is a concern driving donations from large AI companies to open source foundations.
- Open source projects and communities serve as irreplaceable real-world training environments for engineering talent, providing practical, industry-relevant experience and developing judgement and quality control skills that AI-generated code alone cannot replicate.
- AI should be treated as a digital public good, comparable to air and water, particularly in education, where free or low-cost access to AI tools is essential for equitable learning outcomes.
- Specification-driven development, enabled by AI, lowers barriers to participation in open source by allowing contributors to define intent in natural language rather than requiring mastery of specific programming languages, thereby broadening the contributor base.
- Open source code that survives the current AI-driven period of stress will ultimately be in a stronger position, having been reviewed by both human and artificial intelligence, making it more robust and trustworthy than proprietary alternatives.
- Contributors to open source and AI systems are not adequately compensated commercially; building economic incentive mechanisms for contributors could make open source AI ecosystems more sustainable and equitable.
“Richard Stallman asked 'who controls your computer?' in the closed-source software era. But today, with large language models, we must ask: 'who controls your mind and your knowledge?' Open source in the AI area is critical for transparency and giving users the right to understand how mechanisms in big data retrieve and process information.”
“The Apache Software Foundation is experiencing up to 170 pull requests per day on major projects, many suspected to be AI-generated. This is creating a 'tsunami' that risks burning out core maintainers. In response, ASF launched Apache Magpie, an AI-based tool to help review and triage pull requests. After surviving this tsunami, open source will be in a much better state because so many 'digital eyes' will have reviewed the code.”
“Most AI models are distributed as 'open weights' rather than truly open source. The Linux Foundation distinguishes multiple levels: open model, open weights, open data, and potentially a fourth level. At Apache, they never release binaries — they release source code, with the expectation that two people building from the same source get the same result. With AI, this deterministic reproducibility is impossible, even with the same training data and infrastructure, because results will never be identical.”
“When Google trained their models by feeding all the books in a library into training, this is analogous to a person reading all those books — which is fine. But if Google were required to release all the training data, they would be violating copyrights. So sometimes, even if companies genuinely want to publish training data, it is legally impossible to do so.”
“Open source AI is more complex than open source software. It must encompass at minimum the models, the data, the publication reports, and the source code. There is currently no agreed standard on what level of openness is required for each element. Many corporations claim their models are 'open source' — but this is largely a marketing story. Users need to know what they can actually do with a model. My lab is working to build a transparency spectrum to catalogue the real openness level of AI models.”
“In the age of AI, the digital public good is very important. AI can generate content very easily, but how do we return benefits to contributors? In the open source software era, many contributors gave their time without receiving commercial benefit. If we can construct a system where contributing to open source also yields commercial feedback, the system can run even better. The whole world is broken — maybe we can build a new one.”
“Short-term, AI is removing barriers — for example, helping contributors with strong domain knowledge but weak software engineering skills to participate more effectively. But if we scale this up, we will have a lot of AI-generated code going into open source, which will then go back into AI training. This is 'recursive poisoning' of AI datasets, which everybody is afraid of. The donations open source foundations are receiving may be because the companies behind big models are scared about starting to eat their own dog food.”
“Specification-driven development, where you write rules and architecture in plain text (e.g., Markdown), combined with AI, is like 'developing on steroids.' He described writing a specification for an industrial driver in Java and then asking Claude to implement it in Rust — a language he does not know — and receiving a perfectly functioning result. This means being able to formulate what you want without dealing with syntax, which will open many more doors in the future.”
Who controls your mind and knowledge in the age of large language models, and how can open source AI ensure transparency and user rights over AI-driven knowledge systems?
As AI increasingly mediates access to information and knowledge, this question raises fundamental issues about cognitive autonomy, transparency, and the governance of AI systems. It extends Richard Stallman's original question about software control into the AI era and is critical for ensuring democratic access to knowledge.
How can open source foundations like the Apache Software Foundation manage the 'tsunami' of AI-generated pull requests without burning out core maintainers of major projects?
The sustainability of open source communities depends on the wellbeing of their maintainers. With AI generating vast numbers of pull requests (up to 170 per day on some projects), understanding how to triage, review, and manage these contributions without exhausting human contributors is an urgent operational and governance challenge.
What are the different levels of openness in AI models (open model, open weights, open data, etc.), and how can a standardised spectrum or catalogue be developed to help users understand what they can actually do with a given model?
There is currently no agreed standard for what constitutes 'open source AI.' Both speakers highlighted that many companies market models as open source when they are only open weights. Developing a transparency spectrum would help users, researchers, and policymakers make informed decisions about model use, especially for commercial or research purposes.
How can enterprises and industries share or collaborate on training data while protecting proprietary data assets, particularly through synthetic or simulated data generation?
Industrial enterprises treat their data as private assets, yet AI model training benefits from large, diverse datasets. Exploring mechanisms such as synthetic data generation, federated learning, or shared simulation datasets could unlock significant value for domain-specific AI while respecting data ownership.
How should copyright law be reconciled with the need to publish AI training data, given that releasing training data may violate the copyrights of the original content creators?
This is a significant legal and ethical barrier to full openness in AI. Even companies that wish to publish training data may be legally prevented from doing so. Resolving this tension is essential for advancing open source AI and ensuring legal clarity for developers and organisations.
How can contributors to open source projects and AI training datasets receive fair commercial benefit from their contributions, particularly as AI systems increasingly monetise open source work?
Many open source contributors invest significant time without receiving commercial returns, and AI systems trained on their work may generate substantial revenue for corporations. Designing economic models that reward contributors fairly is important for the long-term sustainability and equity of the open source ecosystem.
What standardised data formats are needed for physical and embodied AI (world models), and how can the open source community develop shared dataset formats for physical data collection?
World models and embodied AI require large amounts of physical, spatial, and temporal data, yet there is currently a lack of standardised formats for such datasets. Developing open standards (analogous to the TS file format for time series data) would facilitate data sharing and accelerate progress in physical AI.
Will the recursive use of AI-generated code in open source projects, which then feeds back into AI training data, lead to a degradation of code quality over time ('recursive poisoning'), and how can this be mitigated?
As AI-generated code enters open source repositories and is subsequently used to train future AI models, there is a risk of compounding errors and declining quality. Understanding and addressing this feedback loop is critical for maintaining the integrity of both open source software and AI training datasets.
Can open source foundations such as the Apache Software Foundation provide or facilitate access to computing resources (e.g., GPUs) for AI training, and what funding models could support this?
Access to computing resources is a major barrier to participation in AI development, particularly for open source communities and researchers in lower-resource settings. Exploring whether foundations could pool resources or attract donations from AI companies to support community computing infrastructure is an important governance and sustainability question.
What is the long-term outlook for open source software after surviving the current wave of AI disruption, and how will the nature of software development and contribution change?
Understanding the post-disruption landscape for open source is essential for strategic planning by foundations, universities, and policymakers. This includes questions about new contribution models, the role of specification-driven development, and how open source communities will evolve as AI lowers barriers to participation.
How can AI be made freely accessible for public-interest domains such as education, so that students and teachers are not priced out of using AI tools?
If AI becomes essential digital infrastructure akin to water or air, then equitable access—particularly in education—becomes a matter of public good. Exploring open source and economic models that enable free or low-cost AI access for educational purposes is critical for reducing inequality and supporting sustainable development goals.
How will specification-driven development, enabled by AI, change the nature of open source contribution and lower barriers for domain experts who lack traditional software engineering skills?
Specification-driven development, where contributors describe desired behaviour in natural language and AI generates the corresponding code, could democratise participation in open source projects. Understanding the implications for code quality, architecture, and community governance is an important area for further research.
How should the evolution of software engineering be understood across its three generations—high-level languages, foundation models, and natural language prompting—and what does this mean for future software production and education?
Framing software engineering as having entered a third generation (prompt-based, natural language programming) has significant implications for curricula, industry practice, and the design of open source tools. Further research is needed to understand how these paradigms interact and what skills will be most valuable going forward.
How can open source communities and AI systems together address the distinct legal, ethical, and technical challenges posed by multimodal AI models, particularly vision models involving deep fakes and image copyright?
Vision models present unique challenges around copyright infringement, deep fakes, and identity misuse that go beyond those of language models. Developing open source tools and governance frameworks to detect misuse and enforce safeguards is an urgent area for research and policy development.
