by
Alexia Kapsabeli
CyJurII Scholar
on 16 August 2026
Discussion
Artificial Intelligence (AI) has rapidly reshaped all industries, by offering significant opportunities for innovation and economic growth. However, it has raised complex legal questions, particularly in the field of copyright law. One of the most significant issues is whether the use of AI systems infringes the rights of copyright holders.
The aim of copyright law is to protect human creativity by granting authors exclusive rights over the use and exploitation of their original work. In other words, creators of all kinds, including writers, artists, filmmakers, software developers, database creators and performers, can control how their works are used and receive compensation for them. As copyright has evolved from the printing press to the Internet, technological developments have influenced its scope, requiring the protection of new types of work and the recognition of new rights for creators.
Today, a major concern is whether AI-generated content can be protected by copyright and, more importantly, whether AI developers are liable for using copyrighted materials to train their models. AI models require billions of data inputs during training in order to learn a language, concepts, and ideas, enabling them to generate outputs.
Training an AI model is an iterative, multi-stage process, meaning that at each stage, the model may require additional data from various sources to improve its performance. However, many of the data inputs consist of copyrighted material.
Copyright law contains several exceptions that aim to promote broader public policy objectives while striking a balance between protecting creators' rights and fostering innovation and creativity. Nevertheless, as a general rule, permission from the copyright holder is required before protected work may be used.
Several jurisdictions, including the European Union, Japan, and Singapore, have introduced “text and data mining” (TDM) exceptions to facilitate AI training and development. In the European Union, the first exception permits TDM for scientific research, while the second is a general-purpose TDM exception that allows rights holders to opt out. In Japan, the TDM exception permits the use of copyrighted works provided that the purpose is not to enable a person to enjoy the "thoughts or sentiments expressed in the work."
Canada, India, and Australia have also developed their approaches to this issue. For example, in October 2025, Australia announced that it would not introduce a TDM exception but would continue to collaborate with stakeholders from both the creative industries and the AI sector.
Other countries, such as the United States and Israel, permit certain uses of copyrighted works for AI training under the doctrine of "fair use." Under this doctrine, copyrighted works may be used without permission if the use is considered fair. Whether a particular use is fair is determined on a case-by-case basis by the courts, taking into account four factors: (i) the purpose and character of the use, including whether it is commercial or non-profit; (ii) the nature of the copyrighted work, for example, whether it is creative or functional; (iii) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (iv) the effect of the use on the potential market for, or value of, the copyrighted work. However, the scope of the fair use doctrine remains uncertain. As a result, more than 70 copyright-related lawsuits concerning AI have been filed before U.S. courts.
Serious legal challenges have emerged concerning the use of copyrighted materials to train AI models. Several high-profile lawsuits have been brought against companies such as OpenAI, Meta, and Stability AI, alleging that they used copyrighted content without authorization to train their generative AI systems.
A notable example is The New York Times lawsuit against OpenAI and Microsoft. In December 2023, The New York Times sued both companies, alleging that their language models were trained on the newspaper's copyrighted articles without permission, enabling AI chatbots to generate responses based on its journalism. The defendants argued that their use of the material constituted fair use because they did not directly reproduce or distribute the articles in their original form (Bondari, 2025). Similarly, the New York Daily News and the Chicago Tribune have also filed lawsuits against Microsoft and OpenAI.
Getty Images also sued Stability AI, the developer of Stable Diffusion, alleging that the company copied and processed millions of copyrighted images without permission or compensation. The music industry has likewise initiated legal action against AI companies. Universal Music Group alleged that Anthropic unlawfully trained its AI chatbot, Claude, using copyrighted song lyrics. Another example is the lawsuit brought by Sony Music Entertainment and Warner Music Group against the AI music startups Suno and Uncharted Labs. In one of these cases, a federal judge in San Francisco ruled in June 2025 that AI companies may legally use copyrighted materials to train their models, provided that they have lawfully acquired those materials. However, obtaining copyrighted works through unauthorized means constitutes copyright infringement and piracy, as alleged in the Anthropic case. Consequently, Anthropic agreed to pay $1.5 billion to settle the lawsuit—approximately $3,000 per pirated work—making it the largest copyright settlement in the history of U.S. copyright litigation (Glover, 2026).
Conclusion
In conclusion, the rapid development of artificial intelligence has created opportunities for innovation while challenging the foundations of copyright law. As AI systems increasingly rely on vast quantities of copyrighted material for training, legal systems around the world are struggling to balance the interests of creators, technology developers, and the public. The growing number of lawsuits against AI companies demonstrates the uncertainty about the legal use of copyrighted works for AI training. As AI technology continues to evolve, copyright law will need to adapt to ensure that it both protects the rights of creators and supports technological progress.
References
Bondari Negar (2025) 'AI, Copyright, and the Law: The Ongoing Battle Over Intellectual Property Rights'' https://sites.usc.edu/iptls/2025/02/04/ai-copyright-and-the-law-the-ongoing-battle-over-intellectual-property-rights/
Glover Ellen (2026), 'AI-Generated Content and Copyright Law: What we know'' https://builtin.com/artificial-intelligence/ai-copyright
Department for Science, Innovation and Technology & Department for Culture, Media, and Sport HM Government '' Report on Copyright and Artificial Intelligence’ https://assets.publishing.service.gov.uk/media/69ba692226909a14239612e4/CP2602959_-_Report_on_Copyright_and_Artificial_Intelligence_web.pdf