As publishers, authors and musicians challenge AI companies, courts must decide where technological innovation ends and copyright infringement begins
As publishers, authors and musicians challenge AI companies, courts must decide where technological innovation ends and copyright infringement begins

Publishers, creators brace for fallout as courts clash over AI data mining

From Delhi to Munich and the US, conflicting rulings are reshaping the battle between AI companies and copyright holders

July was packed with cases pivoting on the legality of artificial intelligence (AI) companies helping themselves to data across the world to train their large language models. The lawsuits involved the critical issue of copyright protection, in particular for publishing companies and a range of creative fields such as literature, cinema and music, as AI behemoths such as OpenAI, Google, Anthropic and Meta freely mine data to train their machines. Is it copyright infringement, as courts in Europe have held, or fair use, as US courts have ruled?

The Delhi High Court also weighed in on this critical question in a case filed by news agency ANI (Asian News International) against OpenAI and said in a preliminary ruling that the AI entity’s use of copyright material did not constitute infringement but was fair use, or “fair dealing” under India’s copyright law.

The copyright law and its interpretation here and in the US and the UK in relation to AI are fluid, unlike in the EU, which has a clearly defined AI Act brought in 2024. This codifies new requirements for AI model training under the copyright law and stipulates that reproductions made for AI training must be authorised by rights holders, unless covered by existing copyright exceptions such as text and data mining (TDM) provisions.

Thus, a regional court in Munich had no difficulty in ruling against OpenAI in a case filed by German music copyright organisation GEMA. It stated that the use of copyrighted song lyrics for training generative AI models without a licence violated German copyright law. The court rejected OpenAI’s untenable argument that users, and not the platform, were responsible for any infringement.

Towards the end of July, a US federal judge gave final approval to a record US $1.5 billion class action settlement in a suit filed in 2024 by a large group of authors against Anthropic over the misuse of their books to train its AI chatbot, Claude. They accused Anthropic of retaining and storing their works, totalling nearly half a million pirated books, without authorisation and sued for damages. The court said that AI model training could copy the books without being liable for copyright infringement, but storing pirated, shadow-library text exposed the company to legal challenges.

The confrontation of rights-holders with AI companies over intellectual property rights, which started in 2024, is intensifying. In May, a group of publishers joined forces to take on Meta and its head Mark Zuckerberg over the unauthorised use of their copyrighted books and academic materials to train Llama, the company’s LLM. Meta has rejected the allegations, maintaining that AI training qualifies as fair use under US copyright law.

Refusing ANI's plea for an interim injunction, Justice Amit Bansal said that using copyrighted material to train LLMs falls within the fair dealing exemption for research under the Copyright Act, 1957.

ANI had filed a lawsuit against OpenAI in the Delhi High Court in November 2024, alleging that the US-based company used its published content, some of it behind paywalls, without permission to train its AI models and that ChatGPT attributed fabricated stories to the news agency.

A number of stakeholders, including the Digital News Publishers Association, the Indian Music Industry, the Federation of Indian Publishers, and technology-based start-ups have filed interventions in this suit, turning the case into a major policy question on what can be termed fair use by the AI industry. The outcome will have implications for every creative field from cinema to music, for every publisher and for the future of media since AI systems train on their output without consent or compensation, and then build a commercial product to compete against them.

Views expressed are personal

Responsive Banner
Fact Net
www.fact.net.in