Copyright and AI: a debate in urgent need of reorientation

Copyright and AI: a debate in urgent need of reorientation

After the failure of the Darcos proposal, remaining on the presumption of use of AI training would be illusory, legal actions have already shifted the stakes towards AI outputs.

The rejection of the Darcos proposal in the National Assembly is both a misunderstanding and a revelation. Presented as a last attempt to get the AI ​​giants to negotiate before a tax on their turnover, it was perceived by others as an open door to cascading lawsuits, confusing presumption of use and proven infringement.

The failure of this text in the National Assembly must act as an electric shock. For months, the French debate on copyright and artificial intelligence has been bogged down in a single question: how to prove that a work was used to train a model, in order to loosen the evidentiary vice that stifles authors, publishers, producers and press companies.

However, by making this presumption the cornerstone of all regulation, we have ended up confusing a useful tool with a miracle solution. However, the ground has already changed: the heart of the disputes between rights holders and AI companies no longer lies only in the absorption of works for training, but in their capacity to produce substitutes, to restore protected markers, to capture informational value… and to make the evidence itself elusive.

First move: the release of AI on the market

The first shift is clear: the real issue is no longer the entry of works into the machine, but their exit onto the market. In the United States, major cases against AI players now focus on the systems’ ability to generate content that competes with, replaces, or summarizes the original works and services. In the dispute between several plaintiffs and OpenAI, judges ordered the production of millions of conversation logs, in order to determine whether ChatGPT’s responses constitute a substitute market. Likewise, in the case of The New York Times against OpenAI and Microsoft, the daily newspaper refocused its strategy on direct grievances: the output of the models, their technical architecture and proof of economic damage. By focusing on the presumption of entrainment, we would have risked missing the essential point: the proof that the models generate content substituting the original works, that is to say a replacement market.

Second move: traces in the results

Second major development: the question is no longer “Was this work used for training?”, but “What usable trace does it leave in the results?”. The litigation Getty Images against Stability AI, in the United Kingdom, demonstrated this vividly: the theory according to which the model was an infringing copy has largely failed. On the other hand, the judge found infringements of trademark law when Getty watermarks reappeared in certain generated images. Thus, a rights holder can lose on the metaphysical question (is the model a copy?) but win on the concrete ground: the exploitable traces in the results. In music, the case of Concord Music Group against Anthropic follows a similar logic: it is the lyrics generated, the knowledge of the contentious releases and the deletion of metadata which now structure the files.

Third trip: the press and neighboring rights

Third shift: the press and rights holders shift litigation towards real-time retrieval, automated summaries and conversational interfaces. In Europe, the Like Company case against Google, pending before the Court of Justice of the European Union (CJEU), is not limited to the question of reproduction during training: it also questions communication to the public, the exception of text and data mining, as well as the attribution of productions to the system supplier. For publishers, the real danger lies not in using their articles for practice, but in retrieving, rephrasing, and monetizing them, a practice that ultimately deprives readers of the original source. Here again, a legal approach centered on the presumption of initial use would only have dealt with part of the problem.

Fourth move: data governance

Finally, the fourth shift: the contentious target is no longer only remuneration, but also the withdrawal of data, transparency and the questioning of supply chains. The action initiated in France against Meta by organizations of authors and publishers shows the way: beyond financial compensation, the plaintiffs demand the complete removal of data directories created without authorization. From now on, the debates no longer focus only on proof of the use of works, but on the governance of data itself: which sources are used? Which corpora are used? What rights are reserved? How to exercise your right to object? How to purge illegal data?

Conclusion: see broader

The political lesson of this parliamentary blockage is paradoxical. On the one hand, the Assembly deprived creators and the media of a useful procedural lever. On the other hand, this failure forces us to broaden our vision. The presumption of use is no longer enough. We must now organize proof of the results produced by the models, guarantee the transparency of corpora and supply chains, allow the purging of illegally collected data and sanction the disappearance of metadata and watermarks, these essential traces of the provenance of works. One year before the electoral deadlines, this implementation becomes urgent for the news press, a pillar of our democracies as recently recalled by the Court of Justice of the European Union.

Philippe Schmitt is a lawyer in Paris

Jake Thompson
Jake Thompson
Growing up in Seattle, I've always been intrigued by the ever-evolving digital landscape and its impacts on our world. With a background in computer science and business from MIT, I've spent the last decade working with tech companies and writing about technological advancements. I'm passionate about uncovering how innovation and digitalization are reshaping industries, and I feel privileged to share these insights through MeshedSociety.com.

Leave a Comment