Can a competitor extract data from your AI model to train its own?

In late 2025, Anthropic, the company behind the chatbot Claude, accused, in a letter to the U.S. Senate The Chinese company Alibaba reportedly conducted some 28.8 million conversations with Claude via 25,000 fake accounts over a six-week period. Not to use Claude, but to save the responses as training data for their own, cheaper model. This technique is called model distillation, and the Dutch blog The Digital Gavel recently explained clearly how it works from a technical standpoint. We address the legal question that arises in this context here, based on Belgian law. The short answer: in Belgium, the responses generated by an AI model are, in principle, not subject to copyright and generally not subject to database rights either, so the provider must primarily rely on its terms of use, the doctrine of third-party liability, and fair market practices. Furthermore, anyone who uses fake accounts risks facing criminal charges.

What Model Distillation Is and Why Every AI Lab Does It Themselves

Model distillation involves a teacher and a student. The teacher model is large, expensive, and capable: the development costs for such a model run into the hundreds of millions of dollars. The student model is small and inexpensive. Instead of training the student model on the entire internet itself, the developer asks the teacher model millions of questions, stores each answer, and has the student model learn those question-answer pairs. After millions of repetitions, the student model begins to resemble the teacher surprisingly well.

This technique is quite common. The major AI labs distill their own state-of-the-art models into low-cost “mini” versions that run on a laptop or phone. Moreover, since the emergence of reasoning models—which spell out their thought processes in full in the response text—distillation using standard output has become dramatically cheaper. In early 2025, a team from Stanford and the University of Washington trained a model using barely a thousand sampled questions, TechCrunch reported, a competitive reasoning model requiring less than fifty dollars' worth of computing power; the researchers described their methodology in the s1-paper.

The distinction between product development and a problem therefore lies not in the technology, but in authorization. Anyone who creates their own model is launching a product. Anyone who extracts data from a competitor’s model via its public interface is generally violating the terms of use. The question is what the provider can do about it under Belgian law.

Why Copyright Law Doesn't Help the AI Provider

The first instinct of an aggrieved company is to invoke copyright. That instinct runs aground here. Copyright protection requires a work that is original: an intellectual creation of one’s own that is expressed through the author’s free and creative choices. In established case law, the Court of Cassation defines this as a discernible human creative activity that gives the work a personal stamp. A text generated by a language model without human creative intervention is a mechanical result and therefore not a work.

That’s not a new insight. As we analyzed earlier in our article on the question of whether a painting created by artificial intelligence is protected by copyright, and later about images generated by generative AI, the effectiveness of the protection depends entirely on human input. With 28.8 million automatically generated responses, there is no human input for each response.

For the provider of the teacher model, this presents an awkward situation. His crown jewels—the high-quality answers—are precisely the part of his product that is not protected by exclusive rights. The training data, the model weights, and the surrounding software may indeed be protected, but the distiller does not copy those. He only copies what the model says. You can find more background on this tension on our overview page about Artificial Intelligence and Copyright.

Does database law offer a solution, then?

A second avenue is database law. The Code of Economic Law (CEL) The producer of a database is granted a sui generis right: he may prohibit the extraction and reuse of the entire contents of his database or a substantial part thereof (Art. XI.307 WER). This right protects investment rather than creativity, and it is interpreted broadly. The Court of Justice ruled in the judgment Innoweb (ECJ, December 19, 2013, C-202/12) even held that a search engine that makes a protected database searchable in real time reuses its entire content.

However, there are three issues with applying this to an AI model.

First, the definition. A database is a collection of independent elements that are systematically or methodically arranged and individually accessible (Art. I.13, 6° WER, transposition of the Database Directive). A trained language model is not an ordered collection of retrievable elements: the knowledge is embedded in billions of parameters, and no single element can be retrieved separately. Furthermore, the answers only come into existence at the moment they are generated. Therefore, before the distillation operation begins, there is no existing database from which to retrieve information.

Second, the investment. Protection requires a substantial investment in the acquisition, verification, or presentation of the content, and the Court of Justice expressly excludes investment in the creation of the data from this requirement (see, among others, the judgment British Horseracing Board, ECJ, November 9, 2004, C-203/02). The hundreds of millions invested in the model are costs incurred to be able to create answers. It is precisely that investment in creation that is not taken into account.

Third, and this is often overlooked: the sui generis right is granted only to producers with a connection to the European Union (Art. XI.315 WER). An American provider without a European business presence that meets the conditions cannot invoke this right in Belgium under any circumstances.

If a collection of data falls outside the scope of both copyright and the sui generis right, the Court of Justice ruled in its judgment Ryanair (ECJ, January 15, 2015, C-30/14) that the operator is free to regulate its use by contract. Protection then shifts entirely to the terms of use. More on that in a moment. Anyone who wishes to read about the regime in detail can find the building blocks on our page about database law.

Are the responses from an AI model a trade secret?

The third line of inquiry involves trade secrets. Since the implementation of the Trade Secrets Directive Title 8/1 of Book XI of the WER protects information that is secret, has commercial value because it is secret, and is subject to reasonable confidentiality measures (Art. I.17/1, 1° WER).

This qualification generally applies to the model weights, the training methods, and the model’s internal probability distributions. No provider shares this information through its public channels. But the aggregator does not withhold that internal information. It collects the answers, and those answers are shared with every paying customer. Information that is, by definition, shared with third parties loses its confidential nature.

The question remains whether querying the model on a massive scale constitutes a prohibited form of reverse engineering. Here, too, the law does not automatically protect the provider: observing, researching, and testing a product that has been made available to the public is a lawful means of acquisition (Art. XI.332/4 in conjunction with Art. XI.332/3, § 1, 2° WER), unless a valid contractual restriction provides otherwise. And so, once again, everything comes down to the terms of use. An overview of this protection regime can be found on our page about trade secrets and know-how.

The Terms of Use as the actual line of defense

Every major AI provider prohibits, in its terms of service, the use of its output to train competing models. Anyone who creates an account accepts those terms. The distiller who flouts this prohibition is therefore in breach of contract, with all the consequences that entails: damages, termination of the contract, and account closure.

In practice, the distiller is rarely at the table himself. According to reports, the operation against Claude was carried out through thousands of seemingly normal accounts—each generating an average of 25 questions per day—and through intermediaries in China who resold access at discounts of up to 90 percent, kept track of the conversation logs, and sold those logs as training data.

The Civil Code (CC) codified the doctrine of third-party complicity in 2023: a third party commits a tort when he participates in a party’s failure to fulfill its contractual obligations, while he knew or ought to have known of those obligations (Art. 5.111 of the Civil Code). A company that purchases training logs from a reseller—even though every professional in the industry knows that the provider’s terms of use prohibit such reuse—therefore directly exposes itself to a claim by the provider itself. Although the contract is binding only on the account holder, the third party who knowingly profits from this breach may be held liable on a non-contractual basis.

In addition, the law governing market practices provides an independent basis for action against a competing distiller. Article VI.104 of the WER prohibits any act contrary to fair market practices by which one enterprise harms or may harm the business interests of another enterprise; potential harm is sufficient. Systematically skimming off another party’s product in order to replicate it at a fraction of the cost is a classic example of parasitic free-riding. The supplier may file an injunction against this with the presiding judge of the Commercial Court, who sits as in summary proceedings (Art. XVII.1 WER): a swift procedure on the merits, without a requirement of urgency. You can read more about this concept on our page about unfair market practices.

Where Fake Accounts Enter the Criminal Justice System

The distillation process itself is one thing. The way it is organized is another. When Anthropic introduced an identity verification process requiring a live selfie during account creation, reports indicate that within a few weeks, a black market emerged for AI-generated IDs and deepfake cameras, with accounts being traded for about thirty dollars each.

Under Belgian law, the classification then shifts from breach of contract to a criminal offense. The Criminal Code (Sw.) punishes anyone who gains access to a computer system or remains in it while knowing that they are not authorized to do so, and anyone who, with fraudulent intent, exceeds their access privileges (Art. 550bis, § 1 and § 2 Sw.; as of September 1, 2026, external and internal hacking under Articles 524 and 525 of the new Penal Code). Anyone who gains access using a false identity document after the provider has refused or closed their account knows that they are not authorized. Creating accounts with AI-generated identity documents may also constitute computer fraud: entering data into a computer system in a way that alters its legal significance (Art. 210bis of the Penal Code; effective September 1, 2026, included in the technology-neutral provision on forgery in documents or on other durable media under Article 451 of the new Penal Code). Our page on hacking examines these computer crimes in greater detail.

The Uncomfortable Mirror: AI Labs Train Themselves Using Others' Work

There is an ironic side to the outrage expressed by AI providers. The very same companies that are now speaking of theft trained their teacher models using the work of millions of authors, photographers, and journalists—often without permission. In 2025, Anthropic reached a $1.5 billion settlement with American authors over this issue, who was finally approved on July 20, 2026 by the federal court in California.

Under Belgian and European law, however, such training is not automatically unlawful. The transposition of the DSM Directive introduced an exception for text and data mining: the reproduction of lawfully accessible works for the purpose of text and data mining is permitted, unless the author has expressly reserved his or her rights in an appropriate manner, in the case of online content via machine-readable means (Art. XI.190, 20° WER). For research organizations, there is a separate, non-excludable exception (Art. XI.191/1, § 1, 7° WER). The AI Regulation Furthermore, it requires providers of general-purpose AI models to have a policy in place to identify and respect those opt-outs (Art. 53(1)(c) of the AI Regulation).

The legal symmetry is therefore striking. Training on human-generated data is permitted unless the rights holder objects in a machine-readable format. Training using model output is not prohibited by an exclusive right, but it is contractually excluded. In both cases, the deciding factor is not property rights, but rather the question of who has made which reservation and who has accepted which clause.

Specifically, what does this mean?

For providers of AI models. Don’t rely on copyright or database rights to protect your output; build your defense around the terms of use. Ensure that the prohibition on data extraction in the terms of use is worded unambiguously—including for resellers and API access—and that acceptance of these terms is verifiable on a per-account basis. Document the detection: the pattern of thousands of accounts with similar queries is your evidence of breach of contract, third-party complicity, and claims regarding unfair market practices. Treat the model weights and internal parameters as trade secrets, with contractual and technical confidentiality measures that you can demonstrate.

For companies that want to train a model or purchase derived training data. Distill your own model or work with the provider’s express permission; some providers offer distillation as a paid service. If you purchase datasets from third parties, conduct a provenance check: anyone who acquires training logs that they should know were collected in violation of another party’s terms of use risks a direct liability claim under Article 5.111 of the Dutch Civil Code and a cease-and-desist order for unfair market practices. If you have any doubts about the terms of such a dataset deal, you would be well advised to seek assistance in advance from a lawyer specializing in intellectual property.

Frequently asked questions (FAQ)

Is the text generated by a chatbot protected by copyright?
In principle, no. Copyright requires human creative activity with a personal touch. An answer generated autonomously by a language model is a mechanical result and therefore not a work. Only when a person subsequently edits the output using their own free and creative choices can copyright arise in that edited work.

Can I use the responses from ChatGPT, Claude, or Gemini to train my own AI model?
The terms of service of the major providers expressly prohibit this. If you do so using your own account, you are in breach of contract. If you purchase the data from an intermediary, you risk being held liable as a third-party accomplice to that breach of contract (Art. 5.111 of the Dutch Civil Code) and being subject to an injunction for unfair market practices.

Can an AI provider in Belgium bring criminal charges against someone for model distillation?
Not for the distillation itself, but for the context. Anyone who gains access to the platform using false identity information or after being denied an account may be guilty of unauthorized access to a computer system (Art. 550bis of the Penal Code; effective September 1, 2026, Articles 524–525 of the new Penal Code) and computer fraud (Art. 210bis of the Penal Code; effective September 1, 2026, Article 451 of the new Penal Code).

Conclusion

Model distillation exposes a gap in intellectual property law: the most valuable component of an AI product—the answers themselves—is protected in Belgium neither by copyright nor, generally speaking, by database law. The AI provider’s actual protection lies in its terms of use, reinforced by the third-party liability under Article 5.111 of the Civil Code, injunctive relief for unfair market practices, and—in the case of fake accounts—criminal law. Anyone in Belgium who wishes to train a model or purchase derived training data should therefore first review the contractual chain, not just intellectual property law.


Joris Deene

Mr. Joris Deene is a partner at Everest Attorneys and heads the department of intellectual property, IT law, AI law, data protection, and media law. ICT Legal Guide is that department’s knowledge platform. Joris publishes and teaches on copyright law, trademark law, software law, the GDPR, the AI Act, the DSA, and media law.

Contact

Questions? Need advice?
Contact Attorney Joris Deene.

Phone: 09/280.20.68
E-mail: joris.deene@everest-law.be

Topics