← Back to list

Can and should we protect synthetic data?

Written by Dr. Peter R. Slowinski, J.S.M. (Stanford)

Free flow · 2026-05-28 14:10 · 8 claps · 4.3 min read
#ai #data #synthetic-data #business #law
Open on Medium ↗
Wiki topics: AI · AI · General ⚖️ · Law & Justice

Can and should we protect synthetic data?

Written by Dr. Peter R. Slowinski, J.S.M. (Stanford)

Conference poster by Peter R. Slowinski

Conference poster by Peter R. Slowinski

Have you ever read that synthetic data is detrimental to the quality of AI models that are trained on it? Sometimes, synthetic data, that is, data created by AI systems, is used synonymously with AI slop, that is, AI-generated garbage floating around the internet. But synthetic data is much more than this, and certain synthetic data can be of great value. This value poses the question how investments in synthetic data can be secured through legal means and whether intellectual property law plays a role in this.

What is synthetic data?

The term “synthetic data” is quite broad and describes data originating not from real human interaction with machines or from Internet of Things (IOT) sensors but instead data created by artificial intelligence systems in order to mimic real-life data. This covers of course also things like fake pictures uploaded to social media or music created using AI. The value of such data for training yet another AI is indeed questionable and may lower the quality of AI models trained on such data. Maybe it is a little bit like a human raised entirely on highly processed food. Studies show us that the effects are not entirely positive. But there are circumstances where synthetic data is a good substitute or supplement for real data. It depends on the use case.

Synthetic data as a supplement for real data

Imagine that you have a factory that produces parts for car engines using CNC milling machines. The parts must meet very high demands to perform in the required way. And they need to be checked for production defects. In the past, a worker in the factory would take a manufactured part every so often, look at it from all sides, measure it, look for defects and if such defects were found, sort these parts out and recalibrate the machine. This is time consuming. AI-supported machine vision technologies could be used to check every single part for defects, thus reducing defective parts. There is, however, one limitation. A human worker is capable of abstract thinking and will only need to see a few defects in order to be able to recognize defects that he or she has not seen before. AI is not capable of abstract thinking and requires large sets of training data to “learn” what a defect is and what is perfectly acceptable. Of course, one could use all the pieces that have been sorted out because of defects, photograph them, label the data in the pictures and use this to train the machine. This may turn out to be more costly and time consuming that staying with human quality control. Or, one can use synthetic datasets. A generative AI system, for example a generative adversarial network (GAN) can learn from a small sample of pictures what kinds of defects exist and synthetically create a much larger dataset of such pictures with various kinds of defects. These pictures can then be used to train the machine vision system to detect and flag defective parts. Therefore, synthetic data can supplement real data where collection and labelling of real data is too costly (or in some cases, like the detection of rare medical conditions, impossible).

Autonomous driving

Let’s look at another example: autonomous driving. Here again, costs and the incapacity of AI systems for abstract thinking are limiting factors. By now hundreds and thousands of hours of video material have been collected on our streets to train autonomous cars — which are in fact AI on wheels — to know what driving around is about. But this is still not enough material; some situations simply cannot be filmed in real life either because they occur rarely and would need to be staged or they are simply too risky for humans. But such situations can be created artificially as training material for AI. In a way, you can see it as one AI teaching another one how to drive.

Anonymization

Finally, synthetization techniques can be used to create or enhance anonymity. Medical data from patients are one such use case. In many cases the identity of individual patients is not relevant for training an AI model to recognize patterns that may help in the development of new treatment methods or the detection of anomalies. But removing personal information can be time consuming and, if done manually, imperfect. Using synthetization methods, on the other hand, it is possible not just to remove information but to overwrite it with information that is false but does not affect the data quality. This can enhance anonymization and thus trust.

Protecting the value of the synthetization process

Good synthetization methods are, therefore, of significant value and even if they are cheaper than collection and labelling of real life data, they remain costly. The legal question is how the investments in the synthetization methods and the creation of synthetic data can be secured through legal means and whether the current legal framework is up to the task. Here intellectual property protection could play a role, but whether IP rights actually cover such synthetic data is not clear. The synthetization process itself may be patent-protected — provided that it produces a technical effect and is new and inventive. The synthetic data will probably not be protected by copyright since it is not a human’s personal creation but the output of an AI system. In the EU we also have the sui generis database right. But this was created long before anyone thought about creating synthetic data, and the criteria do not really fit. This leaves the creators of synthetic data with trade secret protection. But how practical is this in cases where a company that is excellent at creating synthetic data wants to sell or license the data to third parties? After all, trade secret protection is not as absolute as patents or copyright.

Do companies require more protection?

In order to establish whether the current protection level is sufficient for the establishment of sustainable business models, I am conducting an empirical analysis of companies’ business models and protection methods. In the end, we will hopefully know if there are sufficient incentives to create synthetic data gold instead of AI slop.

The project “Legal protection of synthetic data for artificial intelligence applications” is funded by the Polish National Research Center (NCN)” OPUS 27 — Project Number: 2024/53/B/HS5/01018.


메타데이터
post_id
f0bb15bb875e
slug
can-and-should-we-protect-synthetic-data-f0bb15bb875e
url
https://medium.com/@kwasn/can-and-should-we-protect-synthetic-data-f0bb15bb875e
canonical_url
https://medium.com/@kwasn/can-and-should-we-protect-synthetic-data-f0bb15bb875e
author_url
https://medium.com/@kwasn
status
ok
fetched_at
2026-06-09 15:37:30