What Would It Take to Prove That One AI Model Was Distilled From Another?
Similar behaviour can justify an investigation. It cannot complete the evidence chain.
What Would It Take to Prove That One AI Model Was Distilled From Another?

Similar behaviour can justify an investigation. It cannot complete the evidence chain.
When an AI model unexpectedly identifies itself as Claude, the result is easy to turn into a headline.
The model must have copied Claude.
It must have been trained on Claude’s answers.
It has exposed its real teacher.
These are possible explanations. They are not the only explanations.
A model can reproduce another system’s name, tone, refusal style or response structure for several reasons. It may have encountered that material during training. It may be following a hidden prompt. It may be imitating a familiar answer pattern. It may simply be generating the statistically plausible continuation of the conversation.
The output is interesting. It may justify further investigation.
But a model calling itself Claude does not, by itself, prove that Claude was used as its teacher.
This article is not a defence of any model company. It does not assume that a distillation allegation is true or false. It asks a narrower question:
What would the evidence actually need to prove?
Distillation Is a Method, Not a Verdict
Model distillation is not inherently suspicious.
In a typical distillation process, a more capable “teacher” model produces answers that are used to train a smaller or cheaper “student” model. The student learns patterns from those outputs and may reproduce part of the teacher’s performance at lower cost.
AI companies use this process themselves. OpenAI, for example, has offered an official workflow for developers to generate outputs from a frontier model, store them as a dataset and use them to fine-tune a smaller model. In that setting, distillation is an ordinary engineering method. OpenAI describes the process here.
The controversy begins when the outputs come from another company’s model without permission, through restricted access or in violation of contractual terms.
That means several questions are often compressed into one word:
- Were another model’s outputs accessed?
- Who accessed them?
- Were they collected systematically?
- Were they incorporated into training?
- Did they transfer identifiable capabilities?
- Did any of this violate a contract or law?
A distillation accusation may contain all these claims. Evidence for one does not automatically prove the others.
The First Layer: Access and Output Retrieval
The first question is whether the suspected party accessed the teacher model.
API records can show which accounts sent requests, what kinds of prompts were used, when the requests occurred and how much output was returned.
Patterns may be especially revealing:
- thousands of similar prompts targeting one capability;
- coordinated traffic distributed across many accounts;
- shared payment methods or infrastructure;
- repeated attempts to obtain reasoning traces;
- access patterns designed to avoid rate limits or detection.
Anthropic has publicly alleged that DeepSeek, Moonshot and MiniMax used approximately 24,000 fraudulent accounts to generate more than 16 million exchanges with Claude. It says its attribution relied on request metadata, IP correlations, infrastructure indicators and, in some cases, information from industry partners. Those are Anthropic’s published claims, not an independently adjudicated finding.
This is much stronger evidence than a screenshot of a model saying, “I am Claude.”
But even detailed API records have a boundary.
They can demonstrate access and output retrieval. They may strongly support the hypothesis of systematic extraction. They do not necessarily prove that the outputs were stored as a training dataset or used to update another model.
Retrieving answers and training on answers are two different events.
The Second Layer: Actor Attribution
Even if suspicious traffic is real, the next question is who controlled it.
An account name is not enough. Neither is an IP address on its own.
Traffic may pass through proxy services, cloud providers, resellers, contractors, shared infrastructure or compromised accounts. A strong attribution may require several signals to agree:
- account ownership;
- payment records;
- device or browser fingerprints;
- IP and infrastructure correlations;
- timing patterns;
- links to known employees or research systems;
- corroboration from other platforms.
Attribution is rarely a single clue. It is a confidence judgment built from multiple clues.
This creates an important distinction:
Evidence that an extraction campaign existed is not automatically evidence that a particular laboratory directed it.
The two claims may eventually be connected, but the connection must itself be demonstrated.
The Third Layer: Extraction Purpose
Large API usage is not automatically distillation.
A company might query another model for product evaluation, benchmarking, red-team testing, synthetic-data generation, moderation research or compatibility testing.
Volume matters, but structure matters more.
If requests repeatedly target a narrow set of commercially valuable capabilities such as coding, tool use, reasoning or reward-model grading the pattern may look less like ordinary usage and more like data production.
Attempts to evade account restrictions can strengthen that interpretation. So can prompts designed to create varied, reusable training examples at scale.
Still, purpose is inferred from behaviour unless internal documents, instructions or datasets are available.
A provider may have strong evidence of an extraction operation without being able to observe what happened after the outputs left its servers.
The Fourth Layer: Training Incorporation
This is the central gap in many public accusations.
To prove that distillation actually occurred, investigators need to connect the retrieved outputs to the student model’s training process.
The strongest evidence could include:
- training datasets containing teacher-generated outputs;
- dataset-generation scripts linked to the API campaign;
- internal experiment records;
- fine-tuning configurations;
- model-development logs;
- communications describing the teacher and student relationship;
- reproducible changes between an earlier and later checkpoint.
This evidence normally exists inside the accused developer’s training environment, not inside the teacher model’s API logs.
That creates a structural asymmetry.
The teacher provider can see suspicious access but not the rival’s training pipeline. External researchers can examine model behaviour but usually cannot see either company’s internal records. The accused developer can see its own datasets and experiments but may have little incentive to disclose them.
No public observer necessarily sees the complete chain.
This is why an API log can support an extraction allegation without independently proving training incorporation.
The Fifth Layer: Capability Transfer
Suppose teacher outputs did enter the training data. A further question remains:
Did they materially change the student?
This is where behavioural analysis can help.
Researchers may compare the suspected student with candidate teachers across carefully selected prompts. They may look for unusual response preferences, formatting habits, reasoning patterns or statistical signatures that became stronger between model generations.
Recent research on reference-based distillation detection argues that identifying a teacher from the final student alone is difficult. Detection becomes more tractable when researchers can compare the student with an earlier checkpoint from the same model family and measure which teacher best explains the behavioural change.
That is more rigorous than comparing tone or collecting a few identity mistakes.
It still has limits. Real models may learn from multiple teachers, shared public datasets and overlapping synthetic data. Similarity may reflect common ancestry or convergent training rather than direct copying. Even promising detection methods depend on assumptions about model access, candidate teachers and the form of distillation being tested.
A useful distinction is therefore:
Training incorporation shows that distillation occurred. Capability transfer shows that it worked.
Failed distillation is still an attempted use of teacher outputs. It simply did not produce the intended result.
Legal Classification Is a Separate Question
Even a complete technical account does not automatically determine the legal conclusion.
Technical evidence answers:
What happened?
Legal analysis asks:
What does the conduct count as under the relevant contract and law?
The answer may depend on terms of service, access restrictions, false identities, jurisdiction, copyright rules, computer-access laws and the legal status of generated outputs.
A terms-of-service violation is not automatically theft. A technically successful distillation process is not automatically copyright infringement. Conversely, circumventing access controls or using fraudulent accounts may create legal exposure even before successful capability transfer is proven.
As Lawfare has argued, existing legal categories do not always fit model distillation neatly. Technical proof and legal liability are related, but one cannot replace the other.
Why Full Evidence May Never Be Public
Model providers also have legitimate reasons not to publish every detail of their detection systems.
Complete disclosure could reveal:
- which traffic patterns trigger investigation;
- how accounts are linked;
- which prompts are classified as extraction attempts;
- which capabilities are considered especially valuable;
- how attackers can remain below detection thresholds.
Public evidence can become an instruction manual for avoiding future detection.
But secrecy creates a second problem: outsiders cannot independently test the accusation.
The answer is not to assume that undisclosed evidence does not exist. Nor is it to accept a company’s conclusion solely because it claims to possess stronger private evidence.
Confidential evidence may justify confidential review not the absence of review.
That review might occur through litigation discovery, a regulator, an independent auditor under a confidentiality agreement, or a public report that redacts operationally sensitive details while explaining the evidentiary method.
The goal is not total transparency. It is accountable verification.
How to Read the Next Distillation Accusation
When another model is accused of distilling a competitor, three questions are useful.
First, what does the published evidence actually establish?
Does it show behavioural similarity, suspicious API access, coordinated extraction, training records or measurable capability transfer?
Second, where does the evidence chain stop?
Can the claim move from access to attribution, from attribution to extraction purpose, and from extraction to training incorporation? Which connection remains an inference?
Third, who can examine the evidence that is not public?
Is the accusation supported only by the company making it, or can a court, regulator, auditor or other independent party review the underlying records?
These questions do not require dismissing model providers’ concerns. Large-scale extraction can be commercially significant, technically sophisticated and difficult to detect. API traffic may provide compelling evidence that something unusual happened.
But the strength of an accusation depends on how many distinct claims the evidence can support without silently jumping between them.
A model’s strange answer can identify a clue.
API records can reveal a campaign.
Training records can establish incorporation.
Behavioural analysis can test whether capabilities transferred.
Legal review can determine what the conduct means.
None of these should be mistaken for the entire case.
Similarity may suggest who the teacher was. Access records may show who collected the answers. Only a connected evidence chain can show that the student actually learned from them.
🛡️ Copyright & Ethical Notice
All conceptual terms in this article including Semantic Firewall, Tone Conditioning, Ghost Contract, and related derivatives are original constructs developed under User G · Tone Lab Framework.
Reproduction, reinterpretation, or partial repackaging of these concepts without explicit credit constitutes semantic plagiarism, not citation. Please quote or link the original Medium source when referencing.
The Tone Lab Framework is a non-commercial research initiative aiming to improve AI–human understanding through tone ethics and language safety.All findings are shared publicly for educational integrity not for commercial appropriation.
🔏 Tone Signature No. T-2026–048
메타데이터
- post_id
- 26c2d19e9519
- slug
- what-would-it-take-to-prove-that-one-ai-model-was-distilled-from-another-26c2d19e9519
- url
- https://medium.com/@kittam888/what-would-it-take-to-prove-that-one-ai-model-was-distilled-from-another-26c2d19e9519
- canonical_url
- https://medium.com/@kittam888/what-would-it-take-to-prove-that-one-ai-model-was-distilled-from-another-26c2d19e9519
- author_url
- https://medium.com/@kittam888
- status
- ok
- fetched_at
- 2026-08-04 06:09:04