← Back to list

The Likeness Screen You’re Running Doesn’t Work

Six words. Zero reference images. Three recognizable likenesses. Welcome to training data liability.

Joseph Desmond Cruel · 2026-05-01 03:37 · 46 claps · 10.0 min read paywalled
#future-of-work #ai #artificial-intelligence #filmmaking #creativity
Open on Medium ↗
Wiki topics: AI · AI · General 🎬 · Film & Television 🥊 · Combat Sports 🏃 · Running & Endurance

The Likeness Screen You’re Running Doesn’t Work

Six words. Zero reference images. Three recognizable likenesses. Welcome to training data liability.

Note yet a Medium Member? Read this article for free here.

Reading through my LinkedIn feed this week, I came upon a story about a visual effects professional who recently documented a concerning pattern. He works in production environments where likeness rights matter. He understands the stakes. He ran a controlled test using Seedance 2.0, a commercial AI image generation platform marketed to creative professionals.

The test was simple. Four generation attempts. Each used the same six-word prompt: “headshot of a blonde haired woman.” No reference image was provided. No actor name was mentioned. No celebrity identifier was included in the text.

The first generation returned a face that matched a recognizable public figure. He ran it through Google Reverse Image Search. The tool confirmed the match.

The second generation returned another recognizable face. He ran the same verification process. Google returned no match. The face was recognizable to him personally, but the only verification tool available to most production teams failed to identify it.

The third generation returned a third recognizable likeness. Again, no verification match from Google.

The fourth generation was blocked by the platform itself. Seedance’s internal content filter flagged the output before delivery. The platform’s own safety system determined the result violated policy.

Three of four got through undetected by the platform’s filter. Two were confirmed by external verification. One was recognizable but unverifiable. All four came from a prompt with zero likeness instruction.

What this demonstrates goes beyond the technical capabilities of the model. It surfaces a structural gap between what AI generation tools can produce and what production teams can defend.

The Contract Nobody Reads

Most Terms of Service agreements for AI generation platforms include a specific clause. The language varies by vendor, but the legal architecture is consistent. The platform states that the model may reproduce copyrighted or protected material during generation. And when it does, the liability for that reproduction transfers to the user.

The reproduction happens inside the vendor’s system. The legal exposure lands in your production pipeline.

This clause exists because the vendors understand what their models learned during training. They know what data went in. They know what patterns the model can reproduce. And they structured the contract to ensure that knowledge remains proprietary while the consequences of outputs become the user’s responsibility.

Reading this clause carefully before onboarding a tool is not paranoia. It is due diligence. The question is not whether the clause is reasonable. The question is whether your production team’s workflow accounts for what the clause actually says.

Three Paradigms, Three Liability Positions

Current AI generation models in production use fall into three categories. Each category carries a different legal position. Understanding which category your tools fall into is the foundation of defensible AI workflows.

Current AI generation models in production use fall into three categories. Each category carries a different legal position.

Current AI generation models in production use fall into three categories. Each category carries a different legal position.

Proprietary with Indemnification

Adobe Firefly represents this category most clearly. The model is closed. Adobe controls and audits the training data. And Adobe provides an explicit vendor warranty: if a likeness or copyright claim arises from Firefly outputs, Adobe has a legal position to defend the user.

This is what indemnification looks like in practice. Adobe reviewed its training corpus. Adobe made a legal determination about what was inside it. Adobe signed a contract that says if someone sues you for what the model outputs, Adobe participates in the defense. That contract has limits. It has exclusions. It has coverage caps. But it exists. The boundary is documented. You can read it before you deploy the tool.

The trade-off is creative control. Firefly outputs tend to be more conservative. The model refuses prompts other platforms allow. The stylistic range is narrower. Adobe locked down the training data to make the warranty possible. You cannot have indemnification and unrestricted creative freedom in the same package. Vendors who promise both should be questioned closely.

For production workflows where client-facing deliverables carry high exposure — theatrical trailers, marketing materials, anything with broad distribution — this trade-off makes sense. You exchange some creative latitude for documented legal coverage.

Proprietary Without Indemnification

Runway, Seedance, Kling, and similar commercial platforms fall here. The models are closed. The training data is undisclosed. And the vendor provides no warranty on outputs.

You cannot audit the training data because you cannot access it. You cannot verify what the model learned. The model may have been trained on scraped likeness data. The Terms of Service does not guarantee it wasn’t. The vendor has legal and technical information you don’t. And the contract says the liability is yours.

This is where most production pipelines currently operate. The models perform well. The outputs meet quality standards. The vendors provide API access or web interfaces. And the vendors provide nothing else. No training manifest. No provenance documentation. No warranty.

The Seedance test demonstrates why this matters. The platform clearly has likeness detection running server-side — the fourth generation was blocked. But three of four got through. The platform’s own safety system caught one output and missed three. If the vendor’s detection system is inconsistent, post-generation screening by production teams will be even less reliable.

The exposure here is quantifiable but not eliminable. You know you’re exposed. You cannot measure how much. The best practice is to limit use of these tools to internal iterations, previsualization, and temporary outputs that will be replaced before client delivery. You accept the exposure in exchange for speed and flexibility, but you don’t deliver these outputs to final.

Open Weights

FLUX, Stable Diffusion XL, and most checkpoints running through ComfyUI workflows fall into this category. The weights are publicly released. The training data is unknown, partially disclosed, or documented in a lineage you cannot fully verify. And there is no vendor liability. The vendor released the weights. What you do with them is entirely your responsibility.

Open weights models have no built-in mechanism preventing training on scraped likeness data. The vendors don’t warrant otherwise because they don’t provide warranties. The checkpoint you download from a community repository like Civitai or Hugging Face carries no provenance guarantee. No license chain. No training audit.

Here’s where it gets concrete. Your artist is running a ComfyUI workflow. They downloaded a checkpoint from Civitai. The checkpoint was trained on a dataset assembled from open web scraping. The dataset may have included Getty Images thumbnails, Instagram portraits, actor headshots indexed by Google. Nobody documented exactly what went in. Nobody reviewed the licensing. Nobody checked for scraped likeness data. The checkpoint has 40,000 downloads and produces high-quality outputs.

Your artist generates a background crowd shot. Twelve faces in the frame. One of them resembles a recognizable actor. Your artist didn’t prompt for that actor. The face appeared because the checkpoint learned it from training data. You deliver the shot. Six months later, the actor’s legal team sends a letter.

You go back to the checkpoint. You check the model card on Civitai. The training data section says “various open datasets.” There is no audit log. There is no license declaration. The vendor who released the weights is not in the lawsuit. You are.

This paradigm offers maximum creative control and zero legal protection. Use it only if your production has the legal capacity to defend a likeness claim without vendor support. Most productions don’t. If you cannot afford the lawsuit, you cannot afford the checkpoint.

Calling this the Paradigm Classification framework gives it a name, but the concept should already be part of every pipeline supervisor’s working vocabulary. The alternative is what the Seedance test revealed: discovering after generation that the only verification tool available — Google Reverse Image Search — returned one wrong answer and provided no answer for two others.

The Verification Gap

Google Reverse Image Search was not designed as a production-grade likeness verification system. It indexes what it indexes. It does not index every face in every training dataset that every AI model was trained on. The coverage gap is not a temporary limitation waiting for a fix. It is a permanent structural condition.

As of writing this article, there is no industry-standard tool for output likeness screening before delivery. Not one that production pipelines are running at scale with reliable results.

The platforms know this. The Terms of Service agreements were written with this knowledge in the room.

The defensible position is not post-generation screening. Screening what already came out of the model is too late. The defensible move is pre-generation model selection. You choose the paradigm before the session starts. You make the liability decision before the first prompt is typed.

What Defensible Workflows Actually Look Like

A defensible AI workflow is not about having the perfect tool when something goes wrong. It is about having the right documentation when someone asks what happened.

A defensible AI workflow is not about having the perfect tool when something goes wrong.

A defensible AI workflow is not about having the perfect tool when something goes wrong.

Document Your Model Provenance

Every checkpoint loaded in a production session is a liability declaration. If you cannot identify the checkpoint’s paradigm, its license status, and the provenance of its training data, you cannot quantify your exposure.

This means you need a manifest. Every checkpoint used in production gets logged before any generation happens:

  • Model name and version
  • Source repository or vendor
  • Download or license acquisition date
  • License status (if declared)
  • Training data provenance (if available)
  • Paradigm classification (indemnified / proprietary / open weights)
  • Risk tier assignment

You write this down before you generate the first frame. You do not reconstruct this during legal discovery after a claim arrives.

The manifest is not bureaucratic overhead. It is the only evidence you have that the exposure was known and managed. Without it, your legal position becomes “we didn’t know what was in the model.” That position does not survive scrutiny.

Document Your Creative Intent

If a prompt returns an unexpected likeness, the only evidence of your intent is your process record. Timestamps. Prompts. Iteration logs. The absence of deliberate likeness instruction is a factual claim. Without documentation, it becomes an unverifiable assertion.

Here’s what this looks like in practice. Your artist is working on a period drama. The prompt reads: “1920s street scene, pedestrians in period clothing, overcast lighting.” The model returns a frame with a face that resembles a living actor. Your artist did not prompt for that actor. Your artist did not use a reference image. The face appeared because the model learned it from training data.

You have two defenses. First, the prompt log proves you did not instruct the model to generate that likeness. Second, the iteration log shows the face was an unintended result, not a deliberate reproduction attempt. Without those logs, you have nothing. The output exists. The likeness exists. The claim arrives. You cannot prove intent.

The documentation standard is straightforward:

  • Session ID for each generation session
  • Checkpoint or model used
  • Full prompt text
  • Timestamp
  • Output file path
  • Artist or operator name
  • If iterations occurred, log each one
  • If multiple checkpoints were used, log the switch

The log should be machine-readable. The log should be stored outside the artist’s local workstation. The log should survive the end of production.

Separate Tools by Risk Tier

Risk tier separation happens at the pipeline level, not at the artist level. You do not leave it to individual artists to decide which tool is appropriate for which shot. You define the boundaries. You enforce them. You log which tier was used for which deliverable.

Tier One: Indemnified models only. Adobe Firefly and similar platforms with explicit vendor warranties. Use this tier for shots that will appear in marketing materials, theatrical trailers, client-facing deliverables, or anything with broad public distribution. The exposure is highest here. The creative constraints are the price of the warranty.

Tier Two: Proprietary without indemnification. Runway, Seedance, Kling. Use this tier for internal iterations, previsualization, temporary outputs that will be replaced before final delivery. You accept the exposure in exchange for speed and creative flexibility. You do not deliver these outputs to clients. You do not include them in final locked frames.

Tier Three: Open weights. FLUX, Stable Diffusion XL, community checkpoints. Use this tier only if your production has the legal and financial capacity to defend a likeness claim without vendor support. If you cannot afford the lawsuit, you cannot afford the checkpoint.

The separation is not punitive. It is protective. When a claim arrives, you have documentation showing that exposure was assessed and managed according to output risk, not ignored or left to chance.

The Liability Architecture

The liability architecture of AI generation tools is not a temporary gap waiting for regulation to close. It is an intentional design. Platforms distributed legal risk downstream to capture value upstream. They built the models. They control the training data. They write the Terms of Service. They retain the indemnity. You receive the output.

Production teams that did not read the Terms of Service carefully are operating under contracts they may not fully understand. That is not an accusation. It is a description of the current state of many AI-enabled production pipelines.

The accidental likeness is the moment that contract becomes visible.

The Training Data Pool

The visual effects professional who ran the Seedance test made one observation worth repeating: “Even through iteration to modify the face of a person, the AI will use its training data to come up with variations. So all variations could be copies of existing people.”

Every variation is drawn from the same pool. The pool is undisclosed. You can iterate endlessly. You are not escaping the training data. You are moving through it.

The checkpoint you load is the contract you sign. The question is whether you understand what is in it before you deploy it in production.

Key Takeaways

  • AI image generation platforms operate under three distinct liability paradigms: proprietary with indemnification, proprietary without indemnification, and open weights — each carrying different legal exposure levels for production teams
  • Post-generation likeness screening is insufficient defense; the only reliable protection is pre-generation model selection based on documented paradigm classification
  • Defensible workflows require three components: checkpoint manifests documenting model provenance, prompt and iteration logs proving creative intent, and risk-tiered tool separation enforced at the pipeline level
  • Terms of Service agreements from most AI platforms explicitly transfer liability for copyrighted or likeness-infringing outputs to the user, regardless of whether the reproduction was intentional

Before You Go

If understanding your AI tool exposure before delivery matters to you:

  1. Give it a clap 👏 to help other VFX supervisors, post producers, and studio heads find practical, field-tested insights on production risk.
  2. **Connect on LinkedIn**: Let’s discuss creative craft and emerging technology.
  3. Follow for more: Get grounded, actionable analysis on AI, film production, and workflow governance — no hedging, no fluff.
  4. **Buy Me a Coffee** ☕ — a small gesture that keeps the research and caffeine flowing.

Your Next Step: Audit which paradigm category each of your current AI generation tools falls into. Document it. That single act changes everything about how you defend your pipeline.


메타데이터
post_id
a7599f0097ef
slug
the-likeness-screen-youre-running-doesn-t-work-a7599f0097ef
url
https://medium.com/@jdcruel/the-likeness-screen-youre-running-doesn-t-work-a7599f0097ef
canonical_url
https://medium.com/@jdcruel/the-likeness-screen-youre-running-doesn-t-work-a7599f0097ef
author_url
https://medium.com/@jdcruel
status
ok
fetched_at
2026-06-14 11:28:49