OSINT future: Where Did This Image Come From?
Author: Berend Watchus. Independent non-profit AI & Cybersecurity Researcher. Publication for OSINT Team, online magazine. May 28, 2026.
OSINT future: Where Did This Image Come From? Building a family tree for AI content — and why it’s risky
Author: Berend Watchus. Independent non-profit AI & Cybersecurity Researcher. Publication for OSINT Team, online magazine. May 28, 2026.

copyright: Ching-Chun Chang and Isao Echizen https://arxiv.org/pdf/2605.27551

https://arxiv.org/pdf/2605.27551

https://arxiv.org/abs/2605.27551
In two earlier pieces for this publication and System Weakness, I argued that mandatory “online passport” identity verification was building something dangerous: permanent, immutable identity-linkage that would become catastrophic the moment it leaked.
[embed][Why 2025's 'Online Passport' Gold Rush Will Get People Blackmailed and Ki77ed Why 2025's 'Online Passport' Gold Rush Will Get People Blackmailed and Ki77ed Author: Berend Watchus [note: nov 10…systemweakness.com](https://systemweakness.com/why-2025s-online-passport-gold-rush-will-get-people-blackmailed-and-ki77ed-3d9d4c1aa19c)
[embed][Why 2025's 'Online Passport' Gold Rush Will Get People Blackmailed and Ki77ed Why 2025's 'Online Passport' Gold Rush Will Get People Blackmailed and Ki77ed Author: Berend Watchus [Publication for…osintteam.blog](https://osintteam.blog/why-2025s-online-passport-gold-rush-will-get-people-blackmailed-and-ki77ed-16ade544319a)
That was late 2025. By spring 2026 the prediction had a body count of breaches behind it — IDMerit’s roughly one billion exposed records, Persona’s age-check system found running 269 distinct identity and watchlist checks behind a simple selfie prompt. The pattern held: collect immutable identity data at scale, link it to sensitive activity, and you have built a blackmail and prosecution engine waiting for a breach.
A new paper proposes the same failure, moved one step upstream — to the moment of creation, where there is no database left to breach at all.
It is called On the Origin of Synthetic Information by Means of Steganographic Inheritance, by Ching-Chun Chang and Isao Echizen. It is elegant and biologically framed, aimed at a real problem: in a flood of synthetic media, how do you trace where an image came from? Their answer borrows from genetics. When a platform generates an image, it hides a trait derived from the parent source invisibly inside the offspring. The trait survives editing. Extract it later, match it against candidates, and you have found the parent. Repeat across generations and you reconstruct a family tree of synthetic images — one original at the root, stylized descendants branching above it.
It is clever work, and I want to take its problems seriously rather than dismiss it. The trouble is that the same design that traces a forged image also traces everyone else, and that second function is the one the paper never really reckons with.
The technical novelty is narrower than it looks
Credit first. Steganography — hiding data inside media — is decades old, and its image form is mature. The authors say so, benchmarking against signal-processing methods from the early 2000s and noting, honestly, that those old techniques held up surprisingly well against modern generative edits. The hiding mechanism is not the contribution. Neither is provenance-at-generation: C2PA, Google’s SynthID, and earlier deep-watermarking systems already embed origin data into content — and two of them sit in the paper’s own baseline table.

One of those baselines is worth pausing on. StegaStamp, from 2019, was built with an explicit ambition: a hidden code robust enough to embed a unique identifier “within every photo on the internet.” Universal traceability was never a worried-about side-effect of this line of work. It was the stated goal, years ago, by authors whose system the new paper builds upon.
So what is actually new is heritability: rather than a fixed tag, the mark is derived from the specific parent, so each generation carries a trace of its ancestor and a multi-step tree can be rebuilt. That invites two questions any analyst should ask. First, what does heritability achieve that a flat unique ID at each step, with parent-child links held in an external registry, does not? The paper asserts the advantage more than it demonstrates it. Second, and worse: the whole scheme depends on the platform cooperating at every step. The authors concede that once content passes through an uncooperative tool, the chain snaps. Hold that concession. It becomes the heart of the problem.
A broken promise versus an open register
Here is the distinction that separates this from my earlier work on verification breaches. The breach was a broken promise. People began under a pseudonym, with secrecy promised; the harm came when that promise failed. Everyone can see that as wrong, because a safeguard collapsed.
Image ancestry is the opposite. It is an open register by design. There is no promise of secrecy to betray, because disclosure is the entire point. It is sold not as a risk but as a virtue: “this image came from X.” Provenance. Attribution. IP protection. And who could be against that?
That is precisely why it is more dangerous. The verification system had to fail to expose you. The ancestry register exposes you by working perfectly, and is applauded for it. We have learned to fear the database that breaks its promise. We have not yet learned to fear the one that never made a promise in the first place — that was built, from the start, to tell everyone where you came from.
Who actually gets traced
Now look at who is standing inside this register.
Midjourney, DALL·E 3, Adobe Firefly, Stable Diffusion, Leonardo — tens of millions of hobbyists, designers, students, marketers, game artists. The paper requires platform cooperation to function. So if it is adopted, traceability is not applied to suspects. It is applied to the entire user base, by default, at the moment of creation.
And it does not stop at AI users. Follow the cooperation logic honestly and the population is everyone who makes a visual. Adobe is on that list twice — Firefly and Photoshop — so every edited photograph is in scope. If cameras, phones, or editing suites embed or inherit the trait, every photographer is traced. If social platforms propagate it — and they are the natural cooperating layer — every upload, repost, meme, and profile picture joins the tree. Every company and organisation with visuals — packaging, marketing, a logo, an internal deck — has its output rendered into a traceable lineage, handing competitors free reconnaissance.
The cooperating parties — camera makers, editing software, generators, social platforms — are exactly the chokepoints every image passes through on its way to the public. There is almost no normal path to an audience that avoids all of them. “Opt out” is not realistically available to anyone operating normally.
And here the cooperation assumption inverts from a weakness into the central horror. From the user’s side, cooperation is the threat. The determined bad actor the system was built to catch simply steps outside the mainstream toolchain and walks away clean. The ordinary person — who had no reason to be traced and no way to opt out — is the one left permanently marked. It is a manhunt that misses the fugitive and tags the bystander.
From lookup to list
There is a step the provenance framing quietly skips. A mark that reliably answers “who made this?” for one image does not stay a single-image tool. Once the answer is extractable at scale, “find everyone who has ever made this kind of image” stops being a manhunt and becomes a database query. The open visual web turns into something you can run a SELECT statement against, indexed by maker.
Two collection logics follow, both ugly. The first is targeting by content: compile the makers of anything an actor cares about — political imagery, dissident material, pornography, religious provocation. This is watchlist logic, and the matching infrastructure already exists; Persona was found running 269 checks, including watchlist and politically-exposed-persons screening. The mark simply adds a new field to match on. The second logic is worse and specific to this medium: collect everything, because you can. Video is expensive to store at population scale. Images are not. The rational move for a bulk collector is not to decide today what matters but to keep all of it — mark and all — and decide later. The innocent image you make this morning sits in the archive until the day the definition of “undesirable” changes around it. That is the “back catalogue” problem from the verification breaches, except every image is pre-linked to its maker at the moment of creation.
The result is a standing, append-only, globally queryable record of who made what — cheap to keep, searchable by category, matchable against any watchlist, mineable retroactively whenever a government or a market or a blackmailer revises what counts as interesting. And note the difference from every breach I have documented before: this record requires no breach. The mark is public by design — you extract it, you do not steal it. The blacklist is not the system failing. It is the system working exactly as specified, pointed at a list.
This will not be the project of one rogue regime. For any serious intelligence service, a globally extractable, maker-attributing mark sitting in public imagery is not an opportunity to weigh — it is low-hanging fruit they are practically mandated to collect. The logic is the oldest in the trade: a capability you decline, your rival builds. The moment one major service compiles the maker-index, every peer must follow or cede the ground. It costs almost nothing — public images, parsed, not networks, breached — and yields a population-scale record of who made what. Nobody whose job is collection leaves that on the table. The system can be sincerely meant for accountability and it changes nothing: from a collector’s chair, it is a free, standing, global registry of everyone who ever made an image worth caring about.
Nor will the collectors only be states. A state service at least operates, in principle, within some structure of law and oversight. The same low-cost, high-yield logic hands the index just as readily to actors bound by nothing: private-intelligence and corporate-espionage firms with a client and an invoice; data brokers adding one more field to merge and resell; organized crime assembling extortion lists; harassment networks and self-appointed morality vigilantes; and, at the smallest scale, the stalker or abusive ex who once lacked the resources to trace a target and now needs only a decoder and an afternoon. For these actors “illegal intelligence practice” is not a deterrent but a description of the work. The barrier that protected people was never that the data didn’t exist; it was that compiling it was hard. This makes it easy. And things that are cheap to do get done — by everyone, not only by those authorized to.
There is a final inversion, the ugliest. Until now we have assumed the target actually made the thing being traced. The mark removes even that. Anyone who has published ten thousand images has handed an adversary ten thousand valid parent-sources. Take one, generate something vile from it, and the heritable trait points back to them — not by forgery, but truthfully, because the nasty image genuinely descended from their photograph. The system cannot distinguish “made by them” from “made from them by an enemy”; it sees descent and reports it as origin. The attacker then chooses where to drop it for maximum harm — into the religious-police jurisdiction the target is visiting, a competitor’s inbox, a journalist’s, an employer’s security team, a community that will turn on them.
And the attacker need not even make the accusation. The system makes it for them. They drop the lineage where it will be found and stand back: don’t take my word for it — look it up, it’s in the official record, it started with this one. The malice vanishes behind the apparent neutrality of a technical trace. It can be amplified by sock-puppet accounts manufacturing the look of corroboration, by recruited real people, or — worst, because it needs no coordination — by ordinary, careless individuals who simply believe a record built to be believed and pass it on. “The official record says you are the source” is very nearly unfalsifiable in the only courts that matter here: a forum, a mob, a border desk, an employer. And in the framing case the record is not even lying. The vile thing truly descended from their image. So when the crowd looks it up, it checks out. The system performs perfectly and confirms the frame. A person is destroyed not by a forgery but by a true trace of a thing they never did.
The danger is in the success, not the failure
A persistent, transformation-resistant, content-bound trait is a provenance signal from one chair and a profiling key from the other. An adversary need not decode the mark to use it — only that it be stable and linkable. “I cannot read who made this, but I can prove the same source made these ten thousand artifacts” is itself the harm. This is why the “keyed, opaque payload” defence fails: keying protects legibility, not linkability, and linkability is where profiling lives. Pseudonymity is not privacy.
The cross-jurisdiction hazard is sharper still. Because the mark is in the content, not held by a gatekeeper, it travels with the artifact — across borders, into jurisdictions where authorship is itself an offence. And control is exactly what no maker has. Once an image is published, shared, or scraped, its creator cannot decide who picks it up, which country it travels to, how it is altered, or where it eventually surfaces. A photographer in Amsterdam, an illustrator in São Paulo, an AI artist anywhere — none can govern the journey of what they made. The mark makes that uncontrollable journey traceable back to them at every stop. As I noted in the 2026 follow-up: the same data means embarrassment in one country and a threat to life in another. The mark is indifferent to which — and so is the route the image takes to get there.
“Settle it in copyright law” — and why that fails
The obvious rebuttal is that this belongs in intellectual-property law — let creators settle provenance there. But that misunderstands the harm. IP law governs ownership and permission: who may use an image, and whether they were allowed to. The danger here has nothing to do with ownership.
Consider a handful of images scraped at random: a luxury car with a Cuban licence plate; topless sunbathers on an anonymous beach; an unidentifiable machine part, perhaps an engine; an abstract oil painting of an inverted cross amid cadavers and blood; a photograph of an art gallery; a group of business figures and politicians; some unremarkable interiors. On their own these images are mute. A car is a car; a gallery is a gallery. Reused with or without permission, they say nothing about anyone.
What a heritable origin mark adds is precisely the story they lack — who made each one, who used it, where it travelled, in what context. Suddenly the mute car is “photographed by a named person on a trip never meant to be public”; the painting is “created by, and circulated among, these specific people”; the business figures are “linked to whoever commissioned and shared this.” The exposure is not in the pixels. It is in the lineage the mark attaches to them. Copyright law has nothing to say about that, because no ownership question is being asked. You cannot settle in an IP court a harm that IP law was never built to see.
And this is not an edge case requiring a sophisticated adversary. Anyone can choose, in five seconds, a handful of utterly ordinary images — a sunbather, an artwork, a political portrait, a religious provocation — each a crime to have made or shared somewhere on the map. The images are mute and the selection is trivial. What is not trivial is the answer the mark supplies to the only question that matters in that jurisdiction: who made this? The creator never chose to send their work there. The mark sends the answer anyway.
It defeats forgettability before anyone weighed the trade
The paper sends the legal and ethical questions to “future multidisciplinary collaboration,” which understates them. I want to avoid invoking a specific erasure statute, because those are full of carve-outs and, in any case, plenty of records are deliberately permanent and legitimate — vehicle ownership, criminal histories, land and company registers. We accept those because a sector argued out the trade and decided persistent traceability served a public good.
That is the point. The question is not whether some law forbids an undeletable lineage trait. It is whether the interests were weighed at all — and they were not. Property registers exist because legislators, insurers, and courts fought it out and struck a balance that is still revisited. This trait proposes permanent, content-bound, heritable traceability across all visual media, by default, before that weighing has been done for this mechanism. It offends the principle of forgettability — that a person can act, let time pass, and not be pursued indefinitely — not after deliberation, but as an engineering default.
Measured against the four safeguards I recommended after the 2026 breaches, the design fails each. Data minimisation: it maximises, embedding a permanent identifier in every artifact. Segmentation: it unifies, linking everything one source makes. Privacy-preserving alternatives: it offers none. Clear cross-border standards: it ships the linkage across every border by default, inside the file.
How it becomes total — without anyone deciding it should
Two opposite desires push toward the same place. The people who most want out of any registry — makers of intimate or sensitive or dangerous-where-they-travel content — cannot stay out, because the toolchain marks them anyway. And the people who want in — anyone with commercial or copyrighted work to protect — walk in voluntarily, because provenance is the very thing being sold to them as protection. Demand for attribution and the impossibility of escape converge on one database. Suddenly everyone’s entire visual footprint is registered, indexed, permanent — the way we record house ownership, ISBNs, DOIs, car registration.
But those are registries of deliberate, discrete, public-by-nature assets, each justified by an interest someone argued out. This is a registry of everything anyone ever made: the holiday snap, the discarded draft, the nine thousand images you never thought twice about. Not a registry of assets but of people — of their complete visual trace — arriving with the boring civic legitimacy of a land registry. That is the genuinely dangerous part. The most total record of human visual activity ever proposed gets to wear the same unremarkable administrative uniform as the Kadaster (The Netherlands’ Cadastre, Land Registry and Mapping Agency — in short Kadaster), and nobody marches against an administrative record. Of course we track who made what. Don’t we track who owns what house?
example:
And it will not stay voluntary, because it need not be mandated — only made the price of protection. Picture the courtroom, a few years on. A photographer, or a model, files an infringement claim over their own work. The court asks the question registration regimes have always asked: was the image registered as an official work, as the statute requires? It is not a novel demand — copyright systems already bar suits over unregistered works and reserve their real remedies for those who registered in time. Port that logic to images and the unregistered photograph becomes unenforceable: anyone may take it, because it has no protected standing. No law need command that every image be marked. A law need only rule that only marked images are defensible — and the outcome is identical, wearing the same reasonable face as the ISBN.
This welds the trap shut. The maker who wanted to stay out for safety learns that staying out means total defenselessness. Safety and rights become mutually exclusive — enroll and be traceable, or abstain and be powerless. There is no third door. And the cruelty lands hardest on the model, asked to prove a right over the image of her own body by first surrendering that image to the official index — permanent, attributed, discoverable in every jurisdiction it ever reaches. To assert control over your own likeness, you must register it into the system that exposes you. The law makes protecting yourself and exposing yourself the same act.
What this ends in
Let me build the case deliberately, because the point is how easily ordinary, separate, lawful lives are fused into one exposed target.
Imagine someone who makes LGBTQ content in the EU, where it is lawful — and who also, quite separately, photographs wildlife and ruins for a living, including lawful trips to countries where the first kind of work is a crime. Two unrelated parts of one ordinary life. The watermark refuses to keep them apart. The lineage asserts they are a source; their innocent in-country photographs, scraped and regenerated by strangers into illegal content, point home to them anyway. A forensic check of the laptop they carry for editing recovers files they deleted but never wiped, because deletion was never erasure. A cover credit, found in seconds, supplies identity and motive — and the professionalism aggravates rather than mitigates, because it means produced and distributed, deliberate and repeated. The interrogator’s line writes itself: we already had enough — you are the source, it is watermarked into the image, the whole lineage points here; and now we have even found the originals on your hardware. It would say this even if the originals were planted, because the record would still confirm them.
Or take something far more ordinary: an OnlyFans model who flies home to visit her parents. Her work is lawful where she lives and makes it. Her family lives somewhere it is not — somewhere a customs officer, a hostile relative, or a local authority with a grudge needs only to look up what the register already says she has made. She has committed no crime in the country she has entered. She has simply gone to see her parents, carrying a phone and a face that a maker-index can resolve. The separation she relied on — work in one country, family in another, never the two systems meeting — is exactly what a content-bound, border-crossing mark dissolves. She did not bring her work with her. The mark brought it for her.
Neither of these is an exotic case assembled to frighten. They are two unremarkable lives — a working photographer, a content creator visiting family — exposed not by anything they did wrong (locally), but by a system that insists every image carry its maker’s name into every place it travels (including later iterations). It could also have nothing to do with pornography or nudity and more with politics etc. Even if you would ignore all photography, still all original AI art and images will have similar problems.
And consider how much of this population’s work is already beyond their reach. A webcam performer, a Patreon or OnlyFans creator, a porn producer does not have a tidy portfolio of a few controlled images. They have hundreds, often thousands — some published deliberately, some sold to subscribers under a promise of limited access, and a great deal stolen and redistributed, without consent, across content-harvesting sites, leak platforms, and fan forums built precisely for that purpose. Every one of those copies, authorised or not, is a valid parent-source. The mark does not introduce the loss of control; that already happened, the day the work was ripped and mirrored. What the mark adds is to turn each of those uncontrolled copies into a homing beacon, resolving every scattered fragment back to one person — and it cannot tell the deliberate post from the stolen leak. To the index, the victim of non-consensual sharing and the willing publisher are the same data point. The leak forums that already sort this content by performer become, with a maker-resolving mark, exactly the pre-built, identity-linked target list of the kind no one ever consented to.
And there is already a market that shows how irreversible this is. A content-protection industry sells creators takedown services on subscription, and these services can remove leaked material from large automated stream-hoarding platforms quickly — sometimes within minutes or hours. But removal is partial and ongoing: new copies reappear, and a recurring fee buys continual cleanup rather than a final resolution. That a whole paid industry exists to chase this, and never finishes, is itself the evidence that “you cannot pull it back” is not rhetoric but a priced, permanent reality. And a maker-resolving mark changes the stakes of everything that is not removed: each remaining copy becomes individually traceable to one named person. The creator pays to clear what they can, while the mark keeps a homing signal on all that stays. As with the verification startups I wrote about in 2025, the persistent exposure is not an accident at the edge of the system — it is a standing condition the system runs on.
So people will reach for the dark. To share a private nude between adults, harming no one, they can no longer use any normal channel, because every normal channel marks and registers what passes through it. To share a political image where that carries risk, the same. To do the unremarkable, safely, they must adopt the operational security of a smuggler — and the open, registered world is left holding only the innocent and the compliant, while the shadow channels fill with criminals and ordinary people fleeing a database, side by side. The system does not shrink the dark. It populates it. And it gives away the whole game: a nude between adults, an opinion, a self-portrait were never the danger. The registry makes them dangerous purely by recording them.
But the likeliest ending is none of these. Most people will not flee, and will not resist. They will do what people do: go along. Accept the terms, use the convenient tool, assume that something offered as protection must be protective, and never trace the implications to their end. The registry will not have to be forced on anyone. It will be adopted, one accepted default at a time, by people doing the reasonable, frictionless thing. And that is exactly how the whole of this article comes to apply to them — not because they were compelled, but because going along was easier than asking where the line led. By the time the consequences arrive, everyone is already inside.
메타데이터
- post_id
- e849515581e6
- slug
- where-did-this-image-come-from-building-a-family-tree-for-ai-content-and-why-its-risky-e849515581e6
- url
- https://osintteam.blog/where-did-this-image-come-from-building-a-family-tree-for-ai-content-and-why-its-risky-e849515581e6
- canonical_url
- https://osintteam.blog/where-did-this-image-come-from-building-a-family-tree-for-ai-content-and-why-its-risky-e849515581e6
- author_url
- https://medium.com/@BerendWatchusIndependent
- status
- ok
- fetched_at
- 2026-06-23 17:05:31