Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM-Driven Code Rendering

In plain words: how can we tell whether a video was painted by a model or drawn by code an AI assistant wrote? And what could a hidden watermark show? (LLM means large language model, the AI behind chat assistants.)

A video model can paint every pixel of a video directly.

Model-generated A paraglider with a red wing stands on a rocky ridge in warm low light, looking out over hazy mountain ranges and layers of cloud.

Screenshot. Frame from Google's Veo 3.1 documentation, where the clip is labelled as generated by Veo 3.1. Source

An AI assistant can instead write code that a renderer turns into frames.

Code + imported assets A 3D animated scene in a green park with low-poly trees. Two blue pillars carry game box art and a ranking number, one labelled Super Smash Bros. Ultimate, 35.8 M, number 20.

Decoded frame. Frame from a video by Dilum Sanjaya (@DilumSanjaya), shown in the official Remotion gallery, which labels the tool Claude Code. Source

Either way, the checker (whoever tests the clip) often gets only the finished pixels, not how the video was made.

What the checker actually receives

clip.mp4
  frame 0001: pixels
  frame 0002: pixels
  frame 0003: pixels
  ...

often missing:
  the prompt
  the code
  the model's hidden data
  a log of how it was rendered
  a signed record

Two ways to make a video

Look at the two frames at the top of this page. One comes from a video that Google labels as generated by its video model. The other comes from a video credited to a creator who worked with code and an AI assistant. You may feel sure which is which. But style is a weak clue: a model can be asked for many looks, and code can display imported photographs. A feeling is not evidence.

This page is a plain-language guide to a research paper about provenance: where a clip came from and how it was made. The paper asks how we could find out how a video was made, and what a watermark could and could not prove. A watermark, or mark for short, is a hidden signal added to a picture or video so that a detector can look for it later. The paper poses questions. It does not report experiments. First, here are the two main ways videos are made today.

Watch the video

Prefer to watch? This 4-minute explainer goes with this page. It was made by the paper's first author and follows the same research agenda in everyday words. You can play it right here.

A frame from the video: two bar charts on crumpled paper, one drawn directly by an AI image model and one drawn from code, with the caption “Two charts. Same numbers, similar look.”
Rethinking Visual Provenance: Two Ways AI Makes Pictures, and How a Hidden Tag Might Help. A 4-minute explainer by Zheng Gao, first author of the paper, posted on YouTube on 6 October 2026. Its description says it is a research agenda (“questions, not results”) and discloses how it was made: drawn from code with AI help, a computer-generated voice and AI-generated characters. Press play to load the video from YouTube; nothing is loaded before that. Watch on YouTube ↗The cover is a frame at about 0:08 from the local render of this video, not a screenshot of the YouTube page. More in the credits.

Route A: a model paints the pixels. Many image and video generators are diffusion models1. They start from random noise and remove the noise step by step, guided by a text prompt.

Often the model works on a compressed internal code, and a decoder2 expands that code into the final pixels. The steps in the middle are numbers inside the model, not instructions a person can read.

Route B: code draws the pixels. An AI assistant can instead write code. A separate tool, the renderer, runs that code and draws every frame. For example, Claude Code is an AI coding assistant that edits files and runs commands on a computer. Remotion3 has a tutorial on using Claude Code to build a video project. The renderer then draws frame 0, frame 1, frame 2, and so on, and combines them into a video file. Here the middle step is readable: it is the code.

In this page, a model “paints” (generates) pixels directly, and code “draws” (renders) them through a renderer. After this section we mostly use the technical words, generate and render.

A video can use both. A clip can have a generated background, a chart drawn by code, and a person's own annotations. When an AI agent4 chooses among these tools, the paper calls it a controller. A controller is not a third kind of video. It is the part that decides which tools to use.

Diagram with an agent at the top choosing between three panels: a, image or video generation through a model, a latent and a decoder; b, code-based rendering through code or SVG and a render step; c, hybrid composition that layers a generated asset, code and a human edit.
Figure 1. Two routes to pixels, and a mix of both. Route A (panel a) goes through a model and a decoder. Route B (panel b) goes through code and a renderer. Panel c shows a video built from generated assets, code and human edits. The panels show steps, not a claim that the routes give identical pixels. Conceptual illustration made with an image model for the paper; not a measurement. This is Figure 1 of the paper (PDF p. 1).

The same chart, made two ways

The easiest place to see the difference is a still picture. Both charts below show the same four invented weekly numbers. The left one was made by an image-generation model. The right one is SVG code5 written by an AI agent and then turned into a picture on our own computer. Both are illustrations for the paper. Neither is a measurement of anything.

A stranger who found these two pictures would have to guess the route from the style. A guess is not a label.

Real examples

Below are frames from real videos. Each caption names the source and says what that source itself reports about how the video was made. We took the labels from the sources, never from how a frame looks. These are screenshots of existing videos, shown only for illustration. We did not run any test on them.

Labelled as generated by a video model

Google's documentation labels these frames as outputs of its Veo 3.1 video model.

Credited to code-based workflows

Each frame here is credited to a video rendered from code. The credits come from the creators, or from the official Remotion gallery. Some of these videos also use imported images or 3D models. For the “still loading” frame, the credit comes from the creator's own report (he is a co-author of this paper), not from the public page.

Key ideas


Three questions: guess, mark, authenticate

Imagine a friend sends you a video clip and asks, “Is this made by AI?” That one question contains three different questions. Each needs different evidence, and confusing them is the easiest way to claim too much.

1. Guess: does it look machine-made?

A detector studies the video and estimates how it was made. Nobody arranged anything beforehand, so the paper calls this passive inference (it also uses “detection” more loosely, for this and for finding a mark). Think of a detective judging a painting by the way it was painted.

Where the comparison breaks: a detective can point to clues and explain them. A detector may give a verdict learned from example videos6, and it can be wrong. The benchmarks we read do not settle whether a detector trained on videos from video models says anything reliable about a video drawn by code.

2. Mark: is a hidden signal in it?

Here someone adds a hidden mark that people do not notice: the creator, or the provider of the tool. Later a detector checks for it and may read a short hidden message. That is a watermark. Think of an invisible-ink stamp checked under ultraviolet light.

Where the comparison breaks: a stamp stays on the paper. A digital mark can be put in at different steps, and later edits can damage it. A detector also needs the right tool or key. And finding nothing does not show that no AI was involved. The clip may come from a tool that never marks, or the mark may have been lost.

3. Authenticate: does a signed record support it?

Here a trusted issuer, such as the company that provides a tool, signs a record that says, for example, “this file came out of our tool.” A digital signature lets anyone check who signed, and a later change to the signed data breaks the check. A standard called C2PA7 describes how such records work. Checking such a record is called authentication. Think of a receipt stamped by an official notary (a person who officially confirms a document).

Where the comparison breaks: a receipt says what the notary saw. Even a valid signature backs only the statement itself. It does not show that the scene is real or that nothing was added later. And a receipt can become separated from the file it describes.

Two small examples from the paper

These examples do not make the evidence useless. They show that a result must always be reported together with the exact test that produced it. A mark is still useful evidence, but what it says depends on what was tested and who had what. A good system should also be allowed to answer “can't tell.”

Key ideas


What does the checker receive?

Imagine you are the checker. In this page, the checker is whoever or whatever examines a clip: a person, or software such as a detector. A clip arrives and you must say something about where it came from. What do you actually have?

Usually it is just the video: a grid of coloured dots, called pixels, for every frame. Sometimes the creator also provides the code. Sometimes there is a signed record. What you can honestly conclude depends on which you have.

Choose what the checker has

Each option below adds just one item to the video. The options are not stacked on top of each other.

What you can say. A matching watermark detector can test whether a mark that was added to these finished frames is present. A detector that guesses how the clip was made can also give an answer, but its guess may be wrong.

What you still cannot say with certainty. Which steps made the clip (the paper treats this as an open question), who wrote any code behind it, or that no AI was involved. Two different pieces of code can draw exactly the same pixels.

The same idea as a grid

Each row is a claim someone might want to make. Each column is what the checker has. This grid is our own plain-language summary of Sections 2, 3 and 5 of the paper. The paper has no such table and does not rate evidence by strength. The shading shows how far each kind of evidence can go. No cell is unconditional: even the best case depends on a stated condition.

Claim
Video only
Video and code
Video and signed record
A mark added to the finished frames is present
Only if a mark was added, it survived the edits, and you have the matching detector.
Only if the same conditions hold (a mark was added, it survived, and you have the detector).
Only if the same conditions hold (a mark was added, it survived, and you have the detector).
A message hidden in the code is present
Only if the renderer passes the message into the pixels. A renderer that ignores it leaves nothing to find.
Can support, if you also know how the message was hidden (a key or a method for reading it).
Not by itself. A record about a file does not contain the code unless it says so.
A specific tool exported this file
Only if that tool also marked the pixels.
Cannot establish. Code does not say where or how it was run.
Only if the signature is valid, you trust the issuer, and the record is still attached to the file.
The scene is real, or the history shown is complete
Cannot establish.
Cannot establish.
Cannot establish. Later changes may be undocumented.

Same picture, different code

Here is the smallest case where information disappears while rendering. Below is a tiny SVG drawing with a comment, a note for human readers. Suppose the creator hides one bit of a message there, 0 or 1. The renderer ignores comments. Below are three versions of the drawing, one after another; with scripts on, they appear as three tabs you can click.

<svg viewBox="0 0 160 100">
  <!-- message: 0 -->
  <rect width="160" height="100" fill="#f6f6f3"/>
  <circle cx="80" cy="50" r="28" fill="#d97757"/>
</svg>

The comment says 0. Someone who receives this file can read the 0. Someone who receives only the picture cannot.

Look at the second version again. A checker who sees only the picture can only guess8. That follows from the pictures being identical, not from anything special about circles.

Two limits show what this example does not claim. First, it depends on a renderer that ignores comments. Some renderers read comments or names, and a number in the code can turn into motion on screen. So the real question is never “is it code?” It is “does the change reach the pixels?” Second, it says nothing general about watermarks. Source-code marks exist, and so do marks in 3D assets that are designed to stay detectable in rendered images. The point is narrow: rendering can erase some differences and keep others.

Diagram in two panels and a bottom strip. Panel a: two different pieces of code, P0 and P1, go through the same renderer and give the same pixels, which a pixels-only observer sees. Panel b: if the rendering step keeps the mark, marked state becomes marked pixels that a detector can read. The strip lists four things a checker may receive: pixels, source or latent state, metadata, and authenticated records.
Figure 3. What reaches the checker depends on the rendering step. Left: two different pieces of code can give the same pixels, so a pixels-only checker cannot tell them apart. Right: a change that alters the picture does reach the checker. The bottom row lists four things a checker may be given. In the picture, “the map” is the rendering step, “observer” is the checker, and P0 and P1 are two different pieces of code. Conceptual illustration made with an image model for the paper. The left side assumes that the two renderings give the same pixels. It does not test this by running anything. This is Figure 2 of the paper (PDF p. 8).

In plain words

Before asking “does the mark survive?”, ask “what does the checker receive?” More evidence helps only when you know exactly what it shows.


Where could a watermark go?

Suppose you run a video company and want every clip to carry a hidden mark. Where would you put it? There are more places than you might expect.

Each route below is a list of steps. For each step, a panel shows what you could mark there, what can remove or limit the mark, and who needs access. With scripts on, you click a step to open its panel. The “what you could mark” lists follow Section 3 and the table of candidate locations (Table 2, Appendix D) of the paper. The “what can remove or limit it” and “who needs access” lines are our plain-language reading of Sections 3 and 5, not results from the paper. Where the paper names no published method for a stage, we say so. The lists show where a mark could sit, not that any product does this.

Method names are examples of published work, and each is linked under Read further. If you would rather start from a concrete example, section 5 shows the real path that Route B can follow.

Route A: a model paints the pixels

What you could mark here
Text that goes into the model, or text a language model writes for it. The paper's table lists “conditions or generated textual content” at this stage. It lists this as a location only and names no method for it.
What can remove or limit it
Rewriting the text. The finished clip does not carry the prompt, so a checker who has only the clip cannot check it.
Who needs access
Whoever controls the request: the person, the AI agent, or the service they use.
Four panels on marking a still image: a, a marked noise pattern goes into a generator; b, latent codes pass through an adapted decoder; c, a raster image passes through an embedder; d, a spatial extractor returns a message for each region of a composite image.
Figure 4. Four things to look at when marking an image. (a) The starting noise. (b) The decoder. (c) The finished pixels. (d) A reader that returns a separate message for each region. The teal areas show where a mark could be placed. They are not measured patterns. The figure draws still images for clarity; the video methods named above work at the same kinds of places. Conceptual illustration made with an image model for the paper. This is Figure 3 of the paper (PDF p. 11).

Route B: code draws the pixels

What you could mark here
Choices in the code the assistant writes. Some methods (for example SWEET and STONE) nudge which words the model picks while it writes code, so the finished code carries a pattern. Others (for example ACW and SrcMarker) rewrite finished code in small marked ways.
What can remove or limit it
Rewriting or reformatting the code. The renderer may also ignore these choices, as in the comment demo above.
Who needs access
Whoever holds the code. A checker needs the code and the shared key or detection function.

The last two steps are the same for both routes. A mark added to finished pixels can be applied to a model's video, to a code-rendered video, or to a camera's. What differs is everything before it. These steps are different places to put a mark. They are not ranked from weak to strong. Several can be used at once, and a checker may not know which step was marked.

Three rows of what someone could do: edit exported pixels, render a marked asset again so the asset mark may persist, or render clean source so that an optional output-only marking step is bypassed.
Figure 5. Three different things someone could do. Top: edit the finished pixels. Middle: draw a marked asset again, where the mark may persist. Bottom: render again from clean code and clean assets, which skips a mark that was added only to the output. This does not show that upstream marks can always be erased. Conceptual illustration made with an image model for the paper. This is Figure 5 of the paper (PDF p. 15).

In plain words

Marks can go in many places. Each place has its own checker, its own access needs and its own ways to be lost. No place is simply the best.


How today's tools make code-made video

A creator types a request: “make a ranking video of the best-selling games, in a 3D scene.” Here is the documented path that request can follow, as the official documents describe it. We read the documents and the public code (a pinned snapshot of Remotion's open-source repository: one fixed commit, accessed 4 October 2026). We did not run any of it.

  1. You describe the video in words. For example, a ranking of the best-selling games, shown in a 3D scene.
  2. Claude Code writes the code. Claude Code is an AI coding assistant that edits files and runs commands. It works inside a Remotion project folder. Remotion's tutorial lists Claude Code among its supported coding agents and walks through creating the project, installing agent “skills” (add-on instruction packages) and starting a live preview (a window that shows the video while you edit).
  3. Claude Code produces code. It says what to draw at each frame number.
  4. A renderer runs the code, frame by frame. In server-side rendering, the renderer packages the project and picks a scene, which Remotion calls a composition. It then computes each frame, captures it, and compresses the frames into a video file.
  5. The video file is exported.

So the assistant supplies the code, and a separate renderer draws the video. Remotion's documents also show a second way of working: a language model returns code as text, which is then compiled and rendered. That example uses a provider other than Anthropic.

A 3D animated scene in a green park with low-poly trees and a chain-link fence. Blue pillars carry game box art, a title, a sales figure and a rank number.
Example. A code-rendered game-ranking video. In a public thread, its creator Dilum Sanjaya describes opening a Remotion project in Claude Code and describing the video they wanted. The thread says a few versions were generated to learn what matters in prompts, and that images and a Sketchfab 3D background (Sketchfab is a site that hosts 3D models) were added later. Decoded frame from the video in the official Remotion gallery, scaled to 1280 px wide, no other edits. The gallery credits @DilumSanjaya and labels the tool Claude Code. Gallery entry · Creator's workflow post · Creator's asset post

Look at what the creator says. The code, the imported images and the 3D background have different origins. Who wrote the code, where each asset came from, and which renderer drew the final frames are three separate questions.

Where the code runs decides who holds what

The same code can run in three places. Each changes who can see the code, the logs and any secret keys. None of them changes who wrote the code.

What marks and credentials exist today

Anthropic's help page (How Claude marks AI-generated content) describes two things: watermarks that some supported models add to the text they write, and C2PA credentials that some supported file types get. Which models and files are covered depends on the page; we did not test any of them. The page also says a positive detection result is evidence, not full proof. It adds that converting a file or taking a screenshot can remove metadata, the extra information stored inside a file.

Code is text, but a finished video is pixels. As the comment demo in section 3 showed, a mark in text does not automatically reach the pixels. None of the documents we read says that a mark in Claude's text is still present in a rendered video. This does not mean it is impossible. It only means we have no source that says it is present.

C2PA credentials are signed records tied to a file. If the attached data is lost, a watermark or fingerprint (a short summary computed from the content) can help find the right record again. C2PA calls this soft binding10. Soft binding helps find records, but it does not replace the signature.

Still pictures also have two routes

Chat assistants can draw still pictures in more than one way too, and a picture you receive may not say which. Claude's custom visuals feature builds diagrams and charts as web code, with SVG or HTML downloads. OpenAI documents both Code Interpreter, which runs Python and can make graph images, and a separate image-generation tool. The Claude Artifacts help page we consulted lists some export formats, and we found no general promise of MP4 export on it.

In plain words

Today's documented path is simple: you ask, an assistant writes code, a renderer draws the frames. The code is a readable set of instructions, but often only the operator has it, not the checker.


Mixed videos: many sources in one frame

The frame below comes from a music video called “still loading”. Its creator, Zheng Gao, is a co-author of this paper. He reports that the video was rendered by code written with Claude's help; we have not checked this independently, and the public Bilibili page does not say how it was made. The scene inside looks like a detailed illustration. So where did that scene come from? The sources do not say.

A young woman with braided blonde hair, a yellow cardigan and a long red skirt walks down a stone lane lit by lanterns, above a harbour town at sunset. A small loading bar and a time label sit in the upper-left corner.
Example. Reported by its creator as code-rendered; where the artwork inside came from is not stated. Frame at 01:11 of “still loading” (Bilibili, 1 October 2026). Its creator, who is also a co-author of this paper, reports that it was rendered by code written with Claude's help; the public page does not state the workflow. Decoded frame, retouched in one small area. In the upper-left corner (about 2% of the frame) an image model was used to paint over the platform's logo and account-name overlay and to redraw the part of the progress bar behind it, so those pixels are generated, not original. The rest is the original decoded frame. Source

This can happen: a video drawn by code can still contain things the code did not create. Examples are pictures, 3D models, voices, or footage from elsewhere. Remotion's Prompt to Video template is a documented case. It combines a generated script, generated images and a generated voiceover, made with OpenAI and ElevenLabs services, into one video. An assistant that can both run code and call an image tool could, in principle, mix the two routes inside one project (the paper treats this as a modelled possibility).

So “model-made” and “code-made” can be too coarse. The paper treats an AI agent as a controller that picks operations. As section 1 said, a controller is not a third kind of video.

A mark on one layer is evidence about that layer only

Picture a lesson video with three layers: a generated background, a chart drawn by code, and a person's circle and arrow. Suppose a detector finds a watermark that belongs to the background. That tells you about the background. It does not tell you who drew the chart, who added the annotations, or whether the lesson is right.

Three inputs, a generated mountain-lake background, a code-drawn line chart and a human-drawn circle with an arrow, are layered into one final frame. A bottom strip lists possible evidence: asset identity, visible contribution and export event.
Figure 6. One composition, several evidence questions. Evidence about the background's identity, about what is still visible, and about the export event are different claims. None of them proves the history of the whole frame. Conceptual illustration made with an image model for the paper. It does not show real detector output. This is Figure 6 of the paper (PDF p. 16).

Three traps follow from this.

One 2026 preprint, TRACE (Gao et al.; its first author, Zheng Gao, is also this paper's first author), takes a different approach. It proposes marking an agent's log of actions instead of the pixels. A mark in the log is evidence about the log, and a checker who has only the clip does not have the log.

In plain words

In a mixed video, a mark on one part is evidence about that part only. Say which layer, and which claim, the evidence is about.


Before anyone says “robust”: seven things to pin down

Someone tells you their watermark is “robust”, which means the mark can still be read after edits. Robust against what? For whom? To support which claim? Without answers, the word means very little.

The paper asks for seven things to be written down before any comparison (Sections 1.3 and 2.2 of the paper). It calls the full set a verification specification. Here they are as plain questions. The small circled letter beside each is the symbol the paper uses.

  1. How was the video made?

    List the steps: inputs, tools, settings, and any randomness. If an AI agent is involved, say how its choices are limited.

  2. What may the party that adds the mark see and change?

    That party might see only the finished pixels, or the code, or a model's inner state. What they can access decides what they can mark.

  3. What does the checker receive?

    Only pixels? Also the code, a record, or a key? You saw in section 3 that this changes what can be concluded.

  4. What can an attacker see and do?

    Someone with only pixels can edit the file. Someone with the code can also rebuild it and render it again. These are different powers.

  5. What must stay true?

    The task the video does (the paper calls this task preservation). A chart must keep its numbers and labels. A lesson must keep its order of events. Looking similar is not enough.

  6. What exactly is being claimed?

    “A mark is present.” “This layer came from that asset.” “This tool exported this file.” Each needs different evidence.

  7. What does an unmarked video look like?

    You need a baseline to measure false alarms: how often a detector says “marked” about a clip that was never marked.

The paper's full specification also fixes how keys are generated, who may ask the issuer to sign things, and who may run the checker. We leave those out here to keep the list short. The benefit is simple. A test that uses only pixels, a test that reads a message from the code, and a signed export record each depend on different assumptions. If any of the seven things changes, you are no longer comparing the same thing. A good setup also lets the checker say “can't tell”.

In plain words

Before you compare two watermarks, check that the same seven things are fixed. Otherwise you are comparing the setups, not the watermarks.


Ten open questions

Imagine a short animation in which bars grow to their final heights, like the bar chart in Figure 2. The numbers and labels must stay right. (The paper's own worked example is smaller: one bar over five frames.) How much else about that animation could a creator change to carry a hidden message, and what would survive editing?

Section 6 of the paper poses ten questions of this kind. The paper does not answer them, and it does not claim that every version of them is unsolved everywhere. Each is written as a small, carefully limited problem with a small starting case where one can check the basic definitions. Below, each one is restated in plain words, with why it matters and a sensible first step. The full versions, with exact conditions, are in the paper. Many of them use the bar animation as a running example: its values and labels must stay correct, so the task limits what may be changed. The four themes in the filter are our own grouping for browsing, not the paper's sections.

Showing 10 of 10 questions

1Can the route that made a clip be told apart, under stated conditions?

In plain words: take clips of the same kinds of scenes, each made wholly by one route (mixed videos are treated separately). Is there a simple rule for when a test can beat guessing, and when it cannot, as the scenes and the choices of code vary? The test may answer “can't tell” only up to a stated limit. The checker may also be given extra information from a fixed menu, each item with a stated cost in trust or disclosure, but never a free, trusted label of the route.

Why it matters: two different pieces of code can give identical pixels. The chance of picking each one may differ between routes, so look-alike clips do not settle the question either way.

A sensible first step: start with very small artificial examples where the best possible test can be computed exactly, and let the test say “can't tell”.

2Which hidden changes are still there after rendering?

In plain words: pick small changes that keep what a scene is for. After rendering, frame sampling and editing, can a checker who has the pixels and a key, but not the original clean video or the scene's coordinates, still tell them apart?

Why it matters: a change that leaves the picture identical gives the checker nothing to find.

A sensible first step: in the bar animation, vary the bars' timing slightly. Ask how many different timings can still be told apart as the video is sampled less often (fewer frames per second).

3From drawing instructions to a picture

In plain words: take a simple drawing language. Which line, colour or timing details can carry a message all the way to the exported picture (the raster image)? Assume the checker does not know which renderer or settings were used.

Why it matters: a mark in the source is not automatically a mark in the picture.

A sensible first step: fix a tiny set of drawing rules and search for one method that works across several renderers, or find a case where none can.

4A fair comparison of places to mark

In plain words: compare marking at each stage the paper lists (the noise or decoder, the code, the assets, the renderer and export, and the finished pixels) under the same task, attacker and checker. When can one place do as well as another?

Why it matters: otherwise “A beats B” may only mean A had more access.

A sensible first step: fix the hidden message size, error rate and false-alarm rate. Then ask whether a pixel mark can imitate an upstream mark (one placed at an earlier step).

5How much can we hide?

In plain words: in the bar animation, how many message bits can survive while the chart stays correct? How does that change with frame count, rounding and edits?

Why it matters: the room the task leaves, not merely the number of possible videos, limits what can hide.

A sensible first step: start with zero errors allowed, where different messages must give different possible videos even after the permitted edits.

6What can someone with the code undo?

In plain words: suppose an attacker can rebuild the video from a marked image or 3D model from an earlier step. How much access and effort does it take to push the average chance of reading the message down to a stated level while the video still does its task?

Why it matters: rendering a marked asset again, swapping in a clean one and rebuilding the code are different powers.

A sensible first step: let the attacker change at most a few timing numbers in a small set of rules, and look for a bound.

7Finding the position in an edited clip

In plain words: after cuts, speed changes or dropped frames, how much of the clip do you need to read the message and know where you are? You must not produce false results by trying many alignments.

Why it matters: reporting the best of many tries inflates false alarms unless the whole search is counted. A repeating pattern alone can leave the absolute position unknowable, though the clip's length or boundaries may still give it away.

A sensible first step: allow one deletion at an unknown spot, with no blending.

8Edits on top of edits

In plain words: a mark survives one allowed edit, then another. Can we state simple, checkable conditions under which the message and the false-alarm guarantees still hold after the pair, without checking every sequence? A counterexample would also be an answer.

Why it matters: each cut may remove up to a quarter of what it receives. Start with 64 frames: the first cut removes 16, and the second removes 12 of the 48 left. Together they remove 28 of 64, more than a quarter of the original.

A sensible first step: allow a few frame insertions, deletions and replacements, then one spatial change. Measure the total change compared with the original clip.

9Is this asset still visible in the video?

In plain words: a layered video uses an original asset, such as an image or 3D model. Suppose the checker has the asset and a record of how the layers were masked (the layout). When can it show the asset is still visible in a frame, keeping both kinds of mistake (missing it, and wrongly claiming it) under control?

Why it matters: layers can be hidden, look-alike patches can come from several assets, and a match does not show who used the asset or when.

A sensible first step: two solid layers, nothing transparent, and no compression. Then add edits.

10Records that support an event without revealing extra information

In plain words: take one trusted issuer. Can a record let a checker accept authorised copies of its output almost always, and wrongly accept anything else only rarely? And can it do this while revealing nothing beyond what the file shows, plus the public information the setup declares, such as the issuer, the kind of event and the verification result? The keys are assumed to stay uncompromised.

Why it matters: hiding a message is not privacy, and a genuine asset credential does not confirm a new export.

A sensible first step: compare a signature on the exact file, a credential kept apart from the file, and a signed record found again through soft binding. Ask when a watermark inside the video actually helps.

In plain words

The ten questions share one idea: say exactly what is changed, what the checker sees, and what is being claimed. Then ask what is possible.


What this page is and is not

Before you take anything from this page, here is exactly what it is.

Screenshots show how their sources label them. They are not evidence that any watermark or detector works.

What it is

What it is not

About the pictures

The video frames are screenshots or decoded frames of existing videos, used here to illustrate. One was supplied by the paper's author. Each caption gives the credit and what the source says about how the video was made. The five concept figures, and the left chart in Figure 2, were made with an image model for the paper. They explain ideas. They are not results and not proofs. One frame has a small retouched area, which its caption describes.

In plain words

This page explains a set of good questions. It does not answer them, and it makes no claim about how well any watermark works.

Words used here

Watermark
A hidden mark added to a picture, video or text so that a detector can find it later. It may also carry a short message. “Mark” is short for watermark.
Provenance
The history of where something came from: who or what made it, with which tools and inputs, and what happened to it afterwards.
LLM
Large language model: the kind of AI behind chat assistants.
Checker
Any person or software that examines a clip to say something about how it was made. What it can conclude depends on what it receives: pixels, code, keys or records.
Detector
A checker that looks for a mark or estimates how a clip was made. It can be wrong.
Passive inference (passive detection)
Guessing how a clip was made from the clip alone. Nobody arranged a mark beforehand, so the answer is an estimate and can be wrong.
Authentication
Checking a signed record from an issuer, such as “this file came out of our tool”, under stated trust assumptions.
Signed record
A statement protected by a digital signature, so that changes to it can be detected.
Issuer
The person or organisation that signs a record, such as the provider of a tool.
Key
A secret value, like a password, used to add a mark or to sign a record. A checker may need the matching key.
Fingerprint
A short code computed from the content of a clip. It can help find the matching record without a hidden mark.
Robust
A mark is robust if a checker can still read it after allowed edits such as compression or cropping. The paper asks: robust against what?
Diffusion model
A generator that starts from random noise and removes it step by step, guided by a prompt, until a picture or video appears.
Latent
A compressed internal code that some generators work on before turning it into pixels.
Decoder
The part of a generator that turns a latent into pixels.
Metadata
Extra information stored inside a file, such as who made it or a signed record. Converting the file can strip it.
Pixels
The tiny coloured dots that make up each frame of a picture or video.
Renderer
Software that follows a description, such as code or an SVG file, and draws the picture or video frames it describes.
Rendering
The act of drawing frames from code or an SVG file. The same code can give different results if the renderer or its settings change.
Operator
The person or company that runs a renderer or a service.
Asset
An image, 3D model, texture or sound file that code loads and uses when it renders a video.
SVG
A text format for drawings. A file lists shapes, colours and text. It can be opened and edited as text, or drawn as a picture.
Remotion
A toolkit that makes videos from React code. The code describes what each frame shows, and Remotion renders it.
Claude Code
An AI coding assistant from Anthropic that can edit files and run commands in a project. Remotion documents how to use it for video projects.
AI agent (controller)
Software, often guided by a language model, that picks tools and runs them. The paper calls it a controller when it chooses among ways of making media.
C2PA and Content Credentials
A standard for signed records about how a piece of media was made or handled. Content Credentials is the common name for such records.
Soft binding
Using a watermark or fingerprint in the content to find the matching C2PA record again when the metadata has been lost. It does not replace the signature.
Hybrid media
A video or image built from several sources, such as generated assets, code-drawn parts and human edits.
Task preservation
Keeping what the video is for. A chart keeps its right numbers and labels. It need not keep every pixel the same.
False alarm
A detector saying “marked” about a clip that was never marked.
Attacker
Anyone who tries to remove, fake or avoid a mark. What an attacker can see and change, such as only pixels or also the source, is part of the setup.
Can't tell (abstain)
To answer “can't tell” instead of yes or no. A good checker is allowed to do this.

Read further

Twenty-five primary sources from the paper's reference list. Each link goes to the source itself: a paper, a standard, or official documentation. Years are the first public (arXiv) year unless a venue is named.

Watermarks for images and video

Marks in code and in 3D assets

Tools and workflows

Provenance, detection and limits

Credits

Example frames

All frames are screenshots or decoded frames, shown to illustrate. We claim nothing about them beyond what each source says. Labels follow the sources and not how a frame looks. X may ask you to log in to open the X links below.

Illustrations

The paper

“Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM-Driven Code Rendering”. Authors: Zheng Gao, Xiaoyu Li, Zhicheng Bao, Yang Song and Jiaojiao Jiang, UNSW Sydney. The PDF prints no date; its sources were accessed on 4 and 5 October 2026, and we date this page October 2026. We describe the paper as a conceptual working draft; the paper itself calls it a conceptual research agenda. Read it as a PDF (46 pages, 15.8 MB).