
An AI provider announces text watermarking. The headline is simple. The conclusion many people may draw from it is not: watermarked means written by AI.
For IZZY, the reason to examine this news is to make a more useful distinction. A text can be processed with AI without being autonomously authored by AI. If people collapse those two events, a technical signal can become a false account of who contributed what.
Imagine a colleague writes a press release. They ask an AI assistant to improve the structure, translate one paragraph and fix the punctuation. Later, a detector reports a watermark.
Did the AI write the release?
That question asks the signal to do more than it can. A text watermark can support a narrower conclusion: a compatible model was probably involved in generating enough of the tested wording for its statistical pattern to be detectable. It is not a percentage meter for human versus AI authorship.
The distinction matters because AI now sits at many points in a content workflow. “Used AI” can mean generating a first draft, rewriting a human draft, translating it, changing five sentences or correcting three commas. Those are different contributions, even when the final document carries the same binary label in a dashboard.
This article explains what generation-time text watermarking is, what a positive or negative result can establish, and which records a publishing team still needs.
Answer in 60 seconds
An AI text watermark is usually not a visible stamp or a hidden character. In the statistical schemes covered here, the model embeds a pattern while choosing the next tokens in its response. A detector with the relevant key or scoring method tests whether the word sequence contains enough evidence of that pattern. (Anthropic, Dathathri et al., Nature)
A detected watermark can support that the watermarked model was likely involved in producing or processing the tested text. It does not tell you:
- who originated the ideas or supplied the source draft;
- what percentage of the document is “human” or “AI”;
- whether the model drafted the text or heavily edited it;
- whether the claims are true, sourced, original or legally usable;
- who owns the text or is responsible for publishing it.
The reverse is also important. No detected mark does not prove human authorship. The text may be too short or factual, lightly proofread, generated by an unmarked or older model, heavily edited after generation, or tested with an incompatible detector.
Anthropic’s current explanation makes the mixed-authorship problem explicit. Its announced Claude watermark can only estimate the likelihood that Claude was partly involved; it cannot distinguish “Claude wrote this” from “Claude heavily edited this”. Anthropic also says a grammar-and-punctuation-only proofread may leave too few changed words for the mark to register, while a translation produced by Claude can carry a watermark because Claude chooses the translated words. (Anthropic)
Treat a watermark as one provenance signal. Keep the draft history, the model-use record, the human review and the publication decision separate.
In this article
- What “AI watermark” can mean
- How a statistical text watermark works
- How IZZY distinguishes AI authorship from AI assistance
- What a positive or negative result can establish
- What the EU AI Act changes, and what it does not
- How the IZZY workflow records AI contribution
- What to do when a detector flags a document
- Conclusion
- Frequently asked questions
- Sources
1. What “AI watermark” can mean
The word watermark is being used for several different mechanisms. They should not be treated as interchangeable.
| Mechanism | What it records or tests | What it does not establish |
|---|---|---|
| Generation-time text watermark | A statistical pattern embedded as the model chooses tokens | The human/AI authorship percentage, factual quality or ownership |
| Content Credential or provenance metadata | Signed statements about an asset’s origin, tools or editing history | That every statement in the credential is true in a wider real-world sense, or that missing metadata means no AI was used |
| Visible AI label | A disclosure presented to a reader | The detailed contribution of each person or tool |
| Post-hoc AI-writing classifier | Stylistic or learned features associated with AI-written text | A provider-specific watermark unless it has the relevant watermark method or key |
| Internal process record | Who used which tool, for what task, with which review and approval | A technical signal embedded in the published text |
Anthropic illustrates the first two mechanisms in the same announcement. It describes a statistical watermark for Claude-generated text, but a C2PA Content Credential for supported image and file formats. The text pattern is embedded in generation; the file credential is signed provenance metadata. (Anthropic, C2PA specification)
C2PA itself is careful about the boundary. Its specification provides a way to attach cryptographically verifiable provenance assertions to an asset. It does not turn provenance into a judgement that the content is “good”, “bad” or factually trustworthy.
This is why “the detector found an AI watermark” is not enough information for a governance decision. First ask: which mechanism, from which provider or standard, tested by which tool?
2. How a statistical text watermark works
An autoregressive language model produces a response one token at a time. At each step, it estimates a probability distribution over possible next tokens and samples from the available choices.
Some choices are tightly constrained. In “2 + 2 =”, replacing “4” with another number changes the answer. Other choices have more freedom. A model may be able to select “grey” or “overcast” without materially changing the sentence.
A generation-time watermark uses those flexible choices to introduce a keyed statistical pattern. The detector later scores the relationship between the observed tokens and the pattern expected under that key. The result is evidence with a confidence threshold, not an invisible sentence saying “this document was written by AI”.
The general approach is established in peer-reviewed work. A 2023 ICML paper described a scheme that promotes a randomised set of candidate tokens and detects the resulting pattern with a statistical test. The 2024 SynthID-Text paper described a different sampling method, keyed scoring and production-scale deployment. Anthropic says its Claude implementation is a version of the SynthID-Text approach. (Kirchenbauer et al., ICML, Dathathri et al., Nature, Anthropic)
Three practical limits follow from the mechanism.
Length matters
Longer passages usually offer more token choices and therefore more evidence. Short samples can leave a detector uncertain. No universal minimum applies across providers, models, languages, prompts and detector thresholds.
The type of text matters
Creative or varied writing offers more alternative phrasings than exact quotations, code, fixed facts or tightly constrained answers. Google DeepMind says its text watermark works best on longer, diverse responses and is less effective where little variation is possible. Anthropic makes the same point for factual passages, proofreading and much code. (Google DeepMind, Anthropic)
Later editing matters
Light editing may preserve enough of the pattern to detect. A substantial rewrite, paraphrase or translation by another system can weaken it. The SynthID-Text paper identifies editing, spoofing and scrubbing as continuing limitations; Anthropic says a complete rewrite can remove its mark. (Dathathri et al., Nature, Anthropic)
These are scheme-level observations, not a guaranteed threshold for a particular document. Until a provider publishes its detector, operating parameters and validation results, do not turn “probably robust to light edits” into a numeric rule.
3. How IZZY distinguishes AI authorship from AI assistance
The following distinction is an IZZY editorial definition. It is not a legal test, an industry standard or a claim that authorship can always be reduced to one rule.
AI authorship: the model is given the topic and creates the content autonomously
In the clearest AI-authorship scenario, a person supplies a topic or short instruction and the model produces the material without a human-designed research framework, source process, decision structure or specialist editorial direction.
The person may still decide whether to use the output. But the substance and wording of the draft were produced by the model from a minimal brief.
AI assistance: people design and own the content system
In an AI-assisted workflow, the human contribution is not limited to correcting a finished model response. People define the problem, build the research and editorial process, set the evidence boundaries, direct the model’s task, review the output and remain responsible for the publication.
The model proposes content inside that system. A specialist then checks, changes, accepts or rejects it.
This means “AI was used” and “AI was the author” are not equivalent statements.
A public version of the IZZY content workflow
The simplified workflow used for an evidence-led article is:
- Select the signal and reader decision. Decide why the topic matters now, who needs the answer and which decision the article should support.
- Check existing coverage. Review current IZZY articles and the live service context so the new piece has a distinct job and does not repeat an existing article.
- Research current questions. Examine recent discussions for language, confusion and objections, while treating community evidence as directional rather than representative.
- Build the evidence base. Open primary, official, peer-reviewed and relevant practitioner sources; separate verified facts from observations and inference.
- Create the claim boundaries. Record what each source supports, what it does not support and which volatile claims will need rechecking.
- Define the editorial position. Set the angle, structure, exclusions, first-party IZZY contribution and the practical reader outcome.
- Use AI for synthesis and structure. Give the model the research and evidence boundaries to organise, compare and stress-test, not to invent an article from a one-line topic.
- Apply specialist editing. Check the reasoning, facts, sources, technical or regulatory nuance, IZZY voice and whether the proposed framework is genuinely useful.
- Run publication QA. Recheck links, metadata, unsupported claims, limitations, conversion path and any need for legal or subject-matter review.
- Keep human responsibility. A named person approves, revises, delays or rejects the article. The model does not make the publication decision.
This is deliberately a public, simplified description. It explains the control points without disclosing IZZY’s internal prompts, qualification logic, research operations or other proprietary methods.
This article began with a founder-selected news signal. The technical and regulatory claims were checked against primary and peer-reviewed sources; the interpretation of authorship versus assistance was then reviewed with the founder; and the article remains subject to human editorial and legal review. It is not presented as a client case or a fully autonomous AI article.
What changes across drafting, editing, translation and proofreading
The original idea behind this article was that even a typo check could watermark human text. The evidence supports a more precise answer: a light proofread can produce too few model-chosen words to create a detectable signal. The probability of detection grows as the model makes more wording decisions.
| Workflow event | How much wording the model chooses | What a watermark result may show | What it cannot tell you |
|---|---|---|---|
| Human supplies a draft; model corrects punctuation and a few errors | Very little | A mark may be too sparse to detect | Whether the proofread happened if no mark appears |
| Human supplies a draft; model line-edits several sentences | Some | A mark may show likely involvement if enough marked choices remain | Which ideas or passages originated with the human |
| Human supplies a draft; model substantially rewrites it | A great deal | A positive result may be consistent with heavy model editing | Whether the model or human should be called the “author” |
| Model creates a full first draft | Most or all output wording | A sufficiently long result may carry strong watermark evidence | Whether the facts are correct, sources were checked or a human later approved it |
| Model translates human-written text | All wording in the target language | Anthropic says its translated output carries its watermark | Who authored the underlying ideas or source-language text |
| Human lightly edits model-generated text | The model’s earlier choices mostly remain | The watermark may survive | The scale or quality of the human review |
| Human or another model completely rewrites the output | Few original choices remain | The original watermark may weaken or disappear | That no AI was involved earlier in the workflow |
| Human text is pasted into an AI chat only for analysis, with no returned text published | None of the published wording | A generated response may be marked, but the untouched source is not rewritten by that act | Whether AI was consulted, unless a separate process record exists |
The table explains why authorship percentages are the wrong output. A detector looks for evidence in the final token sequence. It does not replay the drafting history, compare every version, assign intellectual contribution or decide whether “author”, “editor”, “translator” or “proofreader” is the right description.
Two documents could therefore produce similar detector results while having very different histories:
- one started as a human draft and received a substantial model rewrite;
- another was model-drafted and then carefully reviewed and edited by a person.
The watermark alone cannot reconstruct which path occurred.
4. What a positive or negative result can establish
Use calibrated language. The safest interpretation depends on whether the detector is authentic, compatible with the scheme and applied to enough unaltered text.
For IZZY, the most immediate risk is not the existence of the watermark but its interpretation. A binary result can create false precision about the model’s contribution and become the basis for accusations against an employee, writer, supplier or student. The detector cannot supply the missing drafting history, so the organisation must not pretend that it can.
If a compatible watermark is detected
You may be able to say:
“This test found statistical evidence consistent with the named model or watermarking scheme having generated or processed at least part of this text.”
You should not upgrade that to:
- “AI wrote 100% of this document.”
- “The named person did not write it.”
- “The text was not human-reviewed.”
- “The content is false, plagiarised or low quality.”
- “The publication complies with AI-transparency law.”
- “The provider owns the output.”
Anthropic states that its watermark says nothing about ownership or authorship and cannot identify a specific user, organisation or chat. It is designed to test likely Claude involvement. Those are different claims. (Anthropic)
If no watermark is detected
You may be able to say:
“This detector did not find enough evidence of this particular watermark in the tested sample.”
You should not upgrade that to “a human wrote it”. Possible explanations include:
- the provider or model did not apply that watermark;
- the text predates the model’s marking support;
- the wrong detector or key was used;
- the passage is too short, exact or factual;
- the model only made a few proofreading changes;
- later editing weakened the pattern;
- another model with a different scheme generated the text.
In other words, a negative result has a defined technical scope. It is not a certificate of human authorship.
A detector score is not an authorship score
Detection performance is normally discussed through thresholds, true-positive rates and false-positive rates. The SynthID-Text research, for example, evaluates detection as a function of text length at a fixed false-positive rate. That is useful for measuring a detector. It does not convert a 99% watermark probability into “99% of the words were written by AI”. (Dathathri et al., Nature)
The percentage belongs to the detector’s hypothesis test, not to the division of creative or editorial labour.
5. What the EU AI Act changes, and what it does not
The current attention to text marking is partly regulatory. Article 50(2) of the EU AI Act requires providers of systems that generate synthetic audio, image, video or text to make outputs machine-readable and detectable as artificially generated or manipulated, as far as technically feasible. It also tells providers to account for effectiveness, interoperability, robustness, reliability, implementation cost and the state of the art. (EU AI Act, Article 50)
The provider marking duty is not absolute. Article 50(2) excludes systems to the extent they perform an assistive function for standard editing or do not substantially alter the input data or its semantics. The European Commission’s current guidelines identify standard editing as an example of what can fall outside scope. (European Commission guidelines)
Two distinctions are operationally important.
Provider marking and publisher disclosure are separate
A machine-readable mark is a technical signal added or enabled by the AI provider. A visible label is information presented by a deployer or publisher to an audience. Article 50 gives them different scopes.
For text, the deployer disclosure rule in Article 50(4) concerns AI-generated or manipulated text published to inform the public on matters of public interest. It also contains an exception where the content has undergone human review or editorial control and a person or organisation holds editorial responsibility. The facts of a specific publication still matter.
A watermark is not a compliance decision
A positive mark does not prove that the right visible disclosure was made. A missing mark does not automatically show a provider breached the law. The system, model date, use case, editing function, technical feasibility, provider/deployer role and applicable exceptions all matter.
The Commission published a voluntary Code of Practice to help providers and deployers implement these duties, which have applied since 2 August 2026. The current guidelines and Code should be checked when designing a real workflow. This article explains the evidence boundary; it is not legal advice. (European Commission Code of Practice, European Commission guidelines)
6. How the IZZY workflow records AI contribution
Watermark detection becomes useful when it sits inside the human-owned workflow described above. The following five records make that process reviewable. They are IZZY operating guidance, not a technical standard or legal-compliance certificate.
For each material public asset, keep five records.
1. Signal: what exactly was detected?
Record:
- provider, model and version where known;
- detector name, version and access route;
- date of the test;
- text sample tested;
- raw score, category or threshold result;
- whether the detector is designed for that watermark.
Avoid screenshots without context. A future reviewer needs to understand what the tool actually tested.
2. Event: what did the model do?
Classify the use: generate, expand, rewrite, translate, summarise, line-edit, proofread or analyse. Preserve the prompt or task instruction where policy and confidentiality permit.
This event record supplies the context the watermark lacks. “Proofread punctuation only” and “rewrite for a new audience” should not collapse into the same checkbox.
3. Contribution: what came from people and prior sources?
Keep the original draft, meaningful versions and tracked changes. Name the subject-matter contributor and editor. Record sources used for material factual claims.
This does not produce a perfect authorship percentage either. It creates reviewable provenance for the decision that matters: can the organisation stand behind the final asset?
4. Assurance: what was independently checked?
Record the checks relevant to the content:
- factual and source verification;
- rights, confidentiality and personal-data review;
- brand and accessibility review;
- subject-matter or legal review where risk requires it;
- final human approval and unresolved limitations.
A watermark is not a substitute for any of these controls.
5. Decision: what will be published, labelled or escalated?
Name the accountable approver. Record whether the asset can be published, needs a visible disclosure, must be revised, or requires specialist review. State the reason and the policy or rule applied.
The result is not “AI: yes/no”. It is an evidence-backed publishing decision.
If your organisation already uses an AI pilot policy, add this record to it rather than creating a parallel bureaucracy. Our AI pilot-to-production governance checklist covers the wider allow, review, disclose, restrict and block decisions; the signal-to-decision record handles the narrower watermark-evidence problem.
7. What to do when a detector flags a document
Do not begin by rewriting the text. Preserve the artefact and the result first.
- Confirm the detector. Is it the provider’s detector or a tool validated for that exact scheme? A generic AI-writing classifier is not equivalent.
- Preserve the tested version. Keep the complete text, the sample boundaries and the raw result.
- Check model coverage. Was the named model capable of applying that watermark on the generation date? Anthropic says older Claude model support is being rolled out over time, so the date and version matter.
- Retrieve workflow evidence. Look for the original draft, version history, tracked changes, model task and named reviewers.
- Classify the model’s role. Separate proofreading, line editing, substantial rewriting, translation and full generation.
- Run the checks the mark cannot perform. Verify claims and sources, assess rights and confidentiality, and confirm the responsible editor.
- Apply the relevant policy. Decide on approval, visible disclosure, revision or escalation based on use and risk - not on the detector alone.
- Escalate high-consequence decisions. Do not use a watermark result by itself to make an employment, academic-misconduct, legal or regulatory determination. Those decisions need the applicable process and specialist review.
For a low-risk blog draft, this may take minutes. For investor communications, regulated advice, employment material or public-interest reporting, it should be more formal.
If the team cannot reconstruct the workflow, label that evidence gap honestly. “Could not verify the drafting history” is safer than inventing a precise authorship story from a statistical signal.
The operational conclusion
Text watermarking can make model involvement more detectable. That is useful. It is also a narrower achievement than “solving AI authorship”.
The mark lives in model-generated choices. Authorship lives across ideas, drafts, edits, sources, approvals and responsibility. A detector sees the first layer; your content process must preserve the rest.
At IZZY, we therefore do not define authorship by whether AI touched a document. We ask who designed the content process, selected and verified the evidence, directed the work, edited the result and accepted responsibility for publication. That does not make every AI-assisted document human-authored by default. It makes the contribution question answerable with workflow evidence instead of a binary guess.
For publishing teams, the practical rule is simple:
Use the watermark to open an evidence review, not to close the authorship question.
If you need to map where AI generates, rewrites, translates or approves public content - and define the evidence each step must keep - IZZY’s Advisory work can help turn the workflow into owned controls and decision records.
Map where AI touches your public content
In a 30-minute scoping call we will map where AI generates, rewrites, translates or approves your public content, and define the evidence each step should keep.
Frequently asked questions
No. Providers can use different marking methods, keys and rollout schedules, and some models or deployments may not apply a generation-time watermark. A detector for one scheme cannot automatically detect every other model.
Anthropic announced on 14 August 2026 that future Claude models will generate watermarked text using a version of SynthID-Text, with support for older models to be added over the following months. It also said a detection API would be offered, but implementation details were still being worked out on the research date. Check the model and current provider documentation before assuming a particular response is marked. (Anthropic)
Not necessarily in a detectable way. Anthropic says a grammar-and-punctuation-only proofread may change too few words for its watermark to register. Heavier editing creates more model choices and therefore more room for a signal.
No. A compatible positive result can support likely model involvement in the tested wording. It does not calculate the model’s share of the ideas, source draft or editorial contribution, and it cannot distinguish full drafting from heavy editing on its own.
No. The sample may be short, factual, lightly edited, generated before marking support, produced by another model, extensively rewritten or tested with the wrong detector.
Ordinary copying does not change the token sequence, so it should not by itself remove a pattern embedded in word choices. Editing can weaken the pattern, and a complete rewrite may remove it. This is different from deleting metadata or invisible characters.
Anthropic says its announced watermark contains no identifying information and cannot be traced to a specific person, organisation or chat. Do not generalise that privacy property to every possible marking system without checking its design.
No. In the mechanisms discussed here, a text watermark is embedded statistically during token generation. A C2PA Content Credential is a cryptographically signed provenance record associated with an asset. They can complement one another, but they carry different evidence.
No such universal conclusion follows from Article 50. It separates provider machine-readable marking from deployer disclosure and includes scope conditions and exceptions, including standard editing in the provider duty and human review/editorial responsibility in the public-interest text disclosure rule. Apply the current official guidance to the specific workflow and seek legal advice where the consequence warrants it.
Sources, method and limitations
Researched on 18 August 2026. The core evidence was checked against:
- Anthropic’s primary explanation of Claude’s announced text watermark, published 14 August 2026;
- the peer-reviewed SynthID-Text paper in Nature;
- the foundational ICML/PMLR statistical-watermark paper;
- Google DeepMind’s SynthID text explainer;
- Article 50 of the EU AI Act, the European Commission’s current transparency guidelines and the final Code of Practice;
- the C2PA technical specifications.
Recent community discussions were also reviewed to identify live questions about proofreading, editing, translation and false authorship inferences. They informed the questions addressed, not the technical or legal conclusions. Automated community retrieval returned no ranked candidates, and X, TikTok, Instagram and usable YouTube coverage were unavailable in that pass, so the demand signal is directional rather than representative.
First-party editorial input was collected from the IZZY founder on 18 August 2026. It informed the distinction between autonomous AI authorship and a human-designed AI-assisted content process, the public workflow description and the warning about false contribution claims and accusations. It did not supply or replace the technical and legal evidence. No client case is claimed.
Important limitations:
- Anthropic had announced a forthcoming detector API but had not published its operating details in the sources checked.
- No universal minimum sample length or edit threshold was verified; performance depends on the scheme and text.
- Provider rollout and EU implementation guidance can change.
- Watermark research demonstrates capabilities and limitations under defined tests; it does not make every real-world detector result conclusive.
- The IZZY authorship/assistance distinction and workflow are editorial and operating definitions, not a universal authorship test.
- This article is operational guidance, not legal advice or a determination of authorship, copyright, employment, academic integrity or regulatory compliance.