Claude's Watermark Can't Tell Assistance From Authorship
Why a model-level watermark can turn grammar correction, faithful translation and ghostwriting into the same durable signal: machine involved.
Yesterday, Anthropic announced that supported Claude models will embed invisible watermarks in all generated text. The internet reacted exactly as you would expect: badly.
I ran this piece through Claude for a grammar, spelling and punctuation check. Then I realised what I had just done. Nothing in the argument changed. No sentence was rewritten. But the article may now have a hidden criminal record.
If I used one of Anthropic’s supported post-2 August models, Claude may have embedded an invisible statistical signal in the output:
THE MACHINE IS INVOLVED
The argument quickly became about whether people should be allowed to remove the mark, whether it will work, and whether Claude users should switch models. I wanted to look at the question underneath all of that: what exactly is Anthropic marking, what does the law require, and what happens when a grammar check and a ghostwritten draft can carry the same durable signal?
Readers who want the wider compliance map before getting into this one can start with my EU AI Act cheat sheet, which sets out the current deadline structure. What follows sits inside one provision of it.
TL;DR from an AI governance consultant
Anthropic has over-complied. It appears to mark considerably more than Article 50 requires, which is a lawful choice and the subject of this piece.
Article 50(2) tells providers which outputs they must mark. It places no ceiling on what a provider may choose to mark anyway.
Article 50 draws two lines. Standard editing and faithful translation are exempt. Source code, short outputs and non-human-facing content fall outside scope. Different legal routes, same point: not every AI interaction warrants a durable trace.
Anthropic marks all generated text from supported models, worldwide, including outputs the Commission does not require to be marked.
The mark says that a model was involved, not how much. The Code, which Anthropic has signed, describes an optional way to communicate the proportion and location of AI changes.
The detection half is far less publicly documented than the marking half. Anthropic says technical details are still forthcoming, leaving independent observers unable to assess the detection architecture, access conditions, performance or error profile.
A signal that cannot separate AI generated from AI modified, AI assisted or processed by AI is not neutral. It creates a social category with consequences the provider does not control.
My position is that a provider should stay near the legal minimum where the deception risk this provision targets is not materially engaged, and should explain publicly which objective any broader implementation serves.
What Anthropic has said
Anthropic says that supported Claude models released after 2 August 2026 carry embedded marks in generated text, while supported files receive signed C2PA provenance metadata. The company says the system operates at model level across Claude’s products and cloud distribution surfaces, worldwide, and that it is still adding support for older models.
The limit is Anthropic’s own.
A detected mark may show that Claude processed an output, and not that Claude authored the underlying work or that the signal records the full provenance of the content. That caveat is where this story begins.
If everyone has already read the announcement, here is the part they may have missed: Article 50 does not require every form of model involvement to be treated as the same event.
I think Anthropic has over-complied with Article 50. I think the provision was written as a risk-sensitive, operation-specific obligation, and that a provider should stay near the legal minimum where the law does not ask for marking. What has shipped goes considerably further, and the consequences of that choice fall on people who had no part in making it.
So what does Article 50(2) require, and what does it not
Article 50(2) requires providers of AI systems generating synthetic audio, image, video or text to ensure outputs are marked in a machine-readable format and detectable as artificially generated or manipulated.
The provision then limits itself. To the extent a system performs an assistive function for standard editing, or does not substantially alter the input data or its semantics, the marking obligation does not apply.
Those four words at the front, to the extent, carry a lot of weight.
The limit is scoped to a particular operation and to how much that operation changed, rather than switched on or off for a whole system.
That limit is where the argument starts. Marking rules tell a provider what it has to mark. They say nothing about what it is forbidden from marking. Anthropic can go further if it wants to, and my argument is that it should not without a reason that survives the law’s own logic.
The Code takes a different approach to text, where file metadata is less useful: for its voluntary architecture, it proposes an invisible watermark for longer free-form text, with a 200-token threshold. That threshold belongs to the Code, not to Article 50. The architecture acknowledges modality. It does not require the signal to communicate degree.
Where this stops being abstract is in the Commission’s worked examples.
Standard editing gets defined as preparing existing content for publication, meaning small fixes to readability, grammar, quality and format, with nothing newly generated.
Grammar correction, spellchecking, light stylistic polishing that leaves substance intact, faithful translation: all exempt.
Summaries, and rewriting that shifts style, structure or meaning: not exempt.
Now put Anthropic’s documentation next to that list.
Proofreading and translation are named there, as everyday things people ask Claude to do that can leave a mark on the output. So marking is not required in these cases. Anthropic has decided they are cases where marking happens anyway.
A second point sits buried in the guidance and matters more than it sounds. What counts is the operation, not the machine. A system capable of drafting essays does not forfeit the standard-editing exemption when it fixes a comma. A comma is not a ghostwriter. Article 50 understands that. The product design does not yet seem to.
Two things I should be straight about. Commission guidance is not binding, and the Court of Justice has the final word on any of it. And this boundary between generation, manipulation and assistance was contested long before any guidance appeared, so if you want that argument in full, Nicolaj Feltes is worth your time.
None of which makes Anthropic’s approach unlawful. Article 50 sets a floor, not a ceiling. A lawful choice can still swap the legislature’s calibration for a private company’s default.
Source code
The Guidelines treat source code as outside the scope of the marking obligation: programming, scripting, markup, query and configuration languages, including code comments that form an integral part of the output. Nobody has to mark a YAML file.
That is not a ban on marking it. It is the Commission saying it could not see what marking would achieve here.
Claude Code is one of the surfaces Anthropic lists. What it has not published is anything code-specific: whether a statistical mark reaches code output at all, or how it behaves in a language where changing one token can break a build. Until that documentation exists, the honest answer is that we do not know. I am not going to guess in either direction.
The law differentiates in several places
Article 50 is not a blunt rule. The Commission allows more proportionate approaches in limited, closed or industrial settings, where content is less likely to circulate and the risk of deception is lower.
The point is not that every output deserves the same treatment. It is that context matters: who sees it, what they can do with it, and whether anyone could be misled. That is why a provider should calibrate rather than blanket.
For the full framework, including chatbots, emotion recognition and deployer duties, see my transparency guide.
Transparency is not an AI-slop detector
LinkedIn has added an option to report AI slop. Pangram’s AI-detection tool is now embedded in Substack. Anthropic’s watermarking move lands in the same moment: another technical signal meant to tell us that AI was here.
But “AI was here” is not, on its own, a useful category. Article 50 is a transparency rule, not an AI-slop detector. It responds to a broad problem. Synthetic and materially manipulated content can mislead people, distort the information on which they make decisions, and make it harder to know where information came from in the first place. That is why Recital 133 lists watermarks, metadata identification and provenance methods among the techniques providers may use.
But a broad purpose does not mean unlimited application. Article 50 builds proportionality into the obligation: standard editing and non-substantial alterations sit outside the mandatory marking scope. The law does not ask whether AI was involved in the abstract. It asks when that involvement warrants a durable technical trace.
Anthropic has replaced the law’s operation-level threshold with a model-level default: if Claude generated the text, Claude marked it. The legislature distinguished generation from assistance. The product design treats both as the same event.
Consider how many distinct things the same signal now stands for: AI generated. AI modified. AI assisted. Processed by AI. Machine involved. A signal that cannot separate these categories is not neutral. It may be perfectly serviceable for a provider’s internal purposes, and it creates a social category whose consequences the provider does not control.
The EU icon set is more differentiated on the human-facing side. It includes a basic icon for AI involvement, alongside separate icons for content that is fully AI-generated and for pre-existing human-made content partially modified with AI. The machine-readable layer does not require providers to communicate the proportion, location or nature of the AI intervention. BTW, I can label my article:)
People get categories. Machines get a trace. The problem is that the trace may travel further.
The trust cost of a low-context signal
What follows is a governance risk, not a claim that Anthropic’s watermark has already produced these effects.
There is, however, a familiar pattern in research on AI disclosure. Tell people that identical work was made with AI and they often judge it more harshly: less authentic, less trustworthy, less worth engaging with. The Nuremberg Institute for Market Decisions found this in advertising. Majerova and colleagues report a similar result for AI-labelled marketing content.
That does not prove that an invisible Claude watermark will have the same effect. It does show why the meaning of the mark matters once it becomes visible to a client, reader, platform or institution.
These effects are not uniform across audiences, uses or product categories. They vary with AI literacy, perceived usefulness, the nature of the content and whether human review is visible. That variation strengthens the case for contextual disclosure rather than weakening it, because a generic machine-involved signal does not supply the context people need to interpret the work fairly.
That is the provenance-stigma risk. It is a practical category rather than a legal one: a technical trace of AI involvement becomes a shortcut for low effort, low authenticity, suspect authorship or weak accountability. An ambiguous signal is enough. Suspicion follows, and the person whose work carries the mark has to explain it.
A binary model-level mark does not tell its recipient:
Whether a human or the model authored the underlying work.
How much of the content the model generated or altered.
Whether the intervention was grammar correction, translation, summarisation, rewriting or full drafting.
Whether a translation was faithful to its source.
Whether a human exercised substantive editorial control.
Whether the detection result is reliable, robust or confidence-scored.
Whether the mark survived ordinary editing, copying, translation or format conversion.
Without that context, institutions fall back on crude binaries: human or AI, clean or suspect, authentic or synthetic. That is the sociological counterpart of the one-bit problem running through this piece.
Trust is not just about whether a sentence is true. It also concerns effort, accountability, intention and voice. For a newsletter writer, journalist, consultant, academic or creative, the question can quickly become less “is this accurate?” than “is this really yours?”
That does not make marking the wrong tool. It makes calibration the important question. A signal should say enough to address a real transparency risk, but not so little that it invites recipients to treat every form of AI assistance as the same thing.
Marking, detection and interoperability
Article 50 requires more than a hidden mark. The output must also be detectable. Anthropic has described the marking. It has not yet given the public enough information to assess who can use a detector, how reliable it is, what a positive result means or how someone can challenge it.
Three questions have to be answered before any institution acts on a detection result, and none of them has a public answer yet.
Who gets access to the detector.
What a positive result actually means.
How the person affected can challenge it.
Those are governance questions rather than engineering ones, and they will decide whether this signal does more good than harm long before any benchmark exists.
A recent preprint makes the wider architectural point: provenance is not something providers can reliably bolt onto the end of a human and AI workflow. Read Schmitt, Kruse and co-authors here. (A preprint, not settled law.)
Three dates are circulating, with three different legal statuses. Article 50 applies from 2 August 2026. Systems already on the EU market before that date have until 2 December to catch up on marking and detection. Code signatories then have a separate voluntary interoperability target of 2 February 2027. Anthropic signed the Code, so the clock is real even where its legal source is different.
The robustness tension
The Code asks signatories to test robustness against a demanding list: lexical substitution, homoglyphs, screenshots and format changes; paraphrasing, translation cycles, and character insertion or deletion; even the analogue hole—print, scan and optical character recognition.
Anthropic’s own limitations name heavy editing, paraphrasing, translation and short passages. The overlap is striking. That is not, by itself, evidence of non-compliance. The Code assesses performance holistically, and external benchmarks remain immature.
As GPTZero CTO Alex Cui notes, text watermarking involves trade-offs between robustness, detectability and ease of removal. That is an interested industry view, not evidence about Anthropic’s system. The broader point remains: embedding a mark is not the same as preserving its meaning.
Text watermarks and file-provenance metadata fail differently. A text watermark may survive copying, but depend on a provider-specific detector. C2PA can carry richer provenance, but can disappear through screenshots, re-encoding, format conversion or platform stripping. One can persist without enough context; the other can carry context without reliably surviving the journey.
That leaves an awkward asymmetry. Good-faith users may preserve a watermark through ordinary copying and pasting, while motivated actors can try to weaken it through paraphrasing, translation, substantial editing or re-rendering. The signal may therefore remain most visible in the workflows least associated with deception.
Your vendor’s watermark is not your compliance plan
If you wrap Claude into a product under your own name, you may carry provider duties of your own. If you publish AI-generated or materially manipulated public-interest text, you may have separate deployer disclosure duties.
Provider marking is machine-readable. Deployer disclosure is human-readable. One does not replace the other, and the Guidelines are explicit that deployers cannot rely on the provider’s Article 50(2) marking because those marks are not clear and distinguishable to the people exposed to the content.
Where the deployer duty applies, the exception has two cumulative conditions: human review or editorial control, and a natural or legal person holding editorial responsibility.
For the full Article 50 map, covering chatbots, deep fakes, emotion recognition and public-interest text, see my transparency guide and the Commission’s quick facts page.
Contract and copyright: two plausible exposure pathways
Neither of these is a settled legal outcome, and both need a lawyer rather than a newsletter writer.
The first is contractual. Many writers work under agreements containing no-AI clauses. A mark attaching to a proofread hands a counterparty a ready-made artefact for an allegation of breach, in circumstances where the mark cannot establish how much of the work was machine-produced.
The second is evidentiary in copyright. Where human authorship is contested, a signal indicating machine involvement in a substantially human work supplies an argument the other side did not previously have. The Commission is clear that applying a label or icon has no bearing on eligibility for copyright protection, which is assessed under applicable copyright law.
A Claude mark is not proof of full AI authorship, breach of a no-AI clause, absence of human authorship or ineligibility for copyright protection. Anthropic itself says a detected mark may show content was processed by Claude, and not that Claude was the original author or that the mark records a complete provenance chain.
The governance consequence stands anyway. A detection signal alters bargaining power, creates a disclosure burden and triggers costly disputes even when it proves nothing. Do not treat a detection result as standalone evidence in a disciplinary, academic, contractual, employment or copyright decision, and build contestability and human review in before anyone acts on one.
The governance consequences of marking past the minimum
The central governance problem is that a low-context signal can redistribute trust without redistributing understanding. Once a model-level mark travels into employment, education, publishing, contracting or platform governance, it starts to function as a judgment about authorship while supplying no reliable account of authorship. Employers, clients, publishers, academic integrity offices, platforms and automated moderation systems all have reasons to want a clean answer, and none of them will read the provider’s caveats before acting on one.
It can distort trust in predominantly human work and change the information conditions under which people produce content online. It raises privacy and autonomy concerns where a person’s limited interaction with a model becomes legible to third parties without a clear public-interest need. And it globalises an EU-derived product choice to writers and users outside the EU who had no role in European rulemaking.
Article 50 does not create a general privacy right against watermarking, and I am not claiming it does. These are governance and policy consequences of a broad technical implementation, and they are the reason proportionality in implementation is worth arguing about even where the law permits going further.
Who bears the cost of a blanket mark
The costs will not fall evenly. Non-native English speakers may rely more on grammar correction and translation; disabled users on AI-enabled reading, writing and communication tools; students, freelancers and early-career workers may have less power to explain what a detection result does and does not mean. Technically neutral systems can still distribute their costs unevenly.
What a proportionate implementation would look like
Make the signal carry degree. The Code already describes the mechanism, covering proportion of content generated or manipulated and localisation of AI changes. Moving that from permissive to required is the highest-value change available and needs no new technology.
Distinguish generated from modified in the machine-readable layer, as the deployer-facing icons already do in the human-readable one.
Publish detection documentation alongside marking deployment, including access rules, error rates and confidence semantics, so that institutions acting on results can calibrate what a result means.
Stay at the legal minimum where the Commission has found no marking duty, or explain publicly which additional objective the broader implementation serves.
For organisations, three things are worth doing this quarter. Ask every model vendor for its Article 50 implementation statement and its detection documentation timeline. Map where in your publishing pipeline metadata gets stripped. And write a rule prohibiting the use of watermark detection as a standalone basis for disciplinary or academic sanction, since that is the harm the regulator has already anticipated and the one nobody downstream is preparing for.
I am following how the other Section 1 signatories will implement this, since OpenAI, Meta, Microsoft, Mistral, Cohere, Getty Images, Black Forest Labs, Lenovo and Synthesia signed the same instrument.
The part I cannot resolve
I argued for this obligation before it existed. I have written that voluntary transparency initiatives are self-designed, self-enforced and non-binding, and that when accountability is defined by those being held accountable it becomes reputation management. I still think that.
The statute wrote a threshold. Anthropic chose a broader mark. That may be lawful. It is still a different governance regime, one that can treat assistance, translation and authorship as the same durable event.
The thing that stays with me is smaller than the legal argument. Somewhere in the next few months, a watermark attached after a spelling correction may be treated as evidence of broad AI authorship in a contract dispute, an academic integrity review or an employment process. The risk is not that the mark proves too much. It is that institutions may ask it to. Every document needed to prevent that already exists and is publicly available. The regulator wrote the caveat. The provider wrote the warning. Nobody in between is going to read it.
Nesibe,
📚 Everything I’ve filed so far lives at reports.techletter.co — cheat sheets on the EU AI Act, ISO 42001, agentic AI governance and AI policy structure. Sourced and dated, always.
🔔 New around here? Subscribe and I’ll see you next week.
This article was researched, written and edited by me. The analysis and views are my own. As a non-native English speaker, I used AI only for limited grammar correction and occasional translation support. All final editorial decisions were mine. Editorial responsibility for TechLetter rests with Nesibe Kırış Can.





