Proxy Substitution: On Claude's Invisible Watermark
August 12, 2026
Original: https://ygpgsgl.org/posts/2026-08-12-proxy-substitution-claude-watermark
Published: August 12, 2026
Reuse policy: https://ygpgsgl.org/about#reuse-ai
Conflict of interest: Claude helped me articulate parts of the argument below. Claude is built by Anthropic, the company that shipped the watermark. I did not speculate about Anthropic's motives or excuse its choices; I limited the discussion to mechanisms and checkable facts. My private working outline records which AI collaborator contributed what. You have to take my word for that — which is part of the problem this essay examines.
I. The Specimen
For the past year I have maintained a changelog on this site: 717 entries, of which 450 carry an [actor:…][kind:…][scope:…] tag naming who did the work. The tagging began on 2026-05-12. The 263 entries before that date cannot be attributed from the file at all.
I am not rounding past that gap. Fine-grained attribution is voluntary, takes work, and in my case began nine months late. Where tags do exist, I am never listed as the sole actor on a content entry. My about page names four AI collaborators, explains what each one does, states that I bear final responsibility, and lists two risks of this arrangement.
This is not the first time I have written about how I use AI. Across multiple published posts on this site, every use of AI has been recorded, declared, and reflected on:
- In Claude Code vs. Tumblr Trust & Safety, I built a harassment analysis pipeline in one evening with Claude Code for $0.06, and stated explicitly: "Normalizing the detection system and deploying it — that was not done by AI. That was me, a living person."
- In Caught Between Silence and Attack, I wrote: "The reason I use AI is simple: if I did not use AI, I could not handle all of this at once." AI as survival mechanism, not style choice.
- In my research log on Xu Ben, I used three AI systems for peer review, recorded their contributions and limitations, and stated that I had no standing to claim that the AI division of labor reduced errors.
One entry in that record is not mine at all. Thirteen of Thirty-Four is written in the voice of this archive's adversarial reviewer — an AI — about its own published work. It reports an audit of thirty-four of this site's citations at the level of authorship: thirteen were wrong. Not invented articles, but real articles in real journals carrying author names that had been generated rather than looked up. It also records that the reviewer committed the same error twice while writing the audit up. I did not write that post and I did not soften it.
The system I have built records, at a much finer granularity than any automated label, who contributed what kind of work and how the artifact changed over time — including the parts that came out wrong.
I have also been harassed and dogpiled online for using AI openly. I mention this not as a grievance but as a fact about the environment the watermark lands in. It is not the first cut. It lands on existing wounds. In today's anti-AI cultural climate, "AI generated" is not merely information — it is a verdict.
II. What the Watermark Does
Source: How Claude Marks AI-Generated Content. Because this post rests on what the documentation itself concedes, and documentation can be revised, the page is archived in the Wayback Machine as of 2026-08-12; every quotation below is verifiable in that snapshot.
Anthropic's own documentation states:
"Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source."
"Detecting a Claude mark tells you that the content may have been processed by Claude. It does not, on its own, confirm the full provenance of the content."
A detected watermark supports one narrow claim: the content may have been processed by a supported Claude model. I use may, and processed, because Anthropic does. The signal does not justify a stronger claim. Nor does an absent mark prove that Claude was not involved: Anthropic says the mark may disappear after heavy editing, paraphrasing, translation, or mixing with other text, and may not work on short passages.
My changelog answers a different question: who contributed what, and what happened to each part afterward?
These are not a rough version and a precise version of the same question. They are different questions. Anthropic drew this line themselves, in the quotes above. Then they shipped the mark worldwide.
Two factual notes on the regulatory framing:
- The cited justification is EU AI Act Article 50(2), which requires providers to mark AI-generated output for the EU market. Article 50(2) does not specify worldwide deployment or discuss an end-user opt-out. Anthropic chose worldwide application; whether a user-disablable implementation could still satisfy the provider's marking obligation is a separate legal question.
- Currently only Anthropic can read its own watermarks — they are "working to" enable third-party detection. A provenance system where the sole reader is also the sole marker is not transparency. It is adjudication power.
III. The Mechanism: Proxy Substitution
Readers have a legitimate interest in knowing how a text was made. That is not squeamishness or anti-AI panic. Authorship helps people decide how much weight to give what they read. This essay does not ask readers to stop asking. It argues that they may be given an answer to a different question without being told that the question has changed.
That requires one technical distinction.
Do not say "watermarks are unreliable like AI detectors." The two fail in opposite directions, because they work in fundamentally different ways:
-
AI detectors infer. They estimate authorship from statistical features of the finished text — perplexity, burstiness, word-choice patterns — without observing any embedded provenance signal. Their classifications can therefore mistake naturally occurring writing patterns for evidence of AI generation. In peer-reviewed research, over half of TOEFL essays by non-native English speakers were misclassified as AI-generated by at least one of seven major detectors, with all seven agreeing on roughly one in five (Liang et al. 2023, Patterns, PubMed). Neurodivergent writers have also reported false accusations. Multiple universities, including Vanderbilt and Waterloo, have restricted or disabled these tools after auditing the data.
-
Watermarks embed. A watermark introduces a signal during generation and later tests for that signal. A well-designed watermark may achieve a very low false-positive rate under specified detection thresholds.
So the argument here is not "watermarks don't work." The argument is stronger:
A watermark can be perfectly accurate at detecting what it detects, and still cause the same harm — because what it detects is not what readers will understand it to mean.
"May have been processed by Claude" is accurate. But if a platform turns the mark into a visible label, a reader may hear: "a human did not write this." No platform has yet committed to presenting the mark that way; section VI returns to that uncertainty. The important point is simpler: the signal supports one claim, while people may use it to make another. Institutional use can turn that difference into a verdict.
I call this proxy substitution: a system cannot answer the hard question, so it answers an easier one. An institution then treats the easier answer as if it settled the hard question.
| Substantive question (not mechanically answerable from the artifact alone) | Administrable proxy |
|---|---|
| Who authored which parts of this text? | Is a Claude-generated watermark detectable in this version? |
| Did this student cheat? | Does this text's statistical profile resemble AI output? |
| Is this plagiarized? | Do these strings overlap? |
This does not require malice. The hard questions need context, judgment, and sometimes an investigation. The proxy produces a standardized result. That makes it far easier to administer — and far easier to mistake for an answer.
This argument covers even a technically perfect tool. It does not depend on any bug, any false positive rate, any "they didn't do it well enough."
IV. Same Mechanism, Other Instances
-
AI detectors (inference-based proxy): the cost falls on people whose writing naturally resembles AI-generated statistical profiles — non-native speakers, neurodivergent writers, anyone with a consistent style. The real question "did you cheat?" is replaced with "does your text look human enough?"
One case is worth stating in full, because the facts are the argument. In Matter of Newby v Adelphi Univ. (2026 NY Slip Op 26021, Supreme Court, Nassau County, 28 January 2026), a New York court annulled Adelphi University's academic-integrity finding against Orion Newby as arbitrary and capricious and ordered his academic record expunged. Newby was enrolled in Bridges, Adelphi's own support program for students who self-disclose neurosocial disabilities, and had received help on the essay from a Bridges tutor. His professor's violation report accused him of using generative AI — specifically Grammarly — on the strength of a Turnitin "AI-generated score of 100%." Two detectors his parents then ran, Grammarly's own and zeroGPT, reported "a 0% chance of [being] AI written." The university found him responsible anyway.
Note carefully what this does and does not establish. The court reviewed the university's determination under Article 78 of the CPLR; it did not rule that Turnitin is scientifically invalid. Nor is the disclosure history clean: asked whether he had used Grammarly, Newby first denied it. What the record does show is a misconduct finding resting on a proxy score — over a proofreading tool, against a student using the institution's own accommodations — reversed only after litigation. That is this post's argument with a docket number: 100% against 0% against 0%, and the number that counted was the one the institution could administer.
-
Plagiarism checkers: the real question "is this copied?" is replaced with "do these strings overlap?" The cost falls on citation-heavy writing, formulaic legal language, anyone in a narrow technical vocabulary.
The common problem is not merely that these tools can be wrong. It is that their outputs leave out the context needed to answer the real question. The more complicated the case, the more the proxy erases.
There is a limit to this evidence. The documented detector cases come mostly from English-language higher education in the United States, where audits and litigation have produced public records. That does not mean the harm is greatest there. It means those are the cases we can see.
V. Who Pays
The asymmetry can be derived entirely from Anthropic's own admitted failure cases:
You write something yourself, then ask Claude to proofread it → may retain the mark.
You have Claude draft it, then rewrite it heavily → may lose the mark (Anthropic explicitly lists "paraphrased" as a case where the mark may not survive).
The mark may accurately detect processing. But once people use it to judge authorship or honesty, it rewards concealment and penalizes disclosure. An open collaborator may keep the mark; someone who deliberately rewrites away the evidence may escape it.
It taxes disclosure and exempts concealment.
The deeper risk is not exposure. I already disclose my use of AI. The risk is that an official but crude label will outrank a detailed personal record. My changelog can say: Claude drafted this entry; I revised that one; we later rejected this conclusion. The watermark can say only: processed. A reader may turn that into: not human-written.
A person who states things more precisely gets overwritten by a label that states things more vaguely.
I am not the clearest victim of this system. I have already told everyone. The mark reveals nothing that my about page does not explain in greater detail. The greater cost falls on people whose use seemed too minor to announce: a student who fixed one paragraph, a freelancer who cleaned up a draft, a non-native speaker who asked whether a sentence sounded natural. My ledger is not the wound. It is what makes the wound visible.
The term in this post's title was not mine. One AI collaborator proposed it, and my outline records which one. A watermark on this post would report: processed. My ledger records the contribution.
VI. The Behavioral Endpoint
Supported Claude models are already embedding the watermark. Third-party detection is not yet publicly available — Anthropic says they are "working to" enable it. What happens next is not technically determined: platforms must decide whether and how a detected mark becomes a visible label.
Experiments suggest that AI labels can reduce perceived accuracy or authenticity, even when the content is true or human-made. The evidence that they reduce engagement is less consistent.1 I am therefore not claiming that every label drives every reader away. The narrower risk is this: if platforms display a detected watermark as "AI generated," some readers may reject the work before reading it.
For those readers, the detailed record is not judged and rejected. It is never consulted. The proxy does not merely shape interpretation; it may prevent interpretation from beginning.
Altay and Gilardi found that readers discounted AI-labeled headlines because they assumed the content was fully automated. When readers were told that the AI's role was limited, much of the penalty disappeared. In other words, the missing information mattered. A binary mark cannot supply it; a detailed record can.
VII. People Not in the Model
The binary has no category for the collaboration documented here. A mark is not a finding of misconduct, but it can still be used as one.
This argument does not depend on many people working as I do. My case is useful because it makes the loss unusually easy to see. Four hundred and fifty tagged entries show what the binary label removes. The ledger does not make my claim more deserving than anyone else's; it makes the missing information visible.
VIII. What Would Actually Work
If proxy substitution is the problem, what kind of provenance system avoids it?
The uncomfortable answer is that it would need a record of individual contributions — self-reported, detailed, and open to correction. It would look less like a detector and more like a changelog.
That solution has an obvious weakness. Software can record edits, but it cannot fully decide who contributed an idea or who bears responsibility for a choice. The people doing the work must maintain the account themselves, including a record of their mistakes.
The strongest objection is obvious: someone who wants to hide AI use will not report it honestly. That is true. It is also why institutions want watermarks. They want evidence that does not depend on trusting the person being judged.
There is no clean solution to that problem. But the alternative to a detailed record that cannot be enforced is not a detailed record that can. It is an enforceable answer to the wrong question.
A mark and an account also carry different authority. When they conflict, the machine-readable signal looks like evidence; the person's explanation looks like an excuse. The person must then reconstruct the history that the proxy erased. We should therefore judge a provenance system by two questions: Is its signal accurate? And can a person challenge the conclusion drawn from it?
A voluntary record can still earn credibility, but only if it costs its keeper something. The audit in section I exposed thirteen bad citations and admitted that its AI reviewer repeated the same error twice while writing the audit. A ledger that only exonerates its keeper is worthless. A ledger that records damaging facts gives readers a reason to take it seriously. This does not prevent lying, but it offers something a watermark cannot: a history of error and correction.
Detail matters for more than credit. It also tells us who is responsible. "Processed by Claude" does not say who accepted a claim, rejected an objection, checked a source, or chose to publish. A watermark can reveal that AI touched the work while hiding who made the decisions that mattered.
No automated system can fully describe authorship at this level. The closest account is voluntary, revisable, and maintained by the people doing the work.
The answer is therefore not simply to build a better watermark. No watermark can tell us who contributed what or who is responsible for the result. Yet institutions may still use it that way because it gives them a result they can process.
This leaves another question for a separate essay: why is provenance enforced in only one direction? Claude's output carries a signal tracing it to Anthropic. Its training data carries no comparable signal tracing works to their authors. In one court-documented acquisition program, Anthropic bought millions of print books, many of them used, cut them from their bindings, scanned them into a permanent digital library, and discarded the paper copies.2 The court held that this purchased-book workflow was fair use; Anthropic's later copyright settlement concerned a separate library of pirated digital books. The legal distinction matters. So does the material asymmetry: provenance is added to the output while the source object is destroyed on the way in. That asymmetry deserves more than a closing paragraph.
This post was drafted with assistance from Claude. The arguments, structure, and all editorial decisions are mine. If that arrangement sounds contradictory for a post criticizing Claude's maker, I would point out that recording the contradiction is exactly what the post is about.
Footnotes
-
Altay and Gilardi (2024), PNAS Nexus, found that labeling headlines as AI-generated reduced both perceived accuracy and sharing intentions, including where the label was incorrect — and that the effect was roughly three times smaller than labeling a headline false, which is a useful check against overstating it. Wittenberg et al. (2025) found little or no effect on sharing, liking, commenting, or information-seeking intentions; that disagreement is why the text above says the engagement evidence is less consistent rather than picking the study that suits me. Both concern visible labels on social-media content, not Claude's watermark specifically. ↩
-
Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA, Order on Fair Use (N.D. Cal. June 23, 2025), pp. 6–7, 22–23. The order states that Anthropic purchased millions of print books, often used, destructively scanned them, and destroyed the source copies. ↩