Anthropic has described plans to use invisible marking for text from supported Claude models through an approach related to Google DeepMind’s SynthID-Text, according to reporting on the company’s transparency plans. The aim is not to put a visible “AI-generated” notice beside every sentence, nor to insert hidden characters that can be found by inspecting a document. Instead, the system is intended to leave a statistical pattern in ordinary word choices that a detector can test over a sufficiently long passage.
Reports have differed on whether the capability is already active across particular Claude models and surfaces. The important qualification is that the marking is described as planned or available for supported models and outputs, rather than as a guarantee that every Claude response, legacy model or product surface carries the signal. That scope matters when interpreting either a detection result or an absence of one.
That distinction also matters because a text watermark is a provenance signal, not a definitive account of how a document was made. A positive result is meant to support the conclusion that Claude was involved in processing the text. It does not establish that Claude wrote every word, identify the person or organization that used Claude, or prove that all underlying ideas and source material originated with the model. Just as importantly, an absent signal does not demonstrate that a passage was written entirely by a human or without AI assistance.
The basic idea: biasing plausible next words
Large language models produce text one token at a time. For any point in a sentence, the model assigns different likelihoods to possible next tokens. A watermarking system can work inside that selection process: rather than forcing an obviously unnatural word, it influences the model’s choice among several alternatives that are already plausible in context.
Google DeepMind describes SynthID for text in similar terms. Its system adjusts token probability scores during generation, so the final sequence of choices carries a pattern that can be compared with the pattern expected from marked or unmarked text. To a reader, the output should still read normally; the signal exists in the aggregate distribution of choices rather than in a visible string of characters. Anthropic’s reported plan is a version of that general approach, not necessarily an identical implementation. (deepmind.google)
The practical implication is that there is no single “watermarked word” to highlight. Detection depends on many choices across a passage. A detector with the appropriate method or key looks for whether those choices collectively are more consistent with a marked generation process than an unmarked one. The conclusion is therefore probabilistic: it concerns the likelihood of a provenance signal, not a categorical forensic reconstruction of a document’s history.
Why longer, freer-form writing is more useful
Statistical signals need enough material to accumulate. Reports on Anthropic’s plan say detection can be weaker for short passages and for highly constrained material, including factual text and code. In those settings, a model has fewer genuinely interchangeable next-word options. A technical specification, a code snippet, a quotation or a short answer may provide little room to favor one reasonable alternative over another without risking accuracy or usability.
That is why a watermark should not be understood as an all-purpose classifier that can reliably sort every sentence on the internet into “human” and “AI.” It is more suited to evaluating passages where a supported model had sufficient latitude to generate varied language. Google DeepMind similarly says its text system works best with longer and more diverse responses, such as essays, scripts or alternate versions of an email. (deepmind.google)
Rewriting also changes the equation. Reports on Anthropic’s plan specifically acknowledge that extensive rewriting can remove the signal. That makes the system potentially useful for preserving an initial provenance clue, but not a guarantee that every downstream copy will remain attributable.
Independent research reinforces the need for that caution. A 2025 evaluation published through the Association for Computational Linguistics found that attack robustness remains a central underexplored question across LLM watermarking methods. Separately, a 2025 arXiv preprint assessing SynthID-Text reported that meaning-preserving changes, including paraphrasing and back-translation, could substantially reduce detectability. Neither study measures Anthropic’s undisclosed implementation, and the latter is a preprint rather than peer-reviewed research, but both illustrate why text-watermarking claims should be read as bounded evidence rather than a promise of permanent traceability. (aclanthology.org)
Detection is not authorship detection
The most important operational distinction is between watermark detection and conventional AI-writing detection. A conventional detector examines prose after the fact and tries to infer whether its style resembles machine-generated language. Such systems do not need the model provider to have embedded a signal at generation time.
A generation-time watermark takes a different route. The provider deliberately introduces a pattern while text is produced, then checks for that pattern later. In principle, that makes the result more directly tied to a specific participating generation system. But it also narrows the scope: it cannot identify content made by systems that do not use the same marking method, and it cannot turn a missing mark into proof of human authorship.
For teams building review, publishing or compliance workflows, the appropriate question is therefore not simply, “Is this AI?” A more precise question is, “Does this content contain a detectable provenance signal associated with a supported system, and what does that signal actually justify us in concluding?” A positive result may be useful context for disclosure, moderation or audit processes. It should not, by itself, settle disputes about authorship, originality, intent or policy violations.
Why C2PA is a separate part of the plan
Anthropic is also pairing its text-watermarking plan with C2PA signed provenance metadata for supported PNG, JPG and SVG files, according to reports. That is a different mechanism operating on a different kind of content. C2PA, short for the Coalition for Content Provenance and Authenticity, uses signed metadata to carry information about a media file’s origin and history. The reported plan does not mean that Claude text itself receives C2PA credentials.
The distinction is useful because the two approaches have different strengths and failure modes. Metadata can provide richer, signed context when it remains attached to a file, while an invisible watermark attempts to keep a signal within the generated content itself. Text watermarks, meanwhile, can weaken when wording is substantially changed. In May, OpenAI described a similar layered approach for supported images, combining C2PA metadata with Google’s SynthID watermarking rather than relying on either mechanism alone. That example documents one company’s approach; it does not by itself establish that the entire industry has adopted the same model. (openai.com)
The EU AI Act connection
Anthropic has linked the effort to transparency obligations under the EU AI Act, as reported. Article 50 says providers of AI systems that generate or manipulate synthetic audio, image, video or text should ensure outputs are marked in a machine-readable format and detectable as artificially generated or manipulated, as far as technically feasible. It also frames the technical solution around effectiveness, interoperability, robustness and reliability while recognizing the specific limitations of different types of content. (eur-lex.europa.eu)
That language helps explain both Anthropic’s move and the restraint needed in describing it. A token-level text watermark offers a machine-readable route to detection without changing the visible reading experience. Yet the limitations around short, constrained and heavily rewritten material show why compliance-oriented marking cannot honestly be framed as a perfect or universal answer. The practical value lies in adding a verifiable signal where the system can support one, alongside clear disclosure of what detection—and non-detection—does and does not mean.
For users and organizations, the near-term lesson is straightforward: invisible text watermarks may become an important part of generative-AI provenance, especially as providers respond to transparency rules. But they are best treated as one evidence layer in a broader system of labels, metadata, platform policies and human judgment—not as a conclusive test for whether a piece of writing is “really” human or AI-made.




