tech

How Claude's text watermarking works

Future Claude models will generate text that contains a watermark. This is a way of determining the likelihood that Claude was involved in writing the text, and we, along with several other major AI providers, are implementing this change to comply with the EU AI Act.

How Claude's text watermarking works

TL;DR

  • Future Claude models will incorporate text watermarking to identify AI-generated content, complying with the EU AI Act.
  • The watermarking method does not impact the quality, content, or readability of Claude's outputs and is indistinguishable to readers.
  • It functions by altering the source of randomness in word selection, leaving a detectable pattern without adding characters or tokens.
  • Watermarking is not traceable to individual users or organizations and does not incur additional costs.
  • This change aligns with a broader industry trend, as other major AI providers are also implementing watermarking.
  • The method used is a version of Google DeepMind's SynthID-Text, which analyzes word choices to determine the probability of AI generation.
  • Watermarking has limitations, especially with short texts, heavily edited content, factual passages, and code where exact outputs are crucial.
  • Anthropic will offer a watermark detection API, and for images/files, C2PA content credentials will be used.
  • Editing text can circumvent watermarking, but significant alterations may mean the text is no longer considered AI-generated.
  • Watermarks indicate AI involvement (generation or processing) but do not assign ownership or legal responsibility.