by Tony Karp

The Truth About Watermarks in AI Text

 - The Truth About Watermarks in AI Text - - art  - photography - by Tony Karp - Discovery Technology - Cinematography - The Godfather - Designing the Future -
WHAT IS A WATERMARK?

A watermark is a visible mark on an object. It is used to prove its source. A mark in a piece of paper that you see when you hold it up to light is where this sort of mark got its name. But it can also be a mark on a piece of fine glasswork or a piece of jewelry. If the watermark is hidden or concealed, it serves no purpose.
WHAT IT IS

An AI text watermark is a hidden statistical pattern embedded in the text that a large language model produces. It is not a visible stamp, not a hidden character, and not a piece of metadata attached to a file. It exists entirely in the sequence of word choices the model makes as it writes. A reader cannot see it, a spell-checker cannot find it, and copying the text into a different application does not strip it away. The watermark is, in the most literal sense, invisible to everyone who does not hold the secret key that created it.

The idea is simple in principle. When an AI model writes a sentence, it does not know the next word in advance. It generates a ranked list of candidates, each weighted by probability, and then picks from among the top options. In many cases, several candidates are equally good. The sentence "The results of the study were quite" could reasonably continue with "surprising," "promising," "clear," or "significant," and the reader would not notice the difference. The watermark lives in exactly these low-stakes moments. It nudges the model's random selection process so that, over the course of a full document, a subtle statistical pattern accumulates in the choices that were made. The pattern is meaningless to anyone reading the text. But to someone who holds the key, the pattern is a signal.

The watermark does not alter the meaning, tone, creativity, or factual accuracy of the text. Anthropic's own testing found no measurable impact on the quality of Claude's output. Google DeepMind tested SynthID by serving watermarked text to a portion of Gemini's live user base and comparing satisfaction ratings against an unwatermarked control group. No difference was found. In a controlled study, human raters comparing watermarked and unwatermarked answers side by side could not tell which was which. The watermark is, by design, imperceptible.

There is a useful analogy. Imagine a game of Monopoly where, instead of rolling dice to determine how many spaces each player moves, the players use the digits of pi starting from some randomly chosen position deep in the sequence. For all practical purposes, the moves are still random. The game plays out the same way. But if someone later examined the record of every move and knew where in pi the sequence started, they could verify that the game used pi rather than dice. The game that used pi is, in a sense, watermarked. It is the same for AI-generated text.

NOT TRUE WATERMARKS — IT'S STEGANOGRAPHY

The industry calls these marks "watermarks." The press calls them "watermarks." The EU AI Act calls them "watermarks." But the term is misleading, and the misleading has consequences. What AI text systems actually do is closer to steganography — a much older, much different discipline — and understanding the distinction matters for understanding why the technology behaves the way it does.

Watermarks are usually visible

Start with what the word has meant for the last seven centuries. A watermark, in every context where ordinary people encounter one, is something you can see. Hold a twenty-dollar bill up to the light and there is Andrew Jackson's face, embedded in the paper fibers during manufacturing. Open a stock photograph on Shutterstock and there is a translucent grid of logos tiled across the image, removable only by purchasing a license. Open a PDF that someone marked "CONFIDENTIAL" or "DRAFT" and there is a faint word stamped diagonally across every page. These are all watermarks. They are all visible. That visibility is not a defect — it is the point. The watermark communicates: this is genuine currency, this image is not licensed, this document is a draft.

Even the original watermarks — the ones that gave the technology its name — were visible. Italian papermakers in the thirteenth century embedded wire-formed designs into their paper molds. When the wet pulp settled around the wire, it left a thinner area in the sheet. Hold the finished paper up to the light and the design appeared — a crown, a cross, the maker's initials. The technique was called watermarking because the mark was created while the paper was still in water. But the mark itself was always meant to be found. It was a mark of origin, a guarantee of quality, a sign that said: this sheet came from our mill.

This visibility has been the defining feature of watermarking across every medium. In photography, visible watermarks prevent unauthorized commercial use. In film and television, network logos in the corner of the screen establish broadcast origin. In software, trial-version watermarks ("UNREGISTERED" printed across the output) motivate purchase. In architecture and engineering, drawing stamps identify the issuing firm. In legal documents, embossed seals and holographic stickers serve the same purpose. The common thread is that the watermark is present, apparent, and communicative. It tells you something about the content by being noticed.

The category of invisible watermarks exists, but it is the exception, not the norm. Digital forensic watermarks embedded in images and audio for copyright tracking are invisible, and they represent the one domain where the term "watermark" has been applied to something undetectable by human senses. But even forensic watermarks differ fundamentally from AI text marks in their defining property: they are robust. A forensic image watermark is designed to survive compression, cropping, color adjustment, printing, scanning, and re-digitization. Robustness is the engineering priority. It is what makes the mark useful — it travels with the content through any reasonable transformation and can be recovered on the other end. AI text marks have no such robustness. Paraphrase the text and they vanish.

What steganography actually is

Steganography is something else entirely. The word comes from the Greek steganσs, meaning covered or concealed, and graphein, meaning writing. It is the practice of hiding a message inside a medium that does not appear to contain a message. The Greek historian Herodotus recorded one of the earliest known instances: a plotter of a revolt shaved a slave's head, tattooed a message on his scalp, waited for the hair to grow back, and sent the slave as an apparently ordinary messenger. The message was invisible. Its existence was hidden. That is steganography.

The practice evolved through centuries. In the Roman era, Pliny the Elder described writing with the milk of the tithymalus plant, which dried invisibly on paper but turned brown when heated. During the Renaissance, Giovanni Battista della Porta taught how to write on the inside of an eggshell using a mixture of alum and vinegar, so the message appeared only when the shell was peeled. In the twentieth century, microdots — photographs reduced to the size of a period on a page — were used by intelligence services on both sides of the Cold War. In every case, the defining principle was the same: the medium looks ordinary, the message is undetectable, and only the intended recipient can extract it.

The critical difference between watermarking and steganography lies in what they optimize for. Watermarking optimizes for robustness — the mark should survive attacks and transformations. Steganography optimizes for undetectability — the hidden signal should be impossible to discover without the key. These are different goals, and they lead to different engineering trade-offs. A robust watermark can tolerate some loss of stealth. A good steganographic system sacrifices robustness to remain invisible.

AI text marks are steganographic systems

By this taxonomy, AI text marks are steganographic systems that have been given the wrong name. They embed a hidden statistical signal in the pattern of word choices. The signal is invisible to any reader who does not hold the key. But the signal is not robust. Change the words and the signal degrades. Paraphrase the text and the signal collapses. Pass the output through another language model and the signal is destroyed. A July 2026 study found that paraphrasing removed the mark in 98.3% to 100% of cases across three different methods. This is the behavior of a steganographic system, not a watermarking system. A true watermark, by definition, survives transformation. These marks do not.

The security researcher Sean Goedecke made this connection directly, calling the problem "basically a text steganography problem" — concealing a secret code in a medium that has no noise margin, no spare pixels, no room for anything a human would not notice. The only reason it works at all for AI text is that the AI's word-selection process provides a channel for embedding the signal. But that channel is fragile: change the words and the channel disappears.

Why the name matters

The mislabeling matters because it sets the wrong expectations. When people hear "watermark," they think of something persistent, something that follows the content wherever it goes, something that cannot be peeled off. They think of the watermark on a hundred-dollar bill, which survives washing, folding, copying, and scanning. They think of the Shutterstock logo, which can only be removed by purchasing the image. AI text marks behave like none of these. They are fragile signals hidden in word choices, detectable only by the key-holder, and erasable by anyone who rephrases the text. If the industry called them what they technically are — embedded steganographic signals — the public conversation about their capabilities and limitations would be more accurate from the start.

There is a second dimension to the mislabeling. Traditional watermarks assert ownership. They say: this content belongs to someone. AI text marks do not assert ownership. They assert process. They say: this content was processed by a particular system. That is a meaningful distinction. The content belongs to the user — Anthropic's terms of service are clear on this point. The mark is not a claim of authorship or intellectual property. It is a record of which tool was involved. Calling it a watermark imports connotations of ownership and rights that do not apply. A more precise term would be a provenance signal or a process mark.

None of this diminishes what the technology actually does. The statistical detection is real. The mathematics is sound. The signal, when present and undisturbed, can be identified with high confidence. But the word "watermark" promises robustness that the technology does not deliver, implies visibility that the technology deliberately avoids, and suggests ownership that the technology does not assert. Precision in naming matters, especially when laws, institutional policies, and individual consequences are being built on top of the name.

WHAT IT DOES

A watermark allows the holder of the secret key to answer one specific question: "Was this text likely processed by our AI model?" That is the full extent of what it does. It provides a probabilistic assessment, not a binary verdict. The longer the passage, the more data points are available, and the higher the confidence. On a short paragraph, the answer may be inconclusive. On a full page, the signal becomes strong. On a multi-page document, it can be quite definitive.

The detection process is mechanical and objective. The detector does not read the text for style, vocabulary, or tone. It does not guess whether the writing sounds like AI. It replays the key-holder's mathematical coloring over the sequence of words and counts how many of them landed on the favored side of the split. Without a watermark, the count should be close to chance, roughly fifty-fifty. With a watermark, the count is skewed. The degree of skew, measured against what chance alone would produce, yields a confidence score.

This is categorically different from what AI detection tools like GPTZero, Turnitin's AI module, or similar services do. Those tools analyze writing patterns, looking for statistical tells in sentence structure, vocabulary distribution, and phrasing habits that tend to differ between human and machine writing. They are making educated guesses based on surface features. A watermark check is not a guess. It is a cryptographic test. It requires the secret key. Without the key, the test cannot be run. With the wrong key, the test returns noise.

In practical terms, the watermark creates a chain of accountability. When Anthropic, Google, or any other provider embeds a watermark, they are establishing a verifiable record that their system was involved in producing or processing the text. This matters for regulatory compliance, for platform trust and safety operations, for content moderation at scale, and for any institutional context where the provenance of a document is relevant. The EU AI Act, which took effect on August 2, 2026, requires exactly this: a machine-readable signal that allows AI-generated content to be identified.

The watermark carries no identifying information about the user. It cannot be traced to a specific person, organization, account, or conversation. It is a property of the model's output process, not a tag placed on a particular user's session. Anthropic has been explicit about this: there is nothing in the watermark or its key that would allow anyone to recover any information about who prompted the text.

WHAT IT DOESN'T DO

The limitations of AI text watermarking are substantial, and understanding them is at least as important as understanding how the technology works.

It does not prove authorship

A detected watermark means "this text was likely processed by Claude" (or Gemini, or whichever model's key is being checked). It does not mean "Claude wrote this from scratch." If a person writes an essay and asks Claude to proofread it, the returned text may carry a watermark on the handful of words Claude changed. If a person writes a draft and asks Claude to rewrite it entirely, the result will be heavily watermarked. The watermark cannot distinguish between these cases. Anthropic states this plainly: the watermark can only determine that Claude was "likely involved with the content at some point."

It does not work well on short text

The watermark accumulates its signal through volume. Each word choice is a single data point, and many of those data points carry no signal at all because only one word was a reasonable choice. A two-sentence reply, a short email, a haiku — these do not provide enough material for the statistical test to reach confidence. Detection becomes reliable only as length increases.

It does not work on low-entropy content

Entropy, in this context, means the degree of freedom the model has in choosing words. When the model writes "Isaac Newton's most famous work was called Principia" followed by the only correct next word, "Mathematica," the watermark has nothing to act on. Code is similar: variable names, function calls, syntax, and logical structure are all highly constrained. Lists of facts, mathematical proofs, quotations, and any content where precision requires a specific word rather than a choice among equals will carry little or no watermark signal.

It does not survive heavy editing

The watermark lives in the specific sequence of words the model chose. Change those words and the signal degrades. Light editing generally leaves most of the watermark intact. But a thorough rewrite, where every sentence is restructured and every word is replaced with a synonym, destroys the pattern. The watermark does not live in the meaning of the text. It lives in the exact words.

Research confirms this vulnerability. A July 2026 empirical study found that across 846 paraphrase attempts on three different watermarking methods, every KGW-watermarked and Unigram-watermarked text lost its watermark after paraphrasing, a 100% conditional removal rate. SynthID performed only marginally better at 98.3% removal. A free watermark-stripping tool appeared on GitHub within 24 hours of Anthropic's announcement and gained roughly 72 stars per day.

It does not identify which AI produced the text

Each provider uses its own key. Anthropic's key can only check for Claude's involvement. Google's key can only check for Gemini's involvement. If a piece of text was generated by an open-source model running on someone's laptop, no provider's watermark check will return a positive result. A negative result does not mean the text was written by a human.

It does not prevent misuse

A watermark is a detection tool, not a prevention tool. It does not stop anyone from using AI-generated text. It does not block the output. It does not impose consequences. It creates the possibility that someone, at some later point, might check.

It can be spoofed

An attacker who can query a detection service repeatedly can, in principle, learn enough about the watermarking pattern to forge a positive result on human-written text. This is the spoofing attack. Researchers at the University of Maryland demonstrated this, showing that an attacker could infer hidden signatures without white-box access and create false positives. A false negative lets misconduct pass, but a false positive punishes an innocent person.

WHY WE NEED THEM

The case for AI text watermarks rests on a problem that barely existed three years ago: the internet is filling up with synthetic text, and nobody can tell it apart from the real thing.

Before large language models, the volume of text on the internet was bounded by the number of people willing to write it. That constraint is gone. A single person with access to an AI model can now produce, in an afternoon, more text than a team of writers could generate in a month. The text is grammatically correct, topically relevant, and stylistically unremarkable. It reads fine. It just was not written by a person with knowledge, experience, or a stake in whether the claims are true.

The consequences are already visible. AI-generated news sites, built to attract advertising clicks, have been identified by researchers in the hundreds. SEO content farms use AI to flood search results with competent but hollow articles. AI-generated product reviews, academic papers with fabricated citations, synthetic social media comments deployed in coordinated campaigns, AI-generated emails used in phishing and fraud — the category of "content that looks real but is not what it appears to be" has expanded enormously, and the tools for identifying it have not kept pace.

Style-based AI detection tools tried to fill this gap and largely failed. A University of California, Davis student named Orion Newby became the first person in the United States to win a federal lawsuit over a false AI plagiarism accusation after a professor relied on a detection tool that flagged his human-written work as AI-generated. That case, decided in February 2026, illustrated the core problem: tools that guess from surface patterns produce false positives, and false positives destroy lives.

Watermarks offer something fundamentally different. They do not guess. They check for a specific, deliberately embedded signal. A false positive on a watermark check requires the human-written text to have, by sheer coincidence, reproduced the exact statistical pattern that the secret key would produce. The probability of this is astronomically low. This is the single strongest argument for watermarks: they trade the unreliable guesswork of style-based detection for the mathematical precision of a cryptographic test.

There is also the training data problem. As AI-generated text floods the internet, future models increasingly risk training on the output of earlier models. Researchers have shown that this recursive training leads to model collapse — a progressive degradation of quality and diversity. Watermarks provide a mechanism for identifying AI-generated text in training corpora so it can be filtered out.

The regulatory argument is real. The EU AI Act requires providers to embed machine-readable marks. Watermarks are the mechanism the law requires. Without them, compliance is not possible.

Finally, there is the accountability argument. AI systems are being used to produce text that influences decisions. When that text turns out to be wrong, misleading, or harmful, the question "who produced this?" becomes urgent. A watermark establishes which system was involved. That is the first link in a chain of accountability that, without watermarks, does not exist at all.

WHY WE DON'T NEED THEM

The case against AI text watermarks is not a fringe position. It has been articulated by security researchers, legal scholars, AI safety advocates, civil liberties organizations, industry analysts, and, privately, by at least one major AI company that built the technology and decided not to use it. The arguments are distinct from the practical limitations described earlier. Those limitations describe what watermarks fail to do. The arguments here describe why attempting to do it at all may be the wrong approach.

They punish the compliant and miss everyone else

The most fundamental objection is structural. AI text watermarks can only be embedded by providers who choose to embed them. As of August 2026, the providers who have committed to watermarking — Anthropic, Google, and those who signed the EU Code of Practice — are the same providers who already have acceptable use policies, safety teams, content moderation, and terms of service. The providers who do not watermark — open-source models running on private hardware, models served from jurisdictions that do not enforce EU law, Chinese and Russian systems, fine-tuned variants of open-weight models — are the same providers that operate outside any governance framework.

This creates a two-tier system. Content from compliant providers carries a mark. Content from non-compliant providers carries no mark. The mark does not distinguish between responsible use and irresponsible use — it distinguishes between which tool was used. A thoughtful professional using Claude to draft a report gets marked. A disinformation operation using a stripped open-source model to generate thousands of fake news articles does not. The watermark functions as a compliance tax on the responsible while leaving the irresponsible untouched.

The Center for Data Innovation made this argument directly, calling the EU AI Act's watermarking requirement "a misstep" that risks confusing consumers and "detracting from other efforts to address misinformation and content provenance." The Electronic Frontier Foundation argued that watermarking AI-generated content "won't curb disinformation" because the actors producing disinformation are precisely the ones who will not comply.

They create false confidence in detection

Watermarks are frequently presented as a solution to the AI detection problem. This framing is dangerous because it implies that watermark detection can reliably answer the question "was this written by AI?" It cannot. It can only answer "was this processed by a specific model that embeds watermarks?" A negative result means nothing definitive. The text may have been generated by an unwatermarked model. It may have been watermarked and then paraphrased. It may have been generated by a watermarked model but be too short to carry a detectable signal. The absence of a watermark is not evidence of human authorship.

The RAND Corporation warned that this asymmetry — where the presence of a mark is informative but the absence of a mark is not — creates conditions where "media consumers may be overwhelmed with false negatives" and where "the watermarking algorithm itself can become a vehicle for novel and powerful attacks against the information ecosystem." If institutions treat watermark detection as a reliable screening tool, they will make systematic errors in one direction: clearing all unwatermarked AI content while flagging some human content that incidentally passed through a watermarked system.

They solve a problem that disclosure solves better

The premise of watermarking is that people will not voluntarily disclose their AI use, so the technology must do it for them. But in the contexts where disclosure matters most — professional work, academic submissions, legal filings, journalism — the appropriate mechanism is already policy, not technology. Employers can require AI use disclosure. Universities can establish AI use policies. Journals can require authors to declare AI assistance. Courts can require affidavits about the provenance of submitted documents. These mechanisms are imperfect, but they place the obligation where it belongs: on the person, not on the tool.

Watermarks bypass the human decision entirely. They mark the output regardless of context, regardless of how the output will be used, regardless of whether disclosure is appropriate or relevant. A novelist using Claude to workshop dialogue gets the same mark as a student submitting a generated essay as their own work. A lawyer brainstorming arguments gets the same mark as a scammer drafting phishing emails. The watermark cannot distinguish between these cases because it has no knowledge of the context. Disclosure policies can make those distinctions. Watermarks cannot.

They entrench provider power

Detection requires the provider's key. The provider controls the key. This means the provider is the sole arbiter of whether a piece of text was processed by their system. There is no independent verification, no third-party audit, no way for the public to confirm or challenge a detection result. The watermark, in practice, is a system in which a small number of private companies control the mechanism for determining whether content is AI-generated.

A researcher writing under the title "The Monopoly of Truth" argued that this structure creates a dangerous concentration of epistemic authority: the companies that build AI systems also control the tools for identifying AI output, and they operate both sides without independent oversight. If detection APIs become integrated into hiring processes, content moderation pipelines, academic integrity systems, and legal proceedings, the providers will have inserted themselves into a gatekeeping role over the authenticity of written communication. That is a significant expansion of corporate power over public discourse, and it happened not through any democratic process but through a regulatory mandate that the providers themselves helped shape.

They chill legitimate use

The chilling effect is not theoretical. Within days of Anthropic's watermarking announcement, users reported canceling subscriptions, switching to unwatermarked competitors, and restructuring workflows to avoid Claude's output. Business Insider documented multiple cancellations from paying subscribers who objected not to the watermark's technical impact on quality — which is negligible — but to the principle that their tool was now marking their output without their consent.

The deeper concern is that watermarks stigmatize AI use at exactly the moment when AI is becoming a standard professional tool. Calculators did not stamp "CALCULATOR-ASSISTED" on their output. Spell-checkers did not watermark corrected text. Search engines did not mark the documents people found through them. The implicit message of a watermark is that AI-assisted writing is suspect — that it requires a mark, a flag, a warning. For writers, researchers, professionals, and creative workers who use AI as one tool among many, this framing is both inaccurate and counterproductive. It treats the tool as the author and the human as the bystander, when the reality for most users is the opposite.

The problem may not need a technical solution

Not every problem requires a technical fix. The presence of synthetic text on the internet is a real issue, but it may be better addressed through media literacy, institutional policies, content provenance standards at the platform level, and legal frameworks that target harmful use rather than AI use per se. The C2PA standard for images and files provides provenance without embedding anything in the content itself — it attaches metadata to files, which is removable but also verifiable. A similar approach for text — where AI providers publish verifiable receipts for content they generate, without altering the content — has been proposed by multiple researchers as an alternative that avoids the fragility, chilling effects, and power concentration problems of embedded watermarks.

The question is not whether AI-generated text is a problem. It is. The question is whether embedding a fragile steganographic signal in every sentence a model produces is the right response to that problem, or whether it is a technically impressive solution to the wrong question — one that creates new problems while leaving the original problem largely unsolved.

HOW IT WORKS

Every watermarking scheme in production today operates on the same fundamental principle: the watermark changes the source of randomness used to select words, without changing the range of words the model considers.

The green list / red list method

The foundational approach, published by John Kirchenbauer, Jonas Geiping, Yuxin Wen, and colleagues in 2023, works as follows. At each position in the text where the model is choosing the next word, a hash function — seeded by the secret key and a short window of the preceding words — splits the model's entire vocabulary into two groups, conventionally called the green list and the red list. This split is arbitrary and invisible; it has nothing to do with the words' meanings.

With the split in place, the model's sampling process is gently tilted toward green-list words. The tilt is small. A red-list word can still be chosen. But across hundreds of word choices, the cumulative effect is that green-list words appear slightly more often than chance would predict. This excess is the watermark.

Detection replays the hash function at each word position, re-creates the green/red split, and counts how many actual words fell on the green side. If the count is significantly above the 50% baseline, the text is flagged as watermarked.

SynthID: tournament sampling

Google DeepMind's SynthID-Text, published in Nature in 2024, uses what its creators call tournament sampling. At each word position, several candidate words are drawn from the model's probability distribution. The secret key scores each candidate through a pseudorandom function, and a multi-round tournament selects the winner. Every word's expected selection probability remains exactly what the model originally intended. The watermark introduces no bias in expectation, only a correlation between word choices and the key.

SynthID was open-sourced through Hugging Face's Transformers library in October 2024. Anthropic's implementation for Claude is described as "a version of the SynthID-Text approach."

The Aaronson scheme

Scott Aaronson, working at OpenAI beginning in late 2022, proposed replacing the random number generator used to sample the next word with a deterministic function of the key and the recent word history. Detection works by checking whether the observed sequence of draws is consistent with the key-derived randomness. This approach established the theoretical framework that all subsequent text watermarking schemes built on.

What all three share

The common thread is that the watermark acts only at the moment of word selection, and only in cases where the model has genuine freedom to choose among multiple adequate options. Where only one word is correct, the watermark has nothing to work with. The mark is a property of choices, not of content. No specific word is "the watermark." The watermark is the statistical pattern formed by the ensemble of choices across the full document.

EXAMPLES OF USING WATERMARKS

Regulatory compliance

The immediate catalyst for widespread watermark adoption is Article 50 of the EU AI Act, which became enforceable on August 2, 2026. The law requires providers of generative AI systems serving the EU market to embed machine-readable marks in their outputs. Non-compliance carries fines of up to 15 million euros or 3% of global annual turnover. Roughly 190 organizations signed the EU's Code of Practice. Anthropic chose to apply watermarking globally rather than restricting it to EU users.

Platform content moderation

Social media platforms, news aggregators, and content hosting services can use watermark detection as one signal in their moderation pipelines. A watermark check provides a fast, cheap, and objective first pass that is not based on stylistic guessing.

Academic integrity

Universities have struggled with AI-generated submissions since late 2022. A watermark-based check would be more reliable than a style-based detector because it tests for a specific embedded signal. However, its utility depends on whether the student used a watermarked model, whether the text is long enough, and whether the student paraphrased the output enough to destroy the signal.

Journalism and media provenance

News organizations can use watermark checks to verify whether submitted content was produced by AI. Combined with C2PA metadata on images and files, watermarks form part of a layered provenance system.

Legal and forensic use

In legal proceedings, a watermark check could provide admissible evidence of AI involvement. The evidentiary weight remains untested in most jurisdictions, but the technical foundation is more rigorous than any alternative.

Enterprise content governance

Organizations can use a watermark detection API to audit content pipelines without requiring employees to self-report their AI usage. This is not surveillance of the individual, but an aggregate measurement of AI involvement in organizational output.

Protecting training data

Watermarks give developers a way to identify synthetic text in training corpora and filter it out, preserving the integrity of future training data. Some commentators have suggested this is the most concrete practical motivation behind watermarking.

HOW TO DETECT WATERMARKS

Detection of a true statistical watermark requires the secret key. There is no general-purpose tool that can detect any provider's watermark without cooperation from the provider.

Provider-operated detection

Each AI company that implements watermarking controls the key, and therefore controls detection. Google operates an early-access detection portal for SynthID. Anthropic has announced that a watermark detection API is forthcoming.

What detection checks for

The detector replays the key-seeded hash function across the text and counts the proportion of words that fall on the green side. A passage showing 55% green-list words over 1,500 words would be statistically significant; the same 55% over 50 words would not. The detector outputs one of three verdicts: watermarked, uncertain, or not watermarked.

What detection does not check for

The detector does not analyze writing style. It does not look for AI "tells." It does not compare the text against a database of known AI outputs. It is purely a mathematical test applied to the word sequence using the key.

The distinction from AI detection tools

AI detection tools like GPTZero, Originality.ai, and Turnitin's AI detection module operate without any key. They are probabilistic guesses based on statistical features. A watermark check looks for a specific, deliberate signal that was intentionally embedded. The two approaches answer different questions and should not be treated as interchangeable.

Unicode artifact detection

Some AI outputs contain unusual Unicode characters — invisible spaces, zero-width joiners, homoglyph substitutions. These are not watermarks in the SynthID sense. They are either model artifacts or a crude form of text tagging. This is a separate phenomenon from statistical watermarking. Stripping unusual Unicode characters is trivial.

WHAT COULD POSSIBLY GO WRONG?

The technology works as described. The statistical foundations are sound. The intentions are, for the most part, reasonable. And yet the history of well-intentioned technical systems deployed into messy human institutions is not encouraging.

The accusation problem

A watermark detection result says "this text was likely processed by an AI model." It does not say "this person cheated." But the distance between a detection result and an accusation is shorter than it should be. Since 2023, students have been expelled, suspended, and given failing grades on the basis of AI detection tools known to be unreliable. Non-native English speakers have been disproportionately flagged. The institutional reflex — tool says positive, therefore the student cheated — will not become more careful just because the tool has improved.

The scenario that should concern everyone is the person who writes their own work, asks Claude to check the grammar, and ends up with a document that carries a faint watermark signal. In most cases, the signal will be too weak to trigger detection. But "most cases" is not "all cases."

The provider holds all the cards

The provider embeds the watermark, operates the detector, and defines what the results mean. There is no independent verification. A 2026 paper titled "Watermarking Without Standards Is Not AI Governance" argued that without independent verification, watermarking will remain inadequate for accountability. A July 2026 paper drew a parallel to forensic science failures, warning that deploying watermark evidence without meeting Daubert standards risks "repeating the failures documented by the NAS and PCAST reports. Those failures were measured in wrongful convictions."

The arms race

A watermark-stripping tool appeared on GitHub within 24 hours of Anthropic's announcement. The pattern is the same as every previous attempt to embed persistent marks in digital content: the marks persist for people who do not care enough to remove them, and they are trivially defeated by anyone who does. Detection accuracy drops from 95% to 60–70% under adversarial conditions, and open-source models bypass the system entirely.

The practical result is a system that catches the careless and misses the deliberate, which is exactly backwards from the threat model that motivated the regulation.

The chilling effect

People have already changed their behavior. Claude Max subscribers canceled after the watermark announcement. Writers, journalists, and professionals who use AI as a thinking tool may hesitate if they know the output carries a mark — not because they are doing anything wrong, but because the mark creates a risk of misinterpretation by someone else down the line.

The fragmentation problem

Each provider uses its own key. There is no interoperability. A negative result from all providers does not mean the text was human-written. It means it was not detectably written by any of the specific models whose keys were checked. The gap between that technical statement and the conclusion people will draw from it is where the harm lives.

The liar's dividend

Once watermark detection becomes widespread, anyone accused of wrongdoing based on a document can claim the document was AI-generated. Conversely, anyone who did use AI can point to the absence of a watermark as evidence of innocence — knowing that the watermark may have been stripped. The watermark provides a new dimension of plausible deniability in both directions.

The authoritarian use case

In democratic societies, watermark detection is constrained by law and norms. In authoritarian societies, the same tool serves a different purpose. If a government can check whether a dissident's manifesto was produced with a specific AI model, it gains an additional vector for surveillance. The detection API, when it exists, will be a service that someone queries. In jurisdictions where the government monitors API traffic, the act of checking is itself informative.

WHAT IT MEANS FOR SOMEONE WHO USES CLAUDE EVERY DAY

If you are a person who uses Claude regularly — for writing, research, brainstorming, editing, coding, or any of the hundred other things people use it for — the watermark changes less than you might fear and more than you might expect.

Your text quality has not changed

The watermark does not make Claude's output worse. Internal testing at Anthropic and external testing by Google DeepMind found no measurable impact on output quality. When human raters compared watermarked and unwatermarked text side by side, they could not tell the difference.

Your conversations are not being tracked

The watermark carries no identifying information. It does not encode your name, your organization, your account ID, or any characteristic of your session. If someone detects the watermark, all they learn is that Claude was involved.

Your own writing is mostly safe

If you write a document yourself and ask Claude to proofread it, the watermark can only attach to the words Claude changed. A lightly proofread document will not reach the detection threshold. Anthropic has acknowledged this: "Because nearly all the words are the person's, there's very little (if anything) for the watermark to attach to."

The situation changes as Claude's involvement increases. If you give Claude a rough outline and ask it to write the full draft, that draft is heavily watermarked. The watermark tracks word selection, not intellectual contribution.

The real question is disclosure, not detection

The watermark has shifted the anxiety around AI use from "can anyone tell?" to "does it matter if they can?" The people most affected are those in the gap between acceptable use and plausible deniability. The solution is not watermark removal. It is a clearer personal policy about when and how you use AI, and honest communication about it.

Code is mostly unaffected

Code is highly constrained. Anthropic has said the watermark will have "a negligible effect" on code. If you copy a function Claude wrote and paste it into your codebase, the watermark signal is likely too sparse to detect.

Editing is your friend, and it always was

If you use Claude as a drafting partner — generating a first pass that you then substantially rewrite — the watermark will degrade naturally as you edit. This creates a reasonable alignment: text that you genuinely wrote carries little or no detectable watermark. Text that Claude wrote entirely carries a strong watermark. The watermark, imperfectly but meaningfully, correlates with the degree of AI authorship in the final product.

The landscape is still forming

Anthropic has not yet released its detection API. Nobody outside the company can currently run a watermark check on Claude's output. The EU AI Act prohibits deliberately removing AI markings. For the moment, the practical advice is straightforward. Use Claude the way you have been using it. The watermark does not change the tool's capabilities, does not change your rights to the output, and does not identify you.

A BRIEF HISTORY

The deep roots of watermarking

Watermarking predates computers by centuries. Physical watermarks on paper date to thirteenth-century Italy. Digital watermarking emerged in the late 1980s. Digimarc, founded in 1995, released a watermarking plugin bundled with Adobe Photoshop in 1996. Throughout the late 1990s and 2000s, digital watermarking was applied to images, audio, and video for copyright protection.

Text was the exception

Text, alone among digital media, resisted watermarking. It has no noise margin. Every character is significant. Early academic attempts relied on synonym substitution, whitespace manipulation, and font micro-adjustments. None survived casual scrutiny. Text watermarking remained, for decades, an unsolved problem.

The LLM inflection point

What changed was the realization that an LLM's token-selection process provided a new surface for embedding a signal. Scott Aaronson proposed his scheme at Harvard in November 2022, two weeks before ChatGPT launched. Kirchenbauer et al. published the green list / red list method at ICML in 2023. Google DeepMind published SynthID-Text in Nature in 2024 and open-sourced it through Hugging Face.

The regulatory push

Biden's Executive Order 14110 in October 2023 directed NIST to develop watermarking guidance. The Trump administration revoked it in January 2025. The EU continued on its course: the AI Act's transparency obligations became enforceable on August 2, 2026, and roughly 190 signatories signed the Code of Practice.

The OpenAI hesitation

OpenAI built watermarking technology but chose not to deploy it. An April 2023 internal survey found nearly a third of loyal ChatGPT users would be turned off by watermarking. As of mid-2026, OpenAI had signed the EU Code of Practice but had not publicly deployed a text watermark.

Anthropic's entry

Anthropic announced on August 11, 2026, that Claude would embed text watermarks and attach C2PA provenance metadata to generated files. The announcement triggered user debate, subscription cancellations, and the rapid appearance of watermark-removal tools.

The parallel track: C2PA for files

C2PA embeds cryptographically signed metadata into media files declaring their creation history. It lives in the file's metadata and can be stripped by saving without the metadata. C2PA applies to files, not plain text. It is a complementary technology, not a substitute.

NOTES

The state of AI text watermarking in August 2026 is best understood as a regulatory default, not a technical guarantee. The marks are real. The statistical foundation is sound. The detection, when performed with the key by the provider, works. But the marks are fragile against determined editing, undetectable on short passages, ineffective on low-entropy content, and trivially removed by anyone with access to a paraphrasing tool.

This does not mean watermarks are meaningless. They serve as a speed bump that catches low-effort AI slop. They create a regulatory paper trail that satisfies the EU's transparency requirements. They provide a technical foundation on which future systems might be built. And they establish a norm: AI-generated content should, as a default, carry some mark of its origin.

The analogy to digital rights management is instructive. DRM has been broken within days of every new implementation for two decades. It has never stopped a determined pirate. And it is still standard practice, because the institutional and legal weight sits behind the mark rather than the mark's technical resilience. AI watermarks appear to be following the same trajectory: technically fragile, institutionally durable, and now legally mandated across the world's largest regulatory jurisdiction.

The most important thing to understand about AI text watermarks is not how they work or how to detect them. It is what they mean. They represent the beginning of a provenance infrastructure for AI-generated content — an imperfect, early, and contested beginning, but a beginning nonetheless. Whether that infrastructure evolves into something robust or remains a regulatory checkbox is the question that the next few years will answer.

Sources consulted include Anthropic's official documentation, Google DeepMind's SynthID-Text Nature paper (2024),

Kirchenbauer et al. (ICML 2023), Scott Aaronson's published talks, the EU AI Act and Code of Practice, NIST AI 100-4,

the RAND Corporation, the Electronic Frontier Foundation, the Center for Data Innovation, the Brookings Institution,

and reporting from Ars Technica, Forbes, The Verge, TechCrunch, Business Insider, MIT Technology Review, and others.

Copyright 1956-2026 Tony & Marilyn Karp