Moran StudioMoran Studio

The situation outweighs the reality: OpenAI has launched the text watermarking technology textGrain.

•3 min read
The situation outweighs the reality: OpenAI has launched the text watermarking technology textGrain.

To comply with the regulatory requirements of the EU AI Act, OpenAI announced the launch of the text watermarking technology textGrain.

In August this year, Claude also added text watermarking to its model directly based on Google DeepMind's SynthID-Text technology, due to EU compliance requirements and its "safety persona."

OpenAI had already developed this technology two years ago. At that time, The Wall Street Journal disclosed internal test data showing that the watermark recognition accuracy based on statistical bias was as high as 99.9%.

However, OpenAI's own internal research findings showed that if watermarking were added, nearly one-third of ChatGPT's core users would immediately churn, so it held off on applying it.

Of course, the premise was that competitors hadn't rolled it out. Now that Claude has rolled it out and the EU also has requirements, it followed suit.

However, it is not a full rollout. The technology is enabled by default only for text output from ChatGPT and Codex within the EU. It is disabled by default for global API customers, with an opt-in option provided. As for the detector, ordinary users can hardly use it, and applications are only open to a small number of approved research institutions and expert organizations.

To dispel users' concerns about whether text watermarking affects reasoning performance, OpenAI specifically used Astra to run a round of benchmarks, proving that coding and reasoning performance barely declined after watermarking was added. In fact, Claude has been running it for so long that everyone is used to it.

Judging from the official data provided, textGrain's watermarking technology is still rather fragile.

As long as the text is slightly rewritten or 10% of its synonyms are replaced, the detection rate will drop from 92% to 66%. Replace 25% of the words, and the detection rate drops straight to 17%. For mathematical derivations or rigorously logical content, it can hardly be detected at all.

The detection methods are not public; in reality, people will still use it however they like. The watermark itself can only verify whether content is AI-generated, not who generated it.

As a result, the watermark itself is not strong enough either. Run it through another model, make a few small edits, and polish it a bit, and the watermark is gone.

How should I put it? This isn't just fragile—it's extremely fragile. It can only be said that its role in fending off regulation outweighs its practical effect.