OpenAI will begin adding an invisible watermark to text generated by ChatGPT and Codex in the European Union to comply with the bloc's Artificial Intelligence (AI) Act. The company said it will roll out the feature in the coming weeks.
The AI Act's transparency rules took effect on Aug. 2 and require AI companies to mark AI-generated content so other systems can detect it.
OpenAI said the watermark will reach eligible ChatGPT and Codex users on all plans, but only in the EU.
Developers worldwide who use the OpenAI application programming interface (API) can enable it for certain models starting now.
The feature is off by default on the API, and the company said it will not become a worldwide default in the first phase.
The watermark is not a visible symbol. It slightly changes the model's word choices, creating a pattern that readers cannot see but a dedicated detection system can identify. Because the pattern sits in the words themselves, it remains when text is copied and pasted.
OpenAI said the watermark does not identify the user and that it has not observed a meaningful change in model performance.
Alongside the announcement, OpenAI published a technical report on a method called "textGrain." Researchers from the University of Pennsylvania and Yale University contributed to it.
The report shows how it ranks predictions for the next word in a sentence using a secret key. Hundreds of such small changes together allow the detection system to identify AI-generated text using only the text and the key.
OpenAI's tests show that editing can make the watermark harder to detect. In one test, replacing 10% of the words with synonyms reduced the detection rate from about 92% to 66%. Short texts, answers to math questions and translated texts are also harder to detect.
Because of these limits, only approved researchers and specialist organizations will have access to the detection tool at first. OpenAI said they will help assess the system's reliability and responsible use.
The company also said the absence of a watermark does not prove a text was written by a human. The text may be very short, heavily edited or generated by another company's AI system.
A watermark can show that part of a text was generated or processed by an OpenAI system, but not how much was created, edited or shaped by a human, OpenAI said.