AI Labeling

OpenAI to Watermark ChatGPT Text in the EU

AI labeling, OpenAI, ChatGPT, AI watermarking, textGrain, OpenAI ChatGPT watermarking in the EU, ChatGPT AI watermark textGrain
Facebook
X
LinkedIn
Reddit
WhatsApp

The textGrain technology subtly alters the model’s word choices. According to OpenAI, editing the text can significantly reduce the detection rate.

OpenAI will begin adding invisible watermarks to text generated by ChatGPT and Codex in the European Union in the coming weeks. The technology, called textGrain, subtly changes the model’s word choices to create a statistical pattern that can later be detected. The watermark is not visible when text is read or copied. OpenAI announced the move on October 5.

Ad

The technology is intended to comply with Article 50 of the EU AI Act, which requires providers to make AI-generated content machine-readable as such. OpenAI will not introduce the watermark by default outside the EU. Developers can, however, enable it worldwide for supported models through the API starting immediately. OpenAI will initially make the detector available only to approved researchers and specialist organizations. Applications are now open.

OpenAI’s tests show that the watermark is vulnerable to ordinary text editing. For passages containing 400 tokens, the detection rate fell from around 92 percent to 66 percent when 10 percent of the words were replaced with synonyms. When 25 percent of the words were replaced, the detection rate dropped to 17 percent.

Text length also affects detection. At a target false-positive rate of 1 percent, OpenAI detected the watermark in around 80 percent of 200-token responses to psychology questions. For 400-token responses, the rate rose to around 95 percent. Detection was lower for math questions, where the model has less flexibility in its choice of words.

Ad

The absence of a detected watermark does not prove human authorship.

OpenAI

Text generated with OpenAI tools may be too short, edited, or translated to be reliably detected.

Watermark reveals nothing about the user

According to OpenAI, a detected watermark does not reveal who generated the text or provide information about the user’s account, prompt, or conversation. It also does not indicate how much of the final text was written or edited by a human. OpenAI says the watermark does not affect the quality of responses generated by GPT-6 Astra. Benchmark results remained largely unchanged.

(Editorial Team)

Ad

Artikel zu diesem Thema

Weitere Artikel