4 ms·
So, it seems the watermark is less what it sounds like (a stamp) and more an "imperceptible statistical pattern woven into the choice of words and sentence stru
by taurusnoises 1mo ago
So, it seems the watermark is less what it sounds like (a stamp) and more an "imperceptible statistical pattern woven into the choice of words and sentence structures." Does that mean AI responses will sound even more "AI?" Like, will it become even easier to detect on a read-through because of the word choices and patterning? I see the word "imperceptible" there, but what does this mean in this context? My non-tech brain is kinda breaking here.
- iamacyborg 1mo agoIs it possible for them to sound even more AI?
- dgellow 1mo agoDon’t forget the survivor bias is at play, you don’t see the ones that are good at passing for human written
- iamacyborg 1mo agoSure, but if my experience (I know, I know) is anything to go by, it's increasingly difficult to get Claude to write in anything other than it's own house style.
- 317070 1mo agohttps://arxiv.org/html/2510.20075v6 https://arxiv.org/html/2510.20075v6 It is quite counterintuitive, but you can hide texts the same size as the original text in imperceptible statistics of a text. Compared to that feat, hiding a watermark is very easy.
- joenot443 1mo agoThe information being encoded (the watermark) is the _relative ranking of each token compared to other possibilities_. If our prompt was "Write a positive review for a restaurant" and the response began: "The restaurant " Our next set of predictions might be: [was, had, offers] So we append the rank/index of the next token (0, 1, or 2) onto the secret. Given a long enough response, that secret becomes unique enough to use as a watermark. This obviously relies on having full deterministic access to the LLM itself, i.e. I don't believe it will be possible for users to derive the fingerprint from text that they've generated, only Anthropic will be able to. The immediate objection is that this runs the risk of degrading the quality of the response. I think that's totally valid and I'll be curious how Anthropic handles it. That's my very rough understanding! If someone with more knowledge wants to expand, feel free.
- ozgung 1mo agoDo we need the original prompt to recover the watermark?
- TheOtherHobbes 1mo agoSupposedly you can avoid degradation by using synonyms. But not all words have synonyms. The more concrete and factual the prose, the harder it is to watermark. "Cow" is not a synonym for "cat" and "dark matter" is not a synonym for "galaxy." So the watermark words will be biased towards filler and fluff where invisible substitutions are easier, and the content is less (cough...) load-bearing. The likely outcome is the development of AI watermark strippers which filter out all the twitches and tells that make default AI writing so annoying. Google seem to have given up on SynthID for text for now, so this is likely a harder problem than it looks. My guess is Anthropic announced this to meet regulatory requirements. But they don't have a robust detector, and I seriously doubt they have a robust system that can survive trivial rewriting by a different model.
- dgellow 1mo agoUnlikely, but I really hope so!
- josephg 1mo ago> Like, will it become even easier to detect on a read-through because of the word choices and patterning? It's turned on right now. Can you tell a difference? I can't. How many ways could I write this paragraph and still convey the same idea? Way more than we're aware of. Hundreds? Thousands? Maybe a lot more? The number of semantically similar variants increases exponentially with each word. I suspect anthropic could turn their fingerprinting up or down if they want. If it were turned way up, claude would use weird phrasing but it would take very little text to tell if something were AI generated. If they turned it down, it would seem imperceptible to humans, but you would need a large sample to determine (with high accuracy) that a passage was AI generated. There's probably a very large middle ground where humans can't tell, and where it doesn't take a large text sample to know (with high probability) that some text was AI generated.
- hemkeshr 1mo ago[dead]
- floppydive 1mo agoI think its somewhat the opposite of easy to detect patterns in the watermarked text. Regular authors can be statistically fingerprinted and I think we do a rough version of this ourselves. An LLM with dense watermarking may sound less like one author we have a low opinion of and more like an encyclopedia set made by a mix of authors we have low opinions of.
- xtiansimon 1mo ago> "imperceptible statistical pattern” This is how I’m feeling right now. A “watermark” is as an author’s mark. I struggle to understand how a myriad of different texts will produce the same watermark output. How big does a text have to be to generate this sign? What is the false positive rate (where my own authentic prose—gasp—is falsely accused of being AI). How do you “prove” it’s true? Will Anthropic offer some kind of service? I find ChatGPT to be overly loquacious, and my preference for Claude is the brevity of output. Does this mean I will now have to suffer Claude’s gibbering, too?
- bagelcollie 1mo agoFrom my understanding, AI outputs have patterns due to its text prediction algorithm and AI detectors just recognize those patterns as watermarks. If there used to be 1,000 different ways to write a paragraph, watermarks may reduce it to 300 and that’s still a lot. Humans shouldn’t be able to perceive these changes except the output is really short. https://www.seangoedecke.com/text-ai-watermarks/ https://www.seangoedecke.com/text-ai-watermarks/
- spwa4 1mo agoNope. AI labs have agreed to start inserting those watermarks purposefully.
- dolmen 1mo agoEven more em-dash.