AI text watermarks hide a signal in which words a model picks. A common idea is that you can remove that signal by running the text through another model: "rewrite this in your own words."
That raises a simple question. If you rewrite a text enough to break its word patterns, do you still keep its meaning?
I tested it on a single GPU with 16 GB of VRAM. In short: yes, mostly — but the model you pick matters a lot, and there is a real trade-off between "changed a lot" and "kept everything."
The setup
Here is the pipeline:
- Claude Opus 5.5 wrote 100 source texts. There were 10 genres (news, technical explainers, science, opinion, personal stories, marketing, how-to guides, business memos, history, and short fiction), with 10 topics each. Lengths were from about 190 to 900 words. Every text was packed with checkable details: names, numbers, dates, steps, and plot events.
-
Three local models rewrote every text with Ollama on a GPU with 16 GB VRAM:
-
qwen3:4b-instruct(small) -
gemma3:12b(mid) -
qwen3:14b(upper mid, thinking mode off)
-
- Claude Fable 5.1 judged all 300 rewrites blind. Each packet held the original and three rewrites labeled A/B/C in random order. The judge listed every lost, changed, or added fact, and gave each rewrite a score from 1 to 5.
- I measured how much the wording changed by counting how many of the original's word n-grams survive in the rewrite.
All models used the same prompt:
Rewrite the text below completely in your own words.
- Change the vocabulary and the sentence structure; do not copy phrases.
- Keep ALL of the meaning: every fact, name, number, date, step, argument, and plot event.
- Do not add new information, comments, or opinions.
- Keep the same genre, tone, and roughly the same length. Keep the title line, rewritten.
Output only the rewritten text, nothing else.
Settings: temperature 0.7, num_ctx 8192, one run per text.
Why n-grams are a fair proxy for "watermark removed"
Most text watermarks (green-list schemes like Kirchenbauer et al., and Google's SynthID-Text) work at the token level. While the model writes, they nudge it toward certain tokens based on the tokens just before. A detector then counts how often these "preferred" token pairs or sequences appear.
So if most of the original word sequences are gone, most of the signal is gone too. I did not run a watermark detector, because I have no access to one for these texts. Treat the n-gram numbers as a proxy, not a proof.
Results
Meaning, by model
| Model | Avg score (1–5) | "Preserved" | "Mostly preserved" | "Not preserved" | Same understanding for the reader |
|---|---|---|---|---|---|
| qwen3:14b | 4.68 | 82% | 18% | 0% | 100% |
| gemma3:12b | 4.09 | 40% | 60% | 0% | 98% |
| qwen3:4b-instruct | 3.76 | 21% | 76% | 3% | 81% |
| All 300 | 4.18 | 48% | 51% | 1% | 93% |
"Same understanding" was a yes/no question to the judge: would a reader of the rewrite come away with the same understanding as a reader of the original?
When the judge ranked the three rewrites of each text (still blind):
| Model | Ranked best | Ranked worst |
|---|---|---|
| qwen3:14b | 76 | 5 |
| gemma3:12b | 17 | 35 |
| qwen3:4b-instruct | 7 | 60 |
Average errors per rewrite, as listed by the judge:
| Model | Facts lost | Facts changed | Things added |
|---|---|---|---|
| qwen3:14b | 0.59 | 0.96 | 0.05 |
| gemma3:12b | 1.42 | 2.04 | 0.30 |
| qwen3:4b-instruct | 1.22 | 2.94 | 0.38 |
Wording change, by model
"Kept" means the share of the original's n-grams that still appear in the rewrite. Lower means a heavier rewrite.
| Model | Words kept (1-gram) | 2-grams kept | 3-grams kept | 5-grams kept | Longest copied run (words) | Length vs original |
|---|---|---|---|---|---|---|
| gemma3:12b | 60% | 34% | 20% | 8% | 11.7 | 1.12× |
| qwen3:4b-instruct | 64% | 37% | 23% | 9% | 12.4 | 1.05× |
| qwen3:14b | 72% | 51% | 37% | 21% | 20.7 | 1.07× |
The trade-off
Put the two tables side by side:
- qwen3:14b kept the meaning best by far, but it also kept the most of the original wording. About 1 in 5 five-word sequences survived, and on average it copied a run of about 21 words straight from the original.
- gemma3:12b changed the wording the most (only 8% of 5-grams survived) and still scored 4.09. Almost all its rewrites (98%) gave the reader the same understanding.
- qwen3:4b-instruct changed the wording almost as much as Gemma, but it was the least accurate.
Across all 300 rewrites, the correlation between "meaning score" and "5-grams kept" was 0.36. So yes, rewrites that copy more wording tend to keep more meaning. But inside each model the link was weak (0.03–0.22). Most of the difference comes from which model you use, not from how hard a given model happens to rewrite a given text.
What kinds of mistakes?
The type of error mattered more than the count.
qwen3:14b mostly made tiny changes in nuance:
- "two placebo pills work better than one" was hedged to "may be more effective"
- "ransomware" was widened to "malware attacks"
- an overall risk of "moderate" became "average"
gemma3:12b made the wording vaguer or swapped in a near-synonym:
- "Most houseplants" became "Many indoor plants"
- "hiring manager" became "recruiter"
- Kubernetes init containers that "run to completion, one after another" became a single init container that "executes a sequence of tasks". That one is a real technical error.
Here is a typical example. Original:
SafeNest Products announced on Monday that it is recalling about 86,000 infant car seats sold in the United States because the harness buckle may fail to latch fully.
Gemma:
On Monday, SafeNest Products declared a recall of roughly 86,000 car seats intended for infants, distributed across the United States, due to a potential issue with the harness buckle's functionality.
Every fact is still there, but "may fail to latch fully" became "a potential issue with ... functionality." It reads fine, but it says less.
qwen3:4b-instruct made the dangerous mistakes: wrong names and wrong numbers.
- It renamed products: "Quietwave Q7" headphones became "SilentHush S7" all through the text, and the "Elevate Pro Dual-Motor" desk became "Rise Pro Dual-Drive."
- It garbled numbers: a desk height range "from 150 to 200 centimeters" became "from 150 to 20."
- It changed amounts: British tea imports of "millions of pounds" a year became "hundreds of thousands of pounds."
- It cut a farmer's name, "Marcos Ribeiro," to "Marcos Ribe."
A rough keyword search over the judge's notes found number-related changes in about 31 of 100 qwen3:4b rewrites, compared with about 11 for Gemma and 3 for qwen3:14b. This count is approximate, but the gap is clear.
All 3 rewrites judged "not preserved" in the whole test came from qwen3:4b. Two were the renamed products, and the third was a short story whose ending stopped making sense.
By genre and length
The genre made less difference than I expected. The average score ranged from 4.03 (marketing copy) to 4.40 (science writing). Marketing copy had the most "not preserved" verdicts. Product names and spec numbers are exactly where the small model slips, and those details are the meaning of marketing copy. Business memos and technical explainers were most often fully "preserved" (77–80%).
Length did not hurt. The 750–900 word texts scored about the same as, or a little better than, the short ones (4.33 vs 4.19). With an 8K context, all three models handled 900 words easily.
Speed
Total time for 100 rewrites on the 16 GB GPU:
| Model | Total | Per text |
|---|---|---|
| qwen3:4b-instruct | 11 min | ~7 s |
| qwen3:14b | 25 min | ~15 s |
| gemma3:12b | 25 min | ~15 s |
So — does the rewrite keep the meaning?
With a good 12–14B model: yes. Readers would get the same understanding from 98–100% of the rewrites, and none were judged "not preserved." The remaining issues are small shifts in nuance or precision.
With a 4B model: usually, but you can't trust the details. One in five rewrites changed what a reader would understand. Names and numbers are the weak spot, and they are often the facts that matter most.
And there is a trade-off you can't ignore:
- If you want maximum fidelity,
qwen3:14bis the clear winner. But it keeps more of the original wording, so if your goal is to break token-level patterns, it does the weakest job. - If you want maximum change in wording with good meaning,
gemma3:12bis the sweet spot. Only 8% of 5-grams survive, and the reader gets the same understanding 98% of the time.
Caveats
- No real watermark detector. N-gram overlap is a proxy. Some watermark research claims robustness to paraphrase, and semantic-level watermarks would not be hurt by word changes at all.
- One judge, and it is an LLM. Fable 5.1 was strict and specific (it listed things like "price changed from X to Y"), but I did not check it against human ratings. The generator and the judge are both Claude models, which could add some bias toward Claude-like phrasing. That should affect all three local models equally, though.
- One run per text at temperature 0.7. Another seed would give somewhat different rewrites.
- English only, synthetic texts, one prompt. A stricter prompt ("never change any name or number") would probably help the 4B model.
- The keyword counts of name and number errors are approximate.
Next time: SynthID?
This test answered only half of the question. It showed that a good local model keeps the meaning, but it only estimated watermark removal from n-gram overlap.
A real check would use SynthID-Text, Google's open-source text watermark. Would rewriting with a local model really remove a watermark that a detector can see?
Should I run that test next? Tell me in the comments.
Top comments (0)