How to Implement Text Watermarks for LLM-Generated Content: A Complete Guide

The dive-into-llms repository provides a modular, three-stage pipeline—embedding, detection, and evaluation—to implement text watermarks for LLM-generated content using statistical algorithms like KGW, X-SIR, and SIR.

Text watermarks for LLM-generated content embed invisible statistical signatures during token generation, enabling traceability without degrading text quality. The Lordog/dive-into-llms repository provides a production-ready implementation in documents/chapter5/ that supports multiple watermark algorithms and includes robustness testing against paraphrasing and translation attacks. This guide walks through the complete technical workflow using the official command-line tools and Jupyter notebook.

Overview of the Watermark Pipeline

The implementation organizes text watermarking into three independent stages that can be integrated into larger training or inference systems.

Stage 1: Watermark Embedding

The embedding stage modifies the model's token distribution during generation to insert statistical patterns. The watermarker.py script wraps Hugging Face models and injects watermark-specific logic at each decoding step based on the WATERMARK_METHOD_FLAG variable.

Supported algorithms include:

  • KGW (Kirchenbauer-Green-Geiping-Wen): Splits vocabulary into green/red lists based on previous tokens
  • X-SIR: Cross-token statistical watermarking
  • SIR: Sentence-level information retrieval watermarking

The script outputs JSONL files containing generated samples with embedded markers saved to gen/<model_abbr>/<algorithm>/.

Stage 2: Watermark Detection

Detection quantifies watermark presence using z-scores. The detect_watermark.py script reads generated JSONL files and computes statistical tests to determine if text contains the embedded signature. A high positive z-score indicates strong watermark presence, while scores near zero suggest clean text.

This stage operates independently from generation, allowing you to test arbitrary text strings against your watermark keys.

Stage 3: Watermark Evaluation

Evaluation compares watermarked distributions against clean (human) baselines using ROC curves, AUC calculations, and accuracy metrics. The evaluate_watermark.py script takes z-score files from both watermarked and clean texts to measure detection reliability and false positive rates.

Implementing the Watermark Workflow

Prerequisites and Setup

Ensure you have access to the chapter 5 resources in the repository:

cd documents/chapter5

# Dependencies include transformers, torch, and scipy for statistical tests

Set environment variables for your model and algorithm choice:

MODEL=baichuan-inc/Baichuan-7B
MODEL_ABBR=baichuan7b
WATERMARK_METHOD_FLAG="--watermark_method kgw"

Generating Watermarked Text

Run the watermarker to generate text with embedded signatures:

python watermarker.py \
  --model $MODEL \
  $WATERMARK_METHOD_FLAG \
  --output_file gen/$MODEL_ABBR/kgw/mc4.en.mod.jsonl

This creates mc4.en.mod.jsonl containing watermarked samples in the gen/baichuan7b/kgw/ directory.

Computing Detection Z-Scores

Calculate statistical scores for the generated text:

python detect_watermark.py \
  --detect_file gen/$MODEL_ABBR/kgw/mc4.en.mod.jsonl \
  --output_file gen/$MODEL_ABBR/kgw/mc4.en.mod.z_score.jsonl

For baseline comparison, generate scores for clean human text:

python detect_watermark.py \
  --detect_file gen/$MODEL_ABBR/kgw/mc4.en.hum.jsonl \
  --output_file gen/$MODEL_ABBR/kgw/mc4.en.hum.z_score.jsonl

Evaluating Detection Performance

Compute ROC curves and accuracy metrics by comparing watermarked versus clean distributions:

python evaluate_watermark.py \
  --wm_zscore gen/$MODEL_ABBR/kgw/mc4.en.mod.z_score.jsonl \
  --hm_zscore gen/$MODEL_ABBR/kgw/mc4.en.hum.z_score.jsonl \
  --output_dir eval/$MODEL_ABBR/kgw

Testing Robustness Against Attacks

The implementation includes utilities to test watermark survival after content manipulation.

Paraphrase and Translation Attacks

Use paraphrase.py to rewrite watermarked text via GPT-3.5-turbo, simulating an evasion attempt:

python paraphrase.py \
  --input_file gen/$MODEL_ABBR/kgw/mc4.en.mod.jsonl \
  --output_file gen/$MODEL_ABBR/kgw/mc4.en-zh.mod.jsonl

Re-run detection on attacked content to measure score degradation:

python detect_watermark.py \
  --detect_file gen/$MODEL_ABBR/kgw/mc4.en-zh.mod.jsonl \
  --output_file gen/$MODEL_ABBR/kgw/mc4.en-zh.mod.z_score.jsonl

The watermark.ipynb notebook in documents/chapter5/ demonstrates complete robustness testing workflows, including translation attacks and comparative analysis of z-score distributions before and after rewriting.

Summary

  • Three-stage architecture: Independent CLI tools for embedding (watermarker.py), detection (detect_watermark.py), and evaluation (evaluate_watermark.py)
  • Algorithm flexibility: Native support for KGW, X-SIR, and SIR watermarking methods via --watermark_method flags
  • Statistical detection: Z-score-based quantification of watermark strength with ROC and AUC evaluation metrics
  • Attack robustness: Built-in testing against GPT-3.5-turbo paraphrasing and translation attacks using paraphrase.py
  • Modular integration: JSONL-based I/O enables insertion into existing LLM training and inference pipelines

Frequently Asked Questions

What watermark algorithms does the dive-into-llms implementation support?

The repository supports KGW (Kirchenbauer-Green-Geiping-Wen), X-SIR, and SIR algorithms. You can select these via the --watermark_method flag when running watermarker.py, as documented in documents/chapter5/README.md and demonstrated in watermark.ipynb.

How does the detection script determine if text contains a watermark?

The detect_watermark.py script computes a z-score that measures the statistical deviation of token distributions from expected random behavior. According to the implementation in documents/chapter5/, a high positive z-score indicates strong watermark presence, while scores near zero suggest the text is unmarked.

Can this implementation test watermark robustness against rewriting attacks?

Yes. The repository includes paraphrase.py to simulate attacks using GPT-3.5-turbo for paraphrasing and translation. You can re-run detection on rewritten outputs to measure how z-scores degrade, allowing you to calculate robustness metrics and ROC curves under adversarial conditions.

Where are the main implementation files located?

The core implementation resides in documents/chapter5/, specifically watermark.ipynb (full Jupyter notebook pipeline), watermarker.py (generation), detect_watermark.py (statistical testing), and evaluate_watermark.py (performance metrics). The accompanying README.md provides command-line reference documentation for all supported flags and workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →