Deterministic vs Probabilistic Watermarking in LLMs: Implementation and Detection Methods

Deterministic watermarking embeds hidden signals via fixed token mappings like Huffman coding for exact decoding, while probabilistic watermarking biases sampling distributions and requires statistical Z-score detection to verify the watermark.

Watermarking embeds hidden signals into text generated by large language models (LLMs) to enable later verification of AI-generated content. The Lordog/dive-into-llms repository implements both deterministic and probabilistic approaches, demonstrating how each method balances predictability, robustness, and detection complexity.

Deterministic Watermarking: Exact Token Mapping

Deterministic watermarking creates a fixed, reversible relationship between watermark bits and specific tokens. Given the same input and secret key, the model always selects the same tokens to encode the message.

Huffman Coding Implementation

In the deterministic approach, the repository’s Huffman class builds a Huffman tree from the top-k token probabilities and selects the token whose code matches the current watermark bits. This produces an exact, one-to-one mapping from bits to tokens.

The implementation in documents/chapter7/llm_stega.ipynb defines the embedding logic:

class Huffman():
    def __init__(self, k=None, bits=None):
        self.k = k
        self.bits = bits

    def extract(self, pred_prob, word_index):
        prob, idx = torch.topk(pred_prob, k=self.k, dim=1)
        word_prob = [[i.item(), j.item()] for i, j in zip(idx[0], prob[0])]
        nodes = createNodes([item[1] for item in word_prob])
        root = createHuffmanTree(nodes)
        codes = huffmanEncoding(nodes, root)
        for i, w_p in enumerate(word_prob):
            if w_p[0] == word_index:
                return codes[i]

    def __call__(self, pred_prob):
        prob, idx = torch.topk(pred_prob, k=self.k, dim=1)
        word_prob = [[i.item(), j.item()] for i, j in zip(idx[0], prob[0])]
        nodes = createNodes([item[1] for item in word_prob])
        root = createHuffmanTree(nodes)
        codes = huffmanEncoding(nodes, root)
        for i, code in enumerate(codes):
            bit = self.bits[:len(code)]
            if bit == code:
                bit_word_index = word_prob[i][0]
                self.bits = self.bits[len(code):]
                break
        return torch.LongTensor([bit_word_index]).to(pred_prob.device)

This method traverses the same coding tree during decoding to recover bits directly without statistical inference.

Probabilistic Watermarking: Distribution Biasing

Probabilistic watermarking, exemplified by the KGW (Kirchenbauer et al.) algorithm, influences the sampling distribution rather than selecting specific tokens. The model biases the probability distribution toward a "green list" of tokens, introducing stochastic variability while maintaining statistical detectability.

Fixed-Length Coding and Stochastic Sampling

The repository’s class FLC (Fixed-Length Coding) in documents/chapter7/llm_stega.ipynb maps bits to token indices but operates with a fixed number of bits per token. Unlike the deterministic Huffman approach, this can be combined with stochastic sampling from the biased probability distribution, meaning the exact token may vary across generation runs even with identical watermark bits.

Z-Score Statistical Detection

Because probabilistic watermarking introduces randomness, detection requires statistical analysis rather than direct bit extraction. The detect.py script in Chapter 5 computes a Z-score to quantify how strongly the observed token distribution aligns with the expected bias:

python3 detect.py \
    --base_model $MODEL_NAME \
    --detect_file gen/$MODEL_ABBR/kgw/mc4.en.mod.jsonl \
    --output_file gen/$MODEL_ABBR/kgw/mc4.en.mod.z_score.jsonl \
    $WATERMARK_METHOD_FLAG

This script evaluates whether token frequencies are shifted toward the biased distribution, calculating the statistical significance of the watermark presence.

Critical Differences Between Approaches

Understanding the trade-offs between deterministic and probabilistic watermarking requires examining four key dimensions:

  • Embedding Rule: Deterministic methods use fixed coding schemes (Huffman trees) to map bits to exact tokens. Probabilistic methods alter the sampling distribution, allowing token variation while maintaining statistical biases.

  • Detection Method: Deterministic watermarks decode directly by traversing the coding tree—no statistical tests required. Probabilistic watermarks require computing Z-scores or similar statistical measures over large text chunks to infer watermark presence.

  • Predictability: Deterministic outputs are highly predictable and reproducible given the same inputs. Probabilistic outputs vary between runs due to stochastic sampling, making exact token prediction impossible.

  • Robustness Trade-offs: Deterministic methods offer minimal overhead and simple verification but remain vulnerable to deterministic attacks like paraphrasing. Probabilistic methods survive transformations such as translation or rewriting because the statistical bias persists across token substitutions, though detection requires more data.

Summary

  • Deterministic watermarking uses class Huffman in documents/chapter7/llm_stega.ipynb to create fixed bit-to-token mappings, enabling direct decoding without statistical inference.
  • Probabilistic watermarking biases token distributions via methods like class FLC and requires Z-score computation via detect.py to verify watermark presence statistically.
  • Deterministic approaches prioritize exact reproducibility and low detection overhead, while probabilistic approaches sacrifice predictability for robustness against content transformations.
  • The dive-into-llms repository provides working implementations of both paradigms, demonstrating Huffman coding for deterministic embedding and KGW-style statistical detection for probabilistic schemes.

Frequently Asked Questions

What is the main technical difference between deterministic and probabilistic watermarking?

Deterministic watermarking establishes a fixed algorithmic relationship where specific watermark bits always map to specific tokens (e.g., via Huffman coding), allowing exact extraction. Probabilistic watermarking modifies the probability distribution from which tokens are sampled, meaning the same watermark bits may yield different tokens across generations, requiring statistical aggregation (Z-scores) to detect the embedded signal.

How does the Huffman class implement deterministic watermarking in the dive-into-llms repository?

The Huffman class in documents/chapter7/llm_stega.ipynb constructs a Huffman tree from the top-k token probabilities and selects the token whose binary code matches the current watermark bits. This creates a reversible, deterministic mapping that enables exact bit recovery during decoding by traversing the identical tree structure.

Why does probabilistic watermarking require Z-score detection instead of direct decoding?

Because probabilistic watermarking stochastically samples from a biased distribution rather than selecting predetermined tokens, individual token choices contain noise. The Z-score computation aggregates statistics across many tokens to measure whether the observed distribution significantly deviates from the expected unbiased distribution, confirming watermark presence only when the bias is statistically significant.

Which watermarking method is more resistant to paraphrasing attacks?

Probabilistic watermarking demonstrates greater robustness against paraphrasing and translation attacks. Since the watermark exists as a statistical bias in the token distribution rather than as specific token instances, rewriting the text preserves the underlying distributional shift. Deterministic watermarks, which rely on exact token sequences, are more easily removed by synonym substitution or syntactic restructuring.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →