# Deterministic vs Probabilistic Watermarking in LLMs: Implementation and Detection Methods

> Understand deterministic vs probabilistic watermarking in LLMs. Learn about fixed token mappings and biased sampling distributions for effective watermark implementation and detection.

- Repository: [Tongxin Yuan/dive-into-llms](https://github.com/Lordog/dive-into-llms)
- Tags: deep-dive
- Published: 2026-04-16

---

**Deterministic watermarking embeds hidden signals via fixed token mappings like Huffman coding for exact decoding, while probabilistic watermarking biases sampling distributions and requires statistical Z-score detection to verify the watermark.**

Watermarking embeds hidden signals into text generated by large language models (LLMs) to enable later verification of AI-generated content. The `Lordog/dive-into-llms` repository implements both deterministic and probabilistic approaches, demonstrating how each method balances predictability, robustness, and detection complexity.

## Deterministic Watermarking: Exact Token Mapping

Deterministic watermarking creates a fixed, reversible relationship between watermark bits and specific tokens. Given the same input and secret key, the model always selects the same tokens to encode the message.

### Huffman Coding Implementation

In the deterministic approach, the repository’s `Huffman` class builds a Huffman tree from the top-k token probabilities and selects the token whose code matches the current watermark bits. This produces an exact, one-to-one mapping from bits to tokens.

The implementation in `documents/chapter7/llm_stega.ipynb` defines the embedding logic:

```python
class Huffman():
    def __init__(self, k=None, bits=None):
        self.k = k
        self.bits = bits

    def extract(self, pred_prob, word_index):
        prob, idx = torch.topk(pred_prob, k=self.k, dim=1)
        word_prob = [[i.item(), j.item()] for i, j in zip(idx[0], prob[0])]
        nodes = createNodes([item[1] for item in word_prob])
        root = createHuffmanTree(nodes)
        codes = huffmanEncoding(nodes, root)
        for i, w_p in enumerate(word_prob):
            if w_p[0] == word_index:
                return codes[i]

    def __call__(self, pred_prob):
        prob, idx = torch.topk(pred_prob, k=self.k, dim=1)
        word_prob = [[i.item(), j.item()] for i, j in zip(idx[0], prob[0])]
        nodes = createNodes([item[1] for item in word_prob])
        root = createHuffmanTree(nodes)
        codes = huffmanEncoding(nodes, root)
        for i, code in enumerate(codes):
            bit = self.bits[:len(code)]
            if bit == code:
                bit_word_index = word_prob[i][0]
                self.bits = self.bits[len(code):]
                break
        return torch.LongTensor([bit_word_index]).to(pred_prob.device)

```

This method traverses the same coding tree during decoding to recover bits directly without statistical inference.

## Probabilistic Watermarking: Distribution Biasing

Probabilistic watermarking, exemplified by the KGW (Kirchenbauer et al.) algorithm, influences the sampling distribution rather than selecting specific tokens. The model biases the probability distribution toward a "green list" of tokens, introducing stochastic variability while maintaining statistical detectability.

### Fixed-Length Coding and Stochastic Sampling

The repository’s `class FLC` (Fixed-Length Coding) in `documents/chapter7/llm_stega.ipynb` maps bits to token indices but operates with a fixed number of bits per token. Unlike the deterministic Huffman approach, this can be combined with stochastic sampling from the biased probability distribution, meaning the exact token may vary across generation runs even with identical watermark bits.

### Z-Score Statistical Detection

Because probabilistic watermarking introduces randomness, detection requires statistical analysis rather than direct bit extraction. The [`detect.py`](https://github.com/Lordog/dive-into-llms/blob/main/detect.py) script in Chapter 5 computes a Z-score to quantify how strongly the observed token distribution aligns with the expected bias:

```bash
python3 detect.py \
    --base_model $MODEL_NAME \
    --detect_file gen/$MODEL_ABBR/kgw/mc4.en.mod.jsonl \
    --output_file gen/$MODEL_ABBR/kgw/mc4.en.mod.z_score.jsonl \
    $WATERMARK_METHOD_FLAG

```

This script evaluates whether token frequencies are shifted toward the biased distribution, calculating the statistical significance of the watermark presence.

## Critical Differences Between Approaches

Understanding the trade-offs between deterministic and probabilistic watermarking requires examining four key dimensions:

- **Embedding Rule**: **Deterministic** methods use fixed coding schemes (Huffman trees) to map bits to exact tokens. **Probabilistic** methods alter the sampling distribution, allowing token variation while maintaining statistical biases.

- **Detection Method**: **Deterministic** watermarks decode directly by traversing the coding tree—no statistical tests required. **Probabilistic** watermarks require computing Z-scores or similar statistical measures over large text chunks to infer watermark presence.

- **Predictability**: **Deterministic** outputs are highly predictable and reproducible given the same inputs. **Probabilistic** outputs vary between runs due to stochastic sampling, making exact token prediction impossible.

- **Robustness Trade-offs**: **Deterministic** methods offer minimal overhead and simple verification but remain vulnerable to deterministic attacks like paraphrasing. **Probabilistic** methods survive transformations such as translation or rewriting because the statistical bias persists across token substitutions, though detection requires more data.

## Summary

- **Deterministic watermarking** uses `class Huffman` in `documents/chapter7/llm_stega.ipynb` to create fixed bit-to-token mappings, enabling direct decoding without statistical inference.
- **Probabilistic watermarking** biases token distributions via methods like `class FLC` and requires Z-score computation via [`detect.py`](https://github.com/Lordog/dive-into-llms/blob/main/detect.py) to verify watermark presence statistically.
- Deterministic approaches prioritize exact reproducibility and low detection overhead, while probabilistic approaches sacrifice predictability for robustness against content transformations.
- The `dive-into-llms` repository provides working implementations of both paradigms, demonstrating Huffman coding for deterministic embedding and KGW-style statistical detection for probabilistic schemes.

## Frequently Asked Questions

### What is the main technical difference between deterministic and probabilistic watermarking?

Deterministic watermarking establishes a fixed algorithmic relationship where specific watermark bits always map to specific tokens (e.g., via Huffman coding), allowing exact extraction. Probabilistic watermarking modifies the probability distribution from which tokens are sampled, meaning the same watermark bits may yield different tokens across generations, requiring statistical aggregation (Z-scores) to detect the embedded signal.

### How does the Huffman class implement deterministic watermarking in the dive-into-llms repository?

The `Huffman` class in `documents/chapter7/llm_stega.ipynb` constructs a Huffman tree from the top-k token probabilities and selects the token whose binary code matches the current watermark bits. This creates a reversible, deterministic mapping that enables exact bit recovery during decoding by traversing the identical tree structure.

### Why does probabilistic watermarking require Z-score detection instead of direct decoding?

Because probabilistic watermarking stochastically samples from a biased distribution rather than selecting predetermined tokens, individual token choices contain noise. The Z-score computation aggregates statistics across many tokens to measure whether the observed distribution significantly deviates from the expected unbiased distribution, confirming watermark presence only when the bias is statistically significant.

### Which watermarking method is more resistant to paraphrasing attacks?

**Probabilistic watermarking** demonstrates greater robustness against paraphrasing and translation attacks. Since the watermark exists as a statistical bias in the token distribution rather than as specific token instances, rewriting the text preserves the underlying distributional shift. Deterministic watermarks, which rely on exact token sequences, are more easily removed by synonym substitution or syntactic restructuring.