# How Llama 2 Handles EOS Token Detection During Text Generation

> Discover how Llama 2 detects EOS tokens during text generation. Learn how it efficiently stops sampling and trims output for optimized results.

- Repository: [Meta Llama/llama](https://github.com/meta-llama/llama)
- Tags: internals
- Published: 2026-03-05

---

**Llama 2 detects the end-of-sequence (EOS) token during generation by maintaining a boolean tensor that tracks when each batch sequence emits the tokenizer's `eos_id`, breaking the sampling loop early when all sequences have terminated, then trimming outputs post-generation to remove the EOS marker and subsequent tokens.**

During autoregressive text generation in the `meta-llama/llama` repository, the model must know when to stop producing tokens. **EOS token detection** is handled through a two-phase approach implemented in [`llama/generation.py`](https://github.com/meta-llama/llama/blob/main/llama/generation.py) that ensures efficient early stopping and clean output formatting.

## Runtime EOS Detection in the Generation Loop

The core generation logic continuously monitors for the EOS token after each sampling step. The implementation uses a **boolean tensor** named `eos_reached` to track which batch entries have terminated.

After sampling a new token at position `cur_pos`, the code updates `eos_reached` only for positions that are **not** part of the original prompt (masked by `~input_text_mask[:, cur_pos]`) and where the sampled token equals the tokenizer's EOS ID:

```python
eos_reached |= (~input_text_mask[:, cur_pos]) & (
    next_token == self.tokenizer.eos_id
)
if all(eos_reached):
    break

```

This logic appears in [`llama/generation.py`](https://github.com/meta-llama/llama/blob/main/llama/generation.py) at lines 75–81. The `|=` operator ensures that once a sequence has emitted an EOS token, it remains marked as finished. The generation loop breaks immediately when `all(eos_reached)` returns True, preventing unnecessary computation for completed sequences.

## Post-Generation Trimming of EOS Tokens

After the generation loop terminates, the code performs a cleanup pass to remove the EOS token from the final output. This ensures that the returned text does not contain the end-of-sequence marker.

In [`llama/generation.py`](https://github.com/meta-llama/llama/blob/main/llama/generation.py) at lines 24–28, the implementation checks each token list for the presence of `self.tokenizer.eos_id`. If found, it slices the tokens and corresponding log-probabilities up to (but not including) the EOS position:

```python
if self.tokenizer.eos_id in toks:
    eos_idx = toks.index(self.tokenizer.eos_id)
    toks = toks[:eos_idx]
    probs = probs[:eos_idx] if logprobs else None

```

This **post-generation trimming** guarantees that the EOS ID never appears in the text shown to users, and that probability scores align precisely with the returned tokens.

## Practical Implementation Examples

The EOS detection mechanism operates automatically in both text completion and chat completion modes. When calling `text_completion()`, the model stops as soon as it generates the EOS token:

```python

# Example 1 – Simple text completion

from llama import Llama

llama = Llama.build(
    ckpt_dir="checkpoints",
    tokenizer_path="tokenizer.model",
    max_seq_len=2048,
    max_batch_size=4,
)

# The model will stop when it generates its EOS token.

result = llama.text_completion(
    prompts=["Write a haiku about sunrise:"],
    max_gen_len=100,
)
print(result[0]["generation"])

```

Similarly, in `chat_completion()`, the generation halts upon EOS detection regardless of whether the maximum generation length has been reached:

```python

# Example 2 – Chat completion with early EOS stop

dialog = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Tell me a short joke."},
]

response = llama.chat_completion(
    dialogs=[dialog],
    max_gen_len=50,
)
print(response[0]["generation"]["content"])

```

Both methods rely on the generation logic in [`llama/generation.py`](https://github.com/meta-llama/llama/blob/main/llama/generation.py) and the `eos_id` defined in [`llama/tokenizer.py`](https://github.com/meta-llama/llama/blob/main/llama/tokenizer.py) to determine when to terminate.

## Summary

- **Runtime detection** uses a boolean tensor `eos_reached` that updates when sampled tokens match `self.tokenizer.eos_id`, triggering an early break when all batch sequences have finished.
- **Masking logic** ensures only generated tokens are checked for EOS, ignoring the original prompt positions via `~input_text_mask[:, cur_pos]`.
- **Post-processing** trims the EOS token and subsequent entries from the final output lists, ensuring clean text returned to the user.
- **Key files** involved are [`llama/generation.py`](https://github.com/meta-llama/llama/blob/main/llama/generation.py) (detection loop and trimming) and [`llama/tokenizer.py`](https://github.com/meta-llama/llama/blob/main/llama/tokenizer.py) (EOS ID definition).

## Frequently Asked Questions

### How does Llama 2 prevent the EOS token from appearing in the final output?

After generation completes, the code checks if `self.tokenizer.eos_id` exists in the token list. If present, it finds the index of the first occurrence and slices the sequence up to that position, removing the EOS marker and any tokens that follow before returning the result.

### What happens if different sequences in a batch reach EOS at different times?

The `eos_reached` tensor tracks each batch entry independently using boolean indexing. The generation loop continues until `all(eos_reached)` evaluates to True, meaning every sequence in the batch has emitted an EOS token, ensuring no partial outputs are returned prematurely.

### Where is the EOS token ID defined in the Llama 2 codebase?

The `eos_id` is defined in [`llama/tokenizer.py`](https://github.com/meta-llama/llama/blob/main/llama/tokenizer.py) as part of the tokenizer class. The generation code accesses this value via `self.tokenizer.eos_id` to perform comparisons during the sampling loop and post-generation trimming phases.

### Does Llama 2 support custom stopping criteria beyond the EOS token?

The reference implementation in `meta-llama/llama` specifically implements EOS detection through the `eos_reached` tensor logic in [`llama/generation.py`](https://github.com/meta-llama/llama/blob/main/llama/generation.py). Implementing custom stopping criteria would require modifying the generation loop to check additional conditions alongside the EOS token comparison.