How Llama 2 Handles EOS Token Detection During Text Generation
Llama 2 detects the end-of-sequence (EOS) token during generation by maintaining a boolean tensor that tracks when each batch sequence emits the tokenizer's eos_id, breaking the sampling loop early when all sequences have terminated, then trimming outputs post-generation to remove the EOS marker and subsequent tokens.
During autoregressive text generation in the meta-llama/llama repository, the model must know when to stop producing tokens. EOS token detection is handled through a two-phase approach implemented in llama/generation.py that ensures efficient early stopping and clean output formatting.
Runtime EOS Detection in the Generation Loop
The core generation logic continuously monitors for the EOS token after each sampling step. The implementation uses a boolean tensor named eos_reached to track which batch entries have terminated.
After sampling a new token at position cur_pos, the code updates eos_reached only for positions that are not part of the original prompt (masked by ~input_text_mask[:, cur_pos]) and where the sampled token equals the tokenizer's EOS ID:
eos_reached |= (~input_text_mask[:, cur_pos]) & (
next_token == self.tokenizer.eos_id
)
if all(eos_reached):
break
This logic appears in llama/generation.py at lines 75–81. The |= operator ensures that once a sequence has emitted an EOS token, it remains marked as finished. The generation loop breaks immediately when all(eos_reached) returns True, preventing unnecessary computation for completed sequences.
Post-Generation Trimming of EOS Tokens
After the generation loop terminates, the code performs a cleanup pass to remove the EOS token from the final output. This ensures that the returned text does not contain the end-of-sequence marker.
In llama/generation.py at lines 24–28, the implementation checks each token list for the presence of self.tokenizer.eos_id. If found, it slices the tokens and corresponding log-probabilities up to (but not including) the EOS position:
if self.tokenizer.eos_id in toks:
eos_idx = toks.index(self.tokenizer.eos_id)
toks = toks[:eos_idx]
probs = probs[:eos_idx] if logprobs else None
This post-generation trimming guarantees that the EOS ID never appears in the text shown to users, and that probability scores align precisely with the returned tokens.
Practical Implementation Examples
The EOS detection mechanism operates automatically in both text completion and chat completion modes. When calling text_completion(), the model stops as soon as it generates the EOS token:
# Example 1 – Simple text completion
from llama import Llama
llama = Llama.build(
ckpt_dir="checkpoints",
tokenizer_path="tokenizer.model",
max_seq_len=2048,
max_batch_size=4,
)
# The model will stop when it generates its EOS token.
result = llama.text_completion(
prompts=["Write a haiku about sunrise:"],
max_gen_len=100,
)
print(result[0]["generation"])
Similarly, in chat_completion(), the generation halts upon EOS detection regardless of whether the maximum generation length has been reached:
# Example 2 – Chat completion with early EOS stop
dialog = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Tell me a short joke."},
]
response = llama.chat_completion(
dialogs=[dialog],
max_gen_len=50,
)
print(response[0]["generation"]["content"])
Both methods rely on the generation logic in llama/generation.py and the eos_id defined in llama/tokenizer.py to determine when to terminate.
Summary
- Runtime detection uses a boolean tensor
eos_reachedthat updates when sampled tokens matchself.tokenizer.eos_id, triggering an early break when all batch sequences have finished. - Masking logic ensures only generated tokens are checked for EOS, ignoring the original prompt positions via
~input_text_mask[:, cur_pos]. - Post-processing trims the EOS token and subsequent entries from the final output lists, ensuring clean text returned to the user.
- Key files involved are
llama/generation.py(detection loop and trimming) andllama/tokenizer.py(EOS ID definition).
Frequently Asked Questions
How does Llama 2 prevent the EOS token from appearing in the final output?
After generation completes, the code checks if self.tokenizer.eos_id exists in the token list. If present, it finds the index of the first occurrence and slices the sequence up to that position, removing the EOS marker and any tokens that follow before returning the result.
What happens if different sequences in a batch reach EOS at different times?
The eos_reached tensor tracks each batch entry independently using boolean indexing. The generation loop continues until all(eos_reached) evaluates to True, meaning every sequence in the batch has emitted an EOS token, ensuring no partial outputs are returned prematurely.
Where is the EOS token ID defined in the Llama 2 codebase?
The eos_id is defined in llama/tokenizer.py as part of the tokenizer class. The generation code accesses this value via self.tokenizer.eos_id to perform comparisons during the sampling loop and post-generation trimming phases.
Does Llama 2 support custom stopping criteria beyond the EOS token?
The reference implementation in meta-llama/llama specifically implements EOS detection through the eos_reached tensor logic in llama/generation.py. Implementing custom stopping criteria would require modifying the generation loop to check additional conditions alongside the EOS token comparison.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →