# How to Compute Token Log Probabilities During Llama 2 Generation

> Learn how to compute token log probabilities during Llama 2 generation by enabling the logprobs parameter in generate text completion or chat completion calls for detailed output.

- Repository: [Meta Llama/llama](https://github.com/meta-llama/llama)
- Tags: deep-dive
- Published: 2026-03-05

---

**Set `logprobs=True` when calling `generate()`, `text_completion()`, or `chat_completion()` to receive per-token log-probabilities alongside generated tokens.**

The official `meta-llama/llama` repository provides built-in support to compute token log probabilities during text generation. The `generate()` method in [`llama/generation.py`](https://github.com/meta-llama/llama/blob/main/llama/generation.py) exposes a boolean `logprobs` parameter that calculates the negative log-likelihood for every token using PyTorch’s cross-entropy loss, enabling detailed confidence analysis without requiring model modifications.

## Enabling Log Probability Computation

To compute token log probabilities, pass `logprobs=True` to any of the three primary generation methods. This flag is available in:

- **`generate()`** – The core generation loop in [`llama/generation.py`](https://github.com/meta-llama/llama/blob/main/llama/generation.py) (lines 31-38)
- **`text_completion()`** – High-level helper that forwards the flag to `generate()` (lines 62-70)
- **`chat_completion()`** – Dialog-aware helper with identical log probability support

When enabled, the method returns a tuple containing both the generated token IDs and their corresponding log-probabilities. If `logprobs=False` (the default), the second element of the return tuple is `None`.

## How Log Probabilities Are Computed Internally

The implementation stores and computes log probabilities through a dedicated tensor that mirrors the shape of the generated sequence. The source code follows this exact flow:

### Method Signature and Allocation

The `generate()` method accepts the flag with a default value of `False` and prepares storage when activated:

```python
def generate(
    ...,
    logprobs: bool = False,
    ...
) -> Tuple[List[List[int]], Optional[List[List[float]]]]:

```

When `logprobs` is `True`, the code immediately allocates a floating-point tensor matching the token buffer dimensions (lines 71-73 in [`llama/generation.py`](https://github.com/meta-llama/llama/blob/main/llama/generation.py)):

```python
if logprobs:
    token_logprobs = torch.zeros_like(tokens, dtype=torch.float)

```

### Prompt-Only Sequences

If the input prompt fills the entire sequence length (`min_prompt_len == total_len`), the code computes cross-entropy loss for every prompt position before any generation begins (lines 78-84):

```python
token_logprobs = -F.cross_entropy(
    input=logits.transpose(1, 2),
    target=tokens[:, prev_pos + 1 : cur_pos + 1],
    reduction="none",
    ignore_index=pad_id,
)

```

### Iterative Generation Loop

During token generation, after each new token is sampled, the log probability is calculated for the just-generated token (and any preceding unprocessed positions). This occurs inside the main generation loop (lines 100-106):

```python
token_logprobs[:, prev_pos + 1 : cur_pos + 1] = -F.cross_entropy(
    input=logits.transpose(1, 2),
    target=tokens[:, prev_pos + 1 : cur_pos + 1],
    reduction="none",
    ignore_index=pad_id,
)

```

The use of `reduction="none"` ensures per-token losses are preserved rather than aggregated.

### Return Value Structure

After generation completes, the tensor converts to a standard Python list (lines 214-215) and returns through a conditional tuple (lines 231-232):

```python
token_logprobs = token_logprobs.tolist()
return (out_tokens, out_logprobs if logprobs else None)

```

The high-level helpers `text_completion()` and `chat_completion()` attach these values to their result dictionaries under the key `"logprobs"`, aligning parallel lists of tokens and their scores (lines 74-81).

## Code Examples

### Direct `generate()` Call

Use this approach when you need raw token IDs and maximum control:

```python
from llama import Llama

# Initialize the model

llama = Llama.build(
    ckpt_dir="llama-2-7b-chat/",
    tokenizer_path="tokenizer.model",
    max_seq_len=512,
    max_batch_size=4,
)

prompt = ["Hello, how are you?"]
prompt_ids = [llama.tokenizer.encode(p, bos=True, eos=False) for p in prompt]

# Enable log probability computation

tokens, logprobs = llama.generate(
    prompt_tokens=prompt_ids,
    max_gen_len=50,
    temperature=0.7,
    top_p=0.9,
    logprobs=True,  # <‑‑ Enable log probabilities

    echo=False,
)

print("Generated tokens:", tokens[0])
print("Log-probs per token:", logprobs[0])

```

### Text Completion Helper

For decoded text with automatic token-to-string mapping:

```python
result = llama.text_completion(
    prompts=["Explain quantum entanglement in one sentence."],
    max_gen_len=30,
    temperature=0.5,
    top_p=0.95,
    logprobs=True,  # <‑‑ Enable log probabilities

)

print(result[0]["generation"])  # Decoded text

print(result[0]["tokens"])      # Token strings

print(result[0]["logprobs"])    # Log-probability for each token

```

### Chat Completion Helper

For dialog-based generation with per-token scoring:

```python
dialog = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "What's the weather in Paris?"},
]

chat = llama.chat_completion(
    dialogs=[dialog],
    max_gen_len=20,
    temperature=0.8,
    top_p=0.9,
    logprobs=True,  # <‑‑ Enable log probabilities

)

print(chat[0]["generation"]["content"])  # Assistant reply

print(chat[0]["tokens"])                 # Token strings

print(chat[0]["logprobs"])               # Log-probs list

```

## Key Source Files

- **[`llama/generation.py`](https://github.com/meta-llama/llama/blob/main/llama/generation.py)** – Contains the core `generate()` method, `logprobs` flag handling, and cross-entropy calculation logic
- **[`llama/model.py`](https://github.com/meta-llama/llama/blob/main/llama/model.py)** – Defines the Transformer architecture and forward pass that produces the logits used for probability calculation
- **[`llama/tokenizer.py`](https://github.com/meta-llama/llama/blob/main/llama/tokenizer.py)** – Handles token encoding and decoding between strings and integer IDs
- **[`example_text_completion.py`](https://github.com/meta-llama/llama/blob/main/example_text_completion.py)** – Runnable demonstration showing `logprobs` usage with the text completion API
- **[`example_chat_completion.py`](https://github.com/meta-llama/llama/blob/main/example_chat_completion.py)** – Runnable demonstration showing `logprobs` usage with the chat completion API

## Summary

- **Enable log probabilities** by setting `logprobs=True` in `generate()`, `text_completion()`, or `chat_completion()`
- **Implementation** uses `-F.cross_entropy` with `reduction="none"` to compute per-token negative log-likelihoods from model logits
- **Storage** occurs in a `torch.zeros_like` tensor that matches the token sequence shape, converted to Python lists before returning
- **Return format** is `Tuple[List[List[int]], Optional[List[List[float]]]]` where the second element contains batch-wise lists of per-token log-probabilities (or `None` if disabled)
- **High-level APIs** automatically attach log-probabilities to result dictionaries under the `"logprobs"` key alongside `"tokens"`

## Frequently Asked Questions

### What parameter enables token log probability computation in Llama 2?

Pass `logprobs=True` to the `generate()`, `text_completion()`, or `chat_completion()` methods. This boolean flag defaults to `False` and is defined in the method signature at line 31 of [`llama/generation.py`](https://github.com/meta-llama/llama/blob/main/llama/generation.py).

### How are log probabilities calculated internally?

The implementation applies PyTorch’s `F.cross_entropy` function with `reduction="none"` to the model’s logits. Specifically, it computes `-F.cross_entropy(input=logits.transpose(1, 2), target=tokens, ...)` for each token position, yielding the negative log-likelihood (natural logarithm of the predicted probability).

### What is the return format when logprobs is enabled?

The methods return a tuple where the first element is a list of token ID lists (batch dimension included), and the second element is a list of lists containing float values representing the log probability of each corresponding token. When `logprobs=False`, the second element is `None`.

### Can I retrieve log probabilities for prompt tokens as well as generated tokens?

Yes. When the `echo` parameter is set to `True` alongside `logprobs=True`, the returned lists include both the prompt tokens and the generated tokens with their respective log probabilities computed for all positions.