How to Compute Token Log Probabilities During Llama 2 Generation

Set logprobs=True when calling generate(), text_completion(), or chat_completion() to receive per-token log-probabilities alongside generated tokens.

The official meta-llama/llama repository provides built-in support to compute token log probabilities during text generation. The generate() method in llama/generation.py exposes a boolean logprobs parameter that calculates the negative log-likelihood for every token using PyTorch’s cross-entropy loss, enabling detailed confidence analysis without requiring model modifications.

Enabling Log Probability Computation

To compute token log probabilities, pass logprobs=True to any of the three primary generation methods. This flag is available in:

  • generate() – The core generation loop in llama/generation.py (lines 31-38)
  • text_completion() – High-level helper that forwards the flag to generate() (lines 62-70)
  • chat_completion() – Dialog-aware helper with identical log probability support

When enabled, the method returns a tuple containing both the generated token IDs and their corresponding log-probabilities. If logprobs=False (the default), the second element of the return tuple is None.

How Log Probabilities Are Computed Internally

The implementation stores and computes log probabilities through a dedicated tensor that mirrors the shape of the generated sequence. The source code follows this exact flow:

Method Signature and Allocation

The generate() method accepts the flag with a default value of False and prepares storage when activated:

def generate(
    ...,
    logprobs: bool = False,
    ...
) -> Tuple[List[List[int]], Optional[List[List[float]]]]:

When logprobs is True, the code immediately allocates a floating-point tensor matching the token buffer dimensions (lines 71-73 in llama/generation.py):

if logprobs:
    token_logprobs = torch.zeros_like(tokens, dtype=torch.float)

Prompt-Only Sequences

If the input prompt fills the entire sequence length (min_prompt_len == total_len), the code computes cross-entropy loss for every prompt position before any generation begins (lines 78-84):

token_logprobs = -F.cross_entropy(
    input=logits.transpose(1, 2),
    target=tokens[:, prev_pos + 1 : cur_pos + 1],
    reduction="none",
    ignore_index=pad_id,
)

Iterative Generation Loop

During token generation, after each new token is sampled, the log probability is calculated for the just-generated token (and any preceding unprocessed positions). This occurs inside the main generation loop (lines 100-106):

token_logprobs[:, prev_pos + 1 : cur_pos + 1] = -F.cross_entropy(
    input=logits.transpose(1, 2),
    target=tokens[:, prev_pos + 1 : cur_pos + 1],
    reduction="none",
    ignore_index=pad_id,
)

The use of reduction="none" ensures per-token losses are preserved rather than aggregated.

Return Value Structure

After generation completes, the tensor converts to a standard Python list (lines 214-215) and returns through a conditional tuple (lines 231-232):

token_logprobs = token_logprobs.tolist()
return (out_tokens, out_logprobs if logprobs else None)

The high-level helpers text_completion() and chat_completion() attach these values to their result dictionaries under the key "logprobs", aligning parallel lists of tokens and their scores (lines 74-81).

Code Examples

Direct generate() Call

Use this approach when you need raw token IDs and maximum control:

from llama import Llama

# Initialize the model

llama = Llama.build(
    ckpt_dir="llama-2-7b-chat/",
    tokenizer_path="tokenizer.model",
    max_seq_len=512,
    max_batch_size=4,
)

prompt = ["Hello, how are you?"]
prompt_ids = [llama.tokenizer.encode(p, bos=True, eos=False) for p in prompt]

# Enable log probability computation

tokens, logprobs = llama.generate(
    prompt_tokens=prompt_ids,
    max_gen_len=50,
    temperature=0.7,
    top_p=0.9,
    logprobs=True,  # <‑‑ Enable log probabilities

    echo=False,
)

print("Generated tokens:", tokens[0])
print("Log-probs per token:", logprobs[0])

Text Completion Helper

For decoded text with automatic token-to-string mapping:

result = llama.text_completion(
    prompts=["Explain quantum entanglement in one sentence."],
    max_gen_len=30,
    temperature=0.5,
    top_p=0.95,
    logprobs=True,  # <‑‑ Enable log probabilities

)

print(result[0]["generation"])  # Decoded text

print(result[0]["tokens"])      # Token strings

print(result[0]["logprobs"])    # Log-probability for each token

Chat Completion Helper

For dialog-based generation with per-token scoring:

dialog = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "What's the weather in Paris?"},
]

chat = llama.chat_completion(
    dialogs=[dialog],
    max_gen_len=20,
    temperature=0.8,
    top_p=0.9,
    logprobs=True,  # <‑‑ Enable log probabilities

)

print(chat[0]["generation"]["content"])  # Assistant reply

print(chat[0]["tokens"])                 # Token strings

print(chat[0]["logprobs"])               # Log-probs list

Key Source Files

  • llama/generation.py – Contains the core generate() method, logprobs flag handling, and cross-entropy calculation logic
  • llama/model.py – Defines the Transformer architecture and forward pass that produces the logits used for probability calculation
  • llama/tokenizer.py – Handles token encoding and decoding between strings and integer IDs
  • example_text_completion.py – Runnable demonstration showing logprobs usage with the text completion API
  • example_chat_completion.py – Runnable demonstration showing logprobs usage with the chat completion API

Summary

  • Enable log probabilities by setting logprobs=True in generate(), text_completion(), or chat_completion()
  • Implementation uses -F.cross_entropy with reduction="none" to compute per-token negative log-likelihoods from model logits
  • Storage occurs in a torch.zeros_like tensor that matches the token sequence shape, converted to Python lists before returning
  • Return format is Tuple[List[List[int]], Optional[List[List[float]]]] where the second element contains batch-wise lists of per-token log-probabilities (or None if disabled)
  • High-level APIs automatically attach log-probabilities to result dictionaries under the "logprobs" key alongside "tokens"

Frequently Asked Questions

What parameter enables token log probability computation in Llama 2?

Pass logprobs=True to the generate(), text_completion(), or chat_completion() methods. This boolean flag defaults to False and is defined in the method signature at line 31 of llama/generation.py.

How are log probabilities calculated internally?

The implementation applies PyTorch’s F.cross_entropy function with reduction="none" to the model’s logits. Specifically, it computes -F.cross_entropy(input=logits.transpose(1, 2), target=tokens, ...) for each token position, yielding the negative log-likelihood (natural logarithm of the predicted probability).

What is the return format when logprobs is enabled?

The methods return a tuple where the first element is a list of token ID lists (batch dimension included), and the second element is a list of lists containing float values representing the log probability of each corresponding token. When logprobs=False, the second element is None.

Can I retrieve log probabilities for prompt tokens as well as generated tokens?

Yes. When the echo parameter is set to True alongside logprobs=True, the returned lists include both the prompt tokens and the generated tokens with their respective log probabilities computed for all positions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →