How to Compute Token Log Probabilities During Llama 2 Generation
Set logprobs=True when calling generate(), text_completion(), or chat_completion() to receive per-token log-probabilities alongside generated tokens.
The official meta-llama/llama repository provides built-in support to compute token log probabilities during text generation. The generate() method in llama/generation.py exposes a boolean logprobs parameter that calculates the negative log-likelihood for every token using PyTorch’s cross-entropy loss, enabling detailed confidence analysis without requiring model modifications.
Enabling Log Probability Computation
To compute token log probabilities, pass logprobs=True to any of the three primary generation methods. This flag is available in:
generate()– The core generation loop inllama/generation.py(lines 31-38)text_completion()– High-level helper that forwards the flag togenerate()(lines 62-70)chat_completion()– Dialog-aware helper with identical log probability support
When enabled, the method returns a tuple containing both the generated token IDs and their corresponding log-probabilities. If logprobs=False (the default), the second element of the return tuple is None.
How Log Probabilities Are Computed Internally
The implementation stores and computes log probabilities through a dedicated tensor that mirrors the shape of the generated sequence. The source code follows this exact flow:
Method Signature and Allocation
The generate() method accepts the flag with a default value of False and prepares storage when activated:
def generate(
...,
logprobs: bool = False,
...
) -> Tuple[List[List[int]], Optional[List[List[float]]]]:
When logprobs is True, the code immediately allocates a floating-point tensor matching the token buffer dimensions (lines 71-73 in llama/generation.py):
if logprobs:
token_logprobs = torch.zeros_like(tokens, dtype=torch.float)
Prompt-Only Sequences
If the input prompt fills the entire sequence length (min_prompt_len == total_len), the code computes cross-entropy loss for every prompt position before any generation begins (lines 78-84):
token_logprobs = -F.cross_entropy(
input=logits.transpose(1, 2),
target=tokens[:, prev_pos + 1 : cur_pos + 1],
reduction="none",
ignore_index=pad_id,
)
Iterative Generation Loop
During token generation, after each new token is sampled, the log probability is calculated for the just-generated token (and any preceding unprocessed positions). This occurs inside the main generation loop (lines 100-106):
token_logprobs[:, prev_pos + 1 : cur_pos + 1] = -F.cross_entropy(
input=logits.transpose(1, 2),
target=tokens[:, prev_pos + 1 : cur_pos + 1],
reduction="none",
ignore_index=pad_id,
)
The use of reduction="none" ensures per-token losses are preserved rather than aggregated.
Return Value Structure
After generation completes, the tensor converts to a standard Python list (lines 214-215) and returns through a conditional tuple (lines 231-232):
token_logprobs = token_logprobs.tolist()
return (out_tokens, out_logprobs if logprobs else None)
The high-level helpers text_completion() and chat_completion() attach these values to their result dictionaries under the key "logprobs", aligning parallel lists of tokens and their scores (lines 74-81).
Code Examples
Direct generate() Call
Use this approach when you need raw token IDs and maximum control:
from llama import Llama
# Initialize the model
llama = Llama.build(
ckpt_dir="llama-2-7b-chat/",
tokenizer_path="tokenizer.model",
max_seq_len=512,
max_batch_size=4,
)
prompt = ["Hello, how are you?"]
prompt_ids = [llama.tokenizer.encode(p, bos=True, eos=False) for p in prompt]
# Enable log probability computation
tokens, logprobs = llama.generate(
prompt_tokens=prompt_ids,
max_gen_len=50,
temperature=0.7,
top_p=0.9,
logprobs=True, # <‑‑ Enable log probabilities
echo=False,
)
print("Generated tokens:", tokens[0])
print("Log-probs per token:", logprobs[0])
Text Completion Helper
For decoded text with automatic token-to-string mapping:
result = llama.text_completion(
prompts=["Explain quantum entanglement in one sentence."],
max_gen_len=30,
temperature=0.5,
top_p=0.95,
logprobs=True, # <‑‑ Enable log probabilities
)
print(result[0]["generation"]) # Decoded text
print(result[0]["tokens"]) # Token strings
print(result[0]["logprobs"]) # Log-probability for each token
Chat Completion Helper
For dialog-based generation with per-token scoring:
dialog = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What's the weather in Paris?"},
]
chat = llama.chat_completion(
dialogs=[dialog],
max_gen_len=20,
temperature=0.8,
top_p=0.9,
logprobs=True, # <‑‑ Enable log probabilities
)
print(chat[0]["generation"]["content"]) # Assistant reply
print(chat[0]["tokens"]) # Token strings
print(chat[0]["logprobs"]) # Log-probs list
Key Source Files
llama/generation.py– Contains the coregenerate()method,logprobsflag handling, and cross-entropy calculation logicllama/model.py– Defines the Transformer architecture and forward pass that produces the logits used for probability calculationllama/tokenizer.py– Handles token encoding and decoding between strings and integer IDsexample_text_completion.py– Runnable demonstration showinglogprobsusage with the text completion APIexample_chat_completion.py– Runnable demonstration showinglogprobsusage with the chat completion API
Summary
- Enable log probabilities by setting
logprobs=Trueingenerate(),text_completion(), orchat_completion() - Implementation uses
-F.cross_entropywithreduction="none"to compute per-token negative log-likelihoods from model logits - Storage occurs in a
torch.zeros_liketensor that matches the token sequence shape, converted to Python lists before returning - Return format is
Tuple[List[List[int]], Optional[List[List[float]]]]where the second element contains batch-wise lists of per-token log-probabilities (orNoneif disabled) - High-level APIs automatically attach log-probabilities to result dictionaries under the
"logprobs"key alongside"tokens"
Frequently Asked Questions
What parameter enables token log probability computation in Llama 2?
Pass logprobs=True to the generate(), text_completion(), or chat_completion() methods. This boolean flag defaults to False and is defined in the method signature at line 31 of llama/generation.py.
How are log probabilities calculated internally?
The implementation applies PyTorch’s F.cross_entropy function with reduction="none" to the model’s logits. Specifically, it computes -F.cross_entropy(input=logits.transpose(1, 2), target=tokens, ...) for each token position, yielding the negative log-likelihood (natural logarithm of the predicted probability).
What is the return format when logprobs is enabled?
The methods return a tuple where the first element is a list of token ID lists (batch dimension included), and the second element is a list of lists containing float values representing the log probability of each corresponding token. When logprobs=False, the second element is None.
Can I retrieve log probabilities for prompt tokens as well as generated tokens?
Yes. When the echo parameter is set to True alongside logprobs=True, the returned lists include both the prompt tokens and the generated tokens with their respective log probabilities computed for all positions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →