How to Properly Format Prompts for Pretrained vs Chat-Finetuned Llama 2 Models
Pretrained Llama 2 models accept raw text strings via the text_completion method, while chat-finetuned variants require structured message dictionaries with role and content keys processed through chat_completion.
The meta-llama/llama repository implements distinct inference pathways for base and conversational model variants. Understanding how to properly format prompts for pretrained versus chat-finetuned Llama 2 models ensures correct tokenization and generation behavior according to the source implementation.
Architectural Differences in Prompt Handling
The Llama class defined in llama/model.py exposes two primary generation methods that handle input formatting differently based on the model's training regime.
Pretrained Base Models
Pretrained Llama 2 checkpoints are causal language models trained on raw text corpora. The text_completion method expects a List[str] containing raw prompt strings. According to the implementation in llama/generation.py, these strings are tokenized using the Tokenizer class from llama/tokenizer.py, which automatically appends the EOS token before feeding the sequence into the transformer.
Chat-Finetuned Models
Chat-finetuned Llama 2 models undergo additional fine-tuning on conversational data formatted with special role tokens. The chat_completion method requires a List[Dict[str, str]] where each dictionary contains:
role: One of"system","user", or"assistant"content: The message text
As implemented in llama/generation.py, this method constructs the token sequence by injecting special tokens such as <|system|>, <|user|>, <|assistant|>, and <|end|> between messages before generation.
Code Examples for Each Model Type
Pretrained Text Completion
When working with base models, pass raw text prompts directly to text_completion. This pattern from example_text_completion.py demonstrates the implementation:
from llama import Llama
generator = Llama.build(
ckpt_dir="path/to/pretrained_ckpt",
tokenizer_path="path/to/tokenizer.model",
max_seq_len=128,
max_batch_size=4,
)
prompts = [
"I believe the meaning of life is",
"Simply put, the theory of relativity states that ",
]
results = generator.text_completion(
prompts,
max_gen_len=64,
temperature=0.6,
top_p=0.9,
)
The method returns a list of dictionaries containing the generated text under the generation key.
Chat Completion Format
For conversational models, structure your input as a message list. This example from example_chat_completion.py shows the proper format:
from llama import Llama
chatbot = Llama.build(
ckpt_dir="path/to/llama-2-7b-chat",
tokenizer_path="path/to/tokenizer.model",
max_seq_len=256,
max_batch_size=2,
)
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the capital of France?"},
]
result = chatbot.chat_completion(
messages,
max_gen_len=80,
temperature=0.7,
top_p=0.95,
)
The chat_completion method automatically handles conversion from message dictionaries to the underlying token sequence required by chat-finetuned checkpoints.
Key Implementation Files
Several files in the meta-llama/llama repository define the prompt formatting logic:
llama/model.py: Contains theLlamaclass withtext_completionandchat_completionmethod signaturesllama/generation.py: Implements the generation loops and message-to-token conversion for chat modelsllama/tokenizer.py: Handles tokenization including special token insertion for chat formatsexample_text_completion.py: Demonstrates raw prompt formatting for pretrained modelsexample_chat_completion.py: Shows structured message formatting for chat-finetuned models
Summary
- Pretrained models require raw text strings passed to
text_completion, processed as continuous token streams with automatic EOS handling - Chat-finetuned models require
List[Dict]inputs withroleandcontentkeys passed tochat_completion, which injects special role tokens like<|system|>and<|end|> - Both model types use the same
Llama.buildconstructor but expect different input formats through their respective generation methods - The tokenization and special token handling occurs automatically in
llama/generation.pyandllama/tokenizer.py
Frequently Asked Questions
Can I use raw text prompts with chat-finetuned Llama 2 models?
While technically possible to tokenize raw strings, chat-finetuned models expect the specific role-based token structure provided by chat_completion. Using raw text bypasses the special token formatting that conditions the model for conversational behavior, resulting in suboptimal outputs.
What roles are supported in the chat completion format?
The chat_completion method supports three roles as defined in the chat training data: "system" for setting behavioral context, "user" for human queries, and "assistant" for model responses. These roles map to special tokens <|system|>, <|user|>, and <|assistant|> in the underlying implementation within llama/generation.py.
How does the tokenizer handle different prompt formats?
The Tokenizer class in llama/tokenizer.py processes raw text directly for base models, while llama/generation.py pre-processes chat messages by concatenating role tokens and content before tokenization. This ensures chat models receive the exact token pattern seen during fine-tuning.
Do I need to manually add EOS tokens to prompts?
No. The generation infrastructure automatically appends the EOS token to prompts in text_completion and handles the <|end|> token sequence for chat messages within the chat_completion method. Manual addition of these tokens is unnecessary and may cause generation issues.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →