Llama 2 text_completion vs chat_completion: API Differences and Usage Guide
The text_completion API processes raw text strings for single-turn generation, while chat_completion handles structured multi-turn dialogs with special instruction tags and safety checks.
Llama 2 provides two distinct inference entry points in the meta-llama/llama repository that serve different interaction patterns. Understanding the differences between text_completion and chat_completion is essential for selecting the right approach for your application, as they differ in input formatting, prompt construction, safety validation, and output structure.
Core API Signatures and Parameters
Both methods share several generation parameters but accept fundamentally different input types.
In llama/generation.py, the text_completion method signature is:
def text_completion(
self,
prompts: List[str],
temperature: float = 0.6,
top_p: float = 0.9,
max_gen_len: Optional[int] = None,
logprobs: bool = False,
echo: bool = False,
) -> List[CompletionPrediction]
The chat_completion method signature differs primarily in its input parameter:
def chat_completion(
self,
dialogs: List[Dialog],
temperature: float = 0.6,
top_p: float = 0.9,
max_gen_len: Optional[int] = None,
logprobs: bool = False,
) -> List[ChatPrediction]
Note that chat_completion lacks the echo parameter and requires a Dialog type—a list of message dictionaries with role and content keys.
Input Format and Prompt Construction
The most significant difference lies in how each API processes and tokenizes input.
text_completion: Raw String Processing
The text_completion API treats each prompt as an independent text string. In llama/generation.py (lines 64-66), the tokenization is straightforward:
prompt_tokens = [
self.tokenizer.encode(p, bos=True, eos=False) for p in prompts
]
Each string is encoded with bos=True (beginning-of-sequence) and eos=False, then fed directly to the model without modification. This approach is ideal for single-turn completion tasks where you want the model to continue a text fragment.
chat_completion: Structured Dialog with Special Tags
The chat_completion API implements a conversation format using special instruction tokens. In llama/generation.py (lines 24-31), the code defines constants for these tags:
B_INST, E_INST = "[INST]", "[/INST]"
B_SYS, E_SYS = "<<SYS>>\n", "\n<</SYS>>\n\n"
When processing dialogs, the method constructs a single token stream that:
- Wraps system messages with
<<SYS>>and<</SYS>>tags - Wraps each user turn with
[INST]and[/INST]tags - Maintains alternating user/assistant message structure
The tokenization logic (lines 40-62) builds this formatted string before encoding, ensuring the model receives the instruction-tuned format it was trained on.
Safety Checks and Validation
The chat_completion API includes additional safety guards absent from text_completion.
In llama/generation.py (lines 22-23), the code defines forbidden tags:
SPECIAL_TAGS = [B_INST, E_INST, B_SYS, E_SYS]
The chat_completion method scans every message content for these tags. If detected, it returns UNSAFE_ERROR instead of generating text, preventing prompt injection attacks that could manipulate the instruction format.
Additionally, the method asserts proper role ordering (lines 34-39 and 55-57), validating that messages alternate correctly between user and assistant roles, with an optional system message at the start.
Output Format Differences
The return structures differ to match their input paradigms.
text_completion returns a list of dictionaries with a simple string generation:
[
{
"generation": " the pursuit of happiness and understanding...",
"tokens": [...], # if logprobs=True
"logprobs": [...] # if logprobs=True
}
]
chat_completion returns a list where generation contains a role-based message object:
[
{
"generation": {
"role": "assistant",
"content": "Mayonnaise is made by emulsifying egg yolks..."
},
"tokens": [...], # if logprobs=True
"logprobs": [...] # if logprobs=True
}
]
This structure aligns with modern chat API standards, making it easier to integrate into conversational applications.
Code Examples
Text Completion Example
The example_text_completion.py file demonstrates batch processing of independent prompts:
from llama import Llama
generator = Llama.build(
ckpt_dir="checkpoints/7B",
tokenizer_path="tokenizer.model",
max_seq_len=128,
max_batch_size=4,
)
prompts = [
"I believe the meaning of life is",
"Translate English to French:\nsea otter => ",
]
results = generator.text_completion(
prompts,
max_gen_len=64,
temperature=0.6,
top_p=0.9,
)
for prompt, result in zip(prompts, results):
print(f"> {prompt}\n{result['generation']}\n")
Chat Completion Example
The example_chat_completion.py file shows multi-turn conversation handling:
from llama import Llama, Dialog
from typing import List
generator = Llama.build(
ckpt_dir="checkpoints/7B",
tokenizer_path="tokenizer.model",
max_seq_len=512,
max_batch_size=8,
)
dialogs: List[Dialog] = [
[{"role": "user", "content": "What is the recipe of mayonnaise?"}],
[
{"role": "system", "content": "Always answer in haiku"},
{"role": "user", "content": "Describe the Eiffel Tower"},
],
]
results = generator.chat_completion(
dialogs,
max_gen_len=128,
temperature=0.7,
top_p=0.9,
)
for dialog, result in zip(dialogs, results):
for msg in dialog:
print(f"{msg['role'].capitalize()}: {msg['content']}")
print(f"Assistant: {result['generation']['content']}\n")
Summary
- Input Structure:
text_completionacceptsList[str]for raw text prompts, whilechat_completionrequiresList[Dialog]with structured role-based messages. - Prompt Processing: The text API tokenizes strings directly with
bos=True, eos=False; the chat API constructs formatted strings using[INST],[/INST],<<SYS>>, and<</SYS>>tags. - Safety Validation: Only
chat_completionscans for forbidden special tags and validates role ordering, returningUNSAFE_ERRORfor injection attempts. - Output Format: Text completion returns simple string generations; chat completion returns structured objects with
roleandcontentkeys matching modern chat API standards. - Use Cases: Use
text_completionfor single-turn continuation tasks andchat_completionfor multi-turn conversational agents requiring system instructions and context management.
Frequently Asked Questions
Can I use chat_completion for single-turn prompts?
Yes, you can use chat_completion for single-turn prompts by passing a dialog containing a single user message. However, the API will still wrap your input with [INST] and [/INST] tags and return a structured response object with role and content fields, adding slight overhead compared to text_completion.
Does text_completion support system prompts?
No, text_completion does not support system prompts or role-based messaging. It processes raw text strings directly without the <<SYS>> wrapper tags used by chat_completion. If you need to inject system-level instructions, you must manually prepend them to your prompt strings or switch to chat_completion.
Which API should I use for conversational agents?
Use chat_completion for conversational agents. According to the meta-llama/llama source code, this API handles multi-turn context management, enforces proper user/assistant alternation, and formats dialogs with the special instruction tags ([INST], [/INST]) that align with Llama 2's fine-tuning for chat scenarios.
Are there performance differences between the two APIs?
Both APIs delegate to the same underlying generate method in llama/generation.py, so core inference performance is identical. However, chat_completion incurs minor overhead from dialog validation, special tag insertion, and safety scanning for forbidden tokens (SPECIAL_TAGS), while text_completion provides slightly lower latency for simple batch processing of raw prompts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →