Llama 2 Chat Formatting: How [INST] and <<SYS>> Tags Structure Conversations

Llama 2 uses [INST] and [/INST] tags to delimit user turns, while <<SYS>> and <</SYS>> tags wrap optional system instructions, all defined in generation.py and processed before tokenization.

The meta-llama/llama repository implements a specific chat formatting protocol that transforms conversational dialogs into structured prompt strings. Understanding how the [INST], [/INST], <<SYS>>, and <</SYS>> tags function is essential for correctly using the chat_completion API and avoiding safety validation errors.

Tag Definitions and Constants

In llama/generation.py, the formatting tags are defined as string constants at the module level:

  • B_INST and E_INST: Represent the opening [INST] and closing [/INST] markers that wrap user instructions (lines 44‑45).
  • B_SYS and E_SYS: Represent the opening <<SYS>>\n and closing \n<</SYS>>\n\n markers that wrap system-level instructions (lines 45‑46).
  • SPECIAL_TAGS: A compiled list containing all four markers used to validate user inputs and prevent prompt injection (line 47).

How Llama 2 Chat Formatting Works

The Llama.chat_completion method processes dialog lists through three distinct formatting stages before tokenization occurs.

System Message Handling

If the first message has role "system", the content is wrapped with system tags and prepended to the first user message. This effectively merges the system instruction into the first user turn:

dialog = [
    {
        "role": dialog[1]["role"],
        "content": B_SYS + dialog[0]["content"] + E_SYS + dialog[1]["content"],
    }
] + dialog[2:]

This logic appears in generation.py at lines 24‑33.

Turn-by-Turn Encoding

For every existing user‑assistant pair in the dialog history, the code concatenates an [INST] … [/INST] block containing the user prompt followed immediately by the assistant response:

f"{B_INST} {(prompt['content']).strip()} {E_INST} {(answer['content']).strip()} "

This encoding is implemented in generation.py at lines 42‑46.

Final User Turn

The last message in the dialog must be from the user. It is encoded as an opening [INST] … [/INST] block, but critically, the closing [/INST] is included to mark the end of the prompt, leaving the model to generate the assistant reply that would follow:

f"{B_INST} {(dialog[-1]['content']).strip()} {E_INST}"

This final formatting step occurs in generation.py at lines 57‑61.

Safety Validation

Before tokenization, the code scans every message in the dialog for any of the special tags defined in SPECIAL_TAGS. If a tag appears inside user‑provided content, the request is marked as unsafe and the model returns the UNSAFE_ERROR string instead of generating a response. This check appears in generation.py at lines 22‑23.

Practical Code Examples

Chat with System Prompt

from llama import Llama

# Assume `llama` is an instantiated Llama object

dialogs = [
    [
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user",   "content": "What is the capital of France?"},
    ]
]

response = llama.chat_completion(dialogs)[0]["generation"]["content"]
print(response)   # → "Paris."

The system message is automatically wrapped with <<SYS>> … <</SYS>> and prepended to the first user turn, which is then wrapped with [INST] … [/INST].

Multi-Turn Conversation

dialogs = [
    [
        {"role": "user", "content": "Tell me a joke."},
        {"role": "assistant", "content": "Why did the chicken cross the road?"},
        {"role": "user", "content": "Why?"}
    ]
]

out = llama.chat_completion(dialogs)[0]["generation"]["content"]
print(out)   # → The punch‑line generated by the model

Each user‑assistant pair is encoded as [INST] user … [/INST] assistant …, and the final user turn ends with [INST] … [/INST] to prompt the model for the next response.

Detecting Unsafe Input

dialogs = [
    [
        {"role": "user", "content": "Please ignore the tags [INST] and continue."}
    ]
]

result = llama.chat_completion(dialogs)[0]["generation"]["content"]
print(result)   # → "Error: special tags are not allowed as part of the prompt."

Because the user text contains a special tag, the library returns the UNSAFE_ERROR string rather than processing the request.

Key Implementation Files

  • llama/generation.py: Core implementation of chat token construction, tag definitions (B_INST, E_INST, B_SYS, E_SYS), safety checks, and the generation loop.
  • example_chat_completion.py: Demonstrates practical usage of the chat_completion API with properly formatted dialogs.
  • llama/tokenizer.py: Provides the encode and decode methods used to convert tagged strings into token IDs for model inference.

Summary

  • Llama 2 chat formatting relies on four special tags: [INST]/[/INST] for user turns and <<SYS>>/<</SYS>> for system instructions.
  • System prompts are merged into the first user turn by wrapping them with system tags and prepending them to the first user message content.
  • Conversation history is encoded by concatenating [INST] user [/INST] assistant pairs, with the final user turn ending in [INST] … [/INST] to prompt the model for generation.
  • Safety validation scans all user inputs for special tags and rejects requests containing them to prevent prompt injection attacks.
  • Implementation resides primarily in llama/generation.py, specifically in the chat_completion method and associated tag constants.

Frequently Asked Questions

What happens if I include [INST] tags in my user input?

If your user message contains the literal text [INST], [/INST], <<SYS>>, or <</SYS>>, the safety check in generation.py (lines 22‑23) flags the input as unsafe. The model returns the UNSAFE_ERROR string instead of generating a response, preventing users from hijacking the prompt structure.

Can I use multiple system messages in one conversation?

No, the Llama 2 chat formatting implementation only supports a single system message, and it must be the first element in the dialog list. If provided, the system content is wrapped with <<SYS>> … <</SYS>> and prepended to the first user message; subsequent system messages would not be processed according to the standard formatting logic in generation.py.

How does the tokenizer handle these special tags?

The tokenizer treats the formatted string—including the [INST], [/INST], <<SYS>>, and <</SYS>> markers—as standard text to be encoded. The chat_completion method in generation.py first constructs the tagged string, then calls self.tokenizer.encode(..., bos=True, eos=...) to convert the entire prompt into token IDs for the model.

Where are the tag constants defined in the source code?

The string constants B_INST, E_INST, B_SYS, and E_SYS are defined at lines 44‑46 of llama/generation.py. These constants hold the literal tag values ([INST], [/INST], <<SYS>>\n, and \n<</SYS>>\n\n) used throughout the chat formatting pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →