How to Format Instructions and Use Chat Templates for LLM Input
The train-llm-from-scratch repository provides a tokenizer-agnostic chat template implementation that converts role-based message lists into token sequences and binary loss masks, using plain-text role headers and a single end-of-text token for turn separation.
This minimal implementation, found in src/post_training/chat_template.py, handles the complete pipeline from conversation formatting to training-ready tokenization. It works with any tokenizer that supports byte-pair encoding (BPE) by treating role markers as ordinary text rather than special tokens, allowing the model to learn turn-taking through supervised fine-tuning (SFT).
Understanding the Chat Template Design
The chat template system centers on a single special token—the end-of-text marker—while keeping role identifiers as regular text tokens.
Special Tokens and Role Markers
The only true special token in this implementation is the end-of-text token (<|endoftext|>) with a fixed ID of 50256, defined as EOT_ID in src/post_training/chat_template.py (lines 38-40). This token serves dual purposes: it terminates each turn in the conversation and acts as the generation stop token during inference.
Role headers (<|user|>\n, <|assistant|>\n, <|system|>\n) are implemented as plain strings via USER_HEADER, ASSISTANT_HEADER, and SYSTEM_HEADER (lines 42-45). Because these are ordinary token strings rather than registered special tokens, they are encoded using standard BPE tokenization without requiring custom tokenizer configuration.
Conversation Format Structure
A single complete turn follows this strict pattern:
<|user|>
{user_content}<|endoftext|><|assistant|>
{assistant_content}<|endoftext|>
For inference scenarios where you want the model to generate a response, you append only the assistant header to cue generation. The system prompt, if present, appears at the beginning of the sequence using the <|system|>\n header format.
Encoding Conversations for Training and Inference
The module provides three primary utilities for converting between message dictionaries and model-ready inputs.
The encode_chat Function
The encode_chat function (lines 99-136) processes a list of message dictionaries and returns two aligned lists: token IDs and a binary loss mask. The encoding pipeline executes these steps:
- Encode the role header using
tiktoken.encode_ordinary - Encode the message content
- Append the EOT token (ID 50256)
- Build the mask concurrently, assigning 1 to assistant response tokens (including the final EOT) and 0 to all role headers, user content, and system prompts
from src.post_training.chat_template import encode_chat
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain the difference between SFT and RLHF."},
{"role": "assistant", "content": "Supervised fine-tuning uses labeled data..."}
]
token_ids, loss_mask = encode_chat(messages)
# loss_mask contains 1s only for assistant tokens and final EOT
The encode_prompt Function
For inference workflows, encode_prompt (line 95) returns only token IDs for a prompt that ends with an assistant header, preparing the sequence for autoregressive generation. This function sets add_generation_prompt=True internally, ensuring the model receives the cue to begin generating without any target tokens to mask.
from src.post_training.chat_template import encode_prompt
prompt_ids = encode_prompt(messages)
# Sequence ends with <|assistant|>\n, ready for generation
Rendering and Decoding Utilities
The render_chat function (line 73) produces a human-readable string representation for debugging, while decode (lines 139, 145) safely converts token IDs back to text, stripping any IDs greater than or equal to EOT_ID to prevent junk output.
from src.post_training.chat_template import render_chat
print(render_chat(messages, add_generation_prompt=True))
Preparing Instruction Data
Structure your conversation data as a list of dictionaries with specific role keys before encoding. The system message is optional but recommended for setting behavioral context.
conversation = [
{"role": "system", "content": "You are a concise coding assistant."},
{"role": "user", "content": "Write a Python function to reverse a string."},
{"role": "assistant", "content": "def reverse_string(s):\n return s[::-1]"}
]
When add_generation_prompt=True is passed to encoding functions, the sequence ends immediately after the <|assistant|>\n header, resulting in a loss mask of all zeros since no target tokens exist for prediction.
Applying Loss Masks During Training
The binary loss mask enables selective gradient computation on only the assistant's responses. Integrate this mask with your loss function to ignore predictions on user instructions and role headers.
import torch
import torch.nn.functional as F
# token_ids and loss_mask from encode_chat()
input_ids = torch.tensor([token_ids], dtype=torch.long)
mask = torch.tensor([loss_mask], dtype=torch.float)
logits, _ = model(input_ids) # shape [B, T, vocab_size]
logits = logits[:, :-1, :].reshape(-1, logits.size(-1))
targets = input_ids[:, 1:].reshape(-1)
loss = F.cross_entropy(logits, targets, reduction='none')
masked_loss = (loss * mask[:, 1:].reshape(-1)).mean()
This approach ensures the model learns to generate assistant responses while ignoring the prompt tokens during backpropagation, which is essential for both SFT and reinforcement learning from human feedback (RLHF) workflows.
Summary
- Tokenizer-agnostic design: Role headers (
<|user|>,<|assistant|>,<|system|>) are plain text strings encoded via standard BPE, requiring no special tokenizer configuration. - Single special token: Only
<|endoftext|>(ID 50256) is treated as special, serving as both turn delimiter and generation stop token. - Dual encoding modes: Use
encode_chatfor training (returns IDs + mask) andencode_promptfor inference (returns IDs ending with assistant header). - Binary masking: Loss masks contain 1s only for assistant tokens and the final EOT, enabling selective training on completions while ignoring prompts.
Frequently Asked Questions
What is the difference between encode_chat and encode_prompt?
The encode_chat function processes complete conversations including assistant responses, returning both token IDs and a binary loss mask where assistant tokens are marked with 1s. The encode_prompt function is designed for inference, returning only token IDs for the prompt ending with an assistant header, effectively setting the loss mask to all zeros since no target response exists yet.
Why are role headers implemented as plain text instead of special tokens?
Implementing role headers (<|user|>, <|assistant|>, etc.) as ordinary text strings makes the implementation tokenizer-agnostic. According to the source code in src/post_training/chat_template.py, this design avoids dependencies on custom tokenizer configurations and allows the model to learn turn-taking patterns through supervised fine-tuning rather than relying on hardcoded special token IDs.
How does the loss mask work in this implementation?
The loss mask is a binary array aligned with the token sequence, created in encode_chat (lines 99-136). It assigns 1 to tokens belonging to the assistant's content (including the final EOT token) and 0 to all system prompts, user messages, and role headers. When computing cross-entropy loss, this mask zeroes out gradients for non-assistant tokens, ensuring the model only learns to predict the assistant's responses.
Can I use this chat template with any tokenizer?
Yes, the implementation is specifically designed to work with any BPE-based tokenizer because it relies only on standard text encoding for role headers and the fixed end-of-text token ID 50256 (compatible with GPT-2 and similar tokenizers). The code uses tiktoken.encode_ordinary for the headers, meaning no custom special token registration is required in your tokenizer configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →