How to Configure Prompt Styles in Private-GPT: Llama 2, Llama 3, Mistral, ChatML, and Tag

Private-GPT supports six prompt styles—llama2, llama3, mistral, chatml, tag, and default—that format chat messages into LLM-specific text templates, configured via the prompt_style field in your YAML settings.

Private-GPT is an open-source project that provides a customizable interface for running local large language models. Selecting the correct prompt style ensures that system and user messages are formatted according to the specific template expectations of models like Llama 2, Llama 3, or Mistral. This guide explains the supported formats and the exact configuration steps based on the source code implementation.

Supported Prompt Styles in Private-GPT

Private-GPT implements six distinct prompt styles in private_gpt/components/llm/prompt_helper.py, each extending AbstractPromptStyle to define specific token formatting.

Llama 2

The llama2 style uses the classic Llama 2 chat format with special tokens including <s>, [INST], <<SYS>>, and <</SYS>>. This wraps system prompts within the instruction block and separates user turns with [/INST]. The implementation resides in the Llama2PromptStyle class at line 72.

Llama 3

The llama3 style implements the newer Llama 3 chat format utilizing <|begin_of_text|> and <|eot_id|> tokens to demarcate conversation turns. This format is defined in Llama3PromptStyle at line 142 of prompt_helper.py.

Mistral

The mistral style merges system and user instructions into a single [INST] block wrapped in <s> tags, omitting separate system delimiters. This matches the Mistral instruction format and is implemented in MistralPromptStyle at line 41.

ChatML

The chatml style utilizes OpenAI-compatible ChatML syntax with <|im_start|> and <|im_end|> tokens to separate system, user, and assistant roles. The ChatMLPromptStyle class at line 66 handles this formatting.

Tag

The tag style applies a simple colon-separated format using tags like <|system|>:, <|user|>:, and <|assistant|>: to prefix content. This lightweight approach is implemented in TagPromptStyle at line 8.

Default

Setting default or omitting the field delegates formatting to the underlying llama_index library's role-based implementation without custom templating. The DefaultPromptStyle class at line 50 provides this passthrough behavior.

How to Configure Prompt Styles

You configure the prompt style through the LLMSettings class defined in private_gpt/settings/settings.py at lines 39-46. Set the prompt_style field in your YAML configuration file to one of the six supported identifiers.


# settings.yaml (or environment-specific overlay like settings-ollama.yaml)

llm:
  mode: ollama          # or llamacpp, openai, etc.

  temperature: 0.1
  max_new_tokens: 256
  prompt_style: llama3  # Options: default, llama2, llama3, tag, mistral, chatml

The component reads this value at runtime to determine which formatting class to instantiate.

Runtime Resolution and LLM Injection

The selection process follows a factory pattern that decouples formatting from backend implementation.

First, the LLMComponent class in private_gpt/components/llm/llm_component.py calls get_prompt_style() at line 54, passing the configured string. This factory function, located at line 88 of prompt_helper.py, maps the identifier to the corresponding concrete class.

Next, the component injects the style's formatting callbacks into the LLM backend. For example, when using LlamaCPP, lines 73-75 pass messages_to_prompt and completion_to_prompt from the selected style to the constructor:


# From private_gpt/components/llm/llm_component.py

prompt_style = get_prompt_style(settings.llm.prompt_style)

self.llm = LlamaCPP(
    # ... other parameters ...

    messages_to_prompt=prompt_style.messages_to_prompt,
    completion_to_prompt=prompt_style.completion_to_prompt,
)

This architecture allows you to switch between Llama 2 and Llama 3 formats instantly by changing the configuration value, without modifying the LLM initialization code.

Programmatic Access

You can also access prompt styles programmatically for testing or custom implementations:

from private_gpt.settings import settings
from private_gpt.components.llm.prompt_helper import get_prompt_style
from llama_index.core.llms import ChatMessage, MessageRole

# Retrieve configured style name

style_name = settings().llm.prompt_style  # e.g., "mistral"

# Get the concrete formatter

prompt_style = get_prompt_style(style_name)

# Convert messages to raw prompt

messages = [
    ChatMessage(content="You are an assistant.", role=MessageRole.SYSTEM),
    ChatMessage(content="Explain quantum computing.", role=MessageRole.USER),
]

raw_prompt = prompt_style.messages_to_prompt(messages)
print(raw_prompt)

The get_prompt_style factory at line 88 handles the string-to-class mapping, returning an instance of the appropriate AbstractPromptStyle subclass.

Summary

  • Six built-in styles: llama2, llama3, mistral, chatml, tag, and default are implemented in private_gpt/components/llm/prompt_helper.py.
  • Configuration: Set prompt_style in your YAML config (e.g., settings.yaml) under the llm section.
  • Factory pattern: get_prompt_style() at line 88 maps configuration strings to concrete classes.
  • Runtime injection: LLMComponent passes the style's messages_to_prompt and completion_to_prompt methods to the LLM backend at lines 54 and 73-75 of llm_component.py.
  • Decoupled design: Switching models requires only a configuration change, not code modification.

Frequently Asked Questions

What is the default prompt style if I omit the configuration?

If you omit the prompt_style field, Private-GPT defaults to llama2 as specified in the LLMSettings class at line 39 of private_gpt/settings/settings.py. You can also explicitly set the value to default to use the standard llama_index formatting without custom templates.

Can I use different prompt styles with Ollama or OpenAI-compatible backends?

Yes. The prompt style system is decoupled from the LLM backend. Whether you configure mode: ollama, mode: llamacpp, or mode: openai, the prompt_style setting applies universally. The LLMComponent injects the formatting callbacks regardless of the underlying API, though the specific effect depends on whether the backend accepts custom prompt functions.

Which prompt style should I use for Meta's Llama 3 models?

Select the llama3 style for Meta's Llama 3 models. This format uses <|begin_of_text|> and <|eot_id|> tokens as implemented in Llama3PromptStyle at line 142 of prompt_helper.py, aligning with the model's expected chat template.

How does the ChatML format differ from the Tag format?

ChatML uses XML-like tags with <|im_start|> and <|im_end|> delimiters to wrap role-specific content, while the Tag style uses a simpler colon-separated prefix format like <|system|>: and <|user|>:. ChatML is implemented at line 66 and Tag at line 8 of prompt_helper.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →