# Supported LLM Models in the Prompt Engineering Guide: A Complete Technical Reference

> Discover all 16+ supported LLM models in the Prompt Engineering Guide. Explore GPT-4, LLaMA, Mistral, Gemini, Sora & more. Get technical details on capabilities, limits, & strategies.

- Repository: [DAIR.AI/Prompt-Engineering-Guide](https://github.com/dair-ai/Prompt-Engineering-Guide)
- Tags: api-reference
- Published: 2026-03-03

---

**The Prompt Engineering Guide actively documents 16+ large language models—including GPT-4, LLaMA, Mistral, Mixtral, Gemini, and specialized architectures like Sora and Kimi-K2.5—each with dedicated pages under `pages/models/` that detail capabilities, token limits, and prompt engineering strategies.**

The [dair-ai/Prompt-Engineering-Guide](https://github.com/dair-ai/Prompt-Engineering-Guide) repository serves as the definitive open-source resource for prompt engineering techniques across modern AI architectures. Understanding the **supported LLM models in the Prompt Engineering Guide** enables developers to implement model-specific optimizations, from few-shot prompting with GPT-4 to instruction tuning with open-source alternatives like LLaMA-3 and Mixtral.

## Comprehensive List of Supported Models

As enumerated in `pages/models/collection.en.mdx` and cross-referenced in the README's "Models" section (lines 89-100), the guide maintains dedicated documentation for the following architectures.

### OpenAI Models

- **GPT-4**: Documented in `pages/models/gpt-4.en.mdx`, covering advanced reasoning, code generation, and system message optimization.
- **ChatGPT (`gpt-3.5-turbo`)**: Featured in the ChatGPT-specific documentation, focusing on conversational prompt patterns and cost-effective implementations.

### Meta and Open Source Architectures

- **LLaMA**: Found in `pages/models/llama.en.mdx`, detailing the original foundation model's prompt strategies.
- **LLaMA-3**: Covered in `pages/models/llama-3.en.mdx` with updated instruction formats and safety guidelines.
- **Code Llama**: Specialized variant for programming tasks, referenced in the README's model section.

### Mistral AI Family

- **Mistral-7B**: Documented in `pages/models/mistral-7b.en.mdx`, emphasizing efficient small-model prompting.
- **Mistral-Large**: Advanced reasoning model covered in `pages/models/mistral-large.en.mdx`.
- **Mixtral**: Sparse mixture-of-experts (MoE) architecture detailed in `pages/models/mixtral.en.mdx`.
- **Mixtral-8x22B**: High-parameter MoE implementation in `pages/models/mixtral-8x22b.en.mdx`.

### Specialized and Multimodal Models

- **Sora**: Text-to-video model documentation in `pages/models/sora.en.mdx`.
- **Kimi-K2.5**: Long-context processing model in `pages/models/kimi-k2.5.en.mdx`.
- **Grok-1**: Open-weight architecture documented in `pages/models/grok-1.en.mdx`.

### Additional Supported Architectures

- **Gemini**: Google's multimodal models, referenced in the README model listings.
- **Phi-2**: Microsoft's small language model in `pages/models/phi-2.en.mdx`.
- **OLMo**: AI2's open language model framework in `pages/models/olmo.en.mdx`.
- **Flan-T5 / Flan-UL2**: Instruction-tuned encoder-decoder models listed in the README's model section.

## Repository Structure for Model Documentation

According to the source code, model information is organized through a hierarchical documentation system.

### Central Registry

The `pages/models/collection.en.mdx` file functions as the authoritative index, enumerating every supported model with direct links to their respective documentation pages. This file serves as the primary navigation hub for comparing model capabilities.

### Individual Model Pages

Each LLM resides in its own markdown file following the naming convention `pages/models/<model-name>.en.mdx`. For example:

- `pages/models/gpt-4.en.mdx` contains OpenAI's GPT-4 specific parameters.
- `pages/models/llama-3.en.mdx` covers Meta's latest open-weight architecture.
- `pages/models/mixtral.en.mdx` details the MoE routing mechanisms relevant to prompt engineering.

These files include architecture overviews, context window specifications, token limits, and tailored example prompts.

## Practical Implementation Examples

The guide includes executable code patterns for interacting with supported models. Below are implementation templates derived from the repository's examples.

### OpenAI GPT-4 Integration

For commercial API access to GPT-4, use the `openai` Python client:

```python
import openai
import os

# Set your API key in an environment variable (do NOT hard-code it)

#   export OPENAI_API_KEY="sk-..."

openai.api_key = os.getenv("OPENAI_API_KEY")

response = openai.ChatCompletion.create(
    model="gpt-4",                     # ← supported model

    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user",   "content": "Explain the difference between few-shot and chain-of-thought prompting."}
    ],
    temperature=0.7,
)

print(response["choices"][0]["message"]["content"])

```

### Hugging Face Transformers for LLaMA-2

For open-source models like LLaMA, the guide recommends the `transformers` library:

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_name = "meta-llama/Llama-2-7b-chat-hf"   # ← listed in the guide

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.float16,
    device_map="auto"
)

prompt = "### Instruction:\nWrite a short poem about sunrise.\n### Response:"

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

output = model.generate(
    **inputs,
    max_new_tokens=100,
    temperature=0.8,
    do_sample=True,
    top_p=0.95,
)

print(tokenizer.decode(output[0], skip_special_tokens=True))

```

### vLLM Deployment for Mixtral-8x22B

For high-throughput inference with MoE models like Mixtral, the guide provides `vllm` server configurations:

```bash

# Start vllm server (run once)

vllm serve mixtral-8x22b-instruct \
  --port 8080 \
  --tensor-parallel-size 4

```

```python
import requests
import json

payload = {
    "model": "mixtral-8x22b-instruct",
    "prompt": "Summarize the article titled 'Prompt Engineering for LLMs' in three bullet points.",
    "max_tokens": 150,
    "temperature": 0.6,
}
resp = requests.post("http://localhost:8080/v1/completions", json=payload)
print(json.loads(resp.text)["choices"][0]["text"])

```

## Key Source Files

Understanding the repository layout helps locate specific model information quickly:

- **[`README.md`](https://github.com/dair-ai/Prompt-Engineering-Guide/blob/main/README.md)**: Contains the "Prompt Engineering – Models" navigation section (lines 89-100) linking to all supported architectures.
- **`pages/models/collection.en.mdx`**: Central registry enumerating every model page; the definitive source for the complete supported list.
- **`pages/models/<model>.en.mdx`**: Individual documentation files (e.g., `gpt-4.en.mdx`, `llama.en.mdx`) containing architecture details and prompt examples.
- **`guides/`**: Supplementary materials demonstrating advanced prompt techniques for specific models.
- **`components/`**: React components rendering model metadata in the Next.js frontend.

## Summary

The Prompt Engineering Guide provides comprehensive documentation for **16+ supported LLM models** ranging from commercial APIs to open-source weights. Key takeaways include:

- **Primary Registry**: `pages/models/collection.en.mdx` serves as the complete index of supported models.
- **File Locations**: Individual model documentation follows the pattern `pages/models/<model-name>.en.mdx`.
- **Model Categories**: Coverage includes OpenAI's GPT-4, Meta's LLaMA family, Mistral's MoE architectures, Google's Gemini, and specialized models like Sora and Kimi-K2.5.
- **Implementation Ready**: The repository provides working code examples for OpenAI API, Hugging Face Transformers, and vLLM deployment patterns.
- **Source of Truth**: The README's Models section (lines 89-100) and collection page are synchronized to reflect the currently supported LLM ecosystem.

## Frequently Asked Questions

### How do I find the documentation for a specific LLM in the Prompt Engineering Guide?

Navigate to the `pages/models/` directory and locate the file named `<model>.en.mdx` (e.g., `gpt-4.en.mdx` for GPT-4 or `mixtral.en.mdx` for Mixtral). Alternatively, consult `pages/models/collection.en.mdx`, which maintains an aggregated list of all supported models with direct links to their respective documentation pages.

### Does the guide include code examples for every supported LLM model?

While the guide provides implementation patterns for major categories—including OpenAI API usage, Hugging Face Transformers loading, and vLLM serving—the individual model pages focus on architectural specifics, token limits, and prompt strategies. The `guides/` directory contains additional code demonstrations for applying these models in real-world prompt engineering scenarios.

### What distinguishes Mixtral from Mixtral-8x22B in the documentation?

According to the source files, `pages/models/mixtral.en.mdx` covers the foundational Mixtral 8x7B mixture-of-experts architecture, while `pages/models/mixtral-8x22b.en.mdx` specifically documents the larger 8x22B parameter variant, offering higher capacity and broader context handling for complex reasoning tasks.

### Are video generation models like Sora documented differently than text LLMs?

Yes, `pages/models/sora.en.mdx` is included in the model collection despite Sora being a video generation model, indicating the guide's expansion beyond text-only LLMs. The documentation structure remains consistent—detailing capabilities, limitations, and prompt engineering strategies—but the content focuses on visual generation parameters rather than text token probabilities.