Supported LLM Models in the Prompt Engineering Guide: A Complete Technical Reference

The Prompt Engineering Guide actively documents 16+ large language models—including GPT-4, LLaMA, Mistral, Mixtral, Gemini, and specialized architectures like Sora and Kimi-K2.5—each with dedicated pages under pages/models/ that detail capabilities, token limits, and prompt engineering strategies.

The dair-ai/Prompt-Engineering-Guide repository serves as the definitive open-source resource for prompt engineering techniques across modern AI architectures. Understanding the supported LLM models in the Prompt Engineering Guide enables developers to implement model-specific optimizations, from few-shot prompting with GPT-4 to instruction tuning with open-source alternatives like LLaMA-3 and Mixtral.

Comprehensive List of Supported Models

As enumerated in pages/models/collection.en.mdx and cross-referenced in the README's "Models" section (lines 89-100), the guide maintains dedicated documentation for the following architectures.

OpenAI Models

  • GPT-4: Documented in pages/models/gpt-4.en.mdx, covering advanced reasoning, code generation, and system message optimization.
  • ChatGPT (gpt-3.5-turbo): Featured in the ChatGPT-specific documentation, focusing on conversational prompt patterns and cost-effective implementations.

Meta and Open Source Architectures

  • LLaMA: Found in pages/models/llama.en.mdx, detailing the original foundation model's prompt strategies.
  • LLaMA-3: Covered in pages/models/llama-3.en.mdx with updated instruction formats and safety guidelines.
  • Code Llama: Specialized variant for programming tasks, referenced in the README's model section.

Mistral AI Family

  • Mistral-7B: Documented in pages/models/mistral-7b.en.mdx, emphasizing efficient small-model prompting.
  • Mistral-Large: Advanced reasoning model covered in pages/models/mistral-large.en.mdx.
  • Mixtral: Sparse mixture-of-experts (MoE) architecture detailed in pages/models/mixtral.en.mdx.
  • Mixtral-8x22B: High-parameter MoE implementation in pages/models/mixtral-8x22b.en.mdx.

Specialized and Multimodal Models

  • Sora: Text-to-video model documentation in pages/models/sora.en.mdx.
  • Kimi-K2.5: Long-context processing model in pages/models/kimi-k2.5.en.mdx.
  • Grok-1: Open-weight architecture documented in pages/models/grok-1.en.mdx.

Additional Supported Architectures

  • Gemini: Google's multimodal models, referenced in the README model listings.
  • Phi-2: Microsoft's small language model in pages/models/phi-2.en.mdx.
  • OLMo: AI2's open language model framework in pages/models/olmo.en.mdx.
  • Flan-T5 / Flan-UL2: Instruction-tuned encoder-decoder models listed in the README's model section.

Repository Structure for Model Documentation

According to the source code, model information is organized through a hierarchical documentation system.

Central Registry

The pages/models/collection.en.mdx file functions as the authoritative index, enumerating every supported model with direct links to their respective documentation pages. This file serves as the primary navigation hub for comparing model capabilities.

Individual Model Pages

Each LLM resides in its own markdown file following the naming convention pages/models/<model-name>.en.mdx. For example:

  • pages/models/gpt-4.en.mdx contains OpenAI's GPT-4 specific parameters.
  • pages/models/llama-3.en.mdx covers Meta's latest open-weight architecture.
  • pages/models/mixtral.en.mdx details the MoE routing mechanisms relevant to prompt engineering.

These files include architecture overviews, context window specifications, token limits, and tailored example prompts.

Practical Implementation Examples

The guide includes executable code patterns for interacting with supported models. Below are implementation templates derived from the repository's examples.

OpenAI GPT-4 Integration

For commercial API access to GPT-4, use the openai Python client:

import openai
import os

# Set your API key in an environment variable (do NOT hard-code it)

#   export OPENAI_API_KEY="sk-..."

openai.api_key = os.getenv("OPENAI_API_KEY")

response = openai.ChatCompletion.create(
    model="gpt-4",                     # ← supported model

    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user",   "content": "Explain the difference between few-shot and chain-of-thought prompting."}
    ],
    temperature=0.7,
)

print(response["choices"][0]["message"]["content"])

Hugging Face Transformers for LLaMA-2

For open-source models like LLaMA, the guide recommends the transformers library:

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_name = "meta-llama/Llama-2-7b-chat-hf"   # ← listed in the guide

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.float16,
    device_map="auto"
)

prompt = "### Instruction:\nWrite a short poem about sunrise.\n### Response:"

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

output = model.generate(
    **inputs,
    max_new_tokens=100,
    temperature=0.8,
    do_sample=True,
    top_p=0.95,
)

print(tokenizer.decode(output[0], skip_special_tokens=True))

vLLM Deployment for Mixtral-8x22B

For high-throughput inference with MoE models like Mixtral, the guide provides vllm server configurations:


# Start vllm server (run once)

vllm serve mixtral-8x22b-instruct \
  --port 8080 \
  --tensor-parallel-size 4
import requests
import json

payload = {
    "model": "mixtral-8x22b-instruct",
    "prompt": "Summarize the article titled 'Prompt Engineering for LLMs' in three bullet points.",
    "max_tokens": 150,
    "temperature": 0.6,
}
resp = requests.post("http://localhost:8080/v1/completions", json=payload)
print(json.loads(resp.text)["choices"][0]["text"])

Key Source Files

Understanding the repository layout helps locate specific model information quickly:

  • README.md: Contains the "Prompt Engineering – Models" navigation section (lines 89-100) linking to all supported architectures.
  • pages/models/collection.en.mdx: Central registry enumerating every model page; the definitive source for the complete supported list.
  • pages/models/<model>.en.mdx: Individual documentation files (e.g., gpt-4.en.mdx, llama.en.mdx) containing architecture details and prompt examples.
  • guides/: Supplementary materials demonstrating advanced prompt techniques for specific models.
  • components/: React components rendering model metadata in the Next.js frontend.

Summary

The Prompt Engineering Guide provides comprehensive documentation for 16+ supported LLM models ranging from commercial APIs to open-source weights. Key takeaways include:

  • Primary Registry: pages/models/collection.en.mdx serves as the complete index of supported models.
  • File Locations: Individual model documentation follows the pattern pages/models/<model-name>.en.mdx.
  • Model Categories: Coverage includes OpenAI's GPT-4, Meta's LLaMA family, Mistral's MoE architectures, Google's Gemini, and specialized models like Sora and Kimi-K2.5.
  • Implementation Ready: The repository provides working code examples for OpenAI API, Hugging Face Transformers, and vLLM deployment patterns.
  • Source of Truth: The README's Models section (lines 89-100) and collection page are synchronized to reflect the currently supported LLM ecosystem.

Frequently Asked Questions

How do I find the documentation for a specific LLM in the Prompt Engineering Guide?

Navigate to the pages/models/ directory and locate the file named <model>.en.mdx (e.g., gpt-4.en.mdx for GPT-4 or mixtral.en.mdx for Mixtral). Alternatively, consult pages/models/collection.en.mdx, which maintains an aggregated list of all supported models with direct links to their respective documentation pages.

Does the guide include code examples for every supported LLM model?

While the guide provides implementation patterns for major categories—including OpenAI API usage, Hugging Face Transformers loading, and vLLM serving—the individual model pages focus on architectural specifics, token limits, and prompt strategies. The guides/ directory contains additional code demonstrations for applying these models in real-world prompt engineering scenarios.

What distinguishes Mixtral from Mixtral-8x22B in the documentation?

According to the source files, pages/models/mixtral.en.mdx covers the foundational Mixtral 8x7B mixture-of-experts architecture, while pages/models/mixtral-8x22b.en.mdx specifically documents the larger 8x22B parameter variant, offering higher capacity and broader context handling for complex reasoning tasks.

Are video generation models like Sora documented differently than text LLMs?

Yes, pages/models/sora.en.mdx is included in the model collection despite Sora being a video generation model, indicating the guide's expansion beyond text-only LLMs. The documentation structure remains consistent—detailing capabilities, limitations, and prompt engineering strategies—but the content focuses on visual generation parameters rather than text token probabilities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →