Supported LLM Models in the Prompt Engineering Guide: A Complete Technical Reference
The Prompt Engineering Guide actively documents 16+ large language models—including GPT-4, LLaMA, Mistral, Mixtral, Gemini, and specialized architectures like Sora and Kimi-K2.5—each with dedicated pages under pages/models/ that detail capabilities, token limits, and prompt engineering strategies.
The dair-ai/Prompt-Engineering-Guide repository serves as the definitive open-source resource for prompt engineering techniques across modern AI architectures. Understanding the supported LLM models in the Prompt Engineering Guide enables developers to implement model-specific optimizations, from few-shot prompting with GPT-4 to instruction tuning with open-source alternatives like LLaMA-3 and Mixtral.
Comprehensive List of Supported Models
As enumerated in pages/models/collection.en.mdx and cross-referenced in the README's "Models" section (lines 89-100), the guide maintains dedicated documentation for the following architectures.
OpenAI Models
- GPT-4: Documented in
pages/models/gpt-4.en.mdx, covering advanced reasoning, code generation, and system message optimization. - ChatGPT (
gpt-3.5-turbo): Featured in the ChatGPT-specific documentation, focusing on conversational prompt patterns and cost-effective implementations.
Meta and Open Source Architectures
- LLaMA: Found in
pages/models/llama.en.mdx, detailing the original foundation model's prompt strategies. - LLaMA-3: Covered in
pages/models/llama-3.en.mdxwith updated instruction formats and safety guidelines. - Code Llama: Specialized variant for programming tasks, referenced in the README's model section.
Mistral AI Family
- Mistral-7B: Documented in
pages/models/mistral-7b.en.mdx, emphasizing efficient small-model prompting. - Mistral-Large: Advanced reasoning model covered in
pages/models/mistral-large.en.mdx. - Mixtral: Sparse mixture-of-experts (MoE) architecture detailed in
pages/models/mixtral.en.mdx. - Mixtral-8x22B: High-parameter MoE implementation in
pages/models/mixtral-8x22b.en.mdx.
Specialized and Multimodal Models
- Sora: Text-to-video model documentation in
pages/models/sora.en.mdx. - Kimi-K2.5: Long-context processing model in
pages/models/kimi-k2.5.en.mdx. - Grok-1: Open-weight architecture documented in
pages/models/grok-1.en.mdx.
Additional Supported Architectures
- Gemini: Google's multimodal models, referenced in the README model listings.
- Phi-2: Microsoft's small language model in
pages/models/phi-2.en.mdx. - OLMo: AI2's open language model framework in
pages/models/olmo.en.mdx. - Flan-T5 / Flan-UL2: Instruction-tuned encoder-decoder models listed in the README's model section.
Repository Structure for Model Documentation
According to the source code, model information is organized through a hierarchical documentation system.
Central Registry
The pages/models/collection.en.mdx file functions as the authoritative index, enumerating every supported model with direct links to their respective documentation pages. This file serves as the primary navigation hub for comparing model capabilities.
Individual Model Pages
Each LLM resides in its own markdown file following the naming convention pages/models/<model-name>.en.mdx. For example:
pages/models/gpt-4.en.mdxcontains OpenAI's GPT-4 specific parameters.pages/models/llama-3.en.mdxcovers Meta's latest open-weight architecture.pages/models/mixtral.en.mdxdetails the MoE routing mechanisms relevant to prompt engineering.
These files include architecture overviews, context window specifications, token limits, and tailored example prompts.
Practical Implementation Examples
The guide includes executable code patterns for interacting with supported models. Below are implementation templates derived from the repository's examples.
OpenAI GPT-4 Integration
For commercial API access to GPT-4, use the openai Python client:
import openai
import os
# Set your API key in an environment variable (do NOT hard-code it)
# export OPENAI_API_KEY="sk-..."
openai.api_key = os.getenv("OPENAI_API_KEY")
response = openai.ChatCompletion.create(
model="gpt-4", # ← supported model
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain the difference between few-shot and chain-of-thought prompting."}
],
temperature=0.7,
)
print(response["choices"][0]["message"]["content"])
Hugging Face Transformers for LLaMA-2
For open-source models like LLaMA, the guide recommends the transformers library:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_name = "meta-llama/Llama-2-7b-chat-hf" # ← listed in the guide
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.float16,
device_map="auto"
)
prompt = "### Instruction:\nWrite a short poem about sunrise.\n### Response:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(
**inputs,
max_new_tokens=100,
temperature=0.8,
do_sample=True,
top_p=0.95,
)
print(tokenizer.decode(output[0], skip_special_tokens=True))
vLLM Deployment for Mixtral-8x22B
For high-throughput inference with MoE models like Mixtral, the guide provides vllm server configurations:
# Start vllm server (run once)
vllm serve mixtral-8x22b-instruct \
--port 8080 \
--tensor-parallel-size 4
import requests
import json
payload = {
"model": "mixtral-8x22b-instruct",
"prompt": "Summarize the article titled 'Prompt Engineering for LLMs' in three bullet points.",
"max_tokens": 150,
"temperature": 0.6,
}
resp = requests.post("http://localhost:8080/v1/completions", json=payload)
print(json.loads(resp.text)["choices"][0]["text"])
Key Source Files
Understanding the repository layout helps locate specific model information quickly:
README.md: Contains the "Prompt Engineering – Models" navigation section (lines 89-100) linking to all supported architectures.pages/models/collection.en.mdx: Central registry enumerating every model page; the definitive source for the complete supported list.pages/models/<model>.en.mdx: Individual documentation files (e.g.,gpt-4.en.mdx,llama.en.mdx) containing architecture details and prompt examples.guides/: Supplementary materials demonstrating advanced prompt techniques for specific models.components/: React components rendering model metadata in the Next.js frontend.
Summary
The Prompt Engineering Guide provides comprehensive documentation for 16+ supported LLM models ranging from commercial APIs to open-source weights. Key takeaways include:
- Primary Registry:
pages/models/collection.en.mdxserves as the complete index of supported models. - File Locations: Individual model documentation follows the pattern
pages/models/<model-name>.en.mdx. - Model Categories: Coverage includes OpenAI's GPT-4, Meta's LLaMA family, Mistral's MoE architectures, Google's Gemini, and specialized models like Sora and Kimi-K2.5.
- Implementation Ready: The repository provides working code examples for OpenAI API, Hugging Face Transformers, and vLLM deployment patterns.
- Source of Truth: The README's Models section (lines 89-100) and collection page are synchronized to reflect the currently supported LLM ecosystem.
Frequently Asked Questions
How do I find the documentation for a specific LLM in the Prompt Engineering Guide?
Navigate to the pages/models/ directory and locate the file named <model>.en.mdx (e.g., gpt-4.en.mdx for GPT-4 or mixtral.en.mdx for Mixtral). Alternatively, consult pages/models/collection.en.mdx, which maintains an aggregated list of all supported models with direct links to their respective documentation pages.
Does the guide include code examples for every supported LLM model?
While the guide provides implementation patterns for major categories—including OpenAI API usage, Hugging Face Transformers loading, and vLLM serving—the individual model pages focus on architectural specifics, token limits, and prompt strategies. The guides/ directory contains additional code demonstrations for applying these models in real-world prompt engineering scenarios.
What distinguishes Mixtral from Mixtral-8x22B in the documentation?
According to the source files, pages/models/mixtral.en.mdx covers the foundational Mixtral 8x7B mixture-of-experts architecture, while pages/models/mixtral-8x22b.en.mdx specifically documents the larger 8x22B parameter variant, offering higher capacity and broader context handling for complex reasoning tasks.
Are video generation models like Sora documented differently than text LLMs?
Yes, pages/models/sora.en.mdx is included in the model collection despite Sora being a video generation model, indicating the guide's expansion beyond text-only LLMs. The documentation structure remains consistent—detailing capabilities, limitations, and prompt engineering strategies—but the content focuses on visual generation parameters rather than text token probabilities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →