# llama | Meta Llama | Knowledge Base | Instagit

Inference code for Llama models

GitHub Stars: 59.2k

Repository: https://github.com/meta-llama/llama

---

## Articles

### [How to Optimize Memory Usage When Running Llama 2 Inference with Large Batch Sizes](/meta-llama/llama/how-to-optimize-memory-usage-when-running-llama-2-inference-with-large-batch-sizes)

Optimize Llama 2 memory for large batch sizes. Reduce max_batch_size and use chunked inference to avoid out-of-memory errors and enhance GPU efficiency.

- Tags: performance
- Published: 2026-03-05

### [What Is the Difference Between ColumnParallelLinear and RowParallelLinear in Llama 2?](/meta-llama/llama/what-is-the-difference-between-columnparallellinear-and-rowparallellinear-in-llama-2)

Understand the difference between ColumnParallelLinear and RowParallelLinear in Llama 2. Learn how sharding affects weights and operations for efficient model parallelism.

- Tags: deep-dive
- Published: 2026-03-05

### [How Llama 2 Handles EOS Token Detection During Text Generation](/meta-llama/llama/how-does-llama-2-handle-eos-token-detection-during-the-generation-process)

Discover how Llama 2 detects EOS tokens during text generation. Learn how it efficiently stops sampling and trims output for optimized results.

- Tags: internals
- Published: 2026-03-05

### [How to Properly Format Prompts for Pretrained vs Chat-Finetuned Llama 2 Models](/meta-llama/llama/how-to-properly-format-prompts-for-pretrained-versus-chat-finetuned-llama-2-models)

Learn to format prompts for Llama 2 models. Understand input requirements for pretrained text completion and chat-finetuned message dictionaries to optimize your AI interactions.

- Tags: best-practices
- Published: 2026-03-05

### [How the Attention Mask Functions in Llama 2's Forward Pass](/meta-llama/llama/how-does-the-attention-mask-function-within-the-llama-2-forward-pass)

Discover how the Llama 2 attention mask combines padding and causal masks to enforce constraints and ignore tokens in the forward pass. Optimize your LLM understanding.

- Tags: internals
- Published: 2026-03-05

### [Precomputed Frequency Tensor for Rotary Embeddings in Llama: Purpose and Implementation](/meta-llama/llama/what-is-the-purpose-of-the-precomputed-frequency-tensor-for-rotary-embeddings)

Discover the purpose of Llama's precomputed frequency tensor for rotary embeddings. Learn how freqs_cis optimizes attention layers by avoiding costly trigonometric computations for efficient position encoding.

- Tags: deep-dive
- Published: 2026-03-05

### [How to Compute Token Log Probabilities During Llama 2 Generation](/meta-llama/llama/how-to-compute-token-log-probabilities-during-llama-2-generation)

Learn how to compute token log probabilities during Llama 2 generation by enabling the logprobs parameter in generate text completion or chat completion calls for detailed output.

- Tags: deep-dive
- Published: 2026-03-05

### [Difference Between Greedy Sampling and Nucleus (Top-p) Sampling in Llama](/meta-llama/llama/what-is-the-difference-between-greedy-sampling-and-nucleus-top-p-sampling)

Understand the difference between greedy sampling and nucleus (top-p) sampling in Llama. Discover how top-p sampling creates more diverse and creative text generation.

- Tags: deep-dive
- Published: 2026-03-05

### [How the Llama 2 Inference Loop Handles Variable-Length Prompt Batches](/meta-llama/llama/how-does-the-llama-2-inference-loop-handle-batches-with-variable-length-prompts)

Discover how the Llama 2 inference loop efficiently processes variable-length prompt batches by padding to a common size and using boolean masks to avoid influencing generation.

- Tags: internals
- Published: 2026-03-05

### [What Information is Stored in the params.json File in Llama 2 Checkpoints?](/meta-llama/llama/what-information-is-stored-in-the-params.json-file-within-llama-2-checkpoints)

Understand the params.json file in Llama 2 checkpoints. Discover how it stores model architecture, hidden dimensions, layer counts, and other crucial config data to rebuild transformers.

- Tags: internals
- Published: 2026-03-05

### [How to Resolve the "Model Parallel Size Does Not Match Checkpoint Files" Error in LLaMA](/meta-llama/llama/how-to-resolve-the-error-model-parallel-size-does-not-match-checkpoint-files)

Resolve the model parallel size does not match checkpoint files error in LLaMA by matching WORLD_SIZE to your checkpoint shards. Learn how to properly launch inference with torchrun.

- Tags: how-to-guide
- Published: 2026-03-05

### [Impact of max_batch_size on Llama 2 Memory Allocation: KV Cache Scaling Explained](/meta-llama/llama/what-is-the-impact-of-max_batch_size-on-memory-allocation-for-llama-2)

Understand how max_batch_size impacts Llama 2 memory allocation. Learn how KV cache scaling affects GPU memory consumption in transformer layers.

- Tags: performance
- Published: 2026-03-05

### [Llama 2 Chat Formatting: How [INST] and <<SYS>> Tags Structure Conversations](/meta-llama/llama/how-are-inst-and-sys-tags-used-in-llama-2-chat-formatting)

Master Llama 2 chat formatting. Learn how [INST] and <<SYS>> tags structure conversations with user turns and system instructions for better AI interaction.

- Tags: how-to-guide
- Published: 2026-03-05

### [Why KV Cache Pre-Allocation with max_seq_len Is Necessary for Efficient Llama Inference](/meta-llama/llama/why-is-kv-cache-pre-allocation-necessary-with-max_seq_len)

Discover why KV cache pre-allocation with max_seq_len is essential for efficient Llama inference. Learn how it optimizes memory, avoids dynamic reallocation, and enforces sequence length bounds for faster generation.

- Tags: performance
- Published: 2026-03-05

### [How to Customize Temperature and Top‑P Sampling Parameters During Generation in Llama](/meta-llama/llama/how-to-customize-temperature-and-top_p-sampling-parameters-during-generation)

Discover how to customize temperature and top_p sampling parameters for generation in Llama. Learn to control model output by adjusting these key settings in text generation.

- Tags: how-to-guide
- Published: 2026-03-05

### [Llama 2 text_completion vs chat_completion: API Differences and Usage Guide](/meta-llama/llama/what-are-the-differences-between-llama-2s-text_completion-and-chat_completion-apis)

Understand Llama 2 text completion vs chat completion API differences. Learn how to use these powerful tools for single-turn text generation and multi-turn dialogs effectively.

- Tags: tutorial
- Published: 2026-03-05

### [How Grouped Query Attention (GQA) Functions with n_kv_heads in Llama](/meta-llama/llama/how-does-grouped-query-attention-gqa-function-with-the-n_kv_heads-parameter)

Learn how Grouped Query Attention GQA works with n_kv_heads in Llama 2 to boost inference efficiency by sharing key value heads and reducing memory usage.

- Tags: internals
- Published: 2026-03-05

### [Rotary Position Embeddings (RoPE) in Llama 2: Architecture and Implementation](/meta-llama/llama/what-is-the-role-of-rotary-position-embeddings-rope-in-the-llama-2-transformer-architecture)

Discover how Rotary Position Embeddings (RoPE) in Llama 2 enhance Transformer architecture by encoding token positions with complex rotations for superior sequence extrapolation.

- Tags: internals
- Published: 2026-03-05

### [How to Implement KV Caching in Llama 2 for Efficient Inference](/meta-llama/llama/how-to-implement-kv-caching-in-llama-2-for-efficient-inference)

Discover how Llama 2 implements KV caching in its Attention class for faster inference. Learn to optimize token generation and avoid quadratic recomputation for efficient LLM performance.

- Tags: performance
- Published: 2026-03-05

