llama

Inference code for Llama models

19 articles 59.2k View on GitHub ↗
19 articles
How to Optimize Memory Usage When Running Llama 2 Inference with Large Batch Sizes

Optimize Llama 2 memory for large batch sizes. Reduce max_batch_size and use chunked inference to avoid out-of-memory errors and enhance GPU efficiency.

performance
Mar 5, 2026
What Is the Difference Between ColumnParallelLinear and RowParallelLinear in Llama 2?

Understand the difference between ColumnParallelLinear and RowParallelLinear in Llama 2. Learn how sharding affects weights and operations for efficient model parallelism.

deep-dive
Mar 5, 2026
How Llama 2 Handles EOS Token Detection During Text Generation

Discover how Llama 2 detects EOS tokens during text generation. Learn how it efficiently stops sampling and trims output for optimized results.

internals
Mar 5, 2026
How to Properly Format Prompts for Pretrained vs Chat-Finetuned Llama 2 Models

Learn to format prompts for Llama 2 models. Understand input requirements for pretrained text completion and chat-finetuned message dictionaries to optimize your AI interactions.

best-practices
Mar 5, 2026
How the Attention Mask Functions in Llama 2's Forward Pass

Discover how the Llama 2 attention mask combines padding and causal masks to enforce constraints and ignore tokens in the forward pass. Optimize your LLM understanding.

internals
Mar 5, 2026
Precomputed Frequency Tensor for Rotary Embeddings in Llama: Purpose and Implementation

Discover the purpose of Llama's precomputed frequency tensor for rotary embeddings. Learn how freqs_cis optimizes attention layers by avoiding costly trigonometric computations for efficient position encoding.

deep-dive
Mar 5, 2026
How to Compute Token Log Probabilities During Llama 2 Generation

Learn how to compute token log probabilities during Llama 2 generation by enabling the logprobs parameter in generate text completion or chat completion calls for detailed output.

deep-dive
Mar 5, 2026
Difference Between Greedy Sampling and Nucleus (Top-p) Sampling in Llama

Understand the difference between greedy sampling and nucleus (top-p) sampling in Llama. Discover how top-p sampling creates more diverse and creative text generation.

deep-dive
Mar 5, 2026
How the Llama 2 Inference Loop Handles Variable-Length Prompt Batches

Discover how the Llama 2 inference loop efficiently processes variable-length prompt batches by padding to a common size and using boolean masks to avoid influencing generation.

internals
Mar 5, 2026
What Information is Stored in the params.json File in Llama 2 Checkpoints?

Understand the params.json file in Llama 2 checkpoints. Discover how it stores model architecture, hidden dimensions, layer counts, and other crucial config data to rebuild transformers.

internals
Mar 5, 2026
How to Resolve the "Model Parallel Size Does Not Match Checkpoint Files" Error in LLaMA

Resolve the model parallel size does not match checkpoint files error in LLaMA by matching WORLD_SIZE to your checkpoint shards. Learn how to properly launch inference with torchrun.

how-to-guide
Mar 5, 2026
Impact of max_batch_size on Llama 2 Memory Allocation: KV Cache Scaling Explained

Understand how max_batch_size impacts Llama 2 memory allocation. Learn how KV cache scaling affects GPU memory consumption in transformer layers.

performance
Mar 5, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →