# nanochat | Andrej | Knowledge Base | Instagit

The best ChatGPT that $100 can buy.

GitHub Stars: 45.8k

Repository: https://github.com/karpathy/nanochat

---

## Articles

### [KV Cache Layout Used by NanoChat's Engine for Flash Attention 3 (FA3)](/karpathy/nanochat/nanochat-engine-kv-cache-layout-fa3)

Discover the KV cache layout NanoChat uses for Flash Attention 3 (FA3). Understand the `(n_layers, B, T, H, D)` tensor structure for efficient transformer inference.

- Tags: internals
- Published: 2026-03-10

### [How to Enable FP8 Training in NanoChat: Modes, Configuration, and Examples](/karpathy/nanochat/nanochat-enable-fp8-training-modes)

Learn to enable FP8 training in NanoChat with the --fp8 flag. Explore e4m3, e5m2, and auto modes to optimize for NVIDIA Hopper GPUs and improve training efficiency.

- Tags: how-to-guide
- Published: 2026-03-10

### [How NanoChat's Custom Linear Layer Improves Numerical Precision](/karpathy/nanochat/nanochat-custom-linear-layer-precision)

Discover how NanoChat's custom Linear layer uses float32 master weights for high-precision updates and fast low-precision matrix multiplication.

- Tags: deep-dive
- Published: 2026-03-10

### [How NanoChat Manages Computation Precision (dtype) on Different Hardware](/karpathy/nanochat/nanochat-computation-precision-dtype-hardware)

Discover how NanoChat optimizes computation precision (dtype) for various hardware. Learn about its automatic CUDA detection and manual override options for enhanced performance.

- Tags: deep-dive
- Published: 2026-03-10

### [MuonAdamW vs DistMuonAdamW: Key Differences for Single and Multi-GPU Training in nanochat](/karpathy/nanochat/nanochat-muonadamw-vs-distmuonadamw)

Explore MuonAdamW vs DistMuonAdamW for nanochat training. Understand single GPU vs. multi-GPU optimizations and ZeRO-2 sharding benefits without PyTorch DDP overhead.

- Tags: deep-dive
- Published: 2026-03-10

### [How nanochat Combines Muon and AdamW in a Single Optimizer](/karpathy/nanochat/nanochat-muon-adamw-optimizer-combination)

Discover how nanochat combines Muon and AdamW optimizers by partitioning parameters to accelerate training. Learn about its unique approach to tensor optimization for deep learning models.

- Tags: internals
- Published: 2026-03-10

### [ReLU² Activation Function in NanoChat: Implementation and Usage Guide](/karpathy/nanochat/nanochat-relu-squared-activation)

Discover the ReLU² activation function in NanoChat, defined as (max(0, x))². Learn its efficient implementation and how it amplifies positive signals more than standard ReLU.

- Tags: deep-dive
- Published: 2026-03-10

### [What Is QK Normalization in NanoChat's Attention Mechanism?](/karpathy/nanochat/nanochat-qk-normalization-attention)

Discover QK normalization in nanochat's attention mechanism. Learn how RMS normalization enhances query and key vectors before attention computation for improved performance.

- Tags: deep-dive
- Published: 2026-03-10

### [How Rotary Embeddings Are Implemented in NanoChat: A Complete Guide](/karpathy/nanochat/nanochat-rotary-embeddings-implementation)

Discover how NanoChat implements rotary embeddings RoPE as a drop-in replacement for absolute encodings. See the code and understand the application in this complete guide.

- Tags: how-to-guide
- Published: 2026-03-10

### [Modern Architectural Choices in NanoChat's GPT Model: Efficiency Meets State-of-the-Art](/karpathy/nanochat/nanochat-gpt-model-architecture)

Explore modern architectural choices like Rotary Embeddings, GQA, and FlashAttention-3 in NanoChat's GPT model. Achieve efficient inference on commodity GPUs.

- Tags: architecture
- Published: 2026-03-10

### [Which Depth Delivers GPT-2 Capability in NanoChat: The 26-Layer Benchmark](/karpathy/nanochat/nanochat-depth-gpt2-capability)

Discover the 26-layer benchmark for GPT-2 capability in nanochat. Explore the experimental findings detailing the optimal depth for advanced performance in the nanochat repository.

- Tags: deep-dive
- Published: 2026-03-10

### [Tokens-to-Parameters Ratio in NanoChat’s Compute-Optimal Models](/karpathy/nanochat/nanochat-tokens-parameters-ratio)

Discover the compute-optimal tokens-to-parameters ratio in NanoChat's models. Learn how 10.5 tokens per parameter optimize training for efficient AI.

- Tags: deep-dive
- Published: 2026-03-10

### [What Are the Main Stages of LLM Training Covered by nanochat?](/karpathy/nanochat/nanochat-llm-training-stages)

Nanochat focuses on inference only not LLM training stages. Discover the four pipeline steps for serving pre-trained models: loading weights, tokenizing, forward passes, and streaming output.

- Tags: deep-dive
- Published: 2026-03-10

### [How to Train a LLM on a Single GPU Node with NanoChat](/karpathy/nanochat/how-to-train-llm-single-gpu-nanochat)

Train a Transformer LLM on a single GPU node with NanoChat. Learn to install, configure, and run training scripts for efficient LLM development with compute-optimal scaling.

- Tags: tutorial
- Published: 2026-03-10

