# What Are the Main Stages of LLM Training Covered by nanochat?

> Nanochat focuses on inference only not LLM training stages. Discover the four pipeline steps for serving pre-trained models: loading weights, tokenizing, forward passes, and streaming output.

- Repository: [Andrej/nanochat](https://github.com/karpathy/nanochat)
- Tags: deep-dive
- Published: 2026-03-10

---

**`nanochat` does not actually implement any LLM training stages; it is an inference‑only repository that covers the four core pipeline steps required to serve a pre‑trained model: loading weights, tokenizing input, generating tokens via forward passes, and streaming the output.**

Unlike full training frameworks, Andrej Karpathy’s `nanochat` is deliberately minimal—it demonstrates how to chat with an already‑trained transformer rather than how to train one. The repository assumes you have a finished checkpoint (e.g., from Hugging Face or `nanoGPT`) and focuses exclusively on the runtime workflow that turns prompts into streaming text.

## Why nanochat Skips Traditional Training Stages

Traditional LLM development involves three major training phases: **pre‑training** on massive corpora, **supervised fine‑tuning** (SFT) on instruction data, and **reinforcement learning from human feedback** (RLHF). According to the source structure, `nanochat` contains no loss‑computation loops, no back‑propagation utilities, and no data loaders for training batches. Instead, the project is designed as a lightweight wrapper around a frozen model weights file.

If you need a reference implementation of actual training stages, Karpathy’s `nanoGPT` repository provides the pre‑training and fine‑tuning logic; `nanochat` simply consumes the artifacts produced by such systems.

## The Four Main Pipeline Stages nanochat Implements

While `nanochat` omits training, it does implement the complete inference pipeline across four well‑defined stages. Each stage maps to a specific source file in the repository.

### Model Loading

The first stage downloads and initializes the pre‑trained weights. In [`nanochat/main.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/main.py), the application loads a Hugging Face transformer checkpoint into memory, moves it to the appropriate device (CPU or CUDA), and prepares the model for evaluation mode.

```python

# Conceptual flow from nanochat/main.py

from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("gpt2")
model.eval()  # Inference mode only; no gradient computation

```

This step corresponds to **deployment preparation** rather than training, ensuring the frozen parameters are ready for forward passes.

### Tokenization

Before generation can begin, user prompts must be converted into token IDs. The [`nanochat/tokenizer.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/tokenizer.py) file provides a thin wrapper around the Hugging Face tokenizer, handling the text‑to‑integer mapping and attention mask creation.

```python

# From nanochat/tokenizer.py pattern

tokens = tokenizer.encode(prompt, return_tensors="pt")
input_ids = tokens.to(device)

```

This stage bridges human‑readable text and the numerical input space the transformer expects, but it does not involve any parameter updates or training‑data processing.

### Forward Pass and Sampling

The core generation logic lives in [`nanochat/generation.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/generation.py), specifically within the `generate()` method. This stage runs the forward pass to obtain logits, applies temperature scaling, and samples the next token using top‑k or top‑p filtering.

```python

# Representative pattern from nanochat/generation.py

def generate(self, input_ids, max_new_tokens, temperature=0.8):
    for _ in range(max_new_tokens):
        logits = model(input_ids).logits[:, -1, :] / temperature
        probs = torch.softmax(logits, dim=-1)
        next_token = torch.multinomial(probs, num_samples=1)
        input_ids = torch.cat([input_ids, next_token], dim=1)
        yield next_token

```

This is purely **inference‑time computation**; the weights remain frozen throughout the loop.

### Streaming Output

Finally, [`nanochat/ui.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/ui.py) implements the `stream_response()` function, which decodes generated token IDs back to text in real time and pushes them to the web interface via streaming chunks. This stage handles **output rendering** and user experience, completing the inference pipeline without ever touching training loss functions.

```python

# From nanochat/ui.py conceptual flow

for token_id in generator.generate(input_ids):
    text_chunk = tokenizer.decode(token_id, skip_special_tokens=True)
    yield text_chunk

```

## Complete Inference Example

Here is a practical example showing how `nanochat` orchestrates these four stages to serve a model locally:

```python
from nanochat import NanoChat

# Stage 1: Model Loading

chat = NanoChat(model_name="gpt2", device="cuda")

# Stages 2-4: Tokenize, Generate, and Stream

prompt = "Explain the difference between AI training and inference."
for chunk in chat.stream(prompt, temperature=0.7, max_new_tokens=256):
    print(chunk, end="", flush=True)

```

**Key details:**
- `NanoChat()` initializes the tokenizer and model weights (Stages 1 & 2).
- `chat.stream()` wraps the generation loop and decoding (Stages 3 & 4).
- No optimizer states or gradient buffers are allocated, confirming the absence of training.

## Summary

- **`nanochat` is inference‑only:** It does not implement pre‑training, fine‑tuning, or RLHF stages.
- **Four pipeline stages:** The repository covers **model loading** ([`main.py`](https://github.com/karpathy/nanochat/blob/main/main.py)), **tokenization** ([`tokenizer.py`](https://github.com/karpathy/nanochat/blob/main/tokenizer.py)), **forward pass and sampling** ([`generation.py`](https://github.com/karpathy/nanochat/blob/main/generation.py)), and **streaming output** ([`ui.py`](https://github.com/karpathy/nanochat/blob/main/ui.py)).
- **No back‑propagation:** The codebase explicitly operates in `eval()` mode and yields tokens without computing losses or updating weights.
- **Contrast with nanoGPT:** For actual LLM training stages (pre‑training and fine‑tuning), refer to the `nanoGPT` repository; `nanochat` strictly consumes finished checkpoints.

## Frequently Asked Questions

### Does nanochat support fine‑tuning or RLHF?

No. `nanochat` does not contain training loops, loss functions, or optimizers. It loads frozen model weights and runs inference only. If you require fine‑tuning or reinforcement learning from human feedback, you need a training framework such as `nanoGPT`, `transformers.Trainer`, or alignment libraries like TRL.

### What is the difference between nanochat and nanoGPT?

`nanoGPT` is a minimal implementation of a GPT‑style model that includes **training** code (pre‑training and fine‑tuning), whereas `nanochat` is a minimal **serving** interface that demonstrates how to chat with an existing model. Think of `nanoGPT` as the factory and `nanochat` as the showroom.

### Can I use nanochat with my own fine‑tuned model?

Yes. Since `nanochat` uses standard Hugging Face model loading via `AutoModelForCausalLM.from_pretrained()`, you can point it to any local directory or Hugging Face Hub repository containing a compatible transformer checkpoint, including fine‑tuned variants (e.g., LoRA adapters merged into base weights).

### Does nanochat implement the tokenization stage differently during training?

No, because `nanochat` never enters a training phase. Its tokenization logic (found in [`nanochat/tokenizer.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/tokenizer.py)) is used exclusively to encode prompts for the forward pass and decode output tokens for the user interface. The tokenizer remains frozen and does not update its vocabulary or embeddings.