# Building a User Interface for Interacting with Finetuned LLMs

> Create a user interface for interacting with finetuned LLMs. Explore the rasbt/LLMs-from-scratch repository for a Chainlit-based web interface to chat with GPT-2 models.

- Repository: [Sebastian Raschka/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)
- Tags: how-to-guide
- Published: 2026-05-12

---

**The rasbt/LLMs-from-scratch repository provides a minimal Chainlit-based web interface for chatting with a finetuned GPT-2 model, wrapping the core PyTorch implementation from Chapter 7 in an async HTTP/WebSocket server.**

Building a user interface for interacting with finetuned LLMs transforms training scripts into usable applications. The *LLMs-from-scratch* project demonstrates this by exposing its instruction-finetuned GPT-2 model through a lightweight Python web framework. This implementation connects the 355M-parameter checkpoint produced in Chapter 7 to a browser-based chat widget using fewer than 100 lines of UI-specific code.

## Architecture Overview

The UI follows a clean separation between the web layer and the model inference stack. Four core components work together to process user input and stream back generated text.

### Chainlit Web Server

**Chainlit** handles HTTP and WebSocket connections, rendering the chat widget and routing messages to Python callbacks. The entry point resides in [`ch07/06_user_interface/app.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/06_user_interface/app.py), where the `@chainlit.on_message` decorator registers an async handler that Chainlit invokes for every user submission. The server manages session state automatically and sends responses back to the browser without requiring manual socket management.

### Tokenizer and Model Loading

The system reuses the **tiktoken** BPE encoder to maintain consistency with the original GPT-2 vocabulary. Inside `get_model_and_tokenizer()` (lines 32–55 of [`app.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/app.py)), the code loads the finetuned weights from `gpt2-medium355M-sft.pth` using standard PyTorch APIs:

```python
def get_model_and_tokenizer():
    GPT_CONFIG_355M = {
        "vocab_size": 50257,
        "context_length": 1024,
        "emb_dim": 1024,
        "n_heads": 16,
        "n_layers": 24,
        "drop_rate": 0.0,
        "qkv_bias": True,
    }
    tokenizer = tiktoken.get_encoding("gpt2")
    checkpoint = torch.load(
        Path("..") / "01_main-chapter-code" / "gpt2-medium355M-sft.pth",
        weights_only=True
    )
    model = GPTModel(GPT_CONFIG_355M)
    model.load_state_dict(checkpoint)
    model.to(device)
    return tokenizer, model, GPT_CONFIG_355M

```

This function returns the tokenizer, the initialized `GPTModel` instance (defined in [`pkg/llms_from_scratch/ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch04.py)), and the configuration dictionary required for generation constraints.

### Generation Pipeline

Inference relies on the `generate()` utility from [`pkg/llms_from_scratch/ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py). The pipeline executes greedy decoding with a hard limit of 35 new tokens and respects the 1024-token context window. The flow truncates history automatically and stops at the EOS token ID 50256. Because the generation logic lives in the shared library package, the UI benefits from the same optimizations—such as no-grad inference and efficient token-to-text conversion—used throughout the book’s examples.

## Implementation Walkthrough

The UI logic in [`ch07/06_user_interface/app.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/06_user_interface/app.py) orchestrates prompt templating, tokenization, and response cleaning in a single async function.

### Message Handler Decorator

The `@chainlit.on_message` decorator marks the `main()` coroutine as the entry point for all chat interactions. When a user submits text, Chainlit packages it in a `chainlit.Message` object and passes it to this handler:

```python
@chainlit.on_message
async def main(message: chainlit.Message):
    torch.manual_seed(123)  # Deterministic output for demos

    prompt = f"""Below is an instruction that describes a task. Write a response
    that appropriately completes the request.

    ### Instruction:

    {message.content}
    """

```

The handler wraps the raw user content in an instruction template that matches the formatting used during supervised finetuning (SFT) in Chapter 7.

### Model Inference and Decoding

The handler converts the templated prompt to token IDs using `text_to_token_ids()`, moves the tensor to the target device, and calls the generation helper:

```python
    token_ids = generate(
        model=model,
        idx=text_to_token_ids(prompt, tokenizer).to(device),
        max_new_tokens=35,
        context_size=model_cfg["context_length"],
        eos_id=50256,
    )
    full_response = token_ids_to_text(token_ids, tokenizer)

```

The `generate()` function returns the complete sequence—including the prompt—so the UI must extract only the newly generated content.

### Response Extraction

The `extract_response()` helper (lines 60–62) strips the instruction boilerplate to isolate the model’s answer:

```python
def extract_response(full_text, prompt):
    return full_text[len(prompt):].replace("### Response:", "").strip()

```

After extraction, the handler sends the cleaned string back to the client:

```python
    response = extract_response(full_response, prompt)
    await chainlit.Message(content=response).send()

```

## Running the Interface

To launch the chat UI locally, install the dependencies and invoke Chainlit from the repository root:

```bash
pip install -r setup/02_installing-python-libraries/requirements.txt
chainlit run ch07/06_user_interface/app.py

```

The command starts a development server on `http://localhost:8000`. Opening this URL in a browser presents a chat window connected directly to the finetuned checkpoint. Each message triggers the async pipeline defined in [`app.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/app.py), providing immediate feedback without additional JavaScript or frontend configuration.

## Key Source Files

| File | Purpose |
|------|---------|
| [`ch07/06_user_interface/app.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/06_user_interface/app.py) | Full Chainlit UI implementation for the instruction-finetuned model |
| [`pkg/llms_from_scratch/ch04.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch04.py) | Core `GPTModel` Transformer implementation |
| [`pkg/llms_from_scratch/ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py) | `generate()`, `text_to_token_ids()`, and `token_ids_to_text()` utilities |
| [`ch07/01_main-chapter-code/gpt_instruction_finetuning.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/01_main-chapter-code/gpt_instruction_finetuning.py) | Training script that produces the `gpt2-medium355M-sft.pth` checkpoint loaded by the UI |
| [`ch05/06_user_interface/app_orig.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05/06_user_interface/app_orig.py) | Reference UI using the standard OpenAI GPT-2 checkpoint (pre-finetuning) |

These files demonstrate how the repository bridges educational PyTorch code and production deployment patterns while maintaining full reproducibility.

## Summary

- **Chainlit** provides the web layer, converting async Python functions into interactive chat interfaces via the `@chainlit.on_message` decorator.
- The UI loads a 355M-parameter instruction-finetuned GPT-2 model using PyTorch’s standard `load_state_dict()` in `get_model_and_tokenizer()`.
- Inference reuses the library’s `generate()` helper from [`pkg/llms_from_scratch/ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py), ensuring consistent behavior with the book’s training chapters.
- The `extract_response()` function isolates the model’s answer by removing the prompt template before displaying output to the user.
- The entire application runs with a single `chainlit run` command after installing requirements from the repository’s setup files.

## Frequently Asked Questions

### What framework does the LLMs-from-scratch UI use?

The interface uses **Chainlit**, a Python framework designed specifically for building conversational AI demos. Chainlit handles WebSocket connections, message routing, and the browser widget automatically, allowing developers to focus on model inference logic rather than frontend code.

### Can I use this UI with my own finetuned checkpoint?

Yes. Replace the `checkpoint` path in `get_model_and_tokenizer()` (line 43 of [`ch07/06_user_interface/app.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/06_user_interface/app.py)) with the path to your own `.pth` or `.pt` file. Ensure your model architecture matches the `GPT_CONFIG_355M` dictionary (1024 embedding dimensions, 24 layers, 16 heads) or update the configuration to match your saved weights.

### How does the UI prevent the model from exceeding the context window?

The `generate()` function in [`pkg/llms_from_scratch/ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py) automatically truncates input token IDs to fit within the `context_size` parameter (1024 tokens for this model). When the prompt plus requested `max_new_tokens` exceeds this limit, the function trims from the left, preserving the most recent context.

### Why is `torch.manual_seed(123)` set inside the message handler?

Setting a fixed random seed ensures **deterministic output** during demonstrations, making the UI behavior predictable for tutorials and debugging. In production deployments, you should remove this line to allow the model to sample varied responses, or implement proper temperature-based sampling in the `generate()` call.