Building a User Interface for Interacting with Finetuned LLMs

The rasbt/LLMs-from-scratch repository provides a minimal Chainlit-based web interface for chatting with a finetuned GPT-2 model, wrapping the core PyTorch implementation from Chapter 7 in an async HTTP/WebSocket server.

Building a user interface for interacting with finetuned LLMs transforms training scripts into usable applications. The LLMs-from-scratch project demonstrates this by exposing its instruction-finetuned GPT-2 model through a lightweight Python web framework. This implementation connects the 355M-parameter checkpoint produced in Chapter 7 to a browser-based chat widget using fewer than 100 lines of UI-specific code.

Architecture Overview

The UI follows a clean separation between the web layer and the model inference stack. Four core components work together to process user input and stream back generated text.

Chainlit Web Server

Chainlit handles HTTP and WebSocket connections, rendering the chat widget and routing messages to Python callbacks. The entry point resides in ch07/06_user_interface/app.py, where the @chainlit.on_message decorator registers an async handler that Chainlit invokes for every user submission. The server manages session state automatically and sends responses back to the browser without requiring manual socket management.

Tokenizer and Model Loading

The system reuses the tiktoken BPE encoder to maintain consistency with the original GPT-2 vocabulary. Inside get_model_and_tokenizer() (lines 32–55 of app.py), the code loads the finetuned weights from gpt2-medium355M-sft.pth using standard PyTorch APIs:

def get_model_and_tokenizer():
    GPT_CONFIG_355M = {
        "vocab_size": 50257,
        "context_length": 1024,
        "emb_dim": 1024,
        "n_heads": 16,
        "n_layers": 24,
        "drop_rate": 0.0,
        "qkv_bias": True,
    }
    tokenizer = tiktoken.get_encoding("gpt2")
    checkpoint = torch.load(
        Path("..") / "01_main-chapter-code" / "gpt2-medium355M-sft.pth",
        weights_only=True
    )
    model = GPTModel(GPT_CONFIG_355M)
    model.load_state_dict(checkpoint)
    model.to(device)
    return tokenizer, model, GPT_CONFIG_355M

This function returns the tokenizer, the initialized GPTModel instance (defined in pkg/llms_from_scratch/ch04.py), and the configuration dictionary required for generation constraints.

Generation Pipeline

Inference relies on the generate() utility from pkg/llms_from_scratch/ch05.py. The pipeline executes greedy decoding with a hard limit of 35 new tokens and respects the 1024-token context window. The flow truncates history automatically and stops at the EOS token ID 50256. Because the generation logic lives in the shared library package, the UI benefits from the same optimizations—such as no-grad inference and efficient token-to-text conversion—used throughout the book’s examples.

Implementation Walkthrough

The UI logic in ch07/06_user_interface/app.py orchestrates prompt templating, tokenization, and response cleaning in a single async function.

Message Handler Decorator

The @chainlit.on_message decorator marks the main() coroutine as the entry point for all chat interactions. When a user submits text, Chainlit packages it in a chainlit.Message object and passes it to this handler:

@chainlit.on_message
async def main(message: chainlit.Message):
    torch.manual_seed(123)  # Deterministic output for demos

    prompt = f"""Below is an instruction that describes a task. Write a response
    that appropriately completes the request.

    ### Instruction:

    {message.content}
    """

The handler wraps the raw user content in an instruction template that matches the formatting used during supervised finetuning (SFT) in Chapter 7.

Model Inference and Decoding

The handler converts the templated prompt to token IDs using text_to_token_ids(), moves the tensor to the target device, and calls the generation helper:

    token_ids = generate(
        model=model,
        idx=text_to_token_ids(prompt, tokenizer).to(device),
        max_new_tokens=35,
        context_size=model_cfg["context_length"],
        eos_id=50256,
    )
    full_response = token_ids_to_text(token_ids, tokenizer)

The generate() function returns the complete sequence—including the prompt—so the UI must extract only the newly generated content.

Response Extraction

The extract_response() helper (lines 60–62) strips the instruction boilerplate to isolate the model’s answer:

def extract_response(full_text, prompt):
    return full_text[len(prompt):].replace("### Response:", "").strip()

After extraction, the handler sends the cleaned string back to the client:

    response = extract_response(full_response, prompt)
    await chainlit.Message(content=response).send()

Running the Interface

To launch the chat UI locally, install the dependencies and invoke Chainlit from the repository root:

pip install -r setup/02_installing-python-libraries/requirements.txt
chainlit run ch07/06_user_interface/app.py

The command starts a development server on http://localhost:8000. Opening this URL in a browser presents a chat window connected directly to the finetuned checkpoint. Each message triggers the async pipeline defined in app.py, providing immediate feedback without additional JavaScript or frontend configuration.

Key Source Files

File Purpose
ch07/06_user_interface/app.py Full Chainlit UI implementation for the instruction-finetuned model
pkg/llms_from_scratch/ch04.py Core GPTModel Transformer implementation
pkg/llms_from_scratch/ch05.py generate(), text_to_token_ids(), and token_ids_to_text() utilities
ch07/01_main-chapter-code/gpt_instruction_finetuning.py Training script that produces the gpt2-medium355M-sft.pth checkpoint loaded by the UI
ch05/06_user_interface/app_orig.py Reference UI using the standard OpenAI GPT-2 checkpoint (pre-finetuning)

These files demonstrate how the repository bridges educational PyTorch code and production deployment patterns while maintaining full reproducibility.

Summary

  • Chainlit provides the web layer, converting async Python functions into interactive chat interfaces via the @chainlit.on_message decorator.
  • The UI loads a 355M-parameter instruction-finetuned GPT-2 model using PyTorch’s standard load_state_dict() in get_model_and_tokenizer().
  • Inference reuses the library’s generate() helper from pkg/llms_from_scratch/ch05.py, ensuring consistent behavior with the book’s training chapters.
  • The extract_response() function isolates the model’s answer by removing the prompt template before displaying output to the user.
  • The entire application runs with a single chainlit run command after installing requirements from the repository’s setup files.

Frequently Asked Questions

What framework does the LLMs-from-scratch UI use?

The interface uses Chainlit, a Python framework designed specifically for building conversational AI demos. Chainlit handles WebSocket connections, message routing, and the browser widget automatically, allowing developers to focus on model inference logic rather than frontend code.

Can I use this UI with my own finetuned checkpoint?

Yes. Replace the checkpoint path in get_model_and_tokenizer() (line 43 of ch07/06_user_interface/app.py) with the path to your own .pth or .pt file. Ensure your model architecture matches the GPT_CONFIG_355M dictionary (1024 embedding dimensions, 24 layers, 16 heads) or update the configuration to match your saved weights.

How does the UI prevent the model from exceeding the context window?

The generate() function in pkg/llms_from_scratch/ch05.py automatically truncates input token IDs to fit within the context_size parameter (1024 tokens for this model). When the prompt plus requested max_new_tokens exceeds this limit, the function trims from the left, preserving the most recent context.

Why is torch.manual_seed(123) set inside the message handler?

Setting a fixed random seed ensures deterministic output during demonstrations, making the UI behavior predictable for tutorials and debugging. In production deployments, you should remove this line to allow the model to sample varied responses, or implement proper temperature-based sampling in the generate() call.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →