How to Deploy an LLM Demo on Gradio Spaces: Required Files and Setup Guide

To deploy an LLM demo on Gradio Spaces, you need three essential files: an app.py inference script that builds the Gradio interface, a requirements.txt file specifying dependencies like transformers and torch, and the model checkpoint files (or a Hugging Face Hub reference).

The Lordog/dive-into-llms repository provides a complete tutorial in Chapter 1 demonstrating exactly how to package and deploy fine‑tuned language models to Hugging Face’s Gradio Spaces platform. This guide walks through the minimal file structure required based on the source code analysis of documents/chapter1/README.md and the alternative demo patterns found in Chapter 8.

Essential Files for Gradio Spaces Deployment

Gradio Spaces automatically builds a containerized environment when you push files to a new Space. According to the Chapter 1 tutorial in documents/chapter1/README.md (lines 54‑58), your repository must contain the following three components.

app.py – The Entry Point

The app.py file serves as the entry point that Gradio Spaces executes on startup. This script loads your fine‑tuned model using transformers or torch, defines a prediction function, and wraps it in a gr.Interface object to generate the web UI.

As illustrated in the tutorial screenshot (lines 64‑68 of documents/chapter1/README.md), app.py typically imports gradio as gr, instantiates the model and tokenizer, and launches the interface with iface.launch().

requirements.txt – Dependency Management

The requirements.txt file tells Gradio Spaces which Python packages to install before running app.py. The tutorial explicitly pins specific versions to ensure reproducibility (lines 69‑73 of documents/chapter1/README.md):

transformers==4.30.2
torch==2.0.0
gradio==4.1.0

Missing or incorrect versions in this file cause the Space to fail at launch, so always include the exact versions tested with your model.

Model Checkpoints – The LLM Assets

You must provide the model weights and configuration files. These include pytorch_model.bin (or model.safetensors), config.json, and tokenizer files (vocabulary and merges). The tutorial assumes these are saved locally after training and uploaded alongside app.py (referenced in lines 54‑58). Alternatively, you can modify app.py to download the model from the Hugging Face Hub at runtime using from_pretrained().

Optional Supporting Files

While not strictly required for deployment, you can include additional files to improve usability:

  • utils.py – Helper functions for preprocessing or postprocessing
  • README.md – Documentation explaining how to use the demo
  • Example data – Sample inputs demonstrating model capabilities

These assets help users understand your demo but do not affect the Space’s ability to start.

Complete Implementation Example

Below is a minimal, production‑ready app.py adapted from the patterns in Lordog/dive-into-llms. It loads a causal language model and exposes temperature and token‑limit controls through the Gradio UI:

import gradio as gr
from transformers import AutoModelForCausalLM, AutoTokenizer

# -------------------------------------------------

# Load model & tokenizer (replace with your checkpoint)

# -------------------------------------------------

MODEL_NAME = "gpt2"  # Update to your fine-tuned model path

tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
model = AutoModelForCausalLM.from_pretrained(MODEL_NAME)

def generate(prompt, max_new_tokens=50, temperature=0.7):
    inputs = tokenizer(prompt, return_tensors="pt")
    output_ids = model.generate(
        **inputs,
        max_new_tokens=max_new_tokens,
        temperature=temperature,
        do_sample=True,
    )
    return tokenizer.decode(output_ids[0], skip_special_tokens=True)

# -------------------------------------------------

# Build Gradio UI

# -------------------------------------------------

iface = gr.Interface(
    fn=generate,
    inputs=[
        gr.Textbox(lines=5, label="Prompt"),
        gr.Slider(10, 200, step=10, label="Max new tokens"),
        gr.Slider(0.1, 1.0, step=0.1, label="Temperature")
    ],
    outputs=gr.Textbox(label="Generated text"),
    title="LLM Demo on Gradio Spaces",
    description="Enter a prompt and let the model continue the sentence."
)

if __name__ == "__main__":
    iface.launch()

Pair this with the requirements.txt shown above, place your model checkpoint folder in the same directory (or update MODEL_NAME to a Hugging Face Hub identifier), and push all files to a new Gradio Space.

Repository Structure and Examples

The Lordog/dive-into-llms repository provides concrete references for these patterns:

  • Chapter 1 tutorial: Located in documents/chapter1/README.md, lines 54‑58 explain the model export process and lines 64‑68 show the app.py screenshot for the Gradio interface.
  • Dependency specification: Lines 69‑73 in the same file list the exact transformers and torch versions required.
  • Alternative demo: Chapter 8 (documents/chapter8/README.md, lines 140‑144) references a demo_app.py skeleton for multi‑modal LLM scenarios, demonstrating that the same three‑file pattern applies across different model architectures.

Summary

Deploying an LLM demo on Gradio Spaces requires a minimal but precise file structure:

  • app.py serves as the entry point, containing the gr.Interface definition and model loading logic.
  • requirements.txt pins dependencies like transformers==4.30.2 and torch==2.0.0 to ensure the container builds correctly.
  • Model checkpoint files (pytorch_model.bin, config.json, tokenizer files) provide the inference weights, either uploaded directly or downloaded from Hugging Face Hub.
  • Place these files in a Git repository and push to a new Gradio Space to automatically build and launch your demo.

Frequently Asked Questions

Can I use a model from Hugging Face Hub instead of local files?

Yes. In app.py, replace the local path in AutoModelForCausalLM.from_pretrained() with a Hugging Face Hub model ID (e.g., "meta-llama/Llama-2-7b-chat-hf"). The model will download automatically when the Space starts, though cold‑start times will be longer for large checkpoints.

What happens if I forget to include requirements.txt?

Gradio Spaces will fail to launch with an ModuleNotFoundError because the container environment lacks gradio, transformers, and torch. Always include requirements.txt at the repository root with exact versions tested locally.

How do I test my app.py locally before deploying?

Run python app.py in your local environment after installing dependencies from requirements.txt. The script will start a local web server (usually http://127.0.0.1:7860) where you can verify the model loads and the Gradio interface renders correctly before pushing to Spaces.

Are there size limits for model files on Gradio Spaces?

Yes, free-tier Gradio Spaces have repository size limits (typically 50 GB for the entire repository). If your model exceeds this, use the Hugging Face Hub integration to load weights at runtime rather than uploading them directly to the Space repository.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →