How to Deploy Fine-Tuned LLMs with Gradio Spaces: A Step-by-Step Guide

Deploy fine-tuned LLMs with Gradio Spaces by creating an inference script (app.py) that loads your Hugging Face checkpoint, wrapping it in a gr.Interface, and uploading both the script and a requirements.txt to a new Hugging Face Space.

The Lordog/dive-into-llms repository provides a complete workflow for converting fine-tuned checkpoints into shareable web interfaces. According to the source code in documents/chapter1/README.md, this process involves loading a saved model using the transformers library and exposing it through a lightweight Gradio frontend that runs directly on Hugging Face infrastructure.

Fine-Tune Your Model First

Before deployment, you need a trained checkpoint. The tutorial in Chapter 1 (documents/chapter1/README.md) demonstrates fine-tuning a text classification model using the transformers library. Save your checkpoint to a local directory (e.g., ./model/) containing the standard Hugging Face model files (config.json, pytorch_model.bin, tokenizer vocabulary files).

Ensure your saved checkpoint includes:

Building the Inference Script (app.py)

Create a single Python file named app.py that handles model initialization and prediction logic. As implemented in Lordog/dive-into-llms, this script serves as the entry point for the Gradio Space.

Loading the Model and Tokenizer

Use AutoModelForSequenceClassification and AutoTokenizer to load your checkpoint from the ./model/ directory. Set the model to evaluation mode with .eval() to disable dropout and gradient calculation.

import gradio as gr
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch

model_path = "./model"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForSequenceClassification.from_pretrained(model_path)
model.eval()

Creating the Prediction Function

Define a predict function that tokenizes input text, runs inference, and maps the output logits to human-readable labels. The function must accept a string and return a string or dictionary that Gradio can render.

def predict(text: str) -> str:
    inputs = tokenizer(
        text, 
        return_tensors="pt", 
        truncation=True, 
        max_length=512
    )
    with torch.no_grad():
        logits = model(**inputs).logits
    pred_id = logits.argmax(dim=-1).item()
    
    # Adjust this mapping to match your fine-tuned labels

    label_map = {0: "Negative", 1: "Positive"}
    return label_map.get(pred_id, f"Class {pred_id}")

Configuring the Gradio Interface

Instantiate gr.Interface with your prediction function, specifying input and output components. The inputs parameter accepts a gr.Textbox for multi-line text entry, while outputs="label" displays the classification result.

iface = gr.Interface(
    fn=predict,
    inputs=gr.Textbox(lines=2, placeholder="Enter text to classify..."),
    outputs="label",
    title="Fine-Tuned Text Classifier",
    description="Interactive demo of a custom fine-tuned LLM."
)

if __name__ == "__main__":
    iface.launch()

Preparing Dependencies (requirements.txt)

Gradio Spaces require explicit dependency management. Create a requirements.txt file pinning the exact library versions used during fine-tuning. Based on the tutorial in documents/chapter1/requirements.txt, specify:


transformers==4.30.2
torch==2.0.0
gradio

Pinning versions prevents environment mismatches between your local training setup and the Hugging Face Space runtime.

Creating and Configuring the Gradio Space

Navigate to the Hugging Face Space creation page at huggingface.co/new-space?sdk=gradio (referenced in documents/chapter1/README.md lines 156-158). Select Gradio as the SDK, then initialize the Space.

Upload three items to your Space repository:

  1. app.py – Your inference script
  2. requirements.txt – Dependency specifications
  3. ./model/ directory – Your fine-tuned checkpoint files

After uploading, click "Build". The Space automatically installs dependencies and launches the Gradio UI defined in app.py. Once the build completes, the public URL becomes available for testing and sharing.

Advanced: Multimodal Models

For multimodal fine-tuned models beyond text classification, refer to documents/chapter8/README.md and the demo_app.py example in the repository. These files demonstrate handling multiple input types (text and image) and customizing the Gradio interface layout with blocks and components.

Summary

  • Prepare your checkpoint after fine-tuning, saving all tokenizer and model files to a dedicated directory.
  • Write app.py using transformers.AutoModelForSequenceClassification for loading and gr.Interface for the UI wrapper.
  • Pin dependencies in requirements.txt (e.g., transformers==4.30.2, torch==2.0.0, gradio).
  • Deploy via Hugging Face Spaces by uploading your files to a new Gradio Space and triggering the build process.

Frequently Asked Questions

What hardware does Gradio Spaces use for inference?

Gradio Spaces run on Hugging Face infrastructure, typically providing CPU-based containers for free tiers and GPU options (T4, A10G, A100) for upgraded Spaces. The model loads onto the available device automatically; for CPU-only Spaces, ensure your app.py does not explicitly attempt to move tensors to CUDA unless wrapped in conditional logic.

Can I deploy fine-tuned LLMs with custom architectures?

Yes, but you must include the custom model class definition within app.py or install it as a package in requirements.txt. The AutoModel classes work only with standard architectures; for custom heads or modified transformers, instantiate your model class directly using torch.load() or import your training codebase as a module.

How do I handle large model checkpoints that exceed file size limits?

For checkpoints larger than 10GB, use Hugging Face git-lfs to push the model to a separate model repository, then load it via the Hub path in app.py using from_pretrained("your-username/model-name") instead of a local path. Alternatively, enable the Space to download the model at runtime by adding a download script in app.py before the iface.launch() call.

What is the difference between Gradio Spaces and Hugging Face Inference Endpoints?

Gradio Spaces provide interactive web demos with UI components and automatic SSL hosting, ideal for showcasing models to non-technical users. Inference Endpoints offer dedicated API endpoints without a built-in UI, designed for production traffic and programmatic access. Use Spaces for demos and prototypes; use Inference Endpoints for production API serving.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →