# How to Run Inference with Eagle Using the Gradio Demo Interface

> Easily run inference with the Eagle multimodal model using its Gradio demo interface. Interact with a pretrained model for chat-style conversations. Get started now.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: how-to-guide
- Published: 2026-06-28

---

**The NVlabs/Eagle repository provides a ready-to-use Gradio web interface in [`Eagle/gradio_demo.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/gradio_demo.py) that loads a pretrained Eagle model and exposes a chat-style inference endpoint for multimodal conversations.**

The Eagle multimodal language model from NVIDIA Labs supports vision-language tasks through a streamlined Gradio demo. This guide explains how to launch the web UI, load pretrained checkpoints, and run inference with images and text using the exact implementation found in the source code.

## Prerequisites and Installation

Before launching the interface, clone the repository and install the required dependencies. The demo depends on PyTorch, Transformers, and Gradio libraries specified in the project requirements.

```bash
git clone https://github.com/NVlabs/Eagle.git
cd Eagle
pip install -r Eagle/requirements.txt

```

Ensure you have sufficient GPU memory for your chosen model variant. The demo supports single-GPU and multi-GPU configurations through command-line arguments.

## Launching the Gradio Demo

The entry point [`Eagle/gradio_demo.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/gradio_demo.py) implements a complete argument parser (lines 57-69) that controls model selection, hardware configuration, and generation hyperparameters.

Start the server with a specific checkpoint using the `--model-path` flag:

```bash
python Eagle/gradio_demo.py \
    --model-path NVEagle/Eagle-X5-13B-Chat \
    --num-gpus 1 \
    --temperature 0.2 \
    --max-new-tokens 512 \
    --share

```

Key command-line options include:
- **`--model-path`**: Hugging Face model ID or local path to the Eagle checkpoint
- **`--model-base`**: Optional base model for LoRA fine-tuned weights
- **`--num-gpus`**: Number of GPUs for tensor parallelism
- **`--load-8bit` / `--load-4bit`**: Enable quantization for reduced memory usage
- **`--share`**: Creates a public Gradio tunnel URL accessible from any browser

Upon launch, the script binds to `0.0.0.0:6324` by default and outputs both local and external URLs.

## How the Inference Pipeline Works

The demo implements a four-stage inference flow: argument parsing, model loading, UI construction, and response generation.

### Model Loading via builder.py

At line 76 of [`gradio_demo.py`](https://github.com/NVlabs/Eagle/blob/main/gradio_demo.py), the script calls `load_pretrained_model` from [`Eagle/eagle/model/builder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/builder.py). This function:
- Configures device mapping and dtype (FP16 by default, with optional 8-bit or 4-bit quantization)
- Loads the tokenizer, Vision Tower, and language model (including LoRA weights if present)
- Registers special multimodal tokens (`<image>`, `<im_start>`, `<im_end>`) and resizes token embeddings (lines 44-52 in [`builder.py`](https://github.com/NVlabs/Eagle/blob/main/builder.py))
- Returns the `tokenizer`, `model`, `image_processor`, and context window size

### Image and Text Processing

The Gradio interface constructs a chat UI using `gr.Blocks` and registers two critical callbacks:

**`add_text` function (lines 94-108)**: Processes user input by prepending the `<image>` token when an image is uploaded and formatting the text according to the conversation template.

**`generate` function (lines 125-188)**: Handles the actual inference by:
1. Extracting the formatted prompt from the conversation history
2. Processing images through `process_images` (defined in [`Eagle/mm_utils.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/mm_utils.py))
3. Creating token IDs using `tokenizer_image_token` to handle the `<image>` placeholder
4. Streaming the model's output back to the chat widget token-by-token

### Generation and Streaming

The generation loop respects the temperature, top-p, and max-new-tokens parameters specified at launch. The model runs in evaluation mode with automatic mixed precision, and responses stream in real-time to the Gradio chat interface. Utility buttons for up-voting, down-voting, and regenerating responses are wired to callbacks that manage the UI state.

## Step-by-Step Usage Guide

Once the server starts, you will see output similar to:

```

Running on local URL:  http://0.0.0.0:6324/
External URL:         https://<random-id>.gradio.live/

```

To run inference:

1. **Open the External URL** in your browser if accessing remotely, or the local URL for same-machine access.

2. **Upload an image** (optional) by dragging and dropping into the image box, or click to browse. Supported formats include JPG and PNG.

3. **Type your prompt** in the text box. For example: "Describe what is happening in this image" or "What color is the car?"

4. **Click Send** or press Enter. The `<image>` token is automatically inserted into the conversation context if an image is present.

5. **Watch the response** appear token-by-token in the chat window. Generation stops automatically when the model produces an end-of-sequence token or reaches `--max-new-tokens`.

6. **Adjust parameters** in the sidebar to experiment with different sampling strategies. Lower temperatures (e.g., 0.2) produce more deterministic outputs, while higher values increase creativity.

## Customizing Generation Parameters

The demo exposes several inference knobs through the UI and command line:

- **Temperature**: Controls randomness in sampling (default varies by checkpoint, typically 0.2-0.7)
- **Top-p (nucleus sampling)**: Limits token selection to the smallest set whose cumulative probability exceeds the threshold
- **Max new tokens**: Hard limit on response length to prevent excessive generation
- **Num beams**: Optional beam search for higher quality (slower) decoding

These parameters are passed to the generation call in the `generate` function (lines 125-188) and affect the `model.generate()` call's behavior.

## Summary

- **[`Eagle/gradio_demo.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/gradio_demo.py)** provides a complete, production-ready Gradio interface for Eagle inference
- **Model loading** is handled by `load_pretrained_model` in [`Eagle/eagle/model/builder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/builder.py), which configures quantization and registers special tokens like `<image>`
- **Inference** combines `process_images` and `tokenizer_image_token` from [`Eagle/mm_utils.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/mm_utils.py) to handle multimodal inputs
- **Launch requirements** include specifying `--model-path` and optional GPU/precision flags
- **Default configuration** serves on port 6324 with optional public tunneling via `--share`

## Frequently Asked Questions

### What hardware requirements are needed to run the Eagle Gradio demo?

The hardware requirements depend on the model size and quantization settings. A 13B parameter model requires approximately 26GB of VRAM in FP16 mode, but you can reduce this to roughly 13GB with 8-bit quantization (`--load-8bit`) or 7GB with 4-bit quantization (`--load-4bit`) using bitsandbytes. Multi-GPU setups are supported via the `--num-gpus` flag for tensor parallelism.

### Can I use the Gradio demo with local image files instead of the web interface?

Yes, while the Gradio interface provides a web-based image uploader, the underlying code in [`gradio_demo.py`](https://github.com/NVlabs/Eagle/blob/main/gradio_demo.py) processes standard image formats. For programmatic batch processing without the UI, you would need to modify the script to load images from disk paths rather than the Gradio file component, using the same `process_images` utility from [`Eagle/mm_utils.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/mm_utils.py).

### How do I integrate LoRA weights with the Gradio demo?

To use LoRA fine-tuned weights, specify both `--model-path` (pointing to the LoRA adapter weights) and `--model-base` (pointing to the base Eagle model). The `load_pretrained_model` function in [`Eagle/eagle/model/builder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/builder.py) automatically detects and merges LoRA weights when both arguments are provided, handling the token embedding resizing that occurs with additional special tokens.

### Why does the demo use special tokens like `<im_start>` and `<im_end>`?

These special tokens, defined in [`Eagle/constants.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/constants.py), mark the boundaries of visual content within the language model's context window. The `tokenizer_image_token` function in [`Eagle/mm_utils.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/mm_utils.py) replaces the `<image>` placeholder with the appropriate vision feature representations, while `<im_start>` and `<im_end>` help the model distinguish between visual and textual modalities during the generation process.