How to Run Inference with Eagle Using the Gradio Demo Interface
The NVlabs/Eagle repository provides a ready-to-use Gradio web interface in Eagle/gradio_demo.py that loads a pretrained Eagle model and exposes a chat-style inference endpoint for multimodal conversations.
The Eagle multimodal language model from NVIDIA Labs supports vision-language tasks through a streamlined Gradio demo. This guide explains how to launch the web UI, load pretrained checkpoints, and run inference with images and text using the exact implementation found in the source code.
Prerequisites and Installation
Before launching the interface, clone the repository and install the required dependencies. The demo depends on PyTorch, Transformers, and Gradio libraries specified in the project requirements.
git clone https://github.com/NVlabs/Eagle.git
cd Eagle
pip install -r Eagle/requirements.txt
Ensure you have sufficient GPU memory for your chosen model variant. The demo supports single-GPU and multi-GPU configurations through command-line arguments.
Launching the Gradio Demo
The entry point Eagle/gradio_demo.py implements a complete argument parser (lines 57-69) that controls model selection, hardware configuration, and generation hyperparameters.
Start the server with a specific checkpoint using the --model-path flag:
python Eagle/gradio_demo.py \
--model-path NVEagle/Eagle-X5-13B-Chat \
--num-gpus 1 \
--temperature 0.2 \
--max-new-tokens 512 \
--share
Key command-line options include:
--model-path: Hugging Face model ID or local path to the Eagle checkpoint--model-base: Optional base model for LoRA fine-tuned weights--num-gpus: Number of GPUs for tensor parallelism--load-8bit/--load-4bit: Enable quantization for reduced memory usage--share: Creates a public Gradio tunnel URL accessible from any browser
Upon launch, the script binds to 0.0.0.0:6324 by default and outputs both local and external URLs.
How the Inference Pipeline Works
The demo implements a four-stage inference flow: argument parsing, model loading, UI construction, and response generation.
Model Loading via builder.py
At line 76 of gradio_demo.py, the script calls load_pretrained_model from Eagle/eagle/model/builder.py. This function:
- Configures device mapping and dtype (FP16 by default, with optional 8-bit or 4-bit quantization)
- Loads the tokenizer, Vision Tower, and language model (including LoRA weights if present)
- Registers special multimodal tokens (
<image>,<im_start>,<im_end>) and resizes token embeddings (lines 44-52 inbuilder.py) - Returns the
tokenizer,model,image_processor, and context window size
Image and Text Processing
The Gradio interface constructs a chat UI using gr.Blocks and registers two critical callbacks:
add_text function (lines 94-108): Processes user input by prepending the <image> token when an image is uploaded and formatting the text according to the conversation template.
generate function (lines 125-188): Handles the actual inference by:
- Extracting the formatted prompt from the conversation history
- Processing images through
process_images(defined inEagle/mm_utils.py) - Creating token IDs using
tokenizer_image_tokento handle the<image>placeholder - Streaming the model's output back to the chat widget token-by-token
Generation and Streaming
The generation loop respects the temperature, top-p, and max-new-tokens parameters specified at launch. The model runs in evaluation mode with automatic mixed precision, and responses stream in real-time to the Gradio chat interface. Utility buttons for up-voting, down-voting, and regenerating responses are wired to callbacks that manage the UI state.
Step-by-Step Usage Guide
Once the server starts, you will see output similar to:
Running on local URL: http://0.0.0.0:6324/
External URL: https://<random-id>.gradio.live/
To run inference:
-
Open the External URL in your browser if accessing remotely, or the local URL for same-machine access.
-
Upload an image (optional) by dragging and dropping into the image box, or click to browse. Supported formats include JPG and PNG.
-
Type your prompt in the text box. For example: "Describe what is happening in this image" or "What color is the car?"
-
Click Send or press Enter. The
<image>token is automatically inserted into the conversation context if an image is present. -
Watch the response appear token-by-token in the chat window. Generation stops automatically when the model produces an end-of-sequence token or reaches
--max-new-tokens. -
Adjust parameters in the sidebar to experiment with different sampling strategies. Lower temperatures (e.g., 0.2) produce more deterministic outputs, while higher values increase creativity.
Customizing Generation Parameters
The demo exposes several inference knobs through the UI and command line:
- Temperature: Controls randomness in sampling (default varies by checkpoint, typically 0.2-0.7)
- Top-p (nucleus sampling): Limits token selection to the smallest set whose cumulative probability exceeds the threshold
- Max new tokens: Hard limit on response length to prevent excessive generation
- Num beams: Optional beam search for higher quality (slower) decoding
These parameters are passed to the generation call in the generate function (lines 125-188) and affect the model.generate() call's behavior.
Summary
Eagle/gradio_demo.pyprovides a complete, production-ready Gradio interface for Eagle inference- Model loading is handled by
load_pretrained_modelinEagle/eagle/model/builder.py, which configures quantization and registers special tokens like<image> - Inference combines
process_imagesandtokenizer_image_tokenfromEagle/mm_utils.pyto handle multimodal inputs - Launch requirements include specifying
--model-pathand optional GPU/precision flags - Default configuration serves on port 6324 with optional public tunneling via
--share
Frequently Asked Questions
What hardware requirements are needed to run the Eagle Gradio demo?
The hardware requirements depend on the model size and quantization settings. A 13B parameter model requires approximately 26GB of VRAM in FP16 mode, but you can reduce this to roughly 13GB with 8-bit quantization (--load-8bit) or 7GB with 4-bit quantization (--load-4bit) using bitsandbytes. Multi-GPU setups are supported via the --num-gpus flag for tensor parallelism.
Can I use the Gradio demo with local image files instead of the web interface?
Yes, while the Gradio interface provides a web-based image uploader, the underlying code in gradio_demo.py processes standard image formats. For programmatic batch processing without the UI, you would need to modify the script to load images from disk paths rather than the Gradio file component, using the same process_images utility from Eagle/mm_utils.py.
How do I integrate LoRA weights with the Gradio demo?
To use LoRA fine-tuned weights, specify both --model-path (pointing to the LoRA adapter weights) and --model-base (pointing to the base Eagle model). The load_pretrained_model function in Eagle/eagle/model/builder.py automatically detects and merges LoRA weights when both arguments are provided, handling the token embedding resizing that occurs with additional special tokens.
Why does the demo use special tokens like <im_start> and <im_end>?
These special tokens, defined in Eagle/constants.py, mark the boundaries of visual content within the language model's context window. The tokenizer_image_token function in Eagle/mm_utils.py replaces the <image> placeholder with the appropriate vision feature representations, while <im_start> and <im_end> help the model distinguish between visual and textual modalities during the generation process.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →