How to Deploy Local LLM Services Using vLLM or Ollama: A Complete Guide

You can deploy a local LLM service that automatically selects Ollama for CPU-only environments or vLLM for CUDA-enabled GPUs by running the main.py script from the ai-agent-book repository with the --backend argument.

The bojieli/ai-agent-book repository provides a production-ready implementation for deploying local LLM services using either vLLM or Ollama backends. This guide walks through Experiment 2-I (local_llm_serving), which demonstrates how to launch a small language model like Qwen 3-0.6B locally with automatic backend detection and full ReAct tool-calling support.

Auto-Detecting the Right Backend for Your Hardware

The entry point at chapter2/local_llm_serving/main.py implements intelligent backend selection based on your hardware capabilities. The script uses torch.cuda.is_available() to detect whether a CUDA-capable GPU is present, then automatically chooses the most suitable inference engine.

When to Use Ollama (CPU, macOS, and Windows)

Ollama serves as the default backend when no CUDA GPU is detected. This makes it ideal for macOS (including Apple Silicon), Windows, and Linux CPU-only machines. According to the source code in ai-agent-book, Ollama ships the model as a native server process (ollama serve), and the Python code communicates with it via HTTP endpoints. This backend achieved >100 tokens per second on an Apple M2 chip during benchmarking.

When to Use vLLM (Linux with NVIDIA GPUs)

vLLM is automatically selected on Linux systems with CUDA-enabled GPUs. The script launches the vllm engine in the background and streams tokens through its optimized inference engine. This backend is preferred for GPU-accelerated environments due to its efficient memory management and higher throughput compared to CPU inference. The experiment logs in chapter2/local_llm_serving/runs/exp2-1-qwen3-0.6b-20260730-v2/manifest.json confirm successful parallel tool execution when using the vLLM path.

Running the Local LLM Service

The main.py driver accepts three critical arguments to control the deployment:

  • --backend: Specify "ollama", "vllm", or leave unset for auto-detection
  • --mode: Choose "single" for one-off requests or "stream" for continuous token streaming
  • --task: The prompt text sent to the model (e.g., "What is the weather in Tokyo?")

Installation and Setup

First, install the required dependencies using the pinned requirements in the repository:

uv sync

# Or alternatively:

pip install -r requirements.txt

For Ollama deployments, start the daemon and pull the model:

ollama serve
ollama pull qwen3:0.6b

Executing Single-Shot Queries

To run a one-time inference using the Ollama backend:

uv run python chapter2/local_llm_serving/main.py \
    --backend ollama \
    --mode single \
    --task "What is the weather in Tokyo?"

Enabling Streaming Mode

For real-time token streaming on a Linux machine with an NVIDIA GPU:

uv run python chapter2/local_llm_serving/main.py \
    --backend vllm \
    --mode stream \
    --task "Summarize the latest news about AI."

Understanding the ReAct Loop and Tool Calling

The local_llm_serving implementation follows the same ReAct loop used throughout the ai-agent-book, enabling the model to emit tool calls during inference. When you pass a task requiring external data, the model can invoke functions defined in chapter2/local_llm_serving/tools.py, such as get_current_time or get_current_weather.

The framework automatically executes these tools and returns results to the model before generating the final answer. To trigger tool usage:

uv run python chapter2/local_llm_serving/main.py \
    --backend ollama \
    --mode stream \
    --task "What time is it in Vancouver and what's the weather?"

The script prints the tool's result immediately after execution, followed by the model's synthesized final answer.

Performance Benchmarks and Validation

The repository includes chapter2/local_llm_serving/benchmark.py to measure token-per-second throughput for both backends. Key findings from the experiment manifest (manifest.json) show that the 0.6B parameter Qwen 3 model achieves:

  • >100 tokens/second on Apple M2 (Ollama backend)
  • Parallel tool execution capabilities when using the vLLM GPU backend

These metrics validate that small local models can serve production-like agent workflows with acceptable latency.

Summary

  • The chapter2/local_llm_serving/main.py script automatically selects vLLM for Linux CUDA environments and Ollama for CPU/macOS/Windows systems.
  • Three arguments control the deployment: --backend, --mode, and --task.
  • The implementation supports both single-shot and streaming inference modes.
  • Built-in ReAct loop support enables automatic tool calling and execution via tools.py.
  • Experiment logs confirm >100 tok/s performance on Apple Silicon and efficient parallel execution on GPUs.

Frequently Asked Questions

How does the script choose between vLLM and Ollama automatically?

The script calls torch.cuda.is_available() to detect CUDA-capable GPUs. If True and the system is Linux, it selects vLLM; otherwise, it defaults to Ollama. You can override this by passing --backend ollama or --backend vllm explicitly.

Can I use this implementation on macOS with Apple Silicon?

Yes. Since macOS lacks native CUDA support, the auto-detection logic routes to the Ollama backend, which runs efficiently on Apple Silicon. The experiment specifically validated performance on an Apple M2 chip using the Ollama server.

What models are supported by this local LLM service?

The example uses Qwen 3-0.6B, but any model compatible with your chosen backend works. For Ollama, pull models using ollama pull <model>; for vLLM, specify the Hugging Face model identifier. Ensure your hardware has sufficient RAM or VRAM for the chosen model size.

How does the tool calling mechanism work in this setup?

The ReAct loop in main.py parses the model's output for tool call markers (e.g., get_current_weather). When detected, the script executes the corresponding function from tools.py, captures the return value, and feeds it back to the model before generating the final response. This happens automatically without manual intervention.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →