# How to Deploy Local LLM Services Using vLLM or Ollama: A Complete Guide

> Deploy local LLM services easily using vLLM or Ollama backends. This guide shows you how to set up your environment for efficient AI model deployment on CPUs or GPUs.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: how-to-guide
- Published: 2026-08-23

---

**You can deploy a local LLM service that automatically selects Ollama for CPU-only environments or vLLM for CUDA-enabled GPUs by running the [`main.py`](https://github.com/bojieli/ai-agent-book/blob/main/main.py) script from the ai-agent-book repository with the `--backend` argument.**

The `bojieli/ai-agent-book` repository provides a production-ready implementation for deploying local LLM services using either vLLM or Ollama backends. This guide walks through Experiment 2-I (`local_llm_serving`), which demonstrates how to launch a small language model like Qwen 3-0.6B locally with automatic backend detection and full ReAct tool-calling support.

## Auto-Detecting the Right Backend for Your Hardware

The entry point at [`chapter2/local_llm_serving/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/local_llm_serving/main.py) implements intelligent backend selection based on your hardware capabilities. The script uses `torch.cuda.is_available()` to detect whether a CUDA-capable GPU is present, then automatically chooses the most suitable inference engine.

### When to Use Ollama (CPU, macOS, and Windows)

**Ollama** serves as the default backend when no CUDA GPU is detected. This makes it ideal for macOS (including Apple Silicon), Windows, and Linux CPU-only machines. According to the source code in `ai-agent-book`, Ollama ships the model as a native server process (`ollama serve`), and the Python code communicates with it via HTTP endpoints. This backend achieved **>100 tokens per second** on an Apple M2 chip during benchmarking.

### When to Use vLLM (Linux with NVIDIA GPUs)

**vLLM** is automatically selected on Linux systems with CUDA-enabled GPUs. The script launches the `vllm` engine in the background and streams tokens through its optimized inference engine. This backend is preferred for GPU-accelerated environments due to its efficient memory management and higher throughput compared to CPU inference. The experiment logs in [`chapter2/local_llm_serving/runs/exp2-1-qwen3-0.6b-20260730-v2/manifest.json`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/local_llm_serving/runs/exp2-1-qwen3-0.6b-20260730-v2/manifest.json) confirm successful parallel tool execution when using the vLLM path.

## Running the Local LLM Service

The [`main.py`](https://github.com/bojieli/ai-agent-book/blob/main/main.py) driver accepts three critical arguments to control the deployment:

- `--backend`: Specify `"ollama"`, `"vllm"`, or leave unset for auto-detection
- `--mode`: Choose `"single"` for one-off requests or `"stream"` for continuous token streaming
- `--task`: The prompt text sent to the model (e.g., "What is the weather in Tokyo?")

### Installation and Setup

First, install the required dependencies using the pinned requirements in the repository:

```bash
uv sync

# Or alternatively:

pip install -r requirements.txt

```

For Ollama deployments, start the daemon and pull the model:

```bash
ollama serve
ollama pull qwen3:0.6b

```

### Executing Single-Shot Queries

To run a one-time inference using the Ollama backend:

```bash
uv run python chapter2/local_llm_serving/main.py \
    --backend ollama \
    --mode single \
    --task "What is the weather in Tokyo?"

```

### Enabling Streaming Mode

For real-time token streaming on a Linux machine with an NVIDIA GPU:

```bash
uv run python chapter2/local_llm_serving/main.py \
    --backend vllm \
    --mode stream \
    --task "Summarize the latest news about AI."

```

## Understanding the ReAct Loop and Tool Calling

The `local_llm_serving` implementation follows the same **ReAct loop** used throughout the ai-agent-book, enabling the model to emit tool calls during inference. When you pass a task requiring external data, the model can invoke functions defined in [`chapter2/local_llm_serving/tools.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/local_llm_serving/tools.py), such as `get_current_time` or `get_current_weather`.

The framework automatically executes these tools and returns results to the model before generating the final answer. To trigger tool usage:

```bash
uv run python chapter2/local_llm_serving/main.py \
    --backend ollama \
    --mode stream \
    --task "What time is it in Vancouver and what's the weather?"

```

The script prints the tool's result immediately after execution, followed by the model's synthesized final answer.

## Performance Benchmarks and Validation

The repository includes [`chapter2/local_llm_serving/benchmark.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/local_llm_serving/benchmark.py) to measure token-per-second throughput for both backends. Key findings from the experiment manifest ([`manifest.json`](https://github.com/bojieli/ai-agent-book/blob/main/manifest.json)) show that the 0.6B parameter Qwen 3 model achieves:

- **>100 tokens/second** on Apple M2 (Ollama backend)
- **Parallel tool execution** capabilities when using the vLLM GPU backend

These metrics validate that small local models can serve production-like agent workflows with acceptable latency.

## Summary

- The [`chapter2/local_llm_serving/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/local_llm_serving/main.py) script automatically selects **vLLM** for Linux CUDA environments and **Ollama** for CPU/macOS/Windows systems.
- Three arguments control the deployment: `--backend`, `--mode`, and `--task`.
- The implementation supports both single-shot and streaming inference modes.
- Built-in **ReAct loop** support enables automatic tool calling and execution via [`tools.py`](https://github.com/bojieli/ai-agent-book/blob/main/tools.py).
- Experiment logs confirm >100 tok/s performance on Apple Silicon and efficient parallel execution on GPUs.

## Frequently Asked Questions

### How does the script choose between vLLM and Ollama automatically?

The script calls `torch.cuda.is_available()` to detect CUDA-capable GPUs. If `True` and the system is Linux, it selects vLLM; otherwise, it defaults to Ollama. You can override this by passing `--backend ollama` or `--backend vllm` explicitly.

### Can I use this implementation on macOS with Apple Silicon?

Yes. Since macOS lacks native CUDA support, the auto-detection logic routes to the Ollama backend, which runs efficiently on Apple Silicon. The experiment specifically validated performance on an Apple M2 chip using the Ollama server.

### What models are supported by this local LLM service?

The example uses Qwen 3-0.6B, but any model compatible with your chosen backend works. For Ollama, pull models using `ollama pull <model>`; for vLLM, specify the Hugging Face model identifier. Ensure your hardware has sufficient RAM or VRAM for the chosen model size.

### How does the tool calling mechanism work in this setup?

The ReAct loop in [`main.py`](https://github.com/bojieli/ai-agent-book/blob/main/main.py) parses the model's output for tool call markers (e.g., `get_current_weather`). When detected, the script executes the corresponding function from [`tools.py`](https://github.com/bojieli/ai-agent-book/blob/main/tools.py), captures the return value, and feeds it back to the model before generating the final response. This happens automatically without manual intervention.