How to Set Up a Local Ollama Server for AI Agents: Complete Setup Guide

You can set up a local Ollama server for AI agents by installing Ollama on your platform, launching the daemon with ollama serve, pulling the qwen3:0.6b model, and configuring your agent code to point to the OpenAI-compatible endpoint at http://127.0.0.1:11434.

The bojieli/ai-agent-book repository uses Ollama as its default local LLM backend for agent demonstrations including log sanitization, tool-calling, and active-tool discovery. This guide covers the exact installation steps and configuration parameters found in the source code, ensuring your local AI agents can run entirely offline without external API dependencies.

Installation and Platform Setup

Ollama supports macOS, Windows, and Linux environments. The repository provides specific installation commands for each platform in the local LLM serving documentation.

macOS Installation

For macOS users, the recommended approach uses Homebrew. According to the log-sanitization chapter, you can install via:

brew install ollama

This command appears in chapter3/log-sanitization/README.md and represents the quickest method to get the Ollama binary on macOS systems.

Windows and Linux Installation

Windows users should download the installer directly from the official Ollama website, as referenced in chapter2/local_llm_serving/README.md. Linux users can run the official install script provided by Ollama. Both methods establish the same HTTP API foundation that the AI Agent Book expects.

Starting the Ollama Daemon and Pulling Models

Once installed, you must start the Ollama daemon to expose the local API endpoint on port 11434.

Launch the Server

Run the following command to start the background service:

ollama serve

This launches the local server on the default port 11434. The daemon must remain running to accept POST requests to /v1/chat/completions from your AI agents.

Download the Default Model

The repository defaults to qwen3:0.6b for its demonstrations because it is fast, free, and sufficient for agent workflows. Pull this model using:

ollama pull qwen3:0.6b

As documented in chapter3/log-sanitization/README.md, this model is explicitly referenced in the configuration files for log sanitization and other agent experiments.

Verification and Agent Configuration

Before running agent code, verify the server responds correctly and configure your application to route LLM calls through Ollama.

Verify Server Connectivity

Test that the server is reachable and the model is loaded:

curl -s http://127.0.0.1:11434/api/tags | jq .models[0].name

A successful response returns qwen3:0.6b, confirming the endpoint is active and the model is available.

Configure the Agent Provider

In chapter3/log-sanitization/config.py, the repository defines Ollama-specific settings that route all LLM calls to your local server:

import os

PROVIDER = "ollama"  # Tells the framework to use Ollama backend

OLLAMA_HOST = os.getenv("OLLAMA_HOST", "http://127.0.0.1:11434")
MODEL = "qwen3:0.6b"

The code uses the OPENAI_BASE_URL convention internally, allowing standard OpenAI client libraries to communicate with Ollama's compatible interface. Setting PROVIDER = "ollama" switches the agent framework from cloud APIs to your local instance.

Running AI Agents with the Local Server

With the server running and configuration set, execute agent experiments using the local model. For the log-sanitization demonstration:

python -m chapter3.log_sanitization.main --mode llm

This command invokes the agent pipeline (reasoning → tool calls → output) powered entirely by your local Ollama server rather than external APIs.

Native Client Implementation

The repository includes ollama_native.py, which implements the direct HTTP client used by all experiments. This file contains the chat() and chat_stream() functions that format requests and parse responses from the Ollama API endpoint.

Advanced: Structured Outputs

Ollama supports JSON-structured outputs for deterministic agent processing. To enable this, add the response_format parameter to your API calls:

payload = {
    "model": "qwen3:0.6b",
    "messages": messages,
    "stream": False,
    "response_format": {"type": "json_object"}
}

This feature is particularly useful for downstream processing tasks like PII redaction or tool argument parsing, allowing agents to receive predictable, machine-readable responses.

Summary

  • Install Ollama using platform-specific commands (brew install ollama for macOS, official installers for Windows/Linux) as documented in chapter2/local_llm_serving/README.md.
  • Start the daemon with ollama serve to expose the API on port 11434; verify connectivity via curl to /api/tags.
  • Pull the default model qwen3:0.6b using ollama pull, matching the configuration in chapter3/log-sanitization/config.py.
  • Configure agents by setting PROVIDER = "ollama" and OLLAMA_HOST to point to your local server endpoint.
  • Run experiments using module commands like python -m chapter3.log_sanitization.main --mode llm to execute AI agent workflows offline.

Frequently Asked Questions

What port does the local Ollama server use by default?

The default port is 11434. When you run ollama serve, the daemon binds to http://127.0.0.1:11434, and this is the endpoint you configure in OLLAMA_HOST to route AI agent requests locally.

Can I use a different model than qwen3:0.6b?

Yes. While the AI Agent Book defaults to qwen3:0.6b for its demonstrations, you can replace it with any Ollama-supported model such as llama3.1:8b. Simply run ollama pull <model-name> and update the MODEL variable in your config.py file to match your selection.

How does the agent code switch between cloud and local LLM providers?

The framework checks the PROVIDER variable defined in configuration files like chapter3/log-sanitization/config.py. When set to "ollama", the code routes requests through the Ollama client implementation in ollama_native.py using the OPENAI_BASE_URL pattern; otherwise, it uses standard cloud API endpoints.

Does Ollama support streaming responses for AI agents?

Yes. The ollama_native.py file implements both chat() and chat_stream() functions, allowing agents to receive real-time token streams or complete responses depending on the stream parameter in the request payload.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →