How to Set Up Ollama with OpenClaude Using OpenAI Compatibility: Complete Local LLM Guide
OpenClaude automatically detects Ollama's OpenAI-compatible endpoint at http://localhost:11434 and routes chat requests locally without requiring API keys or cloud credentials.
Setting up Ollama with OpenClaude leverages the built-in OpenAI compatibility shim to run large language models entirely on your local machine. This integration, implemented in the Gitlawb/openclaude repository, eliminates external API dependencies while maintaining full protocol compatibility with standard OpenAI chat completion workflows.
Automatic Discovery and Routing
OpenClaude implements zero-configuration provider detection to identify running Ollama instances through dedicated discovery utilities.
Endpoint Detection Logic
In src/utils/providerDiscovery.ts at line 8, the system declares DEFAULT_OLLAMA_BASE_URL = 'http://localhost:11434'. During initialization, the CLI probes this endpoint to determine if Ollama is available. If you run Ollama on a non-standard port or remote server, set the OLLAMA_BASE_URL environment variable to override this default.
Local Provider Classification
According to architecture/integrations.md at line 81, Ollama is categorized with provider kind local. This classification groups Ollama with other local inference engines like LM Studio, instructing the user interface to bypass API key validation and authentication flows required for cloud providers such as OpenAI or Azure.
OpenAI Shim Architecture
The integration translates between OpenClaude's request format and Ollama's native OpenAI-compatible layer through a specialized adapter service.
Context Length Management
The src/services/api/openaiShim/ollamaAdapter.ts file handles protocol translation at lines 59-61, extracting the desired context length and forwarding it via the standard /v1/chat/completions endpoint. By default, OpenClaude requests 32768 tokens of context, which the adapter passes to Ollama's shim layer for processing.
Request Flow Standardization
Because Ollama implements the OpenAI chat completions schema, OpenClaude sends identical JSON payloads to both cloud and local providers. The adapter ensures streaming content tokens arrive in the same format as OpenAI responses, maintaining compatibility with existing consumption logic without provider-specific branching.
Environment Configuration
Control your local LLM behavior using these environment variables:
OLLAMA_BASE_URL– Specifies custom Ollama server location (default:http://localhost:11434)OPENCLAUDE_OLLAMA_NUM_CTX– Overrides the default context window size (default:32768)
export OLLAMA_BASE_URL=http://remote-server:11434
export OPENCLAUDE_OLLAMA_NUM_CTX=65536
Step-by-Step Setup Guide
Configure your local inference pipeline in under five minutes using these validated commands.
1. Install and Initialize Ollama
Install Ollama from the official distribution, then pull your desired model:
# Install Ollama (macOS/Linux/Windows)
# Reference: docs/quick-start-mac-linux.md#L55-L57
# Pull and run Llama 3.1 8B
ollama run llama3.1:8b
2. Verify Server Configuration
Confirm Ollama is listening and configured for adequate context length:
# Check running models and allocated context
ollama ps
# For larger context windows, restart with custom size
ollama serve --num-ctx 65536
3. Launch OpenClaude
Use the specialized CLI aliases for local Ollama operation:
# Standard local mode (balanced latency)
oc-local
# Optimized low-latency mode
oc-fast
Reference these launch profiles in docs/windows-aliases-and-launchers.md for additional customization options.
4. Connect to Remote Instances
To route requests to a remote Ollama server instead of localhost:
export OLLAMA_BASE_URL=http://192.168.1.100:11434
oc-local
Summary
- Zero Configuration: OpenClaude auto-detects Ollama at
http://localhost:11434viasrc/utils/providerDiscovery.tswithout manual endpoint setup. - No API Keys Required: Local provider classification in
architecture/integrations.mdskips authentication validation entirely. - Native Compatibility: The
ollamaAdapter.tstranslates OpenAI-compatible requests to Ollama's shim at lines 59-61. - Adjustable Context: Default 32768 token context adjustable via
OPENCLAUDE_OLLAMA_NUM_CTXenvironment variable. - Flexible Deployment: Supports both local workstations and remote Ollama instances through
OLLAMA_BASE_URL.
Frequently Asked Questions
How does OpenClaude detect my Ollama installation?
OpenClaude probes the default endpoint defined in src/utils/providerDiscovery.ts at line 8, where DEFAULT_OLLAMA_BASE_URL points to http://localhost:11434. If Ollama responds at this address, the CLI automatically configures itself to use local inference mode without requiring manual provider selection.
Can I use Ollama models with different context lengths?
Yes. Set the OPENCLAUDE_OLLAMA_NUM_CTX environment variable to override the default 32768 token limit, or configure Ollama directly using ollama serve --num-ctx 65536 before launching your model. The adapter in src/services/api/openaiShim/ollamaAdapter.ts passes this parameter through the OpenAI-compatible endpoint.
Is an internet connection or API key required for local Ollama usage?
No. Because Ollama is classified with provider kind local in the architecture configuration, OpenClaude bypasses all API key validation and authentication checks. This enables fully offline operation using only your local hardware resources.
What CLI commands are available for launching OpenClaude with Ollama?
OpenClaude provides oc-local for standard local operation and oc-fast for low-latency optimization. These aliases, documented in docs/windows-aliases-and-launchers.md, automatically detect your Ollama instance and apply the appropriate connection profile without additional flags.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →