How to Use Local LLM Models with Embabel: Ollama, LMStudio, and Docker Setup
Embabel integrates with local LLMs through the embabel-agent-starter-ollama starter, which automatically wires Spring AI's OllamaChatModel into Embabel's model-provider infrastructure for on-premise inference.
The embabel-agent repository provides first-class support for local LLM deployment via Ollama, enabling you to run models like Llama 3 entirely on your own hardware. By using the dedicated Ollama starter, you can replace cloud-based providers with local instances while maintaining full compatibility with Embabel's agent framework and tool-calling capabilities. This approach eliminates external API dependencies and ensures complete data privacy for sensitive workloads.
Architecture Overview
Local LLM integration relies on a thin adapter layer that maps Embabel's abstraction to Spring AI's Ollama client.
- embabel-agent-starter-ollama: The Maven starter that declares
org.springframework.ai:spring-ai-ollamain itspom.xmland registersAgentOllamaAutoConfigurationas a Spring bean. - OllamaNodeProperties: POJOs defined in the autoconfigure module that model node configuration, including
baseUrland placeholderapi-keyvalues. - ConfigurableModelProvider: Resolves LLM names (e.g., "llama3") to the appropriate
OllamaChatModelbean at runtime. - Tool Loop: Handles function calling, retries, and thinking policies unchanged—the loop simply routes requests to the local endpoint instead of cloud APIs.
Sources: The Ollama starter structure is defined in embabel-agent-starters/embabel-agent-starter-ollama/pom.xml, while multi-node configuration is tested in OllamaNodePropertiesTest.kt within the autoconfigure module.
Configure Local LLM Support
Add the Maven Dependency
Include the Ollama starter in your project dependencies to enable local model support.
<!-- pom.xml -->
<dependency>
<groupId>com.embabel.agent</groupId>
<artifactId>embabel-agent-starter-ollama</artifactId>
<version>«current-version»</version>
</dependency>
The starter's POM transitively pulls in Spring AI's Ollama integration and activates AgentOllamaAutoConfiguration via Spring Boot's auto-configuration mechanism.
Configure application.yml
Define your Ollama endpoint and model mapping in src/main/resources/application.yml. The configuration uses OllamaNodeConfig to instantiate the chat model beans.
embabel:
models:
default-llm: llama3 # Default model for all LLM calls
llms:
local: llama3 # Named role accessible via "#local"
ollama:
nodes:
- name: main
baseUrl: http://localhost:11434 # Ollama API endpoint
api-key: not-set # Required by Spring AI but unused by Ollama
The baseUrl points to your local Ollama server (or a Docker container/LMStudio instance exposing the Ollama API). While the api-key field is mandatory in Spring AI's configuration properties, Ollama ignores this value for local inference.
Run Ollama via Docker
Deploy Ollama locally using the official Docker image to serve models on port 11434.
# Start the Ollama container
docker run -d -p 11434:11434 --name ollama ollama/ollama:latest
# Pull your desired model (e.g., Llama 3)
docker exec -it ollama ollama pull llama3
Once the container is running and the model is downloaded, Embabel can communicate with the endpoint at http://localhost:11434 as configured in the YAML above.
Implementing Local Model Calls
Use the EmbabelAi entry point to initialize a PromptRunner configured for your local model. The withLlm() method resolves the model name through ConfigurableModelProvider, which returns the Ollama-backed implementation.
import com.embabel.agent.ai.EmbabelAi
val ai = EmbabelAi()
val runner = ai.withLlm("llama3") // Resolves to OllamaChatModel
// Simple text generation
val answer: String = runner.create("Explain quantum computing", String::class.java)
// Structured output with automatic JSON parsing
data class Weather(val city: String, val temperature: Int)
val weather: Weather = runner.createObject(
"What is the temperature in Paris?",
Weather::class.java
)
Java developers can use the equivalent fluent API: ai.withLlm("llama3").create("prompt", String.class).
Advanced Configuration
Multi-Node Setups
For production deployments with multiple Ollama instances (e.g., GPU-enabled servers), define additional nodes in the ollama.nodes list. Each node requires a unique name and distinct baseUrl.
ollama:
nodes:
- name: gpu-server-1
baseUrl: http://192.168.1.100:11434
api-key: not-set
- name: gpu-server-2
baseUrl: http://192.168.1.101:11434
api-key: not-set
The OllamaNodePropertiesTest.kt test suite in embel-agent-autoconfigure/models/embabel-agent-ollama-autoconfigure/src/test/kotlin/com/embabel/agent/config/models/ollama/ demonstrates validation logic for single-node and multi-node configurations.
Summary
- Add the starter: Include
embabel-agent-starter-ollamain your Maven dependencies to activate Ollama auto-configuration. - Configure endpoints: Set
embabel.models.default-llmandollama.nodes.baseUrlinapplication.ymlto point to your local server. - Handle API keys: Provide a placeholder
api-keyvalue to satisfy Spring AI's configuration requirements, even though Ollama does not use it. - Use standard APIs: Invoke local models via
EmbabelAi.withLlm()andPromptRunner—the tool loop and structured output features work identically to cloud providers. - Scale horizontally: Define multiple nodes in the YAML to distribute load across several Ollama instances.
Frequently Asked Questions
Can I use LMStudio instead of Ollama with Embabel?
Yes. LMStudio can serve as the backend if you configure it to expose an OpenAI-compatible local server and adjust the baseUrl accordingly. However, the embabel-agent-starter-ollama specifically targets Ollama's API structure. For LMStudio, you may need to use the generic Spring AI OpenAI starter with a custom baseUrl pointing to LMStudio's local endpoint (typically http://localhost:1234/v1), as noted in the reference documentation for local LLMs.
Why does the configuration require an API key for local models?
Spring AI's OllamaChatModel configuration properties inherit from a base class that mandates an api-key field. According to the source configuration in the installing guide (getting-started/installing/page.adoc), you must provide a dummy value (e.g., not-set) even though Ollama does not perform authentication for local requests. This satisfies the property binding without affecting functionality.
How do I switch between multiple local models?
Define named roles under embabel.models.llms in your YAML. For example, set local-fast: llama3 and local-coding: codellama. In code, reference these via ai.withLlm("local-fast") or ai.withLlm("local-coding"). The ConfigurableModelProvider resolves these aliases to the corresponding Ollama model names at runtime.
Does tool calling work with local Ollama models?
Yes. Embabel's tool loop implementation operates agnostically to the underlying model provider. As implemented in the core agent framework, the loop handles tool invocation, result injection, and retry logic unchanged—the only difference is that HTTP traffic routes to http://localhost:11434/api/chat rather than OpenAI or Anthropic endpoints. Ensure your local model supports function calling (e.g., Llama 3 or specialized tool-use models) for reliable results.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →