How to Use Local LLM Models with Embabel: Ollama, LMStudio, and Docker Setup

Embabel integrates with local LLMs through the embabel-agent-starter-ollama starter, which automatically wires Spring AI's OllamaChatModel into Embabel's model-provider infrastructure for on-premise inference.

The embabel-agent repository provides first-class support for local LLM deployment via Ollama, enabling you to run models like Llama 3 entirely on your own hardware. By using the dedicated Ollama starter, you can replace cloud-based providers with local instances while maintaining full compatibility with Embabel's agent framework and tool-calling capabilities. This approach eliminates external API dependencies and ensures complete data privacy for sensitive workloads.

Architecture Overview

Local LLM integration relies on a thin adapter layer that maps Embabel's abstraction to Spring AI's Ollama client.

  • embabel-agent-starter-ollama: The Maven starter that declares org.springframework.ai:spring-ai-ollama in its pom.xml and registers AgentOllamaAutoConfiguration as a Spring bean.
  • OllamaNodeProperties: POJOs defined in the autoconfigure module that model node configuration, including baseUrl and placeholder api-key values.
  • ConfigurableModelProvider: Resolves LLM names (e.g., "llama3") to the appropriate OllamaChatModel bean at runtime.
  • Tool Loop: Handles function calling, retries, and thinking policies unchanged—the loop simply routes requests to the local endpoint instead of cloud APIs.

Sources: The Ollama starter structure is defined in embabel-agent-starters/embabel-agent-starter-ollama/pom.xml, while multi-node configuration is tested in OllamaNodePropertiesTest.kt within the autoconfigure module.

Configure Local LLM Support

Add the Maven Dependency

Include the Ollama starter in your project dependencies to enable local model support.

<!-- pom.xml -->
<dependency>
    <groupId>com.embabel.agent</groupId>
    <artifactId>embabel-agent-starter-ollama</artifactId>
    <version>«current-version»</version>
</dependency>

The starter's POM transitively pulls in Spring AI's Ollama integration and activates AgentOllamaAutoConfiguration via Spring Boot's auto-configuration mechanism.

Configure application.yml

Define your Ollama endpoint and model mapping in src/main/resources/application.yml. The configuration uses OllamaNodeConfig to instantiate the chat model beans.

embabel:
  models:
    default-llm: llama3   # Default model for all LLM calls

    llms:
      local: llama3        # Named role accessible via "#local"

ollama:
  nodes:
    - name: main
      baseUrl: http://localhost:11434   # Ollama API endpoint

      api-key: not-set                  # Required by Spring AI but unused by Ollama

The baseUrl points to your local Ollama server (or a Docker container/LMStudio instance exposing the Ollama API). While the api-key field is mandatory in Spring AI's configuration properties, Ollama ignores this value for local inference.

Run Ollama via Docker

Deploy Ollama locally using the official Docker image to serve models on port 11434.


# Start the Ollama container

docker run -d -p 11434:11434 --name ollama ollama/ollama:latest

# Pull your desired model (e.g., Llama 3)

docker exec -it ollama ollama pull llama3

Once the container is running and the model is downloaded, Embabel can communicate with the endpoint at http://localhost:11434 as configured in the YAML above.

Implementing Local Model Calls

Use the EmbabelAi entry point to initialize a PromptRunner configured for your local model. The withLlm() method resolves the model name through ConfigurableModelProvider, which returns the Ollama-backed implementation.

import com.embabel.agent.ai.EmbabelAi

val ai = EmbabelAi()
val runner = ai.withLlm("llama3")  // Resolves to OllamaChatModel

// Simple text generation
val answer: String = runner.create("Explain quantum computing", String::class.java)

// Structured output with automatic JSON parsing
data class Weather(val city: String, val temperature: Int)

val weather: Weather = runner.createObject(
    "What is the temperature in Paris?",
    Weather::class.java
)

Java developers can use the equivalent fluent API: ai.withLlm("llama3").create("prompt", String.class).

Advanced Configuration

Multi-Node Setups

For production deployments with multiple Ollama instances (e.g., GPU-enabled servers), define additional nodes in the ollama.nodes list. Each node requires a unique name and distinct baseUrl.

ollama:
  nodes:
    - name: gpu-server-1
      baseUrl: http://192.168.1.100:11434
      api-key: not-set
    - name: gpu-server-2
      baseUrl: http://192.168.1.101:11434
      api-key: not-set

The OllamaNodePropertiesTest.kt test suite in embel-agent-autoconfigure/models/embabel-agent-ollama-autoconfigure/src/test/kotlin/com/embabel/agent/config/models/ollama/ demonstrates validation logic for single-node and multi-node configurations.

Summary

  • Add the starter: Include embabel-agent-starter-ollama in your Maven dependencies to activate Ollama auto-configuration.
  • Configure endpoints: Set embabel.models.default-llm and ollama.nodes.baseUrl in application.yml to point to your local server.
  • Handle API keys: Provide a placeholder api-key value to satisfy Spring AI's configuration requirements, even though Ollama does not use it.
  • Use standard APIs: Invoke local models via EmbabelAi.withLlm() and PromptRunner—the tool loop and structured output features work identically to cloud providers.
  • Scale horizontally: Define multiple nodes in the YAML to distribute load across several Ollama instances.

Frequently Asked Questions

Can I use LMStudio instead of Ollama with Embabel?

Yes. LMStudio can serve as the backend if you configure it to expose an OpenAI-compatible local server and adjust the baseUrl accordingly. However, the embabel-agent-starter-ollama specifically targets Ollama's API structure. For LMStudio, you may need to use the generic Spring AI OpenAI starter with a custom baseUrl pointing to LMStudio's local endpoint (typically http://localhost:1234/v1), as noted in the reference documentation for local LLMs.

Why does the configuration require an API key for local models?

Spring AI's OllamaChatModel configuration properties inherit from a base class that mandates an api-key field. According to the source configuration in the installing guide (getting-started/installing/page.adoc), you must provide a dummy value (e.g., not-set) even though Ollama does not perform authentication for local requests. This satisfies the property binding without affecting functionality.

How do I switch between multiple local models?

Define named roles under embabel.models.llms in your YAML. For example, set local-fast: llama3 and local-coding: codellama. In code, reference these via ai.withLlm("local-fast") or ai.withLlm("local-coding"). The ConfigurableModelProvider resolves these aliases to the corresponding Ollama model names at runtime.

Does tool calling work with local Ollama models?

Yes. Embabel's tool loop implementation operates agnostically to the underlying model provider. As implemented in the core agent framework, the loop handles tool invocation, result injection, and retry logic unchanged—the only difference is that HTTP traffic routes to http://localhost:11434/api/chat rather than OpenAI or Anthropic endpoints. Ensure your local model supports function calling (e.g., Llama 3 or specialized tool-use models) for reliable results.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →