# How to Use Local LLM Models with Embabel: Ollama, LMStudio, and Docker Setup

> Integrate local LLM models with Embabel using Ollama, LMStudio, or Docker. Learn to set up on-premise inference with our starter for seamless local LLM use.

- Repository: [Embabel/embabel-agent](https://github.com/embabel/embabel-agent)
- Tags: how-to-guide
- Published: 2026-08-09

---

**Embabel integrates with local LLMs through the `embabel-agent-starter-ollama` starter, which automatically wires Spring AI's `OllamaChatModel` into Embabel's model-provider infrastructure for on-premise inference.**

The embabel-agent repository provides first-class support for local LLM deployment via Ollama, enabling you to run models like Llama 3 entirely on your own hardware. By using the dedicated Ollama starter, you can replace cloud-based providers with local instances while maintaining full compatibility with Embabel's agent framework and tool-calling capabilities. This approach eliminates external API dependencies and ensures complete data privacy for sensitive workloads.

## Architecture Overview

Local LLM integration relies on a thin adapter layer that maps Embabel's abstraction to Spring AI's Ollama client.

- **embabel-agent-starter-ollama**: The Maven starter that declares `org.springframework.ai:spring-ai-ollama` in its [`pom.xml`](https://github.com/embabel/embabel-agent/blob/main/pom.xml) and registers `AgentOllamaAutoConfiguration` as a Spring bean.
- **OllamaNodeProperties**: POJOs defined in the autoconfigure module that model node configuration, including `baseUrl` and placeholder `api-key` values.
- **ConfigurableModelProvider**: Resolves LLM names (e.g., "llama3") to the appropriate `OllamaChatModel` bean at runtime.
- **Tool Loop**: Handles function calling, retries, and thinking policies unchanged—the loop simply routes requests to the local endpoint instead of cloud APIs.

*Sources:* The Ollama starter structure is defined in [`embabel-agent-starters/embabel-agent-starter-ollama/pom.xml`](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-starters/embabel-agent-starter-ollama/pom.xml), while multi-node configuration is tested in [`OllamaNodePropertiesTest.kt`](https://github.com/embabel/embabel-agent/blob/main/OllamaNodePropertiesTest.kt) within the autoconfigure module.

## Configure Local LLM Support

### Add the Maven Dependency

Include the Ollama starter in your project dependencies to enable local model support.

```xml
<!-- pom.xml -->
<dependency>
    <groupId>com.embabel.agent</groupId>
    <artifactId>embabel-agent-starter-ollama</artifactId>
    <version>«current-version»</version>
</dependency>

```

The starter's POM transitively pulls in Spring AI's Ollama integration and activates `AgentOllamaAutoConfiguration` via Spring Boot's auto-configuration mechanism.

### Configure application.yml

Define your Ollama endpoint and model mapping in [`src/main/resources/application.yml`](https://github.com/embabel/embabel-agent/blob/main/src/main/resources/application.yml). The configuration uses `OllamaNodeConfig` to instantiate the chat model beans.

```yaml
embabel:
  models:
    default-llm: llama3   # Default model for all LLM calls

    llms:
      local: llama3        # Named role accessible via "#local"

ollama:
  nodes:
    - name: main
      baseUrl: http://localhost:11434   # Ollama API endpoint

      api-key: not-set                  # Required by Spring AI but unused by Ollama

```

The `baseUrl` points to your local Ollama server (or a Docker container/LMStudio instance exposing the Ollama API). While the `api-key` field is mandatory in Spring AI's configuration properties, Ollama ignores this value for local inference.

### Run Ollama via Docker

Deploy Ollama locally using the official Docker image to serve models on port 11434.

```bash

# Start the Ollama container

docker run -d -p 11434:11434 --name ollama ollama/ollama:latest

# Pull your desired model (e.g., Llama 3)

docker exec -it ollama ollama pull llama3

```

Once the container is running and the model is downloaded, Embabel can communicate with the endpoint at `http://localhost:11434` as configured in the YAML above.

## Implementing Local Model Calls

Use the `EmbabelAi` entry point to initialize a `PromptRunner` configured for your local model. The `withLlm()` method resolves the model name through `ConfigurableModelProvider`, which returns the Ollama-backed implementation.

```kotlin
import com.embabel.agent.ai.EmbabelAi

val ai = EmbabelAi()
val runner = ai.withLlm("llama3")  // Resolves to OllamaChatModel

// Simple text generation
val answer: String = runner.create("Explain quantum computing", String::class.java)

// Structured output with automatic JSON parsing
data class Weather(val city: String, val temperature: Int)

val weather: Weather = runner.createObject(
    "What is the temperature in Paris?",
    Weather::class.java
)

```

Java developers can use the equivalent fluent API: `ai.withLlm("llama3").create("prompt", String.class)`.

## Advanced Configuration

### Multi-Node Setups

For production deployments with multiple Ollama instances (e.g., GPU-enabled servers), define additional nodes in the `ollama.nodes` list. Each node requires a unique name and distinct `baseUrl`.

```yaml
ollama:
  nodes:
    - name: gpu-server-1
      baseUrl: http://192.168.1.100:11434
      api-key: not-set
    - name: gpu-server-2
      baseUrl: http://192.168.1.101:11434
      api-key: not-set

```

The [`OllamaNodePropertiesTest.kt`](https://github.com/embabel/embabel-agent/blob/main/OllamaNodePropertiesTest.kt) test suite in `embel-agent-autoconfigure/models/embabel-agent-ollama-autoconfigure/src/test/kotlin/com/embabel/agent/config/models/ollama/` demonstrates validation logic for single-node and multi-node configurations.

## Summary

- **Add the starter**: Include `embabel-agent-starter-ollama` in your Maven dependencies to activate Ollama auto-configuration.
- **Configure endpoints**: Set `embabel.models.default-llm` and `ollama.nodes.baseUrl` in [`application.yml`](https://github.com/embabel/embabel-agent/blob/main/application.yml) to point to your local server.
- **Handle API keys**: Provide a placeholder `api-key` value to satisfy Spring AI's configuration requirements, even though Ollama does not use it.
- **Use standard APIs**: Invoke local models via `EmbabelAi.withLlm()` and `PromptRunner`—the tool loop and structured output features work identically to cloud providers.
- **Scale horizontally**: Define multiple nodes in the YAML to distribute load across several Ollama instances.

## Frequently Asked Questions

### Can I use LMStudio instead of Ollama with Embabel?

Yes. LMStudio can serve as the backend if you configure it to expose an OpenAI-compatible local server and adjust the `baseUrl` accordingly. However, the `embabel-agent-starter-ollama` specifically targets Ollama's API structure. For LMStudio, you may need to use the generic Spring AI OpenAI starter with a custom `baseUrl` pointing to LMStudio's local endpoint (typically `http://localhost:1234/v1`), as noted in the reference documentation for local LLMs.

### Why does the configuration require an API key for local models?

Spring AI's `OllamaChatModel` configuration properties inherit from a base class that mandates an `api-key` field. According to the source configuration in the installing guide (`getting-started/installing/page.adoc`), you must provide a dummy value (e.g., `not-set`) even though Ollama does not perform authentication for local requests. This satisfies the property binding without affecting functionality.

### How do I switch between multiple local models?

Define named roles under `embabel.models.llms` in your YAML. For example, set `local-fast: llama3` and `local-coding: codellama`. In code, reference these via `ai.withLlm("local-fast")` or `ai.withLlm("local-coding")`. The `ConfigurableModelProvider` resolves these aliases to the corresponding Ollama model names at runtime.

### Does tool calling work with local Ollama models?

Yes. Embabel's tool loop implementation operates agnostically to the underlying model provider. As implemented in the core agent framework, the loop handles tool invocation, result injection, and retry logic unchanged—the only difference is that HTTP traffic routes to `http://localhost:11434/api/chat` rather than OpenAI or Anthropic endpoints. Ensure your local model supports function calling (e.g., Llama 3 or specialized tool-use models) for reliable results.