# Performance Considerations for Embabel‑Agent: JVM Tuning, LLM Optimization, and Planner Selection

> Optimize Embabel-Agent performance with JVM tuning, LLM optimization, and smart planner selection. Batch LLM calls, prune action graphs, and tune Netty for responsiveness.

- Repository: [Embabel/embabel-agent](https://github.com/embabel/embabel-agent)
- Tags: performance
- Published: 2026-08-14

---

**Batch LLM calls, prune action graphs, and tune Netty thread pools to keep the Embabel‑Agent framework responsive under load.**

The Embabel‑Agent framework runs on the JVM with Spring Boot, which means its performance is shaped by how you optimize LLM traffic, planning algorithms, and reactive I/O. This guide covers the seven critical layers where bottlenecks emerge and how to mitigate them using first‑party APIs and configuration options.

## Planning and Re‑planning Overhead

The default **Goal‑Oriented Action Planning (GOAP)** planner exhaustively explores action sequences. This grows combinatorially with the number of actions and preconditions, making it unsuitable for latency‑sensitive paths.

**Switch to Utility AI for speed‑critical decisions:**

```kotlin
@Bean
fun planner(): Planner = UtilityAiPlanner()

```

Alternatively, implement a custom **GoalChoiceApprover** to prune the action graph before GOAP evaluates it. The planner is pluggable by design—see the [GOAP planning section](https://github.com/embabel/embabel-agent/blob/main/README.md#120-planning-step-is-pluggable) in the README.

## LLM Calls and Token Usage

Each LLM invocation carries network latency and token cost. The framework provides two key optimizations in [`embabel-agent-common/embabel-agent-ai/src/main/kotlin/com/embabel/common/ai/model/EmbeddingService.kt`](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-common/embabel-agent-ai/src/main/kotlin/com/embabel/common/ai/model/EmbeddingService.kt):

**Batch embeddings to reduce API round‑trips:**

```kotlin
val texts = listOf(
    "The quick brown fox jumps over the lazy dog.",
    "Lorem ipsum dolor sit amet, consectetur adipiscing elit."
)

// Single API call for all texts
val vectors = embeddingService.embedMultiple(texts)

```

**Strip boilerplate before sending files to LLMs** using `WellKnownFileContentTransformers` at [`embabel-agent-api/src/main/kotlin/com/embabel/agent/tools/file/WellKnownFileContentTransformers.kt`](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-api/src/main/kotlin/com/embabel/agent/tools/file/WellKnownFileContentTransformers.kt):

```kotlin
val rawFile = Files.readString(Paths.get("src/main/kotlin/com/example/largeFile.kt"))
val compact = WellKnownFileContentTransformers.stripCommentsAndWhitespace(rawFile)
ai.withDefaultLlm().createObject(compact, MyDomainClass::class.java)

```

## I/O and Concurrency with Reactor Netty

Embabel‑Agent uses **Reactor Netty** for all HTTP traffic, including tool calls. This provides non‑blocking I/O and connection pooling out of the box.

Tune the Netty worker thread count to match your CPU cores:

```properties
reactor.netty.ioWorkerCount=16

```

The [configuration reference](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-docs/src/main/asciidoc/reference/configuration/page.adoc#L670) documents reactive stack settings.

## Tool Injection and Matching Performance

The `ToolInjectionStrategy` and `ToolNotFoundPolicy` implement a two‑tier matching algorithm. You can trade precision for speed using the `minTokenLength` parameter.

In [`ToolInjectionStrategy.kt`](https://github.com/embabel/embabel-agent/blob/main/ToolInjectionStrategy.kt) (line 81) and [`ToolNotFoundPolicy.kt`](https://github.com/embabel/embabel-agent/blob/main/ToolNotFoundPolicy.kt) (line 86):

- Increase `minTokenLength` when you have many short‑lived tools to skip expensive matching
- Decrease it when tool names are short and collisions are unlikely

## Caching Strategies

The embedding service caches vectors for identical texts by default. For production workloads, replace the default cache with **Caffeine** or another high‑performance solution via Spring beans:

```kotlin
@Bean
fun embeddingCache(): Cache<String, FloatArray> = Caffeine.newBuilder()
    .maximumSize(10_000)
    .expireAfterWrite(Duration.ofMinutes(30))
    .build()

```

Caching eliminates redundant embedding latency for repeated queries.

## Observability and Profiling

The repository ships with integrations for **YourKit** and **JProfiler** (see badges in [README.md line 8–9](https://github.com/embabel/embabel-agent/blob/main/README.md#L8)). Enable these in development to identify hotspots.

For production tracing, add the `embabel-agent-starter-observability` starter and annotate methods with `@Tracked`:

```kotlin
@Tracked("processOrder")
fun processOrder(order: Order) {
    // business logic
}

```

Reference: [`embabel-agent-observability/src/main/kotlin/com/embabel/agent/observability/tracing/Tracked.kt`](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-observability/src/main/kotlin/com/embabel/agent/observability/tracing/Tracked.kt)

## Resource‑Level Limits and Profiles

Use Spring profiles to select runtime characteristics:

| Profile | Behavior |
|---------|----------|
| `docker-desktop` | Lighter resource footprint for local development |
| `severance` | Disables debug logging, aggressive GC settings for production |

The [`pom.xml`](https://github.com/embabel/embabel-agent/blob/main/pom.xml) (line 209) declares an *Optimized for Kotlin/Java performance* profile that applies compiler optimizations.

## Summary

- **Batch LLM work** with `embedMultiple()` to cut API latency
- **Trim payloads** using `WellKnownFileContentTransformers` before sending to LLMs
- **Choose planners wisely**—GOAP for thoroughness, Utility AI for speed
- **Tune Netty thread pools** to match hardware and avoid starvation
- **Leverage built‑in profiling** with YourKit/JProfiler and OpenTelemetry tracing

## Frequently Asked Questions

### How do I reduce LLM token costs in Embabel‑Agent?

Use `WellKnownFileContentTransformers.stripCommentsAndWhitespace()` to remove boilerplate from file contents before LLM calls, and batch embedding requests via `embedMultiple()` instead of looping over single texts. Both techniques are implemented in the core AI module.

### When should I switch from GOAP to Utility AI planning?

Switch to `UtilityAiPlanner` when your agent needs sub‑100ms decisions or when the action space exceeds ~50 actions with complex preconditions. GOAP's exhaustive search becomes a bottleneck as combinatorial explosion increases.

### What controls tool matching performance?

The `minTokenLength` parameter in `ToolInjectionStrategy` and `ToolNotFoundPolicy` determines how many tokens must match before a tool candidate is evaluated. Raise this value to skip cheap string matching on short tool names when you have hundreds of ephemeral tools.

### Does Embabel‑Agent support async I/O for tool calls?

Yes—every HTTP call uses Reactor Netty for non‑blocking I/O. Configure `reactor.netty.ioWorkerCount` to scale connection handling, and reference the configuration documentation for reactive stack tuning.