Performance Considerations for Embabel‑Agent: JVM Tuning, LLM Optimization, and Planner Selection

Batch LLM calls, prune action graphs, and tune Netty thread pools to keep the Embabel‑Agent framework responsive under load.

The Embabel‑Agent framework runs on the JVM with Spring Boot, which means its performance is shaped by how you optimize LLM traffic, planning algorithms, and reactive I/O. This guide covers the seven critical layers where bottlenecks emerge and how to mitigate them using first‑party APIs and configuration options.

Planning and Re‑planning Overhead

The default Goal‑Oriented Action Planning (GOAP) planner exhaustively explores action sequences. This grows combinatorially with the number of actions and preconditions, making it unsuitable for latency‑sensitive paths.

Switch to Utility AI for speed‑critical decisions:

@Bean
fun planner(): Planner = UtilityAiPlanner()

Alternatively, implement a custom GoalChoiceApprover to prune the action graph before GOAP evaluates it. The planner is pluggable by design—see the GOAP planning section in the README.

LLM Calls and Token Usage

Each LLM invocation carries network latency and token cost. The framework provides two key optimizations in embabel-agent-common/embabel-agent-ai/src/main/kotlin/com/embabel/common/ai/model/EmbeddingService.kt:

Batch embeddings to reduce API round‑trips:

val texts = listOf(
    "The quick brown fox jumps over the lazy dog.",
    "Lorem ipsum dolor sit amet, consectetur adipiscing elit."
)

// Single API call for all texts
val vectors = embeddingService.embedMultiple(texts)

Strip boilerplate before sending files to LLMs using WellKnownFileContentTransformers at embabel-agent-api/src/main/kotlin/com/embabel/agent/tools/file/WellKnownFileContentTransformers.kt:

val rawFile = Files.readString(Paths.get("src/main/kotlin/com/example/largeFile.kt"))
val compact = WellKnownFileContentTransformers.stripCommentsAndWhitespace(rawFile)
ai.withDefaultLlm().createObject(compact, MyDomainClass::class.java)

I/O and Concurrency with Reactor Netty

Embabel‑Agent uses Reactor Netty for all HTTP traffic, including tool calls. This provides non‑blocking I/O and connection pooling out of the box.

Tune the Netty worker thread count to match your CPU cores:

reactor.netty.ioWorkerCount=16

The configuration reference documents reactive stack settings.

Tool Injection and Matching Performance

The ToolInjectionStrategy and ToolNotFoundPolicy implement a two‑tier matching algorithm. You can trade precision for speed using the minTokenLength parameter.

In ToolInjectionStrategy.kt (line 81) and ToolNotFoundPolicy.kt (line 86):

  • Increase minTokenLength when you have many short‑lived tools to skip expensive matching
  • Decrease it when tool names are short and collisions are unlikely

Caching Strategies

The embedding service caches vectors for identical texts by default. For production workloads, replace the default cache with Caffeine or another high‑performance solution via Spring beans:

@Bean
fun embeddingCache(): Cache<String, FloatArray> = Caffeine.newBuilder()
    .maximumSize(10_000)
    .expireAfterWrite(Duration.ofMinutes(30))
    .build()

Caching eliminates redundant embedding latency for repeated queries.

Observability and Profiling

The repository ships with integrations for YourKit and JProfiler (see badges in README.md line 8–9). Enable these in development to identify hotspots.

For production tracing, add the embabel-agent-starter-observability starter and annotate methods with @Tracked:

@Tracked("processOrder")
fun processOrder(order: Order) {
    // business logic
}

Reference: embabel-agent-observability/src/main/kotlin/com/embabel/agent/observability/tracing/Tracked.kt

Resource‑Level Limits and Profiles

Use Spring profiles to select runtime characteristics:

Profile Behavior
docker-desktop Lighter resource footprint for local development
severance Disables debug logging, aggressive GC settings for production

The pom.xml (line 209) declares an Optimized for Kotlin/Java performance profile that applies compiler optimizations.

Summary

  • Batch LLM work with embedMultiple() to cut API latency
  • Trim payloads using WellKnownFileContentTransformers before sending to LLMs
  • Choose planners wisely—GOAP for thoroughness, Utility AI for speed
  • Tune Netty thread pools to match hardware and avoid starvation
  • Leverage built‑in profiling with YourKit/JProfiler and OpenTelemetry tracing

Frequently Asked Questions

How do I reduce LLM token costs in Embabel‑Agent?

Use WellKnownFileContentTransformers.stripCommentsAndWhitespace() to remove boilerplate from file contents before LLM calls, and batch embedding requests via embedMultiple() instead of looping over single texts. Both techniques are implemented in the core AI module.

When should I switch from GOAP to Utility AI planning?

Switch to UtilityAiPlanner when your agent needs sub‑100ms decisions or when the action space exceeds ~50 actions with complex preconditions. GOAP's exhaustive search becomes a bottleneck as combinatorial explosion increases.

What controls tool matching performance?

The minTokenLength parameter in ToolInjectionStrategy and ToolNotFoundPolicy determines how many tokens must match before a tool candidate is evaluated. Raise this value to skip cheap string matching on short tool names when you have hundreds of ephemeral tools.

Does Embabel‑Agent support async I/O for tool calls?

Yes—every HTTP call uses Reactor Netty for non‑blocking I/O. Configure reactor.netty.ioWorkerCount to scale connection handling, and reference the configuration documentation for reactive stack tuning.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →