How to Unit Test Embabel Agents Using FakeOperationContext and Mock LLMs

Use FakeOperationContext.create() to instantiate an isolated operation context, queue deterministic LLM responses with expectResponse(), and verify prompt construction and interaction tracking via getLlmInvocations() to write fast, deterministic unit tests without external service calls.

The Embabel agent framework in the embabel/embabel-agent repository provides a lightweight operation-context model that captures every LLM interaction. When unit testing agents, you should avoid calling real LLM services to prevent flaky tests and external dependencies. The FakeOperationContext class, located in the embabel-agent-test-support module, provides a complete stub implementation that records invocations and returns pre-configured responses while maintaining API compatibility with the production OperationContext interface.

Creating an Isolated Test Context with FakeOperationContext

Start every test by instantiating a fresh stub context using the static factory method. This guarantees isolation between tests and prevents state leakage.

var context = FakeOperationContext.create();

This method returns an implementation of the OperationContext interface defined in embabel-agent-api/src/main/java/com/embabel/agent/api/OperationContext.java. The stub intercepts all calls to context.ai()—the entry point defined in embabel-agent-api/src/main/java/com/embabel/agent/api/ai/AiFacade.java—allowing you to test the full agent pipeline without network requests.

Queuing Deterministic LLM Responses

Before executing agent logic, define the exact strings the mock LLM should return for each invocation. Use expectResponse(String) to queue responses in FIFO order.

context.expectResponse("confirmed");
context.expectResponse("completed");

When the agent invokes context.ai().withDefaultLlm().createObject(prompt, String.class), the stub returns the queued strings sequentially. This deterministic behavior lets you test specific response handling logic without relying on actual LLM output.

Verifying Agent Behavior and Captured Interactions

After executing the agent, use getLlmInvocations() to retrieve a list of LlmInvocation objects representing each LLM call. Each invocation provides access to the full prompt, interaction metadata, and assigned identifiers.

Key assertion methods include:

  • invocation.getPrompt() – Returns the complete prompt text exactly as the agent sent it.
  • invocation.getInteraction().getId() – Returns the interaction ID set via .withId(String).

The following example from embabel-agent-api/src/test/java/com/embabel/agent/test/unit/FakeOperationContextJavaTest.java demonstrates typical usage patterns:

package com.embabel.agent.test.unit;

import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;

public class FakeOperationContextJavaTest {

    @Test
    public void testInteractionIdAccessFromJava() {
        var context = FakeOperationContext.create();
        context.expectResponse("test result");

        var result = context.ai()
                           .withDefaultLlm()
                           .withId("classify-intent")
                           .createObject("Test prompt", String.class);

        assertEquals("test result", result);
        assertEquals(1, context.getLlmInvocations().size());

        var invocation = context.getLlmInvocations().get(0);
        var interaction = invocation.getInteraction();

        // Verify the ID is accessible from Java
        assertEquals("classify-intent", interaction.getId());
    }

    @Test
    public void testMultipleInvocations() {
        var context = FakeOperationContext.create();
        context.expectResponse("first");
        context.expectResponse("second");

        context.ai().withDefaultLlm().withId("op-1")
               .createObject("First", String.class);
        context.ai().withDefaultLlm().withId("op-2")
               .createObject("Second", String.class);

        assertEquals(2, context.getLlmInvocations().size());
        assertEquals("op-1", context.getLlmInvocations().get(0).getInteraction().getId());
        assertEquals("op-2", context.getLlmInvocations().get(1).getInteraction().getId());
    }

    @Test
    public void testGetPromptReturnsFullContent() {
        var context = FakeOperationContext.create();
        context.expectResponse("result");

        var longPrompt = "This is a very long prompt that should be fully returned without any truncation. "
                       + "It contains multiple sentences and should be preserved in its entirety when calling getPrompt().";

        context.ai().withDefaultLlm().withId("test")
               .createObject(longPrompt, String.class);

        var invocation = context.getLlmInvocations().get(0);
        assertEquals(longPrompt, invocation.getPrompt());
        assertFalse(invocation.getPrompt().contains("UserMessage("));
    }
}

Asserting on Prompt Content

The getPrompt() method returns the raw prompt string without truncation. This allows you to verify that the agent constructs expected prompts correctly, including multi-sentence content or specific formatting. The example above ensures that the prompt does not contain unexpected wrapper strings like UserMessage(.

Tracking Operation Identifiers

When agents label operations using .withId(String), the stub captures these identifiers in the interaction object. Retrieve them via invocation.getInteraction().getId() to verify that your agent correctly labels operations for tracing or branching logic.

Summary

Frequently Asked Questions

Where is the FakeOperationContext class defined?

The class is defined in the embabel-agent-test-support module at embabel-agent-test-support/embabel-agent-test/src/main/java/com/embabel/agent/test/FakeOperationContext.java. It provides a stub implementation of the OperationContext interface used in production code at embabel-agent-api/src/main/java/com/embabel/agent/api/OperationContext.java.

How do I handle tests with multiple sequential LLM calls?

Call expectResponse(String) multiple times before executing the agent. The stub returns queued responses in FIFO order. After execution, assert that context.getLlmInvocations().size() matches the expected call count, and verify each interaction's ID and prompt content individually using the LlmInvocation API.

Can I verify the exact prompt sent to the LLM?

Yes. Each LlmInvocation object returned by getLlmInvocations() provides getPrompt(), which returns the complete prompt string exactly as the agent constructed it. This method preserves full content without truncation, enabling precise assertions on prompt construction logic.

What dependencies are required to use FakeOperationContext?

The FakeOperationContext is available when you depend on the embabel-agent-test artifact, which transitively includes the embabel-agent-test-support module on your test classpath. No additional configuration is required beyond including the standard test dependencies in your build configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →