How to Write Unit Tests for Embabel Agents Using FakeOperationContext

Use FakeOperationContext.create() to instantiate a deterministic stub that queues LLM responses and records every invocation, enabling fast, isolated unit tests without calling external APIs.

When building agents with the embabel/embabel-agent framework, you need a reliable way to verify that your logic constructs the correct prompts and handles LLM outputs appropriately. The FakeOperationContext class, provided by the embabel-agent-test-support module, implements the production OperationContext interface while capturing every interaction for programmatic assertion.

What Is FakeOperationContext?

FakeOperationContext is a test double located in the embabel-agent-test-support module that simulates the runtime behavior of the agent's operation context. Unlike the production implementation, this stub does not invoke real LLM endpoints; instead, it maintains an internal queue of expected responses and a log of all invocations made during the test. This design allows you to write fast, deterministic unit tests that verify your agent's logic in complete isolation.

Setting Up the Test Environment

To use the fake context, ensure your test classpath includes the embabel-agent-test-support artifact. This module is typically available as a test-scoped dependency when working with the embabel-agent project structure. Import the class from the com.embabel.agent.test package:

import com.embabel.agent.test.FakeOperationContext;

The Four-Step Testing Workflow

The typical pattern for testing an Embabel agent follows a predictable cycle: instantiate the context, define expected LLM behavior, execute the agent code, and verify the interactions.

Creating the Stub Context

Begin each test by calling the static factory method to obtain a fresh, isolated context instance:

var context = FakeOperationContext.create();

This method returns a new FakeOperationContext that tracks its own invocations independently of other tests.

Queuing Expected LLM Responses

Before invoking your agent, queue the string responses that the stub should return for subsequent LLM calls:

context.expectResponse("classification: urgent");
context.expectResponse("summary: complete");

Each call to expectResponse adds a response to an internal FIFO queue. When the agent invokes the LLM, the stub dequeues and returns the next queued value.

Executing Agent Logic

Interact with the context using the fluent API that mirrors production code. Access the AI facade through context.ai(), configure the LLM and interaction ID, and invoke the creation method:

var result = context.ai()
    .withDefaultLlm()
    .withId("classify-intent")
    .createObject("Analyze this customer message", String.class);

This call consumes one queued response and records the invocation details internally.

Asserting on Captured Invocations

After execution, retrieve the list of captured invocations to verify behavior:

var invocations = context.getLlmInvocations();
assertEquals(1, invocations.size());

var invocation = invocations.get(0);
assertEquals("classify-intent", invocation.getInteraction().getId());
assertEquals("Analyze this customer message", invocation.getPrompt());

The getLlmInvocations() method returns a list of LlmInvocation objects, each exposing the full prompt text, the interaction metadata (including the ID set via withId), and the response that was returned.

Complete Test Class Example

The following example, adapted from the repository's test suite at embabel-agent-api/src/test/java/com/embabel/agent/test/unit/FakeOperationContextJavaTest.java, demonstrates common assertion patterns including multi-invocation scenarios and prompt content verification:

package com.embabel.agent.test.unit;

import com.embabel.agent.test.FakeOperationContext;
import org.junit.jupiter.api.Test;
import static org.junit.jupiter.api.Assertions.*;

public class AgentUnitTest {

    @Test
    public void testSingleInteractionWithId() {
        var context = FakeOperationContext.create();
        context.expectResponse("test result");

        var result = context.ai()
                           .withDefaultLlm()
                           .withId("classify-intent")
                           .createObject("Test prompt", String.class);

        assertEquals("test result", result);
        assertEquals(1, context.getLlmInvocations().size());

        var invocation = context.getLlmInvocations().get(0);
        assertEquals("classify-intent", invocation.getInteraction().getId());
    }

    @Test
    public void testMultipleOrderedInvocations() {
        var context = FakeOperationContext.create();
        context.expectResponse("first");
        context.expectResponse("second");

        context.ai().withDefaultLlm().withId("op-1")
               .createObject("First prompt", String.class);
        context.ai().withDefaultLlm().withId("op-2")
               .createObject("Second prompt", String.class);

        assertEquals(2, context.getLlmInvocations().size());
        assertEquals("op-1", context.getLlmInvocations().get(0).getInteraction().getId());
        assertEquals("op-2", context.getLlmInvocations().get(1).getInteraction().getId());
    }

    @Test
    public void testPromptContentPreservation() {
        var context = FakeOperationContext.create();
        context.expectResponse("result");

        var longPrompt = "This is a very long prompt that should be fully returned without any truncation. "
                       + "It contains multiple sentences and should be preserved in its entirety.";

        context.ai().withDefaultLlm().withId("test")
               .createObject(longPrompt, String.class);

        var invocation = context.getLlmInvocations().get(0);
        assertEquals(longPrompt, invocation.getPrompt());
    }
}

Key Source Files and API Reference

Understanding the relationship between the fake implementation and the production interfaces helps you write more effective tests. The following files define the core contracts and test utilities:

Summary

  • FakeOperationContext.create() provides an isolated, deterministic stub for unit testing Embabel agents without external LLM dependencies.
  • Use expectResponse() to queue deterministic outputs for each LLM call your agent will make during the test.
  • Access getLlmInvocations() after execution to verify the number of calls, the exact prompts sent, and the interaction IDs assigned via withId().
  • The stub preserves full prompt content without truncation, enabling accurate assertions on complex, multi-line inputs.
  • All functionality resides in the embabel-agent-test-support module, with the core API defined in com.embabel.agent.test.

Frequently Asked Questions

What is the difference between FakeOperationContext and the production OperationContext?

OperationContext is the runtime interface that manages real LLM connections and state persistence, while FakeOperationContext is a test-only implementation that replaces network calls with an in-memory queue and records interactions for verification. The fake version provides identical method signatures but executes entirely within the JVM, making tests fast and deterministic.

How do I test agents that make multiple LLM calls in sequence?

Queue multiple responses using successive calls to context.expectResponse() in the order you expect your agent to consume them. The stub returns responses in FIFO order. After execution, assert that context.getLlmInvocations().size() matches the expected count and verify each invocation's ID and prompt individually.

Can I verify the exact prompt string sent to the LLM?

Yes. Each LlmInvocation object returned by getLlmInvocations() provides a getPrompt() method that returns the complete, unmodified prompt string exactly as passed to createObject(). This includes all whitespace and formatting, allowing you to assert on specific prompt templates or variable substitutions.

Where is FakeOperationContext defined in the source code?

The class is defined in the embabel-agent-test-support module at embabel-agent-test-support/embabel-agent-test/src/main/java/com/embabel/agent/test/FakeOperationContext.java. This module is separate from the core API and is intended solely for test-scoped usage, ensuring that production artifacts do not contain testing utilities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →