How to Implement Streaming Responses with StreamingJacksonOutputConverter in Embabel
Embabel's StreamingJacksonOutputConverter uses Jackson's JsonGenerator to convert reactive LLM chunks into a low-latency JSON stream, enabling low-memory, incremental delivery to clients.
The embabel/embabel-agent repository provides a first-class streaming mechanism for AI agents through the StreamingJacksonOutputConverter class. This component sits at the boundary between the LLM client and the HTTP layer, turning incremental model output into valid, streaming JSON without buffering the entire response. Understanding how to wire this converter into your service and controller layers is essential for building responsive, production-ready agent applications.
Architecture and Data Flow
According to the embabel/embabel-agent source code, the streaming pipeline is composed of four core components:
StreamingJacksonOutputConverter— Located atembabel-agent-common/embabel-agent-ai/src/main/java/com/embabel/common/ai/converters/StreamingJacksonOutputConverter.java, this class consumes a reactive stream of string chunks and writes JSON tokens on-the-fly using Jackson'sJsonGenerator.AgentStreamingResponse— Found inembabel-agent-common/embabel-agent-ai/src/main/java/com/embabel/common/ai/AgentStreamingResponse.java, this wrapper exposes the stream as aFlux<String>(orPublisher<String>) that the HTTP layer can subscribe to directly.AgentController— Inembabel-agent-autoconfigure/embabel-agent-openai-autoconfigure/src/main/java/com/embabel/agent/autoconfigure/openai/AgentController.java, this Spring WebFlux controller maps the streaming endpoint and wires the converter into the request-handling chain.JacksonAutoConfiguration— The auto-configuration class atembabel-agent-autoconfigure/embabel-agent-openai-autoconfigure/src/main/java/org/springframework/boot/jackson/autoconfigure/JacksonAutoConfiguration.javaregisters the converter as a Spring bean when Jackson is present on the classpath.
The data flow follows these steps:
- The controller receives a request and builds a
StructuredOutputobject. - The LLM client (e.g.,
OpenAiChatModel) is invoked withstream = true, returning aFlux<String>of partial messages. - The
StreamingJacksonOutputConvertersubscribes to that flux and feeds each chunk into a JacksonJsonGenerator, emitting JSON fragments such as{ "content": "...", "delta": "…" }as soon as they arrive. - The controller returns the resulting
Flux<String>to the web layer, which serializes it as a chunkedapplication/jsonor Server-Sent Events response.
Because conversion happens per chunk, memory consumption stays minimal and the client can process data immediately.
Implementing the Streaming Pipeline
Auto-Configure the Converter Bean
In most Spring Boot setups, no explicit bean is required. The converter is imported automatically:
// No explicit bean needed if Spring Boot is used; the class is imported via
// @ImportAutoConfiguration({JacksonAutoConfiguration.class, AgentOpenAiAutoConfiguration.class})
Create a Streaming Service
Inject both the chat model and the StreamingJacksonOutputConverter, then pipe the raw LLM flux through the converter:
@Service
public class ChatService {
private final OpenAiChatModel chatModel;
private final StreamingJacksonOutputConverter converter;
public ChatService(OpenAiChatModel chatModel,
StreamingJacksonOutputConverter converter) {
this.chatModel = chatModel;
this.converter = converter;
}
/** Returns a reactive stream of JSON fragments */
public Flux<String> chatStream(String userPrompt) {
Flux<String> rawChunks = chatModel.chat(userPrompt, ChatOptions.builder()
.stream(true) // enable streaming
.build());
return converter.convert(rawChunks);
}
}
Build the Controller Endpoint
Return the Flux<String> directly from a Spring WebFlux controller. Producing APPLICATION_NDJSON_VALUE yields newline-delimited JSON:
@RestController
@RequestMapping("/api")
public class AgentController {
private final ChatService chatService;
public AgentController(ChatService chatService) {
this.chatService = chatService;
}
@PostMapping(value = "/stream", produces = MediaType.APPLICATION_NDJSON_VALUE)
public Flux<String> streamResponse(@RequestBody PromptDto prompt) {
return chatService.chatStream(prompt.getText());
}
}
Consume the Stream from the Browser
A client can read the NDJSON stream incrementally using the Streams API:
fetch('/api/stream', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ text: 'Explain quantum entanglement' })
})
.then(response => {
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
function read() {
return reader.read().then(({ done, value }) => {
if (done) return;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n');
buffer = lines.pop(); // keep incomplete line
lines.forEach(line => {
if (line) console.log(JSON.parse(line));
});
return read();
});
}
return read();
});
Summary
StreamingJacksonOutputConverterin Embabel bridges reactive LLM output and streaming JSON by writing tokens directly through Jackson'sJsonGenerator.- The converter is auto-configured as a Spring bean via
JacksonAutoConfigurationwhen Jackson is on the classpath. - As implemented in
embabel/embabel-agent, the typical flow moves fromOpenAiChatModel→StreamingJacksonOutputConverter→AgentController, returning aFlux<String>for minimal memory footprint. - Clients consume the stream as NDJSON or SSE, parsing each chunk independently for real-time UI updates.
Frequently Asked Questions
What does StreamingJacksonOutputConverter do in Embabel?
StreamingJacksonOutputConverter transforms a reactive stream of raw LLM chunks into a continuous JSON output. It uses Jackson's JsonGenerator to emit valid JSON tokens incrementally, avoiding the need to materialize the entire response in memory.
Do I need to manually register a Spring bean for StreamingJacksonOutputConverter?
No. As shown in the embabel/embabel-agent source code, JacksonAutoConfiguration registers the converter automatically when Jackson is detected on the classpath. You only need to inject it into your service or controller.
How does the converter keep memory usage low during large responses?
The converter processes each chunk as it arrives rather than accumulating the full LLM reply. Because it writes JSON tokens on-the-fly via JsonGenerator, the heap footprint remains nearly constant regardless of response size.
Which content type should I use for a streaming endpoint in Embabel?
The AgentController example in the embabel/embabel-agent repository produces MediaType.APPLICATION_NDJSON_VALUE, which sends newline-delimited JSON. This format lets clients parse one JSON object per line and is ideal for incremental, real-time consumption.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →