How to Handle Streaming Responses from the Chat Method in Open Interpreter

Use interpreter.chat(message, stream=True, display=False) to receive a generator that yields real-time dictionaries containing LLM tokens, code execution output, and system messages as they are generated.

Open Interpreter's conversation loop is built around a generator-style API that enables real-time processing of large language model responses. When you handle streaming responses from the chat method in Open Interpreter, you gain access to an iterator that yields chunks immediately as they arrive from the language model or computer-execution layer, rather than waiting for the entire response to complete. This architecture, implemented in the openinterpreter/open-interpreter repository, supports both the built-in terminal interface and custom integrations requiring immediate feedback.

How Streaming Works in Open Interpreter

The Generator Architecture

The chat() method in interpreter/core/core.py checks the stream parameter at line 188. When stream=True, the call delegates to _streaming_chat(), which returns a generator instead of a final message list. According to the source code, setting display=False ensures the wrapper skips the terminal interface and yields directly from _respond_and_store() at line 58.

Internal Execution Flow

The streaming path follows a precise pipeline through the codebase:

  1. Entry Point: interpreter/core/core.py – chat() method (line 188) routes to _streaming_chat() when streaming is enabled.

  2. Raw LLM Generation: interpreter/core/respond.py – respond() (line 87) communicates with interpreter.llm.run and yields raw chunks as the LLM produces tokens.

  3. Post-Processing: interpreter/core/core.py – _respond_and_store() (lines 17-85) consumes these raw chunks, filters ephemeral messages like active_line or review, assembles new messages, and yields processed dictionaries.

  4. Terminal Interface (Optional): interpreter/terminal_interface/terminal_interface.py – terminal_interface() (lines 62-64) calls interpreter.chat(..., display=False, stream=True) and forwards each chunk to render markdown and code blocks when display=True.

Consuming Streaming Responses

Basic Token-by-Token Processing

To capture the LLM output character-by-character, iterate over the generator returned by chat():

from openinterpreter import Interpreter

interpreter = Interpreter()

for chunk in interpreter.chat(
    "Explain the difference between recursion and iteration.",
    stream=True,
    display=False
):
    # Each chunk is a dict: {"role": "assistant", "type": "message", "content": "..."}

    print(chunk["content"], end="", flush=True)

This pattern yields dictionaries as soon as tokens arrive from the LLM, enabling near-real-time display without waiting for the complete response.

Handling Code Execution Chunks

When the LLM generates code, the stream includes confirmation chunks requiring user approval and output chunks containing execution results:

for chunk in interpreter.chat(
    "Create a file named hello.txt with 'Hello, world!' and show its contents.",
    stream=True,
    display=False
):
    if chunk["type"] == "output":
        print("\n🖥️ Output:", chunk["content"])
    elif chunk["type"] == "confirmation":
        # Auto-run confirmed code (use with caution)

        interpreter.computer.run("bash", chunk["content"], stream=True, display=False)

The type field distinguishes between LLM reasoning (message), code blocks (confirmation), and execution results (output).

Async Integration for Custom UIs

For web applications or WebSocket servers, convert the generator into an async iterator:

import asyncio
from openinterpreter import Interpreter

interpreter = Interpreter()

async def stream_to_client(message):
    for chunk in interpreter.chat(message, stream=True, display=False):
        await websocket.send_json(chunk)

Because chat() returns a standard Python generator, it integrates natively with async frameworks when wrapped appropriately, allowing real-time streaming to front-end clients.

Key Source Files

Understanding these files helps debug streaming behavior or extend functionality:

Summary

  • stream=True transforms interpreter.chat() into a lazy iterator that yields chunks immediately as they arrive from the LLM or execution environment.
  • display=False bypasses the terminal interface, returning raw generator control to your code for custom processing pipelines.
  • Each yielded chunk is a dictionary containing role, type, and content fields, with types including message, confirmation, and output.
  • The streaming architecture flows through core.py → respond.py → core.py → optional terminal_interface.py, with key logic at lines 188, 87, and 58.
  • This pattern supports synchronous loops, conditional execution handling, and async WebSocket integrations.

Frequently Asked Questions

What is the difference between stream=True and display=True?

The stream parameter controls whether chat() returns a generator (when True) or a complete message list (when False). The display parameter determines whether the built-in terminal interface consumes that generator to render rich markdown and code blocks. When building custom integrations, use stream=True with display=False to receive raw chunks directly.

What data structure does each streaming chunk use?

Each chunk is a Python dictionary containing at minimum role (typically "assistant"), type (such as "message", "confirmation", or "output"), and content (the actual text). The _respond_and_store() method in core.py wraps raw LLM outputs with the assistant role before yielding them.

How do I filter out system chunks like active_line markers?

The _respond_and_store() method in interpreter/core/core.py (lines 17-85) automatically filters ephemeral chunks such as active_line and review markers before yielding to your code. When consuming the stream manually, you can additionally check the type field to ignore specific chunk categories.

Can I use streaming with async/await patterns?

Yes. While interpreter.chat() returns a standard Python generator, you can iterate over it within an async function using for chunk in interpreter.chat(...) and await asynchronous operations inside the loop, such as sending data over WebSockets. The underlying interpreter.llm.run implementation supports async generation for compatible LLM backends.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →