How to Build a FastAPI Server with Open Interpreter and Stream Responses
Open Interpreter provides a built-in FastAPI server via the AsyncInterpreter and Server classes that exposes OpenAI-compatible streaming endpoints using NDJSON chunks over HTTP or WebSocket.
The openinterpreter/open-interpreter repository includes a production-ready async server framework that transforms the local code-execution engine into a scalable API service. By wrapping AsyncInterpreter with the Server class, you can deploy Open Interpreter as a FastAPI application capable of streaming incremental LLM output to web clients in real time.
Architecture Overview
Open Interpreter’s server architecture centers on three core components defined in interpreter/core/async_core.py:
AsyncInterpreter– The asynchronous engine that runs LLM prompts, executes generated code, and pushes each output chunk to an internal queue (source lines 44‑85).Server– A FastAPI wrapper that mounts the router, adds optional API-key authentication, and launches a uvicorn process (source lines 51‑100).create_router– A factory function that declares HTTP, WebSocket, and OpenAI-compatible endpoints. The streaming route (/openai/chat/completions) yields NDJSON chunks as they become available (source lines 97‑126).
Installation and Setup
The core package is lightweight; the server requires additional extras that install FastAPI, uvicorn, janus, and starlette:
pip install "open-interpreter[server]"
This command satisfies the imports guarded by the try/except block at the top of async_core.py (source lines 21‑35).
Creating the FastAPI Server
Instantiate AsyncInterpreter, pass it to the Server constructor, and call run() to start listening on 0.0.0.0:8000:
from interpreter.core.async_core import AsyncInterpreter, Server
# 1️⃣ Create the interpreter with any Open-Interpreter settings.
interpreter = AsyncInterpreter()
# 2️⃣ Wrap it in the Server. This wires the router and optional auth middleware.
server = Server(interpreter)
# 3️⃣ Start the uvicorn server.
server.run()
During construction, the Server class initializes a FastAPI instance, invokes create_router(async_interpreter) to build endpoints, and installs an HTTP middleware that validates the X-API-KEY header against an authenticate_function when required (source lines 55‑68).
Streaming Responses via HTTP
The OpenAI-compatible endpoint lives at POST /openai/chat/completions. When the request body contains "stream": true, the route returns a StreamingResponse wrapping the openai_compatible_generator generator. This generator iterates over async_interpreter._respond_and_store(), yielding NDJSON payloads (data: {...}\n\n) immediately as each chunk is produced (source lines 19‑27, 33‑55 and lines 26‑55).
Request format
POST /openai/chat/completions HTTP/1.1
Content-Type: application/json
Accept: text/event-stream
{
"model": "gpt-4",
"messages": [{"role": "user", "content": "List the files in the current directory"}],
"stream": true
}
Python client example (httpx)
import httpx
import json
url = "http://127.0.0.1:8000/openai/chat/completions"
payload = {
"model": "gpt-4",
"messages": [{"role": "user", "content": "Explain the difference between threads and async IO"}],
"stream": True,
}
with httpx.stream("POST", url, json=payload) as response:
for line in response.iter_lines():
if not line:
continue
# Parse NDJSON: b'data: {...}'
data = json.loads(line.decode().removeprefix("data: ").strip())
chunk = data["choices"][0]["delta"]["content"]
print(chunk, end="", flush=True)
cURL example
curl -N -X POST http://127.0.0.1:8000/openai/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4",
"messages": [{"role":"user","content":"What is a FastAPI router?"}],
"stream": true
}'
The -N flag prevents cURL from buffering, displaying each data: line as soon as it arrives from the server.
WebSocket Alternative for Real-Time Interaction
For interactive front-ends, the same server exposes a WebSocket endpoint at /. This handler (lines 38‑84 in async_core.py) uses the LMC-chunk protocol: it reads inbound messages from the client, forwards them to async_interpreter.input, and pushes queued output back via the send_output coroutine. This is the same protocol powering the built-in HTML UI generated by create_router.
Securing the API with Authentication
If the environment variable INTERPRETER_REQUIRE_AUTH is set to "True", the server enforces header-based authentication. The middleware checks the X-API-KEY request header against the value stored in INTERPRETER_API_KEY (allowing all requests when the variable is unset). See the authenticate_function implementation around lines 78‑95 (source).
export INTERPRETER_REQUIRE_AUTH=True
export INTERPRETER_API_KEY=super-secret-key
Clients must then include the header:
curl -H "X-API-KEY: super-secret-key" http://127.0.0.1:8000/...
Summary
- Open Interpreter provides a ready-to-use FastAPI server through
AsyncInterpreterand theServerclass defined ininterpreter/core/async_core.py. - Streaming is implemented via an OpenAI-compatible endpoint that yields NDJSON chunks in real-time, allowing clients to process partial LLM output as it is generated.
- WebSocket support at
/enables bidirectional, interactive sessions using the LMC-chunk protocol. - Authentication is optional and controlled via
INTERPRETER_REQUIRE_AUTHandINTERPRETER_API_KEY, enforced by middleware checking theX-API-KEYheader. - Install server dependencies with
pip install "open-interpreter[server]"to obtain FastAPI, uvicorn, and related libraries.
Frequently Asked Questions
What dependencies are required to run the Open Interpreter FastAPI server?
You must install the server extras using pip install "open-interpreter[server]". This pulls in fastapi, uvicorn, janus, and starlette, which are required by the import block at the top of interpreter/core/async_core.py.
How does the streaming protocol work in the Open Interpreter FastAPI server?
The server uses NDJSON (Newline-Delimited JSON) over HTTP. When you POST to /openai/chat/completions with "stream": true, the endpoint returns a StreamingResponse that wraps openai_compatible_generator. This generator consumes chunks from async_interpreter._respond_and_store() and yields lines formatted as data: {...}\n\n, mirroring the OpenAI SSE specification.
Can I use standard OpenAI client libraries to connect to this server?
Yes. The endpoint at /openai/chat/completions is designed to be OpenAI-compatible. You can point any standard client (including the official OpenAI Python or JavaScript SDKs) to http://localhost:8000 and use the same request/response formats, including streaming parameters.
How do I enable authentication on the Open Interpreter server?
Set the environment variable INTERPRETER_REQUIRE_AUTH=True and define INTERPRETER_API_KEY to your secret token. The server middleware (lines 78‑95 in async_core.py) will then validate the X-API-KEY header on every request, rejecting calls that lack the correct key.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →