# How to Build a FastAPI Server with Open Interpreter and Stream Responses

> Build a FastAPI server with Open Interpreter seamlessly. Learn to expose OpenAI-compatible streaming endpoints using NDJSON chunks for efficient data transfer. Get started now.

- Repository: [Open Interpreter/open-interpreter](https://github.com/openinterpreter/open-interpreter)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Open Interpreter provides a built-in FastAPI server via the `AsyncInterpreter` and `Server` classes that exposes OpenAI-compatible streaming endpoints using NDJSON chunks over HTTP or WebSocket.**

The openinterpreter/open-interpreter repository includes a production-ready async server framework that transforms the local code-execution engine into a scalable API service. By wrapping `AsyncInterpreter` with the `Server` class, you can deploy Open Interpreter as a FastAPI application capable of streaming incremental LLM output to web clients in real time.

## Architecture Overview

Open Interpreter’s server architecture centers on three core components defined in [`interpreter/core/async_core.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/async_core.py):

- **`AsyncInterpreter`** – The asynchronous engine that runs LLM prompts, executes generated code, and pushes each output chunk to an internal queue ([source lines 44‑85](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/async_core.py#L44-L85)).
- **`Server`** – A FastAPI wrapper that mounts the router, adds optional API-key authentication, and launches a **uvicorn** process ([source lines 51‑100](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/async_core.py#L51-L100)).
- **`create_router`** – A factory function that declares HTTP, WebSocket, and OpenAI-compatible endpoints. The streaming route (`/openai/chat/completions`) yields **NDJSON** chunks as they become available ([source lines 97‑126](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/async_core.py#L97-L126)).

## Installation and Setup

The core package is lightweight; the server requires additional extras that install **FastAPI**, **uvicorn**, **janus**, and **starlette**:

```bash
pip install "open-interpreter[server]"

```

This command satisfies the imports guarded by the `try/except` block at the top of [`async_core.py`](https://github.com/openinterpreter/open-interpreter/blob/main/async_core.py) ([source lines 21‑35](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/async_core.py#L21-L35)).

## Creating the FastAPI Server

Instantiate `AsyncInterpreter`, pass it to the `Server` constructor, and call `run()` to start listening on `0.0.0.0:8000`:

```python
from interpreter.core.async_core import AsyncInterpreter, Server

# 1️⃣ Create the interpreter with any Open-Interpreter settings.

interpreter = AsyncInterpreter()

# 2️⃣ Wrap it in the Server. This wires the router and optional auth middleware.

server = Server(interpreter)

# 3️⃣ Start the uvicorn server.

server.run()

```

During construction, the `Server` class initializes a FastAPI instance, invokes `create_router(async_interpreter)` to build endpoints, and installs an HTTP middleware that validates the `X-API-KEY` header against an `authenticate_function` when required ([source lines 55‑68](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/async_core.py#L55-L68)).

## Streaming Responses via HTTP

The **OpenAI-compatible** endpoint lives at `POST /openai/chat/completions`. When the request body contains `"stream": true`, the route returns a `StreamingResponse` wrapping the `openai_compatible_generator` generator. This generator iterates over `async_interpreter._respond_and_store()`, yielding **NDJSON** payloads (`data: {...}\n\n`) immediately as each chunk is produced ([source lines 19‑27, 33‑55](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/async_core.py#L19-L27) and [lines 26‑55](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/async_core.py#L26-L55)).

### Request format

```http
POST /openai/chat/completions HTTP/1.1
Content-Type: application/json
Accept: text/event-stream

{
  "model": "gpt-4",
  "messages": [{"role": "user", "content": "List the files in the current directory"}],
  "stream": true
}

```

### Python client example (httpx)

```python
import httpx
import json

url = "http://127.0.0.1:8000/openai/chat/completions"
payload = {
    "model": "gpt-4",
    "messages": [{"role": "user", "content": "Explain the difference between threads and async IO"}],
    "stream": True,
}

with httpx.stream("POST", url, json=payload) as response:
    for line in response.iter_lines():
        if not line:
            continue
        # Parse NDJSON: b'data: {...}'

        data = json.loads(line.decode().removeprefix("data: ").strip())
        chunk = data["choices"][0]["delta"]["content"]
        print(chunk, end="", flush=True)

```

### cURL example

```bash
curl -N -X POST http://127.0.0.1:8000/openai/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-4",
        "messages": [{"role":"user","content":"What is a FastAPI router?"}],
        "stream": true
      }'

```

The `-N` flag prevents cURL from buffering, displaying each `data:` line as soon as it arrives from the server.

## WebSocket Alternative for Real-Time Interaction

For interactive front-ends, the same server exposes a **WebSocket** endpoint at `/`. This handler (lines 38‑84 in [`async_core.py`](https://github.com/openinterpreter/open-interpreter/blob/main/async_core.py)) uses the LMC-chunk protocol: it reads inbound messages from the client, forwards them to `async_interpreter.input`, and pushes queued output back via the `send_output` coroutine. This is the same protocol powering the built-in HTML UI generated by `create_router`.

## Securing the API with Authentication

If the environment variable `INTERPRETER_REQUIRE_AUTH` is set to `"True"`, the server enforces header-based authentication. The middleware checks the `X-API-KEY` request header against the value stored in `INTERPRETER_API_KEY` (allowing all requests when the variable is unset). See the `authenticate_function` implementation around lines 78‑95 ([source](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/async_core.py#L78-L95)).

```bash
export INTERPRETER_REQUIRE_AUTH=True
export INTERPRETER_API_KEY=super-secret-key

```

Clients must then include the header:

```bash
curl -H "X-API-KEY: super-secret-key" http://127.0.0.1:8000/...

```

## Summary

- **Open Interpreter** provides a ready-to-use FastAPI server through `AsyncInterpreter` and the `Server` class defined in [`interpreter/core/async_core.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/async_core.py).
- **Streaming** is implemented via an OpenAI-compatible endpoint that yields NDJSON chunks in real-time, allowing clients to process partial LLM output as it is generated.
- **WebSocket** support at `/` enables bidirectional, interactive sessions using the LMC-chunk protocol.
- **Authentication** is optional and controlled via `INTERPRETER_REQUIRE_AUTH` and `INTERPRETER_API_KEY`, enforced by middleware checking the `X-API-KEY` header.
- Install server dependencies with `pip install "open-interpreter[server]"` to obtain FastAPI, uvicorn, and related libraries.

## Frequently Asked Questions

### What dependencies are required to run the Open Interpreter FastAPI server?

You must install the server extras using `pip install "open-interpreter[server]"`. This pulls in **fastapi**, **uvicorn**, **janus**, and **starlette**, which are required by the import block at the top of [`interpreter/core/async_core.py`](https://github.com/openinterpreter/open-interpreter/blob/main/interpreter/core/async_core.py).

### How does the streaming protocol work in the Open Interpreter FastAPI server?

The server uses **NDJSON** (Newline-Delimited JSON) over HTTP. When you POST to `/openai/chat/completions` with `"stream": true`, the endpoint returns a `StreamingResponse` that wraps `openai_compatible_generator`. This generator consumes chunks from `async_interpreter._respond_and_store()` and yields lines formatted as `data: {...}\n\n`, mirroring the OpenAI SSE specification.

### Can I use standard OpenAI client libraries to connect to this server?

Yes. The endpoint at `/openai/chat/completions` is designed to be **OpenAI-compatible**. You can point any standard client (including the official OpenAI Python or JavaScript SDKs) to `http://localhost:8000` and use the same request/response formats, including streaming parameters.

### How do I enable authentication on the Open Interpreter server?

Set the environment variable `INTERPRETER_REQUIRE_AUTH=True` and define `INTERPRETER_API_KEY` to your secret token. The server middleware (lines 78‑95 in [`async_core.py`](https://github.com/openinterpreter/open-interpreter/blob/main/async_core.py)) will then validate the `X-API-KEY` header on every request, rejecting calls that lack the correct key.