# How to Run OpenClaude as a Headless gRPC Service: Complete Implementation Guide

> Learn to run OpenClaude as a headless gRPC service with this implementation guide. Expose AgentService for bidirectional streaming of messages, tool executions, and permission prompts.

- Repository: [Gitlawb/openclaude](https://github.com/Gitlawb/openclaude)
- Tags: how-to-guide
- Published: 2026-09-05

---

**OpenClaude can be deployed as a stand-alone gRPC server by executing the [`scripts/start-grpc.ts`](https://github.com/Gitlawb/openclaude/blob/main/scripts/start-grpc.ts) bootstrap script, which initializes the `GrpcServer` class from [`src/grpc/server.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/grpc/server.ts) and exposes the `AgentService` on port 50051 by default, enabling bidirectional streaming of chat messages, tool executions, and permission prompts without the interactive CLI.**

OpenClaude is an open-source AI agent framework that supports headless deployment through a native gRPC interface. Running OpenClaude as a headless gRPC service allows you to integrate its agent capabilities into microservices, automated pipelines, or distributed systems without terminal interaction. This guide covers the complete architecture, source code implementation in the Gitlawb/openclaude repository, and practical client examples.

## Understanding the OpenClaude gRPC Architecture

### Service Definition in Protocol Buffers

The gRPC interface is defined in `src/proto/openclaude.proto`, which according to the source code declares the `AgentService` with a bidirectional streaming RPC named `Chat`. This service accepts a stream of `ChatRequest` messages from clients and returns `ServerMessage` events. The protocol supports incremental text delivery through `text_chunk` events, tool execution signals via `tool_start` and `tool_result`, permission workflows using `action_required`, and termination markers in `done` messages.

### Core Server Implementation

The `GrpcServer` class in [`src/grpc/server.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/grpc/server.ts) implements the service logic by constructing a `grpc.Server`, registering the `AgentService`, and binding the `handleChat` method to manage duplex streams. For each incoming connection, the server instantiates a `QueryEngine`, injecting the working directory, available tools, and built-in agents sourced from [`src/tools/AgentTool/builtInAgents.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/tools/AgentTool/builtInAgents.ts).

The server maintains session state in an in-memory `Map` capped at `MAX_SESSIONS = 1000`. When clients provide a `session_id` in their `ChatRequest`, the server stores accumulated message history (`previousMessages = [...engine.getMessages()]`) and re-injects it on subsequent requests with the same ID, enabling true multi-turn conversations across separate gRPC streams.

## Starting the Headless gRPC Server

To launch the headless service, use the bootstrap script [`scripts/start-grpc.ts`](https://github.com/Gitlawb/openclaude/blob/main/scripts/start-grpc.ts). This script performs the same initialization steps as the CLI—loading configurations, applying provider environment variables, and hydrating secure tokens—before constructing the `GrpcServer` instance.

Prerequisites include Node.js version 22 or higher (or Bun). Install dependencies and start the server:

```bash
bun install
bun run scripts/start-grpc.ts

```

By default, the server binds to **localhost** on port **50051**. Override these defaults using environment variables:

```bash
GRPC_HOST=0.0.0.0 GRPC_PORT=6000 bun run scripts/start-grpc.ts

```

## Implementing gRPC Clients for OpenClaude

Once the server is running, any gRPC-capable client can connect. The protocol supports bidirectional streaming, allowing clients to send initial requests, respond to permission prompts, and issue cancellation signals.

### Node.js Client Example

The following client demonstrates connecting to the service, sending a chat request, and handling the various server message types including auto-approving tool permissions:

```typescript
import * as grpc from '@grpc/grpc-js';
import * as protoLoader from '@grpc/proto-loader';
import path from 'path';

const PROTO_PATH = path.resolve(import.meta.dirname, '../src/proto/openclaude.proto');
const packageDef = protoLoader.loadSync(PROTO_PATH, {
  keepCase: true,
  longs: String,
  enums: String,
  defaults: true,
  oneofs: true,
});
const { openclaude } = grpc.loadPackageDefinition(packageDef) as any;

const client = new openclaude.v1.AgentService(
  'localhost:50051',
  grpc.credentials.createInsecure()
);
const stream = client.Chat();

function sendRequest(message: string, sessionId = '') {
  stream.write({ request: { message, session_id: sessionId } });
}

stream.on('data', (msg: any) => {
  if (msg.text_chunk) {
    process.stdout.write(msg.text_chunk.text);
  } else if (msg.tool_start) {
    console.log('\n[Tool started]:', msg.tool_start.tool_name);
  } else if (msg.tool_result) {
    console.log('[Tool result]:', msg.tool_result.output);
  } else if (msg.action_required) {
    console.log('\n[Permission required]:', msg.action_required.question);
    // Auto-approve for demo purposes
    stream.write({
      input: { prompt_id: msg.action_required.prompt_id, reply: 'yes' },
    });
  } else if (msg.done) {
    console.log('\n[Complete]:', msg.done.full_text);
    stream.end();
  }
});

sendRequest('Explain the difference between promises and async/await', 'demo-session');

```

### Python Client Example

For Python implementations, generate bindings from the proto file using `grpcio-tools`, then implement the streaming logic to handle tool permission workflows:

```python
import grpc
import openclaude_pb2 as oc
import openclaude_pb2_grpc as ocg

def request_generator():
    yield oc.ClientMessage(
        request=oc.ChatRequest(message='Summarize the repository structure', session_id='py-demo')
    )
    
    # Generator will be resumed when we need to send UserInput responses

    while True:
        response = yield
        if response:
            yield response

def main():
    channel = grpc.insecure_channel('localhost:50051')
    stub = ocg.AgentServiceStub(channel)
    
    # Create generator for bidirectional streaming

    gen = request_generator()
    first_request = next(gen)
    
    for resp in stub.Chat(iter([first_request])):
        if resp.HasField('text_chunk'):
            print(resp.text_chunk.text, end='', flush=True)
        elif resp.HasField('action_required'):
            print(f'\n[Permission]: {resp.action_required.question}')
            # Send approval

            gen.send(oc.ClientMessage(
                input=oc.UserInput(
                    prompt_id=resp.action_required.prompt_id, 
                    reply='yes'
                )
            ))
        elif resp.HasField('tool_start'):
            print(f'\n[Tool]: Starting {resp.tool_start.tool_name}')
        elif resp.HasField('done'):
            print('\n[Finished]')
            break

if __name__ == '__main__':
    main()

```

## Managing Sessions and Tool Permissions

### Cross-Stream Session Persistence

The server implements stateful conversation management through the `session_id` field. As implemented in [`src/grpc/server.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/grpc/server.ts), when a client provides a non-empty `session_id`, the server extracts the message history from the `QueryEngine` after each turn and stores it in the sessions Map. On subsequent connections with the same ID, the server re-injects this history into a new `QueryEngine` instance, allowing the agent to maintain context across separate gRPC streams.

### Handling Tool Permission Flows

Before executing a tool, the server emits an `action_required` message containing a unique `prompt_id` and the permission question. The client must respond with a `UserInput` message containing the matching `prompt_id` and a reply string. The `handleChat` method in [`src/grpc/server.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/grpc/server.ts) maintains a pending promise for each active prompt and resolves it based on the client's reply, either proceeding with the tool execution or canceling the operation.

## Graceful Cancellation and Error Handling

The server handles client-initiated cancellation through the `CancelSignal` message type. When received in [`src/grpc/server.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/grpc/server.ts), the server marks the stream as interrupted, invokes `engine.interrupt('grpc_cancel')`, and terminates the duplex stream. The implementation also attaches an `end` handler to the stream that clears any outstanding permission prompts, preventing memory leaks and deadlocks when clients disconnect unexpectedly during permission workflows.

## Summary

- OpenClaude exposes a production-ready gRPC interface via `AgentService` defined in `src/proto/openclaude.proto`, supporting bidirectional streaming through the `Chat` RPC.
- The `GrpcServer` class in [`src/grpc/server.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/grpc/server.ts) manages stream lifecycle, session persistence (capped at 1000 sessions), and tool permission workflows using `prompt_id` correlation.
- Launch the headless service using [`scripts/start-grpc.ts`](https://github.com/Gitlawb/openclaude/blob/main/scripts/start-grpc.ts), configurable via `GRPC_HOST` and `GRPC_PORT` environment variables with defaults of localhost and 50051.
- Clients must implement handlers for `action_required` messages to approve or deny tool executions by returning `UserInput` messages with matching `prompt_id` values.
- The protocol supports graceful shutdown through `CancelSignal` and automatic cleanup of pending permissions on client disconnect.

## Frequently Asked Questions

### What is the default port for OpenClaude's gRPC server?

The default port is **50051**, and the default host is **localhost**. You can override these values by setting the `GRPC_PORT` and `GRPC_HOST` environment variables before running [`scripts/start-grpc.ts`](https://github.com/Gitlawb/openclaude/blob/main/scripts/start-grpc.ts).

### How does OpenClaude maintain conversation state across multiple gRPC connections?

The server maintains an in-memory `Map` of session histories, capped at `MAX_SESSIONS = 1000`. When clients include a `session_id` in their `ChatRequest`, the server preserves the conversation context from the `QueryEngine` and re-injects it on subsequent requests with the same ID, enabling stateful multi-turn dialogues across separate streams.

### Can I programmatically approve all tool calls without user intervention?

Yes, though this reduces security. Implement your client to listen for `action_required` server messages and immediately respond with a `UserInput` containing the provided `prompt_id` and reply value of `yes`. The Node.js and Python examples above demonstrate this auto-approval pattern suitable for trusted environments.

### What happens if a client disconnects during a tool permission prompt?

The server implements an `end` handler on the duplex stream that clears any outstanding permission promises when the connection terminates. This prevents memory leaks and ensures the server remains stable even when clients disconnect unexpectedly, as implemented in the error handling logic of [`src/grpc/server.ts`](https://github.com/Gitlawb/openclaude/blob/main/src/grpc/server.ts).