How to Implement Streaming Responses with the GitHub Copilot SDK
The GitHub Copilot SDK enables real-time streaming of model-generated content via Server-Sent Events (SSE) by setting the streaming option to true, which emits incremental assistant.streaming_delta events as tokens arrive from the model.
Real-time streaming reduces perceived latency in AI applications by delivering tokens as they are generated rather than waiting for complete responses. To implement streaming responses with the GitHub Copilot SDK, you configure the client with a boolean flag and listen for specific event types emitted by the session layer. The SDK handles the underlying Server-Sent Events (SSE) transport, parsing, and type-safe event distribution according to the source code in github/copilot-sdk.
Architecture of Streaming in the Copilot SDK
The streaming implementation spans multiple layers of the SDK, from client initialization to low-level transport handling.
Client Configuration Layer
In nodejs/src/client.ts (approximately lines 1560 and 1780), the CopilotClient constructor accepts a streaming option that propagates to all subsequent requests. When enabled, this flag signals the SDK to request SSE transport from the model endpoint.
Request Payload Structure
The request adapter in test/harness/responsesApiAdapter.ts (line 57) demonstrates how the SDK injects the stream: true field into the JSON payload sent to the model layer. This parameter instructs the upstream API to return a streaming response rather than a complete JSON object.
CAPI Proxy and SSE Transport
The test/harness/replayingCapiProxy.ts file (lines 517-587) contains the proxy logic that detects the stream: true flag and sets the response header Content-Type: text/event-stream. This component either streams cached responses slowly or forwards live responses instantly, maintaining the SSE format required by the client.
Session Event Management
At nodejs/src/session.ts (approximately line 230), the Session object listens for raw SSE chunks from the transport layer. It parses these chunks and emits high-level SDK events including assistant.streaming_delta and assistant.reasoning_delta, allowing applications to process incremental content without handling raw HTTP streams.
Type Definitions for Streaming Events
The generated types in nodejs/src/generated/session-events.ts (lines 3340-3385) define concrete interfaces for streaming events. These TypeScript definitions provide autocompletion and type safety for properties like delta (the text fragment) and cumulative_bytes (total bytes received).
Implementing Streaming Responses
The following patterns demonstrate how to enable and consume streaming responses in production code.
Global Streaming Configuration
Enable streaming for all requests by setting the streaming option when instantiating CopilotClient:
import { CopilotClient } from '@github/copilot-sdk';
const client = new CopilotClient({
// ... other configuration
streaming: true, // Enables SSE for all sessions
});
Consuming Token Deltas
After creating a chat session, register listeners for the assistant.streaming_delta event to receive tokens as they arrive:
const session = await client.createChatSession();
session.on('assistant.streaming_delta', (event) => {
// event.delta contains the newest text fragment
// event.cumulative_bytes tracks total bytes received
process.stdout.write(event.delta);
});
Per-Request Streaming Control
Override the client default for individual requests by passing the streaming flag directly to the send method:
await session.send({
model: 'gpt-4',
messages: [{ role: 'user', content: 'Explain async/await.' }],
streaming: true, // Overrides client default for this call only
});
Detecting Stream Completion
Listen for the assistant.message_stop event to handle the final aggregated message after streaming concludes:
session.on('assistant.message_stop', (msg) => {
console.log('\nFull response:', msg.content);
});
End-to-End Streaming Test Implementation
The SDK's test suite in nodejs/test/e2e/streaming_fidelity.e2e.test.ts validates the streaming pipeline. This pattern demonstrates how to verify that delta events fire correctly:
it('should produce delta events when streaming is enabled', async () => {
const client = await createTestClient({ streaming: true });
const session = await client.createChatSession();
const deltas: string[] = [];
session.on('assistant.streaming_delta', (e) => deltas.push(e.delta));
await session.send({
model: 'gpt-4',
messages: [{ role: 'user', content: 'Hi' }]
});
expect(deltas.length).toBeGreaterThan(0);
});
Summary
- Enable streaming globally via the
streaming: trueoption inCopilotClientconfiguration - The SDK uses Server-Sent Events (SSE) with
Content-Type: text/event-streamas implemented in the CAPI proxy layer - Listen for
assistant.streaming_deltaevents on the session object to process incremental tokens - Per-request streaming overrides are supported via the
streamingparameter insession.send() - Type definitions in
nodejs/src/generated/session-events.tsensure TypeScript safety for streaming payloads
Frequently Asked Questions
What transport protocol does the GitHub Copilot SDK use for streaming?
The SDK uses Server-Sent Events (SSE) via HTTP when streaming is enabled, as indicated by the Content-Type: text/event-stream header set in test/harness/replayingCapiProxy.ts. The low-level RPC layer in nodejs/src/generated/rpc.ts (lines 1150-1400) also supports WebSocket transport where available, though SSE remains the primary method for streaming model responses.
How do I access the cumulative byte count during streaming?
The assistant.streaming_delta event includes a cumulative_bytes property that tracks the total bytes received so far. According to the type definitions in nodejs/src/generated/session-events.ts, this field allows you to monitor download progress or implement custom buffering logic as tokens arrive from the model.
Can I enable streaming for only specific requests?
Yes. While you can set streaming: true in the CopilotClient constructor to enable it globally, you can also pass streaming: true to individual session.send() calls to override the default behavior for specific requests, as shown in the request payload handling at test/harness/responsesApiAdapter.ts.
What is the difference between assistant.streaming_delta and assistant.reasoning_delta?
Both events emit incremental content, but they target different model outputs. The assistant.streaming_delta event carries the primary response text tokens, while assistant.reasoning_delta provides intermediate reasoning steps when supported by the model. Both are defined in nodejs/src/generated/session-events.ts and emitted by the session layer in nodejs/src/session.ts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →