# How to Implement Streaming Responses with Server-Sent Events (SSE) in the Gemini API

> Enable streaming responses with Server-Sent Events SSE in the Gemini API. Learn how to implement chunked HTTP responses for real-time data using curl, Python, or JavaScript.

- Repository: [Google/skills](https://github.com/google/skills)
- Tags: how-to-guide
- Published: 2026-06-11

---

**Set `"stream": true` in your Gemini Interactions API request to enable Server-Sent Events (SSE), which returns chunked HTTP responses with `Content-Type: text/event-stream` that you can consume using curl, the Python SDK, or TypeScript/JavaScript.**

The Gemini Enterprise Agent Platform supports real-time streaming through standard HTTP SSE. According to the `google/skills` repository, implementing this requires only a single boolean flag in your request to switch from synchronous responses to chunked event streams that deliver incremental model output as it is generated.

## How SSE Streaming Works

When you POST to the interactions endpoint with `"stream": true`, the server switches from returning a complete JSON response to streaming incremental updates. As documented in [`skills/cloud/gemini-interactions-api/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gemini-interactions-api/SKILL.md) (lines 64-101), the response headers change to `Transfer-Encoding: chunked` and `Content-Type: text/event-stream`.

Each chunk follows the SSE protocol: a line starting with `data:` followed by a JSON object. The JSON structure includes an **event_type** field (typically `"step"`) and a **steps** array containing the incremental content generated by the model.

The stream terminates when the server sends the final chunk and closes the connection. This stateless approach uses the same `/interactions` endpoint for both synchronous and streaming calls, differing only by the `stream` flag.

## Raw HTTP Implementation with curl

For direct API access without SDKs, use `curl` with the `-N` flag to disable buffering. This allows you to see each `data:` line as it arrives from the server.

```bash
PROJECT_ID="my-gcp-project"
LOCATION="global"
ACCESS_TOKEN=$(gcloud auth print-access-token)
AGENT_ID="projects/${PROJECT_ID}/locations/${LOCATION}/agents/default"

curl -N -X POST "https://aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/${LOCATION}/interactions" \
  -H "Authorization: Bearer ${ACCESS_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
        "agent": "'"${AGENT_ID}"'",
        "stream": true,
        "input": [{
          "role": "user",
          "content": [{ "type": "text", "text": "Write a short story about a robot explorer." }]
        }]
      }'

```

The `-N` flag ensures that `curl` does not buffer the output, printing each SSE line immediately as the server sends it. The output format consists of lines beginning with `data:` followed by JSON objects:

```

data: {"event_type":"step","steps":[{"content":[{"type":"text","text":"..."}]}]}
data: {"event_type":"step","steps":[{"content":[{"type":"text","text":"..."}]}]}

```

## Streaming with the Gen AI SDKs

The Gen AI SDKs abstract the HTTP complexity into iterators, yielding the same JSON objects that the raw SSE endpoint returns.

### Python SDK (google-genai)

```python
from google import genai

# SDK automatically picks up GOOGLE_GENAI_USE_ENTERPRISE, GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION

client = genai.Client()

response = client.interactions.create(
    model="gemini-3-flash-preview",
    input="Explain the concept of quantum computing in two sentences.",
    stream=True,               # <-- enable SSE streaming

)

# Iterate over the streamed chunks

for chunk in response:
    if chunk.steps:
        # The newest step is the last element of the list

        step = chunk.steps[-1]
        text = step.content[0].text
        print(text, end="", flush=True)   # prints as soon as the chunk arrives

print()   # final newline

```

The generator yields partial responses as they arrive from the server, allowing you to render text in real time.

### TypeScript/JavaScript SDK (@google/genai)

```typescript
import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI();   // env vars set as in the docs

const responseStream = await ai.interactions.create({
  model: "gemini-3-flash-preview",
  input: "Give a brief history of the internet.",
  stream: true,                # <-- enable SSE streaming

});

// Async iteration over the streamed chunks
for await (const chunk of responseStream) {
  if (chunk.steps) {
    const step = chunk.steps[chunk.steps.length - 1];
    const text = step.content[0].text;
    process.stdout.write(text);
  }
}
process.stdout.write("\n");

```

The `for await...` loop reads each SSE event as soon as it arrives, reproducing the same real-time experience as `curl`.

## Browser Implementation with the Fetch API

For web applications, consume the stream using the Fetch API and a **TextDecoder** to parse the SSE chunks.

```javascript
const url = `https://aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/global/interactions`;
const body = {
  agent: AGENT_ID,
  stream: true,
  input: [{ role: "user", content: [{ type: "text", text: "Describe a sunset in poetry." }] }]
};

fetch(url, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${ACCESS_TOKEN}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify(body)
}).then(resp => {
  const reader = resp.body.getReader();
  const decoder = new TextDecoder("utf-8");
  function read() {
    return reader.read().then(({done, value}) => {
      if (done) return;
      const chunk = decoder.decode(value);
      // SSE lines are separated by "\n"
      for (const line of chunk.split("\n")) {
        if (line.startsWith("data:")) {
          const json = JSON.parse(line.slice(5));
          const step = json.steps?.[json.steps.length - 1];
          if (step?.content?.[0]?.text) {
            document.body.insertAdjacentText("beforeend", step.content[0].text);
          }
        }
      }
      return read();
    });
  }
  read();
});

```

This implementation processes the stream line-by-line, extracting the JSON payload after the `data:` prefix and appending the text content to the DOM as it arrives.

## Architecture and Performance Considerations

The implementation details in [`skills/cloud/gemini-interactions-api/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gemini-interactions-api/SKILL.md) (lines 438-461) reveal several key architectural decisions that optimize streaming performance:

**Stateless HTTP endpoint** – The same `/interactions` path handles both request types, simplifying routing and load balancing.

**Chunked transfer encoding** – This avoids buffering the entire model output on the server, reducing latency and memory usage for both client and server.

**Back-pressure handling** – The server only transmits the next chunk when the client reads the previous one, providing natural flow control without additional configuration.

**Universal compatibility** – Any HTTP client supporting SSE—including browsers, curl, Python's `requests` library, or Node.js fetch—can consume these streams.

## Summary

- Set `"stream": true` in requests to `https://aiplatform.googleapis.com/v1beta1/projects/<PROJECT_ID>/locations/<LOCATION>/interactions` to enable SSE mode.
- Expect `Content-Type: text/event-stream` and `Transfer-Encoding: chunked` headers in the response.
- Parse lines starting with `data:` containing JSON objects with `event_type` and `steps` fields.
- Use `curl -N` for command-line testing, or SDK iterators for Python and TypeScript applications.
- Leverage standard HTTP clients in browsers with the Fetch API and TextDecoder.
- The same endpoint handles both synchronous and streaming requests; only the `stream` flag differs.

## Frequently Asked Questions

### What is the difference between SSE streaming and the Live API in Gemini?

The SSE approach documented in [`skills/cloud/gemini-interactions-api/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gemini-interactions-api/SKILL.md) uses unidirectional HTTP with chunked transfer encoding, while the Live API referenced in [`skills/cloud/gemini-api/references/live_api.md`](https://github.com/google/skills/blob/main/skills/cloud/gemini-api/references/live_api.md) establishes a bidirectional WebSocket connection. SSE is simpler for one-way streaming from server to client, whereas the Live API supports full-duplex communication for interactive use cases.

### How do I handle back-pressure when consuming SSE streams?

The Gemini API implements natural back-pressure through HTTP's chunked transfer encoding. As noted in the source documentation, the server only sends the next chunk when the client reads the previous one, meaning slow consumers automatically throttle the stream without requiring explicit rate-limiting code.

### Can I use SSE streaming with any HTTP client?

Yes. Any client that understands SSE—including curl, browsers, Python's `requests` library with streaming enabled, or Node.js fetch—can consume these streams. The `google-genai` and `@google/genai` SDKs simply wrap this standard protocol with language-specific iterators for convenience.

### What JSON structure is returned in each SSE chunk?

Each line prefixed with `data:` contains a JSON object with an `event_type` field (typically `"step"`) and a `steps` array. Each step contains a `content` array with text elements. According to the specification in [`skills/cloud/gemini-interactions-api/SKILL.md`](https://github.com/google/skills/blob/main/skills/cloud/gemini-interactions-api/SKILL.md), the incremental content arrives as new elements in the `steps` list, allowing you to append the latest text to your output buffer.