How to Implement Streaming Responses with Server-Sent Events (SSE) in the Gemini API

Set "stream": true in your Gemini Interactions API request to enable Server-Sent Events (SSE), which returns chunked HTTP responses with Content-Type: text/event-stream that you can consume using curl, the Python SDK, or TypeScript/JavaScript.

The Gemini Enterprise Agent Platform supports real-time streaming through standard HTTP SSE. According to the google/skills repository, implementing this requires only a single boolean flag in your request to switch from synchronous responses to chunked event streams that deliver incremental model output as it is generated.

How SSE Streaming Works

When you POST to the interactions endpoint with "stream": true, the server switches from returning a complete JSON response to streaming incremental updates. As documented in skills/cloud/gemini-interactions-api/SKILL.md (lines 64-101), the response headers change to Transfer-Encoding: chunked and Content-Type: text/event-stream.

Each chunk follows the SSE protocol: a line starting with data: followed by a JSON object. The JSON structure includes an event_type field (typically "step") and a steps array containing the incremental content generated by the model.

The stream terminates when the server sends the final chunk and closes the connection. This stateless approach uses the same /interactions endpoint for both synchronous and streaming calls, differing only by the stream flag.

Raw HTTP Implementation with curl

For direct API access without SDKs, use curl with the -N flag to disable buffering. This allows you to see each data: line as it arrives from the server.

PROJECT_ID="my-gcp-project"
LOCATION="global"
ACCESS_TOKEN=$(gcloud auth print-access-token)
AGENT_ID="projects/${PROJECT_ID}/locations/${LOCATION}/agents/default"

curl -N -X POST "https://aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/${LOCATION}/interactions" \
  -H "Authorization: Bearer ${ACCESS_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
        "agent": "'"${AGENT_ID}"'",
        "stream": true,
        "input": [{
          "role": "user",
          "content": [{ "type": "text", "text": "Write a short story about a robot explorer." }]
        }]
      }'

The -N flag ensures that curl does not buffer the output, printing each SSE line immediately as the server sends it. The output format consists of lines beginning with data: followed by JSON objects:


data: {"event_type":"step","steps":[{"content":[{"type":"text","text":"..."}]}]}
data: {"event_type":"step","steps":[{"content":[{"type":"text","text":"..."}]}]}

Streaming with the Gen AI SDKs

The Gen AI SDKs abstract the HTTP complexity into iterators, yielding the same JSON objects that the raw SSE endpoint returns.

Python SDK (google-genai)

from google import genai

# SDK automatically picks up GOOGLE_GENAI_USE_ENTERPRISE, GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION

client = genai.Client()

response = client.interactions.create(
    model="gemini-3-flash-preview",
    input="Explain the concept of quantum computing in two sentences.",
    stream=True,               # <-- enable SSE streaming

)

# Iterate over the streamed chunks

for chunk in response:
    if chunk.steps:
        # The newest step is the last element of the list

        step = chunk.steps[-1]
        text = step.content[0].text
        print(text, end="", flush=True)   # prints as soon as the chunk arrives

print()   # final newline

The generator yields partial responses as they arrive from the server, allowing you to render text in real time.

TypeScript/JavaScript SDK (@google/genai)

import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI();   // env vars set as in the docs

const responseStream = await ai.interactions.create({
  model: "gemini-3-flash-preview",
  input: "Give a brief history of the internet.",
  stream: true,                # <-- enable SSE streaming

});

// Async iteration over the streamed chunks
for await (const chunk of responseStream) {
  if (chunk.steps) {
    const step = chunk.steps[chunk.steps.length - 1];
    const text = step.content[0].text;
    process.stdout.write(text);
  }
}
process.stdout.write("\n");

The for await... loop reads each SSE event as soon as it arrives, reproducing the same real-time experience as curl.

Browser Implementation with the Fetch API

For web applications, consume the stream using the Fetch API and a TextDecoder to parse the SSE chunks.

const url = `https://aiplatform.googleapis.com/v1beta1/projects/${PROJECT_ID}/locations/global/interactions`;
const body = {
  agent: AGENT_ID,
  stream: true,
  input: [{ role: "user", content: [{ type: "text", text: "Describe a sunset in poetry." }] }]
};

fetch(url, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${ACCESS_TOKEN}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify(body)
}).then(resp => {
  const reader = resp.body.getReader();
  const decoder = new TextDecoder("utf-8");
  function read() {
    return reader.read().then(({done, value}) => {
      if (done) return;
      const chunk = decoder.decode(value);
      // SSE lines are separated by "\n"
      for (const line of chunk.split("\n")) {
        if (line.startsWith("data:")) {
          const json = JSON.parse(line.slice(5));
          const step = json.steps?.[json.steps.length - 1];
          if (step?.content?.[0]?.text) {
            document.body.insertAdjacentText("beforeend", step.content[0].text);
          }
        }
      }
      return read();
    });
  }
  read();
});

This implementation processes the stream line-by-line, extracting the JSON payload after the data: prefix and appending the text content to the DOM as it arrives.

Architecture and Performance Considerations

The implementation details in skills/cloud/gemini-interactions-api/SKILL.md (lines 438-461) reveal several key architectural decisions that optimize streaming performance:

Stateless HTTP endpoint – The same /interactions path handles both request types, simplifying routing and load balancing.

Chunked transfer encoding – This avoids buffering the entire model output on the server, reducing latency and memory usage for both client and server.

Back-pressure handling – The server only transmits the next chunk when the client reads the previous one, providing natural flow control without additional configuration.

Universal compatibility – Any HTTP client supporting SSE—including browsers, curl, Python's requests library, or Node.js fetch—can consume these streams.

Summary

  • Set "stream": true in requests to https://aiplatform.googleapis.com/v1beta1/projects/<PROJECT_ID>/locations/<LOCATION>/interactions to enable SSE mode.
  • Expect Content-Type: text/event-stream and Transfer-Encoding: chunked headers in the response.
  • Parse lines starting with data: containing JSON objects with event_type and steps fields.
  • Use curl -N for command-line testing, or SDK iterators for Python and TypeScript applications.
  • Leverage standard HTTP clients in browsers with the Fetch API and TextDecoder.
  • The same endpoint handles both synchronous and streaming requests; only the stream flag differs.

Frequently Asked Questions

What is the difference between SSE streaming and the Live API in Gemini?

The SSE approach documented in skills/cloud/gemini-interactions-api/SKILL.md uses unidirectional HTTP with chunked transfer encoding, while the Live API referenced in skills/cloud/gemini-api/references/live_api.md establishes a bidirectional WebSocket connection. SSE is simpler for one-way streaming from server to client, whereas the Live API supports full-duplex communication for interactive use cases.

How do I handle back-pressure when consuming SSE streams?

The Gemini API implements natural back-pressure through HTTP's chunked transfer encoding. As noted in the source documentation, the server only sends the next chunk when the client reads the previous one, meaning slow consumers automatically throttle the stream without requiring explicit rate-limiting code.

Can I use SSE streaming with any HTTP client?

Yes. Any client that understands SSE—including curl, browsers, Python's requests library with streaming enabled, or Node.js fetch—can consume these streams. The google-genai and @google/genai SDKs simply wrap this standard protocol with language-specific iterators for convenience.

What JSON structure is returned in each SSE chunk?

Each line prefixed with data: contains a JSON object with an event_type field (typically "step") and a steps array. Each step contains a content array with text elements. According to the specification in skills/cloud/gemini-interactions-api/SKILL.md, the incremental content arrives as new elements in the steps list, allowing you to append the latest text to your output buffer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →