How to Use the PrivateGPT Summarization Recipe API: Complete HTTP Endpoint Guide

PrivateGPT exposes the summarization recipe via the /v1/summarize HTTP POST endpoint defined in private_gpt/server/recipes/summarize/summarize_router.py, delegating core logic to SummarizeService in private_gpt/server/recipes/summarize/summarize_service.py to generate document summaries with optional streaming support.

The PrivateGPT summarization recipe API allows developers to condense long texts or ingested documents into concise summaries through a RESTful interface. This endpoint leverages the SummaryIndex with tree-summarize queries to process content efficiently. Whether you need to summarize raw text or context from your vector store, this guide covers the complete implementation details based on the zylon-ai/private-gpt source code.

Endpoint Overview and Architecture

The summarization functionality is mounted at /v1/summarize within the FastAPI application. Request routing is handled by summarize_router defined in private_gpt/server/recipes/summarize/summarize_router.py, which validates incoming payloads and determines whether to return a synchronous response or a streaming output.

Core logic resides in SummarizeService (private_gpt/server/recipes/summarize/summarize_service.py). When invoked, this service constructs a list of nodes to summarize: optionally splitting input text into sentences and retrieving filtered document nodes from the internal docstore when contextual summarization is enabled. The service creates an in-memory SummaryIndex and executes a tree-summarize query against the configured LLM.

Request Payload Parameters

The endpoint accepts a JSON payload with the following fields:

  • text (string, optional): Raw text to summarize. If omitted while use_context is true, only ingested documents are processed.
  • use_context (boolean, default false): When enabled, pulls content from documents stored in the vector store.
  • context_filter (object, optional): Limits which ingested documents are considered. Defined in private_gpt/open_ai/extensions/context_filter.py, it accepts a docs_ids array to specify relevant documents.
  • prompt (string, optional): Overrides the default summarization prompt (DEFAULT_SUMMARIZE_PROMPT in summarize_service.py). Falls back to settings().ui.default_summarization_system_prompt if configured.
  • instructions (string, optional): Additional directives appended to the prompt to guide summary style or content.
  • stream (boolean, default false): When true, returns responses in OpenAI Server-Sent-Events format via to_openai_sse_stream from private_gpt/open_ai/openai_models.py.

Implementation Details

The router retrieves the SummarizeService instance via FastAPI's request state injection (request.state.injector.get(SummarizeService)). For non-streaming requests, it returns a SummarizeResponse containing a summary string. Streaming requests trigger SummarizeService.stream_summarize, which yields chunks transformed into the OpenAI SSE format.

The summarization pipeline combines input text nodes with retrieved document nodes from the docstore, applies the custom or default prompt, and executes the LLM query. This architecture is wired into the main application in private_gpt/launcher.py where summarize_router is included during app initialization.

API Usage Examples

Non-Streaming cURL Request

curl -X POST https://your-private-gpt.example.com/v1/summarize \
  -H "Authorization: Bearer <YOUR_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{
        "text": "Lorem ipsum dolor sit amet, consectetur adipiscing elit.",
        "use_context": false,
        "prompt": "Summarize the following in two sentences:",
        "instructions": "Keep the technical terms."
      }'

Response:

{
  "summary": "Lorem ipsum dolor sit amet, consectetur adipiscing elit. It describes..."
}

Streaming cURL Request

curl -N -X POST https://your-private-gpt.example.com/v1/summarize \
  -H "Authorization: Bearer <YOUR_TOKEN>" \
  -H "Content-Type: application/json" \
  -d '{"text":"Long document…","stream":true}'

Each line follows the OpenAI SSE format:


data: {"id":"...","object":"completion.chunk","choices":[{"delta":{"content":"First part of the summary"}}]}

Python Integration (Non-Streaming)

import httpx

api_url = "https://your-private-gpt.example.com/v1/summarize"
headers = {
    "Authorization": "Bearer <YOUR_TOKEN>",
    "Content-Type": "application/json",
}

payload = {
    "text": "Explain quantum entanglement in simple terms.",
    "use_context": True,
    "context_filter": {"docs_ids": ["doc-123", "doc-456"]},
    "instructions": "Make it understandable for a high-school audience.",
    "stream": False,
}

resp = httpx.post(api_url, headers=headers, json=payload)
print(resp.json()["summary"])

Python Integration (Streaming)

import httpx
import json

def sse_iter(response):
    for line in response.iter_lines():
        if line.startswith(b"data:"):
            data = json.loads(line[5:].strip())
            yield data["choices"][0]["delta"].get("content", "")

with httpx.stream("POST", api_url, headers=headers, 
                  json={"text": "Long text …", "stream": True}) as r:
    for chunk in sse_iter(r):
        print(chunk, end="")

Key Source Files Reference

Summary

  • The PrivateGPT summarization recipe API is available at POST /v1/summarize and implemented in summarize_router.py.
  • Request parameters include text, use_context, context_filter, prompt, instructions, and stream.
  • Implementation uses SummarizeService to build node lists and execute tree-summarize queries via SummaryIndex.
  • Streaming support conforms to OpenAI SSE format through utilities in openai_models.py.
  • Source files are located in private_gpt/server/recipes/summarize/ with configuration managed in settings.py and launcher.py.

Frequently Asked Questions

What is the exact endpoint URL for the PrivateGPT summarization recipe API?

The endpoint is /v1/summarize and accepts POST requests. It is defined in private_gpt/server/recipes/summarize/summarize_router.py and mounted to the main FastAPI application in private_gpt/launcher.py. All summarization requests must target this path with proper authentication headers.

How does the streaming response format work in the summarization API?

When stream is set to true, the endpoint returns a StreamingResponse using OpenAI Server-Sent-Events format. The SummarizeService.stream_summarize method yields chunks that are transformed by to_openai_sse_stream in private_gpt/open_ai/openai_models.py, with each line prefixed by data: and containing JSON with a choices array holding delta content.

Can I filter which ingested documents are used for summarization?

Yes. Set use_context to true and provide a context_filter object containing a docs_ids array. This filter is defined in private_gpt/open_ai/extensions/context_filter.py and limits the document nodes retrieved from the internal docstore during the summarization process.

Where is the default summarization prompt defined in PrivateGPT?

The default prompt is defined as DEFAULT_SUMMARIZE_PROMPT in private_gpt/server/recipes/summarize/summarize_service.py (lines 24-31). Alternatively, you can configure a UI-level default via default_summarization_system_prompt in private_gpt/settings/settings.py (lines 371-374), which the service uses when no custom prompt is provided.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →