How nGPT Implements Streaming Responses and Real-Time Markdown Rendering

nGPT achieves real-time markdown rendering by combining Server-Sent Events (SSE) streaming from OpenAI-compatible APIs with Rich's Live display panels, coordinated through thread-safe callbacks in ngpt/ui/renderers.py.

The nGPT CLI tool provides a seamless streaming experience that displays AI responses as they arrive, formatted as markdown in real-time. This article examines the underlying mechanism for streaming responses and real-time markdown rendering in nGPT, breaking down the architecture into three coordinated components: HTTP-SSE streaming, live UI rendering, and thread-safe coordination.

Core Architecture Overview

nGPT's streaming capability relies on three integrated pieces working together:

  • HTTP-SSE Streaming Layer – Handles Server-Sent Events from the API and forwards chunks via callbacks
  • Live Markdown UI Layer – Manages Rich's Live display for real-time markdown rendering
  • Thread-Safe Coordination – Ensures the spinner, live panel, and terminal output never conflict

HTTP-SSE Streaming Layer

The streaming process begins in ngpt/api/client.py, where the NGPTClient.chat() method handles the HTTP request to OpenAI-compatible endpoints.

When stream=True is passed, the client:

  1. Sends the request with streaming enabled
  2. Reads the response line-by-line as Server-Sent Events (SSE)
  3. Parses each JSON chunk to extract the content field
  4. Accumulates the text and forwards it to the stream_callback function

This design decouples the network layer from the UI layer. The client knows nothing about markdown rendering; it simply delivers text chunks as they arrive from the API.

Live Markdown UI Layer

The visual magic happens in ngpt/ui/renderers.py, specifically in the prettify_streaming_markdown() function. This function sets up a Rich Live display that updates in real-time as new content arrives.

The implementation creates three objects:

  • live_display – A rich.live.Live context manager wrapping a Panel(Markdown(...))
  • update_content – A callback function that rewrites the Markdown renderable and calls live.refresh()
  • setup_spinner – A factory for creating the initial loading spinner

When update_content() receives new text, it:

  1. Truncates content to visible height if necessary
  2. Updates the Markdown renderable with the accumulated text
  3. Calls live.refresh() to force immediate terminal redraw

This creates the illusion of real-time markdown rendering, with headers, bold text, and code fences formatting correctly as they stream in.

Thread-Safe Coordination

Terminal rendering requires careful synchronization to prevent the spinner, live panel, and final output from colliding. nGPT solves this through a global TERMINAL_RENDER_LOCK in ngpt/ui/renderers.py.

The create_spinner_handling_callback() function wraps the original callback to manage the spinner lifecycle:

  1. Before the API request starts, a spinner thread begins running
  2. On the first content chunk, the wrapper stops the spinner and clears its line
  3. Subsequent chunks flow directly to the UI callback

This ensures the user sees a loading indicator during network latency, which seamlessly transitions to the streaming markdown display once content arrives.

Implementation Walkthrough

Here is the complete flow when a user runs ngpt text "Explain quantum computing":

  1. Command Entry – text_mode() in ngpt/cli/modes/text.py parses arguments and determines whether to use plain text or stream-prettify mode.

  2. UI Setup – For prettify mode, it calls prettify_streaming_markdown() to create the Live panel, update callback, and spinner factory.

  3. Spinner Start – A daemon thread starts the spinner via setup_spinner(), displaying a loading indicator while waiting for the API.

  4. API Request – NGPTClient.chat() sends the HTTP request with stream=True and the wrapped callback.

  5. SSE Processing – As the API streams SSE events, each chunk is parsed and the accumulated text is passed to the spinner-handling wrapper.

  6. First Chunk Handling – The wrapper stops the spinner, clears its line, and calls update_content() to start the Live display.

  7. Continuous Updates – Each subsequent chunk updates the Markdown panel via live.refresh(), rendering headers, lists, and code blocks in real-time.

  8. Completion – When the [DONE] marker arrives, update_content(..., complete=True) calls live.stop(), finalizing the display and releasing the terminal.

Code Examples

Minimal Streaming Client (No UI)

from ngpt.api.client import NGPTClient

def my_callback(accumulated):
    # Called after each chunk arrives

    print(accumulated, end='', flush=True)

client = NGPTClient(api_key="sk-...")
client.chat(
    prompt="Explain quantum tunnelling.",
    stream=True,
    stream_callback=my_callback,
)

This example demonstrates the raw streaming mechanism where the client handles SSE parsing and delivers accumulated text to your callback.

Real-Time Markdown UI (Full Implementation)

from ngpt.ui.renderers import prettify_streaming_markdown, create_spinner_handling_callback
from ngpt.api.client import NGPTClient
import threading

# Build the live markdown panel

live, update_md, spinner_factory = prettify_streaming_markdown()

# Start spinner before API call

stop_event = threading.Event()
stop_spinner = spinner_factory(stop_event, color="cyan")

# Wrap callback to handle spinner cleanup

wrapped_cb = create_spinner_handling_callback(
    original_callback=update_md,
    stop_spinner_func=stop_spinner,
    first_content_received_ref=[False],
)

# Execute streaming request

client = NGPTClient()
client.chat(
    prompt="Write a short poem about sunrise in markdown.",
    stream=True,
    stream_callback=wrapped_cb,
)

# Live panel stops automatically when stream completes

Running this snippet produces a Rich panel that updates line-by-line, interpreting Markdown syntax as it streams in.

Key Files

  • ngpt/api/client.py – Implements HTTP requests, SSE parsing, and the stream_callback forwarding mechanism.
  • ngpt/ui/renderers.py – Contains prettify_streaming_markdown(), create_spinner_handling_callback(), and the TERMINAL_RENDER_LOCK for thread-safe terminal access.
  • ngpt/cli/modes/text.py – Orchestrates the text mode command, deciding between plain output and stream-prettify mode.
  • ngpt/ui/tui.py – Provides spinner utilities and terminal helpers used during the streaming flow.
  • ngpt/ui/colors.py – Defines the color scheme used for spinners and panel borders.

Summary

  • nGPT uses Server-Sent Events (SSE) to stream API responses chunk-by-chunk through NGPTClient.chat().
  • The Rich library's Live display renders Markdown in real-time via prettify_streaming_markdown() in ngpt/ui/renderers.py.
  • A global terminal lock and spinner-handling wrapper ensure the loading indicator, live panel, and final output never conflict.
  • The architecture cleanly separates network logic (ngpt/api/client.py) from presentation logic (ngpt/ui/renderers.py), enabling both minimal callbacks and full interactive UIs.

Frequently Asked Questions

How does nGPT handle the transition from loading spinner to streaming text?

nGPT uses create_spinner_handling_callback() in ngpt/ui/renderers.py to wrap the markdown update function. This wrapper tracks whether content has arrived via a first_content_received_ref list. On the first chunk, it calls the stop_spinner_func to clear the loading indicator, then forwards subsequent chunks directly to the live markdown panel.

What prevents the spinner and live markdown panel from writing over each other?

A global TERMINAL_RENDER_LOCK defined in ngpt/ui/renderers.py ensures thread-safe access to the terminal. The spinner runs in a daemon thread while the main thread handles the HTTP stream and Rich's Live display. The lock guarantees that only one component writes to stdout at any moment, preventing visual corruption.

Can I use nGPT's streaming mechanism without the markdown rendering?

Yes. The NGPTClient.chat() method in ngpt/api/client.py accepts any callable as stream_callback. You can provide a simple function that prints raw text or processes it programmatically without importing ngpt/ui/renderers.py or using Rich's Live display.

Why does nGPT truncate content to the visible height during streaming?

The update_content() function in prettify_streaming_markdown() truncates accumulated text to the terminal's visible height to prevent the Rich Live panel from growing indefinitely. This ensures the display remains performant and readable during long streaming sessions, showing only the most recent content that fits the screen.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →