Implementing Streaming Tool Calls with Function Calling: A Python Guide for AI Agents

Implementing streaming tool calls with function calling allows language models to process incremental tool outputs through Python generators, enabling real-time interactions where the LLM consumes partial results via iterators rather than blocking for complete execution.

The ai-engineering-from-scratch repository provides a comprehensive curriculum for building AI systems from first principles, including production-ready patterns for streaming tool interactions. When implementing streaming tool calls with function calling, tools yield chunks of data incrementally while the LLM processes each fragment, enabling sub-second latency for voice assistants and live data pipelines. This approach is detailed across Phase 13 (Tools and Protocols) and Phase 14 (Agent Engineering), with reference implementations available in the production runtimes and voice assistant capstone modules.

Architectural Components

The implementation relies on four distinct architectural layers that coordinate the contract, execution, and consumption of streaming tools.

Tool Interface and Schema

The Tool Interface establishes a JSON-Schema contract that defines available functions and their arguments. This layer validates that the model can request specific tools with structured parameters, as covered in Phase 13, Lesson 01 (The Tool Interface). The schema ensures type safety even when processing incremental streams.

Function Calling Loop

The Function Calling Loop parses the model's tool_calls JSON payload, dispatches execution to the appropriate Python function, and handles return values. As implemented in phases/13-tools-and-protocols/09-function-calling/docs/en.md, this loop must detect whether a tool returns a static value or a streaming iterator to handle partial results correctly.

Streaming API

The Streaming API implements tools as Python generators using yield statements to produce partial results. Rather than returning a single string, these functions emit chunks (tokens, sensor readings, or database rows) as they become available. The curriculum covers this pattern in Phase 13, Lesson 03 (Parallel and Streaming Tool Calls), including provider-specific semantics for OpenAI, Gemini, and Anthropic APIs.

Scheduler and Coordinator

In multimodal systems, a Scheduler arbitrates between concurrent streaming sources including ASR (Automatic Speech Recognition), LLM generation, and TTS (Text-to-Speech). The implementation in phases/19-capstone-projects/03-realtime-voice-assistant/code/main.py demonstrates how this coordinator aligns timestamps and manages turn-taking between multiple simultaneous streams.

Execution Flow for Streaming Tools

When implementing streaming tool calls with function calling, the system follows a deterministic six-stage pipeline:

  1. Prompt Analysis: The LLM receives a user prompt and determines that external data retrieval is necessary.

  2. Tool Call Generation: The model outputs a JSON blob specifying the tool name and arguments, such as {"name": "streaming", "arguments": {"input_text": "Hello world"}}.

  3. Function Dispatch: The dispatcher looks up the registered Python function in the tool registry based on the name field.

  4. Streaming Invocation: The dispatcher executes the tool, which returns a generator iterator. The streaming function in phases/14-agent-engineering/29-production-runtimes/code/main.py yields individual words from the input text incrementally.

  5. Progressive Consumption: The LLM receives each yielded chunk through the function calling loop, updating its context window and potentially generating response tokens before the tool completes.

  6. Result Finalization: When the iterator exhausts, the accumulated result is returned to the user, or the LLM may request additional tool calls based on the streamed data.

Code Implementation

The following examples from the ai-engineering-from-scratch source code demonstrate the concrete implementation of streaming generators and their dispatch.

Defining a Streaming Tool Generator

In phases/14-agent-engineering/29-production-runtimes/code/main.py, the streaming function implements the generator pattern required for incremental output:


# File: phases/14-agent-engineering/29-production-runtimes/code/main.py

def streaming(input_text: str) -> Iterable[str]:
    """
    A simple streaming tool that yields the input text one word at a time.
    The function can be called by the LLM via the standard function‑calling
    loop; each yielded token is sent back to the model as a partial result.
    """
    for word in input_text.split():
        yield f"{word} "

This function implements the streaming runtime shape, one of four production shapes defined in Phase 14, Lesson 29 (Production Runtimes).

The Function Calling Loop

The dispatcher code must handle iterators differently from static values. The generic pattern from the Function Calling lesson demonstrates this logic:

def run_function_call(tool_name: str, args: dict) -> str:
    if tool_name == "streaming":
        # `streaming` returns an iterator; we concatenate its pieces.

        return "".join(streaming(args["input_text"]))
    # … handle other tools …

When the LLM produces a tool call such as:

{
  "name": "streaming",
  "arguments": {"input_text": "Hello world from the AI curriculum"}
}

The run_function_call function executes the generator. In true streaming implementations, each yielded value is sent to the LLM immediately rather than concatenated, enabling real-time, interactive experiences where the model reacts to partial data.

Production Runtime Shapes

The curriculum defines four distinct runtime shapes for tool execution in phases/14-agent-engineering/29-production-runtimes/code/main.py:

  • Request-Response: Standard blocking execution returning complete results as strings
  • Streaming: Generator-based execution yielding Iterable[str] chunks incrementally
  • Queue: Asynchronous message-based processing for background task offloading
  • Event: Reactive callbacks triggered by external system events

When implementing streaming tool calls with function calling, use inspect.isgeneratorfunction() or similar checks to distinguish between immediate returns and iterable streams, ensuring the dispatcher consumes generators appropriately.

Real-World Multimodal Example

The real-time voice assistant capstone in phases/19-capstone-projects/03-realtime-voice-assistant/code/main.py demonstrates advanced streaming coordination. This system manages three concurrent streams:

  • ASR: Continuous speech recognition yielding transcript chunks
  • LLM: Token-by-token response generation
  • TTS: Audio segment production

The scheduler synchronizes these streams, allowing the assistant to begin speaking while the LLM is still reasoning or while external tools are streaming partial results. This architecture achieves sub-second latency by implementing streaming tool calls that do not block the conversation loop.

Summary

Frequently Asked Questions

How does streaming differ from standard function calling in LLM applications?

Standard function calling blocks execution until the tool returns a complete result, while streaming function calling yields partial outputs via Python generators. The LLM consumes each chunk as it arrives, enabling real-time responses for long-running operations. According to phases/13-tools-and-protocols/03-parallel-and-streaming-tool-calls/docs/en.md, this pattern prevents conversation freezing during database queries or sensor data retrieval.

What Python type hints should streaming tools use?

Streaming tools should use return type hints such as Iterable[str], Iterator[Chunk], or Generator[str, None, None] to indicate incremental output to the dispatcher. The reference implementation in phases/14-agent-engineering/29-production-runtimes/code/main.py uses Iterable[str] for the streaming function. These annotations signal that the function requires iteration rather than direct value access.

Can multiple tool calls stream simultaneously?

Yes, parallel streaming tool calls are supported as documented in Phase 13, Lesson 03. When the LLM requests multiple tools concurrently, each invocation runs as an independent generator, and the scheduler coordinates these streams. This allows the model to process results from multiple sources simultaneously, such as querying multiple APIs while maintaining conversation continuity.

How do you handle errors in streaming tool execution?

Error handling requires wrapping generator consumption in try-except blocks within the function calling loop. If an exception occurs mid-stream, the dispatcher should yield a structured error chunk to the LLM before terminating the iterator. The curriculum emphasizes maintaining the JSON-Schema contract by formatting error messages as valid JSON chunks that the LLM can interpret as tool failure rather than system crashes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →