# Implementing Streaming Tool Calls with Function Calling: A Python Guide for AI Agents

> Implement streaming tool calls with function calling in Python. Learn how LLMs process partial tool outputs in real-time using generators for efficient AI agent interactions. Read the guide now!

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-26

---

**Implementing streaming tool calls with function calling allows language models to process incremental tool outputs through Python generators, enabling real-time interactions where the LLM consumes partial results via iterators rather than blocking for complete execution.**

The *ai-engineering-from-scratch* repository provides a comprehensive curriculum for building AI systems from first principles, including production-ready patterns for streaming tool interactions. When implementing streaming tool calls with function calling, tools yield chunks of data incrementally while the LLM processes each fragment, enabling sub-second latency for voice assistants and live data pipelines. This approach is detailed across Phase 13 (Tools and Protocols) and Phase 14 (Agent Engineering), with reference implementations available in the production runtimes and voice assistant capstone modules.

## Architectural Components

The implementation relies on four distinct architectural layers that coordinate the contract, execution, and consumption of streaming tools.

### Tool Interface and Schema

The **Tool Interface** establishes a JSON-Schema contract that defines available functions and their arguments. This layer validates that the model can request specific tools with structured parameters, as covered in Phase 13, Lesson 01 (*The Tool Interface*). The schema ensures type safety even when processing incremental streams.

### Function Calling Loop

The **Function Calling Loop** parses the model's `tool_calls` JSON payload, dispatches execution to the appropriate Python function, and handles return values. As implemented in [`phases/13-tools-and-protocols/09-function-calling/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/13-tools-and-protocols/09-function-calling/docs/en.md), this loop must detect whether a tool returns a static value or a streaming iterator to handle partial results correctly.

### Streaming API

The **Streaming API** implements tools as Python generators using `yield` statements to produce partial results. Rather than returning a single string, these functions emit chunks (tokens, sensor readings, or database rows) as they become available. The curriculum covers this pattern in Phase 13, Lesson 03 (*Parallel and Streaming Tool Calls*), including provider-specific semantics for OpenAI, Gemini, and Anthropic APIs.

### Scheduler and Coordinator

In multimodal systems, a **Scheduler** arbitrates between concurrent streaming sources including ASR (Automatic Speech Recognition), LLM generation, and TTS (Text-to-Speech). The implementation in [`phases/19-capstone-projects/03-realtime-voice-assistant/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/03-realtime-voice-assistant/code/main.py) demonstrates how this coordinator aligns timestamps and manages turn-taking between multiple simultaneous streams.

## Execution Flow for Streaming Tools

When implementing streaming tool calls with function calling, the system follows a deterministic six-stage pipeline:

1. **Prompt Analysis**: The LLM receives a user prompt and determines that external data retrieval is necessary.

2. **Tool Call Generation**: The model outputs a JSON blob specifying the tool name and arguments, such as `{"name": "streaming", "arguments": {"input_text": "Hello world"}}`.

3. **Function Dispatch**: The dispatcher looks up the registered Python function in the tool registry based on the `name` field.

4. **Streaming Invocation**: The dispatcher executes the tool, which returns a generator iterator. The `streaming` function in [`phases/14-agent-engineering/29-production-runtimes/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/29-production-runtimes/code/main.py) yields individual words from the input text incrementally.

5. **Progressive Consumption**: The LLM receives each yielded chunk through the function calling loop, updating its context window and potentially generating response tokens before the tool completes.

6. **Result Finalization**: When the iterator exhausts, the accumulated result is returned to the user, or the LLM may request additional tool calls based on the streamed data.

## Code Implementation

The following examples from the *ai-engineering-from-scratch* source code demonstrate the concrete implementation of streaming generators and their dispatch.

### Defining a Streaming Tool Generator

In [`phases/14-agent-engineering/29-production-runtimes/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/29-production-runtimes/code/main.py), the `streaming` function implements the generator pattern required for incremental output:

```python

# File: phases/14-agent-engineering/29-production-runtimes/code/main.py

def streaming(input_text: str) -> Iterable[str]:
    """
    A simple streaming tool that yields the input text one word at a time.
    The function can be called by the LLM via the standard function‑calling
    loop; each yielded token is sent back to the model as a partial result.
    """
    for word in input_text.split():
        yield f"{word} "

```

This function implements the **streaming runtime shape**, one of four production shapes defined in Phase 14, Lesson 29 (*Production Runtimes*).

### The Function Calling Loop

The dispatcher code must handle iterators differently from static values. The generic pattern from the Function Calling lesson demonstrates this logic:

```python
def run_function_call(tool_name: str, args: dict) -> str:
    if tool_name == "streaming":
        # `streaming` returns an iterator; we concatenate its pieces.

        return "".join(streaming(args["input_text"]))
    # … handle other tools …

```

When the LLM produces a tool call such as:

```json
{
  "name": "streaming",
  "arguments": {"input_text": "Hello world from the AI curriculum"}
}

```

The `run_function_call` function executes the generator. In true streaming implementations, each yielded value is sent to the LLM immediately rather than concatenated, enabling **real-time, interactive experiences** where the model reacts to partial data.

## Production Runtime Shapes

The curriculum defines four distinct runtime shapes for tool execution in [`phases/14-agent-engineering/29-production-runtimes/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/29-production-runtimes/code/main.py):

- **Request-Response**: Standard blocking execution returning complete results as strings
- **Streaming**: Generator-based execution yielding `Iterable[str]` chunks incrementally
- **Queue**: Asynchronous message-based processing for background task offloading
- **Event**: Reactive callbacks triggered by external system events

When implementing streaming tool calls with function calling, use `inspect.isgeneratorfunction()` or similar checks to distinguish between immediate returns and iterable streams, ensuring the dispatcher consumes generators appropriately.

## Real-World Multimodal Example

The real-time voice assistant capstone in [`phases/19-capstone-projects/03-realtime-voice-assistant/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/03-realtime-voice-assistant/code/main.py) demonstrates advanced streaming coordination. This system manages three concurrent streams:

- **ASR**: Continuous speech recognition yielding transcript chunks
- **LLM**: Token-by-token response generation
- **TTS**: Audio segment production

The scheduler synchronizes these streams, allowing the assistant to begin speaking while the LLM is still reasoning or while external tools are streaming partial results. This architecture achieves **sub-second latency** by implementing streaming tool calls that do not block the conversation loop.

## Summary

- **Streaming tool calls** use Python generators (`yield`) to return partial results incrementally, avoiding blocking operations that freeze the LLM context.
- The **function calling loop** must detect iterator types and consume chunks progressively, as demonstrated in [`phases/13-tools-and-protocols/09-function-calling/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/13-tools-and-protocols/09-function-calling/docs/en.md).
- Four **production runtime shapes** exist (request-response, streaming, queue, event), with the streaming shape defined in [`phases/14-agent-engineering/29-production-runtimes/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/29-production-runtimes/code/main.py).
- **Provider-specific implementations** vary between OpenAI, Gemini, and Anthropic APIs, detailed in [`phases/13-tools-and-protocols/02-function-calling-deep-dive/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/13-tools-and-protocols/02-function-calling-deep-dive/docs/en.md).
- **Multimodal schedulers** coordinate concurrent streaming sources (ASR, LLM, TTS) to enable real-time voice assistants with minimal latency.

## Frequently Asked Questions

### How does streaming differ from standard function calling in LLM applications?

Standard function calling blocks execution until the tool returns a complete result, while streaming function calling yields partial outputs via Python generators. The LLM consumes each chunk as it arrives, enabling real-time responses for long-running operations. According to [`phases/13-tools-and-protocols/03-parallel-and-streaming-tool-calls/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/13-tools-and-protocols/03-parallel-and-streaming-tool-calls/docs/en.md), this pattern prevents conversation freezing during database queries or sensor data retrieval.

### What Python type hints should streaming tools use?

Streaming tools should use return type hints such as `Iterable[str]`, `Iterator[Chunk]`, or `Generator[str, None, None]` to indicate incremental output to the dispatcher. The reference implementation in [`phases/14-agent-engineering/29-production-runtimes/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/29-production-runtimes/code/main.py) uses `Iterable[str]` for the `streaming` function. These annotations signal that the function requires iteration rather than direct value access.

### Can multiple tool calls stream simultaneously?

Yes, parallel streaming tool calls are supported as documented in Phase 13, Lesson 03. When the LLM requests multiple tools concurrently, each invocation runs as an independent generator, and the scheduler coordinates these streams. This allows the model to process results from multiple sources simultaneously, such as querying multiple APIs while maintaining conversation continuity.

### How do you handle errors in streaming tool execution?

Error handling requires wrapping generator consumption in try-except blocks within the function calling loop. If an exception occurs mid-stream, the dispatcher should yield a structured error chunk to the LLM before terminating the iterator. The curriculum emphasizes maintaining the JSON-Schema contract by formatting error messages as valid JSON chunks that the LLM can interpret as tool failure rather than system crashes.