How AI Agents Can Expand Their Observation and Action Spaces: Multi-Modal Architecture Guide

AI agents expand their observation and action spaces by integrating multi-modal sensors (voice, vision, robotics) and asynchronous tool frameworks, following the core formula Agent = LLM + Context + Tools.

The ai-agent-book repository by bojieli demonstrates architectural patterns for transforming text-only language models into perceptually aware, physically capable systems. In Chapter 6 – Interaction: Expanding the Observation and Action Spaces, the codebase provides concrete implementations showing how to broaden both what an agent can perceive and what it can execute in real-world environments.

Multi-Modal Observation Patterns

Expanding an agent's observation space requires capturing signals beyond text prompts. The repository organizes four primary modalities into a sensor layer that feeds processed context to the LLM core.

Voice: Streaming and End-to-End Speech

Voice integration enables continuous acoustic input and output, allowing agents to listen and speak in real time. The repository provides two complementary implementations:

  • Streaming ASR (chapter6/streaming-speech/qwen2_streaming.py) chunks raw audio and produces partial transcripts for low-latency response using models like Qwen2-7B-Chat. This approach streams results incrementally rather than waiting for utterance completion.

  • End-to-End Speech (chapter6/end-to-end-speech/speech_model.py) employs MiniCPM-o 4.5 to map raw audio directly to text and back to audio without intermediate text processing, reducing pipeline complexity.

Computer Use: GUI and Browser Control

Computer use capabilities allow agents to perceive visual interfaces and execute click, scroll, and typing actions. The repository implements this through vision-language models:

Robotics: Physical World Interaction

Robotics integration extends observations to physical telemetry and actions to mechanical effectors:

Multimodal Fusion: Phone Agents

The Phone Agent project (chapter6/phone-agent/agent.py) combines WebRTC audio streams, Whisper speech recognition, and language models to create voice-only conversational agents. This demonstrates how fusing multiple modalities (audio streaming, speech-to-text, LLM inference) creates robust input channels for telephony applications.

Expanding Action Spaces with Tools

To broaden the action space, agents require asynchronous execution frameworks that handle long-running operations, interruptions, and parallel tool calls.

Event-Driven Asynchronous Architecture

The repository implements an event-driven asyncio runtime in chapter6/async-agent/runtime.py. This single-threaded loop processes an inbox of prioritized events, enabling:

  • Interruption handling for user corrections mid-execution
  • Cancellation of obsolete tool calls when context changes
  • Parallel execution of independent tools (e.g., browsing while transcribing audio)

The AgentRuntime class manages the lifecycle of tools registered via runtime.register_tool(), queuing function calls as events that the main loop dispatches asynchronously.

MCP Integration for Tool Ecosystems

The Modular Compute Platform (MCP) pattern, demonstrated in chapter6/agent-with-event-trigger/server_fastapi.py, wraps existing tool servers (code interpreters, web browsers) behind a FastAPI gateway. This architecture automatically exposes MCP tools as observable events and callable actions, allowing agents to discover and invoke capabilities dynamically without hardcoded integrations.

The Perception-Action Loop Architecture

The repository structures expanded agents into four distinct layers that form a continuous loop:

  1. Sensor Layer – Captures raw signals (microphone PCM, screenshots, robot telemetry) and preprocesses them through encoders (ASR, vision models).

  2. Context Enrichment – Processesed observations append to the LLM's conversation history as structured messages (e.g., {"role": "user", "content": "<image_data>"}).

  3. Decision Layer – The LLM selects tools via function calling, generating structured JSON actions targeting specific effectors (browser, speech synthesizer, robot controller).

  4. Actuator Layer – Executes the chosen tool (synchronous or async), emitting completion events that feed back into the sensor layer, completing the cycle.

This loop is concretely implemented in chapter6/async-agent/events.py, which defines tool contracts and event schemas that bridge perception and action.

Practical Implementation Examples

Example 1: Streaming Speech Recognition

The following snippet demonstrates low-latency voice input using the streaming ASR implementation:

from streaming_speech.qwen2_streaming import StreamingASR

asr = StreamingASR(model_name="Qwen2-7B-Chat")
audio_stream = get_microphone_stream()  # yields raw PCM chunks

for partial in asr.transcribe(audio_stream):
    # partial is a progressively improving transcript

    print("Partial transcript:", partial)
    # Feed partial results to the LLM on-the-fly for immediate response

Source: chapter6/streaming-speech/qwen2_streaming.py

Example 2: Async Event-Driven Agent with Browser Tool

This example shows how to register and invoke tools within the asynchronous runtime:

import asyncio
from async_agent.runtime import AgentRuntime
from async_agent.events import BrowserTool

async def main():
    runtime = AgentRuntime()
    # Register browsing capability

    runtime.register_tool("browser_open", BrowserTool.open_url)
    
    # Start the event processing loop

    asyncio.create_task(runtime.run())
    
    # Enqueue a complex request that triggers tool use

    await runtime.enqueue_user_message(
        "Search the latest news about AI agents and summarize the top result."
    )
    
    # Allow time for async tool execution

    await asyncio.sleep(5)

asyncio.run(main())

Sources: chapter6/async-agent/runtime.py (core loop) and chapter6/async-agent/events.py (tool definitions)

Summary

  • Agent = LLM + Context + Tools is the foundational formula for expanding capabilities beyond text processing.
  • Multi-modal observation requires dedicated sensor layers for voice (qwen2_streaming.py), vision (main.py), and robotics (teleop.py).
  • Action expansion relies on asynchronous event loops (runtime.py) and modular tool integration (server_fastapi.py) to handle concurrent, interruptible operations.
  • The perception-action loop continuously enriches LLM context with preprocessed sensor data, enabling closed-loop control of browsers, speech systems, and physical robots.

Frequently Asked Questions

How do streaming ASR models reduce latency in voice agents?

Streaming ASR models like those in chapter6/streaming-speech/qwen2_streaming.py process audio chunks incrementally rather than waiting for complete utterances. This allows agents to begin reasoning from partial transcripts immediately, reducing perceived response latency from seconds to milliseconds while the model continues refining its transcription.

What is the difference between computer-use-open-model and native Claude computer use?

The computer-use-open-model implementation (chapter6/computer-use-open-model/main.py) uses open vision-language models (Qwen-VL) to parse screenshots and calculate click coordinates locally. In contrast, the native Claude example (chapter6/claude-computer-use-native/run_weather_task.py) relies on proprietary API tool-calling conventions. Both achieve GUI automation but differ in model hosting and action schema specifications.

How does the event-driven architecture handle physical robot control?

The architecture treats robot commands as asynchronous events in chapter6/async-agent/runtime.py. High-level commands (e.g., "move arm to position") enqueue as events that the runtime dispatches to hardware interfaces in chapter6/xlerobot-teleoperation/teleop.py. State changes from robot sensors feed back as new observation events, creating a closed control loop that respects mechanical latency without blocking the main agent thread.

Can agents integrate existing MCP tool servers without modification?

Yes. The agent-with-event-trigger implementation (chapter6/agent-with-event-trigger/server_fastapi.py) wraps existing MCP servers behind a FastAPI gateway that translates between MCP protocols and the agent's internal event schema. This allows agents to consume external tools (code interpreters, databases) as observable, callable actions without requiring changes to the underlying tool implementation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →