How AI Agents Can Expand Their Observation and Action Spaces: Multi-Modal Architecture Guide
AI agents expand their observation and action spaces by integrating multi-modal sensors (voice, vision, robotics) and asynchronous tool frameworks, following the core formula Agent = LLM + Context + Tools.
The ai-agent-book repository by bojieli demonstrates architectural patterns for transforming text-only language models into perceptually aware, physically capable systems. In Chapter 6 – Interaction: Expanding the Observation and Action Spaces, the codebase provides concrete implementations showing how to broaden both what an agent can perceive and what it can execute in real-world environments.
Multi-Modal Observation Patterns
Expanding an agent's observation space requires capturing signals beyond text prompts. The repository organizes four primary modalities into a sensor layer that feeds processed context to the LLM core.
Voice: Streaming and End-to-End Speech
Voice integration enables continuous acoustic input and output, allowing agents to listen and speak in real time. The repository provides two complementary implementations:
-
Streaming ASR (
chapter6/streaming-speech/qwen2_streaming.py) chunks raw audio and produces partial transcripts for low-latency response using models like Qwen2-7B-Chat. This approach streams results incrementally rather than waiting for utterance completion. -
End-to-End Speech (
chapter6/end-to-end-speech/speech_model.py) employs MiniCPM-o 4.5 to map raw audio directly to text and back to audio without intermediate text processing, reducing pipeline complexity.
Computer Use: GUI and Browser Control
Computer use capabilities allow agents to perceive visual interfaces and execute click, scroll, and typing actions. The repository implements this through vision-language models:
-
The open-model implementation in
chapter6/computer-use-open-model/main.pyuses Qwen-VL to parse screenshots and issue discrete GUI actions (click coordinates, keypress sequences). -
The native Claude example in
chapter6/claude-computer-use-native/run_weather_task.pydemonstrates OpenAI-style tool calls for web browsing, enabling automated research and form interaction.
Robotics: Physical World Interaction
Robotics integration extends observations to physical telemetry and actions to mechanical effectors:
-
Teleoperation (
chapter6/xlerobot-teleoperation/teleop.py) provides a control loop that sends high-level commands to real robot arms and validates state changes through sensor feedback. -
Navigation planning (
chapter6/gemini-xlerobot-navigation/desktop_planner.py) contains contracts and planners for mobile robot navigation, allowing agents to perceive spatial layouts and plan movement paths.
Multimodal Fusion: Phone Agents
The Phone Agent project (chapter6/phone-agent/agent.py) combines WebRTC audio streams, Whisper speech recognition, and language models to create voice-only conversational agents. This demonstrates how fusing multiple modalities (audio streaming, speech-to-text, LLM inference) creates robust input channels for telephony applications.
Expanding Action Spaces with Tools
To broaden the action space, agents require asynchronous execution frameworks that handle long-running operations, interruptions, and parallel tool calls.
Event-Driven Asynchronous Architecture
The repository implements an event-driven asyncio runtime in chapter6/async-agent/runtime.py. This single-threaded loop processes an inbox of prioritized events, enabling:
- Interruption handling for user corrections mid-execution
- Cancellation of obsolete tool calls when context changes
- Parallel execution of independent tools (e.g., browsing while transcribing audio)
The AgentRuntime class manages the lifecycle of tools registered via runtime.register_tool(), queuing function calls as events that the main loop dispatches asynchronously.
MCP Integration for Tool Ecosystems
The Modular Compute Platform (MCP) pattern, demonstrated in chapter6/agent-with-event-trigger/server_fastapi.py, wraps existing tool servers (code interpreters, web browsers) behind a FastAPI gateway. This architecture automatically exposes MCP tools as observable events and callable actions, allowing agents to discover and invoke capabilities dynamically without hardcoded integrations.
The Perception-Action Loop Architecture
The repository structures expanded agents into four distinct layers that form a continuous loop:
-
Sensor Layer – Captures raw signals (microphone PCM, screenshots, robot telemetry) and preprocesses them through encoders (ASR, vision models).
-
Context Enrichment – Processesed observations append to the LLM's conversation history as structured messages (e.g.,
{"role": "user", "content": "<image_data>"}). -
Decision Layer – The LLM selects tools via function calling, generating structured JSON actions targeting specific effectors (browser, speech synthesizer, robot controller).
-
Actuator Layer – Executes the chosen tool (synchronous or async), emitting completion events that feed back into the sensor layer, completing the cycle.
This loop is concretely implemented in chapter6/async-agent/events.py, which defines tool contracts and event schemas that bridge perception and action.
Practical Implementation Examples
Example 1: Streaming Speech Recognition
The following snippet demonstrates low-latency voice input using the streaming ASR implementation:
from streaming_speech.qwen2_streaming import StreamingASR
asr = StreamingASR(model_name="Qwen2-7B-Chat")
audio_stream = get_microphone_stream() # yields raw PCM chunks
for partial in asr.transcribe(audio_stream):
# partial is a progressively improving transcript
print("Partial transcript:", partial)
# Feed partial results to the LLM on-the-fly for immediate response
Source: chapter6/streaming-speech/qwen2_streaming.py
Example 2: Async Event-Driven Agent with Browser Tool
This example shows how to register and invoke tools within the asynchronous runtime:
import asyncio
from async_agent.runtime import AgentRuntime
from async_agent.events import BrowserTool
async def main():
runtime = AgentRuntime()
# Register browsing capability
runtime.register_tool("browser_open", BrowserTool.open_url)
# Start the event processing loop
asyncio.create_task(runtime.run())
# Enqueue a complex request that triggers tool use
await runtime.enqueue_user_message(
"Search the latest news about AI agents and summarize the top result."
)
# Allow time for async tool execution
await asyncio.sleep(5)
asyncio.run(main())
Sources: chapter6/async-agent/runtime.py (core loop) and chapter6/async-agent/events.py (tool definitions)
Summary
- Agent = LLM + Context + Tools is the foundational formula for expanding capabilities beyond text processing.
- Multi-modal observation requires dedicated sensor layers for voice (
qwen2_streaming.py), vision (main.py), and robotics (teleop.py). - Action expansion relies on asynchronous event loops (
runtime.py) and modular tool integration (server_fastapi.py) to handle concurrent, interruptible operations. - The perception-action loop continuously enriches LLM context with preprocessed sensor data, enabling closed-loop control of browsers, speech systems, and physical robots.
Frequently Asked Questions
How do streaming ASR models reduce latency in voice agents?
Streaming ASR models like those in chapter6/streaming-speech/qwen2_streaming.py process audio chunks incrementally rather than waiting for complete utterances. This allows agents to begin reasoning from partial transcripts immediately, reducing perceived response latency from seconds to milliseconds while the model continues refining its transcription.
What is the difference between computer-use-open-model and native Claude computer use?
The computer-use-open-model implementation (chapter6/computer-use-open-model/main.py) uses open vision-language models (Qwen-VL) to parse screenshots and calculate click coordinates locally. In contrast, the native Claude example (chapter6/claude-computer-use-native/run_weather_task.py) relies on proprietary API tool-calling conventions. Both achieve GUI automation but differ in model hosting and action schema specifications.
How does the event-driven architecture handle physical robot control?
The architecture treats robot commands as asynchronous events in chapter6/async-agent/runtime.py. High-level commands (e.g., "move arm to position") enqueue as events that the runtime dispatches to hardware interfaces in chapter6/xlerobot-teleoperation/teleop.py. State changes from robot sensors feed back as new observation events, creating a closed control loop that respects mechanical latency without blocking the main agent thread.
Can agents integrate existing MCP tool servers without modification?
Yes. The agent-with-event-trigger implementation (chapter6/agent-with-event-trigger/server_fastapi.py) wraps existing MCP servers behind a FastAPI gateway that translates between MCP protocols and the agent's internal event schema. This allows agents to consume external tools (code interpreters, databases) as observable, callable actions without requiring changes to the underlying tool implementation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →