How to Use Skills for Real-Time AI Applications: The Complete Streaming Guide
Real-time AI applications rely on streaming skill endpoints that maintain persistent HTTP/2 or WebSocket connections, enabling bidirectional text, audio, and video processing with minimal latency.
The VoltAgent/awesome-agent-skills repository provides specialized streaming skills designed specifically for real-time AI applications such as live chat and bidirectional voice synthesis. These skills abstract provider-specific protocols into a unified async-iterator interface that works across JavaScript, .NET, and Java environments. According to the README.md in the official repository, developers can implement real-time capabilities using official skills for Google Gemini Live and Azure AI Voice Live.
Architecture of Real-Time Streaming Skills
Real-time skills in the Awesome-Agent-Skills collection follow a five-layer architecture that handles persistent connections automatically.
Streaming Transport Layer
All real-time skills expose HTTP/2 or WebSocket endpoints that continuously transmit JSON-encoded events such as input, output, and error. This persistent connection eliminates the overhead of repeated HTTP requests, enabling sub-second latency required for conversational AI.
MCP Server as Connection Proxy
The MCP (Multi-Channel Processor) server acts as a thin proxy that converts standard agent invoke calls into streaming requests. According to the source code analysis, this server handles critical infrastructure concerns including automatic reconnection, back-pressure management, and authentication token refresh.
Skill-Specific Protocol Adapters
Each skill bundles an adapter that maps provider-specific protocols to a generic event schema. For example, the Gemini Live skill translates Google's GenerateContentStream RPC, while Azure Voice Live handles the RealtimeSpeech socket protocol. These adapters ensure your agent code remains provider-agnostic.
The Agent Event Loop
The agent implements a standard event loop pattern: while (msg = await skill.next()) { … }. When the skill yields partial results, the agent processes them immediately, enabling features like typing indicators or audio streaming without waiting for complete responses.
State and Memory Integration
Because agent memory lives in the VoltAgent core, partial results can be stored during streaming and referenced later. This enables "continue where we left off" functionality after network interruptions.
Official Real-Time Skills Available
The README.md file in the VoltAgent/awesome-agent-skills repository catalogs three official streaming skills optimized for real-time applications.
-
Gemini Live API (Google Gemini): Enables bidirectional text/audio streaming for chat-style agents. Referenced at line 136 in the repository's
README.md. -
Azure AI Voice Live (.NET) (Microsoft Azure): Provides real-time, low-latency speech-to-speech capabilities for .NET applications. Documented at line 222.
-
Azure AI Voice Live (Java) (Microsoft Azure): Offers identical voice capabilities for Java agents. Found at line 555 of the
README.md.
Implementation Examples
The following examples demonstrate how to integrate these streaming skills into VoltAgent-based applications. All examples assume you have installed the VoltAgent SDK (npm i @voltagent/core) and configured the appropriate skill packages.
Streaming Text with Gemini Live API (Node.js)
This implementation uses the @voltagent/skills/google-gemini package to create a real-time text chat agent:
import { VoltAgent } from '@voltagent/core';
import { createSkill } from '@voltagent/skills/google-gemini';
// 1️⃣ Initialise the Gemini Live skill (API key is read from the env)
const geminiLive = createSkill('gemini-live-api-dev');
// 2️⃣ Create an agent that uses the skill
const agent = new VoltAgent({
name: 'RealtimeChat',
// … other config (memory, tools, etc.)
});
async function chat() {
// Start a streaming invocation
const stream = await geminiLive.invoke({
prompt: 'You are a helpful assistant. Answer step‑by‑step.',
// optional: `temperature`, `maxOutputTokens`, etc.
});
// Consume partial responses as they arrive
for await (const chunk of stream) {
console.log('Assistant says:', chunk.text); // incremental text
}
}
chat().catch(console.error);
The createSkill function initializes a streaming endpoint, while invoke returns an async iterator that yields partial chunk objects. This pattern enables immediate UI updates as text streams in.
Real-Time Voice with Azure AI Voice Live (.NET)
For speech-to-speech applications, the Azure AI Voice Live skill provides bidirectional audio streaming:
using VoltAgent;
using VoltAgent.Skills.Microsoft.AzureAI;
// Initialise the voice‑live skill (credentials come from environment variables)
var voiceLive = SkillFactory.Create("azure-ai-voicelive-dotnet");
// Create the agent
var agent = new VoltAgent(new VoltAgentOptions { Name = "RealtimeVoiceBot" });
async Task RunConversationAsync()
{
// Start streaming speech synthesis/recognition
await foreach (var evt in voiceLive.InvokeAsync(new VoiceLiveRequest
{
// optional: language, voice model, etc.
Prompt = "Welcome! Ask me anything."
}))
{
switch (evt.Type)
{
case VoiceLiveEventType.AudioChunk:
PlayAudio(evt.AudioData); // Play partial audio buffer
break;
case VoiceLiveEventType.Transcript:
Console.WriteLine($"User said: {evt.Text}");
break;
}
}
}
await RunConversationAsync();
This example uses InvokeAsync to yield a mixed stream of AudioChunk and Transcript events, enabling simultaneous playback of generated speech and recognition of user input.
Unified Pattern for Any Real-Time Skill
All real-time skills in the collection share a consistent async-iterator interface, allowing abstraction across providers:
async def use_realtime_skill(skill, request):
async for event in skill.invoke(request):
handle(event) # render UI, play audio, store in memory, etc.
You only need to pass the appropriate skill object—such as geminiLive or voiceLive—and a request payload. The underlying adapter handles protocol translation automatically.
Repository Structure and Key Files
The Awesome-Agent-Skills repository contains the following key files that support real-time development:
README.md: Contains the complete catalog of official real-time skills with line-specific references (e.g., Gemini Live at line 136, Azure Voice Live at lines 222 and 555) and links to external documentation.LICENSE: Defines usage rights for the skill collection.CONTRIBUTING.md: Explains how to add new real-time skills to the collection.opencode.json: Internal metadata used by the Open-Code environment (not required for end-users).
These files provide the reference documentation, while the actual skill implementations reside in the external repositories linked from the README.md.
Summary
- Real-time AI applications require streaming skill endpoints that maintain persistent HTTP/2 or WebSocket connections rather than standard request-response APIs.
- The MCP server handles connection management, authentication, and reconnection automatically, allowing developers to focus on agent logic.
- Official streaming skills include Gemini Live API and Azure AI Voice Live for both .NET and Java, each providing protocol-specific adapters that map to a unified event schema.
- The async-iterator pattern (
for await...ofin JavaScript,await foreachin C#) enables immediate processing of partial results, supporting real-time UI updates and audio streaming. - All skills follow a consistent architecture where
createSkillinitializes the connection andinvokereturns a stream of events that the agent processes in a standard event loop.
Frequently Asked Questions
What distinguishes real-time skills from standard API skills?
Standard skills use synchronous HTTP requests that return complete responses, while real-time skills maintain persistent connections through HTTP/2 or WebSocket protocols. This enables bidirectional streaming of partial data chunks, reducing latency from seconds to milliseconds for conversational applications.
How does the MCP server handle network interruptions during streaming?
The MCP server implements automatic reconnection logic and back-pressure management. According to the VoltAgent architecture, when a connection drops, the server attempts to reconnect transparently while preserving the agent's memory state, allowing conversations to resume without data loss.
Can I switch between Gemini and Azure Voice skills without rewriting my agent?
Yes. Both skills implement the same invoke method signature and return async iterators with a unified event schema. You can swap createSkill('gemini-live-api-dev') for SkillFactory.Create("azure-ai-voicelive-dotnet") while keeping your agent's event loop unchanged, as the skill-specific adapters handle all protocol translation.
Do real-time skills support state persistence during long-running streams?
Yes. Because agent memory resides in the VoltAgent core rather than the skill adapter, partial results can be stored and referenced throughout the streaming session. This enables features like "continue where we left off" after network glitches or multi-turn conversations that reference earlier parts of the stream.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →