How to Invoke A2A Agents with LiteLLM: A2A Protocol Implementation Guide

LiteLLM exposes A2A (agent-to-agent) calls as first-class API operations through JSON-RPC 2.0 endpoints, supporting both direct agent communication and LLM-as-agent patterns via a unified proxy interface.

LiteLLM, maintained by BerriAI, provides native support for the Agent-to-Agent (A2A) protocol, allowing developers to invoke A2A agents through a standardized JSON-RPC 2.0 interface. Whether you need to communicate directly with agent services or route requests through LiteLLM's proxy for authentication and rate-limiting, the BerriAI/litellm repository implements a complete A2A stack from API endpoints to SDK clients.

Understanding the A2A Architecture

The LiteLLM A2A implementation separates concerns across five distinct layers, each handled by specific modules in the codebase.

Public API Endpoints

The LiteLLM proxy exposes standard A2A routes under the /a2a/ path. The FastAPI router registration in litellm/proxy/agent_endpoints/a2a_endpoints.py (lines 77-91) defines the primary endpoint:

  • POST /a2a/{agent_id}/message/send – Standard message submission
  • Streaming counterpart for real-time responses

These endpoints accept JSON-RPC 2.0 envelopes and validate the required "jsonrpc": "2.0" field before processing.

Request Routing and Call Types

Incoming requests are classified via the CallTypes enum defined in litellm/types/utils.py (lines 838-842). The system maps URL patterns to specific call types:

  • CallTypes.asend_message for asynchronous operations
  • CallTypes.send_message for synchronous operations

This routing table ensures that A2A requests reach the correct handler while maintaining compatibility with LiteLLM's existing request processing pipeline.

The Proxy Handler: invoke_agent_a2a

The core request processing logic resides in the invoke_agent_a2a function within litellm/proxy/agent_endpoints/a2a_endpoints.py (lines 92-150). This handler performs several critical operations:

  1. Validation: Verifies the JSON-RPC envelope structure and required fields
  2. Authentication: Looks up the agent definition via _get_agent and validates the caller's virtual-key permissions
  3. Parameter Extraction: Pulls LiteLLM-specific parameters (including guardrails) from the request body
  4. Routing Decision: Determines whether to forward directly to the downstream agent or route through the LiteLLM completion bridge based on the presence of custom_llm_provider in litellm_params
  5. Tracing: Injects headers including X-LiteLLM-Trace-Id and X-LiteLLM-Agent-Id for observability

A2A SDK Entry Points

For programmatic access, the asend_message function in litellm/a2a_protocol/main.py (lines 8-30) serves as the primary async entry point. The implementation logic branches based on configuration:

  • Direct Flow: When no custom_llm_provider is specified, creates or reuses an A2AClient and calls _execute_a2a_send_with_retry to communicate with the downstream agent
  • Bridge Flow: When custom_llm_provider is present, delegates to _send_message_via_completion_bridge to convert the A2A payload into an LLM completion request

The Completion Bridge

The completion bridge, documented in litellm/a2a_protocol/litellm_completion_bridge/README.md, enables treating any LiteLLM-supported LLM as an A2A agent. When activated via litellm_params, the bridge:

  1. Translates the A2A JSON-RPC payload into an OpenAI-style chat request
  2. Invokes litellm.acompletion with the specified provider and model
  3. Converts the LLM response back into the A2A response format

This allows legacy systems to interact with modern LLMs through the standardized A2A protocol without modifying the client code.

Three Methods to Invoke A2A Agents

LiteLLM supports multiple invocation patterns depending on your infrastructure requirements and authentication needs.

Method 1: Direct A2A SDK Usage

For scenarios requiring direct agent communication without proxy overhead, use the native A2A SDK integration. This approach bypasses LiteLLM's authentication layer and connects straight to the agent service.

from litellm.a2a_protocol import asend_message, create_a2a_client
from a2a.types import SendMessageRequest, MessageSendParams
from uuid import uuid4

# Initialize a raw A2A client that talks directly to the agent service

a2a_client = await create_a2a_client(base_url="http://localhost:2024")

# Build the A2A request payload

request = SendMessageRequest(
    id=str(uuid4()),
    params=MessageSendParams(
        message={
            "role": "user",
            "parts": [{"kind": "text", "text": "Hello from LiteLLM!"}],
            "messageId": uuid4().hex,
        }
    ),
)

# Send the message

response = await asend_message(a2a_client=a2a_client, request=request)

print(response)            # LiteLLM-wrapped A2A response

print(response.message)    # Extracted assistant reply

Key characteristics: No litellm_params required; communicates directly via the agent's /message/send endpoint; suitable for internal networks where authentication is handled elsewhere.

Method 2: Via LiteLLM Proxy with Authentication

When you require authentication, rate-limiting, and spend logging, route requests through the LiteLLM proxy. The proxy validates virtual keys before forwarding to downstream agents.

import httpx
from a2a.client import A2AClient, A2ACardResolver
from a2a.types import SendMessageRequest, MessageSendParams
from uuid import uuid4

# Proxy base URL and virtual key from LiteLLM dashboard

proxy_base = "http://localhost:4000/a2a/my-langgraph-agent"
headers = {"Authorization": "Bearer sk-1234"}

async with httpx.AsyncClient(headers=headers) as httpx_client:
    # Resolve the agent card to get downstream URL and metadata

    resolver = A2ACardResolver(httpx_client=httpx_client, base_url=proxy_base)
    agent_card = await resolver.get_agent_card()

    # Build client pointing at the proxy

    client = A2AClient(httpx_client=httpx_client, agent_card=agent_card)

    # Build request

    request = SendMessageRequest(
        id=str(uuid4()),
        params=MessageSendParams(
            message={
                "role": "user",
                "parts": [{"kind": "text", "text": "Hello through the proxy!"}],
                "messageId": uuid4().hex,
            }
        ),
    )

    # Proxy handles authentication before forwarding to the real agent

    response = await client.send_message(request)
    print(response.message)

Key characteristics: Uses POST /a2a/{agent_id}/message/send endpoint; processes virtual-key authentication via the Authorization header; records token usage for SpendLogs; supports agent card resolution through the proxy.

Method 3: Using the Completion Bridge

To invoke any LiteLLM-supported LLM as if it were an A2A agent, use the completion bridge. This method requires no separate A2A service running.

from litellm.a2a_protocol import asend_message
from a2a.types import SendMessageRequest, MessageSendParams
from uuid import uuid4

request = SendMessageRequest(
    id=str(uuid4()),
    params=MessageSendParams(
        message={
            "role": "user", 
            "parts": [{"kind": "text", "text": "Summarize this"}], 
            "messageId": uuid4().hex
        }
    ),
)

# Route through the bridge to call a LangGraph agent

response = await asend_message(
    request=request,
    api_base="http://localhost:2024",
    litellm_params={"custom_llm_provider": "langgraph", "model": "agent"},
)

print(response.message)   # Bridge returns the LLM's answer in A2A format

Key characteristics: Sets custom_llm_provider in litellm_params to trigger the bridge; internally invokes litellm.acompletion; converts between A2A JSON-RPC and OpenAI chat formats automatically.

Key Implementation Files

The A2A functionality spans several critical files in the BerriAI/litellm repository:

Summary

  • LiteLLM implements the A2A protocol as first-class JSON-RPC 2.0 endpoints at POST /a2a/{agent_id}/message/send, with routing logic defined in litellm/proxy/agent_endpoints/a2a_endpoints.py.

  • The system supports three invocation patterns: direct SDK usage for unauthenticated internal communication, proxy-mediated calls with virtual-key authentication, and the completion bridge for LLM-as-agent scenarios.

  • Request classification occurs through CallTypes in litellm/types/utils.py, while the asend_message function in litellm/a2a_protocol/main.py orchestrates the actual message transmission.

  • When custom_llm_provider is specified in litellm_params, LiteLLM routes the request through its completion bridge, translating A2A payloads into standard LLM completion calls via litellm.acompletion.

  • All proxy requests include tracing headers (X-LiteLLM-Trace-Id, X-LiteLLM-Agent-Id) and support spend logging for cost tracking.

Frequently Asked Questions

What is the A2A protocol in LiteLLM?

The A2A (Agent-to-Agent) protocol in LiteLLM is a JSON-RPC 2.0-based communication standard that enables standardized message exchange between AI agents. Implemented in the BerriAI/litellm repository, it allows developers to invoke agents through uniform endpoints while supporting both direct agent communication and LLM-backed agent simulation through the completion bridge.

How does authentication work for A2A agents in LiteLLM?

Authentication occurs at the proxy layer through virtual keys. When invoking agents via the LiteLLM proxy (Method 2), clients must include an Authorization: Bearer <virtual_key> header. The invoke_agent_a2a handler validates these permissions against the agent definition before forwarding the request to the downstream service, enabling centralized access control and rate limiting.

What is the difference between direct A2A calls and the completion bridge?

Direct A2A calls transmit JSON-RPC messages straight to a running agent service using the A2AClient, requiring no litellm_params configuration. The completion bridge intercepts A2A requests when custom_llm_provider is set in litellm_params, converting them into OpenAI-style chat completions via litellm.acompletion before translating the response back to A2A format. The bridge allows any LiteLLM-supported LLM to behave as an A2A agent without implementing the protocol natively.

How do I enable tracing for A2A agent invocations?

LiteLLM automatically injects tracing headers into A2A requests. When using the SDK or proxy, the system adds X-LiteLLM-Trace-Id and X-LiteLLM-Agent-Id headers to track request flows. These identifiers appear in the _hidden_params of the LiteLLMSendMessageResponse object and are recorded in SpendLogs for comprehensive observability across agent interactions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →