How Tambo AI's Backend Service Orchestrates LLM Interactions: A Deep Dive
Tambo AI's backend service orchestrates LLM interactions through a provider-agnostic pipeline that converts thread messages into streaming AI decisions using the AI SDK, with the AISdkClient handling provider-specific implementations and the runDecisionLoop function managing the orchestration flow.
The tambo-ai/tambo repository implements a sophisticated backend architecture in packages/backend that abstracts LLM provider complexity while enabling real-time streaming of AI-generated decisions, suggestions, and component updates. This article examines how the backend factory, LLM client interface, and decision-loop service collaborate to transform user messages into actionable UI components.
Backend Architecture Overview
The orchestration pipeline follows a five-stage flow that separates transport concerns from business logic:
- Backend factory (
createTamboBackend) instantiates an LLM client (AISdkClient) configured for specific providers (OpenAI, Anthropic, Mistral, Google, Groq, Cerebras). - Decision-loop service formats conversations, injects system prompts, and calls
LLMClient.complete()withstream: true. - AI-SDK client converts requests into AI SDK format, applies token limits, merges custom parameters, and invokes
streamTextorgenerateText. - Streaming deltas transform into
LLMStreamItemobjects containing legacy LLM responses and AG-UI events for the V1 UI streaming API. - Decision-loop service consumes stream items, parses tool calls, extracts UI events, and yields
DecisionStreamItemobjects for platform consumption.
The LLM Client Interface
All LLM interactions flow through a single abstract interface defined in packages/backend/src/services/llm/llm-client.ts. This contract enables provider swapping without modifying downstream services:
// packages/backend/src/services/llm/llm-client.ts
export interface LLMClient {
chainId: string;
complete(params: StreamingCompleteParams): Promise<AsyncIterableIterator<LLMStreamItem>>;
complete(params: CompleteParams): Promise<LLMResponse>;
}
The LLMClient supports both streaming and single-shot completions, returning either an async iterator of LLMStreamItem objects or a complete LLMResponse. Services like the decision-loop, suggestion generation, and thread-name generation depend only on this interface, allowing mock implementations for testing.
Source: llm-client.ts – lines 58‑64
Provider-Agnostic Implementation with AISdkClient
The AISdkClient class in packages/backend/src/services/llm/ai-sdk-client.ts implements the LLMClient interface using the AI SDK to normalize interactions across OpenAI, Anthropic, and other providers.
Provider Factory Selection
The client dynamically selects the correct provider factory using getProviderFromModel():
function getProviderFromModel(model: string, provider: Provider): keyof typeof PROVIDER_FACTORIES {
if (provider === "openai-compatible") return "openai-compatible";
switch (provider) {
case "openai": return "openai";
case "anthropic": return "anthropic";
case "mistral": return "mistral";
case "groq": return "groq";
case "gemini": return "google";
case "cerebras": return "openai-compatible";
default: return "openai";
}
}
This mapping supports OpenAI, Anthropic, Mistral, Google (Gemini), Groq, and OpenAI-compatible endpoints like Cerebras.
Source: ai-sdk-client.ts – lines 98‑124
Request Configuration and Token Management
The complete() method constructs the AI SDK configuration by merging parameters in a specific hierarchy:
- Provider defaults
- Model defaults
- User-provided custom parameters
- Model-specific provider parameters
- OpenAI-compatible extra keys
const baseConfig: AICompleteParams = {
model: modelInstance,
messages: modelMessages,
tools,
toolChoice: params.tool_choice ? this.convertToolChoice(params.tool_choice) : undefined,
...(responseFormat && { responseFormat }),
providerOptions: {
[providerKey]: {
...providerCfg?.providerSpecificParams,
...modelCfg?.modelParamsDefaults,
...modelSpecificProviderParams,
...(providerKey === "openai-compatible" && providerSpecificCustomParams),
},
},
...modelDefaults,
...filteredCustomParams,
};
Token limiting is applied via limitTokens based on model-specific input limits before sending the request.
Source: ai-sdk-client.ts – lines 165‑215
Streaming vs Non-Streaming Execution
The client branches based on the stream parameter:
- Streaming: Calls
streamText()from the AI SDK and processes the result throughhandleStreamingResponse() - Non-streaming: Calls
generateText()and converts the result viaconvertToLLMResponse()
if (params.stream) {
const result = await streamText({ ...baseConfig, abortSignal: params.abortSignal });
return this.handleStreamingResponse(result);
} else {
const result = await generateText(baseConfig);
return this.convertToLLMResponse(result);
}
Source: ai-sdk-client.ts – lines 222‑235
Streaming Response Processing
The handleStreamingResponse() method in ai-sdk-client.ts transforms AI SDK streams into Tambo's internal format.
Converting AI SDK Streams to LLMStreamItem
The method consumes result.fullStream and yields LLMStreamItem objects containing:
llmResponse: A partialLLMResponsemaintaining backward compatibility with OpenAI-style response shapesaguiEvents: An array ofBaseEventobjects for the V1 UI streaming API (text messages, tool calls, reasoning, and custom component events)
The loop detects delta types (text-start, tool-input-start, reasoning-start) and generates message IDs for each stream using generateMessageId().
AG-UI Event Generation
During streaming, the client builds AG-UI events for:
- Text messages:
TEXT_MESSAGE_START,TEXT_MESSAGE_CONTENT,TEXT_MESSAGE_END - Tool calls:
TOOL_CALL_START,TOOL_CALL_ARGS,TOOL_CALL_END - Reasoning:
THINKING_START,THINKING_CONTENT,THINKING_END - Component streaming: Custom events for UI component rendering
Partial tool-call arguments are accumulated to support JSON parsing while streaming, ensuring the UI can display partial results immediately.
for await (const delta of result.fullStream) {
// …switch on delta.type, build aguiEvents, accumulate message / tool args…
yield {
llmResponse: {
message: {
content: accumulatedMessage,
role: "assistant",
tool_calls: toolCallRequest ? [toolCallRequest] : undefined,
refusal: null,
},
reasoning: accumulatedReasoning,
reasoningDurationMS: reasoningStartTimestamp && reasoningEndTimestamp
? reasoningEndTimestamp - reasoningStartTimestamp
: undefined,
index: 0,
logprobs: null,
},
aguiEvents,
};
}
Source: ai-sdk-client.ts – lines 456‑525
The Decision Loop Orchestration
The runDecisionLoop() function in packages/backend/src/services/decision-loop/decision-loop-service.ts serves as the core orchestrator that bridges raw LLM outputs with actionable UI decisions.
Thread Preparation and Tool Injection
Before calling the LLM, the service:
- Filters UI tools using
isUiToolName()to identify component-generating functions - Adds standard parameters to every tool definition for consistent handling
- Builds the system prompt via
generateDecisionLoopPrompt()incorporating custom instructions - Prefetches MCP resources using
prefetchAndCacheResources()to reduce latency
Stream Consumption and Decision Assembly
The function calls llmClient.complete() with stream: true and iterates over each LLMStreamItem:
- Extracts message text and tool calls from the
llmResponse - Parses tool arguments and builds
ToolCallRequestobjects - Extracts UI component IDs from custom AG-UI events for component tracking
- Constructs
LegacyComponentDecisionobjects maintaining backward compatibility - Yields
DecisionStreamItemcontaining both the accumulated decision and rawaguiEvents
export async function* runDecisionLoop(
llmClient: LLMClient,
messages: ThreadMessage[],
strictTools: OpenAI.Chat.Completions.ChatCompletionTool[],
customInstructions: string | undefined,
forceToolChoice: string | undefined,
resourceFetchers: ResourceFetcherMap,
abortSignal?: AbortSignal,
): AsyncIterableIterator<DecisionStreamItem> {
// …prepare tools, system prompt, cached resources…
const responseStream = await llmClient.complete({
messages: promptMessages,
tools: toolsWithStandardParameters,
promptTemplateName: "decision-loop",
promptTemplateParams: systemPromptArgs,
stream: true,
tool_choice: convertToolChoice(forceToolChoice),
abortSignal,
});
for await (const streamItem of responseStream) {
const llmResponse = streamItem.llmResponse;
const message = getLLMResponseMessage(llmResponse);
const toolCall = llmResponse.message?.tool_calls?.[0];
// …parse tool args, build ToolCallRequest, extract UI component ID…
const parsedChunk: Partial<LegacyComponentDecision> = {
role: MessageRole.Assistant,
message: displayMessage,
componentName: isUITool ? toolCall.function.name.slice(UI_TOOLNAME_PREFIX.length) : "",
componentId: isUITool ? (componentId ?? accumulatedDecision.componentId) : undefined,
props: isUITool ? filteredToolArgs : null,
toolCallRequest: clientToolRequest,
toolCallId: toolCall ? getLLMResponseToolCallId(llmResponse) : undefined,
statusMessage,
completionStatusMessage,
reasoning: llmResponse.reasoning ?? undefined,
reasoningDurationMS: llmResponse.reasoningDurationMS ?? undefined,
};
accumulatedDecision = { ...accumulatedDecision, ...parsedChunk };
yield { decision: accumulatedDecision, aguiEvents: streamItem.aguiEvents };
}
}
Source: decision-loop-service.ts – lines 85‑118, 172‑226, 267‑284 and lines 172‑226
End-to-End Integration Example
The following example demonstrates creating a backend instance and running the decision loop to generate UI components:
import { createTamboBackend } from "./tambo-backend";
// 1️⃣ Build backend (LLM client is created inside)
const backend = await createTamboBackend(process.env.OPENAI_API_KEY, "chain-123", "user-456");
// 2️⃣ Prepare conversation and tool list
const messages = [{ role: "user", content: "Show me a chart of sales." }];
const tools = [{ type: "function", function: { name: "show_component_chart", parameters: {} } }];
// 3️⃣ Run the decision loop (streaming)
for await (const { decision, aguiEvents } of backend.runDecisionLoop({
messages,
strictTools: tools,
customInstructions: undefined,
forceToolChoice: undefined,
resourceFetchers: {}, // empty map for this example
})) {
console.log("UI decision:", decision);
// UI can render events directly:
// renderAGUIEvents(aguiEvents);
}
This pattern illustrates how Tambo AI's backend service orchestrates LLM interactions by abstracting provider details while exposing a streaming interface for real-time UI updates.
Summary
- Provider-agnostic architecture: All LLM interactions funnel through the
LLMClientinterface, allowing new providers to be added without modifying decision-loop logic. - Streaming-first design: The backend always requests streaming responses (
stream: true) and re-assembles partial responses into both legacy and AG-UI event formats. - Tool conversion layer: OpenAI-style tool definitions map to AI SDK
ToolSetobjects while preserving function names, descriptions, and JSON schemas. - Hierarchical configuration: Provider defaults cascade through model defaults, user-provided custom parameters, and model-specific provider params, ensuring flexibility with sensible defaults.
- Event-driven UI rendering: By emitting
BaseEvent[]arrays alongside each LLM chunk, the frontend renders text, tool calls, and reasoning incrementally for real-time visualizations.
Frequently Asked Questions
What is the role of the LLMClient interface in Tambo AI's backend?
The LLMClient interface defines the contract for all LLM interactions within Tambo AI's backend service. Located in packages/backend/src/services/llm/llm-client.ts (lines 58‑64), it specifies a complete() method that supports both streaming (AsyncIterableIterator<LLMStreamItem>) and non-streaming (LLMResponse) execution modes. This abstraction allows services like the decision-loop, suggestion generation, and thread-name generation to depend solely on the interface rather than concrete implementations, enabling easy testing with mock clients and seamless provider swaps.
How does Tambo AI handle different LLM providers like OpenAI and Anthropic?
Tambo AI handles multiple providers through the AISdkClient class in packages/backend/src/services/llm/ai-sdk-client.ts, which implements the LLMClient interface using the AI SDK. The getProviderFromModel() function (lines 98‑124) maps provider strings to specific factory functions—such as createOpenAI, createAnthropic, createMistral, createGoogle, and createGroq—while treating Cerebras and custom endpoints as openai-compatible. This architecture ensures that provider-specific authentication, parameter handling, and model instantiation are isolated within the client implementation, while the rest of the backend operates on normalized LLMStreamItem objects.
What is the difference between LLMStreamItem and DecisionStreamItem?
LLMStreamItem and DecisionStreamItem represent different abstraction layers in the orchestration pipeline. The LLMStreamItem (generated by AISdkClient in ai-sdk-client.ts, lines 456‑525) contains raw LLM output formatted as both a legacy LLMResponse object (mirroring OpenAI's structure) and an array of aguiEvents (AG-UI protocol events for UI rendering).
The DecisionStreamItem (produced by runDecisionLoop() in decision-loop-service.ts, lines 267‑284) represents a higher-level abstraction that extracts UI-specific information from the raw LLM output. It contains a LegacyComponentDecision object with parsed component names, props, tool call requests, and reasoning metadata—effectively transforming raw LLM tokens into structured UI decisions that the frontend can render immediately.
How does the decision loop handle streaming tool calls?
The decision loop handles streaming tool calls through incremental parsing within the runDecisionLoop() generator function (located in packages/backend/src/services/decision-loop/decision-loop-service.ts, lines 172‑226). As the function iterates over LLMStreamItem objects from the LLM client, it extracts tool call information from llmResponse.message.tool_calls and accumulates partial arguments across stream chunks.
For UI-specific tools (identified by isUiTool), the function slices the tool name prefix to extract the component name, filters the tool arguments to create component props, and extracts component IDs from custom AG-UI events. This incremental assembly allows the backend to yield partial DecisionStreamItem objects containing incomplete tool calls that get refined as more stream data arrives, enabling the frontend to display "thinking" states and partial component renders before the LLM finishes generating the complete tool call.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →