How Tambo AI's Backend Service Orchestrates LLM Interactions: A Deep Dive

Tambo AI's backend service orchestrates LLM interactions through a provider-agnostic pipeline that converts thread messages into streaming AI decisions using the AI SDK, with the AISdkClient handling provider-specific implementations and the runDecisionLoop function managing the orchestration flow.

The tambo-ai/tambo repository implements a sophisticated backend architecture in packages/backend that abstracts LLM provider complexity while enabling real-time streaming of AI-generated decisions, suggestions, and component updates. This article examines how the backend factory, LLM client interface, and decision-loop service collaborate to transform user messages into actionable UI components.

Backend Architecture Overview

The orchestration pipeline follows a five-stage flow that separates transport concerns from business logic:

  1. Backend factory (createTamboBackend) instantiates an LLM client (AISdkClient) configured for specific providers (OpenAI, Anthropic, Mistral, Google, Groq, Cerebras).
  2. Decision-loop service formats conversations, injects system prompts, and calls LLMClient.complete() with stream: true.
  3. AI-SDK client converts requests into AI SDK format, applies token limits, merges custom parameters, and invokes streamText or generateText.
  4. Streaming deltas transform into LLMStreamItem objects containing legacy LLM responses and AG-UI events for the V1 UI streaming API.
  5. Decision-loop service consumes stream items, parses tool calls, extracts UI events, and yields DecisionStreamItem objects for platform consumption.

The LLM Client Interface

All LLM interactions flow through a single abstract interface defined in packages/backend/src/services/llm/llm-client.ts. This contract enables provider swapping without modifying downstream services:

// packages/backend/src/services/llm/llm-client.ts
export interface LLMClient {
  chainId: string;
  complete(params: StreamingCompleteParams): Promise<AsyncIterableIterator<LLMStreamItem>>;
  complete(params: CompleteParams): Promise<LLMResponse>;
}

The LLMClient supports both streaming and single-shot completions, returning either an async iterator of LLMStreamItem objects or a complete LLMResponse. Services like the decision-loop, suggestion generation, and thread-name generation depend only on this interface, allowing mock implementations for testing.

Source: llm-client.ts – lines 58‑64

Provider-Agnostic Implementation with AISdkClient

The AISdkClient class in packages/backend/src/services/llm/ai-sdk-client.ts implements the LLMClient interface using the AI SDK to normalize interactions across OpenAI, Anthropic, and other providers.

Provider Factory Selection

The client dynamically selects the correct provider factory using getProviderFromModel():

function getProviderFromModel(model: string, provider: Provider): keyof typeof PROVIDER_FACTORIES {
  if (provider === "openai-compatible") return "openai-compatible";
  switch (provider) {
    case "openai": return "openai";
    case "anthropic": return "anthropic";
    case "mistral": return "mistral";
    case "groq": return "groq";
    case "gemini": return "google";
    case "cerebras": return "openai-compatible";
    default: return "openai";
  }
}

This mapping supports OpenAI, Anthropic, Mistral, Google (Gemini), Groq, and OpenAI-compatible endpoints like Cerebras.

Source: ai-sdk-client.ts – lines 98‑124

Request Configuration and Token Management

The complete() method constructs the AI SDK configuration by merging parameters in a specific hierarchy:

  1. Provider defaults
  2. Model defaults
  3. User-provided custom parameters
  4. Model-specific provider parameters
  5. OpenAI-compatible extra keys
const baseConfig: AICompleteParams = {
  model: modelInstance,
  messages: modelMessages,
  tools,
  toolChoice: params.tool_choice ? this.convertToolChoice(params.tool_choice) : undefined,
  ...(responseFormat && { responseFormat }),
  providerOptions: {
    [providerKey]: {
      ...providerCfg?.providerSpecificParams,
      ...modelCfg?.modelParamsDefaults,
      ...modelSpecificProviderParams,
      ...(providerKey === "openai-compatible" && providerSpecificCustomParams),
    },
  },
  ...modelDefaults,
  ...filteredCustomParams,
};

Token limiting is applied via limitTokens based on model-specific input limits before sending the request.

Source: ai-sdk-client.ts – lines 165‑215

Streaming vs Non-Streaming Execution

The client branches based on the stream parameter:

  • Streaming: Calls streamText() from the AI SDK and processes the result through handleStreamingResponse()
  • Non-streaming: Calls generateText() and converts the result via convertToLLMResponse()
if (params.stream) {
  const result = await streamText({ ...baseConfig, abortSignal: params.abortSignal });
  return this.handleStreamingResponse(result);
} else {
  const result = await generateText(baseConfig);
  return this.convertToLLMResponse(result);
}

Source: ai-sdk-client.ts – lines 222‑235

Streaming Response Processing

The handleStreamingResponse() method in ai-sdk-client.ts transforms AI SDK streams into Tambo's internal format.

Converting AI SDK Streams to LLMStreamItem

The method consumes result.fullStream and yields LLMStreamItem objects containing:

  • llmResponse: A partial LLMResponse maintaining backward compatibility with OpenAI-style response shapes
  • aguiEvents: An array of BaseEvent objects for the V1 UI streaming API (text messages, tool calls, reasoning, and custom component events)

The loop detects delta types (text-start, tool-input-start, reasoning-start) and generates message IDs for each stream using generateMessageId().

AG-UI Event Generation

During streaming, the client builds AG-UI events for:

  • Text messages: TEXT_MESSAGE_START, TEXT_MESSAGE_CONTENT, TEXT_MESSAGE_END
  • Tool calls: TOOL_CALL_START, TOOL_CALL_ARGS, TOOL_CALL_END
  • Reasoning: THINKING_START, THINKING_CONTENT, THINKING_END
  • Component streaming: Custom events for UI component rendering

Partial tool-call arguments are accumulated to support JSON parsing while streaming, ensuring the UI can display partial results immediately.

for await (const delta of result.fullStream) {
  // …switch on delta.type, build aguiEvents, accumulate message / tool args…
  yield {
    llmResponse: {
      message: {
        content: accumulatedMessage,
        role: "assistant",
        tool_calls: toolCallRequest ? [toolCallRequest] : undefined,
        refusal: null,
      },
      reasoning: accumulatedReasoning,
      reasoningDurationMS: reasoningStartTimestamp && reasoningEndTimestamp
        ? reasoningEndTimestamp - reasoningStartTimestamp
        : undefined,
      index: 0,
      logprobs: null,
    },
    aguiEvents,
  };
}

Source: ai-sdk-client.ts – lines 456‑525

The Decision Loop Orchestration

The runDecisionLoop() function in packages/backend/src/services/decision-loop/decision-loop-service.ts serves as the core orchestrator that bridges raw LLM outputs with actionable UI decisions.

Thread Preparation and Tool Injection

Before calling the LLM, the service:

  1. Filters UI tools using isUiToolName() to identify component-generating functions
  2. Adds standard parameters to every tool definition for consistent handling
  3. Builds the system prompt via generateDecisionLoopPrompt() incorporating custom instructions
  4. Prefetches MCP resources using prefetchAndCacheResources() to reduce latency

Stream Consumption and Decision Assembly

The function calls llmClient.complete() with stream: true and iterates over each LLMStreamItem:

  1. Extracts message text and tool calls from the llmResponse
  2. Parses tool arguments and builds ToolCallRequest objects
  3. Extracts UI component IDs from custom AG-UI events for component tracking
  4. Constructs LegacyComponentDecision objects maintaining backward compatibility
  5. Yields DecisionStreamItem containing both the accumulated decision and raw aguiEvents
export async function* runDecisionLoop(
  llmClient: LLMClient,
  messages: ThreadMessage[],
  strictTools: OpenAI.Chat.Completions.ChatCompletionTool[],
  customInstructions: string | undefined,
  forceToolChoice: string | undefined,
  resourceFetchers: ResourceFetcherMap,
  abortSignal?: AbortSignal,
): AsyncIterableIterator<DecisionStreamItem> {
  // …prepare tools, system prompt, cached resources…
  const responseStream = await llmClient.complete({
    messages: promptMessages,
    tools: toolsWithStandardParameters,
    promptTemplateName: "decision-loop",
    promptTemplateParams: systemPromptArgs,
    stream: true,
    tool_choice: convertToolChoice(forceToolChoice),
    abortSignal,
  });

  for await (const streamItem of responseStream) {
    const llmResponse = streamItem.llmResponse;
    const message = getLLMResponseMessage(llmResponse);
    const toolCall = llmResponse.message?.tool_calls?.[0];
    // …parse tool args, build ToolCallRequest, extract UI component ID…
    const parsedChunk: Partial<LegacyComponentDecision> = {
      role: MessageRole.Assistant,
      message: displayMessage,
      componentName: isUITool ? toolCall.function.name.slice(UI_TOOLNAME_PREFIX.length) : "",
      componentId: isUITool ? (componentId ?? accumulatedDecision.componentId) : undefined,
      props: isUITool ? filteredToolArgs : null,
      toolCallRequest: clientToolRequest,
      toolCallId: toolCall ? getLLMResponseToolCallId(llmResponse) : undefined,
      statusMessage,
      completionStatusMessage,
      reasoning: llmResponse.reasoning ?? undefined,
      reasoningDurationMS: llmResponse.reasoningDurationMS ?? undefined,
    };
    accumulatedDecision = { ...accumulatedDecision, ...parsedChunk };
    yield { decision: accumulatedDecision, aguiEvents: streamItem.aguiEvents };
  }
}

Source: decision-loop-service.ts – lines 85‑118, 172‑226, 267‑284 and lines 172‑226

End-to-End Integration Example

The following example demonstrates creating a backend instance and running the decision loop to generate UI components:

import { createTamboBackend } from "./tambo-backend";

// 1️⃣ Build backend (LLM client is created inside)
const backend = await createTamboBackend(process.env.OPENAI_API_KEY, "chain-123", "user-456");

// 2️⃣ Prepare conversation and tool list
const messages = [{ role: "user", content: "Show me a chart of sales." }];
const tools = [{ type: "function", function: { name: "show_component_chart", parameters: {} } }];

// 3️⃣ Run the decision loop (streaming)
for await (const { decision, aguiEvents } of backend.runDecisionLoop({
  messages,
  strictTools: tools,
  customInstructions: undefined,
  forceToolChoice: undefined,
  resourceFetchers: {},   // empty map for this example
})) {
  console.log("UI decision:", decision);
  // UI can render events directly:
  // renderAGUIEvents(aguiEvents);
}

This pattern illustrates how Tambo AI's backend service orchestrates LLM interactions by abstracting provider details while exposing a streaming interface for real-time UI updates.

Summary

  • Provider-agnostic architecture: All LLM interactions funnel through the LLMClient interface, allowing new providers to be added without modifying decision-loop logic.
  • Streaming-first design: The backend always requests streaming responses (stream: true) and re-assembles partial responses into both legacy and AG-UI event formats.
  • Tool conversion layer: OpenAI-style tool definitions map to AI SDK ToolSet objects while preserving function names, descriptions, and JSON schemas.
  • Hierarchical configuration: Provider defaults cascade through model defaults, user-provided custom parameters, and model-specific provider params, ensuring flexibility with sensible defaults.
  • Event-driven UI rendering: By emitting BaseEvent[] arrays alongside each LLM chunk, the frontend renders text, tool calls, and reasoning incrementally for real-time visualizations.

Frequently Asked Questions

What is the role of the LLMClient interface in Tambo AI's backend?

The LLMClient interface defines the contract for all LLM interactions within Tambo AI's backend service. Located in packages/backend/src/services/llm/llm-client.ts (lines 58‑64), it specifies a complete() method that supports both streaming (AsyncIterableIterator<LLMStreamItem>) and non-streaming (LLMResponse) execution modes. This abstraction allows services like the decision-loop, suggestion generation, and thread-name generation to depend solely on the interface rather than concrete implementations, enabling easy testing with mock clients and seamless provider swaps.

How does Tambo AI handle different LLM providers like OpenAI and Anthropic?

Tambo AI handles multiple providers through the AISdkClient class in packages/backend/src/services/llm/ai-sdk-client.ts, which implements the LLMClient interface using the AI SDK. The getProviderFromModel() function (lines 98‑124) maps provider strings to specific factory functions—such as createOpenAI, createAnthropic, createMistral, createGoogle, and createGroq—while treating Cerebras and custom endpoints as openai-compatible. This architecture ensures that provider-specific authentication, parameter handling, and model instantiation are isolated within the client implementation, while the rest of the backend operates on normalized LLMStreamItem objects.

What is the difference between LLMStreamItem and DecisionStreamItem?

LLMStreamItem and DecisionStreamItem represent different abstraction layers in the orchestration pipeline. The LLMStreamItem (generated by AISdkClient in ai-sdk-client.ts, lines 456‑525) contains raw LLM output formatted as both a legacy LLMResponse object (mirroring OpenAI's structure) and an array of aguiEvents (AG-UI protocol events for UI rendering).

The DecisionStreamItem (produced by runDecisionLoop() in decision-loop-service.ts, lines 267‑284) represents a higher-level abstraction that extracts UI-specific information from the raw LLM output. It contains a LegacyComponentDecision object with parsed component names, props, tool call requests, and reasoning metadata—effectively transforming raw LLM tokens into structured UI decisions that the frontend can render immediately.

How does the decision loop handle streaming tool calls?

The decision loop handles streaming tool calls through incremental parsing within the runDecisionLoop() generator function (located in packages/backend/src/services/decision-loop/decision-loop-service.ts, lines 172‑226). As the function iterates over LLMStreamItem objects from the LLM client, it extracts tool call information from llmResponse.message.tool_calls and accumulates partial arguments across stream chunks.

For UI-specific tools (identified by isUiTool), the function slices the tool name prefix to extract the component name, filters the tool arguments to create component props, and extracts component IDs from custom AG-UI events. This incremental assembly allows the backend to yield partial DecisionStreamItem objects containing incomplete tool calls that get refined as more stream data arrives, enabling the frontend to display "thinking" states and partial component renders before the LLM finishes generating the complete tool call.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →