Token Tracking and Cost Optimization in ChatDev LLM Agent Calls: A Complete Guide

ChatDev instruments every LLM invocation with a built-in TokenTracker that records input, output, and total token consumption per workflow node, enabling precise cost analysis and budget control through JSON exports and programmatic APIs.

The OpenBMB/ChatDev framework provides comprehensive token tracking capabilities that allow developers to monitor, analyze, and optimize the costs associated with multi-agent LLM workflows. By capturing granular usage data at the provider, model, and node level, ChatDev transforms opaque API costs into actionable metrics for engineering teams managing production agent systems.

Core Architecture of Token Tracking in ChatDev

ChatDev implements token tracking through a layered architecture that spans data structures, runtime context management, and provider-specific instrumentation.

The TokenUsage Dataclass and TokenTracker Class

At the foundation lies the TokenUsage dataclass defined in utils/token_tracker.py. This structure stores raw token counts (prompt_tokens, completion_tokens, total_tokens) alongside metadata including node_id, model_name, provider, and a flexible metadata dictionary for extensibility.

The TokenTracker class—also in utils/token_tracker.py—functions as a singleton-like instance per workflow execution. It maintains:

  • total_usage: Cumulative tokens across all calls
  • node_usages: Aggregated consumption per workflow node
  • model_usages: Breakdown by specific models (e.g., gpt-4, gemini-pro)
  • node_call_counts: Execution frequency tracking to identify loops
  • call_history: Chronological record with execution_number to distinguish repeated node invocations

The tracker provides export_to_file() for JSON serialization and getter methods like get_token_usage() and get_node_execution_count() for runtime inspection.

Runtime Integration and Context Propagation

The tracking system initializes in workflow/runtime/runtime_builder.py, where RuntimeBuilder.build() instantiates TokenTracker(workflow_id=self.graph.name) and injects it into the execution context.

The RuntimeContext dataclass in workflow/runtime/runtime_context.py carries the token_tracker field, making the tracker globally accessible to all node executors without relying on global variables. When AgentExecutor runs a node, it attaches the tracker to the provider configuration via agent_config.token_tracker = self.context.get_token_tracker().

Provider-Level Instrumentation

Each LLM provider integrates with the tracking system through standardized hooks:

OpenAI Provider (runtime/node/agent/providers/openai_provider.py):

  • Extracts token usage via extract_token_usage() from API responses
  • Calls _track_token_usage() to enrich TokenUsage objects with execution context
  • Records usage through token_tracker.record_usage()

Gemini Provider (runtime/node/agent/providers/gemini_provider.py):

  • Implements identical patterns for Google's Gemini API
  • Ensures cross-provider cost comparison capabilities

Both providers automatically capture the node_id, model_name, workflow_id, and provider fields before recording, creating a complete audit trail.

Persistence and Export Mechanisms

After workflow completion, ResultArchiver in workflow/runtime/result_archiver.py calls token_tracker.export_to_file() to generate token_usage_<workflow>.json. This file contains the full aggregated dataset including per-node, per-model, and total usage statistics.

For API consumers, server/routes/execute_sync.py includes token usage in HTTP responses, while runtime/sdk.py exposes the data through the Python SDK.

How Token Tracking Works During Execution

The token flow follows a six-stage pipeline during workflow execution:

  1. Initialization: RuntimeBuilder creates a fresh TokenTracker instance when the workflow graph instantiates
  2. Configuration: The node executor attaches the tracker reference to the provider config before calling the LLM
  3. Request: The provider (OpenAI or Gemini) transmits the prompt to the respective API
  4. Extraction: Upon receiving the response, extract_token_usage() parses prompt_tokens, completion_tokens, and total_tokens
  5. Recording: _track_token_usage() constructs a TokenUsage object enriched with node and model metadata, then invokes token_tracker.record_usage() to update cumulative counters and append to call_history
  6. Export: ResultArchiver writes the aggregated data to token_usage_<workflow>.json and injects usage statistics into the workflow log

This instrumentation adds minimal overhead while capturing complete cost visibility.

Practical Cost Optimization Strategies

ChatDev's token tracking enables several concrete optimization patterns for production deployments.

Identifying Expensive Workflow Nodes

The node_usages dictionary reveals which specific workflow steps consume the most tokens. By calling token_tracker.get_node_usage(node_id), developers can identify bottlenecks such as:

  • Iterative refinement loops that accumulate thousands of tokens
  • Context-heavy nodes that process large system prompts
  • Multi-turn conversation nodes with extended histories

Once identified, expensive nodes can be refactored, cached, or replaced with lighter prompt templates.

Model Selection and Provider Comparison

The model_usages and provider fields enable data-driven model selection. Teams can compare token consumption between gpt-4 and gpt-3.5-turbo, or between OpenAI and Gemini implementations, to optimize the cost-quality trade-off.

For example, if token_tracker.get_token_usage()["model_usages"] shows that a particular node uses 90% of the budget with gpt-4, you can test the same node with gpt-3.5-turbo and compare output quality against the cost savings shown in the tracker.

Preventing Unnecessary Calls

The node_call_counts dictionary helps detect inefficient looping patterns. If a node executes 50 times when only 5 calls were expected, the tracker reveals this immediately through token_tracker.get_node_execution_count(node_id).

Developers can implement circuit breakers:

if token_tracker.get_node_execution_count("search_web") > threshold:
    # Switch to cached results or terminate early

    pass

Code Examples for Token Monitoring

Retrieve Token Usage After Workflow Execution

When using the ChatDev SDK, token statistics return automatically alongside results:

from chatdev.runtime.sdk import ChatDevRuntime

runtime = ChatDevRuntime(workflow_path="my_workflow")
result, token_usage = runtime.run()

print(f"Final result: {result}")
print(f"Total tokens: {token_usage['total_usage']['total_tokens']}")
print(f"Per-node breakdown: {token_usage['node_usages']}")

The run() method internally accesses executor.token_tracker.get_token_usage() as implemented in runtime/sdk.py.

Inspect Per-Model Consumption

Analyze which models drive costs:

usage = token_tracker.get_token_usage()
for model, stats in usage["model_usages"].items():
    print(f"Model {model}: {stats['total_tokens']} tokens")
    print(f"  Input: {stats.get('prompt_tokens', 0)}")
    print(f"  Output: {stats.get('completion_tokens', 0)}")

Export Usage Data Manually

For custom monitoring pipelines:

from utils.token_tracker import TokenTracker

tracker = TokenTracker(workflow_id="custom_analysis")

# ... execute workflow nodes ...

tracker.export_to_file("outputs/usage_report.json")

Implement Runtime Cost Controls

Adjust node behavior based on accumulated costs:


# Check if specific node exceeds budget

if token_tracker.get_node_usage("data_analysis")["total_tokens"] > 10000:
    # Reduce max_tokens for subsequent calls

    node_config.max_tokens = 256
    # Or switch to cheaper model

    node_config.model = "gpt-3.5-turbo"

Summary

  • TokenTracker in utils/token_tracker.py provides comprehensive instrumentation for every LLM call, tracking input, output, and total tokens alongside node, model, and provider metadata.
  • The architecture spans data structures (TokenUsage), runtime context propagation (RuntimeBuilder, RuntimeContext), provider implementations (OpenAI and Gemini), and persistence (ResultArchiver).
  • Cost optimization leverages per-node usage statistics (node_usages), model-level aggregation (model_usages), and execution counting (node_call_counts) to identify expensive patterns and optimize spending.
  • Programmatic access via token_tracker.get_token_usage() enables real-time budget enforcement, while JSON exports (token_usage_<workflow>.json) support post-hoc analysis and CI/CD integration.
  • The SDK's run() method and server endpoints (/execute_sync) expose token data directly to calling applications, making ChatDev suitable for production cost management.

Frequently Asked Questions

How does ChatDev track tokens when using multiple different LLM providers in the same workflow?

ChatDev tracks tokens across providers through standardized instrumentation in each provider implementation. The OpenAI provider in runtime/node/agent/providers/openai_provider.py and the Gemini provider in runtime/node/agent/providers/gemini_provider.py both extract token counts using provider-specific logic (parsing prompt_tokens and completion_tokens for OpenAI, or equivalent fields for Gemini), then record usage through the shared TokenTracker instance. The provider field in each TokenUsage object preserves the source attribution, allowing comparison of costs between OpenAI and Google APIs within the same token_usage_<workflow>.json export.

Can I access token usage data while a workflow is still running, or only after completion?

You can access token usage data during workflow execution through the RuntimeContext. Since TokenTracker maintains cumulative counters in memory, calling context.get_token_tracker().get_token_usage() from within custom node code or external monitoring hooks provides real-time cost visibility. This enables mid-workflow decisions, such as aborting execution if total_usage exceeds a budget threshold, or dynamically switching models based on current consumption rates.

What specific file contains the token usage data after a workflow finishes?

The aggregated token usage persists to token_usage_<workflow>.json in the working directory, generated by ResultArchiver.export() in workflow/runtime/result_archiver.py. The filename includes the workflow identifier set during TokenTracker initialization in RuntimeBuilder.build(). This JSON file contains the complete total_usage object, per-node breakdowns, per-model statistics, chronological call_history, and execution counts.

How can I prevent a single workflow node from consuming excessive tokens?

ChatDev provides several mechanisms for node-level cost control. First, inspect token_tracker.get_node_execution_count(node_id) to detect unexpected loops. Second, monitor token_tracker.get_node_usage(node_id)["total_tokens"] to set thresholds. Third, configure max_tokens parameters in node configurations to cap individual calls. Finally, implement caching logic that checks usage statistics before allowing repeated invocations of expensive nodes like web search or data analysis steps.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →