Token Tracking and Cost Optimization in ChatDev LLM Agent Calls: A Complete Guide
ChatDev instruments every LLM invocation with a built-in TokenTracker that records input, output, and total token consumption per workflow node, enabling precise cost analysis and budget control through JSON exports and programmatic APIs.
The OpenBMB/ChatDev framework provides comprehensive token tracking capabilities that allow developers to monitor, analyze, and optimize the costs associated with multi-agent LLM workflows. By capturing granular usage data at the provider, model, and node level, ChatDev transforms opaque API costs into actionable metrics for engineering teams managing production agent systems.
Core Architecture of Token Tracking in ChatDev
ChatDev implements token tracking through a layered architecture that spans data structures, runtime context management, and provider-specific instrumentation.
The TokenUsage Dataclass and TokenTracker Class
At the foundation lies the TokenUsage dataclass defined in utils/token_tracker.py. This structure stores raw token counts (prompt_tokens, completion_tokens, total_tokens) alongside metadata including node_id, model_name, provider, and a flexible metadata dictionary for extensibility.
The TokenTracker class—also in utils/token_tracker.py—functions as a singleton-like instance per workflow execution. It maintains:
total_usage: Cumulative tokens across all callsnode_usages: Aggregated consumption per workflow nodemodel_usages: Breakdown by specific models (e.g.,gpt-4,gemini-pro)node_call_counts: Execution frequency tracking to identify loopscall_history: Chronological record withexecution_numberto distinguish repeated node invocations
The tracker provides export_to_file() for JSON serialization and getter methods like get_token_usage() and get_node_execution_count() for runtime inspection.
Runtime Integration and Context Propagation
The tracking system initializes in workflow/runtime/runtime_builder.py, where RuntimeBuilder.build() instantiates TokenTracker(workflow_id=self.graph.name) and injects it into the execution context.
The RuntimeContext dataclass in workflow/runtime/runtime_context.py carries the token_tracker field, making the tracker globally accessible to all node executors without relying on global variables. When AgentExecutor runs a node, it attaches the tracker to the provider configuration via agent_config.token_tracker = self.context.get_token_tracker().
Provider-Level Instrumentation
Each LLM provider integrates with the tracking system through standardized hooks:
OpenAI Provider (runtime/node/agent/providers/openai_provider.py):
- Extracts token usage via
extract_token_usage()from API responses - Calls
_track_token_usage()to enrichTokenUsageobjects with execution context - Records usage through
token_tracker.record_usage()
Gemini Provider (runtime/node/agent/providers/gemini_provider.py):
- Implements identical patterns for Google's Gemini API
- Ensures cross-provider cost comparison capabilities
Both providers automatically capture the node_id, model_name, workflow_id, and provider fields before recording, creating a complete audit trail.
Persistence and Export Mechanisms
After workflow completion, ResultArchiver in workflow/runtime/result_archiver.py calls token_tracker.export_to_file() to generate token_usage_<workflow>.json. This file contains the full aggregated dataset including per-node, per-model, and total usage statistics.
For API consumers, server/routes/execute_sync.py includes token usage in HTTP responses, while runtime/sdk.py exposes the data through the Python SDK.
How Token Tracking Works During Execution
The token flow follows a six-stage pipeline during workflow execution:
- Initialization:
RuntimeBuildercreates a freshTokenTrackerinstance when the workflow graph instantiates - Configuration: The node executor attaches the tracker reference to the provider config before calling the LLM
- Request: The provider (OpenAI or Gemini) transmits the prompt to the respective API
- Extraction: Upon receiving the response,
extract_token_usage()parsesprompt_tokens,completion_tokens, andtotal_tokens - Recording:
_track_token_usage()constructs aTokenUsageobject enriched with node and model metadata, then invokestoken_tracker.record_usage()to update cumulative counters and append tocall_history - Export:
ResultArchiverwrites the aggregated data totoken_usage_<workflow>.jsonand injects usage statistics into the workflow log
This instrumentation adds minimal overhead while capturing complete cost visibility.
Practical Cost Optimization Strategies
ChatDev's token tracking enables several concrete optimization patterns for production deployments.
Identifying Expensive Workflow Nodes
The node_usages dictionary reveals which specific workflow steps consume the most tokens. By calling token_tracker.get_node_usage(node_id), developers can identify bottlenecks such as:
- Iterative refinement loops that accumulate thousands of tokens
- Context-heavy nodes that process large system prompts
- Multi-turn conversation nodes with extended histories
Once identified, expensive nodes can be refactored, cached, or replaced with lighter prompt templates.
Model Selection and Provider Comparison
The model_usages and provider fields enable data-driven model selection. Teams can compare token consumption between gpt-4 and gpt-3.5-turbo, or between OpenAI and Gemini implementations, to optimize the cost-quality trade-off.
For example, if token_tracker.get_token_usage()["model_usages"] shows that a particular node uses 90% of the budget with gpt-4, you can test the same node with gpt-3.5-turbo and compare output quality against the cost savings shown in the tracker.
Preventing Unnecessary Calls
The node_call_counts dictionary helps detect inefficient looping patterns. If a node executes 50 times when only 5 calls were expected, the tracker reveals this immediately through token_tracker.get_node_execution_count(node_id).
Developers can implement circuit breakers:
if token_tracker.get_node_execution_count("search_web") > threshold:
# Switch to cached results or terminate early
pass
Code Examples for Token Monitoring
Retrieve Token Usage After Workflow Execution
When using the ChatDev SDK, token statistics return automatically alongside results:
from chatdev.runtime.sdk import ChatDevRuntime
runtime = ChatDevRuntime(workflow_path="my_workflow")
result, token_usage = runtime.run()
print(f"Final result: {result}")
print(f"Total tokens: {token_usage['total_usage']['total_tokens']}")
print(f"Per-node breakdown: {token_usage['node_usages']}")
The run() method internally accesses executor.token_tracker.get_token_usage() as implemented in runtime/sdk.py.
Inspect Per-Model Consumption
Analyze which models drive costs:
usage = token_tracker.get_token_usage()
for model, stats in usage["model_usages"].items():
print(f"Model {model}: {stats['total_tokens']} tokens")
print(f" Input: {stats.get('prompt_tokens', 0)}")
print(f" Output: {stats.get('completion_tokens', 0)}")
Export Usage Data Manually
For custom monitoring pipelines:
from utils.token_tracker import TokenTracker
tracker = TokenTracker(workflow_id="custom_analysis")
# ... execute workflow nodes ...
tracker.export_to_file("outputs/usage_report.json")
Implement Runtime Cost Controls
Adjust node behavior based on accumulated costs:
# Check if specific node exceeds budget
if token_tracker.get_node_usage("data_analysis")["total_tokens"] > 10000:
# Reduce max_tokens for subsequent calls
node_config.max_tokens = 256
# Or switch to cheaper model
node_config.model = "gpt-3.5-turbo"
Summary
- TokenTracker in
utils/token_tracker.pyprovides comprehensive instrumentation for every LLM call, tracking input, output, and total tokens alongside node, model, and provider metadata. - The architecture spans data structures (
TokenUsage), runtime context propagation (RuntimeBuilder,RuntimeContext), provider implementations (OpenAI and Gemini), and persistence (ResultArchiver). - Cost optimization leverages per-node usage statistics (
node_usages), model-level aggregation (model_usages), and execution counting (node_call_counts) to identify expensive patterns and optimize spending. - Programmatic access via
token_tracker.get_token_usage()enables real-time budget enforcement, while JSON exports (token_usage_<workflow>.json) support post-hoc analysis and CI/CD integration. - The SDK's
run()method and server endpoints (/execute_sync) expose token data directly to calling applications, making ChatDev suitable for production cost management.
Frequently Asked Questions
How does ChatDev track tokens when using multiple different LLM providers in the same workflow?
ChatDev tracks tokens across providers through standardized instrumentation in each provider implementation. The OpenAI provider in runtime/node/agent/providers/openai_provider.py and the Gemini provider in runtime/node/agent/providers/gemini_provider.py both extract token counts using provider-specific logic (parsing prompt_tokens and completion_tokens for OpenAI, or equivalent fields for Gemini), then record usage through the shared TokenTracker instance. The provider field in each TokenUsage object preserves the source attribution, allowing comparison of costs between OpenAI and Google APIs within the same token_usage_<workflow>.json export.
Can I access token usage data while a workflow is still running, or only after completion?
You can access token usage data during workflow execution through the RuntimeContext. Since TokenTracker maintains cumulative counters in memory, calling context.get_token_tracker().get_token_usage() from within custom node code or external monitoring hooks provides real-time cost visibility. This enables mid-workflow decisions, such as aborting execution if total_usage exceeds a budget threshold, or dynamically switching models based on current consumption rates.
What specific file contains the token usage data after a workflow finishes?
The aggregated token usage persists to token_usage_<workflow>.json in the working directory, generated by ResultArchiver.export() in workflow/runtime/result_archiver.py. The filename includes the workflow identifier set during TokenTracker initialization in RuntimeBuilder.build(). This JSON file contains the complete total_usage object, per-node breakdowns, per-model statistics, chronological call_history, and execution counts.
How can I prevent a single workflow node from consuming excessive tokens?
ChatDev provides several mechanisms for node-level cost control. First, inspect token_tracker.get_node_execution_count(node_id) to detect unexpected loops. Second, monitor token_tracker.get_node_usage(node_id)["total_tokens"] to set thresholds. Third, configure max_tokens parameters in node configurations to cap individual calls. Finally, implement caching logic that checks usage statistics before allowing repeated invocations of expensive nodes like web search or data analysis steps.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →