# Token Tracking and Cost Optimization in ChatDev LLM Agent Calls: A Complete Guide

> Master token tracking and cost optimization in ChatDev LLM agent calls. Our guide details how ChatDev's TokenTracker enables precise cost analysis and budget control for your workflows.

- Repository: [OpenBMB/ChatDev](https://github.com/OpenBMB/ChatDev)
- Tags: how-to-guide
- Published: 2026-04-01

---

**ChatDev instruments every LLM invocation with a built-in TokenTracker that records input, output, and total token consumption per workflow node, enabling precise cost analysis and budget control through JSON exports and programmatic APIs.**

The OpenBMB/ChatDev framework provides comprehensive token tracking capabilities that allow developers to monitor, analyze, and optimize the costs associated with multi-agent LLM workflows. By capturing granular usage data at the provider, model, and node level, ChatDev transforms opaque API costs into actionable metrics for engineering teams managing production agent systems.

## Core Architecture of Token Tracking in ChatDev

ChatDev implements token tracking through a layered architecture that spans data structures, runtime context management, and provider-specific instrumentation.

### The TokenUsage Dataclass and TokenTracker Class

At the foundation lies the **`TokenUsage`** dataclass defined in [`utils/token_tracker.py`](https://github.com/OpenBMB/ChatDev/blob/main/utils/token_tracker.py). This structure stores raw token counts (`prompt_tokens`, `completion_tokens`, `total_tokens`) alongside metadata including `node_id`, `model_name`, `provider`, and a flexible `metadata` dictionary for extensibility.

The **`TokenTracker`** class—also in [`utils/token_tracker.py`](https://github.com/OpenBMB/ChatDev/blob/main/utils/token_tracker.py)—functions as a singleton-like instance per workflow execution. It maintains:

- **`total_usage`**: Cumulative tokens across all calls
- **`node_usages`**: Aggregated consumption per workflow node
- **`model_usages`**: Breakdown by specific models (e.g., `gpt-4`, `gemini-pro`)
- **`node_call_counts`**: Execution frequency tracking to identify loops
- **`call_history`**: Chronological record with `execution_number` to distinguish repeated node invocations

The tracker provides `export_to_file()` for JSON serialization and getter methods like `get_token_usage()` and `get_node_execution_count()` for runtime inspection.

### Runtime Integration and Context Propagation

The tracking system initializes in [`workflow/runtime/runtime_builder.py`](https://github.com/OpenBMB/ChatDev/blob/main/workflow/runtime/runtime_builder.py), where `RuntimeBuilder.build()` instantiates `TokenTracker(workflow_id=self.graph.name)` and injects it into the execution context.

The **`RuntimeContext`** dataclass in [`workflow/runtime/runtime_context.py`](https://github.com/OpenBMB/ChatDev/blob/main/workflow/runtime/runtime_context.py) carries the `token_tracker` field, making the tracker globally accessible to all node executors without relying on global variables. When `AgentExecutor` runs a node, it attaches the tracker to the provider configuration via `agent_config.token_tracker = self.context.get_token_tracker()`.

### Provider-Level Instrumentation

Each LLM provider integrates with the tracking system through standardized hooks:

**OpenAI Provider** ([`runtime/node/agent/providers/openai_provider.py`](https://github.com/OpenBMB/ChatDev/blob/main/runtime/node/agent/providers/openai_provider.py)):
- Extracts token usage via `extract_token_usage()` from API responses
- Calls `_track_token_usage()` to enrich `TokenUsage` objects with execution context
- Records usage through `token_tracker.record_usage()`

**Gemini Provider** ([`runtime/node/agent/providers/gemini_provider.py`](https://github.com/OpenBMB/ChatDev/blob/main/runtime/node/agent/providers/gemini_provider.py)):
- Implements identical patterns for Google's Gemini API
- Ensures cross-provider cost comparison capabilities

Both providers automatically capture the `node_id`, `model_name`, `workflow_id`, and `provider` fields before recording, creating a complete audit trail.

### Persistence and Export Mechanisms

After workflow completion, **`ResultArchiver`** in [`workflow/runtime/result_archiver.py`](https://github.com/OpenBMB/ChatDev/blob/main/workflow/runtime/result_archiver.py) calls `token_tracker.export_to_file()` to generate `token_usage_<workflow>.json`. This file contains the full aggregated dataset including per-node, per-model, and total usage statistics.

For API consumers, [`server/routes/execute_sync.py`](https://github.com/OpenBMB/ChatDev/blob/main/server/routes/execute_sync.py) includes token usage in HTTP responses, while [`runtime/sdk.py`](https://github.com/OpenBMB/ChatDev/blob/main/runtime/sdk.py) exposes the data through the Python SDK.

## How Token Tracking Works During Execution

The token flow follows a six-stage pipeline during workflow execution:

1. **Initialization**: `RuntimeBuilder` creates a fresh `TokenTracker` instance when the workflow graph instantiates
2. **Configuration**: The node executor attaches the tracker reference to the provider config before calling the LLM
3. **Request**: The provider (OpenAI or Gemini) transmits the prompt to the respective API
4. **Extraction**: Upon receiving the response, `extract_token_usage()` parses `prompt_tokens`, `completion_tokens`, and `total_tokens`
5. **Recording**: `_track_token_usage()` constructs a `TokenUsage` object enriched with node and model metadata, then invokes `token_tracker.record_usage()` to update cumulative counters and append to `call_history`
6. **Export**: `ResultArchiver` writes the aggregated data to `token_usage_<workflow>.json` and injects usage statistics into the workflow log

This instrumentation adds minimal overhead while capturing complete cost visibility.

## Practical Cost Optimization Strategies

ChatDev's token tracking enables several concrete optimization patterns for production deployments.

### Identifying Expensive Workflow Nodes

The **`node_usages`** dictionary reveals which specific workflow steps consume the most tokens. By calling `token_tracker.get_node_usage(node_id)`, developers can identify bottlenecks such as:

- Iterative refinement loops that accumulate thousands of tokens
- Context-heavy nodes that process large system prompts
- Multi-turn conversation nodes with extended histories

Once identified, expensive nodes can be refactored, cached, or replaced with lighter prompt templates.

### Model Selection and Provider Comparison

The **`model_usages`** and **`provider`** fields enable data-driven model selection. Teams can compare token consumption between `gpt-4` and `gpt-3.5-turbo`, or between OpenAI and Gemini implementations, to optimize the cost-quality trade-off.

For example, if `token_tracker.get_token_usage()["model_usages"]` shows that a particular node uses 90% of the budget with `gpt-4`, you can test the same node with `gpt-3.5-turbo` and compare output quality against the cost savings shown in the tracker.

### Preventing Unnecessary Calls

The **`node_call_counts`** dictionary helps detect inefficient looping patterns. If a node executes 50 times when only 5 calls were expected, the tracker reveals this immediately through `token_tracker.get_node_execution_count(node_id)`.

Developers can implement circuit breakers:

```python
if token_tracker.get_node_execution_count("search_web") > threshold:
    # Switch to cached results or terminate early

    pass

```

## Code Examples for Token Monitoring

### Retrieve Token Usage After Workflow Execution

When using the ChatDev SDK, token statistics return automatically alongside results:

```python
from chatdev.runtime.sdk import ChatDevRuntime

runtime = ChatDevRuntime(workflow_path="my_workflow")
result, token_usage = runtime.run()

print(f"Final result: {result}")
print(f"Total tokens: {token_usage['total_usage']['total_tokens']}")
print(f"Per-node breakdown: {token_usage['node_usages']}")

```

The `run()` method internally accesses `executor.token_tracker.get_token_usage()` as implemented in [`runtime/sdk.py`](https://github.com/OpenBMB/ChatDev/blob/main/runtime/sdk.py).

### Inspect Per-Model Consumption

Analyze which models drive costs:

```python
usage = token_tracker.get_token_usage()
for model, stats in usage["model_usages"].items():
    print(f"Model {model}: {stats['total_tokens']} tokens")
    print(f"  Input: {stats.get('prompt_tokens', 0)}")
    print(f"  Output: {stats.get('completion_tokens', 0)}")

```

### Export Usage Data Manually

For custom monitoring pipelines:

```python
from utils.token_tracker import TokenTracker

tracker = TokenTracker(workflow_id="custom_analysis")

# ... execute workflow nodes ...

tracker.export_to_file("outputs/usage_report.json")

```

### Implement Runtime Cost Controls

Adjust node behavior based on accumulated costs:

```python

# Check if specific node exceeds budget

if token_tracker.get_node_usage("data_analysis")["total_tokens"] > 10000:
    # Reduce max_tokens for subsequent calls

    node_config.max_tokens = 256
    # Or switch to cheaper model

    node_config.model = "gpt-3.5-turbo"

```

## Summary

- **TokenTracker** in [`utils/token_tracker.py`](https://github.com/OpenBMB/ChatDev/blob/main/utils/token_tracker.py) provides comprehensive instrumentation for every LLM call, tracking input, output, and total tokens alongside node, model, and provider metadata.
- The architecture spans data structures (`TokenUsage`), runtime context propagation (`RuntimeBuilder`, `RuntimeContext`), provider implementations (OpenAI and Gemini), and persistence (`ResultArchiver`).
- **Cost optimization** leverages per-node usage statistics (`node_usages`), model-level aggregation (`model_usages`), and execution counting (`node_call_counts`) to identify expensive patterns and optimize spending.
- **Programmatic access** via `token_tracker.get_token_usage()` enables real-time budget enforcement, while JSON exports (`token_usage_<workflow>.json`) support post-hoc analysis and CI/CD integration.
- The SDK's `run()` method and server endpoints (`/execute_sync`) expose token data directly to calling applications, making ChatDev suitable for production cost management.

## Frequently Asked Questions

### How does ChatDev track tokens when using multiple different LLM providers in the same workflow?

ChatDev tracks tokens across providers through standardized instrumentation in each provider implementation. The OpenAI provider in [`runtime/node/agent/providers/openai_provider.py`](https://github.com/OpenBMB/ChatDev/blob/main/runtime/node/agent/providers/openai_provider.py) and the Gemini provider in [`runtime/node/agent/providers/gemini_provider.py`](https://github.com/OpenBMB/ChatDev/blob/main/runtime/node/agent/providers/gemini_provider.py) both extract token counts using provider-specific logic (parsing `prompt_tokens` and `completion_tokens` for OpenAI, or equivalent fields for Gemini), then record usage through the shared `TokenTracker` instance. The `provider` field in each `TokenUsage` object preserves the source attribution, allowing comparison of costs between OpenAI and Google APIs within the same `token_usage_<workflow>.json` export.

### Can I access token usage data while a workflow is still running, or only after completion?

You can access token usage data during workflow execution through the `RuntimeContext`. Since `TokenTracker` maintains cumulative counters in memory, calling `context.get_token_tracker().get_token_usage()` from within custom node code or external monitoring hooks provides real-time cost visibility. This enables mid-workflow decisions, such as aborting execution if `total_usage` exceeds a budget threshold, or dynamically switching models based on current consumption rates.

### What specific file contains the token usage data after a workflow finishes?

The aggregated token usage persists to `token_usage_<workflow>.json` in the working directory, generated by `ResultArchiver.export()` in [`workflow/runtime/result_archiver.py`](https://github.com/OpenBMB/ChatDev/blob/main/workflow/runtime/result_archiver.py). The filename includes the workflow identifier set during `TokenTracker` initialization in `RuntimeBuilder.build()`. This JSON file contains the complete `total_usage` object, per-node breakdowns, per-model statistics, chronological `call_history`, and execution counts.

### How can I prevent a single workflow node from consuming excessive tokens?

ChatDev provides several mechanisms for node-level cost control. First, inspect `token_tracker.get_node_execution_count(node_id)` to detect unexpected loops. Second, monitor `token_tracker.get_node_usage(node_id)["total_tokens"]` to set thresholds. Third, configure `max_tokens` parameters in node configurations to cap individual calls. Finally, implement caching logic that checks usage statistics before allowing repeated invocations of expensive nodes like web search or data analysis steps.