# Implementing Custom Callbacks for LLM and Tool Usage Tracking in TradingAgents

> Implement custom callbacks in TradingAgents to track LLM invocations, tool executions, and token counts. Monitor your AI trading agent effortlessly with this provider-agnostic architecture.

- Repository: [Tauric Research/TradingAgents](https://github.com/TauricResearch/TradingAgents)
- Tags: how-to-guide
- Published: 2026-03-23

---

**TradingAgents exposes a provider-agnostic callback architecture that lets you monitor every LLM invocation, tool execution, and token count by injecting a custom LangChain callback handler into the graph pipeline.**

TradingAgents is a modular trading intelligence framework that stitches together multiple LLM providers, a graph-based reasoning engine, and a Rich terminal UI. Implementing custom callbacks for LLM and tool usage tracking in TradingAgents gives you real-time observability into model consumption and function-call overhead without modifying core graph logic.

## The Callback Handler Architecture

The observability layer centers on a single class that subclasses LangChain’s `BaseCallbackHandler`. This design captures metrics at two critical boundaries: when the LLM starts generating (and ends), and when a tool is invoked.

### Subclassing BaseCallbackHandler

The reference implementation, `StatsCallbackHandler`, lives in [`cli/stats_handler.py`](https://github.com/TauricResearch/TradingAgents/blob/main/cli/stats_handler.py) and overrides four key hooks:

- `on_llm_start` – increments the call counter when any text-generation model starts.
- `on_chat_model_start` – handles chat-based providers (OpenAI, Claude, etc.).
- `on_llm_end` – captures token usage statistics from the response metadata.
- `on_tool_start` – increments the tool-call counter before a function executes.

```python

# cli/stats_handler.py

from langchain_core.callbacks import BaseCallbackHandler
import threading

class StatsCallbackHandler(BaseCallbackHandler):
    def __init__(self):
        super().__init__()
        self._lock = threading.Lock()
        self.llm_calls = 0
        self.tool_calls = 0
        self.input_tokens = 0
        self.output_tokens = 0

    def on_llm_start(self, *_, **__):
        with self._lock:
            self.llm_calls += 1

    def on_chat_model_start(self, *_, **__):
        with self._lock:
            self.llm_calls += 1

    def on_tool_start(self, *_, **__):
        with self._lock:
            self.tool_calls += 1

    def get_stats(self):
        with self._lock:
            return {
                "llm_calls": self.llm_calls,
                "tool_calls": self.tool_calls,
                "input_tokens": self.input_tokens,
                "output_tokens": self.output_tokens
            }

```

### Thread-Safety Considerations

Because `TradingAgentsGraph` runs LangGraph in asynchronous streaming mode, the handler uses a `threading.Lock` to prevent race conditions when multiple analyst nodes (research, trader, risk) invoke LLMs concurrently. This makes the single handler instance safe to share across the entire graph.

## Wiring Callbacks into the Pipeline

The callback is created once at the CLI layer, then propagated through three distinct integration points: the graph constructor, the LLM client factories, and the tool-execution propagator.

### CLI Instantiation

In [`cli/main.py`](https://github.com/TauricResearch/TradingAgents/blob/main/cli/main.py), the handler is instantiated before the graph is built. This single object is later passed into the `TradingAgentsGraph` constructor.

```python

# cli/main.py

from cli.stats_handler import StatsCallbackHandler
from tradingagents.graph.trading_graph import TradingAgentsGraph

def run_analysis(selected_analyst_keys, config):
    stats_handler = StatsCallbackHandler()  # ← singleton observer

    
    graph = TradingAgentsGraph(
        selected_analyst_keys,
        config=config,
        debug=True,
        callbacks=[stats_handler],  # ← injection point

    )
    return graph, stats_handler

```

### Graph Injection

The `TradingAgentsGraph` class ([`tradingagents/graph/trading_graph.py`](https://github.com/TauricResearch/TradingAgents/blob/main/tradingagents/graph/trading_graph.py)) accepts an optional `callbacks` list in its constructor (lines 51–63) and stores it as `self.callbacks`. By keeping the list as a graph property, the orchestrator can hand the same callback instance to every downstream component without coupling to specific LLM implementations.

```python

# tradingagents/graph/trading_graph.py (simplified)

class TradingAgentsGraph:
    def __init__(self, analysts, config, debug=False, callbacks=None):
        self.callbacks = callbacks or []
        self.config = config
        # ... analyst initialization

```

### LLM Client Attachment

When the graph builds its two reasoning tiers—`deep_thinking_llm` and `quick_thinking_llm`—it injects the callbacks into the keyword arguments consumed by `create_llm_client`. The factory ([`tradingagents/llm_clients/factory.py`](https://github.com/TauricResearch/TradingAgents/blob/main/tradingagents/llm_clients/factory.py)) forwards these kwargs to the concrete provider client (e.g., OpenAI, Anthropic).

```python

# tradingagents/graph/trading_graph.py (lines 74-80)

llm_kwargs = {"temperature": 0.7, "model": self.config.model}
if self.callbacks:
    llm_kwargs["callbacks"] = self.callbacks

deep_llm = create_llm_client(**llm_kwargs)
quick_llm = create_llm_client(**llm_kwargs, quick_mode=True)

```

### Tool Execution Propagation

Tool calls are tracked separately from LLM calls. The `Propagator` class ([`tradingagents/graph/propagation.py`](https://github.com/TauricResearch/TradingAgents/blob/main/tradingagents/graph/propagation.py)) supplies the same `callbacks` list when invoking the LangGraph execution step (lines 56–67). This ensures that `on_tool_start` fires whenever an analyst node triggers a search, calculation, or portfolio update.

```python

# tradingagents/graph/propagation.py (conceptual)

self.propagator = Propagator(
    callbacks=self.callbacks  # ensures tool nodes are observed

)

```

## Displaying and Persisting Metrics

Once the pipeline is instrumented, you can query the handler at any point to surface metrics.

### Live UI Rendering

The Rich-based terminal UI ([`cli/main.py`](https://github.com/TauricResearch/TradingAgents/blob/main/cli/main.py)) polls `stats_handler.get_stats()` inside the display-update loop and renders the counts in the footer. Because the handler is thread-safe, the UI thread can read counters while background graph nodes are still writing them.

```python

# Inside the UI loop (cli/main.py)

current_stats = stats_handler.get_stats()
footer_text = (
    f"LLM Calls: {current_stats['llm_calls']} | "
    f"Tool Calls: {current_stats['tool_calls']}"
)

```

### Post-Run Analysis

After the graph finishes, the same handler instance can be queried to persist data for offline analysis. The following snippet writes a CSV report containing the final tallies:

```python
import csv
from pathlib import Path

def export_stats(path: Path, handler: StatsCallbackHandler):
    stats = handler.get_stats()
    with path.open("w", newline="") as f:
        writer = csv.writer(f)
        writer.writerow(["metric", "value"])
        for key, val in stats.items():
            writer.writerow([key, val])

```

## Extending the System

Adding new metrics requires only subclassing additional hooks. For example, to track latency or error rates, override `on_llm_error` or `on_tool_end` in your custom handler:

```python
class ExtendedStatsHandler(StatsCallbackHandler):
    def on_llm_error(self, error, **kwargs):
        with self._lock:
            self.error_count += 1
    
    def on_tool_end(self, output, **kwargs):
        # Log tool result size or latency here

        pass

```

Because the graph receives the callback list purely as a constructor argument, any provider-specific or use-case-specific handler works without changes to [`tradingagents/graph/trading_graph.py`](https://github.com/TauricResearch/TradingAgents/blob/main/tradingagents/graph/trading_graph.py).

## Summary

- **Define a thread-safe handler** by subclassing `BaseCallbackHandler` and overriding `on_llm_start`, `on_chat_model_start`, `on_llm_end`, and `on_tool_start` in [`cli/stats_handler.py`](https://github.com/TauricResearch/TradingAgents/blob/main/cli/stats_handler.py).
- **Instantiate once** in [`cli/main.py`](https://github.com/TauricResearch/TradingAgents/blob/main/cli/main.py) and pass the list into `TradingAgentsGraph` via the `callbacks` parameter.
- **Attach automatically** to all LLM clients when [`tradingagents/graph/trading_graph.py`](https://github.com/TauricResearch/TradingAgents/blob/main/tradingagents/graph/trading_graph.py) builds `deep_thinking_llm` and `quick_thinking_llm` through the factory in [`tradingagents/llm_clients/factory.py`](https://github.com/TauricResearch/TradingAgents/blob/main/tradingagents/llm_clients/factory.py).
- **Propagate to tools** via the `Propagator` in [`tradingagents/graph/propagation.py`](https://github.com/TauricResearch/TradingAgents/blob/main/tradingagents/graph/propagation.py) so that function calls are counted alongside LLM invocations.
- **Query anytime** using `get_stats()` to drive live UIs or export post-run reports.

## Frequently Asked Questions

### How do I track token usage rather than just call counts?

Override `on_llm_end` in your handler and inspect the `response` object’s `usage_metadata`. The TradingAgents reference implementation stores `response.usage_metadata.get("input_tokens")` and `output_tokens` into counters protected by the same `threading.Lock` used for call counts.

### Can I use multiple callback handlers simultaneously?

Yes. The `TradingAgentsGraph` constructor accepts a list (`callbacks=[handler1, handler2]`), and LangChain will invoke every hook on every handler in the order provided. This lets you separate concerns—for example, one handler for metrics and another for logging—without merging code.

### What happens if a tool fails? Will `on_tool_start` still fire?

`on_tool_start` fires immediately before the tool executes, so it will increment regardless of success. To track failures, also override `on_tool_error` (or `on_tool_end` and check the output). The handler in [`cli/stats_handler.py`](https://github.com/TauricResearch/TradingAgents/blob/main/cli/stats_handler.py) can be extended with an `error_count` field and locked updates inside `on_tool_error`.

### Is the callback system compatible with async streaming?

Yes. The `threading.Lock` inside `StatsCallbackHandler` makes the counters safe for LangGraph’s asynchronous streaming mode. The handlers are invoked in the event loop, but the lock ensures atomic increments when multiple analyst nodes run concurrently.