Implementing Custom Callbacks for LLM and Tool Usage Tracking in TradingAgents
TradingAgents exposes a provider-agnostic callback architecture that lets you monitor every LLM invocation, tool execution, and token count by injecting a custom LangChain callback handler into the graph pipeline.
TradingAgents is a modular trading intelligence framework that stitches together multiple LLM providers, a graph-based reasoning engine, and a Rich terminal UI. Implementing custom callbacks for LLM and tool usage tracking in TradingAgents gives you real-time observability into model consumption and function-call overhead without modifying core graph logic.
The Callback Handler Architecture
The observability layer centers on a single class that subclasses LangChain’s BaseCallbackHandler. This design captures metrics at two critical boundaries: when the LLM starts generating (and ends), and when a tool is invoked.
Subclassing BaseCallbackHandler
The reference implementation, StatsCallbackHandler, lives in cli/stats_handler.py and overrides four key hooks:
on_llm_start– increments the call counter when any text-generation model starts.on_chat_model_start– handles chat-based providers (OpenAI, Claude, etc.).on_llm_end– captures token usage statistics from the response metadata.on_tool_start– increments the tool-call counter before a function executes.
# cli/stats_handler.py
from langchain_core.callbacks import BaseCallbackHandler
import threading
class StatsCallbackHandler(BaseCallbackHandler):
def __init__(self):
super().__init__()
self._lock = threading.Lock()
self.llm_calls = 0
self.tool_calls = 0
self.input_tokens = 0
self.output_tokens = 0
def on_llm_start(self, *_, **__):
with self._lock:
self.llm_calls += 1
def on_chat_model_start(self, *_, **__):
with self._lock:
self.llm_calls += 1
def on_tool_start(self, *_, **__):
with self._lock:
self.tool_calls += 1
def get_stats(self):
with self._lock:
return {
"llm_calls": self.llm_calls,
"tool_calls": self.tool_calls,
"input_tokens": self.input_tokens,
"output_tokens": self.output_tokens
}
Thread-Safety Considerations
Because TradingAgentsGraph runs LangGraph in asynchronous streaming mode, the handler uses a threading.Lock to prevent race conditions when multiple analyst nodes (research, trader, risk) invoke LLMs concurrently. This makes the single handler instance safe to share across the entire graph.
Wiring Callbacks into the Pipeline
The callback is created once at the CLI layer, then propagated through three distinct integration points: the graph constructor, the LLM client factories, and the tool-execution propagator.
CLI Instantiation
In cli/main.py, the handler is instantiated before the graph is built. This single object is later passed into the TradingAgentsGraph constructor.
# cli/main.py
from cli.stats_handler import StatsCallbackHandler
from tradingagents.graph.trading_graph import TradingAgentsGraph
def run_analysis(selected_analyst_keys, config):
stats_handler = StatsCallbackHandler() # ← singleton observer
graph = TradingAgentsGraph(
selected_analyst_keys,
config=config,
debug=True,
callbacks=[stats_handler], # ← injection point
)
return graph, stats_handler
Graph Injection
The TradingAgentsGraph class (tradingagents/graph/trading_graph.py) accepts an optional callbacks list in its constructor (lines 51–63) and stores it as self.callbacks. By keeping the list as a graph property, the orchestrator can hand the same callback instance to every downstream component without coupling to specific LLM implementations.
# tradingagents/graph/trading_graph.py (simplified)
class TradingAgentsGraph:
def __init__(self, analysts, config, debug=False, callbacks=None):
self.callbacks = callbacks or []
self.config = config
# ... analyst initialization
LLM Client Attachment
When the graph builds its two reasoning tiers—deep_thinking_llm and quick_thinking_llm—it injects the callbacks into the keyword arguments consumed by create_llm_client. The factory (tradingagents/llm_clients/factory.py) forwards these kwargs to the concrete provider client (e.g., OpenAI, Anthropic).
# tradingagents/graph/trading_graph.py (lines 74-80)
llm_kwargs = {"temperature": 0.7, "model": self.config.model}
if self.callbacks:
llm_kwargs["callbacks"] = self.callbacks
deep_llm = create_llm_client(**llm_kwargs)
quick_llm = create_llm_client(**llm_kwargs, quick_mode=True)
Tool Execution Propagation
Tool calls are tracked separately from LLM calls. The Propagator class (tradingagents/graph/propagation.py) supplies the same callbacks list when invoking the LangGraph execution step (lines 56–67). This ensures that on_tool_start fires whenever an analyst node triggers a search, calculation, or portfolio update.
# tradingagents/graph/propagation.py (conceptual)
self.propagator = Propagator(
callbacks=self.callbacks # ensures tool nodes are observed
)
Displaying and Persisting Metrics
Once the pipeline is instrumented, you can query the handler at any point to surface metrics.
Live UI Rendering
The Rich-based terminal UI (cli/main.py) polls stats_handler.get_stats() inside the display-update loop and renders the counts in the footer. Because the handler is thread-safe, the UI thread can read counters while background graph nodes are still writing them.
# Inside the UI loop (cli/main.py)
current_stats = stats_handler.get_stats()
footer_text = (
f"LLM Calls: {current_stats['llm_calls']} | "
f"Tool Calls: {current_stats['tool_calls']}"
)
Post-Run Analysis
After the graph finishes, the same handler instance can be queried to persist data for offline analysis. The following snippet writes a CSV report containing the final tallies:
import csv
from pathlib import Path
def export_stats(path: Path, handler: StatsCallbackHandler):
stats = handler.get_stats()
with path.open("w", newline="") as f:
writer = csv.writer(f)
writer.writerow(["metric", "value"])
for key, val in stats.items():
writer.writerow([key, val])
Extending the System
Adding new metrics requires only subclassing additional hooks. For example, to track latency or error rates, override on_llm_error or on_tool_end in your custom handler:
class ExtendedStatsHandler(StatsCallbackHandler):
def on_llm_error(self, error, **kwargs):
with self._lock:
self.error_count += 1
def on_tool_end(self, output, **kwargs):
# Log tool result size or latency here
pass
Because the graph receives the callback list purely as a constructor argument, any provider-specific or use-case-specific handler works without changes to tradingagents/graph/trading_graph.py.
Summary
- Define a thread-safe handler by subclassing
BaseCallbackHandlerand overridingon_llm_start,on_chat_model_start,on_llm_end, andon_tool_startincli/stats_handler.py. - Instantiate once in
cli/main.pyand pass the list intoTradingAgentsGraphvia thecallbacksparameter. - Attach automatically to all LLM clients when
tradingagents/graph/trading_graph.pybuildsdeep_thinking_llmandquick_thinking_llmthrough the factory intradingagents/llm_clients/factory.py. - Propagate to tools via the
Propagatorintradingagents/graph/propagation.pyso that function calls are counted alongside LLM invocations. - Query anytime using
get_stats()to drive live UIs or export post-run reports.
Frequently Asked Questions
How do I track token usage rather than just call counts?
Override on_llm_end in your handler and inspect the response object’s usage_metadata. The TradingAgents reference implementation stores response.usage_metadata.get("input_tokens") and output_tokens into counters protected by the same threading.Lock used for call counts.
Can I use multiple callback handlers simultaneously?
Yes. The TradingAgentsGraph constructor accepts a list (callbacks=[handler1, handler2]), and LangChain will invoke every hook on every handler in the order provided. This lets you separate concerns—for example, one handler for metrics and another for logging—without merging code.
What happens if a tool fails? Will on_tool_start still fire?
on_tool_start fires immediately before the tool executes, so it will increment regardless of success. To track failures, also override on_tool_error (or on_tool_end and check the output). The handler in cli/stats_handler.py can be extended with an error_count field and locked updates inside on_tool_error.
Is the callback system compatible with async streaming?
Yes. The threading.Lock inside StatsCallbackHandler makes the counters safe for LangGraph’s asynchronous streaming mode. The handlers are invoked in the event loop, but the lock ensures atomic increments when multiple analyst nodes run concurrently.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →