Running Backtests and Evaluating Performance from eval_results Logs in TradingAgents
Backtest the TradingAgents LLM-driven trading system by calling TradingAgentsGraph.propagate() for each target date, then analyze the automatically generated JSON logs in eval_results/<TICKER>/TradingAgentsStrategy_logs/ to calculate returns, win rates, and equity curves.
The TradingAgents repository by TauricResearch implements a multi-agent research pipeline that generates trading decisions through simulated bull/bear debates and risk analysis. After each run, the framework persists the complete agent state—including debate transcripts, risk assessments, and final trade decisions—as structured JSON logs, enabling rigorous quantitative backtesting and performance evaluation without rerunning the LLM inference.
How Backtests Are Executed in TradingAgents
Launching Single-Date Propagation
The core entry point for any backtest is the propagate method in tradingagents/graph/trading_graph.py. This method initializes the LangGraph workflow, executes the multi-agent pipeline, and automatically writes the final state to disk.
from tradingagents.graph.trading_graph import TradingAgentsGraph
graph = TradingAgentsGraph(debug=False) # Creates LLM clients, tool nodes, memories
final_state, signal = graph.propagate(
company_name="NVDA", # Ticker symbol
trade_date="2024-05-10", # YYYY-MM-DD format
)
According to the source code in trading_graph.py, the propagate method handles graph creation (lines 94-99), executes the workflow (lines 194-226), and triggers state persistence (lines 229-239). The method returns both the final AgentState dictionary and a string signal representing the trade decision.
Batch Backtesting Over Date Ranges
For historical analysis across multiple dates, wrap the single-date call in a loop. The graph automatically handles log directory creation on the first invocation.
from datetime import datetime, timedelta
from tradingagents.graph.trading_graph import TradingAgentsGraph
def backtest(ticker: str, start: str, end: str):
g = TradingAgentsGraph(debug=False)
start_dt = datetime.fromisoformat(start)
end_dt = datetime.fromisoformat(end)
cur = start_dt
while cur <= end_dt:
g.propagate(ticker, cur.date().isoformat())
cur += timedelta(days=1) # Daily backtest; adjust for market holidays as needed
The _log_state method (lines 262-270 in trading_graph.py) automatically creates the eval_results/<ticker>/TradingAgentsStrategy_logs/ directory structure and appends each day's state to a cumulative JSON file.
Understanding the eval_results JSON Log Structure
Each file (full_states_log_<YYYY-MM-DD>.json) stores a snapshot of the entire graph state. The top-level keys correspond directly to the AgentState Pydantic model defined in tradingagents/agents/utils/agent_states.py.
Key Data Fields for Performance Analysis
When evaluating backtests, focus on these specific fields within each log entry:
company_of_interest– The ticker symbol processedtrade_date– The ISO date of the backtestmarket_report– Raw market data including closing prices (viaget_stock_datainagent_utils.py)investment_debate_state– Complete bull/bear debate history and the judge's final decisionfinal_trade_decision– The Buy/Sell/Hold signal driving your backtest logicrisk_debate_state– Aggressive/conservative/neutral risk assessment that can weight position sizingtrader_investment_decision– The structured allocation plan proposed before risk review
The JSON serialization occurs in _log_state via json.dump(self.log_states_dict, ...), preserving the exact internal state for reproducible analysis.
Post-Processing Workflow: From Logs to Performance Metrics
Loading and Parsing Log Files
Access the persisted states by walking the eval_results directory structure. The following loader aggregates all daily snapshots for a given ticker:
import json
import pathlib
from datetime import datetime
LOG_ROOT = pathlib.Path("eval_results") / "NVDA" / "TradingAgentsStrategy_logs"
def load_logs() -> list[dict]:
"""Read every JSON log for the ticker and return a list ordered by date."""
logs = []
for file in sorted(LOG_ROOT.glob("full_states_log_*.json")):
with open(file, "r", encoding="utf-8") as f:
logs.append(json.load(f))
return logs
Computing Returns and Win Rates
Extract the final_trade_decision and compare against next-day price movements from the market_report field. This example implements a binary win/loss calculation based on directional accuracy:
def compute_returns(logs: list[dict]) -> list[float]:
"""Simple return: +1 for correct Buy/Sell direction, 0 otherwise."""
returns = []
for log in logs:
close_price = float(log["market_report"]["close"])
decision = log["final_trade_decision"].lower()
next_idx = logs.index(log) + 1
if next_idx < len(logs):
next_close = float(logs[next_idx]["market_report"]["close"])
price_up = next_close > close_price
if (decision == "buy" and price_up) or (decision == "sell" and not price_up):
returns.append(1.0) # Correct prediction
else:
returns.append(0.0) # Incorrect or hold
else:
returns.append(0.0) # No forward data for last day
return returns
Visualizing Equity Curves
Convert the returns list into a cumulative equity curve using matplotlib:
import matplotlib.pyplot as plt
def plot_equity(returns: list[float]):
"""Cumulative equity curve assuming 1 unit risk per trade."""
equity = [0]
for r in returns:
equity.append(equity[-1] + r)
plt.figure(figsize=(10, 4))
plt.plot(equity, label="Equity Curve")
plt.title("Backtest Equity Curve – NVDA")
plt.xlabel("Trade #")
plt.ylabel("Cumulative Wins")
plt.legend()
plt.grid(True)
plt.show()
if __name__ == "__main__":
logs = load_logs()
returns = compute_returns(logs)
win_rate = sum(returns) / len(returns) * 100
print(f"Backtest performed on {len(logs)} days – win rate: {win_rate:.2f}%")
plot_equity(returns)
This workflow leverages the market_report structure returned by get_stock_data in tradingagents/agents/utils/agent_utils.py, ensuring your price data matches exactly what the agents saw during inference.
Extending Your Evaluation Pipeline
Beyond basic win-rate calculations, the eval_results logs support sophisticated analysis:
-
Risk-Weighted Position Sizing – Use the
risk_debate_statefield (containing the judge's confidence scores) to scale position sizes. High-confidence aggressive signals can receive larger allocations than conservative holds. -
Memory Integration – Feed realized returns back into the agents using the
reflect_and_remembermethod (lines 272-288 intrading_graph.py). This enables continuous learning where the LLM-driven agents adapt strategies based on backtest outcomes. -
Parallel Multi-Ticker Analysis – Iterate over multiple subdirectories under
eval_results/using multiprocessing. EachTradingAgentsGraphinstance operates independently, allowing simultaneous backtesting across an entire watchlist. -
Pandas DataFrame Conversion – Convert the logs list to a DataFrame for advanced metrics (Sharpe ratio, maximum drawdown, Calmar ratio) using the structured fields from
AgentState.
Summary
- Execute backtests by calling
TradingAgentsGraph.propagate(ticker, date)for each historical date. - Results automatically serialize to
eval_results/<TICKER>/TradingAgentsStrategy_logs/full_states_log_<DATE>.json. - Each JSON contains the complete
AgentState, including debate transcripts, risk assessments, and the final buy/sell/hold decision. - Calculate performance by comparing
final_trade_decisionagainst next-day price movements frommarket_report. - Extend analysis by incorporating
risk_debate_stateconfidence scores and feeding returns into thereflect_and_remembermemory system.
Frequently Asked Questions
How do I run a backtest for a specific date range in TradingAgents?
Initialize a TradingAgentsGraph instance and loop through your target dates, calling propagate(company_name, trade_date) for each day. The framework automatically creates the eval_results directory structure on the first call. You can skip weekends and holidays by adding conditional logic to your date loop before invoking propagation.
What information is stored in the eval_results JSON files?
Each log file contains a complete snapshot of the AgentState model, including the market_report (price data), investment_debate_state (bull/bear arguments), risk_debate_state (risk assessment), investment_plan (structured allocation), and final_trade_decision (the executable signal). These fields map directly to the Pydantic models in tradingagents/agents/utils/agent_states.py.
How can I calculate realistic P&L from the final_trade_decision logs?
Extract the close price from market_report to establish entry prices, then compare against subsequent closing prices to determine trade outcomes. For realistic P&L, modify the compute_returns function to account for position sizing (using risk_debate_state confidence as a weight), transaction costs, and slippage rather than using the binary win/loss example provided.
Can I use the debate transcripts to weight trade decisions in backtests?
Yes. The investment_debate_state field contains the complete debate history and judge decision. You can parse the confidence scores or reasoning strength within the debate transcript to create position-size multipliers. For example, allocate 2x capital when the bull/bear consensus is unanimous versus 0.5x when the debate is split, enhancing the risk-adjusted returns of your backtest strategy.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →