Running Backtests and Evaluating Performance from eval_results Logs in TradingAgents

Backtest the TradingAgents LLM-driven trading system by calling TradingAgentsGraph.propagate() for each target date, then analyze the automatically generated JSON logs in eval_results/<TICKER>/TradingAgentsStrategy_logs/ to calculate returns, win rates, and equity curves.

The TradingAgents repository by TauricResearch implements a multi-agent research pipeline that generates trading decisions through simulated bull/bear debates and risk analysis. After each run, the framework persists the complete agent state—including debate transcripts, risk assessments, and final trade decisions—as structured JSON logs, enabling rigorous quantitative backtesting and performance evaluation without rerunning the LLM inference.

How Backtests Are Executed in TradingAgents

Launching Single-Date Propagation

The core entry point for any backtest is the propagate method in tradingagents/graph/trading_graph.py. This method initializes the LangGraph workflow, executes the multi-agent pipeline, and automatically writes the final state to disk.

from tradingagents.graph.trading_graph import TradingAgentsGraph

graph = TradingAgentsGraph(debug=False)          # Creates LLM clients, tool nodes, memories

final_state, signal = graph.propagate(
    company_name="NVDA",                         # Ticker symbol

    trade_date="2024-05-10",                     # YYYY-MM-DD format

)

According to the source code in trading_graph.py, the propagate method handles graph creation (lines 94-99), executes the workflow (lines 194-226), and triggers state persistence (lines 229-239). The method returns both the final AgentState dictionary and a string signal representing the trade decision.

Batch Backtesting Over Date Ranges

For historical analysis across multiple dates, wrap the single-date call in a loop. The graph automatically handles log directory creation on the first invocation.

from datetime import datetime, timedelta
from tradingagents.graph.trading_graph import TradingAgentsGraph

def backtest(ticker: str, start: str, end: str):
    g = TradingAgentsGraph(debug=False)
    start_dt = datetime.fromisoformat(start)
    end_dt = datetime.fromisoformat(end)
    
    cur = start_dt
    while cur <= end_dt:
        g.propagate(ticker, cur.date().isoformat())
        cur += timedelta(days=1)   # Daily backtest; adjust for market holidays as needed

The _log_state method (lines 262-270 in trading_graph.py) automatically creates the eval_results/<ticker>/TradingAgentsStrategy_logs/ directory structure and appends each day's state to a cumulative JSON file.

Understanding the eval_results JSON Log Structure

Each file (full_states_log_<YYYY-MM-DD>.json) stores a snapshot of the entire graph state. The top-level keys correspond directly to the AgentState Pydantic model defined in tradingagents/agents/utils/agent_states.py.

Key Data Fields for Performance Analysis

When evaluating backtests, focus on these specific fields within each log entry:

  • company_of_interest – The ticker symbol processed
  • trade_date – The ISO date of the backtest
  • market_report – Raw market data including closing prices (via get_stock_data in agent_utils.py)
  • investment_debate_state – Complete bull/bear debate history and the judge's final decision
  • final_trade_decision – The Buy/Sell/Hold signal driving your backtest logic
  • risk_debate_state – Aggressive/conservative/neutral risk assessment that can weight position sizing
  • trader_investment_decision – The structured allocation plan proposed before risk review

The JSON serialization occurs in _log_state via json.dump(self.log_states_dict, ...), preserving the exact internal state for reproducible analysis.

Post-Processing Workflow: From Logs to Performance Metrics

Loading and Parsing Log Files

Access the persisted states by walking the eval_results directory structure. The following loader aggregates all daily snapshots for a given ticker:

import json
import pathlib
from datetime import datetime

LOG_ROOT = pathlib.Path("eval_results") / "NVDA" / "TradingAgentsStrategy_logs"

def load_logs() -> list[dict]:
    """Read every JSON log for the ticker and return a list ordered by date."""
    logs = []
    for file in sorted(LOG_ROOT.glob("full_states_log_*.json")):
        with open(file, "r", encoding="utf-8") as f:
            logs.append(json.load(f))
    return logs

Computing Returns and Win Rates

Extract the final_trade_decision and compare against next-day price movements from the market_report field. This example implements a binary win/loss calculation based on directional accuracy:

def compute_returns(logs: list[dict]) -> list[float]:
    """Simple return: +1 for correct Buy/Sell direction, 0 otherwise."""
    returns = []
    for log in logs:
        close_price = float(log["market_report"]["close"])
        decision = log["final_trade_decision"].lower()
        
        next_idx = logs.index(log) + 1
        if next_idx < len(logs):
            next_close = float(logs[next_idx]["market_report"]["close"])
            price_up = next_close > close_price
            
            if (decision == "buy" and price_up) or (decision == "sell" and not price_up):
                returns.append(1.0)          # Correct prediction

            else:
                returns.append(0.0)          # Incorrect or hold

        else:
            returns.append(0.0)              # No forward data for last day

    return returns

Visualizing Equity Curves

Convert the returns list into a cumulative equity curve using matplotlib:

import matplotlib.pyplot as plt

def plot_equity(returns: list[float]):
    """Cumulative equity curve assuming 1 unit risk per trade."""
    equity = [0]
    for r in returns:
        equity.append(equity[-1] + r)
    
    plt.figure(figsize=(10, 4))
    plt.plot(equity, label="Equity Curve")
    plt.title("Backtest Equity Curve – NVDA")
    plt.xlabel("Trade #")
    plt.ylabel("Cumulative Wins")
    plt.legend()
    plt.grid(True)
    plt.show()

if __name__ == "__main__":
    logs = load_logs()
    returns = compute_returns(logs)
    win_rate = sum(returns) / len(returns) * 100
    print(f"Backtest performed on {len(logs)} days – win rate: {win_rate:.2f}%")
    plot_equity(returns)

This workflow leverages the market_report structure returned by get_stock_data in tradingagents/agents/utils/agent_utils.py, ensuring your price data matches exactly what the agents saw during inference.

Extending Your Evaluation Pipeline

Beyond basic win-rate calculations, the eval_results logs support sophisticated analysis:

  • Risk-Weighted Position Sizing – Use the risk_debate_state field (containing the judge's confidence scores) to scale position sizes. High-confidence aggressive signals can receive larger allocations than conservative holds.

  • Memory Integration – Feed realized returns back into the agents using the reflect_and_remember method (lines 272-288 in trading_graph.py). This enables continuous learning where the LLM-driven agents adapt strategies based on backtest outcomes.

  • Parallel Multi-Ticker Analysis – Iterate over multiple subdirectories under eval_results/ using multiprocessing. Each TradingAgentsGraph instance operates independently, allowing simultaneous backtesting across an entire watchlist.

  • Pandas DataFrame Conversion – Convert the logs list to a DataFrame for advanced metrics (Sharpe ratio, maximum drawdown, Calmar ratio) using the structured fields from AgentState.

Summary

  • Execute backtests by calling TradingAgentsGraph.propagate(ticker, date) for each historical date.
  • Results automatically serialize to eval_results/<TICKER>/TradingAgentsStrategy_logs/full_states_log_<DATE>.json.
  • Each JSON contains the complete AgentState, including debate transcripts, risk assessments, and the final buy/sell/hold decision.
  • Calculate performance by comparing final_trade_decision against next-day price movements from market_report.
  • Extend analysis by incorporating risk_debate_state confidence scores and feeding returns into the reflect_and_remember memory system.

Frequently Asked Questions

How do I run a backtest for a specific date range in TradingAgents?

Initialize a TradingAgentsGraph instance and loop through your target dates, calling propagate(company_name, trade_date) for each day. The framework automatically creates the eval_results directory structure on the first call. You can skip weekends and holidays by adding conditional logic to your date loop before invoking propagation.

What information is stored in the eval_results JSON files?

Each log file contains a complete snapshot of the AgentState model, including the market_report (price data), investment_debate_state (bull/bear arguments), risk_debate_state (risk assessment), investment_plan (structured allocation), and final_trade_decision (the executable signal). These fields map directly to the Pydantic models in tradingagents/agents/utils/agent_states.py.

How can I calculate realistic P&L from the final_trade_decision logs?

Extract the close price from market_report to establish entry prices, then compare against subsequent closing prices to determine trade outcomes. For realistic P&L, modify the compute_returns function to account for position sizing (using risk_debate_state confidence as a weight), transaction costs, and slippage rather than using the binary win/loss example provided.

Can I use the debate transcripts to weight trade decisions in backtests?

Yes. The investment_debate_state field contains the complete debate history and judge decision. You can parse the confidence scores or reasoning strength within the debate transcript to create position-size multipliers. For example, allocate 2x capital when the bull/bear consensus is unanimous versus 0.5x when the debate is split, enhancing the risk-adjusted returns of your backtest strategy.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →