# Running Backtests and Evaluating Performance from eval_results Logs in TradingAgents

> Backtest TradingAgents LLM trading system and analyze eval_results logs to calculate performance metrics like returns, win rates, and equity curves. Learn how to run and evaluate backtests.

- Repository: [Tauric Research/TradingAgents](https://github.com/TauricResearch/TradingAgents)
- Tags: how-to-guide
- Published: 2026-03-23

---

**Backtest the TradingAgents LLM-driven trading system by calling `TradingAgentsGraph.propagate()` for each target date, then analyze the automatically generated JSON logs in `eval_results/<TICKER>/TradingAgentsStrategy_logs/` to calculate returns, win rates, and equity curves.**

The TradingAgents repository by TauricResearch implements a multi-agent research pipeline that generates trading decisions through simulated bull/bear debates and risk analysis. After each run, the framework persists the complete agent state—including debate transcripts, risk assessments, and final trade decisions—as structured JSON logs, enabling rigorous quantitative backtesting and performance evaluation without rerunning the LLM inference.

## How Backtests Are Executed in TradingAgents

### Launching Single-Date Propagation

The core entry point for any backtest is the `propagate` method in [`tradingagents/graph/trading_graph.py`](https://github.com/TauricResearch/TradingAgents/blob/main/tradingagents/graph/trading_graph.py). This method initializes the LangGraph workflow, executes the multi-agent pipeline, and automatically writes the final state to disk.

```python
from tradingagents.graph.trading_graph import TradingAgentsGraph

graph = TradingAgentsGraph(debug=False)          # Creates LLM clients, tool nodes, memories

final_state, signal = graph.propagate(
    company_name="NVDA",                         # Ticker symbol

    trade_date="2024-05-10",                     # YYYY-MM-DD format

)

```

According to the source code in [`trading_graph.py`](https://github.com/TauricResearch/TradingAgents/blob/main/trading_graph.py), the `propagate` method handles graph creation (lines 94-99), executes the workflow (lines 194-226), and triggers state persistence (lines 229-239). The method returns both the final `AgentState` dictionary and a string signal representing the trade decision.

### Batch Backtesting Over Date Ranges

For historical analysis across multiple dates, wrap the single-date call in a loop. The graph automatically handles log directory creation on the first invocation.

```python
from datetime import datetime, timedelta
from tradingagents.graph.trading_graph import TradingAgentsGraph

def backtest(ticker: str, start: str, end: str):
    g = TradingAgentsGraph(debug=False)
    start_dt = datetime.fromisoformat(start)
    end_dt = datetime.fromisoformat(end)
    
    cur = start_dt
    while cur <= end_dt:
        g.propagate(ticker, cur.date().isoformat())
        cur += timedelta(days=1)   # Daily backtest; adjust for market holidays as needed

```

The `_log_state` method (lines 262-270 in [`trading_graph.py`](https://github.com/TauricResearch/TradingAgents/blob/main/trading_graph.py)) automatically creates the `eval_results/<ticker>/TradingAgentsStrategy_logs/` directory structure and appends each day's state to a cumulative JSON file.

## Understanding the eval_results JSON Log Structure

Each file (`full_states_log_<YYYY-MM-DD>.json`) stores a **snapshot** of the entire graph state. The top-level keys correspond directly to the `AgentState` Pydantic model defined in [`tradingagents/agents/utils/agent_states.py`](https://github.com/TauricResearch/TradingAgents/blob/main/tradingagents/agents/utils/agent_states.py).

### Key Data Fields for Performance Analysis

When evaluating backtests, focus on these specific fields within each log entry:

- **`company_of_interest`** – The ticker symbol processed
- **`trade_date`** – The ISO date of the backtest
- **`market_report`** – Raw market data including closing prices (via `get_stock_data` in [`agent_utils.py`](https://github.com/TauricResearch/TradingAgents/blob/main/agent_utils.py))
- **`investment_debate_state`** – Complete bull/bear debate history and the judge's final decision
- **`final_trade_decision`** – The **Buy/Sell/Hold** signal driving your backtest logic
- **`risk_debate_state`** – Aggressive/conservative/neutral risk assessment that can weight position sizing
- **`trader_investment_decision`** – The structured allocation plan proposed before risk review

The JSON serialization occurs in `_log_state` via `json.dump(self.log_states_dict, ...)`, preserving the exact internal state for reproducible analysis.

## Post-Processing Workflow: From Logs to Performance Metrics

### Loading and Parsing Log Files

Access the persisted states by walking the `eval_results` directory structure. The following loader aggregates all daily snapshots for a given ticker:

```python
import json
import pathlib
from datetime import datetime

LOG_ROOT = pathlib.Path("eval_results") / "NVDA" / "TradingAgentsStrategy_logs"

def load_logs() -> list[dict]:
    """Read every JSON log for the ticker and return a list ordered by date."""
    logs = []
    for file in sorted(LOG_ROOT.glob("full_states_log_*.json")):
        with open(file, "r", encoding="utf-8") as f:
            logs.append(json.load(f))
    return logs

```

### Computing Returns and Win Rates

Extract the `final_trade_decision` and compare against next-day price movements from the `market_report` field. This example implements a binary win/loss calculation based on directional accuracy:

```python
def compute_returns(logs: list[dict]) -> list[float]:
    """Simple return: +1 for correct Buy/Sell direction, 0 otherwise."""
    returns = []
    for log in logs:
        close_price = float(log["market_report"]["close"])
        decision = log["final_trade_decision"].lower()
        
        next_idx = logs.index(log) + 1
        if next_idx < len(logs):
            next_close = float(logs[next_idx]["market_report"]["close"])
            price_up = next_close > close_price
            
            if (decision == "buy" and price_up) or (decision == "sell" and not price_up):
                returns.append(1.0)          # Correct prediction

            else:
                returns.append(0.0)          # Incorrect or hold

        else:
            returns.append(0.0)              # No forward data for last day

    return returns

```

### Visualizing Equity Curves

Convert the returns list into a cumulative equity curve using **matplotlib**:

```python
import matplotlib.pyplot as plt

def plot_equity(returns: list[float]):
    """Cumulative equity curve assuming 1 unit risk per trade."""
    equity = [0]
    for r in returns:
        equity.append(equity[-1] + r)
    
    plt.figure(figsize=(10, 4))
    plt.plot(equity, label="Equity Curve")
    plt.title("Backtest Equity Curve – NVDA")
    plt.xlabel("Trade #")
    plt.ylabel("Cumulative Wins")
    plt.legend()
    plt.grid(True)
    plt.show()

if __name__ == "__main__":
    logs = load_logs()
    returns = compute_returns(logs)
    win_rate = sum(returns) / len(returns) * 100
    print(f"Backtest performed on {len(logs)} days – win rate: {win_rate:.2f}%")
    plot_equity(returns)

```

This workflow leverages the `market_report` structure returned by `get_stock_data` in [`tradingagents/agents/utils/agent_utils.py`](https://github.com/TauricResearch/TradingAgents/blob/main/tradingagents/agents/utils/agent_utils.py), ensuring your price data matches exactly what the agents saw during inference.

## Extending Your Evaluation Pipeline

Beyond basic win-rate calculations, the `eval_results` logs support sophisticated analysis:

- **Risk-Weighted Position Sizing** – Use the `risk_debate_state` field (containing the judge's confidence scores) to scale position sizes. High-confidence aggressive signals can receive larger allocations than conservative holds.

- **Memory Integration** – Feed realized returns back into the agents using the `reflect_and_remember` method (lines 272-288 in [`trading_graph.py`](https://github.com/TauricResearch/TradingAgents/blob/main/trading_graph.py)). This enables continuous learning where the LLM-driven agents adapt strategies based on backtest outcomes.

- **Parallel Multi-Ticker Analysis** – Iterate over multiple subdirectories under `eval_results/` using multiprocessing. Each `TradingAgentsGraph` instance operates independently, allowing simultaneous backtesting across an entire watchlist.

- **Pandas DataFrame Conversion** – Convert the logs list to a DataFrame for advanced metrics (Sharpe ratio, maximum drawdown, Calmar ratio) using the structured fields from `AgentState`.

## Summary

- Execute backtests by calling `TradingAgentsGraph.propagate(ticker, date)` for each historical date.
- Results automatically serialize to `eval_results/<TICKER>/TradingAgentsStrategy_logs/full_states_log_<DATE>.json`.
- Each JSON contains the complete `AgentState`, including debate transcripts, risk assessments, and the final buy/sell/hold decision.
- Calculate performance by comparing `final_trade_decision` against next-day price movements from `market_report`.
- Extend analysis by incorporating `risk_debate_state` confidence scores and feeding returns into the `reflect_and_remember` memory system.

## Frequently Asked Questions

### How do I run a backtest for a specific date range in TradingAgents?

Initialize a `TradingAgentsGraph` instance and loop through your target dates, calling `propagate(company_name, trade_date)` for each day. The framework automatically creates the `eval_results` directory structure on the first call. You can skip weekends and holidays by adding conditional logic to your date loop before invoking propagation.

### What information is stored in the eval_results JSON files?

Each log file contains a complete snapshot of the `AgentState` model, including the `market_report` (price data), `investment_debate_state` (bull/bear arguments), `risk_debate_state` (risk assessment), `investment_plan` (structured allocation), and `final_trade_decision` (the executable signal). These fields map directly to the Pydantic models in [`tradingagents/agents/utils/agent_states.py`](https://github.com/TauricResearch/TradingAgents/blob/main/tradingagents/agents/utils/agent_states.py).

### How can I calculate realistic P&L from the final_trade_decision logs?

Extract the `close` price from `market_report` to establish entry prices, then compare against subsequent closing prices to determine trade outcomes. For realistic P&L, modify the `compute_returns` function to account for position sizing (using `risk_debate_state` confidence as a weight), transaction costs, and slippage rather than using the binary win/loss example provided.

### Can I use the debate transcripts to weight trade decisions in backtests?

Yes. The `investment_debate_state` field contains the complete debate history and judge decision. You can parse the confidence scores or reasoning strength within the debate transcript to create position-size multipliers. For example, allocate 2x capital when the bull/bear consensus is unanimous versus 0.5x when the debate is split, enhancing the risk-adjusted returns of your backtest strategy.