# How to Run Synthetic Tests in OpenSRE: A Complete Guide

> Learn to run synthetic tests in OpenSRE with the CLI or Python module. Evaluate agent performance against static fixtures, streamlining your testing process.

- Repository: [Tracer/opensre](https://github.com/Tracer-Cloud/opensre)
- Tags: how-to-guide
- Published: 2026-04-18

---

**Run synthetic tests in OpenSRE using the CLI command `opensre tests synthetic` or the Python module `tests.synthetic.rds_postgres.run_suite` with the `--mock-grafana` flag to evaluate agent performance against static fixtures without live infrastructure.**

OpenSRE provides a fully automated **synthetic root-cause analysis (RCA) suite** that validates an agent's diagnostic capabilities using pre-built scenarios. This system allows you to benchmark changes to LLM prompts, tool registries, and evidence adapters against deterministic test fixtures—no cloud resources required.

## Understanding the Synthetic Test Architecture

The synthetic testing framework consists of four primary components that work together to simulate production investigations.

### FixtureGrafanaBackend

The **FixtureGrafanaBackend** is a mock Grafana client that serves static metric files from a scenario directory instead of querying live data sources. Located in [`tests/synthetic/mock_grafana_backend/backend.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/synthetic/mock_grafana_backend/backend.py), this component provides the identical API that the real Grafana integration expects, allowing the production pipeline to run unmodified against synthetic data.

### Scenario Loader

The **Scenario Loader**, implemented in [`tests/synthetic/rds_postgres/scenario_loader.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/synthetic/rds_postgres/scenario_loader.py), handles test discovery and preparation. This utility discovers all scenario directories, resolves optional base-scenario inheritance, validates JSON and YAML files, and constructs a `ScenarioFixture` object containing the alert payload, evidence sets, metadata, and answer key required for the test.

### Test Runner and Scoring

The **run_suite.py** file in [`tests/synthetic/rds_postgres/run_suite.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/synthetic/rds_postgres/run_suite.py) serves as the command-line driver and orchestrator. This script iterates over fixtures, injects the mock backend when `--mock-grafana` is set, invokes `run_investigation` from [`app/pipeline/runners.py`](https://github.com/Tracer-Cloud/opensre/blob/main/app/pipeline/runners.py) (the production pipeline entry point), and evaluates results using three scoring functions: `score_result`, `score_trajectory`, and `score_reasoning`. The scoring logic compares agent output against the expected [`answer.yml`](https://github.com/Tracer-Cloud/opensre/blob/main/answer.yml), checking root-cause category, required keywords, evidence sources, and investigation trajectory.

## How to Run Synthetic Tests in OpenSRE

You can execute the synthetic suite through multiple interfaces depending on your CI/CD requirements or debugging needs.

### Using the CLI Shortcut

The fastest method uses the built-in OpenSRE CLI, which automatically enables the mock Grafana backend:

```bash
opensre tests synthetic

```

This command maps to [`opensre/tests/cli.py`](https://github.com/Tracer-Cloud/opensre/blob/main/opensre/tests/cli.py) and ultimately executes `run_suite` with `--mock-grafana` enabled by default, making it ideal for local development and quick validation.

### Running via Python Module

For CI scripts or when you need custom flags, invoke the Python module directly:

```bash
python -m tests.synthetic.rds_postgres.run_suite \
    --mock-grafana \
    --json

```

The `--mock-grafana` flag is essential for CI environments to serve static Grafana data from fixtures rather than attempting to connect to live instances. The `--json` flag emits machine-readable JSON results suitable for automated reporting and dashboards.

### Executing Specific Scenarios

To debug a single test case rather than the entire suite, use the `--scenario` flag:

```bash
python -m tests.synthetic.rds_postgres.run_suite \
    --scenario 006-replication-lag-cpu-redherring \
    --mock-grafana

```

This approach runs only the specified scenario directory, allowing rapid iteration when adjusting prompt engineering for specific failure modes like replication lag with CPU red herrings.

### Running Adversarial Axis 2 Tests

The suite includes adversarial "Axis 2" tests that verify the agent's ability to avoid false positives and handle edge cases. These require the selective backend and run via pytest:

```bash
pytest -m axis2 tests/synthetic/rds_postgres/test_suite_axis2.py -v

```

These tests validate that the agent does not hallucinate root causes when presented with ambiguous or confounding evidence.

## Understanding Test Results

After execution, the runner outputs a per-scenario PASS/FAIL line followed by a summary:

```

PASS 001-healthy category=healthy
FAIL 006-replication-lag-cpu-redherring reason='missing required keywords: [...]'
...
Results: 9/10 passed

```

A **PASS** indicates the agent correctly identified the root cause category, included all required keywords from [`answer.yml`](https://github.com/Tracer-Cloud/opensre/blob/main/answer.yml), and cited the correct evidence sources. A **FAIL** provides the specific discrepancy, such as missing keywords or incorrect trajectory, enabling precise debugging of the investigation pipeline.

## Summary

- **Run synthetic tests in OpenSRE** using `opensre tests synthetic` for quick checks or `python -m tests.synthetic.rds_postgres.run_suite --mock-grafana` for CI pipelines.
- The **FixtureGrafanaBackend** in [`tests/synthetic/mock_grafana_backend/backend.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/synthetic/mock_grafana_backend/backend.py) enables testing without live Grafana instances by serving static fixture data.
- **Scenarios** are loaded by [`tests/synthetic/rds_postgres/scenario_loader.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/synthetic/rds_postgres/scenario_loader.py) and evaluated against [`answer.yml`](https://github.com/Tracer-Cloud/opensre/blob/main/answer.yml) keys using the scoring functions in [`run_suite.py`](https://github.com/Tracer-Cloud/opensre/blob/main/run_suite.py).
- Use `--scenario` to isolate specific test cases and `pytest -m axis2` to execute adversarial validation suites.

## Frequently Asked Questions

### What is a synthetic test in OpenSRE?

A synthetic test in OpenSRE is an automated evaluation that feeds the investigation pipeline a pre-constructed fixture containing a static alert and pre-generated evidence. This allows you to verify that the agent correctly diagnoses specific failure modes—such as replication lag or CPU saturation—without requiring access to live RDS instances or Grafana servers.

### Do I need a live Grafana instance to run synthetic tests?

No. When you run synthetic tests with the `--mock-grafana` flag, the `FixtureGrafanaBackend` intercepts Grafana API calls and returns static metric files from the scenario directory. This mock backend is located in [`tests/synthetic/mock_grafana_backend/backend.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/synthetic/mock_grafana_backend/backend.py) and is essential for running tests in CI environments without cloud credentials.

### How do I add a new synthetic test scenario?

Create a new directory under `tests/synthetic/rds_postgres/scenarios/` containing an [`alert.json`](https://github.com/Tracer-Cloud/opensre/blob/main/alert.json) (the triggering alert), an `evidence/` directory with static data files (metrics, events), and an [`answer.yml`](https://github.com/Tracer-Cloud/opensre/blob/main/answer.yml) (the expected root cause and keywords). The [`scenario_loader.py`](https://github.com/Tracer-Cloud/opensre/blob/main/scenario_loader.py) automatically discovers new directories and validates the schema, making them available to [`run_suite.py`](https://github.com/Tracer-Cloud/opensre/blob/main/run_suite.py) immediately.

### What is the difference between regular and Axis 2 tests?

Regular synthetic tests verify that the agent correctly identifies known root causes when presented with clear evidence. **Axis 2 tests** are adversarial validations that ensure the agent does not hallucinate causes or fall for red herrings (such as high CPU usage during a replication lag incident). These run via `pytest -m axis2` and use selective backend configurations to test the agent's reasoning boundaries.