How to Run Synthetic Tests in OpenSRE: A Complete Guide
Run synthetic tests in OpenSRE using the CLI command opensre tests synthetic or the Python module tests.synthetic.rds_postgres.run_suite with the --mock-grafana flag to evaluate agent performance against static fixtures without live infrastructure.
OpenSRE provides a fully automated synthetic root-cause analysis (RCA) suite that validates an agent's diagnostic capabilities using pre-built scenarios. This system allows you to benchmark changes to LLM prompts, tool registries, and evidence adapters against deterministic test fixtures—no cloud resources required.
Understanding the Synthetic Test Architecture
The synthetic testing framework consists of four primary components that work together to simulate production investigations.
FixtureGrafanaBackend
The FixtureGrafanaBackend is a mock Grafana client that serves static metric files from a scenario directory instead of querying live data sources. Located in tests/synthetic/mock_grafana_backend/backend.py, this component provides the identical API that the real Grafana integration expects, allowing the production pipeline to run unmodified against synthetic data.
Scenario Loader
The Scenario Loader, implemented in tests/synthetic/rds_postgres/scenario_loader.py, handles test discovery and preparation. This utility discovers all scenario directories, resolves optional base-scenario inheritance, validates JSON and YAML files, and constructs a ScenarioFixture object containing the alert payload, evidence sets, metadata, and answer key required for the test.
Test Runner and Scoring
The run_suite.py file in tests/synthetic/rds_postgres/run_suite.py serves as the command-line driver and orchestrator. This script iterates over fixtures, injects the mock backend when --mock-grafana is set, invokes run_investigation from app/pipeline/runners.py (the production pipeline entry point), and evaluates results using three scoring functions: score_result, score_trajectory, and score_reasoning. The scoring logic compares agent output against the expected answer.yml, checking root-cause category, required keywords, evidence sources, and investigation trajectory.
How to Run Synthetic Tests in OpenSRE
You can execute the synthetic suite through multiple interfaces depending on your CI/CD requirements or debugging needs.
Using the CLI Shortcut
The fastest method uses the built-in OpenSRE CLI, which automatically enables the mock Grafana backend:
opensre tests synthetic
This command maps to opensre/tests/cli.py and ultimately executes run_suite with --mock-grafana enabled by default, making it ideal for local development and quick validation.
Running via Python Module
For CI scripts or when you need custom flags, invoke the Python module directly:
python -m tests.synthetic.rds_postgres.run_suite \
--mock-grafana \
--json
The --mock-grafana flag is essential for CI environments to serve static Grafana data from fixtures rather than attempting to connect to live instances. The --json flag emits machine-readable JSON results suitable for automated reporting and dashboards.
Executing Specific Scenarios
To debug a single test case rather than the entire suite, use the --scenario flag:
python -m tests.synthetic.rds_postgres.run_suite \
--scenario 006-replication-lag-cpu-redherring \
--mock-grafana
This approach runs only the specified scenario directory, allowing rapid iteration when adjusting prompt engineering for specific failure modes like replication lag with CPU red herrings.
Running Adversarial Axis 2 Tests
The suite includes adversarial "Axis 2" tests that verify the agent's ability to avoid false positives and handle edge cases. These require the selective backend and run via pytest:
pytest -m axis2 tests/synthetic/rds_postgres/test_suite_axis2.py -v
These tests validate that the agent does not hallucinate root causes when presented with ambiguous or confounding evidence.
Understanding Test Results
After execution, the runner outputs a per-scenario PASS/FAIL line followed by a summary:
PASS 001-healthy category=healthy
FAIL 006-replication-lag-cpu-redherring reason='missing required keywords: [...]'
...
Results: 9/10 passed
A PASS indicates the agent correctly identified the root cause category, included all required keywords from answer.yml, and cited the correct evidence sources. A FAIL provides the specific discrepancy, such as missing keywords or incorrect trajectory, enabling precise debugging of the investigation pipeline.
Summary
- Run synthetic tests in OpenSRE using
opensre tests syntheticfor quick checks orpython -m tests.synthetic.rds_postgres.run_suite --mock-grafanafor CI pipelines. - The FixtureGrafanaBackend in
tests/synthetic/mock_grafana_backend/backend.pyenables testing without live Grafana instances by serving static fixture data. - Scenarios are loaded by
tests/synthetic/rds_postgres/scenario_loader.pyand evaluated againstanswer.ymlkeys using the scoring functions inrun_suite.py. - Use
--scenarioto isolate specific test cases andpytest -m axis2to execute adversarial validation suites.
Frequently Asked Questions
What is a synthetic test in OpenSRE?
A synthetic test in OpenSRE is an automated evaluation that feeds the investigation pipeline a pre-constructed fixture containing a static alert and pre-generated evidence. This allows you to verify that the agent correctly diagnoses specific failure modes—such as replication lag or CPU saturation—without requiring access to live RDS instances or Grafana servers.
Do I need a live Grafana instance to run synthetic tests?
No. When you run synthetic tests with the --mock-grafana flag, the FixtureGrafanaBackend intercepts Grafana API calls and returns static metric files from the scenario directory. This mock backend is located in tests/synthetic/mock_grafana_backend/backend.py and is essential for running tests in CI environments without cloud credentials.
How do I add a new synthetic test scenario?
Create a new directory under tests/synthetic/rds_postgres/scenarios/ containing an alert.json (the triggering alert), an evidence/ directory with static data files (metrics, events), and an answer.yml (the expected root cause and keywords). The scenario_loader.py automatically discovers new directories and validates the schema, making them available to run_suite.py immediately.
What is the difference between regular and Axis 2 tests?
Regular synthetic tests verify that the agent correctly identifies known root causes when presented with clear evidence. Axis 2 tests are adversarial validations that ensure the agent does not hallucinate causes or fall for red herrings (such as high CPU usage during a replication lag incident). These run via pytest -m axis2 and use selective backend configurations to test the agent's reasoning boundaries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →