# End-to-End Testing Scenarios for OpenSRE: Real-World Validation of the Incident Response Stack

> Explore OpenSRE's end-to-end testing scenarios. Validate incident response across AWS Kubernetes and third-party services by injecting failures and verifying root-cause analysis in real infrastructure.

- Repository: [Tracer/opensre](https://github.com/Tracer-Cloud/opensre)
- Tags: testing
- Published: 2026-04-18

---

**OpenSRE validates its complete incident-response platform through comprehensive end-to-end (E2E) tests that deploy real infrastructure, inject controlled failures, and verify root-cause analysis across AWS, Kubernetes, and third-party services.**

The OpenSRE repository maintains a robust E2E testing suite under `tests/e2e/` that ensures the platform correctly investigates incidents across diverse technology stacks. These scenarios test the full workflow from infrastructure deployment through synthetic alert generation to final root-cause verification, covering everything from Prefect ECS Fargate pipelines to GitLab merge requests.

## E2E Testing Architecture and Common Patterns

All E2E scenarios in OpenSRE follow a standardized six-step pattern implemented across the `tests/e2e/` directory.

### Infrastructure Deployment Pattern

Each scenario uses an `infrastructure_sdk` module to provision temporary AWS resources including Lambda functions, ECS clusters, Prefect agents, or PostgreSQL instances. The `Makefile` target `make deploy` orchestrates parallel deployment of all stacks, ensuring isolated test environments for each run.

### Controlled Failure Injection

Tests trigger specific failure modes through HTTP endpoints or CLI commands. For example, the Lambda DAG scenario sends `inject_schema_change=true` to force malformed data writes to S3, while the Prefect scenario triggers `inject_error=true` to simulate pipeline failures.

### Evidence Collection and Synthetic Alerting

After failure injection, tests gather evidence from CloudWatch logs, S3 objects, ECS task metadata, or Kubernetes pod logs. The [`app/utils/alert_factory.py`](https://github.com/Tracer-Cloud/opensre/blob/main/app/utils/alert_factory.py) utility constructs synthetic alerts containing annotations linking to the collected evidence, which the investigation agent then processes.

### Investigation CLI Execution and Assertion

The core validation occurs when `app.cli.investigate.run_investigation_cli` executes under LangSmith tracing. The CLI loads the synthetic alert, fetches linked data, runs configured tools, and returns a root-cause report. Tests assert that the investigation completes without error and identifies the correct root cause.

## Core End-to-End Testing Scenarios

OpenSRE maintains independent E2E suites targeting specific technology stacks, each with dedicated test entry points and infrastructure requirements.

### Prefect ECS Fargate Pipeline Failures

Located at [`tests/e2e/upstream_prefect_ecs_fargate/test_agent_e2e.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/upstream_prefect_ecs_fargate/test_agent_e2e.py), this scenario validates incident response for Prefect flows running on AWS Fargate. The test triggers a pipeline failure via the trigger API, collects evidence from CloudWatch Logs and S3 landing buckets, and verifies the ECS task ARN is correctly关联 with the root cause analysis.

### Upstream/Downstream Lambda DAG

The [`tests/e2e/upstream_lambda/test_agent_e2e.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/upstream_lambda/test_agent_e2e.py) suite tests serverless pipeline failures where a mock DAG receives `inject_schema_change=true` and writes corrupted data to S3. The scenario collects Lambda CloudWatch logs and S3 objects, then validates that the investigation agent correctly identifies the schema mismatch as the root cause.

### Apache Flink on ECS

For stream processing validation, [`tests/e2e/upstream_apache_flink_ecs/test_agent_e2e.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/upstream_apache_flink_ecs/test_agent_e2e.py) deploys Flink jobs on ECS and injects bad schema configurations via the trigger API. The test gathers Flink logs, S3 checkpoint data, and ECS task information to verify the agent can trace stream processing failures.

### Kubernetes and Datadog Integration

The [`tests/e2e/kubernetes/test_local.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/kubernetes/test_local.py) suite covers container orchestration failures using local or EKS clusters. It creates failing pods via `kubectl`, emits Datadog alerts, and tests the agent's ability to correlate Kubernetes pod logs with Datadog events and Grafana metrics.

### Grafana Cloud Validation

Located at [`tests/e2e/grafana_validation/test_grafana_cloud_queries.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/grafana_validation/test_grafana_cloud_queries.py), this scenario validates the agent's integration with Grafana Cloud by querying logs, metrics, and traces APIs. It ensures the investigation pipeline can correctly fetch and interpret observability data from Grafana-managed backends.

### Deployment Platform Testing

The [`tests/e2e/deploy/test_deploy_timing.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/deploy/test_deploy_timing.py) suite validates deployment-time health checks across Vercel, Railway, and AWS. It deploys test applications and polls health endpoints to verify the orchestration layer correctly handles deployment failures and timeouts.

### PostHog Analytics Integration

For product analytics validation, [`tests/e2e/posthog/test_orchestrator.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/posthog/test_orchestrator.py) sends synthetic bounce-rate events to PostHog and verifies the orchestration flow correctly processes analytics-based alerts.

### GitLab Merge Request Workflows

The [`tests/e2e/gitlab/test_cloud.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/gitlab/test_cloud.py) suite creates merge requests via the GitLab API and runs investigations on pipeline failures associated with MRs, validating the agent's integration with GitLab commits, pipelines, and MR notes.

### Synthetic Database Connections

OpenSRE tests database-related incident response through [`tests/e2e/postgresql/test_postgresql_e2e.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/postgresql/test_postgresql_e2e.py) and equivalents for MySQL and MongoDB. These scenarios generate failing alerts referring to database sources, then validate that the agent correctly establishes connections and queries the databases to identify root causes.

## Code Implementation Examples

The E2E tests utilize specific helper functions and CLI entry points to orchestrate the validation workflow.

### Triggering Pipeline Failures

The Prefect ECS scenario demonstrates how to inject failures and capture execution metadata:

```python

# tests/e2e/upstream_prefect_ecs_fargate/test_agent_e2e.py

def trigger_pipeline_failure(run_id: str, trace_id: str) -> dict:
    url = f"{CONFIG['trigger_api_url']}trigger?inject_error=true"
    trigger_timer = StepTimer(
        trace_id=trace_id,
        run_id=run_id,
        run_name="upstream_downstream_pipeline_prefect",
        tool_id="pipeline_trigger",
        tool_name="Pipeline Orchestrator",
        tool_cmd="trigger_prefect_flow",
    )
    response = requests.post(url, timeout=60)
    result = response.json()
    # …extract correlation_id, s3_key, task_arn …

    trigger_timer.finish(exit_code=0, metadata={...})
    return {
        "correlation_id": correlation_id,
        "s3_key": s3_key,
        "task_arn": task_arn,
        "bucket": CONFIG["s3_bucket"],
    }

```

### Constructing Synthetic Alerts

The Lambda DAG scenario uses [`alert_factory.py`](https://github.com/Tracer-Cloud/opensre/blob/main/alert_factory.py) to create investigation-ready alerts:

```python

# tests/e2e/upstream_lambda/test_agent_e2e.py

raw_alert = create_alert(
    pipeline_name="upstream_downstream_pipeline",
    run_name=run_id,
    status="failed",
    timestamp=datetime.now(UTC).isoformat(),
    annotations={
        "s3_bucket": failure_data["bucket"],
        "s3_key": failure_data["s3_key"],
        "correlation_id": failure_data["correlation_id"],
        "error": failure_data["error_message"],
        "lambda_log_group": failure_data["log_group"],
        "function_name": UPSTREAM_DOWNSTREAM_CONFIG["mock_dag_function_name"],
        "context_sources": "s3,lambda,cloudwatch",
    },
)

```

### Executing the Investigation CLI

Tests invoke the core investigation logic under LangSmith tracing:

```python

# tests/e2e/upstream_lambda/test_agent_e2e.py

@traceable(
    run_type="chain",
    name=f"test_lambda_upstream - {raw_alert['alert_id'][:8]}",
    metadata={"alert_id": raw_alert['alert_id']},
)
def run_investigation():
    return run_investigation_cli(
        alert_name="Pipeline failure: upstream_downstream_pipeline",
        pipeline_name="upstream_downstream_pipeline",
        severity="critical",
        raw_alert=raw_alert,
    )

```

## Key Files and Their Roles

| Path | Purpose |
|------|---------|
| [`tests/e2e/upstream_prefect_ecs_fargate/test_agent_e2e.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/upstream_prefect_ecs_fargate/test_agent_e2e.py) | Prefect ECS Fargate E2E flow – triggers failure, gathers CloudWatch logs, runs investigation |
| [`tests/e2e/upstream_lambda/test_agent_e2e.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/upstream_lambda/test_agent_e2e.py) | Lambda DAG E2E flow – injects schema changes, reads Lambda logs, creates alerts |
| [`tests/e2e/upstream_apache_flink_ecs/test_agent_e2e.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/upstream_apache_flink_ecs/test_agent_e2e.py) | Apache Flink ECS failure injection and verification |
| [`tests/e2e/kubernetes/test_local.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/kubernetes/test_local.py) | Local Kubernetes pod failure simulation and Datadog alert handling |
| [`tests/e2e/grafana_validation/test_grafana_cloud_queries.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/grafana_validation/test_grafana_cloud_queries.py) | Grafana Cloud query validation for logs, metrics, and traces |
| [`tests/e2e/deploy/test_deploy_timing.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/deploy/test_deploy_timing.py) | Deployment platform health checks – Vercel, Railway, AWS |
| [`tests/e2e/posthog/test_orchestrator.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/posthog/test_orchestrator.py) | PostHog bounce-rate alert orchestration |
| [`tests/e2e/gitlab/test_cloud.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/gitlab/test_cloud.py) | GitLab merge request and pipeline investigation |
| [`tests/e2e/postgresql/test_postgresql_e2e.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/e2e/postgresql/test_postgresql_e2e.py) | Synthetic PostgreSQL connection and query validation |
| [`app/cli/investigate.py`](https://github.com/Tracer-Cloud/opensre/blob/main/app/cli/investigate.py) | Core CLI entry point for investigation pipeline |
| [`app/utils/alert_factory.py`](https://github.com/Tracer-Cloud/opensre/blob/main/app/utils/alert_factory.py) | Synthetic alert construction utility |
| [`tests/utils/tracer_ingest.py`](https://github.com/Tracer-Cloud/opensre/blob/main/tests/utils/tracer_ingest.py) | StepTimer and tracing utilities for E2E workflows |
| `Makefile` | Orchestrates `deploy` and `test-e2e` targets |
| `tests/e2e/**/infrastructure_sdk/*.py` | Cloud resource provisioning scripts per scenario |

## Summary

OpenSRE's end-to-end testing framework validates the entire incident-response pipeline through real-world failure injection across multiple cloud-native technologies. The testing suite covers:

- **Container orchestration** via Prefect ECS Fargate and Kubernetes scenarios
- **Serverless workflows** through Lambda DAG and Apache Flink ECS tests
- **Observability integrations** with Grafana Cloud, Datadog, and PostHog
- **Deployment platforms** including Vercel, Railway, and AWS
- **Version control and CI/CD** through GitLab merge request workflows
- **Database systems** via synthetic PostgreSQL, MySQL, and MongoDB connections

All scenarios follow a consistent six-step pattern: infrastructure deployment, controlled failure injection, evidence collection, synthetic alert creation, investigation CLI execution, and outcome assertion. These tests run under CI via the `test-e2e` pytest marker, ensuring OpenSRE functions correctly in clean, ephemeral environments.

## Frequently Asked Questions

### What directory contains the end-to-end tests in OpenSRE?

All end-to-end tests reside in the `tests/e2e/` directory at the repository root. Each subdirectory represents a specific technology scenario (such as `upstream_lambda`, `kubernetes`, or `postgresql`) containing its own [`test_agent_e2e.py`](https://github.com/Tracer-Cloud/opensre/blob/main/test_agent_e2e.py) or equivalent test file along with infrastructure provisioning scripts.

### How does OpenSRE inject failures during E2E testing?

OpenSRE injects failures through HTTP API calls or CLI triggers that set specific flags like `inject_error=true` or `inject_schema_change=true`. For example, the Lambda DAG scenario sends a POST request with `inject_schema_change=true` to force the pipeline to write malformed data to S3, while the Prefect scenario triggers flows with error injection enabled to simulate real pipeline failures.

### What evidence sources does OpenSRE collect during E2E investigations?

The E2E tests collect evidence from multiple sources depending on the scenario: CloudWatch logs for Lambda and ECS tasks, S3 objects containing failure data, ECS task metadata and ARNs, Kubernetes pod logs, Datadog events, Grafana Cloud logs/metrics/traces, and database query logs from PostgreSQL, MySQL, or MongoDB instances.

### How does OpenSRE verify that the investigation agent correctly identified the root cause?

Each E2E test asserts that the `run_investigation_cli` function from [`app/cli/investigate.py`](https://github.com/Tracer-Cloud/opensre/blob/main/app/cli/investigate.py) completes without error and returns a root-cause report containing the expected error message or correlation ID. The tests verify that the agent successfully fetched evidence from the annotated sources (such as specific S3 keys or log groups) and correctly traced the injected failure back to its source.