End-to-End Testing Scenarios for OpenSRE: Real-World Validation of the Incident Response Stack

OpenSRE validates its complete incident-response platform through comprehensive end-to-end (E2E) tests that deploy real infrastructure, inject controlled failures, and verify root-cause analysis across AWS, Kubernetes, and third-party services.

The OpenSRE repository maintains a robust E2E testing suite under tests/e2e/ that ensures the platform correctly investigates incidents across diverse technology stacks. These scenarios test the full workflow from infrastructure deployment through synthetic alert generation to final root-cause verification, covering everything from Prefect ECS Fargate pipelines to GitLab merge requests.

E2E Testing Architecture and Common Patterns

All E2E scenarios in OpenSRE follow a standardized six-step pattern implemented across the tests/e2e/ directory.

Infrastructure Deployment Pattern

Each scenario uses an infrastructure_sdk module to provision temporary AWS resources including Lambda functions, ECS clusters, Prefect agents, or PostgreSQL instances. The Makefile target make deploy orchestrates parallel deployment of all stacks, ensuring isolated test environments for each run.

Controlled Failure Injection

Tests trigger specific failure modes through HTTP endpoints or CLI commands. For example, the Lambda DAG scenario sends inject_schema_change=true to force malformed data writes to S3, while the Prefect scenario triggers inject_error=true to simulate pipeline failures.

Evidence Collection and Synthetic Alerting

After failure injection, tests gather evidence from CloudWatch logs, S3 objects, ECS task metadata, or Kubernetes pod logs. The app/utils/alert_factory.py utility constructs synthetic alerts containing annotations linking to the collected evidence, which the investigation agent then processes.

Investigation CLI Execution and Assertion

The core validation occurs when app.cli.investigate.run_investigation_cli executes under LangSmith tracing. The CLI loads the synthetic alert, fetches linked data, runs configured tools, and returns a root-cause report. Tests assert that the investigation completes without error and identifies the correct root cause.

Core End-to-End Testing Scenarios

OpenSRE maintains independent E2E suites targeting specific technology stacks, each with dedicated test entry points and infrastructure requirements.

Prefect ECS Fargate Pipeline Failures

Located at tests/e2e/upstream_prefect_ecs_fargate/test_agent_e2e.py, this scenario validates incident response for Prefect flows running on AWS Fargate. The test triggers a pipeline failure via the trigger API, collects evidence from CloudWatch Logs and S3 landing buckets, and verifies the ECS task ARN is correctly关联 with the root cause analysis.

Upstream/Downstream Lambda DAG

The tests/e2e/upstream_lambda/test_agent_e2e.py suite tests serverless pipeline failures where a mock DAG receives inject_schema_change=true and writes corrupted data to S3. The scenario collects Lambda CloudWatch logs and S3 objects, then validates that the investigation agent correctly identifies the schema mismatch as the root cause.

For stream processing validation, tests/e2e/upstream_apache_flink_ecs/test_agent_e2e.py deploys Flink jobs on ECS and injects bad schema configurations via the trigger API. The test gathers Flink logs, S3 checkpoint data, and ECS task information to verify the agent can trace stream processing failures.

Kubernetes and Datadog Integration

The tests/e2e/kubernetes/test_local.py suite covers container orchestration failures using local or EKS clusters. It creates failing pods via kubectl, emits Datadog alerts, and tests the agent's ability to correlate Kubernetes pod logs with Datadog events and Grafana metrics.

Grafana Cloud Validation

Located at tests/e2e/grafana_validation/test_grafana_cloud_queries.py, this scenario validates the agent's integration with Grafana Cloud by querying logs, metrics, and traces APIs. It ensures the investigation pipeline can correctly fetch and interpret observability data from Grafana-managed backends.

Deployment Platform Testing

The tests/e2e/deploy/test_deploy_timing.py suite validates deployment-time health checks across Vercel, Railway, and AWS. It deploys test applications and polls health endpoints to verify the orchestration layer correctly handles deployment failures and timeouts.

PostHog Analytics Integration

For product analytics validation, tests/e2e/posthog/test_orchestrator.py sends synthetic bounce-rate events to PostHog and verifies the orchestration flow correctly processes analytics-based alerts.

GitLab Merge Request Workflows

The tests/e2e/gitlab/test_cloud.py suite creates merge requests via the GitLab API and runs investigations on pipeline failures associated with MRs, validating the agent's integration with GitLab commits, pipelines, and MR notes.

Synthetic Database Connections

OpenSRE tests database-related incident response through tests/e2e/postgresql/test_postgresql_e2e.py and equivalents for MySQL and MongoDB. These scenarios generate failing alerts referring to database sources, then validate that the agent correctly establishes connections and queries the databases to identify root causes.

Code Implementation Examples

The E2E tests utilize specific helper functions and CLI entry points to orchestrate the validation workflow.

Triggering Pipeline Failures

The Prefect ECS scenario demonstrates how to inject failures and capture execution metadata:


# tests/e2e/upstream_prefect_ecs_fargate/test_agent_e2e.py

def trigger_pipeline_failure(run_id: str, trace_id: str) -> dict:
    url = f"{CONFIG['trigger_api_url']}trigger?inject_error=true"
    trigger_timer = StepTimer(
        trace_id=trace_id,
        run_id=run_id,
        run_name="upstream_downstream_pipeline_prefect",
        tool_id="pipeline_trigger",
        tool_name="Pipeline Orchestrator",
        tool_cmd="trigger_prefect_flow",
    )
    response = requests.post(url, timeout=60)
    result = response.json()
    # …extract correlation_id, s3_key, task_arn …

    trigger_timer.finish(exit_code=0, metadata={...})
    return {
        "correlation_id": correlation_id,
        "s3_key": s3_key,
        "task_arn": task_arn,
        "bucket": CONFIG["s3_bucket"],
    }

Constructing Synthetic Alerts

The Lambda DAG scenario uses alert_factory.py to create investigation-ready alerts:


# tests/e2e/upstream_lambda/test_agent_e2e.py

raw_alert = create_alert(
    pipeline_name="upstream_downstream_pipeline",
    run_name=run_id,
    status="failed",
    timestamp=datetime.now(UTC).isoformat(),
    annotations={
        "s3_bucket": failure_data["bucket"],
        "s3_key": failure_data["s3_key"],
        "correlation_id": failure_data["correlation_id"],
        "error": failure_data["error_message"],
        "lambda_log_group": failure_data["log_group"],
        "function_name": UPSTREAM_DOWNSTREAM_CONFIG["mock_dag_function_name"],
        "context_sources": "s3,lambda,cloudwatch",
    },
)

Executing the Investigation CLI

Tests invoke the core investigation logic under LangSmith tracing:


# tests/e2e/upstream_lambda/test_agent_e2e.py

@traceable(
    run_type="chain",
    name=f"test_lambda_upstream - {raw_alert['alert_id'][:8]}",
    metadata={"alert_id": raw_alert['alert_id']},
)
def run_investigation():
    return run_investigation_cli(
        alert_name="Pipeline failure: upstream_downstream_pipeline",
        pipeline_name="upstream_downstream_pipeline",
        severity="critical",
        raw_alert=raw_alert,
    )

Key Files and Their Roles

Path Purpose
tests/e2e/upstream_prefect_ecs_fargate/test_agent_e2e.py Prefect ECS Fargate E2E flow – triggers failure, gathers CloudWatch logs, runs investigation
tests/e2e/upstream_lambda/test_agent_e2e.py Lambda DAG E2E flow – injects schema changes, reads Lambda logs, creates alerts
tests/e2e/upstream_apache_flink_ecs/test_agent_e2e.py Apache Flink ECS failure injection and verification
tests/e2e/kubernetes/test_local.py Local Kubernetes pod failure simulation and Datadog alert handling
tests/e2e/grafana_validation/test_grafana_cloud_queries.py Grafana Cloud query validation for logs, metrics, and traces
tests/e2e/deploy/test_deploy_timing.py Deployment platform health checks – Vercel, Railway, AWS
tests/e2e/posthog/test_orchestrator.py PostHog bounce-rate alert orchestration
tests/e2e/gitlab/test_cloud.py GitLab merge request and pipeline investigation
tests/e2e/postgresql/test_postgresql_e2e.py Synthetic PostgreSQL connection and query validation
app/cli/investigate.py Core CLI entry point for investigation pipeline
app/utils/alert_factory.py Synthetic alert construction utility
tests/utils/tracer_ingest.py StepTimer and tracing utilities for E2E workflows
Makefile Orchestrates deploy and test-e2e targets
tests/e2e/**/infrastructure_sdk/*.py Cloud resource provisioning scripts per scenario

Summary

OpenSRE's end-to-end testing framework validates the entire incident-response pipeline through real-world failure injection across multiple cloud-native technologies. The testing suite covers:

  • Container orchestration via Prefect ECS Fargate and Kubernetes scenarios
  • Serverless workflows through Lambda DAG and Apache Flink ECS tests
  • Observability integrations with Grafana Cloud, Datadog, and PostHog
  • Deployment platforms including Vercel, Railway, and AWS
  • Version control and CI/CD through GitLab merge request workflows
  • Database systems via synthetic PostgreSQL, MySQL, and MongoDB connections

All scenarios follow a consistent six-step pattern: infrastructure deployment, controlled failure injection, evidence collection, synthetic alert creation, investigation CLI execution, and outcome assertion. These tests run under CI via the test-e2e pytest marker, ensuring OpenSRE functions correctly in clean, ephemeral environments.

Frequently Asked Questions

What directory contains the end-to-end tests in OpenSRE?

All end-to-end tests reside in the tests/e2e/ directory at the repository root. Each subdirectory represents a specific technology scenario (such as upstream_lambda, kubernetes, or postgresql) containing its own test_agent_e2e.py or equivalent test file along with infrastructure provisioning scripts.

How does OpenSRE inject failures during E2E testing?

OpenSRE injects failures through HTTP API calls or CLI triggers that set specific flags like inject_error=true or inject_schema_change=true. For example, the Lambda DAG scenario sends a POST request with inject_schema_change=true to force the pipeline to write malformed data to S3, while the Prefect scenario triggers flows with error injection enabled to simulate real pipeline failures.

What evidence sources does OpenSRE collect during E2E investigations?

The E2E tests collect evidence from multiple sources depending on the scenario: CloudWatch logs for Lambda and ECS tasks, S3 objects containing failure data, ECS task metadata and ARNs, Kubernetes pod logs, Datadog events, Grafana Cloud logs/metrics/traces, and database query logs from PostgreSQL, MySQL, or MongoDB instances.

How does OpenSRE verify that the investigation agent correctly identified the root cause?

Each E2E test asserts that the run_investigation_cli function from app/cli/investigate.py completes without error and returns a root-cause report containing the expected error message or correlation ID. The tests verify that the agent successfully fetched evidence from the annotated sources (such as specific S3 keys or log groups) and correctly traced the injected failure back to its source.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →