# How to Set Up Data Quality Checks Using Great Expectations: A Complete Guide

> Learn how to set up data quality checks with Great Expectations. Define, validate, and document data rules for automated quality gates and HTML reports.

- Repository: [DataExpert.io/data-engineer-handbook](https://github.com/DataExpert-io/data-engineer-handbook)
- Tags: how-to-guide
- Published: 2026-08-09

---

**Great Expectations enables data engineers to define, validate, and document data quality rules through modular expectation suites that automatically generate HTML reports and integrate into CI/CD pipelines for automated quality gates.**

While the DataExpert-io/data-engineer-handbook repository references Great Expectations in its [`README.md`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/README.md) as a recommended data quality tool, it does not ship with a concrete implementation. This guide demonstrates exactly how to set up data quality checks using Great Expectations with the repository's sample event data, providing a production-ready framework for validating data pipelines.

## Understanding the Great Expectations Architecture

Before writing code, you must understand the five core components that comprise a Great Expectations (GE) deployment. This modular architecture separates concerns between data connection, rule definition, execution, persistence, and visualization.

### Core Components

- **Data Source** – GE connects to **Pandas DataFrames**, **Spark DataFrames**, SQL databases, or file-based sources (CSV, Parquet). In the Data Engineer Handbook repository, you will work with the sample file located at `intermediate-bootcamp/materials/3-spark-fundamentals/data/events.csv`.

- **Expectation Suite** – A JSON-serializable collection of data assertions (e.g., column uniqueness, null-rate thresholds, regex patterns). Create these programmatically via `context.create_expectation_suite()` or interactively through Data Docs.

- **Validation Engine** – Executes the suite against a specific **Batch** of data (a slice such as a daily CSV file) and returns a structured result object containing pass/fail status per expectation.

- **Result Store** – Persists validation outcomes to configurable backends (local filesystem, S3, or databases) for auditability. Configure this in [`great_expectations.yml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/great_expectations.yml) under `store_backend`.

- **Data Docs** – Auto-generated static HTML sites created by `context.build_data_docs()` that visualize expectations and historical validation results, enabling self-service data quality monitoring for non-technical stakeholders.

## Prerequisites and Installation

Install Great Expectations and initialize a project context. This creates the `great_expectations/` directory containing configuration files and expectation stores.

```bash
pip install great_expectations

```

```python
import great_expectations as ge

# Initialize the Data Context (creates great_expectations/ directory)

context = ge.data_context.DataContext()

```

## Implementing Data Quality Checks with Sample Data

The following workflow uses the `events.csv` file from the handbook's Spark fundamentals module to demonstrate a complete validation pipeline.

### Step 1: Create an Expectation Suite

Define a named collection of expectations that will serve as your data contract. The `overwrite_existing=True` parameter ensures you can iterate during development.

```python
suite = context.create_expectation_suite(
    expectation_suite_name="events_suite", 
    overwrite_existing=True
)

```

### Step 2: Load Data and Build a Validator

Load the sample CSV into a Pandas DataFrame, then construct a **Batch** and **Validator**. The validator binds your data to the expectation suite and provides methods for defining assertions.

```python
import pandas as pd

# Load the handbook's sample data

df = pd.read_csv(
    "intermediate-bootcamp/materials/3-spark-fundamentals/data/events.csv"
)

# Configure the batch runtime parameters

batch = context.get_batch({
    "datasource_name": "my_pandas_datasource",
    "data_connector_name": "default_runtime_data_connector_name",
    "data_asset_name": "events",
    "runtime_parameters": {"batch_data": df},
    "batch_identifiers": {"default_identifier_name": "events_batch"},
})

# Initialize the validator

validator = context.get_validator(
    batch=batch,
    expectation_suite_name="events_suite",
)

```

### Step 3: Define Expectations

Apply specific data quality rules using built-in expectation methods. These assertions cover schema, completeness, and validity constraints.

```python

# Schema expectations

validator.expect_column_to_exist("event_id")

# Uniqueness constraints

validator.expect_column_values_to_be_unique("event_id")

# Completeness checks

validator.expect_column_values_to_not_be_null("event_timestamp")

# Format validation using regex

validator.expect_column_values_to_match_regex(
    "event_url", 
    r"^https?://.+$"
)

# Persist the suite to the GE store

validator.save_expectation_suite(discard_failed_expectations=False)

```

### Step 4: Validate and Generate Reports

Execute the suite against the batch and produce human-readable documentation. The validation result object contains detailed statistics on pass/fail rates, unexpected values, and partial unexpected counts.

```python

# Execute validation

results = validator.validate()
print(results)

# Generate static HTML documentation

context.build_data_docs()
print("Data Docs URL:", context.get_site_url())

```

## Scaling with Apache Spark

For large datasets, swap the Pandas DataFrame for a Spark DataFrame. The validation API remains identical; only the datasource configuration changes.

```python
from pyspark.sql import SparkSession

spark = SparkSession.builder.appName("GE_Spark").getOrCreate()
spark_df = spark.read.parquet("path/to/large_dataset.parquet")

batch = context.get_batch({
    "datasource_name": "my_spark_datasource",
    "data_connector_name": "default_runtime_data_connector_name",
    "data_asset_name": "events_parquet",
    "runtime_parameters": {"batch_data": spark_df},
    "batch_identifiers": {"default_identifier_name": "events_batch"},
})

# Validator and expectation logic remains unchanged

validator = context.get_validator(batch=batch, expectation_suite_name="events_suite")

```

## Automating Checks in CI/CD Pipelines

Convert your validation workflow into a **Checkpoint**—a reusable configuration object that binds a batch, expectation suite, and result store. Checkpoints return non-zero exit codes on validation failure, enabling automated quality gates.

Create a checkpoint configuration or use the Python API, then execute via CLI:

```bash
great_expectations checkpoint run events_checkpoint

```

Integrate this command into your GitHub Actions, GitLab CI, or Jenkins pipeline to prevent bad data from reaching production tables. When validation fails, the pipeline stops, and Data Docs provide immediate forensic detail on which expectations failed and why.

## Summary

- **Great Expectations** provides a modular architecture separating data connection, rule definition, execution, and reporting.
- Use `context.create_expectation_suite()` to define data contracts and `context.get_validator()` to execute them against batches.
- Reference the handbook's sample data at `intermediate-bootcamp/materials/3-spark-fundamentals/data/events.csv` to practice validation workflows.
- Generate **Data Docs** with `context.build_data_docs()` to create self-service documentation for stakeholders.
- Implement **Checkpoints** in CI/CD pipelines using `great_expectations checkpoint run` to enforce automated data quality gates.

## Frequently Asked Questions

### What is the difference between an Expectation and a Checkpoint?

An **Expectation** is a single assertion about your data (e.g., `expect_column_values_to_be_unique`), while a **Checkpoint** is a configuration object that bundles an expectation suite with a specific batch and result store for automated execution. Checkpoints are designed for production pipelines, whereas expectations define the underlying business rules.

### Can Great Expectations handle petabyte-scale Spark datasets?

Yes. Great Expectations integrates natively with **Apache Spark** through the Spark datasource connector. By using `context.get_batch()` with a Spark DataFrame instead of Pandas, you can validate distributed datasets across clusters while utilizing the same expectation syntax and validation engine.

### How do I store validation results in Amazon S3 instead of locally?

Modify the [`great_expectations.yml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/great_expectations.yml) configuration file to specify an S3 backend for the `validations_store_name` and `site_store_backend`. Change the `store_backend` class to `TupleS3StoreBackend` and provide your bucket name and prefix. This ensures all validation history and Data Docs persist to cloud storage for team access.

### What happens when a data quality check fails during CI/CD execution?

When a checkpoint run encounters failed expectations, the process exits with a non-zero status code, causing the CI/CD pipeline to fail immediately. The generated validation results contain detailed JSON reports identifying which specific rows violated which expectations, allowing engineers to debug data quality issues before they impact downstream analytics.