How to Set Up Data Quality Checks Using Great Expectations: A Complete Guide

Great Expectations enables data engineers to define, validate, and document data quality rules through modular expectation suites that automatically generate HTML reports and integrate into CI/CD pipelines for automated quality gates.

While the DataExpert-io/data-engineer-handbook repository references Great Expectations in its README.md as a recommended data quality tool, it does not ship with a concrete implementation. This guide demonstrates exactly how to set up data quality checks using Great Expectations with the repository's sample event data, providing a production-ready framework for validating data pipelines.

Understanding the Great Expectations Architecture

Before writing code, you must understand the five core components that comprise a Great Expectations (GE) deployment. This modular architecture separates concerns between data connection, rule definition, execution, persistence, and visualization.

Core Components

  • Data Source – GE connects to Pandas DataFrames, Spark DataFrames, SQL databases, or file-based sources (CSV, Parquet). In the Data Engineer Handbook repository, you will work with the sample file located at intermediate-bootcamp/materials/3-spark-fundamentals/data/events.csv.

  • Expectation Suite – A JSON-serializable collection of data assertions (e.g., column uniqueness, null-rate thresholds, regex patterns). Create these programmatically via context.create_expectation_suite() or interactively through Data Docs.

  • Validation Engine – Executes the suite against a specific Batch of data (a slice such as a daily CSV file) and returns a structured result object containing pass/fail status per expectation.

  • Result Store – Persists validation outcomes to configurable backends (local filesystem, S3, or databases) for auditability. Configure this in great_expectations.yml under store_backend.

  • Data Docs – Auto-generated static HTML sites created by context.build_data_docs() that visualize expectations and historical validation results, enabling self-service data quality monitoring for non-technical stakeholders.

Prerequisites and Installation

Install Great Expectations and initialize a project context. This creates the great_expectations/ directory containing configuration files and expectation stores.

pip install great_expectations
import great_expectations as ge

# Initialize the Data Context (creates great_expectations/ directory)

context = ge.data_context.DataContext()

Implementing Data Quality Checks with Sample Data

The following workflow uses the events.csv file from the handbook's Spark fundamentals module to demonstrate a complete validation pipeline.

Step 1: Create an Expectation Suite

Define a named collection of expectations that will serve as your data contract. The overwrite_existing=True parameter ensures you can iterate during development.

suite = context.create_expectation_suite(
    expectation_suite_name="events_suite", 
    overwrite_existing=True
)

Step 2: Load Data and Build a Validator

Load the sample CSV into a Pandas DataFrame, then construct a Batch and Validator. The validator binds your data to the expectation suite and provides methods for defining assertions.

import pandas as pd

# Load the handbook's sample data

df = pd.read_csv(
    "intermediate-bootcamp/materials/3-spark-fundamentals/data/events.csv"
)

# Configure the batch runtime parameters

batch = context.get_batch({
    "datasource_name": "my_pandas_datasource",
    "data_connector_name": "default_runtime_data_connector_name",
    "data_asset_name": "events",
    "runtime_parameters": {"batch_data": df},
    "batch_identifiers": {"default_identifier_name": "events_batch"},
})

# Initialize the validator

validator = context.get_validator(
    batch=batch,
    expectation_suite_name="events_suite",
)

Step 3: Define Expectations

Apply specific data quality rules using built-in expectation methods. These assertions cover schema, completeness, and validity constraints.


# Schema expectations

validator.expect_column_to_exist("event_id")

# Uniqueness constraints

validator.expect_column_values_to_be_unique("event_id")

# Completeness checks

validator.expect_column_values_to_not_be_null("event_timestamp")

# Format validation using regex

validator.expect_column_values_to_match_regex(
    "event_url", 
    r"^https?://.+$"
)

# Persist the suite to the GE store

validator.save_expectation_suite(discard_failed_expectations=False)

Step 4: Validate and Generate Reports

Execute the suite against the batch and produce human-readable documentation. The validation result object contains detailed statistics on pass/fail rates, unexpected values, and partial unexpected counts.


# Execute validation

results = validator.validate()
print(results)

# Generate static HTML documentation

context.build_data_docs()
print("Data Docs URL:", context.get_site_url())

Scaling with Apache Spark

For large datasets, swap the Pandas DataFrame for a Spark DataFrame. The validation API remains identical; only the datasource configuration changes.

from pyspark.sql import SparkSession

spark = SparkSession.builder.appName("GE_Spark").getOrCreate()
spark_df = spark.read.parquet("path/to/large_dataset.parquet")

batch = context.get_batch({
    "datasource_name": "my_spark_datasource",
    "data_connector_name": "default_runtime_data_connector_name",
    "data_asset_name": "events_parquet",
    "runtime_parameters": {"batch_data": spark_df},
    "batch_identifiers": {"default_identifier_name": "events_batch"},
})

# Validator and expectation logic remains unchanged

validator = context.get_validator(batch=batch, expectation_suite_name="events_suite")

Automating Checks in CI/CD Pipelines

Convert your validation workflow into a Checkpoint—a reusable configuration object that binds a batch, expectation suite, and result store. Checkpoints return non-zero exit codes on validation failure, enabling automated quality gates.

Create a checkpoint configuration or use the Python API, then execute via CLI:

great_expectations checkpoint run events_checkpoint

Integrate this command into your GitHub Actions, GitLab CI, or Jenkins pipeline to prevent bad data from reaching production tables. When validation fails, the pipeline stops, and Data Docs provide immediate forensic detail on which expectations failed and why.

Summary

  • Great Expectations provides a modular architecture separating data connection, rule definition, execution, and reporting.
  • Use context.create_expectation_suite() to define data contracts and context.get_validator() to execute them against batches.
  • Reference the handbook's sample data at intermediate-bootcamp/materials/3-spark-fundamentals/data/events.csv to practice validation workflows.
  • Generate Data Docs with context.build_data_docs() to create self-service documentation for stakeholders.
  • Implement Checkpoints in CI/CD pipelines using great_expectations checkpoint run to enforce automated data quality gates.

Frequently Asked Questions

What is the difference between an Expectation and a Checkpoint?

An Expectation is a single assertion about your data (e.g., expect_column_values_to_be_unique), while a Checkpoint is a configuration object that bundles an expectation suite with a specific batch and result store for automated execution. Checkpoints are designed for production pipelines, whereas expectations define the underlying business rules.

Can Great Expectations handle petabyte-scale Spark datasets?

Yes. Great Expectations integrates natively with Apache Spark through the Spark datasource connector. By using context.get_batch() with a Spark DataFrame instead of Pandas, you can validate distributed datasets across clusters while utilizing the same expectation syntax and validation engine.

How do I store validation results in Amazon S3 instead of locally?

Modify the great_expectations.yml configuration file to specify an S3 backend for the validations_store_name and site_store_backend. Change the store_backend class to TupleS3StoreBackend and provide your bucket name and prefix. This ensures all validation history and Data Docs persist to cloud storage for team access.

What happens when a data quality check fails during CI/CD execution?

When a checkpoint run encounters failed expectations, the process exits with a non-zero status code, causing the CI/CD pipeline to fail immediately. The generated validation results contain detailed JSON reports identifying which specific rows violated which expectations, allowing engineers to debug data quality issues before they impact downstream analytics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →