How to Implement Data Quality Checks Using Great Expectations in the Data Engineer Handbook

Great Expectations is an open-source Python library for defining, executing, and documenting data validation tests that you can integrate into the Data Expert-io handbook to create reusable, modular data quality pipelines.

This guide shows you how to implement data quality checks using Great Expectations within the DataEngineer-io/data-engineer-handbook repository structure. While the handbook lists GE as a recommended tool in its Data Quality section (see README.md lines 75-80), it does not yet contain a ready-made implementation—so you'll build one following the same architectural patterns used for Spark and Flink boot-camp modules.

Where Great Expectations Fits in the Handbook

The handbook's README.md explicitly recommends GE for data quality work. You can extend this by creating a lightweight validation pipeline that mirrors the modular job structure found in intermediate-bootcamp/materials/3-spark-fundamentals/src/jobs/.

Following the handbook's design principles of clean, modular, and reproducible data pipelines, you'll keep validation logic isolated from transformation jobs. This approach allows reuse across multiple lessons and provides immediate feedback to learners.

Architectural Design for GE Integration

Data Source Layer

Load raw data using the same utilities the handbook employs in its Spark or Databricks notebooks. The intermediate-bootcamp/materials/6-data-impact-training/data/events.csv file serves as an ideal test dataset.

Expectation Suite

Define suites of expectations in a dedicated expectations/ package. This follows the handbook's pattern of separating configuration from execution logic.

Validation Runner

Create a reusable script at validation/run_validation.py that:

  • Instantiates a GE DataContext
  • Loads datasets as GE Batch objects
  • Executes expectation suites
  • Emits JSON or HTML reports

CI/CD Integration

Hook the validation runner into existing workflows or boot-camp notebooks to provide immediate learner feedback.

Step-by-Step Implementation

1. Add the Dependency

Update your requirements.txt following the pattern in intermediate-bootcamp/materials/5-kpis-and-experimentation/requirements.txt:


# Add to requirements.txt

great-expectations>=0.18

2. Create the Validation Structure

mkdir -p intermediate-bootcamp/materials/6-data-impact-training/validation/reports
touch intermediate-bootcamp/materials/6-data-impact-training/validation/__init__.py

3. Write the Validation Script

The following implementation follows the handbook's folder conventions and uses the boot-camp's events.csv sample data:


# File: intermediate-bootcamp/materials/6-data-impact-training/validation/run_validation.py

import great_expectations as ge
from great_expectations.checkpoint import SimpleCheckpoint
from pathlib import Path


# ------------------------------------------------------------------

# 1️⃣ Initialise a DataContext (in-memory for this example)

# ------------------------------------------------------------------

context = ge.get_context()


# ------------------------------------------------------------------

# 2️⃣ Define a simple Expectation Suite (or load an existing one)

# ------------------------------------------------------------------

suite_name = "events_suite"
if not context.suites.get_suite_names().contains(suite_name):
    suite = context.add_expectation_suite(suite_name)
else:
    suite = context.get_expectation_suite(suite_name)


# ------------------------------------------------------------------

# 3️⃣ Add expectations (column existence, type, non-null, etc.)

# ------------------------------------------------------------------

suite.add_expectation(
    expectation_type="expect_table_columns_to_match_ordered_list",
    kwargs={"column_list": ["event_id", "timestamp", "device_id", "event_type"]},
)
suite.add_expectation(
    expectation_type="expect_column_values_to_not_be_null",
    kwargs={"column": "event_id"},
)
suite.add_expectation(
    expectation_type="expect_column_values_to_be_of_type",
    kwargs={"column": "timestamp", "type_": "datetime"},
)

# Save the suite for reuse

context.save_expectation_suite(suite, overwrite_existing=True)


# ------------------------------------------------------------------

# 4️⃣ Create a Batch (load the CSV used in the boot-camp)

# ------------------------------------------------------------------

data_path = Path(__file__).parents[2] / "data" / "events.csv"
batch = context.get_batch(
    batch_kwargs={"path": str(data_path), "datasource": "filesystem"},
    expectation_suite_name=suite_name,
)


# ------------------------------------------------------------------

# 5️⃣ Run a SimpleCheckpoint to validate and produce a report

# ------------------------------------------------------------------

checkpoint = SimpleCheckpoint(
    name="events_check",
    data_context=context,
    expectation_suite_name=suite_name,
    batch_list=[batch],
)

# Execute validation

result = checkpoint.run()


# ------------------------------------------------------------------

# 6️⃣ Export an HTML report (handy for boot-camp review)

# ------------------------------------------------------------------

report_path = Path(__file__).parent / "reports" / "events_validation.html"
result.get("validation_result").save_expectation_suite(
    expectation_suite_name=suite_name, 
    format="html", 
    path=str(report_path)
)

print(f"✅ Validation complete – report saved to {report_path}")

4. Run the Validation

cd intermediate-bootcamp/materials/6-data-impact-training/validation
python run_validation.py

Key Implementation Details

In-memory vs. file-based context: The example uses ge.get_context() for simplicity. For production work, switch to ge.data_context.DataContext(project_config="great_expectations.yml") to persist configuration.

Programmatic vs. CLI suite creation: The script adds expectations programmatically. Alternatively, use great_expectations suite new to generate JSON suites and commit them to version control.

Report consumption: The generated HTML report opens directly in browsers, giving learners visual feedback on data quality failures—critical for the handbook's educational mission.

Extending to Other Boot-Camp Modules

Apply this same pattern to other handbook datasets:

  • Reuse the validation/ directory structure across intermediate-bootcamp/materials/*
  • Parameterize suite_name and data_path via CLI arguments or environment variables
  • Store baseline expectation suites as JSON files in expectations/ for version control

Key Files in the Repository

File Purpose
README.md Data Quality tools list including GE recommendation
intermediate-bootcamp/materials/6-data-impact-training/data/events.csv Sample dataset for validation testing
intermediate-bootcamp/materials/5-kpis-and-experimentation/requirements.txt Dependency listing template
intermediate-bootcamp/materials/6-data-impact-training/README.md Context for data-impact training module

Summary

  • Great Expectations integrates cleanly with the Data Engineer Handbook's modular architecture
  • Place validation scripts in validation/ directories following existing job patterns
  • Use SimpleCheckpoint for straightforward validation and reporting
  • Target events.csv and similar boot-camp datasets for immediate learner relevance
  • Generate HTML reports stored under reports/ for visual feedback

Frequently Asked Questions

What version of Great Expectations should I use with the handbook?

Use great-expectations>=0.18 as specified in the dependency files. This version provides the get_context() API and SimpleCheckpoint class used in the implementation. Earlier versions may require different initialization patterns.

Can I use Great Expectations with Spark datasets in the handbook?

Yes. Replace the filesystem batch kwargs with Spark DataFrame configuration. The handbook's Spark modules in intermediate-bootcamp/materials/3-spark-fundamentals/ demonstrate DataFrame patterns you can adapt—pass your Spark session to GE's SparkDFExecutionEngine via datasource configuration.

How do I persist expectation suites across notebook sessions?

Switch from in-memory context to file-based: context = ge.data_context.DataContext(project_config="great_expectations.yml"). This creates a great_expectations/ directory with version-controlled suites, or generate suites via CLI (great_expectations suite new) and load them programmatically with context.get_expectation_suite().

Where should validation reports be stored for boot-camp review?

Save reports to validation/reports/ as shown in the example. This location is accessible to learners completing the 6-data-impact-training module and follows the handbook's convention of separating outputs from source code. For CI integration, additionally emit JSON to reports/ at the repository root.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →