How to Implement Data Quality Checks Using Great Expectations in the Data Engineer Handbook
Great Expectations is an open-source Python library for defining, executing, and documenting data validation tests that you can integrate into the Data Expert-io handbook to create reusable, modular data quality pipelines.
This guide shows you how to implement data quality checks using Great Expectations within the DataEngineer-io/data-engineer-handbook repository structure. While the handbook lists GE as a recommended tool in its Data Quality section (see README.md lines 75-80), it does not yet contain a ready-made implementation—so you'll build one following the same architectural patterns used for Spark and Flink boot-camp modules.
Where Great Expectations Fits in the Handbook
The handbook's README.md explicitly recommends GE for data quality work. You can extend this by creating a lightweight validation pipeline that mirrors the modular job structure found in intermediate-bootcamp/materials/3-spark-fundamentals/src/jobs/.
Following the handbook's design principles of clean, modular, and reproducible data pipelines, you'll keep validation logic isolated from transformation jobs. This approach allows reuse across multiple lessons and provides immediate feedback to learners.
Architectural Design for GE Integration
Data Source Layer
Load raw data using the same utilities the handbook employs in its Spark or Databricks notebooks. The intermediate-bootcamp/materials/6-data-impact-training/data/events.csv file serves as an ideal test dataset.
Expectation Suite
Define suites of expectations in a dedicated expectations/ package. This follows the handbook's pattern of separating configuration from execution logic.
Validation Runner
Create a reusable script at validation/run_validation.py that:
- Instantiates a GE
DataContext - Loads datasets as GE
Batchobjects - Executes expectation suites
- Emits JSON or HTML reports
CI/CD Integration
Hook the validation runner into existing workflows or boot-camp notebooks to provide immediate learner feedback.
Step-by-Step Implementation
1. Add the Dependency
Update your requirements.txt following the pattern in intermediate-bootcamp/materials/5-kpis-and-experimentation/requirements.txt:
# Add to requirements.txt
great-expectations>=0.18
2. Create the Validation Structure
mkdir -p intermediate-bootcamp/materials/6-data-impact-training/validation/reports
touch intermediate-bootcamp/materials/6-data-impact-training/validation/__init__.py
3. Write the Validation Script
The following implementation follows the handbook's folder conventions and uses the boot-camp's events.csv sample data:
# File: intermediate-bootcamp/materials/6-data-impact-training/validation/run_validation.py
import great_expectations as ge
from great_expectations.checkpoint import SimpleCheckpoint
from pathlib import Path
# ------------------------------------------------------------------
# 1️⃣ Initialise a DataContext (in-memory for this example)
# ------------------------------------------------------------------
context = ge.get_context()
# ------------------------------------------------------------------
# 2️⃣ Define a simple Expectation Suite (or load an existing one)
# ------------------------------------------------------------------
suite_name = "events_suite"
if not context.suites.get_suite_names().contains(suite_name):
suite = context.add_expectation_suite(suite_name)
else:
suite = context.get_expectation_suite(suite_name)
# ------------------------------------------------------------------
# 3️⃣ Add expectations (column existence, type, non-null, etc.)
# ------------------------------------------------------------------
suite.add_expectation(
expectation_type="expect_table_columns_to_match_ordered_list",
kwargs={"column_list": ["event_id", "timestamp", "device_id", "event_type"]},
)
suite.add_expectation(
expectation_type="expect_column_values_to_not_be_null",
kwargs={"column": "event_id"},
)
suite.add_expectation(
expectation_type="expect_column_values_to_be_of_type",
kwargs={"column": "timestamp", "type_": "datetime"},
)
# Save the suite for reuse
context.save_expectation_suite(suite, overwrite_existing=True)
# ------------------------------------------------------------------
# 4️⃣ Create a Batch (load the CSV used in the boot-camp)
# ------------------------------------------------------------------
data_path = Path(__file__).parents[2] / "data" / "events.csv"
batch = context.get_batch(
batch_kwargs={"path": str(data_path), "datasource": "filesystem"},
expectation_suite_name=suite_name,
)
# ------------------------------------------------------------------
# 5️⃣ Run a SimpleCheckpoint to validate and produce a report
# ------------------------------------------------------------------
checkpoint = SimpleCheckpoint(
name="events_check",
data_context=context,
expectation_suite_name=suite_name,
batch_list=[batch],
)
# Execute validation
result = checkpoint.run()
# ------------------------------------------------------------------
# 6️⃣ Export an HTML report (handy for boot-camp review)
# ------------------------------------------------------------------
report_path = Path(__file__).parent / "reports" / "events_validation.html"
result.get("validation_result").save_expectation_suite(
expectation_suite_name=suite_name,
format="html",
path=str(report_path)
)
print(f"✅ Validation complete – report saved to {report_path}")
4. Run the Validation
cd intermediate-bootcamp/materials/6-data-impact-training/validation
python run_validation.py
Key Implementation Details
In-memory vs. file-based context: The example uses ge.get_context() for simplicity. For production work, switch to ge.data_context.DataContext(project_config="great_expectations.yml") to persist configuration.
Programmatic vs. CLI suite creation: The script adds expectations programmatically. Alternatively, use great_expectations suite new to generate JSON suites and commit them to version control.
Report consumption: The generated HTML report opens directly in browsers, giving learners visual feedback on data quality failures—critical for the handbook's educational mission.
Extending to Other Boot-Camp Modules
Apply this same pattern to other handbook datasets:
- Reuse the
validation/directory structure acrossintermediate-bootcamp/materials/* - Parameterize
suite_nameanddata_pathvia CLI arguments or environment variables - Store baseline expectation suites as JSON files in
expectations/for version control
Key Files in the Repository
| File | Purpose |
|---|---|
README.md |
Data Quality tools list including GE recommendation |
intermediate-bootcamp/materials/6-data-impact-training/data/events.csv |
Sample dataset for validation testing |
intermediate-bootcamp/materials/5-kpis-and-experimentation/requirements.txt |
Dependency listing template |
intermediate-bootcamp/materials/6-data-impact-training/README.md |
Context for data-impact training module |
Summary
- Great Expectations integrates cleanly with the Data Engineer Handbook's modular architecture
- Place validation scripts in
validation/directories following existing job patterns - Use
SimpleCheckpointfor straightforward validation and reporting - Target
events.csvand similar boot-camp datasets for immediate learner relevance - Generate HTML reports stored under
reports/for visual feedback
Frequently Asked Questions
What version of Great Expectations should I use with the handbook?
Use great-expectations>=0.18 as specified in the dependency files. This version provides the get_context() API and SimpleCheckpoint class used in the implementation. Earlier versions may require different initialization patterns.
Can I use Great Expectations with Spark datasets in the handbook?
Yes. Replace the filesystem batch kwargs with Spark DataFrame configuration. The handbook's Spark modules in intermediate-bootcamp/materials/3-spark-fundamentals/ demonstrate DataFrame patterns you can adapt—pass your Spark session to GE's SparkDFExecutionEngine via datasource configuration.
How do I persist expectation suites across notebook sessions?
Switch from in-memory context to file-based: context = ge.data_context.DataContext(project_config="great_expectations.yml"). This creates a great_expectations/ directory with version-controlled suites, or generate suites via CLI (great_expectations suite new) and load them programmatically with context.get_expectation_suite().
Where should validation reports be stored for boot-camp review?
Save reports to validation/reports/ as shown in the example. This location is accessible to learners completing the 6-data-impact-training module and follows the handbook's convention of separating outputs from source code. For CI integration, additionally emit JSON to reports/ at the repository root.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →