# How to Implement Data Quality Checks Using Great Expectations in the Data Engineer Handbook

> Learn to implement data quality checks with Great Expectations. Integrate this powerful Python library into your data pipelines for robust validation and documentation. Perfect for data engineers.

- Repository: [DataExpert.io/data-engineer-handbook](https://github.com/DataExpert-io/data-engineer-handbook)
- Tags: how-to-guide
- Published: 2026-08-06

---

**Great Expectations is an open-source Python library for defining, executing, and documenting data validation tests that you can integrate into the Data Expert-io handbook to create reusable, modular data quality pipelines.**

This guide shows you how to implement **data quality checks using Great Expectations** within the [DataEngineer-io/data-engineer-handbook](https://github.com/DataExpert-io/data-engineer-handbook) repository structure. While the handbook lists GE as a recommended tool in its *Data Quality* section (see [`README.md`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/README.md) lines 75-80), it does not yet contain a ready-made implementation—so you'll build one following the same architectural patterns used for Spark and Flink boot-camp modules.

## Where Great Expectations Fits in the Handbook

The handbook's [`README.md`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/README.md) explicitly recommends GE for data quality work. You can extend this by creating a lightweight validation pipeline that mirrors the modular job structure found in `intermediate-bootcamp/materials/3-spark-fundamentals/src/jobs/`.

Following the handbook's design principles of *clean, modular, and reproducible* data pipelines, you'll keep validation logic isolated from transformation jobs. This approach allows reuse across multiple lessons and provides immediate feedback to learners.

## Architectural Design for GE Integration

### Data Source Layer

Load raw data using the same utilities the handbook employs in its Spark or Databricks notebooks. The `intermediate-bootcamp/materials/6-data-impact-training/data/events.csv` file serves as an ideal test dataset.

### Expectation Suite

Define suites of expectations in a dedicated `expectations/` package. This follows the handbook's pattern of separating configuration from execution logic.

### Validation Runner

Create a reusable script at [`validation/run_validation.py`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/validation/run_validation.py) that:
- Instantiates a GE `DataContext`
- Loads datasets as GE `Batch` objects
- Executes expectation suites
- Emits JSON or HTML reports

### CI/CD Integration

Hook the validation runner into existing workflows or boot-camp notebooks to provide immediate learner feedback.

## Step-by-Step Implementation

### 1. Add the Dependency

Update your [`requirements.txt`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/requirements.txt) following the pattern in [`intermediate-bootcamp/materials/5-kpis-and-experimentation/requirements.txt`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/5-kpis-and-experimentation/requirements.txt):

```bash

# Add to requirements.txt

great-expectations>=0.18

```

### 2. Create the Validation Structure

```bash
mkdir -p intermediate-bootcamp/materials/6-data-impact-training/validation/reports
touch intermediate-bootcamp/materials/6-data-impact-training/validation/__init__.py

```

### 3. Write the Validation Script

The following implementation follows the handbook's folder conventions and uses the boot-camp's `events.csv` sample data:

```python

# File: intermediate-bootcamp/materials/6-data-impact-training/validation/run_validation.py

import great_expectations as ge
from great_expectations.checkpoint import SimpleCheckpoint
from pathlib import Path


# ------------------------------------------------------------------

# 1️⃣ Initialise a DataContext (in-memory for this example)

# ------------------------------------------------------------------

context = ge.get_context()


# ------------------------------------------------------------------

# 2️⃣ Define a simple Expectation Suite (or load an existing one)

# ------------------------------------------------------------------

suite_name = "events_suite"
if not context.suites.get_suite_names().contains(suite_name):
    suite = context.add_expectation_suite(suite_name)
else:
    suite = context.get_expectation_suite(suite_name)


# ------------------------------------------------------------------

# 3️⃣ Add expectations (column existence, type, non-null, etc.)

# ------------------------------------------------------------------

suite.add_expectation(
    expectation_type="expect_table_columns_to_match_ordered_list",
    kwargs={"column_list": ["event_id", "timestamp", "device_id", "event_type"]},
)
suite.add_expectation(
    expectation_type="expect_column_values_to_not_be_null",
    kwargs={"column": "event_id"},
)
suite.add_expectation(
    expectation_type="expect_column_values_to_be_of_type",
    kwargs={"column": "timestamp", "type_": "datetime"},
)

# Save the suite for reuse

context.save_expectation_suite(suite, overwrite_existing=True)


# ------------------------------------------------------------------

# 4️⃣ Create a Batch (load the CSV used in the boot-camp)

# ------------------------------------------------------------------

data_path = Path(__file__).parents[2] / "data" / "events.csv"
batch = context.get_batch(
    batch_kwargs={"path": str(data_path), "datasource": "filesystem"},
    expectation_suite_name=suite_name,
)


# ------------------------------------------------------------------

# 5️⃣ Run a SimpleCheckpoint to validate and produce a report

# ------------------------------------------------------------------

checkpoint = SimpleCheckpoint(
    name="events_check",
    data_context=context,
    expectation_suite_name=suite_name,
    batch_list=[batch],
)

# Execute validation

result = checkpoint.run()


# ------------------------------------------------------------------

# 6️⃣ Export an HTML report (handy for boot-camp review)

# ------------------------------------------------------------------

report_path = Path(__file__).parent / "reports" / "events_validation.html"
result.get("validation_result").save_expectation_suite(
    expectation_suite_name=suite_name, 
    format="html", 
    path=str(report_path)
)

print(f"✅ Validation complete – report saved to {report_path}")

```

### 4. Run the Validation

```bash
cd intermediate-bootcamp/materials/6-data-impact-training/validation
python run_validation.py

```

## Key Implementation Details

**In-memory vs. file-based context**: The example uses `ge.get_context()` for simplicity. For production work, switch to `ge.data_context.DataContext(project_config="great_expectations.yml")` to persist configuration.

**Programmatic vs. CLI suite creation**: The script adds expectations programmatically. Alternatively, use `great_expectations suite new` to generate JSON suites and commit them to version control.

**Report consumption**: The generated HTML report opens directly in browsers, giving learners visual feedback on data quality failures—critical for the handbook's educational mission.

## Extending to Other Boot-Camp Modules

Apply this same pattern to other handbook datasets:

- Reuse the `validation/` directory structure across `intermediate-bootcamp/materials/*`
- Parameterize `suite_name` and `data_path` via CLI arguments or environment variables
- Store baseline expectation suites as JSON files in `expectations/` for version control

## Key Files in the Repository

| File | Purpose |
|------|---------|
| [`README.md`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/README.md) | Data Quality tools list including GE recommendation |
| `intermediate-bootcamp/materials/6-data-impact-training/data/events.csv` | Sample dataset for validation testing |
| [`intermediate-bootcamp/materials/5-kpis-and-experimentation/requirements.txt`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/5-kpis-and-experimentation/requirements.txt) | Dependency listing template |
| [`intermediate-bootcamp/materials/6-data-impact-training/README.md`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/6-data-impact-training/README.md) | Context for data-impact training module |

## Summary

- **Great Expectations integrates cleanly** with the Data Engineer Handbook's modular architecture
- **Place validation scripts** in `validation/` directories following existing job patterns
- **Use `SimpleCheckpoint`** for straightforward validation and reporting
- **Target `events.csv`** and similar boot-camp datasets for immediate learner relevance
- **Generate HTML reports** stored under `reports/` for visual feedback

## Frequently Asked Questions

### What version of Great Expectations should I use with the handbook?

Use `great-expectations>=0.18` as specified in the dependency files. This version provides the `get_context()` API and `SimpleCheckpoint` class used in the implementation. Earlier versions may require different initialization patterns.

### Can I use Great Expectations with Spark datasets in the handbook?

Yes. Replace the filesystem batch kwargs with Spark DataFrame configuration. The handbook's Spark modules in `intermediate-bootcamp/materials/3-spark-fundamentals/` demonstrate DataFrame patterns you can adapt—pass your Spark session to GE's `SparkDFExecutionEngine` via datasource configuration.

### How do I persist expectation suites across notebook sessions?

Switch from in-memory context to file-based: `context = ge.data_context.DataContext(project_config="great_expectations.yml")`. This creates a `great_expectations/` directory with version-controlled suites, or generate suites via CLI (`great_expectations suite new`) and load them programmatically with `context.get_expectation_suite()`.

### Where should validation reports be stored for boot-camp review?

Save reports to `validation/reports/` as shown in the example. This location is accessible to learners completing the `6-data-impact-training` module and follows the handbook's convention of separating outputs from source code. For CI integration, additionally emit JSON to `reports/` at the repository root.