Data Quality Testing and Monitoring: A Practical Guide from the Data Engineer Handbook

A robust data quality strategy combines automated testing embedded in CI/CD pipelines with continuous production monitoring to catch defects early and detect drift.

Ensuring trustworthy data requires a systematic approach to data quality testing and monitoring. According to the DataExpert-io/data-engineer-handbook, modern data teams treat quality checks as first-class components of their infrastructure, integrating validation directly into ETL/ELT workflows and maintaining vigilant downstream observability. This guide walks through the practical implementation patterns documented in the handbook's learning materials.

Define Clear Expectations for Data Quality

Start by documenting the shape and behavior of your data. According to the handbook's methodology, comprehensive expectations should cover:

  • Schema compliance: Column names, data types, and structural integrity
  • Value constraints: Acceptable ranges or enumerations (e.g., age ≥ 0, country ∈ ISO list)
  • Uniqueness constraints: Primary-key validation and duplicate detection
  • Referential integrity: Relationships between tables and foreign key constraints

Embed Tests in the ETL/ELT Pipeline

Treat data tests like code tests by executing them on every build or deployment. The README.md in the DataExpert-io/data-engineer-handbook repository curates several frameworks for implementing these checks [L75-L84].

dbt for Declarative SQL Testing

dbt enables declarative SQL tests that integrate seamlessly into CI/CD pipelines. Define constraints directly in your project configuration:


# models/users.sql (dbt model)

select *
from {{ source('raw', 'users') }}

# tests/users_not_null.yml

version: 2
models:
  - name: users
    columns:
      - name: user_id
        tests:
          - not_null
          - unique

Great Expectations for Python-Based Validation

Great Expectations provides Python-based suites that generate validation reports and integrate with orchestrators like Airflow or Prefect:

import great_expectations as ge

df = ge.read_csv("data/events.csv")
suite = df.validate(
    {
        "expect_column_values_to_not_be_null": {"column": "event_id"},
        "expect_column_values_to_be_in_type_list": {"column": "event_timestamp", "type_list": ["datetime"]},
        "expect_column_values_to_be_between": {"column": "quantity", "min_value": 0, "max_value": 1000},
    }
)
suite.save_expectation_suite("events_suite.json")

Commercial and OSS Platforms

The handbook also references specialized platforms such as Metaplane, Soda, and DQOps, which offer dedicated data quality dashboards and alerting capabilities for organizations requiring enterprise-grade observability.

Implement Continuous Monitoring and Alerting

After deployment, establish downstream monitors that regularly sample data against defined expectations. Modern observability stacks like OpenLineage and Streamdal surface anomalies in near-real-time dashboards and trigger alerts via Slack, PagerDuty, or email.

For custom implementations, Python-based monitors can run on scheduled intervals through Airflow or Prefect:

import pandas as pd
import sqlalchemy as sa
import slack_sdk

engine = sa.create_engine("postgresql://user:pwd@db:5432/warehouse")
df = pd.read_sql("SELECT COUNT(*) AS cnt, AVG(price) AS avg_price FROM sales", engine)

if df["cnt"].iloc[0] == 0 or df["avg_price"].iloc[0] < 0:
    client = slack_sdk.WebClient(token="xoxb-...")
    client.chat_postMessage(
        channel="#data-quality",
        text="🚨 Data quality alert: sales table empty or negative average price!"
    )

Build Feedback Loops for Continuous Improvement

When tests fail or monitors flag anomalies, capture the root cause and update expectations if business rules have legitimately changed. Version-control your test suites to create a self-correcting data quality loop that evolves with your product.

Follow the Data Quality Patterns Learning Path

The intermediate-bootcamp/introduction.md file outlines a dedicated "Data Quality Patterns" bootcamp that provides concrete lectures and labs walking through the end-to-end process of defining, testing, and monitoring data quality [L36-L38]. These hands-on exercises provide ready-to-use code snippets and a disciplined workflow for implementing the patterns discussed above.

Additional resources in the repository include:

Summary

  • Define expectations covering schema, values, uniqueness, and referential integrity before writing tests
  • Embed tests directly into ETL/ELT pipelines using dbt, Great Expectations, or specialized platforms
  • Monitor continuously in production with automated sampling and alerting through observability tools
  • Iterate and version-control test suites to maintain alignment with evolving business rules
  • Leverage the handbook's bootcamp materials for hands-on implementation guidance

Frequently Asked Questions

What is the difference between data testing and data monitoring?

Data testing refers to validation checks that run during development or deployment to catch defects before they reach production, typically integrated into CI/CD pipelines. Data monitoring involves continuously observing production data for drift, anomalies, or failures after deployment, often using sampling techniques and real-time alerting systems.

How do I choose between dbt tests and Great Expectations?

Choose dbt when working primarily in SQL-based warehouses and needing simple, declarative constraints like non-null checks or referential integrity. Choose Great Expectations when requiring complex Python-based validations, rich documentation generation, or sophisticated statistical checks across diverse data formats including CSVs and dataframes.

Where can I find hands-on labs for data quality implementation?

The DataExpert-io/data-engineer-handbook repository contains a "Data Quality Patterns" bootcamp outlined in intermediate-bootcamp/introduction.md that provides structured lectures and labs [L36-L38]. Additionally, the Spark fundamentals homework in intermediate-bootcamp/materials/3-spark-fundamentals/homework/homework_testing.md offers big-data-specific testing exercises.

How do I handle schema changes in data quality tests?

When schemas change legitimately, update your expectation suites to reflect the new business rules and version-control these changes alongside your pipeline code. Implement contract testing between upstream producers and downstream consumers to detect breaking schema changes during the CI/CD phase before they impact production monitors.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →