# How to Implement KPIs and Experimentation Frameworks in Data Engineering

> Learn to implement KPIs and experimentation frameworks. Define metrics, form hypotheses, and use feature flags for A/B testing. Reference the DataExpert-io data-engineer-handbook.

- Repository: [DataExpert.io/data-engineer-handbook](https://github.com/DataExpert-io/data-engineer-handbook)
- Tags: how-to-guide
- Published: 2026-08-08

---

**Implement KPIs and experimentation frameworks by defining leading and lagging metrics, formulating clear null and alternative hypotheses, and integrating feature-flag services like Statsig to randomize users and capture events, as demonstrated in the DataExpert-io/data-engineer-handbook reference implementation.**

The DataExpert-io/data-engineer-handbook provides a practical blueprint for data engineers to build lightweight, metric-driven experimentation systems. This guide walks through the core components of how to implement KPIs and experimentation frameworks using the repository's Flask-based reference architecture and Statsig integration.

## Defining Leading and Lagging KPIs

Effective experimentation starts with separating **leading metrics** (early signals) from **lagging metrics** (business impact). According to the handbook's KPI section in [`intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md), a leading metric might be "Number of users who sign up after a combo deal," while the corresponding lagging metric is "Increase in sign-up revenue"【^1†L12-L15】.

Leading indicators provide rapid feedback on user behavior changes, allowing teams to detect trends before statistical significance on revenue metrics emerges. Lagging indicators validate whether those behavioral changes actually drove business value, typically requiring longer observation periods and batch analytical jobs to calculate.

## Structuring Hypotheses for Data Experiments

Every experiment requires clearly defined null and alternative hypotheses. The handbook references a Spotify case study in [`intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md) that demonstrates this practice across three distinct experiments【^1†L8-L40】.

A well-formed hypothesis structure enables precise metric selection and post-hoc analysis. When you state the expected impact on specific KPIs before launch, you eliminate confirmation bias and establish objective criteria for rollout decisions.

## Implementing Feature Flag Experiments with Statsig

The repository includes a minimal Flask API in [`intermediate-bootcamp/materials/5-kpis-and-experimentation/src/server.py`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/5-kpis-and-experimentation/src/server.py) that integrates **Statsig** for feature-flag-driven experimentation. This implementation shows how to connect UI variations directly to metric collection.

### Initializing the Statsig SDK and User Identification

First, initialize the Statsig client and create a deterministic user identifier based on IP hashing with optional randomization for testing:

```python
import os
import random
from statsig import statsig
from statsig.statsig_user import StatsigUser

API_KEY = os.getenv('STATSIG_API_KEY')
statsig.initialize(API_KEY)

def get_user_id(req):
    # Optional randomization for testing

    random_num = req.args.get('random')
    hash_string = req.remote_addr
    if random_num:
        hash_string = str(random.randint(0, 1_000_000))
    return str(hash(hash_string))

```

This ensures consistent user assignment across sessions while allowing test overrides via query parameters.

### Capturing Leading Metrics via Event Logging

Track user actions that serve as leading indicators by logging events within your application routes. The `/signup` endpoint in [`server.py`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/server.py) demonstrates capturing a page visit event【^1†L42-L47】:

```python
from statsig.statsig_event import StatsigEvent
from flask import request, Flask

app = Flask(__name__)

@app.route('/signup')
def signup():
    user_id = get_user_id(request)
    user = StatsigUser(user_id)
    event = StatsigEvent(user=user, event_name='visited_signup')
    statsig.log_event(event)
    return "This is the signup page"

```

Each `StatsigEvent` payload flows to your data lake, enabling real-time counts of user actions that feed into leading KPI dashboards.

### Randomizing Users and Serving Variants

The core experimentation logic resides in the `/tasks` endpoint, which retrieves experiment configuration and adjusts both UI rendering and backend logic based on the assigned variant【^1†L57-L78】:

```python
@app.route('/tasks')
def get_tasks():
    user_id = get_user_id(request)
    experiment = statsig.get_experiment(StatsigUser(user_id), "button_color_v3")
    
    color = experiment.get("Button Color", "blue")
    paragraph = experiment.get("Paragraph Text", "Data Engineering Boot Camp")
    
    # Filter tasks based on experiment group (example KPI logic)

    filtered = [t for t in tasks if t['id'] % 2 == (0 if color in ('Red', 'Orange') else 1)]
    
    # Render HTML showing the variant and metrics

    return f"""
    <h1>{paragraph}</h1>
    <button style="background-color: {color};">Click Me</button>
    <p>Showing {len(filtered)} tasks</p>
    """

```

This approach ties the feature flag directly to a business logic filter, allowing you to measure how different UI variants affect task engagement—a leading indicator for platform adoption.

## Architectural Layers of the Experimentation Framework

The handbook outlines a four-layer architecture for production experimentation systems:

- **Data Collection Layer**: Captures user actions via `StatsigEvent` objects using `statsig.log_event()` as shown in the `/signup` route. In production, these events stream to systems like Kafka or Kinesis before landing in a data lake.

- **Experiment Management Layer**: Assigns users to variants using `StatsigUser` objects and the `statsig.get_experiment()` method. This ensures randomized, mutually exclusive group assignment across the user base.

- **Metric Calculation Layer**: Computes leading metrics in real-time or near-real-time using streaming aggregates, while lagging metrics like revenue uplift require downstream batch jobs (e.g., Spark or dbt models) that join experiment assignments with transaction tables.

- **Decision Layer**: Compares observed metrics against the pre-defined hypotheses to determine rollout, iteration, or discard. The handbook's homework in [`intermediate-bootcamp/materials/5-kpis-and-experimentation/homework/homework.md`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/5-kpis-and-experimentation/homework/homework.md) prompts learners to "describe 3 experiments … and hypothesize which metrics will be impacted," reinforcing this analytical step【^1†L5-L7】.

## Summary

- **Separate metric types**: Distinguish leading indicators (early user actions) from lagging indicators (revenue impact) to enable rapid iteration while validating business value.
- **Formulate hypotheses first**: Define null and alternative hypotheses before coding, referencing the Spotify case study patterns in the handbook's README.
- **Use feature flags for randomization**: Implement `statsig.get_experiment()` to assign users to variants consistently and isolate the impact of specific changes.
- **Log structured events**: Capture leading metrics via `StatsigEvent` objects at key user interaction points to feed real-time dashboards.
- **Plan for lagging analysis**: Design batch pipelines to calculate business-impact metrics, as real-time systems typically cannot determine revenue attribution immediately.

## Frequently Asked Questions

### What is the difference between leading and lagging KPIs in experimentation?

Leading KPIs measure immediate user actions—such as button clicks or page visits—that indicate engagement shortly after exposure to an experiment variant. Lagging KPIs measure ultimate business outcomes like revenue or retention that may take days or weeks to materialize. The Data Engineer Handbook emphasizes tracking both separately so teams can iterate quickly on leading signals while validating true impact through lagging metrics.

### How does feature flagging enable A/B testing in data engineering?

Feature flagging services like Statsig provide randomized user assignment and variant delivery without requiring separate code deployments. By wrapping logic in `statsig.get_experiment()` calls, data engineers can dynamically alter data pipelines, API responses, or UI elements for specific user segments while maintaining a control group, enabling controlled experiments in production environments.

### What role does Statsig play in the Data Engineer Handbook experimentation framework?

Statsig serves as the experiment randomization and event collection backbone in the handbook's reference implementation. It handles user bucketing through `StatsigUser` objects, variant configuration via the `get_experiment()` method, and metric logging through `log_event()`, allowing the Flask application to focus on business logic rather than statistical infrastructure.

### How do you track business impact metrics in this framework?

Business impact (lagging) metrics are typically computed in downstream analytical systems rather than the experimentation SDK. The Flask application captures raw events and experiment assignments, which are then joined with financial or transactional data in batch jobs (e.g., Spark or SQL models) to calculate metrics like "increase in sign-up revenue" attributed to specific experiment variants.