# How to Design and Measure KPIs for Experimentation: A Complete Data Engineering Guide

> Learn to design and measure KPIs for experimentation with this data engineering guide. Define objectives, track leading and lagging metrics, and instrument your app for effective A/B testing.

- Repository: [DataExpert.io/data-engineer-handbook](https://github.com/DataExpert-io/data-engineer-handbook)
- Tags: how-to-guide
- Published: 2026-08-06

---

**Effective experimentation requires defining clear business objectives, separating leading metrics (early signals) from lagging metrics (ultimate outcomes), and instrumenting your application to capture events at the moment of user interaction.**

Designing robust experimentation frameworks is a core competency for data engineers. This guide walks through the practical implementation in the *DataExpert-io/data-engineer-handbook* repository, which demonstrates KPI-driven A/B testing using a lightweight Flask service integrated with Statsig.

## How to Design KPIs for Experimentation

The handbook's "KPIs and Experimentation" module structures experiment design around four foundational steps documented in [`intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md).

### Define the Business Objective

Start with a concise, measurable goal. The repository uses Spotify-centric examples like "increase sign-up revenue" or "boost podcast engagement" to anchor technical work in product outcomes.

### State Your Hypotheses

Explicitly formulate two competing statements:

- **Null hypothesis (H₀):** The experiment has no effect on the target metric
- **Alternative hypothesis (H₁):** The experiment produces the expected impact

This creates the statistical framework for later analysis.

### Select Leading and Lagging Metrics

**Leading KPIs** provide early, observable signals. The handbook uses "number of users who sign up after launching the Blink combo" as a leading indicator that can be measured within days.

**Lagging KPIs** capture ultimate business value. "Increase in sign-up sale" represents the true north metric that may take weeks or months to stabilize.

### Allocate Test Cells Evenly

Random assignment eliminates selection bias. The repository recommends a **50/50 split** by default to maximize statistical power while maintaining sufficient sample sizes in each variant.

## How to Measure KPIs for Experimentation in Production

The implementation in [`server.py`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/server.py) demonstrates end-to-end instrumentation using Statsig for real-time experiment management.

### Instrument Leading KPI Events

The `/signup` endpoint captures the critical first step in the conversion funnel:

```python
@app.route('/signup')
def signup():
    # Generate a deterministic user identifier

    user_id = str(hash(request.remote_addr))
    statsig_user = StatsigUser(user_id)

    # Record the leading KPI: user reached signup page

    statsig.log_event(
        StatsigEvent(user=statsig_user, event_name='visited_signup')
    )
    return "This is the signup page"

```

This logs the `visited_signup` event (source: [`intermediate-bootcamp/materials/5-kpis-and-experimentation/src/server.py`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/5-kpis-and-experimentation/src/server.py), lines 42-46), creating the raw data needed to compute signup rates per experiment variant.

### Fetch Experiment Variants and Drive UI

The `/tasks` endpoint (lines 57-82) retrieves experiment configuration and applies it to the user experience:

```python
@app.route('/tasks')
def get_tasks():
    user_id = str(hash(request.remote_addr))
    
    # Fetch experiment parameters for this user

    color = statsig.get_experiment(
        StatsigUser(user_id), 
        "button_color_v3"
    ).get("Button Color", "blue")
    
    paragraph_text = statsig.get_experiment(
        StatsigUser(user_id), 
        "button_color_v3"
    ).get("Paragraph Text", "Data Engineering Boot Camp")
    
    # Filter tasks and render variant-specific content

    tasks = create_tasks(int(user_id), str(color))
    return render_template_string(
        html,
        total_tasks=len(tasks),
        tasks=tasks,
        color_1=color,
        paragraph_text=paragraph_text
    )

```

The `statsig.get_experiment()` call deterministically assigns users to variants based on their ID, ensuring consistent experiences across sessions.

### Aggregate KPIs in Your Data Warehouse

Downstream SQL queries transform raw events into actionable metrics:

```sql
-- Leading KPI: signup funnel entry by variant
SELECT
    experiment_variant,
    COUNT(DISTINCT user_id) AS unique_signup_visits,
    COUNT(*) AS total_signup_visits
FROM events
WHERE event_name = 'visited_signup'
  AND timestamp >= '2024-01-01'
GROUP BY experiment_variant;

-- Lagging KPI: revenue attribution to experiment
SELECT
    ea.experiment_variant,
    COUNT(DISTINCT t.user_id) AS converting_users,
    SUM(t.revenue) AS total_revenue,
    AVG(t.revenue) AS revenue_per_user
FROM transactions t
JOIN experiment_assignments ea 
  ON t.user_id = ea.user_id
WHERE ea.experiment_name = 'button_color_v3'
GROUP BY ea.experiment_variant;

```

## Best Practices for KPI-Driven Experimentation

Based on the handbook's implementation, follow this operational checklist:

| Step | Action | Implementation Detail |
|:---|:---|:---|
| **1. Objective** | Write a single-sentence business goal | Aligns engineering work with product outcomes |
| **2. Hypotheses** | Document H₀ and H₁ before launch | Prevents post-hoc rationalization |
| **3. Metrics** | Define ≥1 leading and ≥1 lagging KPI | Enables both early detection and final validation |
| **4. Randomization** | Use deterministic, user-level assignment | `hash(user_id) % 100 < 50` for 50/50 splits |
| **5. Instrumentation** | Log events synchronously at interaction point | `statsig.log_event()` in request handler |
| **6. Pipeline** | Stream events to durable storage | Kafka → Snowflake/BigQuery pattern |
| **7. Analysis** | Apply appropriate statistical tests | t-tests for continuous metrics, chi-square for rates |
| **8. Decision** | Roll out winners, document null results | Prevents experiment pollution and knowledge loss |

## Summary

- **Design KPIs for experimentation** by pairing leading metrics (early, frequent) with lagging metrics (ultimate business impact) and explicitly stating testable hypotheses.
- **Measure KPIs for experimentation** by instrumenting your application to log events at user interaction points, using tools like Statsig for variant assignment and event collection.
- The Data Engineer Handbook provides a complete reference implementation in [`server.py`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/server.py) demonstrating Flask + Statsig integration with deterministic user assignment and synchronous event logging.
- Aggregate raw events in your data warehouse to compute conversion rates, revenue per variant, and statistical significance.

## Frequently Asked Questions

### What is the difference between leading and lagging KPIs in experimentation?

**Leading KPIs** are early indicators you can measure quickly—like page visits or button clicks—that predict eventual success. **Lagging KPIs** are the ultimate business outcomes, such as revenue or customer lifetime value, that take longer to materialize but represent true experimental impact. The handbook's signup example uses `visited_signup` (leading) and `sign-up sale` (lagging) to capture both time horizons.

### Why must experiment cells be allocated evenly?

Even allocation (typically 50/50) maximizes **statistical power**—the probability of detecting a true effect when one exists. Uneven splits reduce the effective sample size in the smaller cell, requiring longer runtimes to reach significance. The handbook's [`intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md) specifies this as default practice unless business constraints demand otherwise.

### How does Statsig enable real-time KPI measurement?

Statsig provides two critical capabilities demonstrated in [`server.py`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/server.py): `get_experiment()` for **deterministic variant assignment** based on user IDs, and `log_event()` for **synchronous event capture**. This combination lets you expose users to treatment immediately while streaming structured events to downstream analytics systems for KPI computation.

### What instrumentation is required to capture leading KPIs?

You must insert **event logging calls at precise interaction points** in your application code. The handbook's implementation logs `visited_signup` directly in the `/signup` route handler, ensuring the event fires exactly when the user engages with the signup flow—neither earlier (inflated counts) nor later (missed sessions due to page abandonment).