How to Design KPIs and Run A/B Testing Experiments in Data Engineering: A Complete Guide

Data engineers design KPIs and run A/B testing experiments by defining leading and lagging metrics, implementing deterministic user assignment via hashing, and logging conversion events through experimentation platforms like Statsig to measure statistical significance.

Designing effective Key Performance Indicators (KPIs) and conducting reliable A/B tests are essential responsibilities for data engineers who want to drive data-informed product decisions. According to the DataExpert-io/data-engineer-handbook repository, this process involves structuring experiments with clear hypotheses, implementing proper randomization logic, and building event collection pipelines that feed into statistical analysis frameworks.

Define Business Goals and Metrics

Every experiment starts with translating business objectives into measurable metrics. The intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md file outlines the Spotify approach, which requires distinguishing between metric types and formalizing hypotheses before writing code.

Leading vs. Lagging Metrics

Leading metrics provide early signals that predict eventual outcomes, such as "number of users who initiate checkout." Lagging metrics represent the ultimate business impact, such as "total revenue from completed purchases." You should define both metric types for every experiment to enable rapid iteration while validating long-term value.

Hypothesis Formulation

Write explicit statistical hypotheses before launching any test:

  • Null hypothesis: The change does not improve the target metric (e.g., "The new combo deal does not improve sign-up revenue").
  • Alternative hypothesis: The change does improve the target metric (e.g., "The new combo deal does improve sign-up revenue").

The repository documents three distinct experiments, each specifying objectives, null and alternative hypotheses, leading and lagging metrics, and a 50/50 test-cell allocation in the README.

Structure Experiments with an Experimentation Platform

Rather than building custom randomization logic, integrate a managed experimentation service such as Statsig to handle user bucketing and event aggregation. This approach ensures statistical rigor while reducing engineering overhead.

User Identification and Deterministic Assignment

Consistent user assignment prevents exposure bias across sessions. The implementation in intermediate-bootcamp/materials/5-kpis-and-experimentation/src/server.py demonstrates deterministic user ID generation:

  • Generate a hash from a stable identifier (e.g., request.remote_addr) to ensure the same user always receives the same variant
  • Fall back to a random integer when the ?random query parameter is present for manual testing purposes
  • Convert the hash to a string to create the user_id passed to the experimentation platform

Event Logging and Conversion Tracking

Capture exposure and conversion events using the platform's SDK. The statsig.log_event() method streams data to the backend for aggregation, enabling automatic calculation of conversion rates and confidence intervals without manual SQL queries.

Implement the Experimentation Pipeline

The Flask application in server.py demonstrates the end-to-end integration pattern for serving experiment-driven UIs and tracking conversions.

Retrieve Experiment Variants

Use statsig.get_experiment() to fetch configuration parameters for the current user session:

from flask import Flask, request
from statsig import statsig
from statsig.statsig_user import StatsigUser
import random, os

statsig.initialize(os.getenv('STATSIG_API_KEY'))
app = Flask(__name__)

@app.route('/tasks')
def get_tasks():
    # Identify user with deterministic hash

    hash_string = request.remote_addr
    if request.args.get('random'):
        hash_string = str(random.randint(0, 1_000_000))
    user_id = str(hash(hash_string))
    
    # Fetch experiment configuration

    experiment = statsig.get_experiment(StatsigUser(user_id), "button_color_v3")
    color = experiment.get("Button Color", "blue")
    paragraph_text = experiment.get("Paragraph Text", "Data Engineering Boot Camp")
    
    return f"""
        <div style="background:{color}">
            <h1>Experiment group: {color}</h1>
            <h5>{paragraph_text}</h5>
        </div>
    """

The code retrieves the Button Color and Paragraph Text variants from the "button_color_v3" experiment and renders the UI accordingly.

Log Conversion Events

Track key actions to measure leading metrics using statsig.log_event():

from statsig.statsig_event import StatsigEvent

@app.route('/signup')
def signup():
    # Consistent user identification logic

    hash_string = request.remote_addr
    if request.args.get('random'):
        hash_string = str(random.randint(0, 1_000_000))
    user_id = str(hash(hash_string))
    
    # Log the conversion event

    statsig_user = StatsigUser(user_id)
    statsig.log_event(
        StatsigEvent(
            user=statsig_user,
            event_name='signup_completed'
        )
    )
    return "Signup complete – event logged"

This pattern enables the experimentation platform to automatically calculate conversion rates for the leading metric "users who complete signup."

Analyze Results and Validate Statistical Significance

After collecting sufficient data, perform statistical analysis to accept or reject the null hypothesis:

  1. Extract metrics from the experimentation platform's aggregated datasets (exposure counts and conversion events)
  2. Apply statistical tests such as t-tests for continuous metrics or chi-square tests for categorical conversion rates to determine significance
  3. Validate lagging metrics only after leading metrics show statistically significant movement, ensuring the sample size is large enough to detect revenue impacts

The requirements.txt file in the repository lists statsig as a dependency, indicating the platform handles the statistical calculations internally while data warehouses like Snowflake or BigQuery store raw event data for offline validation.

Deploy Winning Variants and Iterate

Once analysis confirms the alternative hypothesis with statistical significance:

  • Roll out the winning variant to 100% of users via the feature-flag service
  • Archive experiment documentation including objectives, hypotheses, allocation strategy, and results in version-controlled directories
  • Design follow-up experiments by refining hypotheses or testing adjacent variations based on insights from lagging metric analysis

Summary

  • Define dual metrics: Structure every experiment with leading metrics for rapid feedback and lagging metrics for business impact validation.
  • Use deterministic hashing: Generate consistent user IDs via str(hash(remote_addr)) to maintain variant assignment across sessions.
  • Leverage managed platforms: Integrate services like Statsig using get_experiment() and log_event() rather than building custom randomization logic.
  • Validate statistically: Apply proper statistical tests to reject the null hypothesis before deploying changes to production.
  • Document thoroughly: Store experiment configurations and results in intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md or similar version-controlled documentation.

Frequently Asked Questions

What is the difference between leading and lagging metrics in A/B testing?

Leading metrics are early indicators that change quickly, such as click-through rates or initiation of a signup flow, allowing you to detect trends before the experiment concludes. Lagging metrics are ultimate outcomes like revenue or retention that take longer to materialize but represent true business value. According to the DataExpert-io/data-engineer-handbook methodology, you should monitor leading metrics for early signals while validating decisions against lagging metrics.

How do you ensure users stay in the same A/B test variant across sessions?

Implement deterministic user assignment by hashing a stable identifier such as IP address (request.remote_addr) to generate a consistent user_id string. As shown in intermediate-bootcamp/materials/5-kpis-and-experimentation/src/server.py, the hash function ensures the same input always produces the same output, bucketing the user into the same variant every time they return.

What statistical methods validate A/B test results?

Data engineers typically use t-tests for comparing means of continuous metrics like revenue per user, or chi-square tests for categorical data like conversion rates. The experimentation platform (e.g., Statsig) calculates confidence intervals and p-values automatically, but you should verify these calculations against your data warehouse records before accepting the alternative hypothesis.

How do you track conversions in an A/B testing pipeline?

Track conversions by calling the experimentation platform's event logging method—specifically statsig.log_event(StatsigEvent(user=statsig_user, event_name='signup_completed'))—when users complete the target action. This approach captures the timestamp, user variant, and event metadata required to calculate conversion rates and statistical significance without complex ETL pipelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →