How to Design and Measure KPIs for Experimentation: A Complete Data Engineering Guide
Effective experimentation requires defining clear business objectives, separating leading metrics (early signals) from lagging metrics (ultimate outcomes), and instrumenting your application to capture events at the moment of user interaction.
Designing robust experimentation frameworks is a core competency for data engineers. This guide walks through the practical implementation in the DataExpert-io/data-engineer-handbook repository, which demonstrates KPI-driven A/B testing using a lightweight Flask service integrated with Statsig.
How to Design KPIs for Experimentation
The handbook's "KPIs and Experimentation" module structures experiment design around four foundational steps documented in intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md.
Define the Business Objective
Start with a concise, measurable goal. The repository uses Spotify-centric examples like "increase sign-up revenue" or "boost podcast engagement" to anchor technical work in product outcomes.
State Your Hypotheses
Explicitly formulate two competing statements:
- Null hypothesis (H₀): The experiment has no effect on the target metric
- Alternative hypothesis (H₁): The experiment produces the expected impact
This creates the statistical framework for later analysis.
Select Leading and Lagging Metrics
Leading KPIs provide early, observable signals. The handbook uses "number of users who sign up after launching the Blink combo" as a leading indicator that can be measured within days.
Lagging KPIs capture ultimate business value. "Increase in sign-up sale" represents the true north metric that may take weeks or months to stabilize.
Allocate Test Cells Evenly
Random assignment eliminates selection bias. The repository recommends a 50/50 split by default to maximize statistical power while maintaining sufficient sample sizes in each variant.
How to Measure KPIs for Experimentation in Production
The implementation in server.py demonstrates end-to-end instrumentation using Statsig for real-time experiment management.
Instrument Leading KPI Events
The /signup endpoint captures the critical first step in the conversion funnel:
@app.route('/signup')
def signup():
# Generate a deterministic user identifier
user_id = str(hash(request.remote_addr))
statsig_user = StatsigUser(user_id)
# Record the leading KPI: user reached signup page
statsig.log_event(
StatsigEvent(user=statsig_user, event_name='visited_signup')
)
return "This is the signup page"
This logs the visited_signup event (source: intermediate-bootcamp/materials/5-kpis-and-experimentation/src/server.py, lines 42-46), creating the raw data needed to compute signup rates per experiment variant.
Fetch Experiment Variants and Drive UI
The /tasks endpoint (lines 57-82) retrieves experiment configuration and applies it to the user experience:
@app.route('/tasks')
def get_tasks():
user_id = str(hash(request.remote_addr))
# Fetch experiment parameters for this user
color = statsig.get_experiment(
StatsigUser(user_id),
"button_color_v3"
).get("Button Color", "blue")
paragraph_text = statsig.get_experiment(
StatsigUser(user_id),
"button_color_v3"
).get("Paragraph Text", "Data Engineering Boot Camp")
# Filter tasks and render variant-specific content
tasks = create_tasks(int(user_id), str(color))
return render_template_string(
html,
total_tasks=len(tasks),
tasks=tasks,
color_1=color,
paragraph_text=paragraph_text
)
The statsig.get_experiment() call deterministically assigns users to variants based on their ID, ensuring consistent experiences across sessions.
Aggregate KPIs in Your Data Warehouse
Downstream SQL queries transform raw events into actionable metrics:
-- Leading KPI: signup funnel entry by variant
SELECT
experiment_variant,
COUNT(DISTINCT user_id) AS unique_signup_visits,
COUNT(*) AS total_signup_visits
FROM events
WHERE event_name = 'visited_signup'
AND timestamp >= '2024-01-01'
GROUP BY experiment_variant;
-- Lagging KPI: revenue attribution to experiment
SELECT
ea.experiment_variant,
COUNT(DISTINCT t.user_id) AS converting_users,
SUM(t.revenue) AS total_revenue,
AVG(t.revenue) AS revenue_per_user
FROM transactions t
JOIN experiment_assignments ea
ON t.user_id = ea.user_id
WHERE ea.experiment_name = 'button_color_v3'
GROUP BY ea.experiment_variant;
Best Practices for KPI-Driven Experimentation
Based on the handbook's implementation, follow this operational checklist:
| Step | Action | Implementation Detail |
|---|---|---|
| 1. Objective | Write a single-sentence business goal | Aligns engineering work with product outcomes |
| 2. Hypotheses | Document H₀ and H₁ before launch | Prevents post-hoc rationalization |
| 3. Metrics | Define ≥1 leading and ≥1 lagging KPI | Enables both early detection and final validation |
| 4. Randomization | Use deterministic, user-level assignment | hash(user_id) % 100 < 50 for 50/50 splits |
| 5. Instrumentation | Log events synchronously at interaction point | statsig.log_event() in request handler |
| 6. Pipeline | Stream events to durable storage | Kafka → Snowflake/BigQuery pattern |
| 7. Analysis | Apply appropriate statistical tests | t-tests for continuous metrics, chi-square for rates |
| 8. Decision | Roll out winners, document null results | Prevents experiment pollution and knowledge loss |
Summary
- Design KPIs for experimentation by pairing leading metrics (early, frequent) with lagging metrics (ultimate business impact) and explicitly stating testable hypotheses.
- Measure KPIs for experimentation by instrumenting your application to log events at user interaction points, using tools like Statsig for variant assignment and event collection.
- The Data Engineer Handbook provides a complete reference implementation in
server.pydemonstrating Flask + Statsig integration with deterministic user assignment and synchronous event logging. - Aggregate raw events in your data warehouse to compute conversion rates, revenue per variant, and statistical significance.
Frequently Asked Questions
What is the difference between leading and lagging KPIs in experimentation?
Leading KPIs are early indicators you can measure quickly—like page visits or button clicks—that predict eventual success. Lagging KPIs are the ultimate business outcomes, such as revenue or customer lifetime value, that take longer to materialize but represent true experimental impact. The handbook's signup example uses visited_signup (leading) and sign-up sale (lagging) to capture both time horizons.
Why must experiment cells be allocated evenly?
Even allocation (typically 50/50) maximizes statistical power—the probability of detecting a true effect when one exists. Uneven splits reduce the effective sample size in the smaller cell, requiring longer runtimes to reach significance. The handbook's intermediate-bootcamp/materials/5-kpis-and-experimentation/README.md specifies this as default practice unless business constraints demand otherwise.
How does Statsig enable real-time KPI measurement?
Statsig provides two critical capabilities demonstrated in server.py: get_experiment() for deterministic variant assignment based on user IDs, and log_event() for synchronous event capture. This combination lets you expose users to treatment immediately while streaming structured events to downstream analytics systems for KPI computation.
What instrumentation is required to capture leading KPIs?
You must insert event logging calls at precise interaction points in your application code. The handbook's implementation logs visited_signup directly in the /signup route handler, ensuring the event fires exactly when the user engages with the signup flow—neither earlier (inflated counts) nor later (missed sessions due to page abandonment).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →