Monitoring and Observability for Systems: A Complete Design Guide

The system-design-notes repository contains a dedicated chapter that designs a complete metrics-monitoring and alerting system, covering data collection, transmission, storage, alerting, and visualization at scale.

This repository provides a production-ready blueprint for implementing monitoring and observability for systems that handle millions of metrics. The design focuses exclusively on operational telemetry—CPU utilization, memory consumption, and request rates—while explicitly excluding log aggregation and distributed tracing to maintain architectural clarity.

Core Components of the Monitoring Architecture

The architecture in 20. Metrics Monitoring and Alerting System/README.md organizes observability into five distinct layers. Each layer addresses specific scalability challenges while maintaining low-latency data freshness.

Data Collection: Pull vs. Push Models

The design supports dual ingestion strategies to accommodate diverse infrastructure constraints.

Pull-based collection deploys collectors that periodically scrape /metrics endpoints exposed by services. Service discovery mechanisms—implemented via etcd or ZooKeeper—dynamically maintain lists of active endpoints, ensuring collectors target healthy instances only.

Push-based collection installs agents on individual hosts that forward metrics upstream to load-balanced collectors. This model handles sudden traffic spikes gracefully by buffering data locally during network partitions.

The repository provides a detailed comparison table analyzing trade-offs between these models across debugging capability, firewall traversal complexity, latency characteristics, and data authenticity guarantees.

Data Transmission Pipeline

Raw metrics flow through a resilient queuing layer before persistence. The design specifies Kafka (or equivalent ingestion systems) as the primary buffer between collectors and storage. Stream-processing frameworks such as Flink or Spark consume from these queues to perform real-time aggregation, filtering, and format normalization before writing to the time-series database.

Time-Series Storage Strategy

The documentation emphasizes why dedicated time-series databases (InfluxDB, Prometheus) outperform generic relational stores for observability workloads. Key factors include:

  • Low cardinality requirements: Efficient handling of high-frequency data points with limited unique tag combinations
  • Compression algorithms: Specialized encoding for timestamp-value pairs reduces storage footprint by 90% compared to standard databases
  • Down-sampling techniques: Automatic aggregation of historical data to maintain query performance while reducing costs

The storage tier implements multi-resolution retention: raw data persists for 7 days, 1-minute aggregates for 30 days, and 1-hour aggregates for 1 year.

Alerting and Notification System

Alert rules are expressed declaratively in YAML configuration files. The Alert Manager component evaluates these rules against incoming metrics, performs deduplication to prevent notification storms, and implements exponential backoff for delivery retries.

Processed alerts are published back to Kafka, enabling flexible downstream routing to email, PagerDuty, webhooks, or SMS gateways.

Visualization Layer

Grafana serves as the recommended dashboard interface, with optional caching layers inserted between the visualization tier and time-series databases. This caching reduces database load for frequently accessed operational dashboards without sacrificing data freshness for ad-hoc queries.

Design Specifications and Requirements

The system targets 10 million active metrics with a 1-year retention period. These requirements drive specific architectural decisions:

  • Horizontal scaling of collectors to handle ingestion bursts
  • Separation of hot and cold storage tiers
  • Asynchronous processing pipelines to decollect data collection from alerting latency

The scope explicitly excludes log aggregation and distributed tracing, focusing engineering resources on high-cardinality metric collection and dimensional analysis.

Data Model and Collection Implementation

Metrics follow a standardized line protocol compatible with both InfluxDB and Prometheus. Each data point consists of four tuple elements:

  1. Metric name (e.g., CPU.load)
  2. Tags (key-value pairs for dimensions like host and region)
  3. Timestamp (Unix epoch seconds)
  4. Value (64-bit floating point)

# Sample line-protocol metric (compatible with InfluxDB/Prometheus)

CPU.load host=webserver01,region=us-west 1613707265 62
CPU.load host=webserver02,region=us-west 1613707265 43

Service discovery for pull-based scraping maintains endpoint lists in etcd or ZooKeeper, allowing collectors to automatically adapt to auto-scaling events and deployment rollouts.

Code Examples and Configuration

Alert Rule Definition

Alert configurations use YAML syntax with PromQL-style expressions:


# Example alert rule (YAML) – triggers when a host is down for >5 min

- name: instance_down
  rules:
    - alert: instance_down
      expr: up == 0
      for: 5m
      labels:
        severity: page

Time-Series Querying

For analysis of historical trends, the repository demonstrates Flux queries for exponential moving averages:

-- Flux query (InfluxDB) for a 10‑second exponential moving average
from(db:"telegraf")
  |> range(start:-1h)
  |> filter(fn: (r) => r._measurement == "foo")
  |> exponentialMovingAverage(size:-10s)

Custom Metric Scraping

Implementation examples include Python-based collectors for the pull model:


# Simple Python scraper using the pull model (Prometheus style)

import requests

def scrape_metrics(url):
    resp = requests.get(url, timeout=5)
    resp.raise_for_status()
    return resp.text

metrics = scrape_metrics("http://service-a:9090/metrics")
print(metrics)   # forward to Kafka or write directly to TSDB

Observability in Other System Designs

While the dedicated monitoring chapter provides comprehensive coverage, other designs in the repository reference observability concepts contextually.

The Stock Exchange design (28. Stock Exchange/README.md) implements latency determinism tracking using percentile monitoring (p99, p999) to ensure market fairness and regulatory compliance.

The Payment System chapter (26. Payment System/README.md) discusses replication lag as an observability concern, noting that users may observe inconsistent data during cross-region synchronization windows.

Summary

  • The system-design-notes repository provides a complete architectural blueprint for monitoring and observability for systems at scale
  • The design supports both pull (Prometheus-style) and push (CloudWatch-style) ingestion models with detailed trade-off analysis
  • Multi-resolution storage (raw 7 days → 1 min 30 days → 1 hour 1 year) optimizes cost and query performance
  • Kafka-based pipelines decouple ingestion from processing, enabling real-time alerting and historical analysis
  • Alert Manager implements deduplication, retry logic, and multi-channel delivery through YAML-configured rules
  • Related chapters demonstrate applied observability in financial systems (latency percentiles) and payment platforms (replication monitoring)

Frequently Asked Questions

What is the difference between pull and push models in this monitoring design?

Pull models have collectors scrape /metrics endpoints periodically using service discovery (etcd/ZooKeeper), which simplifies debugging and ensures data authenticity but requires firewall configuration. Push models use host agents that forward metrics to load-balanced collectors, which traverse firewalls more easily and handle sudden spikes through local buffering but require careful agent resource management.

How does the alerting system prevent duplicate notifications?

The Alert Manager component implements deduplication logic that identifies identical alerts across evaluation cycles, groups related alerts to reduce noise, and applies exponential backoff for delivery retries before publishing to Kafka for downstream channels like PagerDuty or email.

Why does the design recommend time-series databases over PostgreSQL for metrics storage?

Time-series databases like InfluxDB or Prometheus provide specialized compression algorithms for sequential data, efficient down-sampling for historical aggregation, and optimized query performance for high-cardinality labels—capabilities that generic relational databases lack without significant operational overhead.

Does the repository cover distributed tracing and log aggregation?

No, the monitoring chapter explicitly excludes distributed tracing and log aggregation to maintain focus on operational metrics (CPU, memory, request rates). These observability pillars are treated as separate architectural concerns requiring different storage and querying strategies.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →