How to Configure CloudWatch Alarms with Anomaly Detection
You configure CloudWatch alarms with anomaly detection by first creating an AWS::CloudWatch::AnomalyDetector resource to establish a two-week baseline, then referencing the detector's synthetic ANOMALY_DETECTION_BAND metric in an alarm using the LessThanLowerOrGreaterThanUpperThreshold comparison operator.
Amazon CloudWatch supports three alarm types—metric, composite, and anomaly-detection alarms—with the latter automatically learning normal patterns and firing only when observed values deviate from expected bands. According to the aws/agent-toolkit-for-aws repository, implementing these alarms requires specific configuration properties and evaluation mechanics documented in skills/core-skills/aws-observability/references/alarms.md. This guide provides the exact architectural patterns and code examples used by automated agents in the toolkit.
Understanding Anomaly Detection Architecture
Anomaly detection alarms consist of two distinct resources working in tandem. First, an AnomalyDetector resource analyzes your metric history to build a two-week baseline model. This detector exposes a synthetic metric called ANOMALY_DETECTION_BAND that represents the expected range of values.
Second, a standard CloudWatch Alarm references this band rather than a static threshold. The alarm evaluates whether the actual metric falls outside the detector's predicted range. The skills/core-skills/aws-observability/SKILL.md file links this reference architecture to higher-level skills, enabling agents to automatically suggest the correct configuration when users request CloudWatch anomaly detection.
Configuration Requirements
Successful implementation requires specific property configurations as defined in the alarms.md reference:
- ComparisonOperator: Must be set to
LessThanLowerOrGreaterThanUpperThresholdto detect deviations in either direction - ThresholdMetricId: Points to the ID of your
ANOMALY_DETECTION_BANDexpression (e.g.,ad1) - EvaluationPeriods: Set to
1because the band already models variance and extended evaluation periods are unnecessary - TreatMissingData: Set to
breachingso that lack of data is considered a breach of the expected pattern - Band Width: The second argument to
ANOMALY_DETECTION_BAND(metric, width)defaults to2and controls how "wide" the normal band appears
Step-by-Step Configuration Process
Create the Anomaly Detector
First, provision the detector that will learn your metric's normal behavior. Specify the metric name, namespace, and statistic (e.g., Sum for Lambda Invocations). The detector automatically trains on historical data and becomes available for alarm creation within minutes.
In skills/core-skills/aws-observability/references/alarms.md, the recommended pattern uses the AWS::CloudWatch::AnomalyDetector resource type with the specific metric dimensions you want to monitor.
Configure the Alarm to Reference the Detector
Create a CloudWatch alarm that consumes the detector's output using metric math. Define an expression ANOMALY_DETECTION_BAND(m1, 2) where m1 references your source metric. Set the alarm's ThresholdMetricId to the ID of this expression (commonly ad1).
Use the LessThanLowerOrGreaterThanUpperThreshold comparison operator to trigger when values fall either below the lower bound or above the upper bound of the predicted band.
Handle Missing Data and Band Width
For anomaly-detection workloads, configure TreatMissingData: breaching to ensure that gaps in telemetry are treated as anomalies. Adjust the band width parameter (the 2 in ANOMALY_DETECTION_BAND(m1, 2)) only if your workload requires tighter or looser tolerance; the default value of 2 works for most standard workloads.
Implementation Examples
AWS CLI
The following commands create the detector first, then the alarm that references it:
aws cloudwatch put-anomaly-detector \
--metric-name Invocations \
--namespace AWS/Lambda \
--stat Sum
aws cloudwatch put-metric-alarm \
--alarm-name Lambda-Invocations-Anomaly \
--comparison-operator LessThanLowerOrGreaterThanUpperThreshold \
--evaluation-periods 1 \
--metrics '[{"Id":"ad1","Expression":"ANOMALY_DETECTION_BAND(m1,2)"},
{"Id":"m1","MetricStat":{"Metric":{"Namespace":"AWS/Lambda","MetricName":"Invocations"},"Period":86400,"Stat":"Sum"}}]' \
--threshold-metric-id ad1 \
--treat-missing-data breaching
The first command establishes the baseline model, while the second creates the alarm that evaluates the band.
CloudFormation (YAML)
Declare both resources in your template as shown in skills/core-skills/aws-observability/references/alarms.md:
Resources:
InvocationsAnomalyDetector:
Type: AWS::CloudWatch::AnomalyDetector
Properties:
MetricName: Invocations
Namespace: AWS/Lambda
Stat: Sum
InvocationsAnomalyAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
ComparisonOperator: LessThanLowerOrGreaterThanUpperThreshold
EvaluationPeriods: 1
Metrics:
- Expression: ANOMALY_DETECTION_BAND(m1, 2)
Id: ad1
- Id: m1
MetricStat:
Metric:
MetricName: Invocations
Namespace: AWS/Lambda
Period: 86400
Stat: Sum
ThresholdMetricId: ad1
TreatMissingData: breaching
AWS CDK (TypeScript)
The CDK provides type-safe constructs that mirror the CloudFormation resources:
import { AnomalyDetector, Alarm, ComparisonOperator, TreatMissingData, Metric, MathExpression } from 'aws-cdk-lib/aws-cloudwatch';
import { Duration } from 'aws-cdk-lib';
// Create the detector
new AnomalyDetector(this, 'InvocationsDetector', {
metricName: 'Invocations',
namespace: 'AWS/Lambda',
statistic: 'Sum',
});
// Create the alarm consuming the detector's band
new Alarm(this, 'InvocationsAnomalyAlarm', {
metric: new MathExpression({
expression: 'ANOMALY_DETECTION_BAND(m1, 2)',
usingMetrics: {
m1: new Metric({
namespace: 'AWS/Lambda',
metricName: 'Invocations',
period: Duration.days(1),
statistic: 'Sum',
}),
},
}),
comparisonOperator: ComparisonOperator.LESS_THAN_LOWER_OR_GREATER_THAN_UPPER_THRESHOLD,
evaluationPeriods: 1,
treatMissingData: TreatMissingData.BREACHING,
});
This pattern provides compile-time validation while generating the same underlying CloudFormation resources.
Cost Considerations and Service Limits
Anomaly-detection alarms incur higher costs than standard metric alarms due to the computational overhead of band evaluation. The service imposes a limit of 1,000 band-evaluations per second per account, as documented in skills/core-skills/aws-observability/references/alarms.md.
If you approach this limit, consider consolidating alarms or using standard threshold alarms for high-cardinality metrics. The repository also references specialized implementations in skills/specialized-skills/database-skills/amazon-elasticache/references/monitoring/alarm-packs.md for service-specific anomaly detection patterns.
Summary
- Two-resource architecture: You must create both an
AnomalyDetectorand anAlarmto implement anomaly detection - Required operator: Use
LessThanLowerOrGreaterThanUpperThresholdto detect band breaches - Single evaluation period: Set
EvaluationPeriodsto1because the band already accounts for variance - Missing data handling: Configure
TreatMissingData: breachingto treat missing telemetry as anomalies - Default band width: The standard width of
2inANOMALY_DETECTION_BANDworks for most workloads - Service limits: Account for the 1,000 band-evaluations per second limit when designing at scale
Frequently Asked Questions
What is the difference between a standard CloudWatch alarm and an anomaly detection alarm?
A standard alarm compares metrics against static thresholds you define, while an anomaly detection alarm uses machine learning to establish a dynamic baseline. The anomaly detector automatically adjusts for time-of-day and seasonal patterns, firing only when metrics deviate from the predicted ANOMALY_DETECTION_BAND as implemented in the aws/agent-toolkit-for-aws reference architecture.
Why must EvaluationPeriods be set to 1 for anomaly detection alarms?
The ANOMALY_DETECTION_BAND metric math expression already incorporates variance modeling over time, effectively smoothing short-term fluctuations. Setting EvaluationPeriods greater than 1 would add unnecessary delay to anomaly detection without improving accuracy, since the band itself represents the expected range of normal behavior.
How does the ANOMALY_DETECTION_BAND width parameter affect sensitivity?
The second argument to ANOMALY_DETECTION_BAND(m1, width) defines how many standard deviations the band extends from the predicted value. A value of 2 (the default) captures approximately 95% of normal variations. Increasing this number reduces false positives by widening the "normal" range, while decreasing it creates a tighter band that triggers more aggressively on minor deviations.
What happens if data is missing for an anomaly detection alarm?
By default, missing data is treated as missing, which means the alarm maintains its previous state. For anomaly detection use cases, the recommended configuration is TreatMissingData: breaching, which causes the alarm to trigger when telemetry stops arriving. This prevents silent failures where an application stops reporting metrics but the alarm remains in an OK state.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →