How to Send Log Files from EC2 to S3: CloudWatch Agent and Lambda Pipeline Guide
Use the CloudWatch Agent to stream logs from EC2 to CloudWatch Logs, then configure a subscription filter with Lambda or Kinesis Data Firehose to automatically deliver them to Amazon S3 for durable, cost-effective storage.
Sending log files from an Amazon EC2 instance to S3 is a fundamental DevOps pattern for centralized archiving and compliance. According to the litu54/DevOps-Interview-Guide repository—specifically the interview questions in Amazon/DevOps_Consultant_1.md—candidates must demonstrate a production-ready pipeline that handles collection, buffering, and delivery without requiring custom background daemons on every instance.
The Recommended Architecture: CloudWatch Agent to S3
The most reliable pattern follows a collection → buffering → delivery workflow using native AWS services. Instead of running cron jobs that upload files directly, you deploy the unified CloudWatch Agent on EC2 to stream logs to CloudWatch Logs, then use a subscription filter to forward events to a Lambda function that writes them to S3.
This architecture decouples log generation from storage, provides automatic retries with exponential backoff, and eliminates the need to manage credentials on the instance.
Step 1: Install the Unified CloudWatch Agent
Deploy the amazon-cloudwatch-agent package on your Amazon Linux 2/2023 or Ubuntu instance. The agent runs as a systemd service and reads log files without locking them, supporting active rotation scenarios.
Your EC2 instance must have an IAM role attached that includes the cloudwatch:PutLogEvents permission. Never embed access keys in the configuration; rely on the instance metadata service for credential retrieval.
Step 2: Configure the Agent for File Collection
Create a JSON configuration file—typically /opt/aws/amazon-cloudwatch-agent/etc/amazon-cloudwatch-agent.json—that defines which files to monitor and where to send them.
{
"logs": {
"logs_collected": {
"files": {
"collect_list": [
{
"file_path": "/var/log/app.log",
"log_group_name": "app-logs",
"log_stream_name": "{instance_id}"
}
]
}
}
}
}
The {instance_id} placeholder automatically resolves to the EC2 instance ID, creating separate log streams per host for easier troubleshooting. Restart the agent after updating the configuration: sudo /opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl -a fetch-config -m ec2 -s -c file:config.json.
Step 3: Create a CloudWatch Logs Subscription Filter
Once logs arrive in the CloudWatch Logs group, create a subscription filter that triggers your Lambda function in near real-time. Use the AWS CLI to establish the connection:
aws logs put-subscription-filter \
--log-group-name app-logs \
--filter-name ToS3 \
--filter-pattern "" \
--destination-arn arn:aws:lambda:us-east-1:123456789012:function:LogToS3Lambda
An empty filter pattern ("") forwards all log events. For high-throughput scenarios, substitute the Lambda ARN with a Kinesis Data Firehose ARN to batch records before S3 delivery, reducing invocation costs.
Step 4: Lambda Function for S3 Delivery
The Lambda function receives batched log events via the data field in the event payload, compresses them for cost savings, and uses s3:PutObject to store them. Below is a Node.js implementation that handles Base64 decoding and GZIP compression:
const AWS = require('aws-sdk');
const zlib = require('zlib');
const s3 = new AWS.S3();
exports.handler = async (event) => {
const bucket = process.env.TARGET_BUCKET;
const timestamp = new Date().toISOString();
// CloudWatch Logs sends data base64-encoded
const payload = Buffer.from(event.awslogs.data, 'base64');
const parsed = JSON.parse(zlib.gunzipSync(payload));
// Compress for S3 storage
const compressed = zlib.gzipSync(JSON.stringify(parsed));
const key = `logs/${timestamp}.json.gz`;
await s3.putObject({
Bucket: bucket,
Key: key,
Body: compressed,
ContentEncoding: 'gzip',
ContentType: 'application/json'
}).promise();
console.log(`Wrote ${key} to ${bucket}`);
};
Deploy this function with an execution role containing s3:PutObject permissions on the target bucket. Set the TARGET_BUCKET environment variable to your S3 bucket name.
Alternative: Direct AWS CLI Transfer with Cron
For simple, low-volume use cases—such as legacy applications or one-off debugging—you can schedule a cron job that invokes the AWS CLI directly. This method requires no CloudWatch or Lambda resources but lacks buffering, automatic retries, and centralized monitoring.
Add the following to /etc/crontab to upload logs every 15 minutes with timestamped filenames:
*/15 * * * * root /usr/local/bin/aws s3 cp /var/log/app.log s3://my-bucket/logs/$(hostname)-$(date +\%Y\%m\%d\%H\%M).log
Warning: This approach requires the AWS CLI to be configured with credentials or an IAM role, and it does not handle log rotation gracefully. If the file is moved or truncated during upload, you risk data loss or corruption.
Security and Lifecycle Management
IAM Roles: The EC2 instance role only needs cloudwatch:PutLogEvents and logs:DescribeLogGroups. The Lambda execution role requires s3:PutObject on the destination bucket. Avoid using IAM user access keys on instances.
Data Retention: Configure S3 lifecycle rules to transition objects to S3 Glacier after 90 days and delete them after one year. This reduces storage costs while maintaining compliance with audit requirements.
Encryption: Enable default encryption on the S3 bucket using AWS KMS (SSE-KMS) to protect log data at rest. CloudWatch Logs and S3 both support TLS 1.2 for data in transit.
Summary
- CloudWatch Agent provides the most reliable method to send log files from EC2 to S3, handling buffering, retries, and credential management automatically.
- A subscription filter bridges CloudWatch Logs and S3, enabling near real-time delivery without persistent connections from the instance.
- Lambda functions can transform and compress logs before storage, reducing S3 costs and enabling custom formatting (JSON, Parquet).
- For simple scenarios,
aws s3 cpwith cron works but lacks the scalability and observability of the CloudWatch pipeline. - Always use IAM roles instead of hardcoded credentials, and implement S3 lifecycle policies to manage long-term costs.
Frequently Asked Questions
Can I use Kinesis Data Firehose instead of Lambda to deliver logs to S3?
Yes. Kinesis Data Firehose is often preferred for high-volume streams because it batches records and handles compression automatically without requiring you to manage Lambda concurrency. Configure the subscription filter destination ARN to point to your Firehose delivery stream, which then writes to S3 in configurable buffer intervals (e.g., every 5 minutes or 128 MB).
What IAM permissions does the EC2 instance need to send logs to CloudWatch?
The EC2 instance role requires only logs:CreateLogGroup, logs:CreateLogStream, and logs:PutLogEvents. For the CloudWatch Agent specifically, cloudwatch:PutMetricData is also needed if collecting system metrics. No S3 permissions are required on the instance when using the CloudWatch → Lambda → S3 pipeline.
How does the CloudWatch Agent handle log rotation?
The agent monitors file descriptors, not just filenames, so it continues reading from the same inode even when a log rotation utility (like logrotate) renames the file. It automatically detects new files created at the original path and begins streaming them immediately without losing events or duplicating data.
Is there a cost difference between using CloudWatch Logs versus direct S3 uploads?
CloudWatch Logs incurs ingestion charges per GB and storage costs for the retention period you configure. However, by exporting to S3 via subscription filters and setting CloudWatch retention to 1 day (or deleting immediately), you minimize long-term CloudWatch storage costs. S3 storage is significantly cheaper than CloudWatch Logs storage, making the pipeline cost-effective for archival use cases despite the slight added complexity.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →