How S3 Remote Session Storage Enables Distributed Deployments in AgentsView
AgentsView uses Amazon S3 as a shared session store, allowing multiple nodes to read the same agent session files while maintaining isolated machine namespaces and efficient local caching.
The open-source AgentsView project stores AI-agent session data locally in SQLite by default, but for production-scale deployments, it can read session files directly from Amazon S3. This architecture, implemented in the kenn-io/agentsview repository, eliminates the need for a central coordinator by treating S3 as the source of truth while keeping each node’s local database synchronized through intelligent metadata checks and temporary file handling.
Machine-Scoped Identifiers for Collision Avoidance
When multiple nodes fetch the same session from a shared S3 bucket, AgentsView prevents ID collisions by prefixing the session ID with the originating machine name.
In internal/sync/s3.go, the s3SessionIDPrefix helper (lines 15-20) constructs identifiers using the format machine~id. For example, a session originating from node-01 becomes node-01~codex:123e4567-e89b-12d3-a456-426614174000. This namespace isolation ensures that when several instances import identical S3 objects, each maintains a distinct database row while still referencing the same underlying data.
Safe Temporary Path Handling
Security considerations govern how S3 objects are materialized locally. The safeS3TempRelPath function (lines 69-100) strips the s3:// scheme and transforms the object key into a safe relative path suitable for temporary directories.
This function validates each path component to prevent directory-traversal attacks, ensuring that even maliciously crafted S3 keys cannot write files outside the designated temporary location. The sanitized path is then used to create a temporary file that exists only for the duration of the parsing operation.
On-Demand Materialization Without Persistent Local Copies
The core of the S3 pipeline is processS3Session, which implements a download-parse-cleanup cycle that leaves no permanent local artifacts.
The function (lines 44-68) downloads the S3 object to a temporary file, passes it to the standard per-agent parser, and removes the file when the function returns. Specifically, lines 73-87 handle the download-to-temp-file logic, ensuring the node never retains a permanent copy of the raw S3 object. This approach keeps local storage minimal while allowing the node to parse complex session formats from Codex, Claude, and other providers.
Metadata-Driven Sync Efficiency
Before fetching any object, the engine checks whether the local SQLite cache already contains the current version. The shouldSkipFileWithPrefix function compares the S3 object's size, modification time, and an optional sourceFingerprint against the local database record.
If the metadata matches, the engine skips the download entirely and reuses the existing SQLite row. This optimization reduces network traffic and enables efficient synchronization across dozens of nodes polling the same bucket, making the architecture suitable for high-scale distributed deployments.
Hydrating Ancillary Session Data
Certain AI providers require additional files beyond the main session log. For OpenAI Codex sessions, hydrateS3CodexSessionIndex pulls the session index file and stores it locally so the parser can resolve human-readable session names. For Anthropic Claude sessions, hydrateS3ClaudeToolResults retrieves tool-call result data.
Both helpers are invoked immediately after the temporary file is written (see lines 72-84 of processS3Session), ensuring the parser has access to all necessary context before processing begins.
Preserving the True Source URI
After parsing completes, the engine overwrites the parsed Session.File.Path with the original S3 URI. This assignment occurs in lines 96-100 of processS3Session.
By storing the remote location rather than the temporary local path, the database maintains a reference to the canonical source. This ensures that subsequent operations—such as re-syncing, exporting, or cross-referencing—know exactly where the data originated, even if the temporary file has long since been deleted.
Force-Replace Strategy for Consistency
When S3 object metadata changes—indicating a modified session file—or when ancillary data like Codex indices or Claude tool results are refreshed, the engine sets forceReplace = true. This flag forces an update to the SQLite row, overwriting cached data with the fresh content.
This mechanism guarantees that all nodes in a distributed deployment eventually converge on the latest session state, even if they poll the bucket at different intervals or join the cluster after updates have occurred.
Configuring S3 Remote Session Storage
To enable distributed deployment, point the AgentsView configuration to an S3 URI:
// config.yaml
data_dir: "s3://my-agentsview-bucket/sessions/"
Programmatically, initialize the sync engine with a machine identifier and S3 source:
package main
import (
"context"
"log"
"go.kenn.io/agentsview/internal/sync"
)
func main() {
// Initialize engine for node "node-01" with S3 backend
engine, err := sync.NewEngine(sync.Config{
Machine: "node-01",
Source: "s3://my-agentsview-bucket/sessions/",
})
if err != nil {
log.Fatalf("engine init: %v", err)
}
// Run synchronization
if err := engine.SyncAll(context.Background()); err != nil {
log.Fatalf("sync failed: %v", err)
}
}
After syncing, verify the remote source is preserved:
sess, err := db.GetSession(ctx, "node-01~codex:123e4567-e89b-12d3-a456-426614174000")
if err != nil {
log.Fatal(err)
}
// Output: s3://my-agentsview-bucket/sessions/codex/123e4567-e89b-12d3-a456-426614174000.json
fmt.Printf("Session source: %s\n", sess.File.Path)
Summary
- S3 remote session storage in AgentsView treats Amazon S3 as a shared source of truth, enabling horizontal scaling without a central coordinator.
- Machine-scoped identifiers (
machine~id) prevent collisions when multiple nodes import identical session files from the same bucket. - Safe temporary paths (
safeS3TempRelPath) and automatic cleanup ensure security and minimal local storage usage. - Metadata-driven skipping (
shouldSkipFileWithPrefix) reduces network traffic by comparing S3 metadata against local cache before downloading. - Ancillary hydration functions handle provider-specific requirements like Codex session indices and Claude tool results.
- Source URI preservation maintains the original S3 path in the database, ensuring traceability across distributed nodes.
- Force-replace logic guarantees convergence when session files or their dependencies change.
Frequently Asked Questions
How does AgentsView handle session ID conflicts in distributed deployments?
AgentsView prefixes every session ID with the machine name using the s3SessionIDPrefix function in internal/sync/s3.go. This creates unique identifiers like node-01~session-uuid and node-02~session-uuid, allowing multiple nodes to store references to the same underlying S3 object without database collisions.
Does AgentsView keep permanent local copies of S3 session files?
No. The processS3Session function downloads S3 objects to temporary files only for the duration of parsing. The temporary file is removed when the function returns, leaving only the parsed data in the local SQLite database. This design minimizes disk usage on each node while maintaining fast query performance.
What happens if an S3 session file is modified after initial import?
The sync engine compares the S3 object's size, modification time, and optional fingerprint against the local record. If any metadata differs, the engine sets forceReplace = true and re-downloads the file, updating the SQLite row with the new content. This ensures all nodes eventually synchronize to the latest version of the session data.
Can AgentsView sync from S3 buckets containing millions of session files?
Yes. The metadata-driven skipping mechanism (shouldSkipFileWithPrefix) ensures that only changed or new files are downloaded. Each node maintains its own SQLite cache of session metadata, making the architecture horizontally scalable without increasing load on the S3 bucket proportionally to the number of polling nodes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →