# How S3 Remote Session Storage Enables Distributed Deployments in AgentsView

> Learn how S3 remote session storage in AgentsView enables distributed deployments. Discover shared session access, isolated namespaces, and efficient local caching for scalable solutions.

- Repository: [Kenn Software/agentsview](https://github.com/kenn-io/agentsview)
- Tags: architecture
- Published: 2026-07-04

---

**AgentsView uses Amazon S3 as a shared session store, allowing multiple nodes to read the same agent session files while maintaining isolated machine namespaces and efficient local caching.**

The open-source AgentsView project stores AI-agent session data locally in SQLite by default, but for production-scale deployments, it can read session files directly from Amazon S3. This architecture, implemented in the `kenn-io/agentsview` repository, eliminates the need for a central coordinator by treating S3 as the source of truth while keeping each node’s local database synchronized through intelligent metadata checks and temporary file handling.

## Machine-Scoped Identifiers for Collision Avoidance

When multiple nodes fetch the same session from a shared S3 bucket, AgentsView prevents ID collisions by prefixing the session ID with the originating machine name.

In [`internal/sync/s3.go`](https://github.com/kenn-io/agentsview/blob/main/internal/sync/s3.go), the `s3SessionIDPrefix` helper (lines 15-20) constructs identifiers using the format `machine~id`. For example, a session originating from `node-01` becomes `node-01~codex:123e4567-e89b-12d3-a456-426614174000`. This namespace isolation ensures that when several instances import identical S3 objects, each maintains a distinct database row while still referencing the same underlying data.

## Safe Temporary Path Handling

Security considerations govern how S3 objects are materialized locally. The `safeS3TempRelPath` function (lines 69-100) strips the `s3://` scheme and transforms the object key into a safe relative path suitable for temporary directories.

This function validates each path component to prevent directory-traversal attacks, ensuring that even maliciously crafted S3 keys cannot write files outside the designated temporary location. The sanitized path is then used to create a temporary file that exists only for the duration of the parsing operation.

## On-Demand Materialization Without Persistent Local Copies

The core of the S3 pipeline is `processS3Session`, which implements a **download-parse-cleanup** cycle that leaves no permanent local artifacts.

The function (lines 44-68) downloads the S3 object to a temporary file, passes it to the standard per-agent parser, and removes the file when the function returns. Specifically, lines 73-87 handle the download-to-temp-file logic, ensuring the node never retains a permanent copy of the raw S3 object. This approach keeps local storage minimal while allowing the node to parse complex session formats from Codex, Claude, and other providers.

## Metadata-Driven Sync Efficiency

Before fetching any object, the engine checks whether the local SQLite cache already contains the current version. The `shouldSkipFileWithPrefix` function compares the S3 object's size, modification time, and an optional `sourceFingerprint` against the local database record.

If the metadata matches, the engine skips the download entirely and reuses the existing SQLite row. This optimization reduces network traffic and enables efficient synchronization across dozens of nodes polling the same bucket, making the architecture suitable for high-scale distributed deployments.

## Hydrating Ancillary Session Data

Certain AI providers require additional files beyond the main session log. For OpenAI Codex sessions, `hydrateS3CodexSessionIndex` pulls the session index file and stores it locally so the parser can resolve human-readable session names. For Anthropic Claude sessions, `hydrateS3ClaudeToolResults` retrieves tool-call result data.

Both helpers are invoked immediately after the temporary file is written (see lines 72-84 of `processS3Session`), ensuring the parser has access to all necessary context before processing begins.

## Preserving the True Source URI

After parsing completes, the engine overwrites the parsed `Session.File.Path` with the original S3 URI. This assignment occurs in lines 96-100 of `processS3Session`.

By storing the remote location rather than the temporary local path, the database maintains a reference to the canonical source. This ensures that subsequent operations—such as re-syncing, exporting, or cross-referencing—know exactly where the data originated, even if the temporary file has long since been deleted.

## Force-Replace Strategy for Consistency

When S3 object metadata changes—indicating a modified session file—or when ancillary data like Codex indices or Claude tool results are refreshed, the engine sets `forceReplace = true`. This flag forces an update to the SQLite row, overwriting cached data with the fresh content.

This mechanism guarantees that all nodes in a distributed deployment eventually converge on the latest session state, even if they poll the bucket at different intervals or join the cluster after updates have occurred.

## Configuring S3 Remote Session Storage

To enable distributed deployment, point the AgentsView configuration to an S3 URI:

```go
// config.yaml
data_dir: "s3://my-agentsview-bucket/sessions/"

```

Programmatically, initialize the sync engine with a machine identifier and S3 source:

```go
package main

import (
	"context"
	"log"

	"go.kenn.io/agentsview/internal/sync"
)

func main() {
	// Initialize engine for node "node-01" with S3 backend
	engine, err := sync.NewEngine(sync.Config{
		Machine: "node-01",
		Source:  "s3://my-agentsview-bucket/sessions/",
	})
	if err != nil {
		log.Fatalf("engine init: %v", err)
	}

	// Run synchronization
	if err := engine.SyncAll(context.Background()); err != nil {
		log.Fatalf("sync failed: %v", err)
	}
}

```

After syncing, verify the remote source is preserved:

```go
sess, err := db.GetSession(ctx, "node-01~codex:123e4567-e89b-12d3-a456-426614174000")
if err != nil {
    log.Fatal(err)
}
// Output: s3://my-agentsview-bucket/sessions/codex/123e4567-e89b-12d3-a456-426614174000.json
fmt.Printf("Session source: %s\n", sess.File.Path)

```

## Summary

- **S3 remote session storage** in AgentsView treats Amazon S3 as a shared source of truth, enabling horizontal scaling without a central coordinator.
- **Machine-scoped identifiers** (`machine~id`) prevent collisions when multiple nodes import identical session files from the same bucket.
- **Safe temporary paths** (`safeS3TempRelPath`) and automatic cleanup ensure security and minimal local storage usage.
- **Metadata-driven skipping** (`shouldSkipFileWithPrefix`) reduces network traffic by comparing S3 metadata against local cache before downloading.
- **Ancillary hydration** functions handle provider-specific requirements like Codex session indices and Claude tool results.
- **Source URI preservation** maintains the original S3 path in the database, ensuring traceability across distributed nodes.
- **Force-replace logic** guarantees convergence when session files or their dependencies change.

## Frequently Asked Questions

### How does AgentsView handle session ID conflicts in distributed deployments?

AgentsView prefixes every session ID with the machine name using the `s3SessionIDPrefix` function in [`internal/sync/s3.go`](https://github.com/kenn-io/agentsview/blob/main/internal/sync/s3.go). This creates unique identifiers like `node-01~session-uuid` and `node-02~session-uuid`, allowing multiple nodes to store references to the same underlying S3 object without database collisions.

### Does AgentsView keep permanent local copies of S3 session files?

No. The `processS3Session` function downloads S3 objects to temporary files only for the duration of parsing. The temporary file is removed when the function returns, leaving only the parsed data in the local SQLite database. This design minimizes disk usage on each node while maintaining fast query performance.

### What happens if an S3 session file is modified after initial import?

The sync engine compares the S3 object's size, modification time, and optional fingerprint against the local record. If any metadata differs, the engine sets `forceReplace = true` and re-downloads the file, updating the SQLite row with the new content. This ensures all nodes eventually synchronize to the latest version of the session data.

### Can AgentsView sync from S3 buckets containing millions of session files?

Yes. The metadata-driven skipping mechanism (`shouldSkipFileWithPrefix`) ensures that only changed or new files are downloaded. Each node maintains its own SQLite cache of session metadata, making the architecture horizontally scalable without increasing load on the S3 bucket proportionally to the number of polling nodes.