Implementing Session Deduplication in ADR Using filter_entries_by_existing_files

The filter_entries_by_existing_files method in Sensor/adr_sensor/observer.py provides an idempotent deduplication mechanism that compares incoming agent session timestamps against existing output files, ensuring only new or updated sessions are processed during repeated sensor runs.

ADR (Agent-Driven Research) is an open-source framework from the uber/ADR repository that records agent interactions as discrete sessions stored in JSON format. When re-running the sensor on previously processed datasets, unfiltered execution duplicates work and risks overwriting existing outputs. The filter_entries_by_existing_files utility eliminates redundant processing by scanning existing adr.*.json files and retaining only AgentEvent entries with newer timestamps.

How filter_entries_by_existing_files Works

The deduplication logic resides in Sensor/adr_sensor/observer.py and operates through three coordinated steps:

Discovery of Existing Sessions

The helper method _get_existing_session_files scans the target output directory for files matching the pattern adr.*.json. For each file discovered, it extracts the session identifier and timestamp using parse_timestamp_from_filename. The method maintains a mapping of existing_files[session_id] containing only the most recent file per session, ensuring comparisons against the latest known state.

Timestamp Comparison and Normalization

For every incoming AgentEvent entry, filter_entries_by_existing_files checks if the session_id exists in the existing files map. If the session is unseen, the entry is immediately retained. For existing sessions, the method normalizes both timestamps using normalize_timestamp from Sensor/adr_sensor/utils/timestamp_utils, strips microseconds, and performs an inequality comparison. Only entries with timestamps greater than the existing file's timestamp are retained.

CLI Integration Point

The command-line interface in Sensor/adr_sensor/cli.py invokes this filter at line 119 before writing output, creating an idempotent pipeline where re-executing the sensor on identical input produces no duplicate files.

Implementation Examples

Basic CLI Usage

When processing sessions through the ADR CLI, deduplication happens automatically:

from adr_sensor.observer import SessionObserver
from adr_sensor.cli import parse_args

args = parse_args()
entries = load_entries(args.input)          # load raw AgentEvent objects

# Deduplicate against existing output files

entries = SessionObserver().filter_entries_by_existing_files(
    entries, output_dir=args.output_dir
)
SessionObserver().save_sessions_to_individual_files(entries, args.output_dir)

Manual Deduplication in Custom Scripts

For programmatic use outside the CLI, instantiate the observer directly:

from pathlib import Path
from adr_sensor.observer import SessionObserver
from adr_sensor.parsers.claude_parser import ClaudeParser

# Parse a directory of raw Claude logs

parser = ClaudeParser()
entries = parser.parse_dir(Path("raw_logs"))

# Choose an output location

out_dir = Path("adr_output")

# Apply deduplication

obs = SessionObserver()
new_entries = obs.filter_entries_by_existing_files(entries, out_dir)

# Persist only the new or updated sessions

saved = obs.save_sessions_to_individual_files(new_entries, out_dir)
print(f"Saved {len(saved)} new session files.")

Internal Filtering Logic

The core comparison logic strips microseconds to ensure consistent datetime matching:

def filter_entries_by_existing_files(self, entries, output_dir=None):
    existing_files = self._get_existing_session_files(output_dir)
    filtered = []
    for entry in entries:
        sid = entry.session_id
        if sid not in existing_files:
            filtered.append(entry)
            continue
        # Compare timestamps (microseconds stripped)

        if normalize_timestamp(entry.timestamp).replace(microsecond=0) > \
           normalize_timestamp(existing_files[sid]["timestamp"]).replace(microsecond=0):
            filtered.append(entry)
    return filtered

Key Source Files

Understanding the complete deduplication architecture requires familiarity with these components:

Summary

  • filter_entries_by_existing_files enables idempotent session processing by comparing timestamps against existing adr.*.json files.
  • The method utilizes _get_existing_session_files to build a session map and normalize_timestamp for accurate datetime comparison.
  • Integration with Sensor/adr_sensor/cli.py ensures automatic deduplication during CLI execution.
  • Microsecond stripping ensures filesystem timestamp inconsistencies do not cause false positives.

Frequently Asked Questions

How does ADR handle timestamp normalization for deduplication?

ADR uses the normalize_timestamp utility from Sensor/adr_sensor/utils/timestamp_utils.py to convert various timestamp representations into uniform datetime objects. During comparison, microseconds are stripped using .replace(microsecond=0) to ensure filesystem-level timestamp variations do not trigger unnecessary reprocessing.

What file pattern does filter_entries_by_existing_files search for?

The _get_existing_session_files method scans the output directory for files matching the glob pattern adr.*.json. It extracts session identifiers and timestamps from these filenames to build the comparison map used by the filtering logic.

Can I use filter_entries_by_existing_files outside the ADR CLI?

Yes. Instantiate SessionObserver directly and call filter_entries_by_existing_files(entries, output_dir) to manually deduplicate a list of AgentEvent objects against an existing directory. This is useful for custom batch processing scripts or integration with external data pipelines.

Where are the deduplication tests located?

Unit tests verifying the filter behavior reside in Sensor/tests/test_observer.py at lines 128-158. These tests confirm that the method correctly removes unchanged sessions while retaining entries with newer timestamps.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →