# Implementing Session Deduplication in ADR Using filter_entries_by_existing_files

> Learn how to implement session deduplication in ADR using filter_entries_by_existing_files. Ensure only new or updated sessions are processed for efficient sensor runs.

- Repository: [Uber Open Source/ADR](https://github.com/uber/ADR)
- Tags: how-to-guide
- Published: 2026-08-07

---

**The `filter_entries_by_existing_files` method in [`Sensor/adr_sensor/observer.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/observer.py) provides an idempotent deduplication mechanism that compares incoming agent session timestamps against existing output files, ensuring only new or updated sessions are processed during repeated sensor runs.**

ADR (Agent-Driven Research) is an open-source framework from the uber/ADR repository that records agent interactions as discrete sessions stored in JSON format. When re-running the sensor on previously processed datasets, unfiltered execution duplicates work and risks overwriting existing outputs. The `filter_entries_by_existing_files` utility eliminates redundant processing by scanning existing `adr.*.json` files and retaining only **AgentEvent** entries with newer timestamps.

## How filter_entries_by_existing_files Works

The deduplication logic resides in [`Sensor/adr_sensor/observer.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/observer.py) and operates through three coordinated steps:

### Discovery of Existing Sessions

The helper method `_get_existing_session_files` scans the target output directory for files matching the pattern `adr.*.json`. For each file discovered, it extracts the session identifier and timestamp using `parse_timestamp_from_filename`. The method maintains a mapping of `existing_files[session_id]` containing only the most recent file per session, ensuring comparisons against the latest known state.

### Timestamp Comparison and Normalization

For every incoming `AgentEvent` entry, `filter_entries_by_existing_files` checks if the `session_id` exists in the existing files map. If the session is unseen, the entry is immediately retained. For existing sessions, the method normalizes both timestamps using `normalize_timestamp` from `Sensor/adr_sensor/utils/timestamp_utils`, strips microseconds, and performs an inequality comparison. Only entries with timestamps **greater than** the existing file's timestamp are retained.

### CLI Integration Point

The command-line interface in [`Sensor/adr_sensor/cli.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/cli.py) invokes this filter at line 119 before writing output, creating an idempotent pipeline where re-executing the sensor on identical input produces no duplicate files.

## Implementation Examples

### Basic CLI Usage

When processing sessions through the ADR CLI, deduplication happens automatically:

```python
from adr_sensor.observer import SessionObserver
from adr_sensor.cli import parse_args

args = parse_args()
entries = load_entries(args.input)          # load raw AgentEvent objects

# Deduplicate against existing output files

entries = SessionObserver().filter_entries_by_existing_files(
    entries, output_dir=args.output_dir
)
SessionObserver().save_sessions_to_individual_files(entries, args.output_dir)

```

### Manual Deduplication in Custom Scripts

For programmatic use outside the CLI, instantiate the observer directly:

```python
from pathlib import Path
from adr_sensor.observer import SessionObserver
from adr_sensor.parsers.claude_parser import ClaudeParser

# Parse a directory of raw Claude logs

parser = ClaudeParser()
entries = parser.parse_dir(Path("raw_logs"))

# Choose an output location

out_dir = Path("adr_output")

# Apply deduplication

obs = SessionObserver()
new_entries = obs.filter_entries_by_existing_files(entries, out_dir)

# Persist only the new or updated sessions

saved = obs.save_sessions_to_individual_files(new_entries, out_dir)
print(f"Saved {len(saved)} new session files.")

```

### Internal Filtering Logic

The core comparison logic strips microseconds to ensure consistent datetime matching:

```python
def filter_entries_by_existing_files(self, entries, output_dir=None):
    existing_files = self._get_existing_session_files(output_dir)
    filtered = []
    for entry in entries:
        sid = entry.session_id
        if sid not in existing_files:
            filtered.append(entry)
            continue
        # Compare timestamps (microseconds stripped)

        if normalize_timestamp(entry.timestamp).replace(microsecond=0) > \
           normalize_timestamp(existing_files[sid]["timestamp"]).replace(microsecond=0):
            filtered.append(entry)
    return filtered

```

## Key Source Files

Understanding the complete deduplication architecture requires familiarity with these components:

- **[`Sensor/adr_sensor/observer.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/observer.py)** – Contains `filter_entries_by_existing_files`, `_get_existing_session_files`, and filename extraction utilities.
- **[`Sensor/adr_sensor/cli.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/cli.py)** – Invokes the filter before exporting sessions at line 119.
- **[`Sensor/adr_sensor/utils/timestamp_utils.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/utils/timestamp_utils.py)** – Provides `normalize_timestamp` and `parse_timestamp_from_filename` for consistent temporal comparisons.
- **[`Sensor/tests/test_observer.py`](https://github.com/uber/ADR/blob/main/Sensor/tests/test_observer.py)** – Unit tests verifying deduplication behavior at lines 128-158.

## Summary

- **`filter_entries_by_existing_files`** enables idempotent session processing by comparing timestamps against existing `adr.*.json` files.
- The method utilizes `_get_existing_session_files` to build a session map and `normalize_timestamp` for accurate datetime comparison.
- Integration with [`Sensor/adr_sensor/cli.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/cli.py) ensures automatic deduplication during CLI execution.
- Microsecond stripping ensures filesystem timestamp inconsistencies do not cause false positives.

## Frequently Asked Questions

### How does ADR handle timestamp normalization for deduplication?

ADR uses the `normalize_timestamp` utility from [`Sensor/adr_sensor/utils/timestamp_utils.py`](https://github.com/uber/ADR/blob/main/Sensor/adr_sensor/utils/timestamp_utils.py) to convert various timestamp representations into uniform datetime objects. During comparison, microseconds are stripped using `.replace(microsecond=0)` to ensure filesystem-level timestamp variations do not trigger unnecessary reprocessing.

### What file pattern does filter_entries_by_existing_files search for?

The `_get_existing_session_files` method scans the output directory for files matching the glob pattern `adr.*.json`. It extracts session identifiers and timestamps from these filenames to build the comparison map used by the filtering logic.

### Can I use filter_entries_by_existing_files outside the ADR CLI?

Yes. Instantiate `SessionObserver` directly and call `filter_entries_by_existing_files(entries, output_dir)` to manually deduplicate a list of `AgentEvent` objects against an existing directory. This is useful for custom batch processing scripts or integration with external data pipelines.

### Where are the deduplication tests located?

Unit tests verifying the filter behavior reside in [`Sensor/tests/test_observer.py`](https://github.com/uber/ADR/blob/main/Sensor/tests/test_observer.py) at lines 128-158. These tests confirm that the method correctly removes unchanged sessions while retaining entries with newer timestamps.