# How Metrics Are Tracked in Harvey-Labs metrics.json and How Document Coverage Is Calculated

> Discover how Harvey-Labs tracks run metadata, execution stats, and document coverage in metrics.json. Learn the formula for calculating document coverage.

- Repository: [Harvey/harvey-labs](https://github.com/harveyai/harvey-labs)
- Tags: how-to-guide
- Published: 2026-08-11

---

**Harvey-Labs tracks run metadata, execution statistics, tool usage, and document-level metrics in a [`metrics.json`](https://github.com/harveyai/harvey-labs/blob/main/metrics.json) file, where document coverage is calculated as `documents_read / total_documents` based on deduplicated file access logs.**

The `harvey-labs` repository provides a benchmarking harness for evaluating large language models on document-based tasks. After every run, it persists detailed telemetry to [`metrics.json`](https://github.com/harveyai/harvey-labs/blob/main/metrics.json), including a precise measurement of how thoroughly the model explored available source documents. This article explains the full metrics schema and the exact mechanism behind document coverage calculation.

## What Metrics Are Stored in metrics.json

The [`metrics.json`](https://github.com/harveyai/harvey-labs/blob/main/metrics.json) file is written at the end of each run in [`harness/run.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/run.py) (lines 64-78). The data breaks down into five logical categories:

| Category | Fields | Source |
|----------|--------|--------|
| **Run identification** | `model`, `task`, `run_id`, `completed_at` | [`harness/run.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/run.py) |
| **Execution statistics** | `turn_count`, `input_tokens`, `output_tokens`, `total_tokens`, `wall_clock_seconds`, `finished_cleanly` | [`harness/run.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/run.py) |
| **Tool usage counters** | `bash_commands`, `files_written`, `files_edited`, `glob_searches`, `grep_searches` | [`harness/tools.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/tools.py) via `get_metrics()` |
| **Document-level metrics** | `total_documents`, `documents_read`, `documents_read_list`, `documents_read`, `documents_skipped`, `documents_skipped_list` | [`harness/tools.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/tools.py) via `get_metrics()` |
| **Optional adapter metrics** | e.g., `costs`, `api_calls` | Task-specific adapters like [`adapters/openai.py`](https://github.com/harveyai/harvey-labs/blob/main/adapters/openai.py) |

The `total_tokens` field aggregates `input_tokens + output_tokens`. The `finished_cleanly` boolean indicates whether the run terminated normally or via timeout/error.

## How Document Coverage Is Calculated

Document coverage in harvey-labs derives from two distinct operations in [`harness/tools.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/tools.py):

1. **Establishing the total document pool** — At initialization, the `Tools` class scans the task's `documents/` directory recursively, collecting every file path. The count becomes `total_documents` (lines 46-52).

2. **Recording actual file reads** — Each `read_file` tool invocation appends the accessed path to `self.files_read`. After completion, `get_metrics()` (lines 53-58) deduplicates this list to produce `documents_read_list` and `documents_read`.

The final coverage ratio follows this simple division:

```python
document_coverage = documents_read / total_documents

```

### Deduplication and Skipped Files

The `get_metrics()` method handles list hygiene carefully. It:

- Removes duplicate reads using `dict.fromkeys()` to preserve order while deduplicating
- Computes `documents_skipped_list` as the set difference between all available documents and unique reads
- Returns both counts and full lists for traceability

This implementation ensures that re-reading the same file does not inflate coverage metrics.

## Document Metrics Implementation in harness/tools.py

The core logic resides in the `Tools.get_metrics()` method:

```python

# From harness/tools.py — document metrics collection

def get_metrics(self) -> dict:
    # Enumerate all available documents in the task directory

    all_documents = sorted(
        str(f.relative_to(self.documents_dir))
        for f in self.documents_dir.rglob("*") if f.is_file()
    )
    
    # Deduplicate read history while preserving order

    unique_reads = list(dict.fromkeys(self.files_read))
    
    # Identify never-accessed documents

    skipped = [f for f in all_documents if f not in unique_reads]

    return {
        "documents_read": len(unique_reads),
        "documents_read_list": unique_reads,
        "documents_skipped": len(skipped),
        "documents_skipped_list": skipped,
        "total_documents": len(all_documents),
        # ... additional tool counters omitted ...

    }

```

The `self.files_read` list is populated incrementally throughout the run whenever the `read_file` tool executes.

## Reading and Using metrics.json

After a run completes, the metrics file appears in the run's output directory:

```python

# Example: loading and inspecting run metrics

import json
from pathlib import Path

run_dir = Path("results/2024-04-01-run-001")
metrics = json.loads((run_dir / "metrics.json").read_text())

print(f"Model: {metrics['model']}")
print(f"Task: {metrics['task']}")
print(f"Turns: {metrics['turn_count']}")
print(f"Tokens: {metrics['total_tokens']}")
print(f"Document coverage: {metrics['documents_read']}/{metrics['total_documents']}")
print(f"Coverage %: {metrics['documents_read'] / metrics['total_documents']:.1%}")

# Inspect which files were never read

print("Skipped files:", metrics["documents_skipped_list"][:5])

```

This pattern enables automated analysis of model behavior across benchmark runs.

## Where Coverage Is Displayed

The coverage ratio surfaces in two interfaces:

- **Terminal summary** — [`harness/run.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/run.py) (lines 88-90) prints metrics including document counts after run completion
- **Interactive playback** — [`utils/playback.py`](https://github.com/harveyai/harvey-labs/blob/main/utils/playback.py) (lines 45-61) renders a coverage indicator in the terminal UI (hardcoded to 62 for tutorial demonstrations)

Both interfaces consume the same [`metrics.json`](https://github.com/harveyai/harvey-labs/blob/main/metrics.json) data, ensuring consistency between automated benchmarks and human review.

## Summary

- **harvey-labs metrics.json** contains five metric categories: run identification, execution statistics, tool usage, document-level data, and optional adapter-specific fields
- **Document coverage** equals `documents_read / total_documents`, calculated from deduplicated file access logs against the complete document directory scan
- The **counting mechanism** operates in [`harness/tools.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/tools.py) at initialization (total pool) and during tool execution (actual reads)
- **Deduplication** prevents double-counting repeated file accesses via `dict.fromkeys()` ordering
- Results persist to disk and appear in both the terminal summary and interactive playback UI

## Frequently Asked Questions

### What happens if a document is read multiple times during a run?

Only the first read counts toward `documents_read`. The `get_metrics()` method deduplicates `self.files_read` using `list(dict.fromkeys(self.files_read))`, which preserves order while removing duplicates. This ensures coverage metrics reflect unique document exploration rather than access frequency.

### Where is the metrics.json file located after a run completes?

The file path derives from the run directory configured at execution time. By default, runs write to `results/{timestamp}-{run_id}/metrics.json`. The exact path is constructed in [`harness/run.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/run.py) where `json.dump(metrics, f, indent=2)` executes (lines 64-78).

### Can I add custom metrics to the harvey-labs metrics.json without modifying core code?

Yes. The architecture supports adapter-specific metrics through optional fields populated by task adapters. For example, [`adapters/openai.py`](https://github.com/harveyai/harvey-labs/blob/main/adapters/openai.py) injects `costs` and `api_calls`. Create a custom adapter inheriting from the base interface and override the metrics return value to include additional fields alongside the standard set.

### Why might documents_read plus documents_skipped not equal total_documents?

In normal operation, `documents_read + documents_skipped == total_documents` holds by construction. If you observe a mismatch, verify that `self.documents_dir` was not modified during execution and that all file paths in `documents_read_list` use relative formatting matching the `all_documents` scan. The calculation in `get_metrics()` uses simple list comprehension: `skipped = [f for f in all_documents if f not in unique_reads]`.