How Metrics Are Tracked in Harvey-Labs metrics.json and How Document Coverage Is Calculated

Harvey-Labs tracks run metadata, execution statistics, tool usage, and document-level metrics in a metrics.json file, where document coverage is calculated as documents_read / total_documents based on deduplicated file access logs.

The harvey-labs repository provides a benchmarking harness for evaluating large language models on document-based tasks. After every run, it persists detailed telemetry to metrics.json, including a precise measurement of how thoroughly the model explored available source documents. This article explains the full metrics schema and the exact mechanism behind document coverage calculation.

What Metrics Are Stored in metrics.json

The metrics.json file is written at the end of each run in harness/run.py (lines 64-78). The data breaks down into five logical categories:

Category Fields Source
Run identification model, task, run_id, completed_at harness/run.py
Execution statistics turn_count, input_tokens, output_tokens, total_tokens, wall_clock_seconds, finished_cleanly harness/run.py
Tool usage counters bash_commands, files_written, files_edited, glob_searches, grep_searches harness/tools.py via get_metrics()
Document-level metrics total_documents, documents_read, documents_read_list, documents_read, documents_skipped, documents_skipped_list harness/tools.py via get_metrics()
Optional adapter metrics e.g., costs, api_calls Task-specific adapters like adapters/openai.py

The total_tokens field aggregates input_tokens + output_tokens. The finished_cleanly boolean indicates whether the run terminated normally or via timeout/error.

How Document Coverage Is Calculated

Document coverage in harvey-labs derives from two distinct operations in harness/tools.py:

  1. Establishing the total document pool — At initialization, the Tools class scans the task's documents/ directory recursively, collecting every file path. The count becomes total_documents (lines 46-52).

  2. Recording actual file reads — Each read_file tool invocation appends the accessed path to self.files_read. After completion, get_metrics() (lines 53-58) deduplicates this list to produce documents_read_list and documents_read.

The final coverage ratio follows this simple division:

document_coverage = documents_read / total_documents

Deduplication and Skipped Files

The get_metrics() method handles list hygiene carefully. It:

  • Removes duplicate reads using dict.fromkeys() to preserve order while deduplicating
  • Computes documents_skipped_list as the set difference between all available documents and unique reads
  • Returns both counts and full lists for traceability

This implementation ensures that re-reading the same file does not inflate coverage metrics.

Document Metrics Implementation in harness/tools.py

The core logic resides in the Tools.get_metrics() method:


# From harness/tools.py — document metrics collection

def get_metrics(self) -> dict:
    # Enumerate all available documents in the task directory

    all_documents = sorted(
        str(f.relative_to(self.documents_dir))
        for f in self.documents_dir.rglob("*") if f.is_file()
    )
    
    # Deduplicate read history while preserving order

    unique_reads = list(dict.fromkeys(self.files_read))
    
    # Identify never-accessed documents

    skipped = [f for f in all_documents if f not in unique_reads]

    return {
        "documents_read": len(unique_reads),
        "documents_read_list": unique_reads,
        "documents_skipped": len(skipped),
        "documents_skipped_list": skipped,
        "total_documents": len(all_documents),
        # ... additional tool counters omitted ...

    }

The self.files_read list is populated incrementally throughout the run whenever the read_file tool executes.

Reading and Using metrics.json

After a run completes, the metrics file appears in the run's output directory:


# Example: loading and inspecting run metrics

import json
from pathlib import Path

run_dir = Path("results/2024-04-01-run-001")
metrics = json.loads((run_dir / "metrics.json").read_text())

print(f"Model: {metrics['model']}")
print(f"Task: {metrics['task']}")
print(f"Turns: {metrics['turn_count']}")
print(f"Tokens: {metrics['total_tokens']}")
print(f"Document coverage: {metrics['documents_read']}/{metrics['total_documents']}")
print(f"Coverage %: {metrics['documents_read'] / metrics['total_documents']:.1%}")

# Inspect which files were never read

print("Skipped files:", metrics["documents_skipped_list"][:5])

This pattern enables automated analysis of model behavior across benchmark runs.

Where Coverage Is Displayed

The coverage ratio surfaces in two interfaces:

  • Terminal summary — harness/run.py (lines 88-90) prints metrics including document counts after run completion
  • Interactive playback — utils/playback.py (lines 45-61) renders a coverage indicator in the terminal UI (hardcoded to 62 for tutorial demonstrations)

Both interfaces consume the same metrics.json data, ensuring consistency between automated benchmarks and human review.

Summary

  • harvey-labs metrics.json contains five metric categories: run identification, execution statistics, tool usage, document-level data, and optional adapter-specific fields
  • Document coverage equals documents_read / total_documents, calculated from deduplicated file access logs against the complete document directory scan
  • The counting mechanism operates in harness/tools.py at initialization (total pool) and during tool execution (actual reads)
  • Deduplication prevents double-counting repeated file accesses via dict.fromkeys() ordering
  • Results persist to disk and appear in both the terminal summary and interactive playback UI

Frequently Asked Questions

What happens if a document is read multiple times during a run?

Only the first read counts toward documents_read. The get_metrics() method deduplicates self.files_read using list(dict.fromkeys(self.files_read)), which preserves order while removing duplicates. This ensures coverage metrics reflect unique document exploration rather than access frequency.

Where is the metrics.json file located after a run completes?

The file path derives from the run directory configured at execution time. By default, runs write to results/{timestamp}-{run_id}/metrics.json. The exact path is constructed in harness/run.py where json.dump(metrics, f, indent=2) executes (lines 64-78).

Can I add custom metrics to the harvey-labs metrics.json without modifying core code?

Yes. The architecture supports adapter-specific metrics through optional fields populated by task adapters. For example, adapters/openai.py injects costs and api_calls. Create a custom adapter inheriting from the base interface and override the metrics return value to include additional fields alongside the standard set.

Why might documents_read plus documents_skipped not equal total_documents?

In normal operation, documents_read + documents_skipped == total_documents holds by construction. If you observe a mismatch, verify that self.documents_dir was not modified during execution and that all file paths in documents_read_list use relative formatting matching the all_documents scan. The calculation in get_metrics() uses simple list comprehension: skipped = [f for f in all_documents if f not in unique_reads].

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →