How to Debug Failed or Incomplete Agent Runs in Harvey-Labs: A Complete Guide to Logs and Output Files
Debug failed or incomplete agent runs in Harvey-Labs by analyzing the JSON-L transcript file generated by _log_turn and _log_tool in harness/agent_loop.py, combined with the output directory contents managed by ToolExecutor.
Harvey-Labs provides a comprehensive logging system for its autonomous agent loop. When a run fails, hangs, or produces unexpected results, you can reconstruct the exact execution path using structured transcript files and the output directory. This guide walks through the debugging workflow based on the actual source code implementation in the harveyai/harvey-labs repository.
Understanding the Two Critical Debug Artefacts
Every agent run produces two persistent artefacts you need to examine:
| Artefact | Location | Contents |
|---|---|---|
Transcript file (*.jsonl) |
Specified via transcript_path parameter |
Complete record of every model turn, tool call, arguments, and result previews |
| Output directory | Passed to ToolExecutor(output_dir=...) |
All files created by tools, intermediate artefacts, and sandbox outputs |
According to the Harvey-Labs source code, the agent loop writes transcript entries through two dedicated logging functions: _log_turn and _log_tool. These append structured JSON lines to your specified path.
Step 1: Enable Full Transcript Logging
To capture debug information, you must explicitly provide a transcript_path when invoking run_agent():
from harness.agent_loop import run_agent
from harness.tools import ToolExecutor
tool_executor = ToolExecutor(
documents_dir="data/documents",
output_dir="runs/exp_001/output",
workspace_dir="runs/exp_001/workspace",
)
result = run_agent(
adapter=my_model_adapter, # OpenAI, Anthropic, etc.
system_prompt=SYSTEM_PROMPT,
user_prompt=USER_PROMPT,
tool_executor=tool_executor,
max_turns=200,
transcript_path="runs/exp_001/transcript.jsonl", # Critical for debugging
)
The transcript_path argument triggers the logging system. Each turn produces a JSON object containing:
turn: Sequential turn numberrole:"assistant","tool", or"system"text: Model's text outputtool_calls: List of tools the model requestedtoken_usage: Input and output token countsresult_preview: First 1000 characters of tool results (see_log_toolimplementation)
Step 2: Locate Your Output Directory Structure
The ToolExecutor class receives output_dir during initialization and stores it in self.output_dir. All tool writes are rooted here:
# Inside ToolExecutor.__init__
self.output_dir = Path(output_dir).resolve()
Tools create files relative to this directory. If your agent generates a memo, it appears at output_dir / "memo.docx". For debugging, organize runs into isolated directories:
runs/
└── exp_001/
├── transcript.jsonl # Complete execution log
├── output/ # Tool-generated files
│ ├── memo.docx
│ └── financial_model.xlsx
└── workspace/ # Sandboxed working files
Step 3: Summarize Transcript with Built-in Utilities
Harvey-Labs includes [utils/playback.py](https://github.com/harveyai/harvey-labs/blob/main/utils/playback.py) specifically for analyzing completed or failed runs.
Quick Summary Command
from utils.playback import print_transcript_summary, playback_output_dir
from pathlib import Path
# One-line summary of turn count, tool calls, and errors
print_transcript_summary("runs/exp_001/transcript.jsonl")
# Tree view of all output files with sizes
playback_output_dir(Path("runs/exp_001/output"))
The print_transcript_summary function reports:
- Total turns executed
- Number of tool invocations
- Any error-containing result previews
This immediately reveals whether the run completed normally or terminated early.
Step 4: Deep Inspection of Individual Turns
For detailed analysis, iterate through transcript entries:
from utils.playback import load_transcript
for entry in load_transcript("runs/exp_001/transcript.jsonl"):
print(f"Turn {entry['turn']} | Role: {entry['role']}")
if entry["role"] == "assistant":
print(f" Model output: {entry['text'][:200]}...")
elif entry["role"] == "tool":
print(f" Tool: {entry['tool_name']}")
print(f" Arguments: {entry['arguments']}")
print(f" Result preview: {entry['result_preview'][:500]}")
# Detect error indicators in tool responses
if "error" in entry.get("result_preview", "").lower():
print(" ⚠️ ERROR DETECTED")
The load_transcript generator yields parsed JSON objects line-by-line without loading the entire file into memory.
Step 5: Detect Context Window Failures
The agent loop catches context overflow specifically. In [harness/agent_loop.py](https://github.com/harveyai/harvey-labs/blob/main/harness/agent_loop.py#L66-L73), the code detects prompt is too long or context_length_exceeded errors:
# From the source: detection and flagging of context overflow
except Exception as e:
if "prompt is too long" in str(e) or "context_length_exceeded" in str(e):
print("Context overflow: prompt too long")
result = {
"content": None,
"tool_calls": None,
"token_usage": None,
"context_overflow": True,
}
Check your result object:
if result.get("context_overflow"):
print("Run aborted: Model context window exceeded")
print(f"Finished cleanly: {result.get('finished_cleanly')}")
This indicates you need to reduce prompt size, compress conversation history, or switch to a model with larger context.
Step 6: Analyze Tool-Level Metrics
After any run, examine ToolExecutor.get_metrics():
metrics = tool_executor.get_metrics()
print(f"Tool calls made: {metrics['tool_calls']}")
print(f"Total tool time: {metrics['tool_time_seconds']:.2f}s")
Unexpectedly low call counts suggest:
- Premature termination (
max_turnstoo low) - Model refusing to use tools
- Early exception in loop
Compare metrics against your transcript summary to identify discrepancies.
Common Failure Patterns and Solutions
| Symptom | Transcript Indicator | Solution |
|---|---|---|
| Run ends mid-task | finished_cleanly: false, fewer turns than max_turns |
Check for exceptions in final turn; increase max_turns |
| Tool errors visible | result_preview contains "error" |
Fix tool implementation or input validation |
| No files in output directory | Zero tool calls in metrics | Review system prompt; model may not understand tool availability |
| Context overflow flagged | context_overflow: true |
Trim system prompt, use conversation compression, or upgrade model |
| Run appears successful but output missing | Tool reports success but file not found | Check ToolExecutor write permissions and path handling |
Complete Debugging Script
Combine all techniques into a reusable diagnostic:
#!/usr/bin/env python3
"""
Debug script for Harvey-Labs agent runs.
Usage: python debug_run.py runs/exp_001/transcript.jsonl
"""
import sys
import json
from pathlib import Path
from utils.playback import load_transcript, playback_output_dir
def diagnose_run(transcript_path: str):
transcript_path = Path(transcript_path)
output_dir = transcript_path.parent / "output"
print("=" * 60)
print(f"Analyzing: {transcript_path}")
print("=" * 60)
# Load and scan transcript
turns = 0
tool_calls = 0
errors = []
context_overflow = False
for entry in load_transcript(transcript_path):
turns += 1
if entry.get("role") == "tool":
tool_calls += 1
preview = entry.get("result_preview", "")
if preview and "error" in preview.lower():
errors.append({
"turn": entry.get("turn"),
"tool": entry.get("tool_name"),
"preview": preview[:300]
})
# Check for context overflow in metadata turns
if entry.get("context_overflow"):
context_overflow = True
# Report findings
print(f"\nTotal turns: {turns}")
print(f"Tool calls: {tool_calls}")
print(f"Context overflow: {context_overflow}")
if errors:
print(f"\n⚠️ Errors detected ({len(errors)}):")
for err in errors[:5]: # Show first 5
print(f" Turn {err['turn']} ({err['tool']}): {err['preview'][:100]}...")
# Output directory status
print(f"\nOutput directory contents:")
if output_dir.exists():
playback_output_dir(output_dir)
else:
print(f" ⚠️ Directory not found: {output_dir}")
if __name__ == "__main__":
diagnose_run(sys.argv[1])
Summary
- Enable transcripts: Always pass
transcript_pathtorun_agent()to capture the execution record via_log_turnand_log_tool - Use playback utilities: [
utils/playback.py](https://github.com/harveyai/harvey-labs/blob/main/utils/playback.py) providesprint_transcript_summaryandplayback_output_dirfor rapid diagnosis - Check context overflow: The loop explicitly flags this condition in results from [
harness/agent_loop.py](https://github.com/harveyai/harvey-labs/blob/main/harness/agent_loop.py#L66-L73) - Verify tool metrics:
ToolExecutor.get_metrics()reveals actual tool execution counts versus expectations - Inspect result previews: Every tool response includes a truncated preview in the transcript for quick error spotting without opening full output files
Frequently Asked Questions
Where does Harvey-Labs store the transcript file?
The transcript is written to the path you specify via the transcript_path parameter in run_agent(). If you omit this parameter, no transcript is generated and debugging capabilities are severely limited.
How do I know if my agent run failed due to context window limits?
Check the context_overflow field in your result dictionary. The agent loop in [harness/agent_loop.py](https://github.com/harveyai/harvey-labs/blob/main/harness/agent_loop.py#L66-L73) catches "prompt is too long" and "context_length_exceeded" errors, sets this flag to True, and terminates the run with finished_cleanly: false.
Can I replay a transcript to see exactly what the agent did?
Yes. Use utils/playback.load_transcript() to iterate through every turn, or print_transcript_summary() for a condensed view. Each entry contains the model's reasoning, tool arguments, and result previews sufficient to reconstruct the decision chain.
What if a tool reports success but the expected file is missing?
Examine the result_preview field in the transcript for that tool call. The preview is truncated to 1000 characters, but often reveals path errors, permission failures, or sandbox restrictions. Also verify ToolExecutor was initialized with the correct output_dir and that the directory exists with write permissions.
How can I programmatically detect errors across many runs?
Use summarize_transcript() from utils/playback which returns a dictionary with turns, tool_calls, and an errors list containing any entries where "error" appears in the result preview. Batch-process transcripts to identify failed runs automatically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →