harvey-labs
A benchmark built to evaluate and improve agent capabilities for supporting legal work.
Discover how Harvey AI's reasoning effort parameters (low, medium, high, max) impact model behavior. Improve complex problem-solving by understanding token allocation and latency trade-offs.
How to Interpret Criterion-Level Results in Harvey-Labs `scores.json`Decode harvey-labs scores.json criterion-level results. Understand verdict and reasoning for task pass or fail to gain deep insights into your test outcomes.
How Harvey‑Labs Prevents Symlink Escape Attacks During Glob and Grep OperationsDiscover how Harvey Labs protects against symlink escape attacks in glob and grep operations. Learn how real path resolution secures your sandbox environment.
How Harvey-Labs Comparison Dashboards Generate Aggregate Metrics Across Multiple RunsDiscover how Harvey-Labs comparison dashboards aggregate metrics across multiple runs by scanning results, deduplicating, grouping by model, and computing pooled pass rates.
How Skill Scripts Are Copied Into the Workspace and Made Available to Agents in Harvey-LabsLearn how Harvey-Labs copies skill scripts into a workspace and makes them available to agents for bash execution. Discover the process of exposing scripts via mounted sandbox directories.
Environment Variables Harvey-Labs Supports and How .env Auto-Loading WorksDiscover supported environment variables in harvey-labs and understand how .env auto-loading simplifies configuration using python-dotenv for seamless access via os.getenv().
Harvey-Labs Document Parsing Pipeline: How AI Agents Extract Text from .docx, .pptx, .xlsx, and .pdf FilesDiscover the Harvey-Labs document parsing pipeline that uses AI agents to extract text from docx, pptx, xlsx, and pdf files. Learn how it securely handles binary files in an isolated container.
How to Configure Custom Shell Command Timeouts in Harvey AI for Long-Running OperationsLearn how to configure custom shell command timeouts in Harvey AI. Use the CLI flag or Sandbox() to manage long-running operations and prevent infinite loops.
How Harvey AI Handles Context Window Overflow and Token Limit Exceedances: 3-Layer Protection SystemDiscover how Harvey AI tackles context window overflow with its 3-layer protection system, including token capping, runtime accounting, and graceful failure.
How All-Pass Rubric Scoring Works in Harvey-Labs: A Complete Technical GuideDiscover how all-pass rubric scoring in Harvey-Labs works. This guide details the LLM judge's binary pass/fail evaluation for task criteria and its pipeline propagation.
How to Debug Failed or Incomplete Agent Runs in Harvey-Labs: A Complete Guide to Logs and Output FilesDebug failed Harvey-Labs agent runs by analyzing JSON-L transcripts and output files from ToolExecutor. Master agent loop logs for complete troubleshooting.
How Metrics Are Tracked in Harvey-Labs metrics.json and How Document Coverage Is CalculatedDiscover how Harvey-Labs tracks run metadata, execution stats, and document coverage in metrics.json. Learn the formula for calculating document coverage.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →