How the `cav trial` Command Performs A/B Testing and Generates Comparative Reports

The cav trial command executes sequential baseline and modified configuration runs, persists granular telemetry to a SQLite database, and renders a side‑by‑side markdown report that highlights differences in latency, token usage, and model output.

The cav trial command is a core feature of the JuliusBrussee/caveman CLI framework that enables developers to locally validate AI provider configurations with full observability. When operating in A/B testing mode, the command orchestrates dual trial sessions, stores detailed metrics in a local database, and generates diff‑based comparative reports using built‑in templating engines.

Understanding the Trial Architecture and Telemetry Capture

Local Engine Execution and Data Collection

When you invoke cav trial, the CLI instantiates a local engine that processes a single request against either a configured provider or a local fixture. According to docs/testing-session-modes.md, the engine captures comprehensive telemetry that includes raw request payloads, timing data for each rewriter stage, token usage statistics, cost estimates, and the final model output. The system serializes this data to a local SQLite database named caveman.db, creating a persistent record for subsequent analysis and reporting.

Database Schema for A/B Testing

The trial system utilizes a relational schema managed by the @caveman-io/engine-cli package. As documented in docs/technical/cli-reference.md, the runTrial function writes request payloads to the trial_payloads table and high‑level metrics to the trials table. When running comparison mode, the compareTrials utility creates a dedicated comparison record that links the two distinct trial IDs, establishing the relational foundation for delta calculations.

Running A/B Tests with the --compare Flag

The A/B testing workflow is initiated by appending the --compare flag to the standard trial invocation. When you execute cav trial --config baseline.yml --compare modified.yml, the CLI performs the following sequence:

  1. Baseline Execution – Runs the first trial using the specified baseline configuration and tags the record as "baseline"
  2. Comparison Execution – Immediately executes a second trial with the modified configuration and tags it as "comparison"
  3. Persistence – Stores both trial records in the same SQLite database instance
  4. Linking – Invokes the compareTrials utility to generate a unique comparison record that foreign‑keys the two trial IDs together

This automated sequencing ensures that both runs operate under identical environmental conditions while maintaining distinct identities for differential analysis.

Generating Comparative Reports

Rendering with markdown-it and Custom Templates

The cav trial report subcommand retrieves the comparison record from the database, merges the two result sets, and calculates metric deltas. As specified in docs/technical/cli-reference.md, the rendering pipeline uses the markdown-it library combined with a custom template located at templates/trial-report.md. The generated document includes a summary table of request‑level metrics and a dedicated "Δ‑Δ" section that displays the calculated differences between the baseline and comparison runs for latency, token counts, and cost estimates.

Text Diff Visualization with diff-match-patch

For analyzing changes in model‑generated content, the report integrates the diff-match-patch npm package. This library produces human‑readable, line‑by‑line diffs that highlight insertions and deletions in the textual output, making it straightforward to identify semantic drift between configuration versions without manual text comparison.

Output Formats and Storage Locations

By default, the comparative report is printed to stdout and simultaneously written as a markdown file to the caveman-reports/ directory (which is automatically added to .gitignore). For distribution‑ready documentation, you can export the report to PDF using the --output pdf flag, which processes the markdown through a headless Chromium renderer.

Implementation Details from the Source Code

The underlying execution logic resides in the @caveman-io/engine-cli dependency rather than the main repository. The runTrial function handles the serialization of telemetry to the trial_payloads and trials tables, while compareTrials manages the creation of the relational comparison record. The Caveman CLI repository provides the orchestration layer that sequences these function calls and invokes the rendering pipeline as documented in docs/technical/cli-reference.md.

Practical Code Examples


# Execute a single baseline trial with full telemetry

cav trial --config baseline.yml

# Run an A/B comparison between baseline and modified configurations

cav trial --config baseline.yml --compare modified.yml

# Generate a markdown report from a specific comparison ID

cav trial report --compare-id 42

# Export the comparative report to PDF format

cav trial report --compare-id 42 --output pdf

Summary

  • The cav trial command captures granular request telemetry—including latency, token usage, and rewriter timing—into a local SQLite database
  • The --compare flag triggers sequential execution of baseline and modified configurations, automatically tagging runs and creating relational comparison records
  • The cav trial report command uses markdown-it with custom templates and diff-match-patch for visualizing textual differences
  • Generated reports default to the caveman-reports/ directory with optional PDF export via headless Chromium
  • Core execution functions runTrial and compareTrials are provided by the @caveman-io/engine-cli package dependency

Frequently Asked Questions

How does the cav trial command associate two separate runs as an A/B test?

The command stores each trial in the trials table of caveman.db and then creates a unique comparison record that links the baseline trial ID to the comparison trial ID. This relational structure allows the report generator to retrieve paired datasets and calculate metric deltas.

What specific metrics are compared in the generated report?

The report analyzes latency, token consumption counts, and cost estimates, presenting the differences in a "Δ‑Δ" section. It also provides a line‑by‑line diff of the model‑generated text to highlight behavioral changes between the two configuration versions.

Can trial reports be generated for single runs without a comparison?

Yes. The cav trial report command can render reports for individual trials; however, the comparative "Δ‑Δ" section and diff visualization require a valid comparison record created via the --compare flag during the trial phase.

Where is the report template located and can it be customized?

The default template resides at templates/trial-report.md within the project structure. Because the rendering engine uses markdown-it with standard markdown templates, you can modify this file or specify an alternative template path to customize the report layout and styling.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →