# Complete Guide to Ponytail Commands: benchmark-local, agentic run, judge, and complete

> Explore Ponytail commands: benchmark-local, agentic run, judge, and complete. This guide details each script to help you optimize your local AI benchmarks and agentic workflows.

- Repository: [DietrichGebert/ponytail](https://github.com/DietrichGebert/ponytail)
- Tags: api-reference
- Published: 2026-08-29

---

**Ponytail provides four primary CLI commands—`benchmark-local`, `agentic run`, `agentic judge`, and `agentic complete`—all implemented as standalone Python scripts in the `benchmarks/` directory using the standard library `argparse` module.**

If you are exploring what commands are available in Ponytail, you will find a minimalist command-line interface designed for benchmarking and evaluating agentic AI workloads. The DietrichGebert/ponytail repository exposes its functionality through discrete entry-point scripts rather than a monolithic binary, keeping the codebase small and focused.

## Overview of Available Ponytail Commands

Ponytail ships with four user-facing commands, each residing in its own module under the `benchmarks/` package. Every command follows a consistent pattern: parse arguments with `argparse.ArgumentParser`, load input data, execute the core routine, and emit results in JSON or human-readable format.

- **`benchmark-local`**: Runs local latency benchmarks against Ollama models
- **`agentic run`**: Executes parallel agentic tasks and writes results to JSON
- **`agentic judge`**: Scores generated answers against reference data
- **`agentic complete`**: Convenience wrapper that chains `run` and `judge` into one step

Each command exposes a `main()` function that serves as the entry point when the script is invoked directly.

## benchmark-local Command

The `benchmark-local` command is your entry point for performance testing local LLM inference.

### Implementation Details

In [`benchmarks/benchmark-local.py`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/benchmark-local.py), the `main()` function constructs an `argparse.ArgumentParser` to accept flags such as `--model`, `--prompt`, and `--repeat`. The script initializes an Ollama client, dispatches the prompt the specified number of times, and calculates timing statistics and token throughput.

This module does not depend on external CLI frameworks; it relies solely on Python’s standard library.

### Usage Example

```bash
python -m benchmarks.benchmark-local \
    --model llama3:8b \
    --prompt "Write a short poem about cats." \
    --repeat 5

```

The tool prints average latency, total execution time, and token usage statistics to stdout after completing all repetitions.

## agentic run Command

Use `agentic run` when you need to evaluate a batch of prompts in parallel.

### Implementation Details

The source file [`benchmarks/agentic/run.py`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/agentic/run.py) implements a worker-pool pattern. It accepts a `--tasks` JSON file containing prompt objects, distributes work across configurable `--workers` processes, and aggregates results into a structured JSON output file.

Key parameters include:
- `--tasks`: Path to the input JSON array of task definitions
- `--output`: Destination file for the results
- `--workers`: Integer controlling concurrency (default varies by system)

### Usage Example

```bash
python -m benchmarks.agentic.run \
    --tasks tasks.json \
    --output run-results.json \
    --workers 4

```

Each object in [`tasks.json`](https://github.com/DietrichGebert/ponytail/blob/main/tasks.json) should specify at minimum a prompt string; the script invokes the configured model for each entry and records the generated output.

## agentic judge Command

After generating answers, you typically need automated scoring against ground-truth references.

### Implementation Details

Located at [`benchmarks/agentic/judge.py`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/agentic/judge.py), this command loads the JSON produced by `agentic run` alongside a `--reference` JSON file. It compares each generated answer to its corresponding reference entry, computes per-task scores, and writes a detailed scoring report to the path specified by `--output`.

The scoring logic is implemented in the script’s `main()` function, allowing easy inspection or modification of evaluation criteria.

### Usage Example

```bash
python -m benchmarks.agentic.judge \
    --answers run-results.json \
    --reference reference.json \
    --output scores.json

```

The resulting [`scores.json`](https://github.com/DietrichGebert/ponytail/blob/main/scores.json) contains metric values for every task, enabling quantitative analysis of model performance.

## agentic complete Command

For rapid iteration, Ponytail offers a single-step evaluation pipeline.

### Implementation Details

[`benchmarks/agentic/complete.py`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/agentic/complete.py) acts as a thin orchestration layer. Its `main()` function internally invokes the logic from `agentic run` followed immediately by `agentic judge`, passing through the appropriate file paths. This eliminates the need for manual intermediate file management when you want end-to-end results.

### Usage Example

```bash
python -m benchmarks.agentic.complete \
    --tasks tasks.json \
    --reference reference.json \
    --output final-report.json

```

Executing this one command produces the same [`final-report.json`](https://github.com/DietrichGebert/ponytail/blob/main/final-report.json) that you would obtain by running `agentic run` and `agentic judge` sequentially.

## Architecture and Design Philosophy

According to the [`AGENTS.md`](https://github.com/DietrichGebert/ponytail/blob/main/AGENTS.md) file in the repository root, Ponytail follows a "lazy senior dev" philosophy. Instead of a complex CLI framework like Click or Typer, each command is a self-contained script using `argparse`. This design keeps dependencies minimal and allows users to import command logic directly into other Python programs.

The modular structure means you can:
- Pipe shell commands between benchmark stages
- Import `main()` from any script into a custom workflow
- Modify argument parsers without side effects on other commands

## Summary

- Ponytail exposes **four distinct commands**: `benchmark-local`, `agentic run`, `agentic judge`, and `agentic complete`.
- All commands are implemented as standalone scripts in `benchmarks/` using standard library `argparse`.
- **[`benchmark-local.py`](https://github.com/DietrichGebert/ponytail/blob/main/benchmark-local.py)** handles local Ollama benchmarking with timing statistics.
- **[`agentic/run.py`](https://github.com/DietrichGebert/ponytail/blob/main/agentic/run.py)** executes parallel task workloads and outputs JSON results.
- **[`agentic/judge.py`](https://github.com/DietrichGebert/ponytail/blob/main/agentic/judge.py)** evaluates generated answers against reference data.
- **[`agentic/complete.py`](https://github.com/DietrichGebert/ponytail/blob/main/agentic/complete.py)** chains run and judge for one-shot evaluation workflows.
- No external CLI frameworks are required; entry points are the `main()` functions in each file.

## Frequently Asked Questions

### How do I install Ponytail to use these commands?

Ponytail is designed to run directly from the repository without installation. Clone the DietrichGebert/ponytail repository and invoke scripts using `python -m benchmarks.<script>`. Ensure you have Python 3.x and the dependencies listed in the project’s requirements installed in your environment.

### Can I run Ponytail commands programmatically from Python?

Yes. Because each command is a standalone module with a `main()` function, you can import `main` from `benchmarks.benchmark-local`, `benchmarks.agentic.run`, or any other module and invoke it programmatically. Alternatively, instantiate the `argparse.ArgumentParser` logic directly for finer control.

### What dependencies are required for the Ponytail CLI?

The CLI relies primarily on the Python standard library, specifically the `argparse` module for interface handling. Specific commands may require external libraries for HTTP communication (such as the Ollama Python client) or JSON processing, but the CLI framework itself has no heavy dependencies.

### How does the agentic complete command differ from running run and judge separately?

The `agentic complete` command in [`benchmarks/agentic/complete.py`](https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/agentic/complete.py) is a convenience wrapper that executes the `agentic run` workflow followed immediately by `agentic judge` using the intermediate files generated in memory or temporary storage. It produces identical output to running the two steps manually but reduces command-line boilerplate and file management overhead.