Complete Guide to Ponytail Commands: benchmark-local, agentic run, judge, and complete

Ponytail provides four primary CLI commands—benchmark-local, agentic run, agentic judge, and agentic complete—all implemented as standalone Python scripts in the benchmarks/ directory using the standard library argparse module.

If you are exploring what commands are available in Ponytail, you will find a minimalist command-line interface designed for benchmarking and evaluating agentic AI workloads. The DietrichGebert/ponytail repository exposes its functionality through discrete entry-point scripts rather than a monolithic binary, keeping the codebase small and focused.

Overview of Available Ponytail Commands

Ponytail ships with four user-facing commands, each residing in its own module under the benchmarks/ package. Every command follows a consistent pattern: parse arguments with argparse.ArgumentParser, load input data, execute the core routine, and emit results in JSON or human-readable format.

  • benchmark-local: Runs local latency benchmarks against Ollama models
  • agentic run: Executes parallel agentic tasks and writes results to JSON
  • agentic judge: Scores generated answers against reference data
  • agentic complete: Convenience wrapper that chains run and judge into one step

Each command exposes a main() function that serves as the entry point when the script is invoked directly.

benchmark-local Command

The benchmark-local command is your entry point for performance testing local LLM inference.

Implementation Details

In benchmarks/benchmark-local.py, the main() function constructs an argparse.ArgumentParser to accept flags such as --model, --prompt, and --repeat. The script initializes an Ollama client, dispatches the prompt the specified number of times, and calculates timing statistics and token throughput.

This module does not depend on external CLI frameworks; it relies solely on Python’s standard library.

Usage Example

python -m benchmarks.benchmark-local \
    --model llama3:8b \
    --prompt "Write a short poem about cats." \
    --repeat 5

The tool prints average latency, total execution time, and token usage statistics to stdout after completing all repetitions.

agentic run Command

Use agentic run when you need to evaluate a batch of prompts in parallel.

Implementation Details

The source file benchmarks/agentic/run.py implements a worker-pool pattern. It accepts a --tasks JSON file containing prompt objects, distributes work across configurable --workers processes, and aggregates results into a structured JSON output file.

Key parameters include:

  • --tasks: Path to the input JSON array of task definitions
  • --output: Destination file for the results
  • --workers: Integer controlling concurrency (default varies by system)

Usage Example

python -m benchmarks.agentic.run \
    --tasks tasks.json \
    --output run-results.json \
    --workers 4

Each object in tasks.json should specify at minimum a prompt string; the script invokes the configured model for each entry and records the generated output.

agentic judge Command

After generating answers, you typically need automated scoring against ground-truth references.

Implementation Details

Located at benchmarks/agentic/judge.py, this command loads the JSON produced by agentic run alongside a --reference JSON file. It compares each generated answer to its corresponding reference entry, computes per-task scores, and writes a detailed scoring report to the path specified by --output.

The scoring logic is implemented in the script’s main() function, allowing easy inspection or modification of evaluation criteria.

Usage Example

python -m benchmarks.agentic.judge \
    --answers run-results.json \
    --reference reference.json \
    --output scores.json

The resulting scores.json contains metric values for every task, enabling quantitative analysis of model performance.

agentic complete Command

For rapid iteration, Ponytail offers a single-step evaluation pipeline.

Implementation Details

benchmarks/agentic/complete.py acts as a thin orchestration layer. Its main() function internally invokes the logic from agentic run followed immediately by agentic judge, passing through the appropriate file paths. This eliminates the need for manual intermediate file management when you want end-to-end results.

Usage Example

python -m benchmarks.agentic.complete \
    --tasks tasks.json \
    --reference reference.json \
    --output final-report.json

Executing this one command produces the same final-report.json that you would obtain by running agentic run and agentic judge sequentially.

Architecture and Design Philosophy

According to the AGENTS.md file in the repository root, Ponytail follows a "lazy senior dev" philosophy. Instead of a complex CLI framework like Click or Typer, each command is a self-contained script using argparse. This design keeps dependencies minimal and allows users to import command logic directly into other Python programs.

The modular structure means you can:

  • Pipe shell commands between benchmark stages
  • Import main() from any script into a custom workflow
  • Modify argument parsers without side effects on other commands

Summary

  • Ponytail exposes four distinct commands: benchmark-local, agentic run, agentic judge, and agentic complete.
  • All commands are implemented as standalone scripts in benchmarks/ using standard library argparse.
  • benchmark-local.py handles local Ollama benchmarking with timing statistics.
  • agentic/run.py executes parallel task workloads and outputs JSON results.
  • agentic/judge.py evaluates generated answers against reference data.
  • agentic/complete.py chains run and judge for one-shot evaluation workflows.
  • No external CLI frameworks are required; entry points are the main() functions in each file.

Frequently Asked Questions

How do I install Ponytail to use these commands?

Ponytail is designed to run directly from the repository without installation. Clone the DietrichGebert/ponytail repository and invoke scripts using python -m benchmarks.<script>. Ensure you have Python 3.x and the dependencies listed in the project’s requirements installed in your environment.

Can I run Ponytail commands programmatically from Python?

Yes. Because each command is a standalone module with a main() function, you can import main from benchmarks.benchmark-local, benchmarks.agentic.run, or any other module and invoke it programmatically. Alternatively, instantiate the argparse.ArgumentParser logic directly for finer control.

What dependencies are required for the Ponytail CLI?

The CLI relies primarily on the Python standard library, specifically the argparse module for interface handling. Specific commands may require external libraries for HTTP communication (such as the Ollama Python client) or JSON processing, but the CLI framework itself has no heavy dependencies.

How does the agentic complete command differ from running run and judge separately?

The agentic complete command in benchmarks/agentic/complete.py is a convenience wrapper that executes the agentic run workflow followed immediately by agentic judge using the intermediate files generated in memory or temporary storage. It produces identical output to running the two steps manually but reduces command-line boilerplate and file management overhead.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →