# Plugin-eval CLI Commands: Complete Reference for OpenAI Plugin Evaluation

> Explore the plugin-eval CLI commands: start, analyze, benchmark, and more. This reference helps you evaluate and optimize your OpenAI plugins with a local workflow.

- Repository: [OpenAI/plugins](https://github.com/openai/plugins)
- Tags: api-reference
- Published: 2026-09-10

---

**The `plugin-eval` CLI provides eight core commands—`start` (alias `guide`), `analyze`, `explain-budget`, `measurement-plan` (alias `recommend-measures`), `init-benchmark`, `benchmark`, `report`, and `compare`—that enable developers to evaluate, benchmark, and optimize OpenAI plugins through a local-first workflow.**

The `plugin-eval` command-line tool is distributed as part of the official [`openai/plugins`](https://github.com/openai/plugins) repository, offering skill and plugin authors a dedicated toolkit for static analysis and performance measurement. According to the source code in [`src/cli.js`](https://github.com/openai/plugins/blob/main/src/cli.js), the CLI parses the first positional argument as the command name and delegates execution to specialized handlers that generate workflow guides, analyze token budgets, and execute comparative benchmarks.

## Core Plugin-eval CLI Commands

The complete command surface is defined in the `usage()` function at [`src/cli.js#L10-L36`](https://github.com/openai/plugins/blob/main/plugins/plugin-eval/src/cli.js#L10) and implemented in the dispatch block spanning [`src/cli.js#L25-L118`](https://github.com/openai/plugins/blob/main/plugins/plugin-eval/src/cli.js#L25). Each command targets a specific phase of the plugin evaluation lifecycle.

### Workflow Generation: `start` and `guide`

The **`start`** command (aliased as **`guide`**) generates an interactive, chat-first workflow guide for a skill or plugin. This command accepts optional flags including `--request` to seed the guide with a specific evaluation goal and `--format` to specify output rendering. The implementation resides at [`src/cli.js#L59-L68`](https://github.com/openai/plugins/blob/main/plugins/plugin-eval/src/cli.js#L59).

### Static Analysis: `analyze`

The **`analyze`** command performs static analysis on a target plugin or skill directory, calculating metric packs and observed usage patterns. It returns a rich JSON payload containing detailed metadata about the codebase structure and token consumption characteristics. This logic is handled at [`src/cli.js#L25-L38`](https://github.com/openai/plugins/blob/main/plugins/plugin-eval/src/cli.js#L25) and relies on core utilities in `src/core/analyzePath`.

### Budget Explanation: `explain-budget`

Use **`explain-budget`** to generate a detailed cost-breakdown projection for the target plugin. This command analyzes the projected token usage and produces a formatted report showing anticipated API costs. The implementation at [`src/cli.js#L71-L78`](https://github.com/openai/plugins/blob/main/plugins/plugin-eval/src/cli.js#L71) interfaces with the budget calculation modules in `src/core/`.

### Measurement Planning: `measurement-plan` and `recommend-measures`

The **`measurement-plan`** command (aliased as **`recommend-measures`**) creates a strategic plan for capturing real-world usage metrics such as token consumption and latency. When provided with `--observed-usage` pointing to a JSONL file of historical data, the command generates tailored recommendations for instrumentation. See the handler at [`src/cli.js#L80-L88`](https://github.com/openai/plugins/blob/main/plugins/plugin-eval/src/cli.js#L80).

### Benchmark Initialization: `init-benchmark`

Before running performance tests, use **`init-benchmark`** to scaffold a benchmark configuration file. This command accepts a `--model` parameter (e.g., `gpt-4o`) and an `--output` path to define the test parameters. The configuration generator is implemented at [`src/cli.js#L91-L99`](https://github.com/openai/plugins/blob/main/plugins/plugin-eval/src/cli.js#L91).

### Benchmark Execution: `benchmark`

The **`benchmark`** command executes the actual performance test using a previously generated configuration file. It supports `--usage-out` for writing raw token consumption data and `--result-out` for structured results. The execution engine at [`src/cli.js#L102-L115`](https://github.com/openai/plugins/blob/main/plugins/plugin-eval/src/cli.js#L102) orchestrates calls to `src/core/runBenchmark`.

### Result Reporting: `report`

Use **`report`** to render a previously generated result JSON file into human-readable formats. The command supports `--format` options including `json`, `markdown`, and `html`, utilizing renderers located in `src/renderers/`. Implementation details are found at [`src/cli.js#L40-L47`](https://github.com/openai/plugins/blob/main/plugins/plugin-eval/src/cli.js#L40).

### Comparative Analysis: `compare`

The **`compare`** command performs a differential analysis between two result files—typically representing "before" and "after" states of plugin optimization. It outputs the comparison in the selected format, highlighting regressions or improvements in token usage and execution time. The diff logic resides at [`src/cli.js#L49-L56`](https://github.com/openai/plugins/blob/main/plugins/plugin-eval/src/cli.js#L49).

### Standard Help Flags

The CLI responds to `--help` or `-h` flags by printing the usage text defined in the `usage()` function. When invoked without a command, the tool defaults to displaying this help documentation.

## Source Code Architecture

The `plugin-eval` CLI is organized into modular directories that separate concerns:

- **[`src/cli.js`](https://github.com/openai/plugins/blob/main/src/cli.js)** – Contains the command-line parser, argument validation, and the dispatch block that routes commands to their respective handlers.
- **`src/core/`** – Houses the heavy-lifting implementations including `analyzePath`, `runBenchmark`, `buildWorkflowGuide`, and budget calculation engines.
- **`src/renderers/`** – Manages output formatting for the supported export formats (JSON, Markdown, HTML).
- **[`src/lib/files.js`](https://github.com/openai/plugins/blob/main/src/lib/files.js)** – Provides utility functions for reading configuration files and writing result outputs used across multiple commands.

## Practical Usage Examples

Generate an interactive workflow guide for a skill:

```bash
plugin-eval start ./skills/my-skill --request "evaluate this skill" --format markdown

```

Run static analysis and persist the JSON output:

```bash
plugin-eval analyze ./plugins/my-plugin --output analysis.json

```

Display a token cost breakdown:

```bash
plugin-eval explain-budget ./plugins/my-plugin --format markdown

```

Create a measurement plan using historical usage data:

```bash
plugin-eval measurement-plan ./plugins/my-plugin \
    --observed-usage usage.jsonl --format markdown

```

Initialize a benchmark configuration targeting GPT-4o:

```bash
plugin-eval init-benchmark ./plugins/my-plugin \
    --model gpt-4o --output benchmark.json

```

Execute the benchmark and capture usage metrics:

```bash
plugin-eval benchmark ./plugins/my-plugin \
    --config benchmark.json \
    --usage-out usage.jsonl \
    --result-out result.json

```

Render results as an HTML report:

```bash
plugin-eval report result.json --format html --output report.html

```

Compare two benchmark runs to measure optimization impact:

```bash
plugin-eval compare baseline.json result.json --format markdown

```

## Summary

- **`start`/`guide`** generates interactive workflow guides for plugin development.
- **`analyze`** performs static code analysis and returns JSON metric packs.
- **`explain-budget`** projects token costs for the target implementation.
- **`measurement-plan`/`recommend-measures`** creates instrumentation strategies based on observed usage.
- **`init-benchmark`** scaffolds configuration files for performance testing.
- **`benchmark`** executes the actual performance measurement against configured models.
- **`report`** renders result files into JSON, Markdown, or HTML formats.
- **`compare`** differentiates between two benchmark results to track improvements.

## Frequently Asked Questions

### How do I view all available plugin-eval commands?

Run `plugin-eval --help` or `plugin-eval -h` to display the complete usage text. As implemented in `src/cli.js#L12-L18`, the CLI automatically prints this documentation when invoked without a command or when the help flag is explicitly provided.

### What is the difference between `init-benchmark` and `benchmark`?

The **`init-benchmark`** command creates a configuration JSON file that defines test parameters and target models, while **`benchmark`** executes the actual evaluation using that configuration. You must run `init-benchmark` first to generate the config file that `benchmark` consumes via its `--config` parameter.

### Can I compare two different versions of the same plugin?

Yes. Use the **`compare`** command with two result JSON files from separate `benchmark` runs. For example: `plugin-eval compare v1-results.json v2-results.json --format markdown`. This executes the diff logic at `src/cli.js#L49-L56` to highlight changes in token usage, latency, or other captured metrics.

### What output formats does plugin-eval support?

According to the rendering implementations in `src/renderers/`, the CLI supports three output formats: **json** (structured data), **markdown** (human-readable documentation), and **html** (rich web reports). Specify your preferred format using the `--format` flag available on most commands.