Plugin-eval CLI Commands: Complete Reference for OpenAI Plugin Evaluation

The plugin-eval CLI provides eight core commands—start (alias guide), analyze, explain-budget, measurement-plan (alias recommend-measures), init-benchmark, benchmark, report, and compare—that enable developers to evaluate, benchmark, and optimize OpenAI plugins through a local-first workflow.

The plugin-eval command-line tool is distributed as part of the official openai/plugins repository, offering skill and plugin authors a dedicated toolkit for static analysis and performance measurement. According to the source code in src/cli.js, the CLI parses the first positional argument as the command name and delegates execution to specialized handlers that generate workflow guides, analyze token budgets, and execute comparative benchmarks.

Core Plugin-eval CLI Commands

The complete command surface is defined in the usage() function at src/cli.js#L10-L36 and implemented in the dispatch block spanning src/cli.js#L25-L118. Each command targets a specific phase of the plugin evaluation lifecycle.

Workflow Generation: start and guide

The start command (aliased as guide) generates an interactive, chat-first workflow guide for a skill or plugin. This command accepts optional flags including --request to seed the guide with a specific evaluation goal and --format to specify output rendering. The implementation resides at src/cli.js#L59-L68.

Static Analysis: analyze

The analyze command performs static analysis on a target plugin or skill directory, calculating metric packs and observed usage patterns. It returns a rich JSON payload containing detailed metadata about the codebase structure and token consumption characteristics. This logic is handled at src/cli.js#L25-L38 and relies on core utilities in src/core/analyzePath.

Budget Explanation: explain-budget

Use explain-budget to generate a detailed cost-breakdown projection for the target plugin. This command analyzes the projected token usage and produces a formatted report showing anticipated API costs. The implementation at src/cli.js#L71-L78 interfaces with the budget calculation modules in src/core/.

Measurement Planning: measurement-plan and recommend-measures

The measurement-plan command (aliased as recommend-measures) creates a strategic plan for capturing real-world usage metrics such as token consumption and latency. When provided with --observed-usage pointing to a JSONL file of historical data, the command generates tailored recommendations for instrumentation. See the handler at src/cli.js#L80-L88.

Benchmark Initialization: init-benchmark

Before running performance tests, use init-benchmark to scaffold a benchmark configuration file. This command accepts a --model parameter (e.g., gpt-4o) and an --output path to define the test parameters. The configuration generator is implemented at src/cli.js#L91-L99.

Benchmark Execution: benchmark

The benchmark command executes the actual performance test using a previously generated configuration file. It supports --usage-out for writing raw token consumption data and --result-out for structured results. The execution engine at src/cli.js#L102-L115 orchestrates calls to src/core/runBenchmark.

Result Reporting: report

Use report to render a previously generated result JSON file into human-readable formats. The command supports --format options including json, markdown, and html, utilizing renderers located in src/renderers/. Implementation details are found at src/cli.js#L40-L47.

Comparative Analysis: compare

The compare command performs a differential analysis between two result files—typically representing "before" and "after" states of plugin optimization. It outputs the comparison in the selected format, highlighting regressions or improvements in token usage and execution time. The diff logic resides at src/cli.js#L49-L56.

Standard Help Flags

The CLI responds to --help or -h flags by printing the usage text defined in the usage() function. When invoked without a command, the tool defaults to displaying this help documentation.

Source Code Architecture

The plugin-eval CLI is organized into modular directories that separate concerns:

  • src/cli.js – Contains the command-line parser, argument validation, and the dispatch block that routes commands to their respective handlers.
  • src/core/ – Houses the heavy-lifting implementations including analyzePath, runBenchmark, buildWorkflowGuide, and budget calculation engines.
  • src/renderers/ – Manages output formatting for the supported export formats (JSON, Markdown, HTML).
  • src/lib/files.js – Provides utility functions for reading configuration files and writing result outputs used across multiple commands.

Practical Usage Examples

Generate an interactive workflow guide for a skill:

plugin-eval start ./skills/my-skill --request "evaluate this skill" --format markdown

Run static analysis and persist the JSON output:

plugin-eval analyze ./plugins/my-plugin --output analysis.json

Display a token cost breakdown:

plugin-eval explain-budget ./plugins/my-plugin --format markdown

Create a measurement plan using historical usage data:

plugin-eval measurement-plan ./plugins/my-plugin \
    --observed-usage usage.jsonl --format markdown

Initialize a benchmark configuration targeting GPT-4o:

plugin-eval init-benchmark ./plugins/my-plugin \
    --model gpt-4o --output benchmark.json

Execute the benchmark and capture usage metrics:

plugin-eval benchmark ./plugins/my-plugin \
    --config benchmark.json \
    --usage-out usage.jsonl \
    --result-out result.json

Render results as an HTML report:

plugin-eval report result.json --format html --output report.html

Compare two benchmark runs to measure optimization impact:

plugin-eval compare baseline.json result.json --format markdown

Summary

  • start/guide generates interactive workflow guides for plugin development.
  • analyze performs static code analysis and returns JSON metric packs.
  • explain-budget projects token costs for the target implementation.
  • measurement-plan/recommend-measures creates instrumentation strategies based on observed usage.
  • init-benchmark scaffolds configuration files for performance testing.
  • benchmark executes the actual performance measurement against configured models.
  • report renders result files into JSON, Markdown, or HTML formats.
  • compare differentiates between two benchmark results to track improvements.

Frequently Asked Questions

How do I view all available plugin-eval commands?

Run plugin-eval --help or plugin-eval -h to display the complete usage text. As implemented in src/cli.js#L12-L18, the CLI automatically prints this documentation when invoked without a command or when the help flag is explicitly provided.

What is the difference between init-benchmark and benchmark?

The init-benchmark command creates a configuration JSON file that defines test parameters and target models, while benchmark executes the actual evaluation using that configuration. You must run init-benchmark first to generate the config file that benchmark consumes via its --config parameter.

Can I compare two different versions of the same plugin?

Yes. Use the compare command with two result JSON files from separate benchmark runs. For example: plugin-eval compare v1-results.json v2-results.json --format markdown. This executes the diff logic at src/cli.js#L49-L56 to highlight changes in token usage, latency, or other captured metrics.

What output formats does plugin-eval support?

According to the rendering implementations in src/renderers/, the CLI supports three output formats: json (structured data), markdown (human-readable documentation), and html (rich web reports). Specify your preferred format using the --format flag available on most commands.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →