How to Use the plugin-eval Tool to Evaluate Plugin Quality

The plugin-eval tool is a dual-purpose Node.js CLI and Codex plugin bundle that performs static analysis, budget calculations, coverage checks, and live benchmarking to evaluate plugin quality.

The plugin-eval tool resides in the openai/plugins repository and provides a unified framework for assessing both skills (chat-first components) and plugins (full Codex bundles). It separates the chat-first front-end from the command-line back-end to deliver comprehensive quality reports through deterministic analysis and optional live execution testing.

Core Architecture of the plugin-eval Tool

The architecture distinguishes between the CLI infrastructure and the Codex plugin interface. The CLI entry point at scripts/plugin-eval.js bootstraps the application by invoking runCli from src/cli.js. This command dispatcher maps sub-commands—such as analyze, explain-budget, and benchmark—to their respective core functions.

The analysis pipeline in src/core/analyze.js orchestrates the entire evaluation workflow. It delegates target resolution to src/core/target.js, which detects whether the supplied path is a skill (containing SKILL.md) or a plugin (containing .codex-plugin/plugin.json). The pipeline then invokes specialized evaluators:

Budget profiling occurs in src/core/budget.js and src/core/baseline.js, which calculate token budgets (trigger, invoke, and deferred) and compare them against established baselines. The presentation layer in src/core/presentation.js shapes the final JSON payload and renders Markdown or HTML output, while src/core/workflow-guide.js generates actionable next-step recommendations. The metric-pack extension system in src/core/metric-packs.js allows external modules to contribute additional checks.

The Codex plugin side lives under the skills/ directory, forwarding natural-language requests to the CLI while preserving a chat-first user experience. The plugin manifest at .codex-plugin/plugin.json enables discovery within the Codex ecosystem.

Running the plugin-eval Tool from the Command Line

You can execute the plugin-eval tool without global installation by invoking the Node.js script directly.

Local Execution Without Installation

Navigate to the repository root and run:

node ./scripts/plugin-eval.js analyze ./my-plugin --format markdown

This command executes static analysis on the target path. The --format markdown flag produces a human-readable report, while --format json yields the raw result object for programmatic processing.

Installing as a Global Command

For frequent use, link the package to your global PATH:

npm link
plugin-eval --help

After linking, invoke the tool directly using commands like:

plugin-eval start ./my-skill --request "What should I fix first?" --format markdown

The start command acts as a chat-first router that explains why a request maps to a specific workflow and displays the first concrete command to execute.

Complete plugin-eval Workflow for Quality Assessment

A typical evaluation follows a structured progression from initial guidance through detailed benchmarking.

1. Start with a Workflow Guide

Begin with the start command to receive a concise evaluation plan:

plugin-eval start ./my-plugin --request "Evaluate this plugin." --format markdown

This leverages src/core/workflow-guide.js to generate context-aware next actions.

2. Run Static Analysis

Execute deterministic analysis to validate code structure, manifest correctness, and coverage metrics:

plugin-eval analyze ./my-plugin --format markdown

This invokes the core pipeline from src/core/analyze.js and outputs findings regarding code quality, budget compliance, and coverage status.

3. Explain Budget Constraints

Before running live tests, review token budget allocations:

plugin-eval explain-budget ./my-plugin --format markdown

This displays the calculated trigger, invoke, and deferred budgets against their baselines.

4. Initialize Benchmark Configuration

Generate a starter configuration under .plugin-eval/:

plugin-eval init-benchmark ./my-plugin

Edit the generated benchmark-config.json to customize prompts and test parameters.

5. Execute Live Benchmarking

Run real Codex executions in an isolated workspace to measure actual performance:

plugin-eval benchmark ./my-plugin --format markdown

This command performs live testing distinct from the static analysis phase.

6. Generate and Compare Reports

Convert raw JSON results into formatted documents:

plugin-eval report result.json --format markdown

Compare before and after states to measure improvement:

plugin-eval compare before.json after.json --format markdown

Using plugin-eval as a Codex Plugin

Once symlinked into ~/plugins or a workspace-local plugins/ folder, invoke the tool directly from chat:


$plugin-eval Evaluate this skill.
$plugin-eval What should I fix first?
$plugin-eval Help me benchmark this plugin.

The skill implementations reside under skills/plugin-eval/ and forward requests to the underlying CLI commands, as defined in skills/plugin-eval/agents/openai.yaml.

Programmatic Integration and Extensions

Integrate the plugin-eval tool into custom Node.js scripts or extend its functionality with custom metric packs.

Importing the Analysis Library

Import public exports from src/index.js for programmatic use:

import { analyzePath } from "./plugins/plugin-eval/src/index.js";

(async () => {
  const result = await analyzePath("./plugins/plugin-eval");
  console.log(result.summary);
})();

Creating Custom Metric Packs

Extend evaluation capabilities by creating a metric pack manifest and runner:

// my-metric-pack/manifest.json
{
  "name": "my-metric-pack",
  "description": "Checks for deprecated APIs",
  "run": "./run.js"
}

Implement the check logic in the runner file:

// run.js
export async function run(target) {
  const { createCheck } = await import("../plugins/plugin-eval/src/core/schema.js");
  const checks = [];
  
  // Scan files for deprecated patterns
  checks.push(createCheck({ 
    id: "deprecated-api", 
    category: "best-practice", 
    severity: "warning", 
    status: "warn", 
    message: "Uses old-lib", 
    remediation: ["Upgrade to new-lib"] 
  }));
  
  return { checks };
}

Invoke the analysis with your custom metric pack:

plugin-eval analyze ./my-plugin --metric-pack-manifests ./my-metric-pack/manifest.json

Summary

  • The plugin-eval tool functions as both a Node.js CLI and a Codex plugin bundle for comprehensive quality assessment.
  • Static analysis runs via analyze and explain-budget commands, while benchmark enables live execution testing.
  • The architecture separates target resolution (src/core/target.js), budget profiling (src/core/budget.js), and specialized evaluators for code, coverage, plugins, and skills.
  • You can extend functionality through custom metric packs using the --metric-pack-manifests flag.
  • Both skills (with SKILL.md) and plugins (with .codex-plugin/plugin.json) are automatically detected and evaluated according to their specific requirements.

Frequently Asked Questions

How does the plugin-eval tool distinguish between a skill and a plugin?

The tool uses src/core/target.js to inspect the target directory. It identifies a skill by the presence of a SKILL.md file and a plugin by the presence of .codex-plugin/plugin.json. This distinction determines which validation rules from src/evaluators/skill.js or src/evaluators/plugin.js are applied during analysis.

What coverage report formats does plugin-eval support?

According to src/evaluators/coverage.js, the tool parses standard coverage formats including lcov.info (LCOV) and coverage.xml (Cobertura). These files are automatically detected and analyzed during the static analysis phase to calculate coverage metrics.

Can I use plugin-eval without installing it globally?

Yes. You can run the tool directly using Node.js: node ./scripts/plugin-eval.js [command] [path]. Global installation via npm link is optional and only necessary if you want to use the plugin-eval command from any directory without specifying the full path to the script.

What are token budgets in the context of plugin-eval?

Token budgets refer to the three categories calculated by src/core/budget.js: trigger (initial activation cost), invoke (execution cost), and deferred (asynchronous operation cost). The tool compares these calculated values against baselines in src/core/baseline.js to flag potential performance or cost issues before deployment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →