How to Use the plugin-eval Tool to Evaluate Plugin Quality
The plugin-eval tool is a dual-purpose Node.js CLI and Codex plugin bundle that performs static analysis, budget calculations, coverage checks, and live benchmarking to evaluate plugin quality.
The plugin-eval tool resides in the openai/plugins repository and provides a unified framework for assessing both skills (chat-first components) and plugins (full Codex bundles). It separates the chat-first front-end from the command-line back-end to deliver comprehensive quality reports through deterministic analysis and optional live execution testing.
Core Architecture of the plugin-eval Tool
The architecture distinguishes between the CLI infrastructure and the Codex plugin interface. The CLI entry point at scripts/plugin-eval.js bootstraps the application by invoking runCli from src/cli.js. This command dispatcher maps sub-commands—such as analyze, explain-budget, and benchmark—to their respective core functions.
The analysis pipeline in src/core/analyze.js orchestrates the entire evaluation workflow. It delegates target resolution to src/core/target.js, which detects whether the supplied path is a skill (containing SKILL.md) or a plugin (containing .codex-plugin/plugin.json). The pipeline then invokes specialized evaluators:
src/evaluators/code.jsperforms static analysis and lint-style checks.src/evaluators/coverage.jsparses coverage reports (lcov.info,coverage.xml).src/evaluators/plugin.jsandsrc/evaluators/skill.jsvalidate manifest-specific rules.
Budget profiling occurs in src/core/budget.js and src/core/baseline.js, which calculate token budgets (trigger, invoke, and deferred) and compare them against established baselines. The presentation layer in src/core/presentation.js shapes the final JSON payload and renders Markdown or HTML output, while src/core/workflow-guide.js generates actionable next-step recommendations. The metric-pack extension system in src/core/metric-packs.js allows external modules to contribute additional checks.
The Codex plugin side lives under the skills/ directory, forwarding natural-language requests to the CLI while preserving a chat-first user experience. The plugin manifest at .codex-plugin/plugin.json enables discovery within the Codex ecosystem.
Running the plugin-eval Tool from the Command Line
You can execute the plugin-eval tool without global installation by invoking the Node.js script directly.
Local Execution Without Installation
Navigate to the repository root and run:
node ./scripts/plugin-eval.js analyze ./my-plugin --format markdown
This command executes static analysis on the target path. The --format markdown flag produces a human-readable report, while --format json yields the raw result object for programmatic processing.
Installing as a Global Command
For frequent use, link the package to your global PATH:
npm link
plugin-eval --help
After linking, invoke the tool directly using commands like:
plugin-eval start ./my-skill --request "What should I fix first?" --format markdown
The start command acts as a chat-first router that explains why a request maps to a specific workflow and displays the first concrete command to execute.
Complete plugin-eval Workflow for Quality Assessment
A typical evaluation follows a structured progression from initial guidance through detailed benchmarking.
1. Start with a Workflow Guide
Begin with the start command to receive a concise evaluation plan:
plugin-eval start ./my-plugin --request "Evaluate this plugin." --format markdown
This leverages src/core/workflow-guide.js to generate context-aware next actions.
2. Run Static Analysis
Execute deterministic analysis to validate code structure, manifest correctness, and coverage metrics:
plugin-eval analyze ./my-plugin --format markdown
This invokes the core pipeline from src/core/analyze.js and outputs findings regarding code quality, budget compliance, and coverage status.
3. Explain Budget Constraints
Before running live tests, review token budget allocations:
plugin-eval explain-budget ./my-plugin --format markdown
This displays the calculated trigger, invoke, and deferred budgets against their baselines.
4. Initialize Benchmark Configuration
Generate a starter configuration under .plugin-eval/:
plugin-eval init-benchmark ./my-plugin
Edit the generated benchmark-config.json to customize prompts and test parameters.
5. Execute Live Benchmarking
Run real Codex executions in an isolated workspace to measure actual performance:
plugin-eval benchmark ./my-plugin --format markdown
This command performs live testing distinct from the static analysis phase.
6. Generate and Compare Reports
Convert raw JSON results into formatted documents:
plugin-eval report result.json --format markdown
Compare before and after states to measure improvement:
plugin-eval compare before.json after.json --format markdown
Using plugin-eval as a Codex Plugin
Once symlinked into ~/plugins or a workspace-local plugins/ folder, invoke the tool directly from chat:
$plugin-eval Evaluate this skill.
$plugin-eval What should I fix first?
$plugin-eval Help me benchmark this plugin.
The skill implementations reside under skills/plugin-eval/ and forward requests to the underlying CLI commands, as defined in skills/plugin-eval/agents/openai.yaml.
Programmatic Integration and Extensions
Integrate the plugin-eval tool into custom Node.js scripts or extend its functionality with custom metric packs.
Importing the Analysis Library
Import public exports from src/index.js for programmatic use:
import { analyzePath } from "./plugins/plugin-eval/src/index.js";
(async () => {
const result = await analyzePath("./plugins/plugin-eval");
console.log(result.summary);
})();
Creating Custom Metric Packs
Extend evaluation capabilities by creating a metric pack manifest and runner:
// my-metric-pack/manifest.json
{
"name": "my-metric-pack",
"description": "Checks for deprecated APIs",
"run": "./run.js"
}
Implement the check logic in the runner file:
// run.js
export async function run(target) {
const { createCheck } = await import("../plugins/plugin-eval/src/core/schema.js");
const checks = [];
// Scan files for deprecated patterns
checks.push(createCheck({
id: "deprecated-api",
category: "best-practice",
severity: "warning",
status: "warn",
message: "Uses old-lib",
remediation: ["Upgrade to new-lib"]
}));
return { checks };
}
Invoke the analysis with your custom metric pack:
plugin-eval analyze ./my-plugin --metric-pack-manifests ./my-metric-pack/manifest.json
Summary
- The plugin-eval tool functions as both a Node.js CLI and a Codex plugin bundle for comprehensive quality assessment.
- Static analysis runs via
analyzeandexplain-budgetcommands, whilebenchmarkenables live execution testing. - The architecture separates target resolution (
src/core/target.js), budget profiling (src/core/budget.js), and specialized evaluators for code, coverage, plugins, and skills. - You can extend functionality through custom metric packs using the
--metric-pack-manifestsflag. - Both skills (with
SKILL.md) and plugins (with.codex-plugin/plugin.json) are automatically detected and evaluated according to their specific requirements.
Frequently Asked Questions
How does the plugin-eval tool distinguish between a skill and a plugin?
The tool uses src/core/target.js to inspect the target directory. It identifies a skill by the presence of a SKILL.md file and a plugin by the presence of .codex-plugin/plugin.json. This distinction determines which validation rules from src/evaluators/skill.js or src/evaluators/plugin.js are applied during analysis.
What coverage report formats does plugin-eval support?
According to src/evaluators/coverage.js, the tool parses standard coverage formats including lcov.info (LCOV) and coverage.xml (Cobertura). These files are automatically detected and analyzed during the static analysis phase to calculate coverage metrics.
Can I use plugin-eval without installing it globally?
Yes. You can run the tool directly using Node.js: node ./scripts/plugin-eval.js [command] [path]. Global installation via npm link is optional and only necessary if you want to use the plugin-eval command from any directory without specifying the full path to the script.
What are token budgets in the context of plugin-eval?
Token budgets refer to the three categories calculated by src/core/budget.js: trigger (initial activation cost), invoke (execution cost), and deferred (asynchronous operation cost). The tool compares these calculated values against baselines in src/core/baseline.js to flag potential performance or cost issues before deployment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →