How to Implement Custom Evaluation Metrics for Cosmos 3 Generated Videos Using Cosmos Evaluator

You can extend Cosmos Evaluator by creating a Python class that inherits from BaseMetric, registering it in the metrics.yaml manifest, and running the evaluator with the --metrics flag to score videos generated by Cosmos 3 using your own domain-specific logic.

Cosmos Evaluator is the official evaluation companion for the NVIDIA Cosmos repository, designed to score the physical plausibility and temporal consistency of videos generated by Cosmos 3. When the built-in metrics do not cover your specific domain requirements, you can implement custom evaluation metrics for Cosmos 3 generated videos by extending the evaluator's plugin architecture. This approach allows you to inject proprietary scoring logic—such as sharpness calculations, audio-visual sync, or physics-based energy conservation—without modifying the core evaluator codebase.

Understanding the Cosmos Evaluator Architecture

The Cosmos Evaluator repository is listed as part of the Cosmos 3 ecosystem in the root README.md at line 667. The evaluation workflow demonstrated in the Physics IQ benchmark at evaluation/cosmos3/Physics_IQ/README.md illustrates how the evaluator processes generated video directories and produces quantitative reports.

At the core of the plugin system is the abstract BaseMetric class defined in cosmos_evaluator/metrics/base.py. Any custom metric must inherit from this base class and implement the contract expected by the evaluator's discovery mechanism. The evaluator discovers available metrics through a manifest file—either metrics.yaml or metrics.json—which maps metric names to their fully-qualified Python class paths.

Creating a Custom Metric Class

Inheriting from BaseMetric

To create a valid metric, define a class that inherits from BaseMetric and implements the required interface. The class must provide a name property, a compute method that accepts a video path and returns a scalar float, and an optional description property for documentation.

Implementing Required Methods

The compute(self, video_path: str, **kwargs) method contains your evaluation logic. This method receives the path to a generated video file and returns a numeric score. Higher or lower values indicate better quality depending on your metric's definition.

Below is a complete implementation of a custom sharpness metric that calculates the average Laplacian variance across video frames:


# file: my_metrics/sharpness.py

from cosmos_evaluator.metrics.base import BaseMetric
import cv2
import numpy as np

class SharpnessScore(BaseMetric):
    """Compute an average Laplacian-based sharpness score for a video."""
    
    @property
    def name(self) -> str:
        return "sharpness"

    @property
    def description(self) -> str:
        return "Higher values indicate a sharper, less blurry video."

    def compute(self, video_path: str, **kwargs) -> float:
        cap = cv2.VideoCapture(video_path)
        laplacian_vals = []
        
        while True:
            ret, frame = cap.read()
            if not ret:
                break
            gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)
            lap = cv2.Laplacian(gray, cv2.CV_64F).var()
            laplacian_vals.append(lap)
            
        cap.release()
        return float(np.mean(laplacian_vals))

Registering Your Metric in the Manifest

After implementing your class, you must register it in cosmos_evaluator/metrics.yaml (or metrics.json) to make it discoverable. Add an entry specifying the metric name and the fully-qualified class path:


# file: cosmos_evaluator/metrics.yaml

metrics:
  - name: "sharpness"
    class_path: "my_metrics.sharpness.SharpnessScore"
  # … keep the existing built-in entries …

This registration step decouples your implementation from the core evaluator code, allowing the tool to instantiate your metric dynamically at runtime.

Running the Evaluator with Custom Metrics

Once registered, invoke the evaluator from the command line using the --metrics flag to specify which metrics to run. You can pass a single metric name or a comma-separated list of metrics.


# Assume generated videos are in ./outputs/video/

cosmos-evaluator \
  --input-dir ./outputs/video/ \
  --metrics sharpness

If you omit the --metrics flag, the evaluator runs all discovered metrics defined in the manifest. The command produces an evaluation_report.json file containing scores for each video processed.

Consuming Evaluation Results

The evaluator outputs a JSON report mapping each video filename to its metric scores. Parse this report to extract quantitative feedback on your Cosmos 3 generated videos:

import json

with open("evaluation_report.json") as f:
    report = json.load(f)

for vid, scores in report["videos"].items():
    print(f"{vid}: Sharpness = {scores['sharpness']:.2f}")

This workflow enables automated benchmarking pipelines where your custom evaluation metrics for Cosmos 3 generated videos drive iterative improvements to your generation prompts or model configurations.

Summary

  • Inherit from BaseMetric located in cosmos_evaluator/metrics/base.py to create valid metric classes that the evaluator recognizes.

  • Implement the name property and compute method to return scalar scores based on video analysis, with optional description for documentation.

  • Register in metrics.yaml using the fully-qualified class path to make your metric discoverable without modifying core evaluator code.

  • Run via command line with --metrics flag targeting your Cosmos 3 output directory, then parse the resulting JSON report for quantitative analysis.

Frequently Asked Questions

What methods must a custom metric implement?

Your class must implement the name property returning a string identifier and the compute(self, video_path, **kwargs) method returning a float. You may also implement a description property to document the metric's purpose. These requirements are enforced by the BaseMetric abstract class in the evaluator source.

Where do I register new metrics for Cosmos Evaluator?

Register new metrics in the cosmos_evaluator/metrics.yaml or cosmos_evaluator/metrics.json manifest file. Each entry requires a name field and a class_path field pointing to your fully-qualified Python class (e.g., "my_metrics.sharpness.SharpnessScore").

Can I run multiple custom metrics simultaneously?

Yes. Pass a comma-separated list to the --metrics flag when running the evaluator, or omit the flag entirely to run all registered metrics against your Cosmos 3 generated videos. The resulting JSON report will contain scores for all specified metrics.

What format does the evaluation report use?

The evaluator outputs a JSON file containing a top-level object with a videos key. This maps each video filename to an object containing metric names and their corresponding numeric scores, allowing programmatic consumption by downstream analysis tools.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →