# How to Implement Custom Evaluation Metrics for Cosmos 3 Generated Videos Using Cosmos Evaluator

> Learn how to implement custom evaluation metrics for Cosmos 3 videos with Cosmos Evaluator. Extend functionality by creating Python classes and registering them in metrics yaml for tailored video scoring.

- Repository: [NVIDIA Corporation/cosmos](https://github.com/NVIDIA/cosmos)
- Tags: how-to-guide
- Published: 2026-06-14

---

**You can extend Cosmos Evaluator by creating a Python class that inherits from `BaseMetric`, registering it in the [`metrics.yaml`](https://github.com/NVIDIA/cosmos/blob/main/metrics.yaml) manifest, and running the evaluator with the `--metrics` flag to score videos generated by Cosmos 3 using your own domain-specific logic.**

Cosmos Evaluator is the official evaluation companion for the NVIDIA Cosmos repository, designed to score the physical plausibility and temporal consistency of videos generated by Cosmos 3. When the built-in metrics do not cover your specific domain requirements, you can implement custom evaluation metrics for Cosmos 3 generated videos by extending the evaluator's plugin architecture. This approach allows you to inject proprietary scoring logic—such as sharpness calculations, audio-visual sync, or physics-based energy conservation—without modifying the core evaluator codebase.

## Understanding the Cosmos Evaluator Architecture

The Cosmos Evaluator repository is listed as part of the Cosmos 3 ecosystem in the root [`README.md`](https://github.com/NVIDIA/cosmos/blob/main/README.md) at line 667. The evaluation workflow demonstrated in the Physics IQ benchmark at [`evaluation/cosmos3/Physics_IQ/README.md`](https://github.com/NVIDIA/cosmos/blob/main/evaluation/cosmos3/Physics_IQ/README.md) illustrates how the evaluator processes generated video directories and produces quantitative reports.

At the core of the plugin system is the abstract `BaseMetric` class defined in [`cosmos_evaluator/metrics/base.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_evaluator/metrics/base.py). Any custom metric must inherit from this base class and implement the contract expected by the evaluator's discovery mechanism. The evaluator discovers available metrics through a manifest file—either [`metrics.yaml`](https://github.com/NVIDIA/cosmos/blob/main/metrics.yaml) or [`metrics.json`](https://github.com/NVIDIA/cosmos/blob/main/metrics.json)—which maps metric names to their fully-qualified Python class paths.

## Creating a Custom Metric Class

### Inheriting from BaseMetric

To create a valid metric, define a class that inherits from `BaseMetric` and implements the required interface. The class must provide a `name` property, a `compute` method that accepts a video path and returns a scalar float, and an optional `description` property for documentation.

### Implementing Required Methods

The `compute(self, video_path: str, **kwargs)` method contains your evaluation logic. This method receives the path to a generated video file and returns a numeric score. Higher or lower values indicate better quality depending on your metric's definition.

Below is a complete implementation of a custom sharpness metric that calculates the average Laplacian variance across video frames:

```python

# file: my_metrics/sharpness.py

from cosmos_evaluator.metrics.base import BaseMetric
import cv2
import numpy as np

class SharpnessScore(BaseMetric):
    """Compute an average Laplacian-based sharpness score for a video."""
    
    @property
    def name(self) -> str:
        return "sharpness"

    @property
    def description(self) -> str:
        return "Higher values indicate a sharper, less blurry video."

    def compute(self, video_path: str, **kwargs) -> float:
        cap = cv2.VideoCapture(video_path)
        laplacian_vals = []
        
        while True:
            ret, frame = cap.read()
            if not ret:
                break
            gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)
            lap = cv2.Laplacian(gray, cv2.CV_64F).var()
            laplacian_vals.append(lap)
            
        cap.release()
        return float(np.mean(laplacian_vals))

```

## Registering Your Metric in the Manifest

After implementing your class, you must register it in [`cosmos_evaluator/metrics.yaml`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_evaluator/metrics.yaml) (or [`metrics.json`](https://github.com/NVIDIA/cosmos/blob/main/metrics.json)) to make it discoverable. Add an entry specifying the metric name and the fully-qualified class path:

```yaml

# file: cosmos_evaluator/metrics.yaml

metrics:
  - name: "sharpness"
    class_path: "my_metrics.sharpness.SharpnessScore"
  # … keep the existing built-in entries …

```

This registration step decouples your implementation from the core evaluator code, allowing the tool to instantiate your metric dynamically at runtime.

## Running the Evaluator with Custom Metrics

Once registered, invoke the evaluator from the command line using the `--metrics` flag to specify which metrics to run. You can pass a single metric name or a comma-separated list of metrics.

```bash

# Assume generated videos are in ./outputs/video/

cosmos-evaluator \
  --input-dir ./outputs/video/ \
  --metrics sharpness

```

If you omit the `--metrics` flag, the evaluator runs all discovered metrics defined in the manifest. The command produces an [`evaluation_report.json`](https://github.com/NVIDIA/cosmos/blob/main/evaluation_report.json) file containing scores for each video processed.

## Consuming Evaluation Results

The evaluator outputs a JSON report mapping each video filename to its metric scores. Parse this report to extract quantitative feedback on your Cosmos 3 generated videos:

```python
import json

with open("evaluation_report.json") as f:
    report = json.load(f)

for vid, scores in report["videos"].items():
    print(f"{vid}: Sharpness = {scores['sharpness']:.2f}")

```

This workflow enables automated benchmarking pipelines where your custom evaluation metrics for Cosmos 3 generated videos drive iterative improvements to your generation prompts or model configurations.

## Summary

- **Inherit from `BaseMetric`** located in [`cosmos_evaluator/metrics/base.py`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_evaluator/metrics/base.py) to create valid metric classes that the evaluator recognizes.

- **Implement the `name` property and `compute` method** to return scalar scores based on video analysis, with optional `description` for documentation.

- **Register in [`metrics.yaml`](https://github.com/NVIDIA/cosmos/blob/main/metrics.yaml)** using the fully-qualified class path to make your metric discoverable without modifying core evaluator code.

- **Run via command line** with `--metrics` flag targeting your Cosmos 3 output directory, then parse the resulting JSON report for quantitative analysis.

## Frequently Asked Questions

### What methods must a custom metric implement?

Your class must implement the `name` property returning a string identifier and the `compute(self, video_path, **kwargs)` method returning a float. You may also implement a `description` property to document the metric's purpose. These requirements are enforced by the `BaseMetric` abstract class in the evaluator source.

### Where do I register new metrics for Cosmos Evaluator?

Register new metrics in the [`cosmos_evaluator/metrics.yaml`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_evaluator/metrics.yaml) or [`cosmos_evaluator/metrics.json`](https://github.com/NVIDIA/cosmos/blob/main/cosmos_evaluator/metrics.json) manifest file. Each entry requires a `name` field and a `class_path` field pointing to your fully-qualified Python class (e.g., `"my_metrics.sharpness.SharpnessScore"`).

### Can I run multiple custom metrics simultaneously?

Yes. Pass a comma-separated list to the `--metrics` flag when running the evaluator, or omit the flag entirely to run all registered metrics against your Cosmos 3 generated videos. The resulting JSON report will contain scores for all specified metrics.

### What format does the evaluation report use?

The evaluator outputs a JSON file containing a top-level object with a `videos` key. This maps each video filename to an object containing metric names and their corresponding numeric scores, allowing programmatic consumption by downstream analysis tools.