# How VSS Evaluators Use LLM-Judge Functionality to Score Trajectory QA Report Quality

> Discover how VSS evaluators use LLM-judge to automatically score trajectory QA reports. Learn about single-field comparisons and dynamic discovery for quality assessment.

- Repository: [NVIDIA AI Blueprints/video-search-and-summarization](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization)
- Tags: how-to-guide
- Published: 2026-05-15

---

**VSS evaluators leverage the LLM-Judge metric to automatically grade trajectory QA reports through single-field comparisons and dynamic field discovery, defaulting to this method when no specific evaluation strategy is configured.**

The **Video Search & Summarization (VSS)** platform from the NVIDIA AI Blueprints repository (`video-search-and-summarization`) implements a sophisticated evaluation pipeline that scores generated reports against reference data. At the heart of this system lies the **LLM-Judge** functionality, which provides both deterministic field scoring and flexible discovery capabilities for evolving report schemas.

## Understanding the LLM-Judge Architecture in VSS

The evaluation architecture centers on the `ReportEvaluator` class, which orchestrates quality assessment through pluggable metrics. The **`LLMJudgeMetric`** serves as the primary evaluation engine, capable of comparing reference and generated values using custom prompts while supporting batch evaluation for unspecified fields.

The metric implements two distinct operational modes:

- **Single-field evaluation** (`evaluate`): Compares individual reference and generated values with reasoning support
- **Dynamic field discovery** (`evaluate_with_field_discovery`): Identifies and scores fields not explicitly defined in the evaluation configuration

These capabilities enable VSS to maintain evaluation coverage even when report structures change dynamically during trajectory generation.

## Two Evaluation Modes: Single-Field vs. Dynamic Discovery

### Single-Field Comparison with Reasoning

The standard evaluation path processes explicit field configurations through the `evaluate` method. When `ReportEvaluator.evaluate_tree` encounters a configured field, it invokes the LLM-Judge with a formatted prompt containing both the reference value and the generated candidate.

```python

# agent/src/vss_agents/evaluators/report_evaluator/field_evaluators/llm_judge.py

@register_metric("llm_judge")
class LLMJudgeMetric(EvaluationMetric):
    def evaluate(self, reference: Any, generated: Any, **kwargs) -> FieldEvaluation:
        # Formats prompt with thinking tags for reasoning extraction

        prompt = self._format_prompt(reference, generated)
        # Returns structured score ∈ [0, 1]

```

The metric optionally injects **thinking tags** to capture the LLM's reasoning process before extracting the final numeric score, providing transparency into evaluation decisions.

### Field Discovery for Unspecified Trajectory Fields

When the configuration enables `allow_dynamic_field_discovery`, the evaluator handles fields present in the generated report but absent from the explicit configuration. The `evaluate_with_field_discovery` method processes these dynamic fields in batch:

```python

# Inside evaluate_tree (evaluate.py)

if allow_dynamic_discovery:
    actual_unspecified = set(actual_dict.keys()) - set(explicit_fields)
    if actual_unspecified:
        llm_judge = cast("LLMJudgeMetric", self.metric_instances["llm_judge"])
        eval_results = await llm_judge.evaluate_with_field_discovery(
            reference_section=reference,
            actual_section=actual_dict,
            unspecified_fields=list(actual_unspecified),
        )

```

This method requires the `multi_field_discovery_prompt` (validated at lines 83-85 in [`llm_judge.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/llm_judge.py)) and returns a mapping of field names to scores and reference mappings, which the evaluator transforms into `EvaluationScore` instances marked with the method `"llm_judge_with_field_discovery"`.

## Configuration and Registration Pipeline

### YAML Configuration Loading

Evaluation behavior is controlled through YAML configuration files parsed into **`EvalMetricsConfig`** models defined in [`agent/src/vss_agents/evaluators/report_evaluator/eval_config_models.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/evaluators/report_evaluator/eval_config_models.py). This configuration specifies which metrics to instantiate and whether to enable dynamic field discovery:

```python

# Example configuration structure

metric_cfg = _load_eval_metrics_yaml("configs/eval_metrics.yaml")

```

### Metric Registration with Decorators

The `LLMJudgeMetric` registers itself in the metric registry using the `@register_metric` decorator pattern, making it available by the string key `"llm_judge"`:

```python

# vss_agents/evaluators/report_evaluator/field_evaluators/llm_judge.py

@register_metric("llm_judge")
class LLMJudgeMetric(EvaluationMetric):
    def __init__(self, llm, single_field_comparison_prompt, multi_field_discovery_prompt):
        self.llm = llm
        self.single_field_comparison_prompt = single_field_comparison_prompt
        self.multi_field_discovery_prompt = multi_field_discovery_prompt

```

During evaluator initialization, `ReportEvaluator` receives a dictionary of metric instances, including the instantiated LLM-Judge with its configured prompts and underlying language model.

## The Evaluation Orchestration Flow

### Default Method Fallback

The `ReportEvaluator.evaluate_tree` method implements a critical fallback mechanism. When traversing the report tree, if a field's configuration specifies no evaluation method, the system automatically defaults to LLM-Judge:

```python

# Inside evaluate_tree (evaluate.py) lines 57-59

if (method := config.method) is None:
    method = "llm_judge"          # <-- default fallback

    logger.debug(f"No method specified for '{'.'.join(path)}', defaulting to llm_judge")

```

This ensures consistent evaluation coverage without requiring explicit method declarations for every field in the schema.

### Recursive Tree Evaluation

The evaluator processes reports recursively, handling both scalar values and nested sections. For each node, it retrieves the corresponding `FieldConfig`, determines the appropriate metric (defaulting to LLM-Judge when unspecified), and delegates scoring to the metric instance. Discovery-enabled fields trigger the batch evaluation path at lines 99-107 of [`evaluate.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/evaluate.py).

### VLM Field Score Aggregation

Beyond general report quality, the system calculates specialized scores for visual-language model outputs. When `include_vlm_output` is enabled, the evaluator aggregates scores from VLM-related fields:

```python

# compute VLM field score (evaluate_item)

if self.include_vlm_output and self.vlm_related_fields:
    for section_name in self.vlm_related_fields:
        section_eval = result.field_scores.get(section_name)
        if section_eval and section_eval.section_score is not None:
            vlm_scores.append(section_eval.section_score)
vlm_field_score = sum(vlm_scores) / len(vlm_scores) if vlm_scores else None

```

This produces the **`vlm_field_score`** exposed in `ExtendedEvalOutputItem` (lines 50-58), allowing separate tracking of multimodal generation quality distinct from the overall `average_score`.

## Implementation Examples

### Initializing the ReportEvaluator

```python
metric_cfg = _load_eval_metrics_yaml("configs/eval_metrics.yaml")
metric_instances = {
    "llm_judge": LLMJudgeMetric(
        llm=my_chat_model,
        single_field_comparison_prompt=SINGLE_PROMPT,
        multi_field_discovery_prompt=MULTI_PROMPT,
    ),
    # other metrics …

}
report_evaluator = ReportEvaluator(
    config=metric_cfg,
    metric_instances=metric_instances,
    object_store_client=store,
    report_url_pattern=r"https?://.*/([^/]+\.json)",
    reference_base_dir="reference_reports",
    include_vlm_output=True,
    vlm_related_fields=["VLM_summary", "VLM_insights"],
)

```

### Processing Dynamic Field Discovery Results

```python

# Transforming discovery results into EvaluationScore objects

for field_name, eval_data in eval_results.items():
    score = EvaluationScore(
        field_name=field_name,
        score=eval_data["score"],
        method="llm_judge_with_field_discovery",
        reference_field=eval_data.get("reference_field")
    )
    field_scores.append(score)

```

## Summary

- **VSS evaluators** use `LLMJudgeMetric` as the primary mechanism for assessing trajectory QA report quality, implementing both single-field and batch evaluation modes.
- The system defaults to `"llm_judge"` when `FieldConfig.method` is unspecified, ensuring comprehensive coverage without explicit configuration.
- **Dynamic field discovery** via `evaluate_with_field_discovery` allows scoring of evolving report schemas through batch LLM prompts that identify and evaluate unspecified fields.
- **VLM-specific scoring** tracks multimodal output quality separately from overall report scores through configurable field aggregation.
- Key files include [`field_evaluators/llm_judge.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/field_evaluators/llm_judge.py) (metric implementation), [`evaluate.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/evaluate.py) (orchestration logic), and [`eval_config_models.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/eval_config_models.py) (configuration schemas).

## Frequently Asked Questions

### What is the default evaluation method when no method is specified in VSS?

When `FieldConfig.method` is `None`, the `ReportEvaluator.evaluate_tree` method automatically defaults to `"llm_judge"` according to the fallback logic at lines 57-59 of [`agent/src/vss_agents/evaluators/report_evaluator/evaluate.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/evaluators/report_evaluator/evaluate.py). This ensures that all fields receive LLM-based quality scoring even without explicit metric assignment in the configuration YAML.

### How does dynamic field discovery work in LLM-Judge?

Dynamic field discovery activates when `allow_dynamic_field_discovery` is enabled in the configuration. The evaluator identifies fields present in the generated report but absent from the explicit field list, then invokes `LLMJudgeMetric.evaluate_with_field_discovery` with the reference section, generated section, and list of unspecified field names. The LLM returns a JSON mapping containing scores and reference field mappings for each discovered field.

### What is the purpose of VLM-related field scores?

The **VLM field score** provides targeted quality assessment for visual-language model outputs (such as video summaries or visual insights) separately from the overall report average. When `include_vlm_output` is enabled, the evaluator calculates this metric by averaging scores across configured VLM-related fields like `"VLM_summary"` and `"VLM_insights"`, exposing it as `vlm_field_score` in the `ExtendedEvalOutputItem` result object.

### How is the LLM-Judge metric registered in the VSS evaluator system?

The `LLMJudgeMetric` class decorates itself with `@register_metric("llm_judge")` in [`agent/src/vss_agents/evaluators/report_evaluator/field_evaluators/llm_judge.py`](https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/blob/main/agent/src/vss_agents/evaluators/report_evaluator/field_evaluators/llm_judge.py). This registration pattern adds the metric to a global registry, allowing `ReportEvaluator` to instantiate and invoke it by string key during tree traversal based on the evaluation configuration.