How VSS Evaluators Use LLM-Judge Functionality to Score Trajectory QA Report Quality
VSS evaluators leverage the LLM-Judge metric to automatically grade trajectory QA reports through single-field comparisons and dynamic field discovery, defaulting to this method when no specific evaluation strategy is configured.
The Video Search & Summarization (VSS) platform from the NVIDIA AI Blueprints repository (video-search-and-summarization) implements a sophisticated evaluation pipeline that scores generated reports against reference data. At the heart of this system lies the LLM-Judge functionality, which provides both deterministic field scoring and flexible discovery capabilities for evolving report schemas.
Understanding the LLM-Judge Architecture in VSS
The evaluation architecture centers on the ReportEvaluator class, which orchestrates quality assessment through pluggable metrics. The LLMJudgeMetric serves as the primary evaluation engine, capable of comparing reference and generated values using custom prompts while supporting batch evaluation for unspecified fields.
The metric implements two distinct operational modes:
- Single-field evaluation (
evaluate): Compares individual reference and generated values with reasoning support - Dynamic field discovery (
evaluate_with_field_discovery): Identifies and scores fields not explicitly defined in the evaluation configuration
These capabilities enable VSS to maintain evaluation coverage even when report structures change dynamically during trajectory generation.
Two Evaluation Modes: Single-Field vs. Dynamic Discovery
Single-Field Comparison with Reasoning
The standard evaluation path processes explicit field configurations through the evaluate method. When ReportEvaluator.evaluate_tree encounters a configured field, it invokes the LLM-Judge with a formatted prompt containing both the reference value and the generated candidate.
# agent/src/vss_agents/evaluators/report_evaluator/field_evaluators/llm_judge.py
@register_metric("llm_judge")
class LLMJudgeMetric(EvaluationMetric):
def evaluate(self, reference: Any, generated: Any, **kwargs) -> FieldEvaluation:
# Formats prompt with thinking tags for reasoning extraction
prompt = self._format_prompt(reference, generated)
# Returns structured score ∈ [0, 1]
The metric optionally injects thinking tags to capture the LLM's reasoning process before extracting the final numeric score, providing transparency into evaluation decisions.
Field Discovery for Unspecified Trajectory Fields
When the configuration enables allow_dynamic_field_discovery, the evaluator handles fields present in the generated report but absent from the explicit configuration. The evaluate_with_field_discovery method processes these dynamic fields in batch:
# Inside evaluate_tree (evaluate.py)
if allow_dynamic_discovery:
actual_unspecified = set(actual_dict.keys()) - set(explicit_fields)
if actual_unspecified:
llm_judge = cast("LLMJudgeMetric", self.metric_instances["llm_judge"])
eval_results = await llm_judge.evaluate_with_field_discovery(
reference_section=reference,
actual_section=actual_dict,
unspecified_fields=list(actual_unspecified),
)
This method requires the multi_field_discovery_prompt (validated at lines 83-85 in llm_judge.py) and returns a mapping of field names to scores and reference mappings, which the evaluator transforms into EvaluationScore instances marked with the method "llm_judge_with_field_discovery".
Configuration and Registration Pipeline
YAML Configuration Loading
Evaluation behavior is controlled through YAML configuration files parsed into EvalMetricsConfig models defined in agent/src/vss_agents/evaluators/report_evaluator/eval_config_models.py. This configuration specifies which metrics to instantiate and whether to enable dynamic field discovery:
# Example configuration structure
metric_cfg = _load_eval_metrics_yaml("configs/eval_metrics.yaml")
Metric Registration with Decorators
The LLMJudgeMetric registers itself in the metric registry using the @register_metric decorator pattern, making it available by the string key "llm_judge":
# vss_agents/evaluators/report_evaluator/field_evaluators/llm_judge.py
@register_metric("llm_judge")
class LLMJudgeMetric(EvaluationMetric):
def __init__(self, llm, single_field_comparison_prompt, multi_field_discovery_prompt):
self.llm = llm
self.single_field_comparison_prompt = single_field_comparison_prompt
self.multi_field_discovery_prompt = multi_field_discovery_prompt
During evaluator initialization, ReportEvaluator receives a dictionary of metric instances, including the instantiated LLM-Judge with its configured prompts and underlying language model.
The Evaluation Orchestration Flow
Default Method Fallback
The ReportEvaluator.evaluate_tree method implements a critical fallback mechanism. When traversing the report tree, if a field's configuration specifies no evaluation method, the system automatically defaults to LLM-Judge:
# Inside evaluate_tree (evaluate.py) lines 57-59
if (method := config.method) is None:
method = "llm_judge" # <-- default fallback
logger.debug(f"No method specified for '{'.'.join(path)}', defaulting to llm_judge")
This ensures consistent evaluation coverage without requiring explicit method declarations for every field in the schema.
Recursive Tree Evaluation
The evaluator processes reports recursively, handling both scalar values and nested sections. For each node, it retrieves the corresponding FieldConfig, determines the appropriate metric (defaulting to LLM-Judge when unspecified), and delegates scoring to the metric instance. Discovery-enabled fields trigger the batch evaluation path at lines 99-107 of evaluate.py.
VLM Field Score Aggregation
Beyond general report quality, the system calculates specialized scores for visual-language model outputs. When include_vlm_output is enabled, the evaluator aggregates scores from VLM-related fields:
# compute VLM field score (evaluate_item)
if self.include_vlm_output and self.vlm_related_fields:
for section_name in self.vlm_related_fields:
section_eval = result.field_scores.get(section_name)
if section_eval and section_eval.section_score is not None:
vlm_scores.append(section_eval.section_score)
vlm_field_score = sum(vlm_scores) / len(vlm_scores) if vlm_scores else None
This produces the vlm_field_score exposed in ExtendedEvalOutputItem (lines 50-58), allowing separate tracking of multimodal generation quality distinct from the overall average_score.
Implementation Examples
Initializing the ReportEvaluator
metric_cfg = _load_eval_metrics_yaml("configs/eval_metrics.yaml")
metric_instances = {
"llm_judge": LLMJudgeMetric(
llm=my_chat_model,
single_field_comparison_prompt=SINGLE_PROMPT,
multi_field_discovery_prompt=MULTI_PROMPT,
),
# other metrics …
}
report_evaluator = ReportEvaluator(
config=metric_cfg,
metric_instances=metric_instances,
object_store_client=store,
report_url_pattern=r"https?://.*/([^/]+\.json)",
reference_base_dir="reference_reports",
include_vlm_output=True,
vlm_related_fields=["VLM_summary", "VLM_insights"],
)
Processing Dynamic Field Discovery Results
# Transforming discovery results into EvaluationScore objects
for field_name, eval_data in eval_results.items():
score = EvaluationScore(
field_name=field_name,
score=eval_data["score"],
method="llm_judge_with_field_discovery",
reference_field=eval_data.get("reference_field")
)
field_scores.append(score)
Summary
- VSS evaluators use
LLMJudgeMetricas the primary mechanism for assessing trajectory QA report quality, implementing both single-field and batch evaluation modes. - The system defaults to
"llm_judge"whenFieldConfig.methodis unspecified, ensuring comprehensive coverage without explicit configuration. - Dynamic field discovery via
evaluate_with_field_discoveryallows scoring of evolving report schemas through batch LLM prompts that identify and evaluate unspecified fields. - VLM-specific scoring tracks multimodal output quality separately from overall report scores through configurable field aggregation.
- Key files include
field_evaluators/llm_judge.py(metric implementation),evaluate.py(orchestration logic), andeval_config_models.py(configuration schemas).
Frequently Asked Questions
What is the default evaluation method when no method is specified in VSS?
When FieldConfig.method is None, the ReportEvaluator.evaluate_tree method automatically defaults to "llm_judge" according to the fallback logic at lines 57-59 of agent/src/vss_agents/evaluators/report_evaluator/evaluate.py. This ensures that all fields receive LLM-based quality scoring even without explicit metric assignment in the configuration YAML.
How does dynamic field discovery work in LLM-Judge?
Dynamic field discovery activates when allow_dynamic_field_discovery is enabled in the configuration. The evaluator identifies fields present in the generated report but absent from the explicit field list, then invokes LLMJudgeMetric.evaluate_with_field_discovery with the reference section, generated section, and list of unspecified field names. The LLM returns a JSON mapping containing scores and reference field mappings for each discovered field.
What is the purpose of VLM-related field scores?
The VLM field score provides targeted quality assessment for visual-language model outputs (such as video summaries or visual insights) separately from the overall report average. When include_vlm_output is enabled, the evaluator calculates this metric by averaging scores across configured VLM-related fields like "VLM_summary" and "VLM_insights", exposing it as vlm_field_score in the ExtendedEvalOutputItem result object.
How is the LLM-Judge metric registered in the VSS evaluator system?
The LLMJudgeMetric class decorates itself with @register_metric("llm_judge") in agent/src/vss_agents/evaluators/report_evaluator/field_evaluators/llm_judge.py. This registration pattern adds the metric to a global registry, allowing ReportEvaluator to instantiate and invoke it by string key during tree traversal based on the evaluation configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →