Agent Performance Success Metrics and Completion Criteria in Agency-Agents

Agent performance success metrics and completion criteria in the Agency-Agents repository are defined through standardized "Your Success Metrics" sections in each agent's markdown definition file, specifying quantitative KPIs such as percentage improvements, automation coverage rates, and adoption thresholds that determine mission completion.

The msitarzewski/agency-agents repository implements a structured approach to defining autonomous agent roles, where every agent's performance is measured against explicitly defined completion criteria. Each agent definition follows a repeatable template that concludes with a "Your Success Metrics" section, enabling programmatic extraction and automated KPI monitoring across diverse agent types from workflow optimizers to developer advocates.

How Agency-Agents Defines Agent Performance Success Metrics

The repository treats every autonomous role as a self-contained knowledge work-stream, with success metrics embedded directly in the agent's definition file rather than external configuration.

The Standardized Template Structure

Every agent definition markdown file in the repository follows a consistent architectural pattern:

  1. Metadata front-matter (--- name: … description: … color: … ---)
  2. Personality block defining communication style
  3. Core mission statement
  4. Critical rules governing behavior
  5. Deliverables specifying tangible outputs
  6. Your Success Metrics section containing bullet-listed completion criteria

This structure appears across all domains, from testing/testing-workflow-optimizer.md to specialized/specialized-developer-advocate.md, enabling the orchestration layer in specialized/agents-orchestrator.md to parse goals programmatically and feed them into monitoring dashboards.

Quantitative KPIs and Measurable Outcomes

The "Your Success Metrics" sections rely on concrete, percentage-based indicators rather than vague qualitative goals. Common metric categories include:

  • Quantitative improvement: Target percentage gains on key performance indicators (e.g., 30-60% improvement in cycle-time)
  • Automation coverage: Portion of routine work automated with reliability (typically 50-80%)
  • Error/rework reduction: Desired cutback of defects (50-90% reduction targets)
  • Adoption/uptake: Rate of stakeholder acceptance (80-95% adoption within defined timeframes)
  • Employee/user satisfaction: Change in satisfaction scores (2-4 point lifts or 20-30% improvements)

Completion Criteria Across Different Agent Types

While the template remains consistent, specific completion criteria vary by agent specialization to reflect domain-specific outcomes.

Process Optimization Agents

The Workflow Optimizer agent defined in testing/testing-workflow-optimizer.md specifies rigorous operational metrics:

  • 40% average improvement in process completion time across optimized workflows
  • 60% of routine tasks automated with reliable performance and error handling
  • 75% reduction in process-related errors and rework through systematic improvement
  • 90% successful adoption rate for optimized processes within 6 months
  • 30% improvement in employee satisfaction scores for optimized workflows

These metrics enable automated validation through the agent's internal self._track_quality_success_metrics() method, ensuring generated plans meet quantitative targets before deployment.

Developer Experience Agents

The Developer Advocate agent in specialized/specialized-developer-advocate.md focuses on activation and time-to-value metrics:

  • Time-to-first-success for new developers ≤ 15 minutes
  • Developer activation rate and documentation satisfaction scores
  • Community engagement metrics including forum participation and content sharing rates

These criteria measure the agent's effectiveness in reducing friction for technical adoption, with specific thresholds that trigger completion status.

Project Management Agents

The Project Shepherd agent defined in project-management/project-management-project-shepherd.md emphasizes risk mitigation and stakeholder alignment:

  • 90% of identified risks successfully mitigated before impacting project outcomes
  • Stakeholder satisfaction scores and on-time delivery rates
  • Team velocity maintenance or improvement during process changes

This agent's completion criteria focus on proactive risk management rather than reactive issue resolution, requiring systematic tracking of risk states throughout the project lifecycle.

Programmatically Extracting Success Metrics from Agent Definitions

The standardized markdown structure enables automated extraction of completion criteria for monitoring and reporting pipelines.

Extracting Metrics from Individual Agent Files

The following Python function parses any agent definition file to retrieve its success metrics using regular expressions:

import re
from pathlib import Path

def load_success_metrics(md_path: Path) -> list[str]:
    """
    Reads a .md file and returns the bullet points that appear
    under the heading "Your Success Metrics".
    """
    content = md_path.read_text(encoding="utf-8")
    # Find the section heading (allow for optional emojis)

    match = re.search(
        r"##\s+🎯?\s*Your Success Metrics\s*(.*?)(?:\n##|\Z)", 
        content, 
        re.S
    )
    if not match:
        return []
    section = match.group(1)
    # Capture markdown list items that start with a dash

    bullets = re.findall(r"^\s*-\s+(.*)", section, re.M)
    return [b.strip() for b in bullets]

# Usage -------------------------------------------------

md_file = Path(
    "testing/testing-workflow-optimizer.md"
)
metrics = load_success_metrics(md_file)
print("\n".join(metrics))

Output:


40% average improvement in process completion time across optimized workflows
60% of routine tasks automated with reliable performance and error handling
75% reduction in process-related errors and rework through systematic improvement
90% successful adoption rate for optimized processes within 6 months
30% improvement in employee satisfaction scores for optimized workflows

Aggregating Metrics Across All Agents

For orchestration and cross-agent analysis, you can normalize metrics from multiple agent definitions into a structured DataFrame:

import pandas as pd
from pathlib import Path

agent_paths = Path(".").rglob("*.md")
records = []

for p in agent_paths:
    if "Your Success Metrics" not in p.read_text():
        continue
    metrics = load_success_metrics(p)
    records.append({"agent": p.stem, "metrics": metrics})

df = pd.DataFrame(records)
print(df.head())

This aggregation enables the orchestration layer referenced in specialized/agents-orchestrator.md to perform automated KPI checks, generate monitoring dashboards, and trigger alerts when agents deviate from their defined completion criteria.

Summary

  • Agent performance success metrics and completion criteria in the Agency-Agents repository are defined declaratively in each agent's markdown definition file under the "Your Success Metrics" section.
  • The repository enforces a standardized template across all agent types, including metadata frontmatter, personality definitions, core missions, and quantifiable success metrics that enable programmatic extraction.
  • Quantitative KPIs dominate the completion criteria, with specific percentage targets for improvement (30-60%), automation coverage (50-80%), error reduction (50-90%), and adoption rates (80-95%).
  • Domain-specific variations exist across agent types: Workflow Optimizers focus on process efficiency metrics, Developer Advocates track time-to-first-success, and Project Shepherds monitor risk mitigation rates.
  • Programmatic access to these metrics is supported through the consistent markdown structure, enabling automated KPI tracking, orchestration layer integration, and cross-agent performance analysis.

Frequently Asked Questions

How are success metrics structured in the Agency-Agents repository?

Success metrics are structured as bullet-listed completion criteria within a dedicated "Your Success Metrics" section at the end of each agent's markdown definition file. This section follows a standardized template that includes quantitative KPIs with specific percentage targets, automation coverage rates, and adoption thresholds. The consistent formatting allows the orchestration layer to parse these metrics programmatically using regular expressions, as implemented in the repository's agent monitoring utilities.

What types of quantitative KPIs are used to measure agent performance?

The repository utilizes several categories of quantitative KPIs including percentage improvements in process efficiency (typically 30-60%), automation coverage rates (50-80% of routine tasks), error and rework reduction percentages (50-90%), adoption rates among stakeholders (80-95% within defined timeframes), and satisfaction score improvements (20-30% uplift). For specialized agents, domain-specific metrics like time-to-first-success (≤15 minutes for developer-facing agents) and risk mitigation rates (≥90% for project management agents) are also employed.

Can success metrics be extracted programmatically from agent definition files?

Yes, success metrics can be extracted programmatically due to the standardized markdown structure used across all agent definitions. The repository follows a consistent pattern where the "Your Success Metrics" section appears as a level-2 heading with optional emoji prefixes, followed by bullet points using dash notation. Python scripts utilizing regular expressions can parse these sections to extract quantitative targets, enabling automated KPI tracking, integration with monitoring dashboards like Grafana or Power BI, and cross-agent performance comparisons through pandas DataFrames.

How do completion criteria differ between agent specializations?

While all agents share the same template structure, completion criteria vary significantly by domain. Process optimization agents like the Workflow Optimizer focus on operational metrics such as 40% improvement in completion time and 60% task automation. Developer experience agents prioritize activation metrics including ≤15 minute time-to-first-success and documentation satisfaction scores. Project management agents emphasize risk mitigation (90% of risks mitigated before impact) and stakeholder alignment. Support agents track customer satisfaction and proactive outreach rates, while design agents measure iteration speed and visual quality approval rates.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →