How cluster.Scorer Determines Routing Decisions in the WorkWeave Router

The WorkWeave cluster.Scorer determines routing decisions by embedding the request prompt, filtering candidates through provider bindings and feature constraints, scoring models across the top-P nearest clusters using a quality-cost-speed blend, and selecting the highest-scoring candidate.

The cluster.Scorer is the central decision engine inside the WorkWeave open-source router. Implemented in internal/router/cluster/scorer.go, it transforms an incoming request into a concrete routing decision through a deterministic 11-step pipeline that balances latency, cost, and quality.

Request Preprocessing and Candidate Filtering

The scorer first sanitizes the request and narrows the candidate pool before any vector computation begins.

Prompt Truncation and Provider-Aware Eligibility

Every routing decision starts with input sanitization. The scorer calls TailTruncate (lines 37–47 in scorer.go) to enforce the cfg.MaxPromptChars limit (default 1024 characters), ensuring the embedding step remains computationally bounded. Simultaneously, the scorer evaluates EnabledProviders from the request. If specified, it narrows the full model registry (s.models) to only those providers, resolving each model’s binding via RequestBindings.resolve (lines 24–36). This prevents the scorer from considering models that are technically deployed but not wired to the current request path.

Model Exclusion and Feature-Based Black-Lists

After provider filtering, the scorer applies request-level constraints. It processes ExcludedModels (or an AllowedModels allow-list) to prune the set, raising an error if the pool empties (lines 54–80). Next, it applies soft feature filters using catalog sets: ToolUseLowSet, AgenticLowSet, and ImageUnsupportedSet. These filters drop models that cannot satisfy tool-use, agentic behavior, or image-input requirements. Each filter is soft—if applying a filter would empty the pool, the scorer skips it to maintain availability (lines 84–106). Finally, if PreferredModels are specified, the scorer prepares a decaying additive bonus via priorityBonusFor (lines 72–90) to boost preferred candidates during final selection.

Embedding Generation and Cluster Selection

With a filtered candidate set, the scorer converts the text prompt into a vector and identifies relevant model clusters.

Prompt Embedding with Timeout Protection

The truncated prompt is sent to the configured Embedder interface. The scorer enforces cfg.EmbedTimeout (default 1.5 seconds) to prevent embedding latency from degrading the request path (lines 70–92). If the embedder fails or times out, the scorer aborts the routing decision early, surfacing ErrClusterUnavailable or a provider-specific error.

Top-P Nearest Cluster Retrieval

The resulting embedding vector is compared against all cluster centroids in the loaded bundle. The scorer invokes topPNearest (lines 60–90) to select the Top-P most similar clusters based on cosine similarity. Only models residing within these top-P clusters are considered in the final scoring phase, dramatically reducing the search space from the global registry to a contextually relevant subset.

Multi-Dimensional Score Calculation

For v2 bundles, the scorer computes a composite score across three axes: quality, cost, and speed.

Routing Knobs and Configuration Overrides

Request-level overrides such as QualityBias, Alpha, and SpeedWeight are validated and merged with the bundle’s default knobs (lines 99–170). The scorer calibrates these dials so that quality-bias adjustments produce uniform-alpha breakpoints, allowing operators to tune the latency-vs-quality trade-off without retraining the cluster model.

The blendScoresV2 Algorithm

Inside blendScoresV2 (lines 106–150), the scorer iterates over every model in every top-P cluster and calculates:

  • Quality: Normalized per-cluster rankings derived from the bundle’s training data.
  • Cost: Derived from the model-axis price table (input/output token costs).
  • Speed: Calculated as TTFT (Time To First Token) plus TPS (Tokens Per Second).

These axes are blended using the active knob weights (Alpha, SpeedWeight, OutputCostRatio). The scorer then applies optional additive terms, including subscription subsidies and the previously computed per-installation preference bonus. The result is a scalar score for every eligible model.

Model Selection and Decision Construction

The final phase selects the winner and packages the metadata.

Argmax Selection and Runner-Up Tracking

The scorer executes an argmax operation (lines 93–110) to identify the model with the highest blended score. It also records the runner-up model to support later exploration or fallback logic. If all models score equally or no candidates remain after filtering, the scorer returns ErrNoEligibleProvider.

Decision Metadata and Logging

The selected model is packaged into a router.Decision struct (lines 119–150). This structure includes the chosen model ID, provider, a human-readable reason string, and rich metadata: the embedding vector, selected cluster IDs, and the full candidate score breakdown. The decision is logged for observability before being returned to the proxy layer.

Implementation Example

The following Go example demonstrates direct usage of the scorer to obtain a routing decision:

import (
    "context"
    "workweave/router/internal/router"
    "workweave/router/internal/router/cluster"
    "workweave/router/internal/providers"
)

// Initialize with a loaded bundle and configured embedder.
cfg := cluster.DefaultConfig()
scorer, _ := cluster.NewScorer(bundle, cfg, embedder, map[string]struct{}{
    providers.ProviderAnthropic: {}, providers.ProviderOpenAI: {},
})

// Construct a request (mirrors HTTP request shape).
req := router.Request{
    PromptText:       "Explain quantum tunnelling in one paragraph.",
    RequestedModel:   "",                // empty = let the scorer pick
    EnabledProviders: nil,               // use all wired providers
    ExcludedModels:   nil,
    HasTools:         false,
    HasImages:        false,
}

// Execute routing.
decision, err := scorer.Route(context.Background(), req)
if err != nil {
    // handle ErrClusterUnavailable, ErrNoEligibleProvider, etc.
}
fmt.Printf("Chosen model: %s (provider %s)\n", decision.Model, decision.Provider)

You can inspect the decision metadata to debug why a specific model was selected:

meta := decision.Metadata
fmt.Printf("Embedding vector length: %d\n", len(meta.Embedding))
fmt.Printf("Top-P cluster IDs: %v\n", meta.ClusterIDs)
fmt.Printf("Score breakdown: %+v\n", meta.CandidateScores)

Summary

  • The cluster.Scorer in internal/router/cluster/scorer.go implements an 11-step pipeline for routing decisions.
  • Input sanitization enforces MaxPromptChars (1024) and embeds with a 1.5s timeout.
  • Candidate filtering uses provider bindings, exclusion lists, and soft feature black-lists (ToolUseLowSet, ImageUnsupportedSet).
  • Cluster selection retrieves the Top-P nearest clusters to the prompt embedding.
  • Scoring blends quality, cost, and speed using blendScoresV2 with request-level knob overrides.
  • Selection uses argmax on the blended scores and returns a router.Decision with full metadata.

Frequently Asked Questions

What happens if the prompt embedding times out?

If the embedder does not return a vector within cfg.EmbedTimeout (default 1.5 seconds), the scorer aborts the routing decision and returns an error. According to the source code in scorer.go (lines 70–92), this protects the request path from hanging on slow or unavailable embedding providers, allowing the caller to fall back to a default model or retry.

How does the scorer handle requests that require tool-use or image inputs?

The scorer checks the request’s HasTools and HasImages flags against catalog black-lists (ToolUseLowSet, AgenticLowSet, ImageUnsupportedSet) in scorer.go (lines 84–106). Models that cannot support these features are filtered out. However, these filters are soft: if applying them would empty the candidate pool, the scorer skips the filter to maintain availability, ensuring the request can still be served by a suboptimal model rather than failing entirely.

Can I bias the routing decision toward cheaper or faster models?

Yes. For v2 bundles, the scorer accepts request-level routing knobs such as QualityBias, SpeedWeight, and Alpha. These are validated and merged with bundle defaults (lines 99–170). Increasing SpeedWeight or decreasing Alpha shifts the blendScoresV2 calculation to favor models with lower TTFT and higher TPS, while adjusting QualityBias recalibrates the breakpoints between quality tiers.

What information is included in the routing decision metadata?

The router.Decision struct includes the selected model ID, provider, and a human-readable reason string. Its Metadata field contains the full embedding vector, the list of Top-P cluster IDs considered, and a detailed score breakdown for every candidate (CandidateScores). This data is logged according to scorer.go (lines 119–150) and can be used for debugging, billing reconciliation, or offline analysis of routing quality.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →