Understanding the AvengersPro Clustering Approach in WorkWeave Router: A Deep Dive into Content-Aware Model Routing

The AvengersPro clustering approach in WorkWeave Router is a content-aware routing engine that embeds incoming requests, compares them against pre-computed cluster centroids, and selects the optimal AI model using a top-P argmax strategy, with all core logic implemented in the internal/router/cluster package.

WorkWeave Router uses AvengersPro as its primary routing engine to intelligently distribute inference requests across multiple models. This deterministic clustering system implements a content-aware design based on established research patterns, where every request is embedded and scored against pre-trained centroids to ensure optimal model selection. The implementation prioritizes safety and auditability by surfacing all failures as explicit HTTP errors rather than silent fallbacks.

Core Architecture of AvengersPro Clustering

The AvengersPro clustering system is organized around several immutable artifacts and strict validation rules. According to the workweave/router source code, the architecture separates concerns between embedding generation, centroid storage, and ranking aggregation.

The Bundle System

A Bundle represents a frozen artifact version containing all necessary routing data: centroids, rankings, a model registry, and optional metadata. Each bundle resides under artifacts/v<X.Y>/ and is referenced by the symbolic artifacts/latest pointer in the repository.

In internal/router/cluster/artifacts.go, bundles are loaded as versioned artifacts that include:

  • A centroids.bin file containing the K × dim matrix
  • Ranking tables mapping clusters to model scores
  • A metadata.yaml file with embedder specifications

By default, only the bundle at artifacts/latest is loaded. However, setting ROUTER_CLUSTER_BUILD_ALL_VERSIONS=true enables loading of all committed bundles, allowing per-request version pinning via the x-weave-cluster-version header.

The Embedder Interface

The Embedder is a pluggable ONNX-backed embedding model (such as Jina v2 or Qwen3 Embedding-0.6B) defined in internal/router/cluster/embedder.go. The embedder ID and output dimension must exactly match the bundle's training embedder, or the system refuses to initialize.

This strict coupling ensures that request embeddings exist in the same vector space as the pre-computed centroids. The embedder abstraction allows for both production ONNX sessions and stub implementations (built with -tags no_onnx) for testing environments.

Centroids and Similarity Scoring

Centroids are stored as a K × dim matrix in internal/router/cluster/centroids.go. When a request arrives, the router projects the embedding vector onto these centroids to obtain a similarity score vector of length K.

The scoring implementation uses efficient vector operations to compute cosine similarity or dot product similarities (depending on the bundle version) between the request embedding and each cluster center. This projection step determines which content clusters the request most closely resembles.

Rankings and Quality Means

For each cluster k ∈ [0, K), the system maintains a mapping from model name to performance scores. In internal/router/cluster/rankings.go, these tables determine the probability that a specific model will successfully serve requests within that cluster.

Version 2 bundles introduce a separate quality_means table for α-blend routing, allowing more nuanced probability calculations based on historical performance metrics. The rankings tables are immutable once loaded, ensuring deterministic routing decisions for identical request embeddings.

Request Flow and Scoring Mechanism

The request lifecycle follows a strict pipeline from embedding generation to final model selection, implemented primarily in internal/router/cluster/scorer.go.

Step-by-Step Request Processing

When router.Service receives a request, it delegates to the cluster scorer following this sequence:

  1. Embedding Generation: The payload is handed to cluster.NewEmbedder, which runs the ONNX model and returns a dense []float32 vector. If embedding fails or times out, the system immediately returns ErrClusterUnavailable mapped to HTTP 503.

  2. Centroid Projection: Scorer.Score projects the vector onto the centroids, yielding a similarity score for each of the K clusters.

  3. Top-P Selection: The Top-P runtime knob (configured via Config.TopP, defaulting to 4) selects the highest-scoring P clusters. This bounds the search space and caps latency by limiting the number of ranking tables consulted.

  4. Model Aggregation: For each selected cluster, the router consults the pre-computed ranking table to produce model-level scores. The system aggregates these scores to find the best-performing model across the selected clusters.

  5. Argmax Selection: The router picks the model with the highest aggregated score. If no model is eligible, it returns ErrNoEligibleProvider resulting in a 4xx error.

The Scorer Implementation

The Scorer struct (lines 49-73 of scorer.go) encapsulates all routing logic:

func (s *Scorer) Score(ctx context.Context, req router.Request) (router.Decision, error) {
    // 1️⃣ Embed the request
    vec, err := s.embed.Embed(ctx, req.Prompt)
    if err != nil {
        return router.Decision{}, cluster.ErrClusterUnavailable
    }

    // 2️⃣ Compute similarity to each centroid
    scores := s.centroids.Score(vec) // → []float64 length K

    // 3️⃣ Keep the top‑P clusters
    topIdx := topPIndices(scores, s.cfg.TopP)

    // 4️⃣ Aggregate model scores from the selected clusters
    bestModel, bestScore := "", -math.MaxFloat64
    for _, k := range topIdx {
        for model, w := range s.rankings[k] {
            if w > bestScore {
                bestScore = w
                bestModel = model
            }
        }
    }

    // 5️⃣ Return the routing decision
    return router.Decision{
        Model: bestModel,
        Score: bestScore,
    }, nil
}

Validation and Safety Checks

During initialization in cluster.NewScorer, the system performs strict validation to prevent runtime mismatches:

if embed.ID() != bundle.EmbedderID() {
    return nil, fmt.Errorf("cluster %s: bundle declares embedder %q but runtime embedder is %q",
        bundle.Version, bundle.EmbedderID(), embed.ID())
}
if embed.Dim() != bundle.Centroids.Dim {
    return nil, fmt.Errorf("cluster %s: embedder dim %d != centroids dim %d",
        bundle.Version, embed.Dim(), bundle.Centroids.Dim)
}
if bundle.Centroids.K < cfg.TopP {
    return nil, fmt.Errorf("cluster %s: K=%d < TopP=%d", bundle.Version, bundle.Centroids.K, cfg.TopP)
}

These checks guarantee load-bearing invariants: a bundle trained with a specific embedder cannot be scored with a mismatched embedder, and clustering hyper-parameters are frozen at bundle creation time.

Configuration and Runtime Behavior

The AvengersPro clustering approach exposes several runtime controls that balance accuracy against latency and resource constraints.

Top-P Selection and Latency Bounds

The Top-P parameter (defined in scorer.go lines 33-36) limits the argmax operation to the best P clusters. By default set to 4, this parameter provides a trade-off between routing accuracy and computational latency. Lower values reduce the ranking lookup time but may miss optimal models in edge-case clusters.

Concurrently, Config.EmbedTimeout (default 1500ms) defines the maximum time allowed for the embedding operation. If the embedder does not return within this window, the scorer aborts with ErrClusterUnavailable, triggering an HTTP 503 response. This aggressive timeout prevents cascade failures when the embedding service degrades.

Versioning and Multi-Bundle Support

Unlike single-version systems, AvengersPro supports side-by-side bundle deployment. The composition root in cmd/router/main.go wires the scorer:

func buildClusterScorer() (router.Router, error) {
    // Load the latest artifact bundle (artifacts/latest → v0.xx)
    bundle, err := cluster.LoadBundleFromArtifacts()
    if err != nil { return nil, err }

    // Construct the embedder (ONNX or stub)
    embedder, err := cluster.NewEmbedder()
    if err != nil { return nil, err }

    // Use production defaults (TopP=4, MaxPromptChars=1024, EmbedTimeout=1500ms)
    cfg := cluster.DefaultConfig()

    // Wire the scorer
    return cluster.NewScorer(bundle, cfg, embedder, allProviders())
}

To enable multi-version evaluation (useful for smoke testing or gradual rollouts), set the environment variable:

export ROUTER_CLUSTER_BUILD_ALL_VERSIONS=true
go build -tags ORT -o router ./cmd/router

This allows per-request version selection via headers:

curl -H "x-weave-cluster-version: v0.35" http://localhost:8080/v1/chat/completions

Architectural Position and Dependencies

The internal/router/cluster package implements the router.Router interface as an inner-ring component. It maintains zero dependencies on presentation or adapter packages, importing only pure-Go utilities from observability, router/catalog, and router/policy.

All I/O operations—including ONNX session creation and artifact file reads—are isolated in the adapter layer (internal/router/cluster/embedder_*.go). This hexagonal architecture ensures the core scoring logic remains testable and independent of infrastructure concerns.

Summary

  • AvengersPro clustering in WorkWeave Router implements content-aware routing through embedding-based centroid comparison and top-P argmax selection.
  • The system relies on immutable Bundles containing centroids, rankings, and metadata, loaded from internal/router/cluster/artifacts.go.
  • Strict validation in NewScorer ensures embedder compatibility and dimension matching, preventing runtime vector space mismatches.
  • Top-P selection (default 4) bounds latency by limiting cluster evaluation, while EmbedTimeout (1500ms) prevents hanging requests.
  • Unlike heuristic routers, AvengersPro never falls back to default models; failures surface as HTTP 503 errors for immediate visibility.
  • Multi-version support via ROUTER_CLUSTER_BUILD_ALL_VERSIONS enables A/B testing and gradual rollouts using the x-weave-cluster-version header.

Frequently Asked Questions

What makes AvengersPro different from heuristic routers?

AvengersPro replaces heuristic rules with deterministic, data-driven clustering. While heuristic routers might use regex patterns or keyword matching, AvengersPro embeds requests into a learned vector space and selects models based on pre-computed performance centroids. Crucially, it never falls back to default models on failure—any embedding timeout, dimension mismatch, or missing rankings results in an immediate HTTP 503 error, making production regressions visible rather than hidden.

How does the embedder validation prevent runtime errors?

During scorer initialization in internal/router/cluster/scorer.go, the system verifies that the runtime embedder's ID and output dimension exactly match the bundle's training specifications. If embed.ID() differs from bundle.EmbedderID() or embed.Dim() differs from bundle.Centroids.Dim, construction fails with a descriptive error. This prevents silent degradation that would occur if requests were embedded in a different vector space than the centroids.

What happens when the embedding timeout is exceeded?

If the embedder does not return within Config.EmbedTimeout (default 1500ms), the Score method returns ErrClusterUnavailable, which the HTTP handler translates to an HTTP 503 Service Unavailable response. This hard timeout prevents resource exhaustion when the ONNX runtime or embedding service experiences degradation, allowing load balancers to route requests to healthy instances.

Can I use multiple cluster versions simultaneously?

Yes. By setting ROUTER_CLUSTER_BUILD_ALL_VERSIONS=true, the router loads all committed bundles from the artifacts/ directory rather than only artifacts/latest. You can then pin specific requests to particular bundle versions using the x-weave-cluster-version header. This enables blue-green deployments, canary testing, and backward compatibility verification without separate deployments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →