# Understanding the AvengersPro Clustering Approach in WorkWeave Router: A Deep Dive into Content-Aware Model Routing

> Discover WorkWeave Router's AvengersPro clustering approach. This content-aware engine routes requests to optimal AI models efficiently using vector embeddings and cluster centroids. Learn more.

- Repository: [Weave/router](https://github.com/workweave/router)
- Tags: deep-dive
- Published: 2026-08-30

---

**The AvengersPro clustering approach in WorkWeave Router is a content-aware routing engine that embeds incoming requests, compares them against pre-computed cluster centroids, and selects the optimal AI model using a top-P argmax strategy, with all core logic implemented in the `internal/router/cluster` package.**

WorkWeave Router uses AvengersPro as its primary routing engine to intelligently distribute inference requests across multiple models. This deterministic clustering system implements a content-aware design based on established research patterns, where every request is embedded and scored against pre-trained centroids to ensure optimal model selection. The implementation prioritizes safety and auditability by surfacing all failures as explicit HTTP errors rather than silent fallbacks.

## Core Architecture of AvengersPro Clustering

The AvengersPro clustering system is organized around several immutable artifacts and strict validation rules. According to the `workweave/router` source code, the architecture separates concerns between embedding generation, centroid storage, and ranking aggregation.

### The Bundle System

A **Bundle** represents a frozen artifact version containing all necessary routing data: centroids, rankings, a model registry, and optional metadata. Each bundle resides under `artifacts/v<X.Y>/` and is referenced by the symbolic `artifacts/latest` pointer in the repository.

In [`internal/router/cluster/artifacts.go`](https://github.com/workweave/router/blob/main/internal/router/cluster/artifacts.go), bundles are loaded as versioned artifacts that include:
- A `centroids.bin` file containing the K × dim matrix
- Ranking tables mapping clusters to model scores
- A [`metadata.yaml`](https://github.com/workweave/router/blob/main/metadata.yaml) file with embedder specifications

By default, only the bundle at `artifacts/latest` is loaded. However, setting `ROUTER_CLUSTER_BUILD_ALL_VERSIONS=true` enables loading of all committed bundles, allowing per-request version pinning via the `x-weave-cluster-version` header.

### The Embedder Interface

The **Embedder** is a pluggable ONNX-backed embedding model (such as Jina v2 or Qwen3 Embedding-0.6B) defined in [`internal/router/cluster/embedder.go`](https://github.com/workweave/router/blob/main/internal/router/cluster/embedder.go). The embedder ID and output dimension must exactly match the bundle's training embedder, or the system refuses to initialize.

This strict coupling ensures that request embeddings exist in the same vector space as the pre-computed centroids. The embedder abstraction allows for both production ONNX sessions and stub implementations (built with `-tags no_onnx`) for testing environments.

### Centroids and Similarity Scoring

**Centroids** are stored as a `K × dim` matrix in [`internal/router/cluster/centroids.go`](https://github.com/workweave/router/blob/main/internal/router/cluster/centroids.go). When a request arrives, the router projects the embedding vector onto these centroids to obtain a similarity score vector of length `K`.

The scoring implementation uses efficient vector operations to compute cosine similarity or dot product similarities (depending on the bundle version) between the request embedding and each cluster center. This projection step determines which content clusters the request most closely resembles.

### Rankings and Quality Means

For each cluster `k ∈ [0, K)`, the system maintains a mapping from model name to performance scores. In [`internal/router/cluster/rankings.go`](https://github.com/workweave/router/blob/main/internal/router/cluster/rankings.go), these tables determine the probability that a specific model will successfully serve requests within that cluster.

Version 2 bundles introduce a separate `quality_means` table for α-blend routing, allowing more nuanced probability calculations based on historical performance metrics. The rankings tables are immutable once loaded, ensuring deterministic routing decisions for identical request embeddings.

## Request Flow and Scoring Mechanism

The request lifecycle follows a strict pipeline from embedding generation to final model selection, implemented primarily in [`internal/router/cluster/scorer.go`](https://github.com/workweave/router/blob/main/internal/router/cluster/scorer.go).

### Step-by-Step Request Processing

When `router.Service` receives a request, it delegates to the cluster scorer following this sequence:

1. **Embedding Generation**: The payload is handed to `cluster.NewEmbedder`, which runs the ONNX model and returns a dense `[]float32` vector. If embedding fails or times out, the system immediately returns `ErrClusterUnavailable` mapped to HTTP 503.

2. **Centroid Projection**: `Scorer.Score` projects the vector onto the centroids, yielding a similarity score for each of the `K` clusters.

3. **Top-P Selection**: The **Top-P** runtime knob (configured via `Config.TopP`, defaulting to 4) selects the highest-scoring `P` clusters. This bounds the search space and caps latency by limiting the number of ranking tables consulted.

4. **Model Aggregation**: For each selected cluster, the router consults the pre-computed ranking table to produce model-level scores. The system aggregates these scores to find the best-performing model across the selected clusters.

5. **Argmax Selection**: The router picks the model with the highest aggregated score. If no model is eligible, it returns `ErrNoEligibleProvider` resulting in a 4xx error.

### The Scorer Implementation

The `Scorer` struct (lines 49-73 of [`scorer.go`](https://github.com/workweave/router/blob/main/scorer.go)) encapsulates all routing logic:

```go
func (s *Scorer) Score(ctx context.Context, req router.Request) (router.Decision, error) {
    // 1️⃣ Embed the request
    vec, err := s.embed.Embed(ctx, req.Prompt)
    if err != nil {
        return router.Decision{}, cluster.ErrClusterUnavailable
    }

    // 2️⃣ Compute similarity to each centroid
    scores := s.centroids.Score(vec) // → []float64 length K

    // 3️⃣ Keep the top‑P clusters
    topIdx := topPIndices(scores, s.cfg.TopP)

    // 4️⃣ Aggregate model scores from the selected clusters
    bestModel, bestScore := "", -math.MaxFloat64
    for _, k := range topIdx {
        for model, w := range s.rankings[k] {
            if w > bestScore {
                bestScore = w
                bestModel = model
            }
        }
    }

    // 5️⃣ Return the routing decision
    return router.Decision{
        Model: bestModel,
        Score: bestScore,
    }, nil
}

```

### Validation and Safety Checks

During initialization in `cluster.NewScorer`, the system performs strict validation to prevent runtime mismatches:

```go
if embed.ID() != bundle.EmbedderID() {
    return nil, fmt.Errorf("cluster %s: bundle declares embedder %q but runtime embedder is %q",
        bundle.Version, bundle.EmbedderID(), embed.ID())
}
if embed.Dim() != bundle.Centroids.Dim {
    return nil, fmt.Errorf("cluster %s: embedder dim %d != centroids dim %d",
        bundle.Version, embed.Dim(), bundle.Centroids.Dim)
}
if bundle.Centroids.K < cfg.TopP {
    return nil, fmt.Errorf("cluster %s: K=%d < TopP=%d", bundle.Version, bundle.Centroids.K, cfg.TopP)
}

```

These checks guarantee **load-bearing invariants**: a bundle trained with a specific embedder cannot be scored with a mismatched embedder, and clustering hyper-parameters are frozen at bundle creation time.

## Configuration and Runtime Behavior

The AvengersPro clustering approach exposes several runtime controls that balance accuracy against latency and resource constraints.

### Top-P Selection and Latency Bounds

The **Top-P** parameter (defined in [`scorer.go`](https://github.com/workweave/router/blob/main/scorer.go) lines 33-36) limits the argmax operation to the best `P` clusters. By default set to 4, this parameter provides a trade-off between routing accuracy and computational latency. Lower values reduce the ranking lookup time but may miss optimal models in edge-case clusters.

Concurrently, `Config.EmbedTimeout` (default 1500ms) defines the maximum time allowed for the embedding operation. If the embedder does not return within this window, the scorer aborts with `ErrClusterUnavailable`, triggering an HTTP 503 response. This aggressive timeout prevents cascade failures when the embedding service degrades.

### Versioning and Multi-Bundle Support

Unlike single-version systems, AvengersPro supports side-by-side bundle deployment. The composition root in [`cmd/router/main.go`](https://github.com/workweave/router/blob/main/cmd/router/main.go) wires the scorer:

```go
func buildClusterScorer() (router.Router, error) {
    // Load the latest artifact bundle (artifacts/latest → v0.xx)
    bundle, err := cluster.LoadBundleFromArtifacts()
    if err != nil { return nil, err }

    // Construct the embedder (ONNX or stub)
    embedder, err := cluster.NewEmbedder()
    if err != nil { return nil, err }

    // Use production defaults (TopP=4, MaxPromptChars=1024, EmbedTimeout=1500ms)
    cfg := cluster.DefaultConfig()

    // Wire the scorer
    return cluster.NewScorer(bundle, cfg, embedder, allProviders())
}

```

To enable multi-version evaluation (useful for smoke testing or gradual rollouts), set the environment variable:

```bash
export ROUTER_CLUSTER_BUILD_ALL_VERSIONS=true
go build -tags ORT -o router ./cmd/router

```

This allows per-request version selection via headers:

```bash
curl -H "x-weave-cluster-version: v0.35" http://localhost:8080/v1/chat/completions

```

## Architectural Position and Dependencies

The `internal/router/cluster` package implements the `router.Router` interface as an inner-ring component. It maintains zero dependencies on presentation or adapter packages, importing only pure-Go utilities from `observability`, `router/catalog`, and `router/policy`.

All I/O operations—including ONNX session creation and artifact file reads—are isolated in the adapter layer (`internal/router/cluster/embedder_*.go`). This hexagonal architecture ensures the core scoring logic remains testable and independent of infrastructure concerns.

## Summary

- **AvengersPro clustering** in WorkWeave Router implements content-aware routing through embedding-based centroid comparison and top-P argmax selection.
- The system relies on immutable **Bundles** containing centroids, rankings, and metadata, loaded from [`internal/router/cluster/artifacts.go`](https://github.com/workweave/router/blob/main/internal/router/cluster/artifacts.go).
- Strict validation in `NewScorer` ensures embedder compatibility and dimension matching, preventing runtime vector space mismatches.
- **Top-P** selection (default 4) bounds latency by limiting cluster evaluation, while **EmbedTimeout** (1500ms) prevents hanging requests.
- Unlike heuristic routers, AvengersPro never falls back to default models; failures surface as HTTP 503 errors for immediate visibility.
- Multi-version support via `ROUTER_CLUSTER_BUILD_ALL_VERSIONS` enables A/B testing and gradual rollouts using the `x-weave-cluster-version` header.

## Frequently Asked Questions

### What makes AvengersPro different from heuristic routers?

AvengersPro replaces heuristic rules with deterministic, data-driven clustering. While heuristic routers might use regex patterns or keyword matching, AvengersPro embeds requests into a learned vector space and selects models based on pre-computed performance centroids. Crucially, it never falls back to default models on failure—any embedding timeout, dimension mismatch, or missing rankings results in an immediate HTTP 503 error, making production regressions visible rather than hidden.

### How does the embedder validation prevent runtime errors?

During scorer initialization in [`internal/router/cluster/scorer.go`](https://github.com/workweave/router/blob/main/internal/router/cluster/scorer.go), the system verifies that the runtime embedder's ID and output dimension exactly match the bundle's training specifications. If `embed.ID()` differs from `bundle.EmbedderID()` or `embed.Dim()` differs from `bundle.Centroids.Dim`, construction fails with a descriptive error. This prevents silent degradation that would occur if requests were embedded in a different vector space than the centroids.

### What happens when the embedding timeout is exceeded?

If the embedder does not return within `Config.EmbedTimeout` (default 1500ms), the `Score` method returns `ErrClusterUnavailable`, which the HTTP handler translates to an HTTP 503 Service Unavailable response. This hard timeout prevents resource exhaustion when the ONNX runtime or embedding service experiences degradation, allowing load balancers to route requests to healthy instances.

### Can I use multiple cluster versions simultaneously?

Yes. By setting `ROUTER_CLUSTER_BUILD_ALL_VERSIONS=true`, the router loads all committed bundles from the `artifacts/` directory rather than only `artifacts/latest`. You can then pin specific requests to particular bundle versions using the `x-weave-cluster-version` header. This enables blue-green deployments, canary testing, and backward compatibility verification without separate deployments.