Routing Algorithms in the libsy Crate: Complete Guide to LLM Traffic Orchestration in Switchyard

The libsy crate implements nine distinct routing algorithms—including Passthrough, Random, FallThrough, and LlmTaskClassifier—that implement the Algorithm trait to dispatch requests across LLM backends with support for composition, session affinity, and judge-based capability routing.

The libsy crate serves as the routing engine for NVIDIA's Switchyard, an open-source LLM gateway that manages traffic across multiple model backends. All routing logic in the crate follows a unified trait-based architecture defined in crates/libsy/src/core/algorithm.rs, enabling seamless composition of simple deterministic routers with sophisticated judge-backed classifiers.

Core Routing Primitives

The foundation of Switchyard's routing system rests on three basic algorithms that handle straightforward dispatch scenarios without complex decision trees.

Passthrough: Fixed-Target Routing

The Passthrough algorithm provides deterministic, single-target routing ideal for diagnostics and fixed-model deployments. Defined in crates/libsy/src/algorithms/passthrough.rs, this algorithm always selects the configured ModelId regardless of request content.

use switchyard::libsy::algorithms::Passthrough;
use switchyard::libsy::core::ModelId;

// Route all traffic to a specific model
let router = Passthrough::new(ModelId::from("meta/llama3-70b"));

Noop: Synthetic Test Responses

For unit testing and integration validation, the Noop algorithm (located in crates/libsy/src/algorithms/noop.rs) returns a hard-coded synthetic response without invoking any backend. This eliminates external dependencies during test execution while maintaining the full routing pipeline structure.

use switchyard::libsy::algorithms::Noop;

// Returns synthetic data for testing
let test_router = Noop {};

Random: Weighted Load Distribution

The Random algorithm in crates/libsy/src/algorithms/rand.rs implements stateless weighted selection using WeightedIndex distribution. It supports optional seeding for reproducible routing decisions in test environments.

use switchyard::libsy::algorithms::Random;
use switchyard::libsy::core::ModelId;

let targets = vec![
    ModelId::from("efficient/8b"),
    ModelId::from("capable/70b")
];
let weights = vec![1.0, 3.0];  // 70b model receives 3x traffic

let router = Random::new(targets, Some(weights), Some(42))?; // Seed 42 for reproducibility

Orchestration and Composition

Complex routing scenarios require combining multiple algorithms into cohesive pipelines. Switchyard provides three compositional structures that wrap or chain primitive algorithms.

FallThrough: The Central Pipeline

FallThrough serves as the primary orchestrator in crates/libsy/src/algorithms/fall_through.rs. It executes a sequential pipeline: first running a Classifier to analyze the request, optionally applying Processors (such as affinity checks), then dispatching to the selected target. This design enables A/B testing and health-aware routing by combining any classifier with any processor.

use switchyard::libsy::algorithms::FallThrough;
use switchyard::libsy::algorithms::util::affinity::AffinityRouter;
use std::sync::Arc;

let pipeline = FallThrough::<()>::new(vec![
    ModelId::from("efficient"),
    ModelId::from("capable")
])
.with_name("production_router")
.with_classifier(Arc::new(my_classifier))
.with_processor(Arc::new(AffinityRouter::new()));

Stage: Multi-Stage Routing Chains

The Stage algorithm (defined in crates/libsy/src/algorithms/stage.rs) chains multiple sub-algorithms into sequential stages. Requests flow through each stage in order, allowing complex workflows such as pre-filtering, classification, and post-processing in a single routing configuration.

Subagent: Hierarchical Delegation

Located in crates/libsy/src/algorithms/subagent.rs, the Subagent algorithm enables nested routing graphs by delegating decisions to subordinate Algorithm instances. This hierarchical approach supports multi-tenant deployments where different request types require entirely separate routing topologies.

Intelligent Classification and Gating

Advanced deployments require content-aware routing that inspects prompts to determine optimal backend selection.

LlmTaskClassifier: Judge-Backed Routing

The LlmTaskClassifier in crates/libsy/src/algorithms/llm_class.rs represents the most sophisticated routing strategy, using an LLM "judge" to classify requests. It supports three operational modes:

  • Capability routing: Directs simple queries to efficient models and complex queries to capable models based on confidence thresholds
  • Escalation routing: Attempts cheaper models first, escalating to expensive models only on failure
  • Custom-schema routing: Applies user-defined JSON schema policies for granular control
use switchyard::libsy::algorithms::LlmTaskClassifier;
use switchyard::libsy::algorithms::LlmTaskClassifierConfig;

let classifier = LlmTaskClassifier::new(
    LlmTaskClassifierConfig::Capability {
        judge_target: ModelId::from("judge/8b"),
        efficient_target: ModelId::from("efficient/8b"),
        capable_target: ModelId::from("capable/70b"),
        config: TaskClassifierConfig {
            base_threshold: 0.5,
            ..Default::default()
        },
    },
)?;

AdvisorGate: Turn-Based Flow Control

The AdvisorGate algorithm (found in crates/libsy/src/algorithms/advisor_gate.rs) implements a gating mechanism that pauses or resumes routing based on turn-based policies. This proves essential for rate-limiting scenarios and conversational state management where routing decisions depend on dialogue position.

Session Management and Utilities

AffinityRouter: Session Persistence

While not a standalone routing algorithm, the AffinityRouter utility in crates/libsy/src/algorithms/util/affinity.rs functions as a processor that maintains session affinity. Once a backend is selected for a session identifier, subsequent requests with the same ID route to the identical target, reducing classifier invocation overhead and maintaining context consistency.

The utility modules in crates/libsy/src/algorithms/util/ provide shared functionality across all algorithms, including target_selector for backend lookup, llm_judge for evaluation prompts, and robustness for failure handling.

Summary

  • The Algorithm trait in crates/libsy/src/core/algorithm.rs unifies all routing implementations with a consistent async interface.
  • Passthrough, Noop, and Random provide deterministic, test, and probabilistic routing primitives.
  • FallThrough serves as the compositional backbone, sequencing classifiers and processors into executable pipelines.
  • LlmTaskClassifier enables intelligent judge-backed routing for capability-based tier selection and automatic escalation.
  • Stage and Subagent support complex multi-stage and hierarchical routing topologies.
  • AffinityRouter maintains session stickiness to optimize performance across stateful conversations.

Frequently Asked Questions

How do I configure a weighted random distribution across multiple models?

Use the Random algorithm from crates/libsy/src/algorithms/rand.rs and provide a weights vector corresponding to your target models. The algorithm uses WeightedIndex to ensure statistical distribution matches your specified ratios, with optional seeding for deterministic testing environments.

What is the difference between FallThrough and Stage routing?

FallThrough (in crates/libsy/src/algorithms/fall_through.rs) executes a classifier-processor-target pipeline for single-pass decision making, while Stage (in crates/libsy/src/algorithms/stage.rs) chains complete routing algorithms sequentially. Use FallThrough for composing classification logic within one decision, and Stage for multi-hop routing where each stage might fundamentally transform the request.

Can I implement custom routing logic without modifying the crate?

Yes. Implement the Algorithm trait defined in crates/libsy/src/core/algorithm.rs for custom routers, or implement the Classifier trait for custom decision logic that plugs into the existing FallThrough orchestrator. The trait-based design allows external crates to define algorithms that the Switchyard server can load via configuration.

How does session affinity interact with intelligent classifiers?

The AffinityRouter utility operates as a Processor within the FallThrough pipeline. When configured, it intercepts requests after classification but before backend selection, checking for existing session mappings. If a session exists, it bypasses the classifier result and routes to the previous backend, reducing latency and maintaining consistency for conversational workloads.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →