How Magnitude Ensures Model ID Consistency Across Provider Formats Using Resilient Symbols

Magnitude normalizes every provider-specific model identifier into a canonical atom stream using the atomizeModelId function and a resilient symbol-based DSL, enabling consistent classification regardless of format variations.

Magnitude solves the model ID consistency problem by decoupling raw provider strings from downstream classification logic. When you work with AI providers like Bedrock, Fireworks, or local GGUF files, the same logical model appears in wildly different formats. This article explains how Magnitude's classifier system transforms these messy identifiers into a unified representation using resilient symbols that tolerate variation in separators, capitalization, and versioning schemes.

The Atomization Pipeline

At the heart of Magnitude's approach is the model-ID atomizer implemented in packages/providers/src/classifier/atomizer.ts. The atomizeModelId function performs three deterministic operations on every raw identifier:

  1. Strips provider noise — removes path prefixes, file extensions, version suffixes, and artifact markers
  2. Normalizes case — converts the entire string to lower-case
  3. Splits into typed atoms — segments the cleaned string into a sequence of labeled tokens

This atom stream becomes the universal intermediate representation that all matching logic consumes.

// Example: normalizing a Bedrock-style ID
import { atomizeModelId } from "@magnitudedev/providers/src/classifier/atomizer";

const raw = "meta.llama3-1-8b-instruct-v1:0";
const atoms = atomizeModelId(raw);
// → [{type:'lit',value:'meta'}, {type:'sep',value:'.'}, {type:'lit',value:'llama3'},
//    {type:'sep',value:'-'}, {type:'lit',value:'1'}, {type:'sep',value:'-'},
//    {type:'lit',value:'8b'}, {type:'sep',value:'-'}, {type:'lit',value:'instruct'},
//    {type:'dot',value:'-'}, {type:'lit',value:'v1'}, {type:'sep',value:':'},
//    {type:'lit',value:'0'}]

Resilient Symbol Types in the DSL

The symbol-based DSL defined in packages/providers/src/classifier/symbols.ts provides six atom types designed to absorb real-world variation:

  • lit — exact literal text (case-insensitive). Normalizes to lower-case so "Llama-3" and "llama_3" match identically.
  • sep — separator flexibility. Matches -, _, :, or their absence, handling "model-v1", "model_v1", or "modelv1" uniformly.
  • dot — decimal point equivalence. Treats both "3.5" and "3p5" as the same version component.
  • num — any all-digit atom. Matches version numbers without constraining length.
  • ver — full version pattern. Accepts digit[.p]digit, single digits, or the dot/p-equivalent forms.
  • opt — optional literal text. Enables optional capability markers like "q4" in "q4-gguf".

These symbols make the pattern language resilient by design rather than requiring exhaustive enumeration of format permutations.

Pattern Matching with Extra Atom Tolerance

The matching engine in packages/providers/src/classifier/matcher.ts walks the atom stream left-to-right against registered family patterns. Critically, it allows extra atoms to be ignored — size tokens, dates, capability markers, and provider-specific noise that survive atomization do not prevent a match.

// Example: matching the atom stream against a family pattern
import { match } from "@magnitudedev/providers/src/classifier/matcher";
import { lit, sep, dot, num, ver, opt } from "@magnitudedev/providers/src/classifier/symbols";

const family = {
  familyId: "llama",
  patterns: [
    {
      pattern: [lit("llama"), sep(), ver()],
      priority: 10,
    },
  ],
};

const result = match(atoms, [family]);
// → { familyId: "llama", priority: 10 }

This tolerance means the same logical model ID produces identical classification results whether sourced from:

  • meta.llama3-1-8b-instruct-v1:0 (Bedrock format)
  • accounts/fireworks/models/llama-v3p1-8b-instruct (custom endpoint)
  • glm-5.2.gguf (local file)

Typed Model IDs at the System Boundary

While atomization handles internal consistency, Magnitude also provides a typed wrapper for external interfaces. The ProviderModelIdSchema in packages/ai/src/provider/model.ts creates validated model identifiers that flow through the system:

// Using the ProviderModelIdSchema to create a typed ID (later atomized)
import { ProviderModelIdSchema } from "@magnitudedev/ai";

const providerId = ProviderModelIdSchema.make("meta-llama/Llama-3.3-70B-Instruct");

This schema establishes the contract between provider-specific inputs and the internal atomization pipeline.

Key Source Files

File Role
packages/providers/src/classifier/atomizer.ts Strips provider noise, lower-cases, and splits into typed atoms
packages/providers/src/classifier/symbols.ts Defines the resilient DSL symbols (lit, sep, dot, num, ver, opt)
packages/providers/src/classifier/matcher.ts Executes pattern matching on the atom stream, ignoring irrelevant atoms
packages/ai/src/provider/model.ts Declares ProviderModelIdSchema, the typed wrapper used throughout the system

Summary

Frequently Asked Questions

What problem does Magnitude's model ID atomization solve?

Provider-specific formats encode the same logical model in incompatible ways. Bedrock uses dot-delimited paths with version suffixes, Fireworks uses hyphenated slugs, and local files include extensions. Without normalization, classification logic would need provider-specific branches. Atomization collapses these variations into a single format that downstream code can process uniformly.

How do resilient symbols handle version number differences?

The ver symbol accepts digit[.p]digit patterns, single digits, and either . or p as decimal separators. This means "3.5", "3p5", "3-5", and "35" (when context makes the digits unambiguous) all match the same version pattern. The dot symbol specifically bridges the gap between conventional decimal notation and the "p" notation used in some model naming conventions.

Why does the matcher ignore extra atoms instead of requiring exact matches?

Real-world model IDs contain variable metadata: quantization levels (q4, q8), context sizes (32k, 128k), dates, and provider-specific tags. These are meaningful for some purposes but irrelevant for family classification. By ignoring unmatched atoms after satisfying the pattern, Magnitude avoids maintaining exhaustive exclusion lists while remaining precise about core identity.

Where does the atomization happen in Magnitude's request flow?

Atomization occurs before any classification or catalog lookup. When a model ID enters the system—whether through ProviderModelIdSchema.make() or direct provider configuration—the atomizeModelId function transforms it immediately. This ensures catalogs, resolvers, and UI components always operate on the normalized representation without duplicating normalization logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →