How to Add a Custom LLM Provider to SkillSpector's Provider Registry
To add a custom LLM provider to SkillSpector, implement the LLMProvider protocol in a new sub-package under src/skillspector/providers/, wire it into the _select_active_provider() function in __init__.py, and optionally register CLI capabilities in _agent_cli.py if your provider wraps a local binary.
SkillSpector (NVIDIA/SkillSpector) routes all LLM traffic through a centralized provider registry that supports both HTTP-based APIs and local CLI binaries. When you add a custom LLM provider to SkillSpector's provider registry, you enable the framework to resolve models, manage token budgets, and authenticate against your own backend using the same security hardening applied to built-in providers like OpenAI and Anthropic.
Understanding the Provider Registry Architecture
SkillSpector maintains two distinct registry mechanisms:
- Provider selector (
src/skillspector/providers/__init__.py): Maps theSKILLSPECTOR_PROVIDERenvironment variable to a concrete class implementing theLLMProviderprotocol. This controls which backend handles chat completions. - CLI provider registry (
src/skillspector/providers/_agent_cli.py): Contains the_REGISTRYdictionary that defines how to spawn local binaries (e.g.,claude,codex), parse their output, and perform health checks.
Most custom providers only need to interact with the first registry. You only need the CLI registry if your provider executes a local executable rather than making HTTP calls.
Step 1: Create the Provider Package
Create a new sub-package under src/skillspector/providers/. The directory structure should include an optional model_registry.yaml for token-budget metadata:
src/skillspector/providers/mycustom/
│ __init__.py
│ provider.py
│ model_registry.yaml # optional
The __init__.py file should expose your provider class:
# src/skillspector/providers/mycustom/__init__.py
from .provider import MyCustomProvider # noqa: F401
Step 2: Implement the Provider Protocols
Your provider class must satisfy the LLMProvider protocol defined in src/skillspector/providers/base.py. This is accomplished by implementing four optional protocols: ModelMetadataProvider, CredentialsProvider, ChatModelProvider, and LLMProvider.
Below is a complete implementation for an HTTP-based provider that uses a hypothetical REST endpoint:
# src/skillspector/providers/mycustom/provider.py
from __future__ import annotations
import os
from pathlib import Path
from typing import Tuple, Optional
from langchain_core.language_models.chat_models import BaseChatModel
from langchain_community.chat_models import ChatOpenAI
from ..base import (
ModelMetadataProvider,
CredentialsProvider,
ChatModelProvider,
LLMProvider,
)
from ..registry import lookup_context_length, lookup_max_output_tokens
class MyCustomProvider(
ModelMetadataProvider, CredentialsProvider, ChatModelProvider, LLMProvider
):
"""Custom LLM provider talking to https://my.custom.api/v1/chat."""
DEFAULT_MODEL = "mycustom-large"
SLOT_DEFAULTS = {"default": "mycustom-large", "analysis": "mycustom-medium"}
# ------------------------------------------------------------------
# ModelMetadataProvider implementation
# ------------------------------------------------------------------
def get_context_length(self, model: str) -> Optional[int]:
yaml_path = str(Path(__file__).with_name("model_registry.yaml"))
return lookup_context_length(yaml_path, model)
def get_max_output_tokens(self, model: str) -> Optional[int]:
yaml_path = str(Path(__file__).with_name("model_registry.yaml"))
return lookup_max_output_tokens(yaml_path, model)
def resolve_model(self, slot: str = "default") -> str:
env_model = os.getenv("SKILLSPECTOR_MODEL", "").strip()
if env_model:
return env_model
return self.SLOT_DEFAULTS.get(slot, self.DEFAULT_MODEL)
# ------------------------------------------------------------------
# CredentialsProvider implementation
# ------------------------------------------------------------------
def resolve_credentials(self) -> Optional[Tuple[str, Optional[str]]]:
key = os.getenv("MYCUSTOM_API_KEY", "").strip()
if not key:
return None
base = os.getenv("MYCUSTOM_BASE_URL", None)
return key, base
# ------------------------------------------------------------------
# ChatModelProvider implementation
# ------------------------------------------------------------------
def create_chat_model(
self,
model: str,
*,
max_tokens: int,
timeout: float | None = 120,
) -> Optional[BaseChatModel]:
creds = self.resolve_credentials()
if creds is None:
return None
api_key, base_url = creds
return ChatOpenAI(
model=model,
api_key=api_key,
base_url=base_url,
max_tokens=max_tokens,
timeout=timeout,
)
Metadata and Token Budgets
SkillSpector uses model_registry.yaml to enforce token limits. Create this file next to provider.py:
# src/skillspector/providers/mycustom/model_registry.yaml
models:
mycustom-large:
context_length: 16384
max_output_tokens: 4096
mycustom-medium:
context_length: 8192
max_output_tokens: 2048
The helper functions lookup_context_length() and lookup_max_output_tokens() from src/skillspector/providers/registry.py parse this file automatically.
Credential Resolution
The resolve_credentials() method must return a tuple of (api_key, base_url) or None. Returning None signals SkillSpector to fall back to the default OpenAI provider, allowing graceful degradation when your custom environment variables are absent.
Step 3: Register in the Provider Selector
Open src/skillspector/providers/__init__.py and extend _select_active_provider() to recognize your provider:
# src/skillspector/providers/__init__.py (excerpt)
def _select_active_provider() -> LLMProvider:
"""Construct the active provider based on SKILLSPECTOR_PROVIDER."""
name = os.environ.get("SKILLSPECTOR_PROVIDER", "").strip().lower()
# ... existing branches ...
if name == "mycustom":
from .mycustom import MyCustomProvider
return MyCustomProvider()
# ... fallback handling ...
Users can now activate your provider:
export SKILLSPECTOR_PROVIDER=mycustom
export MYCUSTOM_API_KEY=sk-abc123
skill-spector run /path/to/repo
Step 4: Add CLI Support (Optional)
If your provider runs a local binary instead of an HTTP API, you must register it in src/skillspector/providers/_agent_cli.py. This requires three helper functions and a CliSpec entry in the _REGISTRY dictionary.
Add the following above the _REGISTRY definition:
# src/skillspector/providers/_agent_cli.py (excerpt)
import subprocess
from typing import List
def _build_mycustom_cli_argv(
binary: str, model: str, max_output_tokens: int
) -> List[str]:
"""Construct the argument list for the binary."""
return [binary, "--model", model, "--max-tokens", str(max_output_tokens)]
def _parse_mycustom_cli_output(raw: str) -> str:
"""Parse stdout from the binary."""
text = raw.strip()
if not text:
raise AgentCLIError("mycustom_cli returned empty output")
return text
def _mycustom_cli_auth_check(binary: str) -> tuple[bool, Optional[str]]:
"""Verify the binary is executable and functional."""
try:
result = subprocess.run(
[binary, "--version"], capture_output=True, timeout=5
)
return (result.returncode == 0, None if result.returncode == 0 else "binary not functional")
except Exception as exc:
return (False, f"auth check failed: {exc}")
# Add to the existing _REGISTRY dict
_REGISTRY: dict[str, CliSpec] = {
# ... existing entries ...
"mycustom_cli": CliSpec(
binary="mycustom_cli_binary",
build_argv=_build_mycustom_cli_argv,
parse_output=_parse_mycustom_cli_output,
auth_check=_mycustom_cli_auth_check,
),
}
Then create a thin wrapper class inheriting from AgentCLIProviderBase (see existing providers like claude_cli for the pattern).
Verification and Testing
After installation, verify the integration with a manual smoke test:
export SKILLSPECTOR_PROVIDER=mycustom
export MYCUSTOM_API_KEY=test-key
skill-spector analyze /path/to/code
You should see log output confirming selection:
INFO skillspector.providers.__init__: Selecting provider 'mycustom'
INFO skillspector.providers.mycustom.provider: Created ChatOpenAI model for mycustom-large
For CLI providers, expect logs from skillspector.providers._agent_cli showing the executed argument list.
Summary
- Create a sub-package in
src/skillspector/providers/<name>/containing__init__.pyandprovider.py. - Implement the four protocols from
src/skillspector/providers/base.py:ModelMetadataProvider,CredentialsProvider,ChatModelProvider, andLLMProvider. - Return
Nonefromresolve_credentials()when required environment variables are missing to enable fallback behavior. - Extend
_select_active_provider()insrc/skillspector/providers/__init__.pyto mapSKILLSPECTOR_PROVIDERvalues to your class. - Bundle a
model_registry.yamlto expose token limits to SkillSpector's budgeting system. - Register CLI binaries in
src/skillspector/providers/_agent_cli.pyusing aCliSpecentry with build, parse, and auth-check functions.
Frequently Asked Questions
What protocols must my custom provider implement?
Your provider must implement the four protocols defined in src/skillspector/providers/base.py: ModelMetadataProvider for token budgets and model resolution, CredentialsProvider for API key management, ChatModelProvider for instantiating LangChain chat models, and LLMProvider as the aggregate marker interface. These protocols ensure SkillSpector can resolve models, authenticate, and manage context windows uniformly across all backends.
How does SkillSpector handle missing credentials for my custom provider?
When resolve_credentials() returns None, SkillSpector automatically falls back to the default OpenAI provider. This allows the tool to remain functional even when your custom provider's environment variables (e.g., MYCUSTOM_API_KEY) are not configured, matching the graceful degradation behavior implemented in create_chat_model() within src/skillspector/providers/__init__.py.
Can I integrate a local CLI tool instead of an HTTP API?
Yes. For CLI-based providers, inherit from AgentCLIProviderBase and register three helper functions in src/skillspector/providers/_agent_cli.py: _build_<name>_argv to construct the command line, _parse_<name>_output to handle stdout, and <name>_auth_check to verify binary health. Expose these through a CliSpec entry in the _REGISTRY dictionary, which the framework uses to manage process spawning and security hardening.
Where does the provider read model token limits?
Token limits are read from an optional model_registry.yaml file bundled in your provider directory. Implement get_context_length() and get_max_output_tokens() to call the helper functions in src/skillspector/providers/registry.py, which parse the YAML and cache the results. This allows SkillSpector to enforce token budgets during analysis planning.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →