How to Configure Custom Model Providers in LangExtract: A Complete Guide
You configure custom model providers in LangExtract by implementing the BaseLanguageModel interface, registering your provider with the @lx.providers.registry.register decorator, and exposing it via a pyproject.toml entry point under the langextract.providers group.
LangExtract, Google's extensible extraction framework, uses a plugin-based architecture to support diverse model backends. Whether you need to integrate a proprietary API or a local inference server, understanding how to configure custom model providers lets you extend LangExtract without modifying core source code.
Understanding LangExtract's Provider Architecture
LangExtract discovers and loads model backends through two distinct layers that work together to resolve model IDs at runtime.
Provider Discovery Layer
The discovery layer builds a registry mapping provider names to import specifications. Located in langextract/plugins.py, this module exposes available_providers() and get_provider_class() to enumerate built-in, optional, and third-party providers.
Built-in providers like gemini and ollama reside in the _BUILTINS dictionary, while optional providers (e.g., openai) live in _OPTIONAL_BUILTINS and load only when dependencies are present. Third-party packages expose providers via entry points under the langextract.providers group, which _discovered() merges into the registry. The allow_override flag controls whether third-party plugins can replace built-in implementations.
Runtime Registration Layer
The runtime layer in langextract/providers/router.py handles lazy resolution of model IDs to concrete provider classes. This layer maintains a list of _Entry objects that map regex patterns to loader functions, avoiding circular imports by deferring class imports until needed.
Key functions include register() for immediate registration, register_lazy() for deferred loading via import-path strings, resolve() for pattern-based matching, and resolve_provider() for name-based lookup. When resolve() scans entries sorted by priority, it returns the first provider whose regex matches the supplied model ID.
Step-by-Step Guide to Configure Custom Model Providers
Follow this workflow to integrate a custom backend into LangExtract.
Step 1: Scaffold Your Provider Package
Use the provided generator script to create a compliant package structure:
python scripts/create_provider_plugin.py MyProvider --with-schema
This creates langextract_myprovider/ with provider.py, schema.py, and test files, plus a pyproject.toml template.
Step 2: Implement the Provider Class
Inherit from lx.inference.BaseLanguageModel and implement the infer method. Store configuration in __init__:
# langextract_myprovider/provider.py
import os
import langextract as lx
@lx.providers.registry.register(r'^mycorp', priority=10)
class MyProviderLanguageModel(lx.inference.BaseLanguageModel):
"""Custom provider integrating MyCorp's inference API."""
def __init__(self, model_id: str, api_key: str = None, **kwargs):
super().__init__()
self.model_id = model_id
self.api_key = api_key or os.getenv("MYCORP_API_KEY")
def infer(self, batch_prompts, **kwargs):
for prompt in batch_prompts:
# Replace with actual API integration
result = f"Processed: {prompt[:50]}..."
yield [lx.inference.ScoredOutput(score=1.0, output=result)]
Step 3: Register Model ID Patterns
The @lx.providers.registry.register decorator accepts a regex pattern and priority. The router uses this to match model IDs starting with your prefix. Higher priority values take precedence when multiple patterns match.
For deferred loading to avoid heavy imports at startup, use register_lazy:
from langextract import providers
providers.registry.register_lazy(
r'^mycorp',
target="langextract_myprovider.provider:MyProviderLanguageModel",
priority=10,
)
Step 4: Configure the Entry Point
Expose your provider via pyproject.toml so langextract.plugins can discover it:
[project.entry-points."langextract.providers"]
myprovider = "langextract_myprovider.provider:MyProviderLanguageModel"
Install in editable mode to test changes:
pip install -e .
Step 5: Load and Test Your Provider
Explicitly load plugins before use (automatic in production via factory.create_model, but manual loading helps in tests):
import langextract as lx
lx.providers.load_plugins_once()
config = lx.factory.ModelConfig(
model_id="mycorp-xyz-001",
provider="MyProviderLanguageModel",
provider_kwargs={"api_key": "YOUR_API_KEY"},
)
model = lx.factory.create_model(config)
for output in model.infer(["Extract entities from this text"]):
print(output[0].output)
Implementing Schema Support for Structured Output
To enable structured extraction, subclass lx.schema.BaseSchema and expose it via get_schema_class():
# langextract_myprovider/schema.py
import langextract as lx
class MyProviderSchema(lx.schema.BaseSchema):
"""Schema defining structured output format."""
@classmethod
def from_examples(cls, examples_data, attribute_suffix="_attributes"):
return cls({"type": "object", "properties": {}})
def to_provider_config(self):
return {"response_schema": self._schema_dict}
@property
def supports_strict_mode(self) -> bool:
return True
Update the provider to reference this class:
class MyProviderLanguageModel(lx.inference.BaseLanguageModel):
@classmethod
def get_schema_class(cls):
return MyProviderSchema
LangExtract forwards schema configuration through provider_kwargs when instantiating the model.
Resolving Providers at Runtime
The router resolves providers through two mechanisms:
Pattern Matching: resolve(model_id) scans registered entries by priority, returning the first provider whose regex matches the model ID. For example, mycorp-xyz-001 matches the pattern ^mycorp.
Name Matching: resolve_provider(provider_name) searches for exact matches or class name containment, allowing you to specify myprovider or MyProviderLanguageModel interchangeably.
Both functions raise InferenceConfigError when resolution fails, indicating whether you need to install a missing plugin or correct the provider name.
Summary
- LangExtract uses a two-layer plugin system: discovery via
langextract/plugins.pyand runtime resolution vialangextract/providers/router.py. - Register providers using
@lx.providers.registry.registerwith regex patterns and priorities, orregister_lazyfor deferred loading. - Expose providers via
pyproject.tomlentry points under thelangextract.providersgroup. - Implement structured output by subclassing
lx.schema.BaseSchemaand exposing it throughget_schema_class(). - Test your integration by calling
lx.providers.load_plugins_once()and creating models vialx.factory.ModelConfigandlx.factory.create_model.
Frequently Asked Questions
What is the difference between provider discovery and runtime registration?
Provider discovery, handled in langextract/plugins.py, builds a static map of provider names to import specifications using entry points and built-in dictionaries. Runtime registration, managed in langextract/providers/router.py, maintains a dynamic registry of regex patterns to loader functions that resolve specific model IDs to provider classes lazily, avoiding circular imports and heavy dependencies until actually needed.
How do I avoid circular imports when registering my provider?
Use the register_lazy function instead of the @register decorator. This accepts a string import path (e.g., "langextract_myprovider.provider:MyProviderLanguageModel") and defers the actual import until the router first attempts to resolve a matching model ID. This pattern keeps your provider module lightweight and prevents import cycles during LangExtract initialization.
Can I override built-in providers like Gemini or Ollama with my custom implementation?
By default, LangExtract prevents third-party plugins from overriding built-in providers to maintain backwards compatibility. The allow_override parameter in available_providers() controls this behavior. While the default setting protects built-ins, you can technically enable overrides if you modify the discovery logic, though this is generally discouraged to avoid breaking core functionality.
How does LangExtract match a model ID to my custom provider?
The router in langextract/providers/router.py matches model IDs using the regex pattern you supplied during registration. When you call resolve(model_id), the router iterates through registered entries sorted by priority (highest first) and returns the first provider whose pattern regex matches the supplied model ID. For example, if you register with r'^mycorp', any model ID starting with "mycorp" will resolve to your provider.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →