How the LLM Interceptor System Works for Logging and Observability in memU
The LLM interceptor system in memU provides a hook-based framework that intercepts every LLM call through LLMClientWrapper, enabling structured logging and observability by executing registered callbacks before, after, or on-error of each request without modifying core client logic.
The NevaMind-AI/memU repository implements a lightweight interceptor framework designed specifically for LLM observability. This system allows developers to inject custom logging, metrics collection, and monitoring into any LLM workflow while keeping the core client code clean and testable. By leveraging the LLMInterceptorRegistry and snapshot-based execution model, the library ensures that observability code never interferes with the primary LLM request flow.
Core Architecture of the LLM Interceptor System
The architecture centers around three primary components that work together to intercept and observe LLM calls throughout their lifecycle.
LLMInterceptorRegistry and Registration
The LLMInterceptorRegistry class in src/memu/llm/wrapper.py serves as the central hub for managing interceptors. It maintains three ordered lists:
_before: Interceptors that run prior to the LLM request_after: Interceptors that run after successful completion_on_error: Interceptors that execute when exceptions occur
Registration methods register_before(), register_after(), and register_on_error() return an LLMInterceptorHandle object. This handle provides a dispose() method for dynamic removal of interceptors during runtime. Each interceptor is stored internally as an _LLMInterceptor dataclass containing an ID, callable, optional name, priority value, registration order, and filter predicate.
Execution Hooks: Before, After, and On-Error
Interceptors hook into the execution pipeline at three distinct phases defined in LLMClientWrapper._invoke():
- Before: Receives
(ctx: LLMCallContext, request_view: LLMRequestView) - After: Receives
(ctx, request_view, response_view: LLMResponseView, usage: LLMUsage) - On-error: Receives
(ctx, request_view, error: Exception, usage: LLMUsage)
These rich context objects provide comprehensive visibility into provider details, model names, token usage, latency metrics, and request/response content without exposing sensitive underlying SDK internals.
Execution Flow and Lifecycle
The execution flow begins when client code calls LLMClientWrapper.chat() or similar methods. Inside _invoke(), the wrapper first builds an LLMCallContext and takes an immutable snapshot of the current registry state via self._registry.snapshot(). This snapshot pattern ensures that interceptors registered during execution do not affect the current call.
The private methods run_before(), run_after(), and run_on_error() iterate over the snapshot's interceptor tuples. Each interceptor undergoes evaluation through _should_run_interceptor(), which checks filter predicates against the current call context and status. Valid interceptors execute via _safe_invoke_interceptor(), which handles both synchronous and asynchronous callables by detecting awaitables and awaiting them when necessary.
Execution order follows specific rules: before interceptors run in ascending priority order, while after and on-error interceptors execute in reverse priority order. This reversal ensures that cleanup or logging interceptors registered later can properly handle resources established by earlier interceptors.
Filtering and Prioritization
The LLM interceptor system supports sophisticated filtering through the where= parameter during registration. Filters can be:
- LLMCallFilter instances specifying allowed providers, models, operations, or step IDs
- Mapping objects interpreted as filter criteria
- Custom callables receiving the context and returning boolean values
Priority values (numeric, lower runs earlier) resolve execution sequence, with registration order (order) serving as a tiebreaker. The _coerce_filter() method in src/memu/llm/wrapper.py normalizes filter inputs, while _should_run_interceptor() evaluates them against the current LLMCallContext and call status ("success" or "error").
Implementation Examples
Basic Logging Interceptor
Register interceptors to emit structured logs for every LLM call:
from memu.llm.wrapper import LLMInterceptorRegistry, LLMClientWrapper
import logging
logger = logging.getLogger("llm_observability")
def log_before(ctx, request):
logger.info(
"LLM request – provider=%s model=%s tokens=%s",
ctx.provider,
ctx.model,
request.input_chars
)
def log_after(ctx, request, response, usage):
logger.info(
"LLM response – status=%s latency_ms=%s",
usage.status,
usage.latency_ms
)
registry = LLMInterceptorRegistry()
registry.register_before(log_before, name="logging_before")
registry.register_after(log_after, name="logging_after")
wrapped_client = LLMClientWrapper(openai_client, registry=registry)
Model-Specific Filtering
Apply interceptors only to specific models using LLMCallFilter:
from memu.llm.wrapper import LLMCallFilter
gpt_filter = LLMCallFilter(models={"gpt-4o-mini"})
def track_metrics(ctx, request):
# Emit Prometheus or OpenTelemetry metrics
metrics_counter.labels(model=ctx.model).inc()
registry.register_before(
track_metrics,
name="metrics_gpt4o",
where=gpt_filter
)
Dynamic Interceptor Management
Remove interceptors programmatically using the returned handle:
handle = registry.register_after(log_after, name="temporary_logger")
# Later in the application lifecycle
handle.dispose() # Safely removes the interceptor from the registry
Strict Mode for Development
Enable strict mode during debugging to propagate interceptor exceptions:
strict_registry = LLMInterceptorRegistry(strict=True)
def validation_interceptor(ctx, request):
assert request.input_chars > 0, "Empty request detected"
strict_registry.register_before(validation_interceptor)
# Exceptions in interceptors now halt the LLM request instead of being logged
Error Handling and Strict Mode
The _safe_invoke_interceptor() method implements defensive execution semantics. By default (strict=False), exceptions within interceptors are caught and logged without disrupting the LLM request flow. This default behavior ensures observability code never compromises application reliability.
When instantiated with strict=True, the registry re-raises interceptor exceptions through _safe_invoke_interceptor(), immediately halting execution. This mode proves valuable during development and testing when interceptor logic must be validated rigorously. The LLMInterceptorRegistry constructor accepts this boolean flag to configure error handling behavior globally for all registered interceptors.
Summary
- The LLM interceptor system in memU uses
LLMInterceptorRegistryto manage ordered lists of before, after, and on-error hooks insrc/memu/llm/wrapper.py. - Snapshots of the registry ensure consistent interceptor sets throughout each LLM call execution, preventing mid-call registration side effects.
- Rich context objects (
LLMCallContext,LLMRequestView,LLMResponseView,LLMUsage) provide comprehensive observability data to interceptors. - Filtering and prioritization allow precise control over which interceptors execute for specific providers, models, or call statuses.
- Safe invocation with configurable strict mode balances reliability (default) against debugging requirements.
Frequently Asked Questions
What data do interceptors receive during execution?
Interceptors receive rich view objects constructed by LLMClientWrapper. Before interceptors receive LLMCallContext (containing provider, model, operation metadata) and LLMRequestView (input content and token estimates). After interceptors additionally receive LLMResponseView (output content) and LLMUsage (latency, token counts, status). On-error interceptors receive the exception object instead of the response view, enabling error classification and alerting.
How does the interceptor system handle failures?
By default, the system operates in non-strict mode where _safe_invoke_interceptor() catches and logs all exceptions without propagating them to the caller. This ensures logging or metrics code never breaks the primary LLM functionality. When LLMInterceptorRegistry is instantiated with strict=True, interceptor exceptions are re-raised immediately, halting the request for debugging purposes.
Can interceptors be dynamically removed after registration?
Yes. The register_before(), register_after(), and register_on_error() methods return an LLMInterceptorHandle object. Calling handle.dispose() invokes the registry's remove() method, immediately unregistering that specific interceptor. This supports dynamic observability configurations where logging or debugging interceptors are toggled based on runtime conditions or feature flags.
What is the difference between priority and order in interceptors?
Priority is a numeric value specified during registration where lower numbers execute earlier. Order is an internal auto-incrementing counter tracking registration sequence. When priorities are equal, order resolves ties. Additionally, while before interceptors execute in ascending priority order, after and on-error interceptors execute in reverse priority order to support proper resource cleanup patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →