# How the LLM Interceptor System Works for Logging and Observability in memU

> Discover how the LLM interceptor system in memU enables structured logging and observability. Learn how it hooks into LLM calls with callbacks without altering core client logic.

- Repository: [NevaMind AI/memU](https://github.com/nevamind-ai/memu)
- Tags: deep-dive
- Published: 2026-02-19

---

**The LLM interceptor system in memU provides a hook-based framework that intercepts every LLM call through `LLMClientWrapper`, enabling structured logging and observability by executing registered callbacks before, after, or on-error of each request without modifying core client logic.**

The NevaMind-AI/memU repository implements a lightweight interceptor framework designed specifically for LLM observability. This system allows developers to inject custom logging, metrics collection, and monitoring into any LLM workflow while keeping the core client code clean and testable. By leveraging the `LLMInterceptorRegistry` and snapshot-based execution model, the library ensures that observability code never interferes with the primary LLM request flow.

## Core Architecture of the LLM Interceptor System

The architecture centers around three primary components that work together to intercept and observe LLM calls throughout their lifecycle.

### LLMInterceptorRegistry and Registration

The `LLMInterceptorRegistry` class in [`src/memu/llm/wrapper.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/wrapper.py) serves as the central hub for managing interceptors. It maintains three ordered lists:

- `_before`: Interceptors that run prior to the LLM request
- `_after`: Interceptors that run after successful completion
- `_on_error`: Interceptors that execute when exceptions occur

Registration methods `register_before()`, `register_after()`, and `register_on_error()` return an `LLMInterceptorHandle` object. This handle provides a `dispose()` method for dynamic removal of interceptors during runtime. Each interceptor is stored internally as an `_LLMInterceptor` dataclass containing an ID, callable, optional name, priority value, registration order, and filter predicate.

### Execution Hooks: Before, After, and On-Error

Interceptors hook into the execution pipeline at three distinct phases defined in `LLMClientWrapper._invoke()`:

1. **Before**: Receives `(ctx: LLMCallContext, request_view: LLMRequestView)`
2. **After**: Receives `(ctx, request_view, response_view: LLMResponseView, usage: LLMUsage)`
3. **On-error**: Receives `(ctx, request_view, error: Exception, usage: LLMUsage)`

These rich context objects provide comprehensive visibility into provider details, model names, token usage, latency metrics, and request/response content without exposing sensitive underlying SDK internals.

## Execution Flow and Lifecycle

The execution flow begins when client code calls `LLMClientWrapper.chat()` or similar methods. Inside `_invoke()`, the wrapper first builds an `LLMCallContext` and takes an immutable **snapshot** of the current registry state via `self._registry.snapshot()`. This snapshot pattern ensures that interceptors registered during execution do not affect the current call.

The private methods `run_before()`, `run_after()`, and `run_on_error()` iterate over the snapshot's interceptor tuples. Each interceptor undergoes evaluation through `_should_run_interceptor()`, which checks filter predicates against the current call context and status. Valid interceptors execute via `_safe_invoke_interceptor()`, which handles both synchronous and asynchronous callables by detecting awaitables and awaiting them when necessary.

Execution order follows specific rules: **before** interceptors run in ascending priority order, while **after** and **on-error** interceptors execute in reverse priority order. This reversal ensures that cleanup or logging interceptors registered later can properly handle resources established by earlier interceptors.

## Filtering and Prioritization

The LLM interceptor system supports sophisticated filtering through the `where=` parameter during registration. Filters can be:

- **LLMCallFilter** instances specifying allowed providers, models, operations, or step IDs
- **Mapping objects** interpreted as filter criteria
- **Custom callables** receiving the context and returning boolean values

Priority values (numeric, lower runs earlier) resolve execution sequence, with registration order (`order`) serving as a tiebreaker. The `_coerce_filter()` method in [`src/memu/llm/wrapper.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/wrapper.py) normalizes filter inputs, while `_should_run_interceptor()` evaluates them against the current `LLMCallContext` and call status (`"success"` or `"error"`).

## Implementation Examples

### Basic Logging Interceptor

Register interceptors to emit structured logs for every LLM call:

```python
from memu.llm.wrapper import LLMInterceptorRegistry, LLMClientWrapper
import logging

logger = logging.getLogger("llm_observability")

def log_before(ctx, request):
    logger.info(
        "LLM request – provider=%s model=%s tokens=%s",
        ctx.provider,
        ctx.model,
        request.input_chars
    )

def log_after(ctx, request, response, usage):
    logger.info(
        "LLM response – status=%s latency_ms=%s",
        usage.status,
        usage.latency_ms
    )

registry = LLMInterceptorRegistry()
registry.register_before(log_before, name="logging_before")
registry.register_after(log_after, name="logging_after")

wrapped_client = LLMClientWrapper(openai_client, registry=registry)

```

### Model-Specific Filtering

Apply interceptors only to specific models using `LLMCallFilter`:

```python
from memu.llm.wrapper import LLMCallFilter

gpt_filter = LLMCallFilter(models={"gpt-4o-mini"})

def track_metrics(ctx, request):
    # Emit Prometheus or OpenTelemetry metrics

    metrics_counter.labels(model=ctx.model).inc()

registry.register_before(
    track_metrics, 
    name="metrics_gpt4o", 
    where=gpt_filter
)

```

### Dynamic Interceptor Management

Remove interceptors programmatically using the returned handle:

```python
handle = registry.register_after(log_after, name="temporary_logger")

# Later in the application lifecycle

handle.dispose()  # Safely removes the interceptor from the registry

```

### Strict Mode for Development

Enable strict mode during debugging to propagate interceptor exceptions:

```python
strict_registry = LLMInterceptorRegistry(strict=True)

def validation_interceptor(ctx, request):
    assert request.input_chars > 0, "Empty request detected"

strict_registry.register_before(validation_interceptor)

# Exceptions in interceptors now halt the LLM request instead of being logged

```

## Error Handling and Strict Mode

The `_safe_invoke_interceptor()` method implements defensive execution semantics. By default (`strict=False`), exceptions within interceptors are caught and logged without disrupting the LLM request flow. This default behavior ensures observability code never compromises application reliability.

When instantiated with `strict=True`, the registry re-raises interceptor exceptions through `_safe_invoke_interceptor()`, immediately halting execution. This mode proves valuable during development and testing when interceptor logic must be validated rigorously. The `LLMInterceptorRegistry` constructor accepts this boolean flag to configure error handling behavior globally for all registered interceptors.

## Summary

- The **LLM interceptor system** in memU uses `LLMInterceptorRegistry` to manage ordered lists of before, after, and on-error hooks in [`src/memu/llm/wrapper.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/wrapper.py).
- **Snapshots** of the registry ensure consistent interceptor sets throughout each LLM call execution, preventing mid-call registration side effects.
- **Rich context objects** (`LLMCallContext`, `LLMRequestView`, `LLMResponseView`, `LLMUsage`) provide comprehensive observability data to interceptors.
- **Filtering and prioritization** allow precise control over which interceptors execute for specific providers, models, or call statuses.
- **Safe invocation** with configurable strict mode balances reliability (default) against debugging requirements.

## Frequently Asked Questions

### What data do interceptors receive during execution?

Interceptors receive rich view objects constructed by `LLMClientWrapper`. Before interceptors receive `LLMCallContext` (containing provider, model, operation metadata) and `LLMRequestView` (input content and token estimates). After interceptors additionally receive `LLMResponseView` (output content) and `LLMUsage` (latency, token counts, status). On-error interceptors receive the exception object instead of the response view, enabling error classification and alerting.

### How does the interceptor system handle failures?

By default, the system operates in non-strict mode where `_safe_invoke_interceptor()` catches and logs all exceptions without propagating them to the caller. This ensures logging or metrics code never breaks the primary LLM functionality. When `LLMInterceptorRegistry` is instantiated with `strict=True`, interceptor exceptions are re-raised immediately, halting the request for debugging purposes.

### Can interceptors be dynamically removed after registration?

Yes. The `register_before()`, `register_after()`, and `register_on_error()` methods return an `LLMInterceptorHandle` object. Calling `handle.dispose()` invokes the registry's `remove()` method, immediately unregistering that specific interceptor. This supports dynamic observability configurations where logging or debugging interceptors are toggled based on runtime conditions or feature flags.

### What is the difference between priority and order in interceptors?

**Priority** is a numeric value specified during registration where lower numbers execute earlier. **Order** is an internal auto-incrementing counter tracking registration sequence. When priorities are equal, order resolves ties. Additionally, while before interceptors execute in ascending priority order, after and on-error interceptors execute in reverse priority order to support proper resource cleanup patterns.