# SDK vs HTTP Client Backends for LLM Integration in memU: A Complete Comparison

> Compare SDK and HTTP client backends for LLM integration in memU. Understand which backend best suits your needs for OpenAI compatible APIs.

- Repository: [NevaMind AI/memU](https://github.com/nevamind-ai/memu)
- Tags: deep-dive
- Published: 2026-02-19

---

**The SDK backend wraps the official OpenAI Python SDK for full feature support with OpenAI models only, while the HTTP client backend uses httpx to support any OpenAI-compatible API including OpenRouter, Doubao, and Grok through provider-specific implementations.**

NevaMind-AI/memU abstracts Large Language Model access through dual **SDK and HTTP client backends for LLM integration**, offering either the convenience of the official OpenAI SDK or the flexibility of a generic HTTP client. These interchangeable architectures allow developers to switch between providers without rewriting application logic, selecting the approach that best matches their target endpoint requirements.

## Core Implementation Differences

### SDK Backend Architecture

The SDK backend delegates all request handling to the official OpenAI Python SDK via `openai.AsyncOpenAI`. Implemented in [`src/memu/llm/openai_sdk.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/openai_sdk.py) within the `OpenAISDKClient` class, this approach provides typed request objects, automatic retries, and built-in streaming capabilities. It depends on the `openai` Python package and currently supports only the OpenAI API.

### HTTP Client Architecture

The HTTP client backend leverages **httpx** to issue raw HTTP requests, implementing the `HTTPLLMClient` class in [`src/memu/llm/http_client.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/http_client.py). This backend constructs payloads and parses responses explicitly, loading provider-specific backends from `src/memu/llm/backends/*` such as `OpenAILLMBackend`, `DoubaoLLMBackend`, and others. It requires only `httpx` and lightweight backend classes, eliminating heavy SDK dependencies.

## Provider Support and Extensibility

The SDK backend restricts you to official OpenAI models, relying on the vendor's SDK for new provider support. In contrast, the HTTP client backend works with any OpenAI-compatible API endpoint, including OpenRouter, Doubao, and Grok.

Adding a new provider to the HTTP backend requires only implementing a small backend class that supplies endpoint URLs and payload helpers. Extending the SDK backend necessitates waiting for vendor SDK updates or writing complex wrappers.

## Configuration and Backend Selection

The system selects the backend through the `client_backend` configuration option defined in [`src/memu/app/settings.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/settings.py). The default value is `"sdk"`, but you can specify `"httpx"` for the HTTP client.

In [`src/memu/app/service.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/service.py), the instantiation logic routes to the appropriate client:

```python

# src/memu/app/service.py

if backend == "sdk":
    from memu.llm.openai_sdk import OpenAISDKClient
    client = OpenAISDKClient(...)
elif backend == "httpx":
    from memu.llm.http_client import HTTPLLMClient
    client = HTTPLLMClient(...)

```

## Feature Set Comparison

Both backends expose the same high-level methods including `chat`, `summarize`, `vision`, `embed`, and `transcribe`. However, their capabilities differ:

- **SDK backend**: Provides full-fledged SDK features including typed request objects, automatic retries, streaming, built-in audio transcription, vision, and embeddings through the official client.
- **HTTP client backend**: Offers the same high-level interface but with explicit implementation control. You can swap endpoints or providers via `endpoint_overrides`, customize request timeouts, and inject custom headers.

## Practical Implementation Examples

### Using the SDK Backend

Configure `"client_backend": "sdk"` to instantiate `OpenAISDKClient` from [`src/memu/llm/openai_sdk.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/openai_sdk.py):

```python
from memu.llm.openai_sdk import OpenAISDKClient

client = OpenAISDKClient(
    base_url="https://api.openai.com/v1",
    api_key="YOUR_OPENAI_KEY",
    chat_model="gpt-4o-mini",
    embed_model="text-embedding-3-large",
)

# Chat completion

reply, raw = await client.chat("Explain quantum computing.")
print(reply)

# Generate embeddings

vectors, _ = await client.embed(["hello world", "memU"])
print(vectors)

```

### Using the HTTP Client Backend

Set `"client_backend": "httpx"` to use `HTTPLLMClient` from [`src/memu/llm/http_client.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/http_client.py) with any OpenAI-compatible provider:

```python
from memu.llm.http_client import HTTPLLMClient

client = HTTPLLMClient(
    base_url="https://openrouter.ai/api/v1",
    api_key="YOUR_OPENROUTER_KEY",
    chat_model="meta-llama/Meta-Llama-3.1-8B-Instruct",
    provider="openrouter",          # Options: "openai", "doubao", "grok", etc.

)

# Chat completion

reply, raw = await client.chat("Summarize the plot of Inception.")
print(reply)

# Embeddings via OpenRouter

embeds, _ = await client.embed(["first sentence", "second sentence"])
print(embeds)

```

## When to Use Each Backend

Choose the **SDK backend** when you have an OpenAI API key and require the convenience of the official client with built-in retry logic, streaming, and full feature support.

Select the **HTTP client backend** when you need to call non-OpenAI LLMs that follow the OpenAI HTTP specification, such as OpenRouter, Doubao, or Grok. This option suits scenarios requiring custom request timeouts, specific header injections, or `endpoint_overrides` that the official SDK does not expose.

## Summary

- The **SDK backend** in [`src/memu/llm/openai_sdk.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/openai_sdk.py) wraps `openai.AsyncOpenAI` and supports only OpenAI APIs with full SDK features.
- The **HTTP client backend** in [`src/memu/llm/http_client.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/llm/http_client.py) uses `httpx` to support any OpenAI-compatible provider through modular backends in `src/memu/llm/backends/*`.
- Configure the backend via the `client_backend` setting in [`src/memu/app/settings.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/settings.py), with instantiation logic handled in [`src/memu/app/service.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/service.py).
- The SDK backend requires the `openai` package, while the HTTP client depends only on `httpx` and lightweight backend classes.
- Both backends expose identical high-level methods (`chat`, `embed`, `vision`, `transcribe`), but the HTTP client offers greater extensibility for custom providers.

## Frequently Asked Questions

### Can I switch between backends without changing my application code?

Yes. Both `OpenAISDKClient` and `HTTPLLMClient` implement the same interface with methods like `chat`, `embed`, and `vision`. You only need to change the `client_backend` configuration value in your JSON or YAML config file, and [`src/memu/app/service.py`](https://github.com/NevaMind-AI/memU/blob/main/src/memu/app/service.py) will instantiate the correct client automatically.

### Which backend supports streaming responses?

The SDK backend provides streaming through the official OpenAI SDK's built-in capabilities. The HTTP client backend also supports streaming, but implements it explicitly within the `HTTPLLMClient` class rather than delegating to an external SDK.

### How do I add a custom LLM provider to memU?

You must use the HTTP client backend. Create a new provider class in `src/memu/llm/backends/` that defines endpoint URLs and payload helpers, similar to existing implementations for Doubao or Grok. The SDK backend does not support custom providers without writing a complex wrapper around the vendor's SDK.

### Does the HTTP client backend support vision and embedding models?

Yes. The `HTTPLLMClient` class provides `vision()` and `embed()` methods that function identically to the SDK backend. These methods construct the appropriate JSON payloads and parse responses for any provider that supports the OpenAI-compatible vision or embeddings API format.