# How Graphify Handles Multiple LLM Backends: Architecture and Implementation

> Discover how Graphify seamlessly integrates multiple LLM backends. It auto-detects providers via environment variables and uses a unified OpenAI-compatible interface for cloud, self-hosted, and local models without code changes.

- Repository: [Graphify Labs/graphify](https://github.com/Graphify-Labs/graphify)
- Tags: architecture
- Published: 2026-07-16

---

**Graphify auto-detects available LLM providers through environment variables and routes all requests through a unified OpenAI-compatible interface, supporting cloud APIs, self-hosted models, and local CLI tools without code changes.**

Graphify-Labs/graphify implements a sophisticated multi-backend architecture that eliminates vendor lock-in by allowing seamless switching between Claude, Gemini, OpenAI, Ollama, and other providers. The system centralizes all LLM interactions in [`graphify/llm.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/llm.py), using a declarative registry and environment-based detection to determine which provider to use. This design enables zero-configuration deployment while maintaining safety through standardized request formatting and content validation.

## The Backend Registry Architecture

### Declarative Configuration in BACKENDS

At the core of Graphify's multi-backend support lies a global constant `BACKENDS` defined in [`graphify/llm.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/llm.py) (lines 59-127). This dictionary maps each supported provider to its configuration metadata, including base URL, default model, required environment variables, pricing details, token limits, temperature defaults, and vision support flags.

The registry currently supports nine distinct backends:

- **Claude** (Anthropic)
- **Gemini** (Google)
- **OpenAI**
- **Ollama** (self-hosted)
- **Kimi** (Moonshot)
- **DeepSeek**
- **Azure OpenAI**
- **Bedrock** (AWS)
- **claude-cli** (local CLI)

Each entry follows a consistent schema that enables the rest of the system to interact with any provider through a generic interface.

## Automatic Backend Detection

### Environment Variable Precedence

The `detect_backend()` function in [`graphify/llm.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/llm.py) implements automatic provider discovery by scanning for API key environment variables. The function checks variables in a strict precedence order:

1. `GEMINI_API_KEY` or `GOOGLE_API_KEY`
2. `ANTHROPIC_API_KEY`
3. `OPENAI_API_KEY`
4. `OLLAMA_HOST` (or default localhost)
5. `DEEPSEEK_API_KEY`
6. `MOONSHOT_API_KEY` (Kimi)
7. `AZURE_OPENAI_API_KEY`
8. `AWS_ACCESS_KEY_ID` (Bedrock)

When `detect_backend()` finds the first matching key, it returns the corresponding backend identifier string. If no keys are present, the function raises a clear `ValueError` indicating that an API key must be configured. This precedence order prioritizes Gemini when available, reflecting the project's preference for Google models.

The helper function `_get_backend_api_key(name)` handles provider-specific key resolution, including support for multiple variable names (such as accepting either `GEMINI_API_KEY` or `GOOGLE_API_KEY` for Google backends).

## Configuration Overrides and Model Selection

### Customizing Model Names

While each backend defines a default model in the `BACKENDS` registry, Graphify allows runtime overrides through environment variables. Users can set `GRAPHIFY_<BACKEND>_MODEL` (e.g., `GRAPHIFY_GEMINI_MODEL` or `GRAPHIFY_OPENAI_MODEL`) to specify alternative model names.

The internal `_resolve_model_name` function checks for these overrides, preferring the generic `GRAPHIFY_<BACKEND>_MODEL` variable when both specific and generic variants exist. This enables testing of preview models or switching to different parameter counts without modifying source code.

### Temperature and Token Limits

Graphify standardizes generation parameters across providers through the `_resolve_temperature` helper. The system reads `GRAPHIFY_LLM_TEMPERATURE` from the environment, validates that the value is numeric, and applies it uniformly regardless of which backend is active.

Token limits are extracted from the `BACKENDS` configuration entry via `max_tokens` (or `max_completion_tokens` for legacy configurations). These limits propagate through to the request payload, ensuring providers respect output length constraints.

## Unified Request Dispatch

### The OpenAI-Compatible Layer

All non-CLI backends route through `_call_openai_compat`, a universal dispatch function that translates Graphify's internal representation into provider-specific HTTP requests. This function constructs a standardized **user message** that concatenates file contents, wrapping each source in `<untrusted_source path="...">` XML delimiters for safety isolation.

The function posts JSON payloads to the provider's configured `base_url`, then validates responses against size limits (`_LLM_JSON_MAX_BYTES = 10 MiB`) and JSON schema requirements before returning structured data. This abstraction allows Graphify to treat Claude, Gemini, and Azure OpenAI as interchangeable services despite their differing native APIs.

## Parallel Processing and Local Execution

### Concurrent File Processing

For large codebases, `extract_corpus_parallel` slices files into manageable chunks using `FileSlice` objects and dispatches them concurrently via `ThreadPoolExecutor`. The function accepts heterogeneous input types—strings, `pathlib.Path` objects, and explicit `FileSlice` instances—ensuring flexibility for callers while maximizing throughput through parallel backend requests.

### Local Claude-CLI Backend

When `backend="claude-cli"` is specified, Graphify bypasses HTTP entirely and invokes the local `claude` binary directly. The `_call_claude_cli` function passes prompts via the `-p --output-format json` flags, reading stdout for responses. This backend requires no API key management (costs accrue to the user's Anthropic subscription) and functions entirely offline after initial setup.

## Code Examples

### Automatic Backend Detection

```python
import os
from graphify import llm

# Set any supported API key

os.environ["OPENAI_API_KEY"] = "sk-..."

backend = llm.detect_backend()  # Returns "openai"

api_key = llm._get_backend_api_key(backend)
print(f"Selected {backend} backend")

```

### Overriding Model Selection

```python
os.environ["GEMINI_API_KEY"] = "gemini-key"
os.environ["GRAPHIFY_GEMINI_MODEL"] = "gemini-1.5-pro-preview-0514"

# Uses the preview model despite default being gemini-1.5-flash

llm.extract_files_direct([Path("doc.md")], backend="gemini", root=Path("."))

```

### Parallel Extraction with Mixed Input Types

```python
from pathlib import Path
from graphify.llm import FileSlice, extract_corpus_parallel

files = [
    "legacy.py",                    # string path

    Path("modern.py"),              # pathlib.Path

    FileSlice("huge.log", 0, 5000)  # sliced chunk

]

result = extract_corpus_parallel(
    files,
    backend="ollama",
    root=Path("."),
    max_concurrency=4
)

```

### Using the Local Claude-CLI Backend

```python

# Ensure `claude` binary is in PATH

result = llm.extract_files_direct(
    [Path("source.py")],
    backend="claude-cli",
    root=Path(".")
)

# No API key required; billed to Claude subscription

```

## Summary

- **Zero-code configuration**: Export any supported API key and Graphify automatically selects the appropriate backend via `detect_backend()`.
- **Predictable precedence**: The detection order favors Gemini, then Claude, then OpenAI, ensuring consistent behavior across environments.
- **Unified interface**: `_call_openai_compat` standardizes all cloud providers behind a single OpenAI-compatible request layer.
- **Safety by default**: All file content wraps in `<untrusted_source>` delimiters with a 10 MiB response size cap.
- **Flexible execution**: Supports both cloud APIs and local CLI tools (claude-cli) without architectural changes.
- **Concurrent processing**: `extract_corpus_parallel` handles mixed path types and large files through automatic slicing and ThreadPoolExecutor dispatch.

## Frequently Asked Questions

### How does Graphify decide which LLM backend to use?

Graphify uses the `detect_backend()` function in [`graphify/llm.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/llm.py) to scan for environment variables in a fixed precedence order: Gemini keys are checked first, followed by Claude, OpenAI, Ollama, DeepSeek, Kimi, Azure, and Bedrock. The first API key found determines the active backend. If no keys are present, the function raises a `ValueError` prompting for configuration.

### Can I use multiple LLM backends simultaneously in one Graphify session?

While `detect_backend()` returns a single active backend, you can explicitly pass different backend names to individual function calls. For example, you could call `extract_files_direct` with `backend="openai"` for one batch and `backend="claude"` for another within the same Python process. However, each individual extraction operation uses one specific backend.

### What is the Claude-CLI backend and when should I use it?

The `claude-cli` backend invokes the local Anthropic Claude command-line tool instead of making HTTP API calls. Available when `backend="claude-cli"` is specified, this mode requires no API key management and processes requests locally using the user's existing Claude subscription. Use this backend for offline operation, unlimited personal usage (subject to Claude CLI terms), or when avoiding cloud API latency.

### How does Graphify handle large files that exceed token limits?

Graphify addresses large files through the `FileSlice` abstraction and `extract_corpus_parallel` function. Large inputs are automatically sliced into chunks that respect the `max_tokens` or `max_completion_tokens` limits defined in the `BACKENDS` registry. These chunks are processed concurrently via `ThreadPoolExecutor`, with results aggregated before returning to the caller.