How Graphify Handles Multiple LLM Backends: Architecture and Implementation

Graphify auto-detects available LLM providers through environment variables and routes all requests through a unified OpenAI-compatible interface, supporting cloud APIs, self-hosted models, and local CLI tools without code changes.

Graphify-Labs/graphify implements a sophisticated multi-backend architecture that eliminates vendor lock-in by allowing seamless switching between Claude, Gemini, OpenAI, Ollama, and other providers. The system centralizes all LLM interactions in graphify/llm.py, using a declarative registry and environment-based detection to determine which provider to use. This design enables zero-configuration deployment while maintaining safety through standardized request formatting and content validation.

The Backend Registry Architecture

Declarative Configuration in BACKENDS

At the core of Graphify's multi-backend support lies a global constant BACKENDS defined in graphify/llm.py (lines 59-127). This dictionary maps each supported provider to its configuration metadata, including base URL, default model, required environment variables, pricing details, token limits, temperature defaults, and vision support flags.

The registry currently supports nine distinct backends:

  • Claude (Anthropic)
  • Gemini (Google)
  • OpenAI
  • Ollama (self-hosted)
  • Kimi (Moonshot)
  • DeepSeek
  • Azure OpenAI
  • Bedrock (AWS)
  • claude-cli (local CLI)

Each entry follows a consistent schema that enables the rest of the system to interact with any provider through a generic interface.

Automatic Backend Detection

Environment Variable Precedence

The detect_backend() function in graphify/llm.py implements automatic provider discovery by scanning for API key environment variables. The function checks variables in a strict precedence order:

  1. GEMINI_API_KEY or GOOGLE_API_KEY
  2. ANTHROPIC_API_KEY
  3. OPENAI_API_KEY
  4. OLLAMA_HOST (or default localhost)
  5. DEEPSEEK_API_KEY
  6. MOONSHOT_API_KEY (Kimi)
  7. AZURE_OPENAI_API_KEY
  8. AWS_ACCESS_KEY_ID (Bedrock)

When detect_backend() finds the first matching key, it returns the corresponding backend identifier string. If no keys are present, the function raises a clear ValueError indicating that an API key must be configured. This precedence order prioritizes Gemini when available, reflecting the project's preference for Google models.

The helper function _get_backend_api_key(name) handles provider-specific key resolution, including support for multiple variable names (such as accepting either GEMINI_API_KEY or GOOGLE_API_KEY for Google backends).

Configuration Overrides and Model Selection

Customizing Model Names

While each backend defines a default model in the BACKENDS registry, Graphify allows runtime overrides through environment variables. Users can set GRAPHIFY_<BACKEND>_MODEL (e.g., GRAPHIFY_GEMINI_MODEL or GRAPHIFY_OPENAI_MODEL) to specify alternative model names.

The internal _resolve_model_name function checks for these overrides, preferring the generic GRAPHIFY_<BACKEND>_MODEL variable when both specific and generic variants exist. This enables testing of preview models or switching to different parameter counts without modifying source code.

Temperature and Token Limits

Graphify standardizes generation parameters across providers through the _resolve_temperature helper. The system reads GRAPHIFY_LLM_TEMPERATURE from the environment, validates that the value is numeric, and applies it uniformly regardless of which backend is active.

Token limits are extracted from the BACKENDS configuration entry via max_tokens (or max_completion_tokens for legacy configurations). These limits propagate through to the request payload, ensuring providers respect output length constraints.

Unified Request Dispatch

The OpenAI-Compatible Layer

All non-CLI backends route through _call_openai_compat, a universal dispatch function that translates Graphify's internal representation into provider-specific HTTP requests. This function constructs a standardized user message that concatenates file contents, wrapping each source in <untrusted_source path="..."> XML delimiters for safety isolation.

The function posts JSON payloads to the provider's configured base_url, then validates responses against size limits (_LLM_JSON_MAX_BYTES = 10 MiB) and JSON schema requirements before returning structured data. This abstraction allows Graphify to treat Claude, Gemini, and Azure OpenAI as interchangeable services despite their differing native APIs.

Parallel Processing and Local Execution

Concurrent File Processing

For large codebases, extract_corpus_parallel slices files into manageable chunks using FileSlice objects and dispatches them concurrently via ThreadPoolExecutor. The function accepts heterogeneous input types—strings, pathlib.Path objects, and explicit FileSlice instances—ensuring flexibility for callers while maximizing throughput through parallel backend requests.

Local Claude-CLI Backend

When backend="claude-cli" is specified, Graphify bypasses HTTP entirely and invokes the local claude binary directly. The _call_claude_cli function passes prompts via the -p --output-format json flags, reading stdout for responses. This backend requires no API key management (costs accrue to the user's Anthropic subscription) and functions entirely offline after initial setup.

Code Examples

Automatic Backend Detection

import os
from graphify import llm

# Set any supported API key

os.environ["OPENAI_API_KEY"] = "sk-..."

backend = llm.detect_backend()  # Returns "openai"

api_key = llm._get_backend_api_key(backend)
print(f"Selected {backend} backend")

Overriding Model Selection

os.environ["GEMINI_API_KEY"] = "gemini-key"
os.environ["GRAPHIFY_GEMINI_MODEL"] = "gemini-1.5-pro-preview-0514"

# Uses the preview model despite default being gemini-1.5-flash

llm.extract_files_direct([Path("doc.md")], backend="gemini", root=Path("."))

Parallel Extraction with Mixed Input Types

from pathlib import Path
from graphify.llm import FileSlice, extract_corpus_parallel

files = [
    "legacy.py",                    # string path

    Path("modern.py"),              # pathlib.Path

    FileSlice("huge.log", 0, 5000)  # sliced chunk

]

result = extract_corpus_parallel(
    files,
    backend="ollama",
    root=Path("."),
    max_concurrency=4
)

Using the Local Claude-CLI Backend


# Ensure `claude` binary is in PATH

result = llm.extract_files_direct(
    [Path("source.py")],
    backend="claude-cli",
    root=Path(".")
)

# No API key required; billed to Claude subscription

Summary

  • Zero-code configuration: Export any supported API key and Graphify automatically selects the appropriate backend via detect_backend().
  • Predictable precedence: The detection order favors Gemini, then Claude, then OpenAI, ensuring consistent behavior across environments.
  • Unified interface: _call_openai_compat standardizes all cloud providers behind a single OpenAI-compatible request layer.
  • Safety by default: All file content wraps in <untrusted_source> delimiters with a 10 MiB response size cap.
  • Flexible execution: Supports both cloud APIs and local CLI tools (claude-cli) without architectural changes.
  • Concurrent processing: extract_corpus_parallel handles mixed path types and large files through automatic slicing and ThreadPoolExecutor dispatch.

Frequently Asked Questions

How does Graphify decide which LLM backend to use?

Graphify uses the detect_backend() function in graphify/llm.py to scan for environment variables in a fixed precedence order: Gemini keys are checked first, followed by Claude, OpenAI, Ollama, DeepSeek, Kimi, Azure, and Bedrock. The first API key found determines the active backend. If no keys are present, the function raises a ValueError prompting for configuration.

Can I use multiple LLM backends simultaneously in one Graphify session?

While detect_backend() returns a single active backend, you can explicitly pass different backend names to individual function calls. For example, you could call extract_files_direct with backend="openai" for one batch and backend="claude" for another within the same Python process. However, each individual extraction operation uses one specific backend.

What is the Claude-CLI backend and when should I use it?

The claude-cli backend invokes the local Anthropic Claude command-line tool instead of making HTTP API calls. Available when backend="claude-cli" is specified, this mode requires no API key management and processes requests locally using the user's existing Claude subscription. Use this backend for offline operation, unlimited personal usage (subject to Claude CLI terms), or when avoiding cloud API latency.

How does Graphify handle large files that exceed token limits?

Graphify addresses large files through the FileSlice abstraction and extract_corpus_parallel function. Large inputs are automatically sliced into chunks that respect the max_tokens or max_completion_tokens limits defined in the BACKENDS registry. These chunks are processed concurrently via ThreadPoolExecutor, with results aggregated before returning to the caller.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →