How Apache Maka Connects to Different Model Providers and Manages Concurrent Connections
Maka uses a pluggable adapter architecture that abstracts every LLM provider behind a unified TypeScript interface, enabling parallel streaming calls across multiple connections while centralizing error handling and usage tracking.
The Apache Maka project decouples provider-specific networking logic from its core execution engine through a sophisticated runtime system. By implementing a factory pattern combined with a protocol-based adapter design, Maka allows developers to simultaneously interact with OpenAI, Anthropic, Azure, and local models using identical method signatures.
The Pluggable Model Adapter Architecture
Maka’s runtime treats every LLM provider as an interchangeable component. This design eliminates vendor lock-in and simplifies the addition of new model endpoints without modifying core business logic.
Connection Catalog and User Configuration
User-provided credentials and endpoint definitions reside in connection-catalog.json. Each entry stores the provider identifier, authentication data, and optional per-model parameters like temperature or max-tokens. The runtime loads this catalog during initialization to determine available connections and their respective capabilities.
When the Desktop UI or CLI adds a new model account, the system persists the configuration to this JSON catalog, making it immediately available to the ModelFactory without requiring application restarts.
Model Factory and Adapter Instantiation
Located at packages/runtime/src/model-factory.ts, the ModelFactory serves as the instantiation layer. When a turn begins, the factory receives a connection identifier from the Runtime Host, reads the corresponding catalog entry, and selects the appropriate concrete adapter class.
The factory supports multiple adapter types including OpenAIAdapter, AnthropicAdapter, and LocalAdapter. It injects credentials and provider-specific metadata into the selected class before returning a fully configured instance to the caller.
The Model Adapter Interface
All provider-specific implementations conform to a strict protocol defined in model-protocol.ts, ensuring consistent behavior regardless of the underlying LLM service.
Standardized Provider Methods
Every adapter implements three core methods defined in packages/runtime/src/model-adapter.ts:
call(messages, options)– Transmits the conversation history to the provider and returns either a fully resolved response or a streaming handle.close()– Disposes of underlying HTTP or WebSocket resources to prevent connection leaks.onError– Provides a callback mechanism for provider-specific failures.
Adapters also expose an onData event for streaming scenarios, emitting chunks as they arrive from the provider’s endpoint.
Error Normalization
Provider-specific error codes undergo transformation into unified ModelError objects within the adapter layer. For example, OpenAI rate-limit responses and Anthropic quota exceeded messages both map to standardized error types that the Desktop renderer displays using logic in apps/desktop/src/renderer/model-connection-errors.ts. This normalization prevents UI components from handling vendor-specific edge cases.
Managing Multiple Concurrent Connections
The ModelRuntime class in packages/runtime/src/model-runtime.ts orchestrates simultaneous interactions with multiple providers through a connection pool pattern.
Parallel Execution and Promise Aggregation
The runtime maintains a Map<connectionId, ModelAdapter> to track active instances. For parallel inference across different models, the runtime invokes call() on multiple adapters and aggregates results using Promise.all():
// Execute calls to OpenAI and local Llama simultaneously
const [openAIResult, localResult] = await Promise.all([
runtime.call('openai-01', messages),
runtime.call('local-llama', messages),
]);
This approach allows benchmark comparisons or ensemble reasoning strategies where multiple models process identical prompts concurrently.
Streaming Aggregation and Ordering
When operating in streaming mode, the runtime subscribes to each adapter’s onData event. It merges chunks into a single logical turn while preserving per-connection ordering. The aggregation layer handles backpressure and ensures that slow providers do not block the delivery of faster streams.
Usage Accounting and Ledger Persistence
Each adapter reports token consumption upon completion. The runtime aggregates these metrics and persists them to the model-call-ledger located at packages/storage/src/model-call-ledger.ts. This ledger stores per-call timestamps, token counts, and connection identifiers, enabling cost tracking and audit trails for enterprise deployments.
Concurrency Controls and Provider Limits
The runtime enforces connection constraints using metadata generated by the refresh:model-metadata script. This data, stored in model-metadata.generated.ts, defines per-provider concurrency caps and rate limits.
When requests exceed available slots, the runtime queues excess calls until capacity frees. Adapters catch rate-limit errors and bubble them up as ModelError objects, allowing the UI to display retry timers or fallback options without crashing the execution context.
Dynamic Metadata Refresh
Maka’s build pipeline includes a refresh:model-metadata script that pulls the latest provider specifications from models.dev. This process regenerates TypeScript metadata files, ensuring the runtime always recognizes current model capabilities, pricing tiers, and concurrency constraints without manual code updates.
Practical Implementation Example
The following TypeScript demonstrates how client applications interact with Maka’s runtime to execute parallel model calls:
// Load connection definitions
import { loadConnectionCatalog } from '@maka/storage';
const catalog = await loadConnectionCatalog();
// Initialize runtime with catalog
import { ModelRuntime } from '@maka/runtime';
const runtime = new ModelRuntime(catalog);
// Parallel inference across providers
const [openAIResult, localResult] = await Promise.all([
runtime.call('openai-01', [{ role: 'user', content: 'Explain quantum entanglement' }]),
runtime.call('local-llama', [{ role: 'user', content: 'Explain quantum entanglement' }]),
]);
console.log('OpenAI:', openAIResult.message);
console.log('Local:', localResult.message);
// Streaming with event handlers
const stream = runtime.stream('openai-01', [{ role: 'user', content: 'Write a haiku' }]);
stream.on('data', chunk => process.stdout.write(chunk.content));
stream.on('end', () => console.log('\nStream complete'));
Summary
- Connection Catalog (
connection-catalog.json) persists user credentials and endpoint configurations separately from application code. - ModelFactory (
packages/runtime/src/model-factory.ts) instantiates provider-specific adapters based on catalog entries. - ModelAdapter (
packages/runtime/src/model-adapter.ts) enforces a uniform interface across all LLM providers through standardizedcall(),close(), and error handling methods. - ModelRuntime (
packages/runtime/src/model-runtime.ts) manages a Map of active adapters, supports parallel execution viaPromise.all(), and aggregates streaming responses. - ModelCallLedger (
packages/storage/src/model-call-ledger.ts) records token usage for auditing and cost management. - Dynamic Refresh scripts keep provider metadata current without requiring code changes.
Frequently Asked Questions
How does Maka handle authentication for different providers?
Maka stores API keys, endpoints, and local binary paths in connection-catalog.json. The ModelFactory reads these entries during adapter instantiation and injects credentials into the appropriate adapter class, ensuring provider-specific authentication headers are set without exposing secrets to the core runtime logic.
Can Maka stream responses from multiple models simultaneously?
Yes. The ModelRuntime maintains independent adapter instances in a Map and subscribes to each adapter’s onData event. It merges streaming chunks while preserving per-connection ordering, allowing real-time comparison of outputs from OpenAI, Anthropic, and local models within the same conversation turn.
What happens when a provider rate limit is reached?
Adapters catch provider-specific rate-limit errors and normalize them into ModelError objects. The runtime either queues subsequent calls if concurrency slots are full or bubbles the error to the UI layer, where model-connection-errors.ts renders appropriate retry guidance without terminating the application session.
How does Maka add support for new model providers?
Developers create a new class implementing the ModelAdapter interface in packages/runtime/src/model-adapter.ts, handling the provider’s specific HTTP or WebSocket protocols. After registering the adapter in the ModelFactory, users simply add a new entry to connection-catalog.json with the provider’s endpoint and credentials, requiring no changes to existing business logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →