# Core Components of Onyx: A Technical Breakdown of the Multi-Tenant AI Platform

> Discover the core components of Onyx, a multi-tenant AI platform. Explore its FastAPI, Next.js, Celery, PostgreSQL, Redis, and Vespa architecture for enterprise RAG.

- Repository: [Onyx/onyx](https://github.com/onyx-dot-app/onyx)
- Tags: technical-breakdown
- Published: 2026-03-28

---

**Onyx is a modular, multi-tenant platform built on a FastAPI backend, Next.js frontend, nine specialized Celery workers, PostgreSQL/Redis persistence, and Vespa vector search, designed for enterprise RAG and AI-augmented chat.**

The `onyx-dot-app/onyx` repository delivers an open-source AI answer engine that combines document ingestion, vector search, and large language model integration. Understanding the core components of Onyx is essential for developers deploying self-hosted instances or extending the platform's connector ecosystem. The architecture separates concerns between synchronous API handling, asynchronous background processing, and modern web interfaces.

## Backend Architecture: The FastAPI Core

At the heart of Onyx lies a **FastAPI application** that serves as the main HTTP API server. This component handles authentication, chat orchestration, document indexing endpoints, and all user-facing REST APIs.

### Main Application Entry Point

The FastAPI app is instantiated in [`backend/onyx/main.py`](https://github.com/onyx-dot-app/onyx/blob/main/backend/onyx/main.py) (lines 13‑17), where routers, middleware, and telemetry are wired together. This single application instance exposes the entire backend surface area, from user management to search queries, while maintaining async request handling for high concurrency.

### Authentication and Authorization Layer

Security is implemented through a dedicated **auth subsystem** located in `backend/onyx/auth/`. The platform leverages **FastAPI-Users** to provide OAuth2, SAML, OIDC, and API-key/PAT authentication, alongside role-based access control (RBAC) that secures multi-tenant deployments.

## Asynchronous Processing: The Celery Worker Ecosystem

Onyx delegates long-running tasks to a sophisticated **Celery background worker** architecture defined in [`backend/onyx/background/celery_app.py`](https://github.com/onyx-dot-app/onyx/blob/main/backend/onyx/background/celery_app.py). These thread-based workers handle connector synchronization, document ingestion, pruning, knowledge-graph construction, and health monitoring without blocking the main API.

### The Nine Worker Types

According to the source documentation in [`AGENTS.md`](https://github.com/onyx-dot-app/onyx/blob/main/AGENTS.md) (lines 29‑78), Onyx ships with nine distinct worker types:
- **Primary** – handles core background tasks and default queues
- **Docfetching** – schedules and executes connector sync jobs
- **Docprocessing** – processes raw documents for embedding
- **Light** – handles lightweight, high-frequency tasks
- **Heavy** – manages resource-intensive operations
- **KG-Processing** – builds and maintains knowledge graphs
- **Monitoring** – performs health checks and system telemetry
- **User-File-Processing** – handles ad-hoc file uploads
- **Beat** – the scheduler that triggers periodic tasks

### Starting the Primary Worker

To launch the primary worker for background task processing:

```bash

# Activate the virtual environment

source .venv/bin/activate

# Run the primary worker (handles core background tasks)

celery -A backend.onyx.background.celery_app worker \
  -Q default -c 4 --pool=threads \
  --loglevel=INFO

```

This command initializes the worker pool defined in [`backend/onyx/background/celery_app.py`](https://github.com/onyx-dot-app/onyx/blob/main/backend/onyx/background/celery_app.py), which serves as the entry point for all task queues.

## Data Storage and Retrieval Layer

Onyx persists state across multiple specialized storage systems optimized for their respective workloads.

### PostgreSQL and Redis

**PostgreSQL** serves as the persistent relational store for user data, document metadata, chat history, and task state. **Redis** functions as both the Celery message broker and a caching layer for session management and rate limiting, as documented in the technology stack section (lines 136‑141).

### Vespa Vector Database

The **Vespa vector database** stores embedded document chunks and metadata for hybrid search and retrieval-augmented generation (RAG). This specialized storage backend enables high-performance nearest-neighbor search across millions of document chunks while supporting filtering and ranking operations.

## Integration and Ingestion Systems

### LLM Integration Layer

Onyx abstracts large language model interactions through adapters built on **LiteLLM** and **LangChain** (lines 142‑144). This layer provides unified interfaces for generation, embeddings, and tool use across providers including OpenAI, Anthropic, Gemini, Ollama, and vLLM, allowing seamless swapping of underlying models without code changes.

### Connector Framework

The platform includes **over 40 source integrations** located in `backend/onyx/connectors/`. These connectors feed documents from cloud applications, file stores, databases, and APIs into the ingestion pipeline. Each connector implements a standard interface that the **docfetching worker** schedules and executes.

To implement a custom connector:

```python

# backend/onyx/connectors/filesystem/connector.py

from onyx.connectors.interfaces import BaseConnector
from onyx.connectors.models import Document

class FileSystemConnector(BaseConnector):
    def fetch(self) -> list[Document]:
        # Scan a directory, read files, return Document objects

        pass

```

## Frontend and User Interface

### Next.js Application Architecture

The **frontend UI** is a Next.js 15+ application built with React 18, TypeScript, and Tailwind CSS (lines 139‑140). Located in `web/src/app/`, this layer provides the chat interface, administrative console, and embeddable widgets that communicate with the FastAPI backend via REST endpoints.

### Opal Design System

UI consistency is maintained through the **Opal design system**, a component library housed in `web/lib/opal/`. This internal library provides reusable building blocks for new interface elements, ensuring visual coherence across the application's pages and components.

## Summary

- **FastAPI backend** ([`backend/onyx/main.py`](https://github.com/onyx-dot-app/onyx/blob/main/backend/onyx/main.py)) serves as the central API server handling all synchronous requests
- **Nine Celery worker types** manage asynchronous tasks from document ingestion to knowledge-graph processing
- **PostgreSQL and Redis** provide persistent storage and message brokering respectively
- **Vespa** operates as the specialized vector database for semantic search
- **LiteLLM and LangChain adapters** enable flexible LLM provider integration
- **40+ connectors** in `backend/onyx/connectors/` facilitate data ingestion from external sources
- **Next.js frontend** with the Opal design system delivers the user interface

## Frequently Asked Questions

### What programming languages and frameworks does Onyx use?

Onyx is built primarily in **Python** for the backend using FastAPI and Celery, with **TypeScript** and **Next.js** powering the frontend. The vector database uses Vespa (configured via Java/JSON), while Redis and PostgreSQL handle data persistence. This polyglot stack optimizes each layer for its specific workload.

### How does Onyx handle authentication and multi-tenancy?

The platform implements authentication through **FastAPI-Users** with support for OAuth2, SAML, OIDC, and API keys, located in `backend/onyx/auth/`. RBAC policies enforce tenant isolation, ensuring that document collections, chat histories, and API credentials remain segregated between organizations in multi-tenant deployments.

### Can I extend Onyx with custom data connectors?

Yes. The connector framework in `backend/onyx/connectors/` uses a standard `BaseConnector` interface that requires implementing a `fetch()` method returning `Document` objects. New connectors are automatically picked up by the **docfetching** Celery worker and can be scheduled through the admin interface without modifying core platform code.

### What is the role of Vespa in the Onyx architecture?

**Vespa** serves as the dedicated vector database and search engine, storing embedded document chunks alongside metadata for hybrid retrieval. Unlike general-purpose databases, Vespa is optimized for high-throughput approximate nearest neighbor search combined with filtering, making it ideal for RAG applications that require both semantic similarity and exact metadata matching.