Main Features of Onyx: Self-Hostable Gen-AI Platform with Enterprise Search
Onyx is a self-hostable, open-source Gen-AI platform that combines custom AI agents, retrieval-augmented generation (RAG), and 40+ enterprise connectors into a modular architecture with fine-grained permissioning.
The onyx-dot-app/onyx repository delivers an MIT-licensed alternative to proprietary AI solutions, offering both a Community Edition and an Enterprise Edition with advanced scalability features. Its architecture separates concerns into specialized Celery workers, enabling horizontal scaling from single-container deployments to Kubernetes clusters handling tens of millions of documents.
Custom AI Agents and Multi-Tool Control Plane (MCP)
Onyx allows users to build custom AI agents with unique instructions, dedicated knowledge bases, and configurable actions. According to the repository's main documentation in README.md, agents can be extended with external tools including web browsing, API calls, and code execution.
The MCP (Multi-Tool Control Plane) provides the execution layer for these agents. As implemented in README.md line 53, MCP gives agents the ability to invoke external tools, run bash commands, manipulate files, or call third-party APIs under fine-grained permission controls. This creates a secure boundary between the LLM and the host system while enabling powerful automation workflows.
Retrieval-Augmented Generation (RAG) and Enterprise Search
At the core of Onyx is a RAG (Retrieval-Augmented Generation) system that combines hybrid search with knowledge-graph capabilities over ingested documents. The platform stores document metadata in Postgres and indexes content in Vespa, enabling high-performance semantic and keyword search across millions of files.
The enterprise search feature mirrors original access controls from source applications, ensuring users only see documents they are authorized to access. As noted in README.md line 81, this permission-aware indexing scales to tens of millions of documents while maintaining security boundaries from the original source apps like Google Drive, Confluence, or SharePoint.
40+ Built-in Connectors for Workplace Applications
Onyx includes native connectors for over 40 workplace applications, eliminating the need for custom ETL pipelines. The platform supports major productivity suites including Google Drive, Slack, Confluence, Linear, and Jira, as documented in README.md line 33.
Each connector operates through a dedicated ingestion pipeline:
- Doc-fetching workers pull content via OAuth or API keys
- Doc-processing workers extract text, metadata, and permissions
- Indexing workers vectorize and store data in Vespa
This separation prevents any single connector from blocking the ingestion queue and allows independent scaling of fetch vs. processing capacity.
Chat UI and Embeddable Widget
The platform provides a lightweight, self-hosted chat interface that supports streaming responses from any LLM. For external integration, Onyx offers an embeddable widget described in widget/README.md, allowing organizations to inject the chat interface into existing web applications or internal tools.
Additionally, the extensions/chrome/README.md demonstrates how agents can be exposed to end-users through browser extensions, extending Onyx capabilities directly into the user's workflow without leaving their current context.
Sandboxed Code Execution with Craft
The Craft feature provides sandboxed code execution environments for AI agents. According to web/src/app/craft/README.md line 38, agents can run code in isolated Next.js or Python sandboxes to generate documents, dashboards, presentations, and other artifacts.
This sandboxing ensures that agent-generated code cannot affect the host system while still allowing powerful automation tasks like data analysis, report generation, and visual content creation.
Scalable Backend Architecture with Celery Workers
Onyx employs a modular Celery-based architecture that isolates workloads across specialized worker types. As detailed in AGENTS.md line 25, the backend includes:
- Primary workers handling chat orchestration and API requests
- Doc-fetching workers managing external API connections
- Doc-processing workers handling text extraction and chunking
- Indexing workers managing vectorization and Vespa updates
- Pruning workers cleaning up stale or deleted documents
Each worker type is configurable and thread-based for stability, allowing operators to scale specific bottlenecks independently. For example, heavy ingestion days can trigger additional doc-processing workers without affecting chat response latency.
Community and Enterprise Editions
The repository offers two distinct editions:
Community Edition (MIT License) provides full access to agents, RAG, connectors, and the chat UI suitable for small to medium deployments.
Enterprise Edition adds advanced capabilities including enhanced RBAC (Role-Based Access Control), custom model routing, and premium connector support for large-scale organizational deployments, as noted in README.md line 97.
Summary
- Onyx is a self-hostable Gen-AI platform offering enterprise search, custom agents, and RAG capabilities in a single open-source package.
- The MCP (Multi-Tool Control Plane) enables agents to safely execute external tools, bash commands, and API calls under strict permission controls.
- 40+ connectors ingest content from workplace apps while preserving original document permissions and access controls.
- The modular Celery architecture separates document fetching, processing, and indexing into horizontally scalable worker pools.
- Sandboxed execution via the Craft feature allows agents to generate code, documents, and dashboards without compromising host security.
- Both Community (MIT) and Enterprise editions are available, with the latter providing advanced RBAC and premium integrations.
Frequently Asked Questions
What is the difference between Onyx Community and Enterprise Edition?
The Community Edition is open-source under the MIT license and includes core features like AI agents, RAG search, all 40+ connectors, and the chat widget. The Enterprise Edition adds advanced role-based access control (RBAC), custom model routing capabilities, and premium connector support designed for large organizations with complex compliance requirements.
How does Onyx handle document security and permissions?
Onyx implements source-aware permissioning that mirrors access controls from the original applications (such as Google Drive or Confluence). When documents are ingested through connectors, their permission metadata is preserved in Postgres and enforced at query time, ensuring users only retrieve search results from documents they are authorized to view in the source system.
Can Onyx run completely offline or in air-gapped environments?
Yes, Onyx can operate entirely offline when deployed via Docker Compose, Helm, or Terraform. The platform supports self-hosted LLMs and does not require external API calls, making it suitable for air-gapped deployments where data cannot leave the organizational network. However, external web search capabilities (Google PSE, Exa, Serper) require internet connectivity unless replaced with internal search indexes.
What types of AI models does Onyx support?
Onyx is model-agnostic and works with any LLM that provides an OpenAI-compatible API endpoint, including self-hosted models like Llama, Mistral, or custom enterprise deployments. The Chat UI and widget support streaming responses from these models, and the agent framework in README.md line 48 can be configured to route different tasks to different model endpoints based on capability requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →