How to Integrate Onyx with Other Tools: A Complete Enterprise Search Integration Guide
Integrate Onyx with external services by implementing custom connectors for data ingestion, consuming the REST API for programmatic search, or extending Celery workers to push indexed documents to downstream systems.
Onyx is an open-source Gen-AI and enterprise search platform designed to unify data from disparate sources. To integrate Onyx with other tools, developers leverage a modular connector framework for data ingestion and a Next.js-based REST API for programmatic access. This guide covers the technical implementation patterns found in the onyx-dot-app/onyx repository, including specific file paths and code examples for building robust integrations.
Onyx Integration Architecture Overview
Onyx connects to external tools through three primary mechanisms: connectors for data ingestion, a REST API for data retrieval, and background workers for asynchronous processing.
Connectors are Python classes that subclass base types defined in backend/onyx/connectors/interfaces.py, such as LoadConnector for full data loads, PollConnector for incremental synchronization, or CheckpointedConnector for resumable large-scale ingestion. These implementations live in backend/onyx/connectors/ and handle authentication, data fetching, and format normalization.
The platform exposes HTTP endpoints via the Next.js frontend (located in web/src/app/) and Python API routes in backend/onyx/server/api/. Background task processing relies on Celery workers—including docfetching, docprocessing, light, and heavy queues—defined in backend/onyx/background/apps/.
How to Build a Custom Connector
To add a new SaaS integration (for example, Notion or Jira), you must implement the connector interface, register it in the factory, and configure the UI.
Step 1: Implement the Connector Class
Create a new file in backend/onyx/connectors/ (e.g., notion_connector.py) that extends the appropriate base class. A PollConnector for incremental sync requires three key methods:
from onyx.connectors.base import PollConnector
from typing import List, Dict
class NotionConnector(PollConnector):
def __init__(self, workspace: str):
self.workspace = workspace
def load_credentials(self, cred: Dict[str, str]) -> None:
"""Store OAuth token from UI configuration."""
self.token = cred["notion_token"]
def poll_source(self, start_ts: float, end_ts: float) -> List[Dict]:
"""Fetch pages modified between start_ts and end_ts."""
# Call Notion API with self.token
pass
def load_from_state(self) -> List[Dict]:
"""Perform full sync for initial load."""
# Fetch all pages in workspace
pass
Step 2: Register in the Connector Factory
Open backend/onyx/connectors/factory.py and add your class to the mapping near line 33:
from onyx.connectors.notion_connector import NotionConnector
CONNECTOR_MAP = {
# existing mappings...
"notion": NotionConnector,
}
Step 3: Configure the Frontend UI
Add connector metadata to web/src/lib/connectors/connectors.ts (around line 79) so administrators can enter credentials via the "Add Connector" page. This TypeScript configuration defines the input fields and validation rules for the Notion workspace and token.
Step 4: Write Integration Tests
Create test files in backend/tests/daily/connectors/notion/ following the pattern in backend/tests/daily/connectors/confluence/. Run the suite with:
python -m dotenv -f .env run -- pytest backend/tests/daily/connectors/notion
How to Consume Onyx Search Results
External applications can query indexed data through the REST API. First, generate an API key in the admin UI (Settings → API Keys) or retrieve it from your .env file.
Search API Request
import os
import requests
API_URL = "http://localhost:3000/api/search"
API_KEY = os.getenv("ONYX_API_KEY")
def onyx_search(query: str):
headers = {"Authorization": f"Bearer {API_KEY}"}
params = {"q": query}
response = requests.get(API_URL, headers=headers, params=params)
response.raise_for_status()
return response.json()["results"]
# Example usage
for hit in onyx_search("quarterly revenue"):
print(f"- {hit['title']} ({hit['source']})")
The endpoint returns document IDs, content snippets, and relevance scores. See backend/onyx/server/api/search.py for the full route implementation.
Trigger Manual Indexing
To force a connector refresh outside the regular schedule, call the admin endpoint:
curl -X POST "http://localhost:3000/api/connector/42/run" \
-H "Authorization: Bearer $ONYX_API_KEY"
This queues a task in the docfetching worker defined in backend/onyx/background/apps/docfetching.py.
How to Push Data to Downstream Systems
For real-time synchronization from Onyx to external tools, extend the light worker in backend/onyx/background/apps/light.py. Create a task that subscribes to the docprocessing queue and POSTs document events to your webhook:
# In backend/onyx/background/apps/light.py
@app.task
def notify_external_system(doc_id: str, content: str):
requests.post(
"https://your-tool.com/webhook",
json={"id": doc_id, "text": content}
)
Deploy this worker using the same Celery configuration as the standard light worker to ensure it processes document events immediately after chunking and embedding.
Critical Source Files for Integration
| Component | File Path | Purpose |
|---|---|---|
| Connector Base Classes | backend/onyx/connectors/interfaces.py |
Defines LoadConnector, PollConnector, CheckpointedConnector |
| Connector Factory | backend/onyx/connectors/factory.py |
Registration mapping for all connector types |
| Connector Docs | backend/onyx/connectors/README.md |
Architecture overview and implementation guidelines |
| Fetch Workers | backend/onyx/background/apps/docfetching.py |
Schedules and executes connector data retrieval |
| Processing Workers | backend/onyx/background/apps/docprocessing.py |
Handles chunking, embedding, and storage |
| Light Workers | backend/onyx/background/apps/light.py |
Runs lightweight tasks and custom webhooks |
| Beat Scheduler | backend/onyx/background/apps/beat.py |
Configures periodic tasks like check_for_indexing |
| Search API | backend/onyx/server/api/search.py |
Search endpoint implementation |
| Connector API | backend/onyx/server/api/connector.py |
Connector management endpoints |
| Frontend Config | web/src/lib/connectors/connectors.ts |
UI schema for credential input |
| Test Patterns | backend/tests/daily/connectors/confluence/ |
Example connector test implementation |
Summary
- Integrate Onyx with other tools using three patterns: custom connectors for ingestion, REST API for retrieval, and Celery workers for downstream publishing.
- Implement connectors by subclassing
PollConnectororLoadConnectorinbackend/onyx/connectors/, then register them infactory.pyand add UI config inweb/src/lib/connectors/connectors.ts. - Authenticate API requests using Bearer tokens in the
Authorizationheader against endpoints defined inbackend/onyx/server/api/. - Trigger manual indexing via
POST /api/connector/{id}/runor wait for the Celery beat scheduler inbackend/onyx/background/apps/beat.pyto runcheck_for_indexingevery 15 seconds. - Extend
backend/onyx/background/apps/light.pyto push document events to external webhooks.
Frequently Asked Questions
How do I authenticate API requests to Onyx?
Generate an API key through the administrative UI under Settings → API Keys, or use the default token specified in your .env file. Include this token in the Authorization header as Bearer <token> when calling endpoints like /api/search or /api/connector/{id}/run.
What is the difference between PollConnector and LoadConnector?
PollConnector is designed for incremental synchronization; you implement poll_source(start_ts, end_ts) to fetch only data modified within a specific timeframe. LoadConnector is used for full data loads through the load_from_state() method, typically for initial ingestion or sources without timestamp-based filtering. Both base classes are defined in backend/onyx/connectors/interfaces.py.
How do I trigger a manual re-index of a connector?
Send a POST request to /api/connector/{id}/run with a valid API key, or dispatch the task directly to the Celery docfetching queue. The platform also runs the check_for_indexing task automatically every 15 seconds via the beat worker in backend/onyx/background/apps/beat.py to handle scheduled refreshes.
Can I embed Onyx search into my existing web application?
Yes. You can iframe the hosted Next.js application (http://localhost:3000) or reuse the React component library located in web/src/components/. For programmatic integration, call the search API from your backend using the Python client pattern shown above and render the JSON results in your native UI.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →