How to Integrate Onyx with Other Tools: A Complete Enterprise Search Integration Guide

Integrate Onyx with external services by implementing custom connectors for data ingestion, consuming the REST API for programmatic search, or extending Celery workers to push indexed documents to downstream systems.

Onyx is an open-source Gen-AI and enterprise search platform designed to unify data from disparate sources. To integrate Onyx with other tools, developers leverage a modular connector framework for data ingestion and a Next.js-based REST API for programmatic access. This guide covers the technical implementation patterns found in the onyx-dot-app/onyx repository, including specific file paths and code examples for building robust integrations.

Onyx Integration Architecture Overview

Onyx connects to external tools through three primary mechanisms: connectors for data ingestion, a REST API for data retrieval, and background workers for asynchronous processing.

Connectors are Python classes that subclass base types defined in backend/onyx/connectors/interfaces.py, such as LoadConnector for full data loads, PollConnector for incremental synchronization, or CheckpointedConnector for resumable large-scale ingestion. These implementations live in backend/onyx/connectors/ and handle authentication, data fetching, and format normalization.

The platform exposes HTTP endpoints via the Next.js frontend (located in web/src/app/) and Python API routes in backend/onyx/server/api/. Background task processing relies on Celery workers—including docfetching, docprocessing, light, and heavy queues—defined in backend/onyx/background/apps/.

How to Build a Custom Connector

To add a new SaaS integration (for example, Notion or Jira), you must implement the connector interface, register it in the factory, and configure the UI.

Step 1: Implement the Connector Class

Create a new file in backend/onyx/connectors/ (e.g., notion_connector.py) that extends the appropriate base class. A PollConnector for incremental sync requires three key methods:

from onyx.connectors.base import PollConnector
from typing import List, Dict

class NotionConnector(PollConnector):
    def __init__(self, workspace: str):
        self.workspace = workspace

    def load_credentials(self, cred: Dict[str, str]) -> None:
        """Store OAuth token from UI configuration."""
        self.token = cred["notion_token"]

    def poll_source(self, start_ts: float, end_ts: float) -> List[Dict]:
        """Fetch pages modified between start_ts and end_ts."""
        # Call Notion API with self.token

        pass

    def load_from_state(self) -> List[Dict]:
        """Perform full sync for initial load."""
        # Fetch all pages in workspace

        pass

Step 2: Register in the Connector Factory

Open backend/onyx/connectors/factory.py and add your class to the mapping near line 33:

from onyx.connectors.notion_connector import NotionConnector

CONNECTOR_MAP = {
    # existing mappings...

    "notion": NotionConnector,
}

Step 3: Configure the Frontend UI

Add connector metadata to web/src/lib/connectors/connectors.ts (around line 79) so administrators can enter credentials via the "Add Connector" page. This TypeScript configuration defines the input fields and validation rules for the Notion workspace and token.

Step 4: Write Integration Tests

Create test files in backend/tests/daily/connectors/notion/ following the pattern in backend/tests/daily/connectors/confluence/. Run the suite with:

python -m dotenv -f .env run -- pytest backend/tests/daily/connectors/notion

How to Consume Onyx Search Results

External applications can query indexed data through the REST API. First, generate an API key in the admin UI (Settings → API Keys) or retrieve it from your .env file.

Search API Request

import os
import requests

API_URL = "http://localhost:3000/api/search"
API_KEY = os.getenv("ONYX_API_KEY")

def onyx_search(query: str):
    headers = {"Authorization": f"Bearer {API_KEY}"}
    params = {"q": query}
    response = requests.get(API_URL, headers=headers, params=params)
    response.raise_for_status()
    return response.json()["results"]

# Example usage

for hit in onyx_search("quarterly revenue"):
    print(f"- {hit['title']} ({hit['source']})")

The endpoint returns document IDs, content snippets, and relevance scores. See backend/onyx/server/api/search.py for the full route implementation.

Trigger Manual Indexing

To force a connector refresh outside the regular schedule, call the admin endpoint:

curl -X POST "http://localhost:3000/api/connector/42/run" \
     -H "Authorization: Bearer $ONYX_API_KEY"

This queues a task in the docfetching worker defined in backend/onyx/background/apps/docfetching.py.

How to Push Data to Downstream Systems

For real-time synchronization from Onyx to external tools, extend the light worker in backend/onyx/background/apps/light.py. Create a task that subscribes to the docprocessing queue and POSTs document events to your webhook:


# In backend/onyx/background/apps/light.py

@app.task
def notify_external_system(doc_id: str, content: str):
    requests.post(
        "https://your-tool.com/webhook",
        json={"id": doc_id, "text": content}
    )

Deploy this worker using the same Celery configuration as the standard light worker to ensure it processes document events immediately after chunking and embedding.

Critical Source Files for Integration

Component File Path Purpose
Connector Base Classes backend/onyx/connectors/interfaces.py Defines LoadConnector, PollConnector, CheckpointedConnector
Connector Factory backend/onyx/connectors/factory.py Registration mapping for all connector types
Connector Docs backend/onyx/connectors/README.md Architecture overview and implementation guidelines
Fetch Workers backend/onyx/background/apps/docfetching.py Schedules and executes connector data retrieval
Processing Workers backend/onyx/background/apps/docprocessing.py Handles chunking, embedding, and storage
Light Workers backend/onyx/background/apps/light.py Runs lightweight tasks and custom webhooks
Beat Scheduler backend/onyx/background/apps/beat.py Configures periodic tasks like check_for_indexing
Search API backend/onyx/server/api/search.py Search endpoint implementation
Connector API backend/onyx/server/api/connector.py Connector management endpoints
Frontend Config web/src/lib/connectors/connectors.ts UI schema for credential input
Test Patterns backend/tests/daily/connectors/confluence/ Example connector test implementation

Summary

  • Integrate Onyx with other tools using three patterns: custom connectors for ingestion, REST API for retrieval, and Celery workers for downstream publishing.
  • Implement connectors by subclassing PollConnector or LoadConnector in backend/onyx/connectors/, then register them in factory.py and add UI config in web/src/lib/connectors/connectors.ts.
  • Authenticate API requests using Bearer tokens in the Authorization header against endpoints defined in backend/onyx/server/api/.
  • Trigger manual indexing via POST /api/connector/{id}/run or wait for the Celery beat scheduler in backend/onyx/background/apps/beat.py to run check_for_indexing every 15 seconds.
  • Extend backend/onyx/background/apps/light.py to push document events to external webhooks.

Frequently Asked Questions

How do I authenticate API requests to Onyx?

Generate an API key through the administrative UI under Settings → API Keys, or use the default token specified in your .env file. Include this token in the Authorization header as Bearer <token> when calling endpoints like /api/search or /api/connector/{id}/run.

What is the difference between PollConnector and LoadConnector?

PollConnector is designed for incremental synchronization; you implement poll_source(start_ts, end_ts) to fetch only data modified within a specific timeframe. LoadConnector is used for full data loads through the load_from_state() method, typically for initial ingestion or sources without timestamp-based filtering. Both base classes are defined in backend/onyx/connectors/interfaces.py.

How do I trigger a manual re-index of a connector?

Send a POST request to /api/connector/{id}/run with a valid API key, or dispatch the task directly to the Celery docfetching queue. The platform also runs the check_for_indexing task automatically every 15 seconds via the beat worker in backend/onyx/background/apps/beat.py to handle scheduled refreshes.

Can I embed Onyx search into my existing web application?

Yes. You can iframe the hosted Next.js application (http://localhost:3000) or reuse the React component library located in web/src/components/. For programmatic integration, call the search API from your backend using the Python client pattern shown above and render the JSON results in your native UI.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →