How to Ingest Data from Databricks into Semantica's Knowledge Graph

TLDR: Use the DatabricksIngestor class from semantica/ingest/databricks_ingestor.py to extract tables or SQL queries from Databricks Unity Catalog and convert them into Semantica documents using export_as_documents.

Semantica provides a dedicated Databricks ingestion module within the semantica-agi/semantica repository that handles the complete lifecycle of extracting data from Databricks lakehouses and converting it into the internal document format used by the knowledge-graph pipeline. The implementation supports both Personal Access Token (PAT) and OAuth M2M authentication while managing connections to Databricks SQL warehouses and Unity Catalog.

Core Ingestion Components

The ingestion architecture consists of three primary Python classes that work together to move data from Databricks Delta Lake into your knowledge graph.

DatabricksConnector

The DatabricksConnector class manages low-level connections to Databricks SQL warehouses and the Unity Catalog workspace client. Located in [semantica/ingest/databricks_ingestor.py](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/databricks_ingestor.py#L79-L133), this component verifies that databricks-sdk and databricks-sql-connector packages are installed and handles authentication automatically.

The connector reads credentials from environment variables—DATABRICKS_HOST, DATABRICKS_TOKEN, etc.—when not provided directly as arguments. It supports both Personal Access Token (PAT) and OAuth M2M (service-principal) authentication methods.

DatabricksIngestor

The DatabricksIngestor class provides the high-level API for data extraction. Implemented in the same file (lines 305–617), it exposes methods including ingest_table(), ingest_query(), list_catalogs(), list_schemas(), list_tables(), get_table_schema(), and get_table_lineage().

This class wraps raw query results in the DatabricksData dataclass and drives progress tracking for long-running jobs via the internal progress tracker utility.

DatabricksData and Document Export

The DatabricksData dataclass (lines 64–77) serves as a container for rows, column lists, schema identifiers, and ingestion timestamps. Call export_as_documents() on the ingestor instance to convert DatabricksData into Semantica's generic document structure, attaching provenance metadata such as source: "databricks", catalog, schema, and table names.

Authentication Methods

Semantica supports two authentication patterns for Databricks connections, both handled transparently by the DatabricksConnector.

Personal Access Token (PAT)

Pass the token directly to the DatabricksIngestor constructor along with the host and HTTP path:

ingestor = DatabricksIngestor(
    host="https://adb-12345.azuredatabricks.net",
    token="dapi-xxxxxxxxxxxxxxxxxxxx",
    http_path="/sql/1.0/warehouses/1111111111111111",
    catalog="main",
    schema="default",
)

The connector passes this token directly to databricks_sql.connect.

OAuth M2M (Service Principal)

For OAuth machine-to-machine authentication, the connector builds a credentials_provider callable using databricks.sdk.core.Config and oauth_service_principal. This is necessary because the SQL connector does not accept client_id and client_secret as plain keyword arguments. The bearer token is obtained automatically and refreshed as needed.

Ingestion Workflow

Follow this architectural flow to ingest data from Databricks into Semantica's knowledge graph:

  1. Initialize the connector – Create a DatabricksIngestor instance with your workspace credentials.
  2. Execute SQL or read tables – Use ingest_table() for full-table extraction or ingest_query() for custom SQL with parameters.
  3. Fetch metadata (optional) – Query Unity Catalog using get_table_schema() or get_table_lineage() to capture column-level lineage.
  4. Convert rows – The ingestor coerces datetime and bytes objects to JSON-serializable primitives in _convert_rows.
  5. Export documents – Call export_as_documents() to generate knowledge-graph ready documents with attached provenance.

Practical Code Examples

Basic Table Ingestion

Extract the first 10,000 rows from a Delta table and convert them to documents:

from semantica.ingest import DatabricksIngestor

ingestor = DatabricksIngestor(
    host="https://adb-12345.azuredatabricks.net",
    token="dapi-xxxxxxxxxxxxxxxxxxxx",
    http_path="/sql/1.0/warehouses/1111111111111111",
    catalog="main",
    schema="default",
)

# Ingest with automatic limit

data = ingestor.ingest_table("customers", limit=10000)

# Convert to Semantica document format

documents = ingestor.export_as_documents(data, text_fields=["name", "notes"])

Custom SQL with Parameters

Execute parameterized queries and handle large result sets with batching:

query = """
SELECT customer_id, purchase_total, purchase_date
FROM sales.orders
WHERE purchase_date >= :start_date
"""

params = {"start_date": "2024-01-01"}

data = ingestor.ingest_query(query, parameters=params, batch_size=5000)
documents = ingestor.export_as_documents(data, text_fields=["purchase_total"])

Metadata and Lineage Discovery

Explore Unity Catalog structure and fetch column-level lineage before ingestion:

catalogs = ingestor.list_catalogs()
schemas = ingestor.list_schemas(catalog="analytics")
tables = ingestor.list_tables(catalog="analytics", schema="raw")

# Get schema details

schema_info = ingestor.get_table_schema(
    "events", 
    catalog="analytics", 
    schema="raw"
)

# Fetch lineage including column dependencies

lineage = ingestor.get_table_lineage(
    "events", 
    catalog="analytics", 
    schema="raw", 
    include_column_lineage=True
)

Context Manager Pattern

Ensure connections close automatically using the context manager interface:

with DatabricksIngestor(
    host="https://adb-12345.azuredatabricks.net",
    token="dapi-xxxxxxxxxxxxxxxxxxxx",
    http_path="/sql/1.0/warehouses/1111111111111111",
    catalog="main",
    schema="default"
) as ingestor:
    data = ingestor.ingest_table("orders")
    documents = ingestor.export_as_documents(data)

# Connection closed automatically here

Summary

  • The DatabricksIngestor class in semantica/ingest/databricks_ingestor.py provides the primary interface for extracting data from Databricks Unity Catalog and Delta Lake.
  • Authentication supports both Personal Access Tokens and OAuth M2M service principals, automatically reading from environment variables when available.
  • Document conversion happens via export_as_documents(), which attaches provenance metadata and converts rows into the knowledge-graph format.
  • Metadata methods like get_table_lineage() and list_catalogs() enable full discovery of Unity Catalog assets before ingestion.
  • The implementation handles connection pooling, progress tracking, and automatic resource cleanup via context managers.

Frequently Asked Questions

How does Semantica handle large Databricks tables?

The ingest_table() and ingest_query() methods support pagination through the batch_size parameter, which streams results via the SQL cursor to prevent memory overflow. The DatabricksIngestor also integrates with semantica/utils/progress_tracker.py to emit status updates during long-running extractions.

What dependencies are required for Databricks ingestion?

You must install databricks-sdk and databricks-sql-connector. The DatabricksConnector class verifies these packages at initialization and raises an ImportError with specific installation instructions if they are missing.

Can I ingest data from multiple catalogs in one session?

Yes. A single DatabricksIngestor instance can query any catalog accessible by the provided credentials. Simply change the catalog and schema parameters in method calls like ingest_table() or use fully-qualified table references (catalog.schema.table) in your SQL queries.

How is data type conversion handled?

The ingestor automatically coerces non-serializable types in the _convert_rows method: datetime objects convert to ISO format strings, and bytes decode to UTF-8. This ensures the resulting documents are JSON-serializable for downstream vector stores.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →