# How to Ingest Data from Databricks into Semantica's Knowledge Graph

> Easily ingest data from Databricks into Semantica's knowledge graph. Learn how to use the DatabricksIngestor class to transform your data into Semantica documents efficiently.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-08

---

**TLDR:** Use the `DatabricksIngestor` class from [`semantica/ingest/databricks_ingestor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/databricks_ingestor.py) to extract tables or SQL queries from Databricks Unity Catalog and convert them into Semantica documents using `export_as_documents`.

Semantica provides a dedicated Databricks ingestion module within the `semantica-agi/semantica` repository that handles the complete lifecycle of extracting data from Databricks lakehouses and converting it into the internal document format used by the knowledge-graph pipeline. The implementation supports both Personal Access Token (PAT) and OAuth M2M authentication while managing connections to Databricks SQL warehouses and Unity Catalog.

## Core Ingestion Components

The ingestion architecture consists of three primary Python classes that work together to move data from Databricks Delta Lake into your knowledge graph.

### DatabricksConnector

The **`DatabricksConnector`** class manages low-level connections to Databricks SQL warehouses and the Unity Catalog workspace client. Located in [[`semantica/ingest/databricks_ingestor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/databricks_ingestor.py)](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/databricks_ingestor.py#L79-L133), this component verifies that `databricks-sdk` and `databricks-sql-connector` packages are installed and handles authentication automatically.

The connector reads credentials from environment variables—`DATABRICKS_HOST`, `DATABRICKS_TOKEN`, etc.—when not provided directly as arguments. It supports both Personal Access Token (PAT) and OAuth M2M (service-principal) authentication methods.

### DatabricksIngestor

The **`DatabricksIngestor`** class provides the high-level API for data extraction. Implemented in the same file (lines 305–617), it exposes methods including `ingest_table()`, `ingest_query()`, `list_catalogs()`, `list_schemas()`, `list_tables()`, `get_table_schema()`, and `get_table_lineage()`.

This class wraps raw query results in the `DatabricksData` dataclass and drives progress tracking for long-running jobs via the internal progress tracker utility.

### DatabricksData and Document Export

The **`DatabricksData`** dataclass (lines 64–77) serves as a container for rows, column lists, schema identifiers, and ingestion timestamps. Call `export_as_documents()` on the ingestor instance to convert `DatabricksData` into Semantica's generic document structure, attaching provenance metadata such as `source: "databricks"`, catalog, schema, and table names.

## Authentication Methods

Semantica supports two authentication patterns for Databricks connections, both handled transparently by the `DatabricksConnector`.

### Personal Access Token (PAT)

Pass the token directly to the `DatabricksIngestor` constructor along with the host and HTTP path:

```python
ingestor = DatabricksIngestor(
    host="https://adb-12345.azuredatabricks.net",
    token="dapi-xxxxxxxxxxxxxxxxxxxx",
    http_path="/sql/1.0/warehouses/1111111111111111",
    catalog="main",
    schema="default",
)

```

The connector passes this token directly to `databricks_sql.connect`.

### OAuth M2M (Service Principal)

For OAuth machine-to-machine authentication, the connector builds a `credentials_provider` callable using `databricks.sdk.core.Config` and `oauth_service_principal`. This is necessary because the SQL connector does not accept `client_id` and `client_secret` as plain keyword arguments. The bearer token is obtained automatically and refreshed as needed.

## Ingestion Workflow

Follow this architectural flow to ingest data from Databricks into Semantica's knowledge graph:

1. **Initialize the connector** – Create a `DatabricksIngestor` instance with your workspace credentials.
2. **Execute SQL or read tables** – Use `ingest_table()` for full-table extraction or `ingest_query()` for custom SQL with parameters.
3. **Fetch metadata (optional)** – Query Unity Catalog using `get_table_schema()` or `get_table_lineage()` to capture column-level lineage.
4. **Convert rows** – The ingestor coerces datetime and bytes objects to JSON-serializable primitives in `_convert_rows`.
5. **Export documents** – Call `export_as_documents()` to generate knowledge-graph ready documents with attached provenance.

## Practical Code Examples

### Basic Table Ingestion

Extract the first 10,000 rows from a Delta table and convert them to documents:

```python
from semantica.ingest import DatabricksIngestor

ingestor = DatabricksIngestor(
    host="https://adb-12345.azuredatabricks.net",
    token="dapi-xxxxxxxxxxxxxxxxxxxx",
    http_path="/sql/1.0/warehouses/1111111111111111",
    catalog="main",
    schema="default",
)

# Ingest with automatic limit

data = ingestor.ingest_table("customers", limit=10000)

# Convert to Semantica document format

documents = ingestor.export_as_documents(data, text_fields=["name", "notes"])

```

### Custom SQL with Parameters

Execute parameterized queries and handle large result sets with batching:

```python
query = """
SELECT customer_id, purchase_total, purchase_date
FROM sales.orders
WHERE purchase_date >= :start_date
"""

params = {"start_date": "2024-01-01"}

data = ingestor.ingest_query(query, parameters=params, batch_size=5000)
documents = ingestor.export_as_documents(data, text_fields=["purchase_total"])

```

### Metadata and Lineage Discovery

Explore Unity Catalog structure and fetch column-level lineage before ingestion:

```python
catalogs = ingestor.list_catalogs()
schemas = ingestor.list_schemas(catalog="analytics")
tables = ingestor.list_tables(catalog="analytics", schema="raw")

# Get schema details

schema_info = ingestor.get_table_schema(
    "events", 
    catalog="analytics", 
    schema="raw"
)

# Fetch lineage including column dependencies

lineage = ingestor.get_table_lineage(
    "events", 
    catalog="analytics", 
    schema="raw", 
    include_column_lineage=True
)

```

### Context Manager Pattern

Ensure connections close automatically using the context manager interface:

```python
with DatabricksIngestor(
    host="https://adb-12345.azuredatabricks.net",
    token="dapi-xxxxxxxxxxxxxxxxxxxx",
    http_path="/sql/1.0/warehouses/1111111111111111",
    catalog="main",
    schema="default"
) as ingestor:
    data = ingestor.ingest_table("orders")
    documents = ingestor.export_as_documents(data)

# Connection closed automatically here

```

## Summary

- The **`DatabricksIngestor`** class in [`semantica/ingest/databricks_ingestor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/databricks_ingestor.py) provides the primary interface for extracting data from Databricks Unity Catalog and Delta Lake.
- **Authentication** supports both Personal Access Tokens and OAuth M2M service principals, automatically reading from environment variables when available.
- **Document conversion** happens via `export_as_documents()`, which attaches provenance metadata and converts rows into the knowledge-graph format.
- **Metadata methods** like `get_table_lineage()` and `list_catalogs()` enable full discovery of Unity Catalog assets before ingestion.
- The implementation handles **connection pooling**, **progress tracking**, and **automatic resource cleanup** via context managers.

## Frequently Asked Questions

### How does Semantica handle large Databricks tables?

The `ingest_table()` and `ingest_query()` methods support pagination through the `batch_size` parameter, which streams results via the SQL cursor to prevent memory overflow. The `DatabricksIngestor` also integrates with [`semantica/utils/progress_tracker.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/utils/progress_tracker.py) to emit status updates during long-running extractions.

### What dependencies are required for Databricks ingestion?

You must install `databricks-sdk` and `databricks-sql-connector`. The `DatabricksConnector` class verifies these packages at initialization and raises an `ImportError` with specific installation instructions if they are missing.

### Can I ingest data from multiple catalogs in one session?

Yes. A single `DatabricksIngestor` instance can query any catalog accessible by the provided credentials. Simply change the `catalog` and `schema` parameters in method calls like `ingest_table()` or use fully-qualified table references (`catalog.schema.table`) in your SQL queries.

### How is data type conversion handled?

The ingestor automatically coerces non-serializable types in the `_convert_rows` method: `datetime` objects convert to ISO format strings, and `bytes` decode to UTF-8. This ensures the resulting documents are JSON-serializable for downstream vector stores.