How to Use `get_table_lineage()` in Semantica with Unity Catalog

get_table_lineage() queries Databricks Unity Catalog REST endpoints to fetch upstream and downstream table relationships, returning a dictionary with dependency lists and optional column-level mappings.

The Semantica library provides first-class integration with Databricks Unity Catalog for automated data lineage discovery. By calling get_table_lineage() from the DatabricksIngestor class, you can programmatically trace data flow across your lakehouse architecture without manually parsing SQL or audit logs.

Understanding the get_table_lineage() Implementation

Source Location and Architecture

The get_table_lineage() method is implemented in semantica/ingest/databricks_ingestor.py (lines 825-860). According to the Semantica source code, this method constructs authenticated requests to the Unity Catalog REST API, validating the provided schema before executing the lineage query.

Method Signature and Return Structure

The method accepts a fully-qualified table name and an optional boolean flag for column granularity:

  • table_name: String in format catalog.schema.table (e.g., main.default.customers)
  • include_column_lineage: Boolean (default False) to retrieve per-column upstream/downstream mappings

The function returns a dictionary containing:

  • upstream: List of fully-qualified table names that feed data into the target
  • downstream: List of fully-qualified table names that consume data from the target
  • columns (conditional): When include_column_lineage=True, a mapping of column names to their respective upstream and downstream dependencies

Prerequisites: Authenticating with Databricks

Before invoking get_table_lineage(), you must initialize a WorkspaceClient with valid credentials for your Unity Catalog-enabled workspace. The client class is exposed through semantica.integrations.databricks.

from semantica.integrations.databricks import WorkspaceClient

ws_client = WorkspaceClient(
    host="https://adb-<workspace-id>.azuredatabricks.net",
    token="YOUR_DATABRICKS_PERSONAL_ACCESS_TOKEN"
)

Retrieving Table Lineage

Basic Table-Level Lineage

Instantiate the DatabricksIngestor with your authenticated client and call get_table_lineage() using the three-part table identifier:

from semantica.ingest.databricks_ingestor import DatabricksIngestor

ingestor = DatabricksIngestor(ws_client)
lineage = ingestor.get_table_lineage("my_catalog.my_schema.my_table")

print("Upstream:", lineage["upstream"])
print("Downstream:", lineage["downstream"])

As implemented in the test suite at tests/test_databricks_ingestor.py (lines 649-744), the method validates that the catalog and schema exist before invoking the Unity Catalog API.

Column-Level Lineage Retrieval

To capture granular dependencies at the column level, set include_column_lineage=True:

col_lineage = ingestor.get_table_lineage(
    "my_catalog.my_schema.my_table",
    include_column_lineage=True
)

# Access column-specific upstream dependencies

print(col_lineage["columns"]["id"]["upstream"])

The column-level mapping logic is exercised in the unit tests at lines 679-738 of tests/test_databricks_ingestor.py, verifying that individual column references resolve to their source tables.

Schema Validation and Error Handling

The ingestor performs strict validation of the table identifier before contacting Unity Catalog. According to the validation tests at lines 743-744 in tests/test_databricks_ingestor.py, the method checks that the specified catalog and schema exist in the workspace, raising a ValueError if the three-part name is malformed or if the table is not registered in Unity Catalog.

Summary

  • get_table_lineage() resides in semantica/ingest/databricks_ingestor.py and interfaces directly with Databricks Unity Catalog REST endpoints.
  • You must provide fully-qualified table names using the catalog.schema.table format.
  • The method returns dictionaries containing upstream and downstream table lists, with optional columns mappings when include_column_lineage=True.
  • Schema validation occurs automatically before API invocation, as verified in the unit test suite.
  • Authentication requires a WorkspaceClient initialized with valid Databricks host and token credentials.

Frequently Asked Questions

What authentication does get_table_lineage() require?

You must pass an authenticated WorkspaceClient instance to the DatabricksIngestor constructor. The client handles OAuth token negotiation or personal access token validation with your Unity Catalog workspace, as defined in integrations/databricks/__init__.py.

Does the method support column-level data lineage?

Yes. When you pass include_column_lineage=True to get_table_lineage(), the response includes a columns dictionary mapping each column name to its specific upstream and downstream dependencies. This functionality is tested in tests/test_databricks_ingestor.py (lines 679-738).

What table naming convention should I use?

Always use the three-part Unity Catalog identifier: catalog.schema.table. The method validates this format against your workspace metadata before executing the REST call, ensuring the catalog and schema exist as confirmed by the validation logic at lines 743-744 of the test file.

Can I use get_table_lineage() with non-Unity Catalog metastores?

No. The implementation specifically targets Unity Catalog REST endpoints. If you attempt to query tables stored in legacy Hive metastores without Unity Catalog registration, the schema validation step will fail and raise a ValueError indicating the table is not found.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →