How to Use `get_table_lineage()` in Semantica with Unity Catalog
get_table_lineage() queries Databricks Unity Catalog REST endpoints to fetch upstream and downstream table relationships, returning a dictionary with dependency lists and optional column-level mappings.
The Semantica library provides first-class integration with Databricks Unity Catalog for automated data lineage discovery. By calling get_table_lineage() from the DatabricksIngestor class, you can programmatically trace data flow across your lakehouse architecture without manually parsing SQL or audit logs.
Understanding the get_table_lineage() Implementation
Source Location and Architecture
The get_table_lineage() method is implemented in semantica/ingest/databricks_ingestor.py (lines 825-860). According to the Semantica source code, this method constructs authenticated requests to the Unity Catalog REST API, validating the provided schema before executing the lineage query.
Method Signature and Return Structure
The method accepts a fully-qualified table name and an optional boolean flag for column granularity:
table_name: String in formatcatalog.schema.table(e.g.,main.default.customers)include_column_lineage: Boolean (defaultFalse) to retrieve per-column upstream/downstream mappings
The function returns a dictionary containing:
upstream: List of fully-qualified table names that feed data into the targetdownstream: List of fully-qualified table names that consume data from the targetcolumns(conditional): Wheninclude_column_lineage=True, a mapping of column names to their respectiveupstreamanddownstreamdependencies
Prerequisites: Authenticating with Databricks
Before invoking get_table_lineage(), you must initialize a WorkspaceClient with valid credentials for your Unity Catalog-enabled workspace. The client class is exposed through semantica.integrations.databricks.
from semantica.integrations.databricks import WorkspaceClient
ws_client = WorkspaceClient(
host="https://adb-<workspace-id>.azuredatabricks.net",
token="YOUR_DATABRICKS_PERSONAL_ACCESS_TOKEN"
)
Retrieving Table Lineage
Basic Table-Level Lineage
Instantiate the DatabricksIngestor with your authenticated client and call get_table_lineage() using the three-part table identifier:
from semantica.ingest.databricks_ingestor import DatabricksIngestor
ingestor = DatabricksIngestor(ws_client)
lineage = ingestor.get_table_lineage("my_catalog.my_schema.my_table")
print("Upstream:", lineage["upstream"])
print("Downstream:", lineage["downstream"])
As implemented in the test suite at tests/test_databricks_ingestor.py (lines 649-744), the method validates that the catalog and schema exist before invoking the Unity Catalog API.
Column-Level Lineage Retrieval
To capture granular dependencies at the column level, set include_column_lineage=True:
col_lineage = ingestor.get_table_lineage(
"my_catalog.my_schema.my_table",
include_column_lineage=True
)
# Access column-specific upstream dependencies
print(col_lineage["columns"]["id"]["upstream"])
The column-level mapping logic is exercised in the unit tests at lines 679-738 of tests/test_databricks_ingestor.py, verifying that individual column references resolve to their source tables.
Schema Validation and Error Handling
The ingestor performs strict validation of the table identifier before contacting Unity Catalog. According to the validation tests at lines 743-744 in tests/test_databricks_ingestor.py, the method checks that the specified catalog and schema exist in the workspace, raising a ValueError if the three-part name is malformed or if the table is not registered in Unity Catalog.
Summary
get_table_lineage()resides insemantica/ingest/databricks_ingestor.pyand interfaces directly with Databricks Unity Catalog REST endpoints.- You must provide fully-qualified table names using the
catalog.schema.tableformat. - The method returns dictionaries containing
upstreamanddownstreamtable lists, with optionalcolumnsmappings wheninclude_column_lineage=True. - Schema validation occurs automatically before API invocation, as verified in the unit test suite.
- Authentication requires a
WorkspaceClientinitialized with valid Databricks host and token credentials.
Frequently Asked Questions
What authentication does get_table_lineage() require?
You must pass an authenticated WorkspaceClient instance to the DatabricksIngestor constructor. The client handles OAuth token negotiation or personal access token validation with your Unity Catalog workspace, as defined in integrations/databricks/__init__.py.
Does the method support column-level data lineage?
Yes. When you pass include_column_lineage=True to get_table_lineage(), the response includes a columns dictionary mapping each column name to its specific upstream and downstream dependencies. This functionality is tested in tests/test_databricks_ingestor.py (lines 679-738).
What table naming convention should I use?
Always use the three-part Unity Catalog identifier: catalog.schema.table. The method validates this format against your workspace metadata before executing the REST call, ensuring the catalog and schema exist as confirmed by the validation logic at lines 743-744 of the test file.
Can I use get_table_lineage() with non-Unity Catalog metastores?
No. The implementation specifically targets Unity Catalog REST endpoints. If you attempt to query tables stored in legacy Hive metastores without Unity Catalog registration, the schema validation step will fail and raise a ValueError indicating the table is not found.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →