OpenMetadata Architecture Components: A Deep Dive into the Four-Layer Design
OpenMetadata follows a schema-first, four-layer architecture separating metadata schemas, storage, APIs, and ingestion into modular, extensible components.
The OpenMetadata platform (available at open-metadata/OpenMetadata) provides a unified metadata management system supporting 80+ data connectors. Its architecture cleanly separates concerns across four core layers, enabling teams to extend, customize, and scale their metadata operations. This guide examines each component with concrete code references from the repository.
Metadata Schemas: The Single Source of Truth
The Metadata Schemas layer defines the canonical data models that govern every entity in OpenMetadata. These JSON schemas establish a shared vocabulary for tables, dashboards, pipelines, users, and their relationships.
Schema-First Code Generation
All schemas reside in openmetadata-spec/src/main/resources/json/schema/. The build process generates type-safe code across three languages from these definitions:
| Language | Generated Output | Example Class |
|---|---|---|
| Java | POJOs | org.openmetadata.schema.entity.data.Table |
| Python | Pydantic models | metadata.generated.schema.entity.data.table.Table |
| TypeScript | Interfaces | generated/entity/data/table.ts |
Connection Schema Example
The MySQL connector schema demonstrates how connection parameters are defined:
// openmetadata-spec/src/main/resources/json/schema/entity/services/connections/mysqlConnection.json
{
"$id": "mysqlConnection",
"type": "object",
"properties": {
"username": { "type": "string" },
"password": { "type": "string" },
"hostPort": { "type": "string" },
"database": { "type": "string" }
},
"required": ["username", "password", "hostPort"]
}
Running make generate propagates this schema to all three language targets, ensuring perfect type consistency across the stack.
Metadata Store: Graph- Oriented Persistence
The Metadata Store persists the metadata graph, lineage relationships, tags, ownership records, and governance artifacts. OpenMetadata supports PostgreSQL and MySQL as backing stores, using a relational schema optimized for graph traversal.
Database Schema and Migrations
The store's structure is defined through Flyway migrations in bootstrap/sql/:
-- Example migration creating the table entity store
CREATE TABLE IF NOT EXISTS table_entity (
id VARCHAR(36) GENERATED ALWAYS AS (json ->> 'id') STORED PRIMARY KEY,
name VARCHAR(256) GENERATED ALWAYS AS (json ->> 'name') STORED,
json JSON NOT NULL,
updatedAt BIGINT GENERATED ALWAYS AS ((json ->> 'updatedAt')::bigint) STORED,
updatedBy VARCHAR(256) GENERATED ALWAYS AS (json ->> 'updatedBy') STORED,
deleted BOOLEAN GENERATED ALWAYS AS ((json ->> 'deleted')::boolean) STORED
);
The JDBI DAO layer in openmetadata-service/src/main/java/org/openmetadata/service/jdbi/ provides type-safe access to these tables, abstracting SQL operations behind Java interfaces.
Data Access Pattern
The store uses a JSON-first approach: entities are stored as JSON documents with generated columns extracting frequently queried fields for indexing. This balances schema flexibility with query performance.
Metadata APIs: RESTful Service Layer
The Metadata APIs expose all platform capabilities through REST endpoints built on Dropwizard. These services handle CRUD operations, full-text search, lineage queries, and governance workflows.
API Structure and Resources
Controllers are organized by entity type in openmetadata-service/src/main/java/org/openmetadata/service/resources/:
| Resource Package | Handles |
|---|---|
databases/ |
Tables, databases, database schemas |
dashboards/ |
Dashboard services, dashboards, charts |
pipelines/ |
Pipeline services, pipelines, tasks |
search/ |
Elasticsearch/OpenSearch integration |
lineage/ |
Lineage graph queries and mutations |
Table Resource Example
The TableResource.java controller demonstrates standard API patterns:
// From openmetadata-service/src/main/java/org/openmetadata/service/resources/databases/TableResource.java
@Path("/v1/tables")
@Tag(name = "Tables", description = "APIs for managing table entities.")
public class TableResource extends EntityResource<Table, TableRepository> {
@POST
@Operation(
summary = "Create a table",
description = "Create a new table under an existing database schema."
)
public Response create(
@Context UriInfo uriInfo,
@Context SecurityContext securityContext,
@Valid CreateTable create
) {
Table table = getTable(create, securityContext.getUserPrincipal().getName());
return create(uriInfo, securityContext, table);
}
}
API Authentication and Authorization
All endpoints enforce JWT-based authentication and policy-based authorization using OpenMetadata's attribute-based access control (ABAC) system. The securityContext parameter propagates user identity through the call stack.
Ingestion Framework: Pluggable Metadata Pipeline Engine
The Ingestion Framework is a Python-based engine that extracts metadata from external systems and writes it to the Metadata Store via the APIs. It supports 80+ connectors spanning databases, data warehouses, dashboards, BI tools, pipeline orchestrators, and messaging systems.
Core Workflow Architecture
The ingestion pipeline follows a source → processor → sink pattern defined in ingestion/src/metadata/workflow/metadata.py:
# From ingestion/src/metadata/workflow/metadata.py
class MetadataWorkflow(BaseWorkflow):
"""
Core workflow for metadata ingestion.
"""
def __init__(self, config: WorkflowConfig):
self.config = config
self.source: Source = self._build_source()
self.sink: Sink = self._build_sink()
self.processor: Optional[Processor] = self._build_processor()
def execute(self) -> None:
"""Execute the full ingestion pipeline."""
for record in self.source.extract():
# Transform through optional processor chain
if self.processor:
record = self.processor.process(record)
# Push to metadata store via REST API
self.sink.write(record)
Connector Implementation Structure
Connectors are organized by system type in ingestion/src/metadata/ingestion/source/:
| Directory | Connector Examples |
|---|---|
database/ |
MySQL, PostgreSQL, Snowflake, BigQuery, Redshift |
dashboard/ |
Tableau, Looker, PowerBI, Metabase, Superset |
pipeline/ |
Airflow, Dagster, Prefect, Fivetran, dbt |
messaging/ |
Kafka, Pulsar, Redpanda |
Schema-First Connector Design
All connectors inherit from base classes that enforce the schema-first pattern. The MySQL connector demonstrates this implementation:
# Simplified excerpt from ingestion/src/metadata/ingestion/source/database/mysql/metadata.py
from metadata.generated.schema.entity.data.table import Table
from metadata.generated.schema.entity.services.connections.mysqlConnection import (
MysqlConnection,
)
from metadata.ingestion.source.database.common_db_source import CommonDbSource
class MysqlSource(CommonDbSource):
"""
MySQL metadata extractor following the schema-first pattern.
"""
connection: MysqlConnection # Type-safe Pydantic model from generated schema
def __init__(self, config: WorkflowConfig, metadata_config: OpenMetadataConnection):
super().__init__(config, metadata_config)
self.connection = self.service_connection # Pydantic-validated connection
def query_table_names(self) -> Iterable[str]:
"""Query MySQL information_schema for table names."""
with self.connection.client() as client:
result = client.execute(
"SELECT table_name FROM information_schema.tables WHERE table_schema = %s",
(self.connection.database,)
)
return [row[0] for row in result]
def yield_table(self, table_name: str) -> Iterable[Table]:
"""
Construct a Table entity matching the canonical schema.
Returns a Pydantic Table model ready for API ingestion.
"""
columns = self._get_columns(table_name)
table = Table(
name=table_name,
database=self.context.database,
columns=columns,
description=self._get_table_comment(table_name)
)
yield table
The Table model returned by yield_table() is the same Pydantic class generated from the JSON schema, ensuring end-to-end type consistency from MySQL's information_schema through to the REST API payload.
Summary
OpenMetadata's four-layer architecture delivers a schema-first, extensible platform for unified metadata management:
- Metadata Schemas (
openmetadata-spec/) — JSON definitions that generate type-safe code across Java, Python, and TypeScript - Metadata Store (
bootstrap/sql/) — Relational graph storage using PostgreSQL or MySQL with Flyway migrations - Metadata APIs (
openmetadata-service/) — Dropwizard REST services exposing CRUD, search, lineage, and governance operations - Ingestion Framework (
ingestion/) — Python pipeline engine with 80+ connectors following a common source → processor → sink pattern
This separation enables teams to extend connectors, customize schemas, or integrate with existing infrastructure without modifying core platform code.
Frequently Asked Questions
What database does OpenMetadata use for storing metadata?
OpenMetadata supports PostgreSQL and MySQL as the primary metadata store. The database schema is managed through Flyway migrations located in bootstrap/sql/, with separate directories for PostgreSQL and MySQL dialects. The store uses a JSON-first design where entities are stored as JSON documents with generated columns for indexed fields.
How does OpenMetadata ensure type safety across different programming languages?
OpenMetadata uses a schema-first code generation approach. Canonical JSON schemas in openmetadata-spec/src/main/resources/json/schema/ serve as the single source of truth. The build process generates Java POJOs, Python Pydantic models, and TypeScript interfaces from these schemas. This guarantees that a Table entity in the Python ingestion framework, Java API service, and React UI all share identical field definitions and validation rules.
What is the ingestion workflow pattern in OpenMetadata?
The ingestion framework follows a source → processor → sink pipeline pattern implemented in ingestion/src/metadata/workflow/metadata.py. The source extracts metadata from external systems (databases, dashboards, pipelines), an optional processor transforms records, and the sink writes to the OpenMetadata REST API. All 80+ connectors in ingestion/src/metadata/ingestion/source/ inherit from common base classes that enforce this pattern and the schema-first design.
How are the REST APIs organized in OpenMetadata?
REST controllers are organized by entity type in openmetadata-service/src/main/java/org/openmetadata/service/resources/. Key packages include databases/ for tables and schemas, dashboards/ for BI tools, pipelines/ for orchestrators, search/ for Elasticsearch integration, and lineage/ for data lineage operations. All controllers extend EntityResource and enforce JWT authentication with policy-based authorization through the SecurityContext parameter.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →