OpenMetadata Architecture Components: A Deep Dive into the Four-Layer Design

OpenMetadata follows a schema-first, four-layer architecture separating metadata schemas, storage, APIs, and ingestion into modular, extensible components.

The OpenMetadata platform (available at open-metadata/OpenMetadata) provides a unified metadata management system supporting 80+ data connectors. Its architecture cleanly separates concerns across four core layers, enabling teams to extend, customize, and scale their metadata operations. This guide examines each component with concrete code references from the repository.


Metadata Schemas: The Single Source of Truth

The Metadata Schemas layer defines the canonical data models that govern every entity in OpenMetadata. These JSON schemas establish a shared vocabulary for tables, dashboards, pipelines, users, and their relationships.

Schema-First Code Generation

All schemas reside in openmetadata-spec/src/main/resources/json/schema/. The build process generates type-safe code across three languages from these definitions:

Language Generated Output Example Class
Java POJOs org.openmetadata.schema.entity.data.Table
Python Pydantic models metadata.generated.schema.entity.data.table.Table
TypeScript Interfaces generated/entity/data/table.ts

Connection Schema Example

The MySQL connector schema demonstrates how connection parameters are defined:

// openmetadata-spec/src/main/resources/json/schema/entity/services/connections/mysqlConnection.json
{
  "$id": "mysqlConnection",
  "type": "object",
  "properties": {
    "username": { "type": "string" },
    "password": { "type": "string" },
    "hostPort": { "type": "string" },
    "database": { "type": "string" }
  },
  "required": ["username", "password", "hostPort"]
}

Running make generate propagates this schema to all three language targets, ensuring perfect type consistency across the stack.


Metadata Store: Graph- Oriented Persistence

The Metadata Store persists the metadata graph, lineage relationships, tags, ownership records, and governance artifacts. OpenMetadata supports PostgreSQL and MySQL as backing stores, using a relational schema optimized for graph traversal.

Database Schema and Migrations

The store's structure is defined through Flyway migrations in bootstrap/sql/:

-- Example migration creating the table entity store
CREATE TABLE IF NOT EXISTS table_entity (
    id VARCHAR(36) GENERATED ALWAYS AS (json ->> 'id') STORED PRIMARY KEY,
    name VARCHAR(256) GENERATED ALWAYS AS (json ->> 'name') STORED,
    json JSON NOT NULL,
    updatedAt BIGINT GENERATED ALWAYS AS ((json ->> 'updatedAt')::bigint) STORED,
    updatedBy VARCHAR(256) GENERATED ALWAYS AS (json ->> 'updatedBy') STORED,
    deleted BOOLEAN GENERATED ALWAYS AS ((json ->> 'deleted')::boolean) STORED
);

The JDBI DAO layer in openmetadata-service/src/main/java/org/openmetadata/service/jdbi/ provides type-safe access to these tables, abstracting SQL operations behind Java interfaces.

Data Access Pattern

The store uses a JSON-first approach: entities are stored as JSON documents with generated columns extracting frequently queried fields for indexing. This balances schema flexibility with query performance.


Metadata APIs: RESTful Service Layer

The Metadata APIs expose all platform capabilities through REST endpoints built on Dropwizard. These services handle CRUD operations, full-text search, lineage queries, and governance workflows.

API Structure and Resources

Controllers are organized by entity type in openmetadata-service/src/main/java/org/openmetadata/service/resources/:

Resource Package Handles
databases/ Tables, databases, database schemas
dashboards/ Dashboard services, dashboards, charts
pipelines/ Pipeline services, pipelines, tasks
search/ Elasticsearch/OpenSearch integration
lineage/ Lineage graph queries and mutations

Table Resource Example

The TableResource.java controller demonstrates standard API patterns:

// From openmetadata-service/src/main/java/org/openmetadata/service/resources/databases/TableResource.java

@Path("/v1/tables")
@Tag(name = "Tables", description = "APIs for managing table entities.")
public class TableResource extends EntityResource<Table, TableRepository> {

    @POST
    @Operation(
        summary = "Create a table",
        description = "Create a new table under an existing database schema."
    )
    public Response create(
        @Context UriInfo uriInfo,
        @Context SecurityContext securityContext,
        @Valid CreateTable create
    ) {
        Table table = getTable(create, securityContext.getUserPrincipal().getName());
        return create(uriInfo, securityContext, table);
    }
}

API Authentication and Authorization

All endpoints enforce JWT-based authentication and policy-based authorization using OpenMetadata's attribute-based access control (ABAC) system. The securityContext parameter propagates user identity through the call stack.


Ingestion Framework: Pluggable Metadata Pipeline Engine

The Ingestion Framework is a Python-based engine that extracts metadata from external systems and writes it to the Metadata Store via the APIs. It supports 80+ connectors spanning databases, data warehouses, dashboards, BI tools, pipeline orchestrators, and messaging systems.

Core Workflow Architecture

The ingestion pipeline follows a source → processor → sink pattern defined in ingestion/src/metadata/workflow/metadata.py:


# From ingestion/src/metadata/workflow/metadata.py

class MetadataWorkflow(BaseWorkflow):
    """
    Core workflow for metadata ingestion.
    """
    
    def __init__(self, config: WorkflowConfig):
        self.config = config
        self.source: Source = self._build_source()
        self.sink: Sink = self._build_sink()
        self.processor: Optional[Processor] = self._build_processor()
    
    def execute(self) -> None:
        """Execute the full ingestion pipeline."""
        for record in self.source.extract():
            # Transform through optional processor chain

            if self.processor:
                record = self.processor.process(record)
            
            # Push to metadata store via REST API

            self.sink.write(record)

Connector Implementation Structure

Connectors are organized by system type in ingestion/src/metadata/ingestion/source/:

Directory Connector Examples
database/ MySQL, PostgreSQL, Snowflake, BigQuery, Redshift
dashboard/ Tableau, Looker, PowerBI, Metabase, Superset
pipeline/ Airflow, Dagster, Prefect, Fivetran, dbt
messaging/ Kafka, Pulsar, Redpanda

Schema-First Connector Design

All connectors inherit from base classes that enforce the schema-first pattern. The MySQL connector demonstrates this implementation:


# Simplified excerpt from ingestion/src/metadata/ingestion/source/database/mysql/metadata.py

from metadata.generated.schema.entity.data.table import Table
from metadata.generated.schema.entity.services.connections.mysqlConnection import (
    MysqlConnection,
)
from metadata.ingestion.source.database.common_db_source import CommonDbSource

class MysqlSource(CommonDbSource):
    """
    MySQL metadata extractor following the schema-first pattern.
    """
    
    connection: MysqlConnection  # Type-safe Pydantic model from generated schema

    
    def __init__(self, config: WorkflowConfig, metadata_config: OpenMetadataConnection):
        super().__init__(config, metadata_config)
        self.connection = self.service_connection  # Pydantic-validated connection

    
    def query_table_names(self) -> Iterable[str]:
        """Query MySQL information_schema for table names."""
        with self.connection.client() as client:
            result = client.execute(
                "SELECT table_name FROM information_schema.tables WHERE table_schema = %s",
                (self.connection.database,)
            )
            return [row[0] for row in result]
    
    def yield_table(self, table_name: str) -> Iterable[Table]:
        """
        Construct a Table entity matching the canonical schema.
        Returns a Pydantic Table model ready for API ingestion.
        """
        columns = self._get_columns(table_name)
        table = Table(
            name=table_name,
            database=self.context.database,
            columns=columns,
            description=self._get_table_comment(table_name)
        )
        yield table

The Table model returned by yield_table() is the same Pydantic class generated from the JSON schema, ensuring end-to-end type consistency from MySQL's information_schema through to the REST API payload.


Summary

OpenMetadata's four-layer architecture delivers a schema-first, extensible platform for unified metadata management:

  • Metadata Schemas (openmetadata-spec/) — JSON definitions that generate type-safe code across Java, Python, and TypeScript
  • Metadata Store (bootstrap/sql/) — Relational graph storage using PostgreSQL or MySQL with Flyway migrations
  • Metadata APIs (openmetadata-service/) — Dropwizard REST services exposing CRUD, search, lineage, and governance operations
  • Ingestion Framework (ingestion/) — Python pipeline engine with 80+ connectors following a common source → processor → sink pattern

This separation enables teams to extend connectors, customize schemas, or integrate with existing infrastructure without modifying core platform code.


Frequently Asked Questions

What database does OpenMetadata use for storing metadata?

OpenMetadata supports PostgreSQL and MySQL as the primary metadata store. The database schema is managed through Flyway migrations located in bootstrap/sql/, with separate directories for PostgreSQL and MySQL dialects. The store uses a JSON-first design where entities are stored as JSON documents with generated columns for indexed fields.

How does OpenMetadata ensure type safety across different programming languages?

OpenMetadata uses a schema-first code generation approach. Canonical JSON schemas in openmetadata-spec/src/main/resources/json/schema/ serve as the single source of truth. The build process generates Java POJOs, Python Pydantic models, and TypeScript interfaces from these schemas. This guarantees that a Table entity in the Python ingestion framework, Java API service, and React UI all share identical field definitions and validation rules.

What is the ingestion workflow pattern in OpenMetadata?

The ingestion framework follows a source → processor → sink pipeline pattern implemented in ingestion/src/metadata/workflow/metadata.py. The source extracts metadata from external systems (databases, dashboards, pipelines), an optional processor transforms records, and the sink writes to the OpenMetadata REST API. All 80+ connectors in ingestion/src/metadata/ingestion/source/ inherit from common base classes that enforce this pattern and the schema-first design.

How are the REST APIs organized in OpenMetadata?

REST controllers are organized by entity type in openmetadata-service/src/main/java/org/openmetadata/service/resources/. Key packages include databases/ for tables and schemas, dashboards/ for BI tools, pipelines/ for orchestrators, search/ for Elasticsearch integration, and lineage/ for data lineage operations. All controllers extend EntityResource and enforce JWT authentication with policy-based authorization through the SecurityContext parameter.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →