# OpenMetadata Architecture Components: A Deep Dive into the Four-Layer Design

> Explore the four-layer OpenMetadata architecture including schemas, storage, APIs, and ingestion. Understand the modular design powering modern metadata management.

- Repository: [OpenMetadata/OpenMetadata](https://github.com/open-metadata/OpenMetadata)
- Tags: deep-dive
- Published: 2026-04-23

---

**OpenMetadata follows a schema-first, four-layer architecture separating metadata schemas, storage, APIs, and ingestion into modular, extensible components.**

The OpenMetadata platform (available at `open-metadata/OpenMetadata`) provides a unified metadata management system supporting 80+ data connectors. Its architecture cleanly separates concerns across four core layers, enabling teams to extend, customize, and scale their metadata operations. This guide examines each component with concrete code references from the repository.

---

## Metadata Schemas: The Single Source of Truth

The **Metadata Schemas** layer defines the canonical data models that govern every entity in OpenMetadata. These JSON schemas establish a shared vocabulary for tables, dashboards, pipelines, users, and their relationships.

### Schema-First Code Generation

All schemas reside in `openmetadata-spec/src/main/resources/json/schema/`. The build process generates type-safe code across three languages from these definitions:

| Language | Generated Output | Example Class |
|----------|-----------------|---------------|
| Java | POJOs | `org.openmetadata.schema.entity.data.Table` |
| Python | Pydantic models | `metadata.generated.schema.entity.data.table.Table` |
| TypeScript | Interfaces | [`generated/entity/data/table.ts`](https://github.com/open-metadata/OpenMetadata/blob/main/generated/entity/data/table.ts) |

### Connection Schema Example

The MySQL connector schema demonstrates how connection parameters are defined:

```json
// openmetadata-spec/src/main/resources/json/schema/entity/services/connections/mysqlConnection.json
{
  "$id": "mysqlConnection",
  "type": "object",
  "properties": {
    "username": { "type": "string" },
    "password": { "type": "string" },
    "hostPort": { "type": "string" },
    "database": { "type": "string" }
  },
  "required": ["username", "password", "hostPort"]
}

```

Running `make generate` propagates this schema to all three language targets, ensuring perfect type consistency across the stack.

---

## Metadata Store: Graph- Oriented Persistence

The **Metadata Store** persists the metadata graph, lineage relationships, tags, ownership records, and governance artifacts. OpenMetadata supports **PostgreSQL** and **MySQL** as backing stores, using a relational schema optimized for graph traversal.

### Database Schema and Migrations

The store's structure is defined through Flyway migrations in `bootstrap/sql/`:

```sql
-- Example migration creating the table entity store
CREATE TABLE IF NOT EXISTS table_entity (
    id VARCHAR(36) GENERATED ALWAYS AS (json ->> 'id') STORED PRIMARY KEY,
    name VARCHAR(256) GENERATED ALWAYS AS (json ->> 'name') STORED,
    json JSON NOT NULL,
    updatedAt BIGINT GENERATED ALWAYS AS ((json ->> 'updatedAt')::bigint) STORED,
    updatedBy VARCHAR(256) GENERATED ALWAYS AS (json ->> 'updatedBy') STORED,
    deleted BOOLEAN GENERATED ALWAYS AS ((json ->> 'deleted')::boolean) STORED
);

```

The **JDBI DAO layer** in `openmetadata-service/src/main/java/org/openmetadata/service/jdbi/` provides type-safe access to these tables, abstracting SQL operations behind Java interfaces.

### Data Access Pattern

The store uses a **JSON-first** approach: entities are stored as JSON documents with generated columns extracting frequently queried fields for indexing. This balances schema flexibility with query performance.

---

## Metadata APIs: RESTful Service Layer

The **Metadata APIs** expose all platform capabilities through REST endpoints built on **Dropwizard**. These services handle CRUD operations, full-text search, lineage queries, and governance workflows.

### API Structure and Resources

Controllers are organized by entity type in `openmetadata-service/src/main/java/org/openmetadata/service/resources/`:

| Resource Package | Handles |
|-----------------|---------|
| `databases/` | Tables, databases, database schemas |
| `dashboards/` | Dashboard services, dashboards, charts |
| `pipelines/` | Pipeline services, pipelines, tasks |
| `search/` | Elasticsearch/OpenSearch integration |
| `lineage/` | Lineage graph queries and mutations |

### Table Resource Example

The [`TableResource.java`](https://github.com/open-metadata/OpenMetadata/blob/main/TableResource.java) controller demonstrates standard API patterns:

```java
// From openmetadata-service/src/main/java/org/openmetadata/service/resources/databases/TableResource.java

@Path("/v1/tables")
@Tag(name = "Tables", description = "APIs for managing table entities.")
public class TableResource extends EntityResource<Table, TableRepository> {

    @POST
    @Operation(
        summary = "Create a table",
        description = "Create a new table under an existing database schema."
    )
    public Response create(
        @Context UriInfo uriInfo,
        @Context SecurityContext securityContext,
        @Valid CreateTable create
    ) {
        Table table = getTable(create, securityContext.getUserPrincipal().getName());
        return create(uriInfo, securityContext, table);
    }
}

```

### API Authentication and Authorization

All endpoints enforce **JWT-based authentication** and **policy-based authorization** using OpenMetadata's attribute-based access control (ABAC) system. The `securityContext` parameter propagates user identity through the call stack.

---

## Ingestion Framework: Pluggable Metadata Pipeline Engine

The **Ingestion Framework** is a Python-based engine that extracts metadata from external systems and writes it to the Metadata Store via the APIs. It supports **80+ connectors** spanning databases, data warehouses, dashboards, BI tools, pipeline orchestrators, and messaging systems.

### Core Workflow Architecture

The ingestion pipeline follows a **source → processor → sink** pattern defined in [`ingestion/src/metadata/workflow/metadata.py`](https://github.com/open-metadata/OpenMetadata/blob/main/ingestion/src/metadata/workflow/metadata.py):

```python

# From ingestion/src/metadata/workflow/metadata.py

class MetadataWorkflow(BaseWorkflow):
    """
    Core workflow for metadata ingestion.
    """
    
    def __init__(self, config: WorkflowConfig):
        self.config = config
        self.source: Source = self._build_source()
        self.sink: Sink = self._build_sink()
        self.processor: Optional[Processor] = self._build_processor()
    
    def execute(self) -> None:
        """Execute the full ingestion pipeline."""
        for record in self.source.extract():
            # Transform through optional processor chain

            if self.processor:
                record = self.processor.process(record)
            
            # Push to metadata store via REST API

            self.sink.write(record)

```

### Connector Implementation Structure

Connectors are organized by system type in `ingestion/src/metadata/ingestion/source/`:

| Directory | Connector Examples |
|-----------|------------------|
| `database/` | MySQL, PostgreSQL, Snowflake, BigQuery, Redshift |
| `dashboard/` | Tableau, Looker, PowerBI, Metabase, Superset |
| `pipeline/` | Airflow, Dagster, Prefect, Fivetran, dbt |
| `messaging/` | Kafka, Pulsar, Redpanda |

### Schema-First Connector Design

All connectors inherit from base classes that enforce the **schema-first pattern**. The MySQL connector demonstrates this implementation:

```python

# Simplified excerpt from ingestion/src/metadata/ingestion/source/database/mysql/metadata.py

from metadata.generated.schema.entity.data.table import Table
from metadata.generated.schema.entity.services.connections.mysqlConnection import (
    MysqlConnection,
)
from metadata.ingestion.source.database.common_db_source import CommonDbSource

class MysqlSource(CommonDbSource):
    """
    MySQL metadata extractor following the schema-first pattern.
    """
    
    connection: MysqlConnection  # Type-safe Pydantic model from generated schema

    
    def __init__(self, config: WorkflowConfig, metadata_config: OpenMetadataConnection):
        super().__init__(config, metadata_config)
        self.connection = self.service_connection  # Pydantic-validated connection

    
    def query_table_names(self) -> Iterable[str]:
        """Query MySQL information_schema for table names."""
        with self.connection.client() as client:
            result = client.execute(
                "SELECT table_name FROM information_schema.tables WHERE table_schema = %s",
                (self.connection.database,)
            )
            return [row[0] for row in result]
    
    def yield_table(self, table_name: str) -> Iterable[Table]:
        """
        Construct a Table entity matching the canonical schema.
        Returns a Pydantic Table model ready for API ingestion.
        """
        columns = self._get_columns(table_name)
        table = Table(
            name=table_name,
            database=self.context.database,
            columns=columns,
            description=self._get_table_comment(table_name)
        )
        yield table

```

The `Table` model returned by `yield_table()` is the **same Pydantic class** generated from the JSON schema, ensuring end-to-end type consistency from MySQL's `information_schema` through to the REST API payload.

---

## Summary

OpenMetadata's four-layer architecture delivers a **schema-first, extensible platform** for unified metadata management:

- **Metadata Schemas** (`openmetadata-spec/`) — JSON definitions that generate type-safe code across Java, Python, and TypeScript
- **Metadata Store** (`bootstrap/sql/`) — Relational graph storage using PostgreSQL or MySQL with Flyway migrations
- **Metadata APIs** (`openmetadata-service/`) — Dropwizard REST services exposing CRUD, search, lineage, and governance operations
- **Ingestion Framework** (`ingestion/`) — Python pipeline engine with 80+ connectors following a common source → processor → sink pattern

This separation enables teams to extend connectors, customize schemas, or integrate with existing infrastructure without modifying core platform code.

---

## Frequently Asked Questions

### What database does OpenMetadata use for storing metadata?

OpenMetadata supports **PostgreSQL** and **MySQL** as the primary metadata store. The database schema is managed through **Flyway migrations** located in `bootstrap/sql/`, with separate directories for PostgreSQL and MySQL dialects. The store uses a **JSON-first design** where entities are stored as JSON documents with generated columns for indexed fields.

### How does OpenMetadata ensure type safety across different programming languages?

OpenMetadata uses a **schema-first code generation** approach. Canonical **JSON schemas** in `openmetadata-spec/src/main/resources/json/schema/` serve as the single source of truth. The build process generates **Java POJOs**, **Python Pydantic models**, and **TypeScript interfaces** from these schemas. This guarantees that a `Table` entity in the Python ingestion framework, Java API service, and React UI all share identical field definitions and validation rules.

### What is the ingestion workflow pattern in OpenMetadata?

The ingestion framework follows a **source → processor → sink** pipeline pattern implemented in [`ingestion/src/metadata/workflow/metadata.py`](https://github.com/open-metadata/OpenMetadata/blob/main/ingestion/src/metadata/workflow/metadata.py). The **source** extracts metadata from external systems (databases, dashboards, pipelines), an optional **processor** transforms records, and the **sink** writes to the OpenMetadata REST API. All 80+ connectors in `ingestion/src/metadata/ingestion/source/` inherit from common base classes that enforce this pattern and the schema-first design.

### How are the REST APIs organized in OpenMetadata?

REST controllers are organized by **entity type** in `openmetadata-service/src/main/java/org/openmetadata/service/resources/`. Key packages include `databases/` for tables and schemas, `dashboards/` for BI tools, `pipelines/` for orchestrators, `search/` for Elasticsearch integration, and `lineage/` for data lineage operations. All controllers extend `EntityResource` and enforce **JWT authentication** with **policy-based authorization** through the `SecurityContext` parameter.