How to Add a New Connector to OpenMetadata: A Complete Developer Guide

To add a new connector to OpenMetadata, you implement a Python Source class that yields metadata entities, define a JSON Schema for configuration, and register the connector with the ingestion framework—enabling automatic UI, CLI, and CI integration.

OpenMetadata's ingestion framework uses a schema-first, plug-in architecture that supports databases, dashboards, pipelines, messaging systems, and more. Each connector lives as a self-contained Python module with clear interfaces for connection handling, entity generation, and capability declaration. This guide walks through the complete workflow based on the actual source code structure in the open-metadata/OpenMetadata repository.

Choose Your Service Type and Base Class

OpenMetadata categorizes connectors by service type, each with a dedicated base class in ingestion/src/metadata/ingestion/source/:

Service Type Example Sources Base Class
Database MySQL, Postgres, Snowflake DatabaseServiceSource (via CommonDbSourceService)
Dashboard Tableau, Looker DashboardServiceSource
Pipeline Airflow, Dagster PipelineServiceSource
Messaging Kafka, Redpanda MessagingServiceSource
ML Model SageMaker, MLflow MlModelServiceSource
Storage S3, GCS StorageServiceSource

The MySQL connector demonstrates this pattern in ingestion/src/metadata/ingestion/source/database/mysql/metadata.py, where it implements DatabaseServiceSource to yield table and column metadata. For implementation standards, reference skills/standards/main.md in the repository.

Scaffold the Connector Structure

OpenMetadata provides a CLI helper that generates boilerplate files for new connectors. This accelerates development and ensures consistent directory structure.

Run the scaffold command in your development environment:

./run.sh scaffold-connector \
    --service-type <service_type> \
    --name <connector_name>

Parameters:

  • service_type — One of the serviceType enum values: database, dashboard, pipeline, messaging, mlmodel, storage
  • connector_name — Lower-case identifier, e.g., mydb, customdash

The scaffolder creates this directory structure under ingestion/src/metadata/ingestion/source/<service_type>/<connector_name>/:

File Purpose
connection.py Low-level client/SDK wrapper with connection testing
metadata.py Source class that yields OpenMetadata entities
service_spec.py ServiceSpec for framework registration
models.py (optional) Pydantic models for connector-specific configuration

The scaffolder logic is documented in skills/commands/scaffold-connector.md.

Define the JSON Schema for Configuration

All connectors require a JSON Schema that drives both UI form generation and configuration validation. These schemas live in openmetadata-spec/src/main/resources/json/schema/services/<service_type>/.

Create a new schema file named <connector_name>.json. For example, MySQL's schema is at openmetadata-spec/src/main/resources/json/schema/services/database/mysql.json.

Key implementation requirements:

  • Use $ref to reference common authentication blocks (OAuth, username/password, SSL) rather than redefining them
  • Declare connector-specific properties such as includeViews, maxResults, or custom timeout values
  • Follow the connection type patterns documented in skills/connector-building/references/connection-type-guide.md

After defining the schema, run the generator to produce Pydantic models:

make generate

This outputs Python models to ingestion/src/metadata/ingestion/source/<service_type>/<connector_name>/models.py, ensuring type-safe configuration handling throughout the connector.

Implement the ServiceSpec

The service_spec.py file registers your connector with the ingestion engine and declares its capabilities. This class bridges configuration, connection, and source implementation.

Create service_spec.py with this structure:

from pydantic import Field
from metadata.ingestion.source.database.common_db_source import DatabaseServiceSource
from metadata.ingestion.source.database.mypg.connection import MyPgConnection
from metadata.ingestion.source.database.mypg.metadata import MyPgSource
from metadata.ingestion.api.common import WorkflowSource

class MyPgSpec(WorkflowSource):
    config: MyPgConnection = Field(..., description="MyPG connector config")
    serviceName: str = Field(..., description="Service name")
    sourceType: str = "MyPg"
    
    # Capability flags — enable only for implemented features

    supportsMetadataExtraction: bool = True
    supportsUsage: bool = True
    supportsLineage: bool = False

Required fields:

  • config — Pydantic model for connection parameters
  • serviceName — Display name for service instances
  • sourceType — Unique identifier matching schema name
  • Capability booleans — Controls which pipeline stages execute

Export the spec class in __init__.py for framework discovery. The complete ServiceSpec template is in skills/standards/service_spec.md.

Build the Connection Wrapper

The connection.py module provides low-level client access and connection health verification. It accepts the generated Pydantic config and exposes standard methods for the ingestion framework.

For SQLAlchemy-based database connectors, implement this pattern:

from sqlalchemy import create_engine
from metadata.ingestion.source.database.common_db_source import CommonDbSourceService

class MyPgConnection(CommonDbSourceService):
    def get_connection(self):
        url = self.get_connection_url()
        return create_engine(url, **self.connection_options)

    def test_connection(self):
        """Validate connectivity before pipeline execution"""
        engine = self.get_connection()
        with engine.connect() as conn:
            conn.execute("SELECT 1")

Key methods:

  • get_connection() — Returns live client object (engine, SDK client, etc.)
  • test_connection() — Lightweight validation called by CLI --validate flag

For reference, see the MySQL implementation in ingestion/src/metadata/ingestion/source/database/mysql/connection.py.

Implement the Source Class for Entity Generation

The metadata.py module contains the Source class that transforms external system metadata into OpenMetadata entities. This class inherits from the service-type base and implements generator methods for each supported entity type.

For database connectors, override yield_table to emit Table entities:

from typing import Iterable
from metadata.ingestion.source.database.common_db_source import CommonDbSourceService
from metadata.generated.schema.entity.data.table import Table

class MyPgSource(CommonDbSourceService):
    def yield_table(self, schema_name: str) -> Iterable[Table]:
        engine = self.connection.get_connection()
        with engine.connect() as conn:
            rows = conn.execute(
                f"""
                SELECT tablename 
                FROM pg_catalog.pg_tables 
                WHERE schemaname = '{schema_name}'
                """
            )
            for row in rows:
                yield Table(
                    name=row["tablename"],
                    database=self.context.get_database(),
                    schema=self.context.get_schema(),
                )

Generator methods to implement by service type:

  • Database: yield_table, yield_view, yield_lineage (optional)
  • Dashboard: yield_dashboard, yield_chart
  • Pipeline: yield_pipeline, yield_task
  • Messaging: yield_topic, yield_subscription

The MySQL source in ingestion/src/metadata/ingestion/source/database/mysql/metadata.py demonstrates complete implementation patterns for all supported operations.

Register the Connector with the Framework

Two registration steps enable discovery across the platform:

1. Enum Registration

Add the connector name to the serviceType enum in openmetadata-spec/src/main/resources/json/schema/services/service_type.json. This populates the Add Service wizard dropdown in the UI.

2. Python Discovery Import

Export the ServiceSpec in ingestion/src/metadata/ingestion/source/__init__.py:

from metadata.ingestion.source.database.mypg.service_spec import MyPgSpec

The complete registration procedure is documented in skills/standards/registration.md.

Add Unit and Integration Tests

OpenMetadata requires real-behavior tests over heavy mocking. Test patterns vary by connector type:

CLI Tests

Subclass CliDBBase.TestSuite for database connectors or CliCommonDB.TestSuite for CommonDbSourceService implementations:


# ingestion/tests/unit/cli_mypg/test_cli_mypg.py

from ingestion.tests.unit.test_cli_db import CliDBBase

class TestMyPg(CliDBBase.TestSuite):
    @classmethod
    def get_connector_name(cls) -> str:
        return "mypg"

Copy from ingestion/tests/unit/cli_mysql/test_cli_mysql.py as a template.

Source Tests

Instantiate the Source with a mock connection (e.g., SQLite in-memory engine) and assert yield_* returns expected entities:

def test_yield_table():
    source = MyPgSource(config, metadata)
    tables = list(source.yield_table("public"))
    assert len(tables) > 0
    assert all(isinstance(t, Table) for t in tables)

Update CI for End-to-End Validation

Add your connector to the GitHub Actions matrix in .github/workflows/py-cli-e2e-tests.yml:

strategy:
  matrix:
    database:
      - mysql
      - postgres
      - mypg  # your new connector

This ensures nightly runs validate your connector against real infrastructure.

Verify UI Integration

Complete the development cycle by confirming UI functionality:

  1. Start the UI locally: yarn start in openmetadata-ui/...
  2. Navigate to Settings → Services → Add Service
  3. Confirm your connector appears in the service type dropdown
  4. Test configuration form rendering for all custom fields
  5. Verify save operations persist correctly

Summary

Adding a new connector to OpenMetadata follows a structured, schema-first approach:

  • Select the service type and inherit from the appropriate base class (DatabaseServiceSource, DashboardServiceSource, etc.)
  • Scaffold with CLI using ./run.sh scaffold-connector to generate consistent boilerplate
  • Define JSON Schema in openmetadata-spec/ for UI-driven configuration
  • Implement three core Python modules: connection.py (client wrapper), metadata.py (entity generators), and service_spec.py (capability registration)
  • Register the connector in both the serviceType enum and Python discovery imports
  • Add tests using real-behavior patterns with CliDBBase or CliCommonDB test suites
  • Enable CI validation by adding to .github/workflows/py-cli-e2e-tests.yml
  • Verify UI integration through local testing of the Add Service wizard

Frequently Asked Questions

What is the fastest way to start building a new OpenMetadata connector?

Use the scaffold CLI command: ./run.sh scaffold-connector --service-type database --name <your_connector>. This generates the complete directory structure including connection.py, metadata.py, service_spec.py, and test boilerplate, ensuring you follow the canonical pattern from the start.

Can I add a connector for a custom internal data source not listed in OpenMetadata's service types?

Yes. The framework supports any source that can be queried programmatically. Select the closest service type (e.g., database for SQL-based sources, pipeline for orchestration tools) and implement the required generator methods. For truly novel categories, you may need to extend the base class hierarchy in ingestion/src/metadata/ingestion/source/.

How does OpenMetadata validate connector configurations before runtime?

Configuration validation uses JSON Schema definitions in openmetadata-spec/. When you run make generate, Pydantic models are created from these schemas. The ingestion framework validates incoming YAML configurations against these models at startup. Additionally, the test_connection() method in your connection.py provides live connectivity verification through the CLI --validate flag.

What testing requirements must I meet before submitting a connector PR?

OpenMetadata requires real-behavior tests over mocking. For database connectors, subclass CliDBBase.TestSuite or CliCommonDB.TestSuite and implement get_connector_name(). Include source-level tests that instantiate your Source with a mock connection (e.g., SQLite in-memory) and assert that yield_* generators return expected entities. Finally, add your connector to the CI matrix in .github/workflows/py-cli-e2e-tests.yml for nightly end-to-end validation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →