# How to Add a New Connector to OpenMetadata: A Complete Developer Guide

> Learn how to add a new connector to OpenMetadata with this developer guide. Implement a Python Source class and JSON Schema for seamless UI, CLI, and CI integration.

- Repository: [OpenMetadata/OpenMetadata](https://github.com/open-metadata/OpenMetadata)
- Tags: how-to-guide
- Published: 2026-04-23

---

**To add a new connector to OpenMetadata, you implement a Python Source class that yields metadata entities, define a JSON Schema for configuration, and register the connector with the ingestion framework—enabling automatic UI, CLI, and CI integration.**

OpenMetadata's ingestion framework uses a **schema-first, plug-in architecture** that supports databases, dashboards, pipelines, messaging systems, and more. Each connector lives as a self-contained Python module with clear interfaces for connection handling, entity generation, and capability declaration. This guide walks through the complete workflow based on the actual source code structure in the `open-metadata/OpenMetadata` repository.

## Choose Your Service Type and Base Class

OpenMetadata categorizes connectors by service type, each with a dedicated base class in `ingestion/src/metadata/ingestion/source/`:

| Service Type | Example Sources | Base Class |
|-------------|-----------------|------------|
| **Database** | MySQL, Postgres, Snowflake | `DatabaseServiceSource` (via `CommonDbSourceService`) |
| **Dashboard** | Tableau, Looker | `DashboardServiceSource` |
| **Pipeline** | Airflow, Dagster | `PipelineServiceSource` |
| **Messaging** | Kafka, Redpanda | `MessagingServiceSource` |
| **ML Model** | SageMaker, MLflow | `MlModelServiceSource` |
| **Storage** | S3, GCS | `StorageServiceSource` |

The MySQL connector demonstrates this pattern in [`ingestion/src/metadata/ingestion/source/database/mysql/metadata.py`](https://github.com/open-metadata/OpenMetadata/blob/main/ingestion/src/metadata/ingestion/source/database/mysql/metadata.py), where it implements `DatabaseServiceSource` to yield table and column metadata. For implementation standards, reference [`skills/standards/main.md`](https://github.com/open-metadata/OpenMetadata/blob/main/skills/standards/main.md) in the repository.

## Scaffold the Connector Structure

OpenMetadata provides a CLI helper that generates boilerplate files for new connectors. This accelerates development and ensures consistent directory structure.

Run the scaffold command in your development environment:

```bash
./run.sh scaffold-connector \
    --service-type <service_type> \
    --name <connector_name>

```

Parameters:
- `service_type` — One of the `serviceType` enum values: `database`, `dashboard`, `pipeline`, `messaging`, `mlmodel`, `storage`
- `connector_name` — Lower-case identifier, e.g., `mydb`, `customdash`

The scaffolder creates this directory structure under `ingestion/src/metadata/ingestion/source/<service_type>/<connector_name>/`:

| File | Purpose |
|------|---------|
| [`connection.py`](https://github.com/open-metadata/OpenMetadata/blob/main/connection.py) | Low-level client/SDK wrapper with connection testing |
| [`metadata.py`](https://github.com/open-metadata/OpenMetadata/blob/main/metadata.py) | Source class that yields OpenMetadata entities |
| [`service_spec.py`](https://github.com/open-metadata/OpenMetadata/blob/main/service_spec.py) | ServiceSpec for framework registration |
| [`models.py`](https://github.com/open-metadata/OpenMetadata/blob/main/models.py) (optional) | Pydantic models for connector-specific configuration |

The scaffolder logic is documented in [`skills/commands/scaffold-connector.md`](https://github.com/open-metadata/OpenMetadata/blob/main/skills/commands/scaffold-connector.md).

## Define the JSON Schema for Configuration

All connectors require a **JSON Schema** that drives both UI form generation and configuration validation. These schemas live in `openmetadata-spec/src/main/resources/json/schema/services/<service_type>/`.

Create a new schema file named `<connector_name>.json`. For example, MySQL's schema is at [`openmetadata-spec/src/main/resources/json/schema/services/database/mysql.json`](https://github.com/open-metadata/OpenMetadata/blob/main/openmetadata-spec/src/main/resources/json/schema/services/database/mysql.json).

Key implementation requirements:

- Use `$ref` to reference common authentication blocks (OAuth, username/password, SSL) rather than redefining them
- Declare connector-specific properties such as `includeViews`, `maxResults`, or custom timeout values
- Follow the connection type patterns documented in [`skills/connector-building/references/connection-type-guide.md`](https://github.com/open-metadata/OpenMetadata/blob/main/skills/connector-building/references/connection-type-guide.md)

After defining the schema, run the generator to produce Pydantic models:

```bash
make generate

```

This outputs Python models to `ingestion/src/metadata/ingestion/source/<service_type>/<connector_name>/models.py`, ensuring type-safe configuration handling throughout the connector.

## Implement the ServiceSpec

The [`service_spec.py`](https://github.com/open-metadata/OpenMetadata/blob/main/service_spec.py) file registers your connector with the ingestion engine and declares its capabilities. This class bridges configuration, connection, and source implementation.

Create [`service_spec.py`](https://github.com/open-metadata/OpenMetadata/blob/main/service_spec.py) with this structure:

```python
from pydantic import Field
from metadata.ingestion.source.database.common_db_source import DatabaseServiceSource
from metadata.ingestion.source.database.mypg.connection import MyPgConnection
from metadata.ingestion.source.database.mypg.metadata import MyPgSource
from metadata.ingestion.api.common import WorkflowSource

class MyPgSpec(WorkflowSource):
    config: MyPgConnection = Field(..., description="MyPG connector config")
    serviceName: str = Field(..., description="Service name")
    sourceType: str = "MyPg"
    
    # Capability flags — enable only for implemented features

    supportsMetadataExtraction: bool = True
    supportsUsage: bool = True
    supportsLineage: bool = False

```

Required fields:
- `config` — Pydantic model for connection parameters
- `serviceName` — Display name for service instances
- `sourceType` — Unique identifier matching schema name
- Capability booleans — Controls which pipeline stages execute

Export the spec class in [`__init__.py`](https://github.com/open-metadata/OpenMetadata/blob/main/__init__.py) for framework discovery. The complete ServiceSpec template is in [`skills/standards/service_spec.md`](https://github.com/open-metadata/OpenMetadata/blob/main/skills/standards/service_spec.md).

## Build the Connection Wrapper

The [`connection.py`](https://github.com/open-metadata/OpenMetadata/blob/main/connection.py) module provides low-level client access and connection health verification. It accepts the generated Pydantic config and exposes standard methods for the ingestion framework.

For SQLAlchemy-based database connectors, implement this pattern:

```python
from sqlalchemy import create_engine
from metadata.ingestion.source.database.common_db_source import CommonDbSourceService

class MyPgConnection(CommonDbSourceService):
    def get_connection(self):
        url = self.get_connection_url()
        return create_engine(url, **self.connection_options)

    def test_connection(self):
        """Validate connectivity before pipeline execution"""
        engine = self.get_connection()
        with engine.connect() as conn:
            conn.execute("SELECT 1")

```

Key methods:
- `get_connection()` — Returns live client object (engine, SDK client, etc.)
- `test_connection()` — Lightweight validation called by CLI `--validate` flag

For reference, see the MySQL implementation in [`ingestion/src/metadata/ingestion/source/database/mysql/connection.py`](https://github.com/open-metadata/OpenMetadata/blob/main/ingestion/src/metadata/ingestion/source/database/mysql/connection.py).

## Implement the Source Class for Entity Generation

The [`metadata.py`](https://github.com/open-metadata/OpenMetadata/blob/main/metadata.py) module contains the **Source** class that transforms external system metadata into OpenMetadata entities. This class inherits from the service-type base and implements generator methods for each supported entity type.

For database connectors, override `yield_table` to emit `Table` entities:

```python
from typing import Iterable
from metadata.ingestion.source.database.common_db_source import CommonDbSourceService
from metadata.generated.schema.entity.data.table import Table

class MyPgSource(CommonDbSourceService):
    def yield_table(self, schema_name: str) -> Iterable[Table]:
        engine = self.connection.get_connection()
        with engine.connect() as conn:
            rows = conn.execute(
                f"""
                SELECT tablename 
                FROM pg_catalog.pg_tables 
                WHERE schemaname = '{schema_name}'
                """
            )
            for row in rows:
                yield Table(
                    name=row["tablename"],
                    database=self.context.get_database(),
                    schema=self.context.get_schema(),
                )

```

Generator methods to implement by service type:
- **Database**: `yield_table`, `yield_view`, `yield_lineage` (optional)
- **Dashboard**: `yield_dashboard`, `yield_chart`
- **Pipeline**: `yield_pipeline`, `yield_task`
- **Messaging**: `yield_topic`, `yield_subscription`

The MySQL source in [`ingestion/src/metadata/ingestion/source/database/mysql/metadata.py`](https://github.com/open-metadata/OpenMetadata/blob/main/ingestion/src/metadata/ingestion/source/database/mysql/metadata.py) demonstrates complete implementation patterns for all supported operations.

## Register the Connector with the Framework

Two registration steps enable discovery across the platform:

### 1. Enum Registration

Add the connector name to the `serviceType` enum in [`openmetadata-spec/src/main/resources/json/schema/services/service_type.json`](https://github.com/open-metadata/OpenMetadata/blob/main/openmetadata-spec/src/main/resources/json/schema/services/service_type.json). This populates the **Add Service** wizard dropdown in the UI.

### 2. Python Discovery Import

Export the ServiceSpec in [`ingestion/src/metadata/ingestion/source/__init__.py`](https://github.com/open-metadata/OpenMetadata/blob/main/ingestion/src/metadata/ingestion/source/__init__.py):

```python
from metadata.ingestion.source.database.mypg.service_spec import MyPgSpec

```

The complete registration procedure is documented in [`skills/standards/registration.md`](https://github.com/open-metadata/OpenMetadata/blob/main/skills/standards/registration.md).

## Add Unit and Integration Tests

OpenMetadata requires **real-behavior tests** over heavy mocking. Test patterns vary by connector type:

### CLI Tests

Subclass `CliDBBase.TestSuite` for database connectors or `CliCommonDB.TestSuite` for `CommonDbSourceService` implementations:

```python

# ingestion/tests/unit/cli_mypg/test_cli_mypg.py

from ingestion.tests.unit.test_cli_db import CliDBBase

class TestMyPg(CliDBBase.TestSuite):
    @classmethod
    def get_connector_name(cls) -> str:
        return "mypg"

```

Copy from [`ingestion/tests/unit/cli_mysql/test_cli_mysql.py`](https://github.com/open-metadata/OpenMetadata/blob/main/ingestion/tests/unit/cli_mysql/test_cli_mysql.py) as a template.

### Source Tests

Instantiate the Source with a mock connection (e.g., SQLite in-memory engine) and assert `yield_*` returns expected entities:

```python
def test_yield_table():
    source = MyPgSource(config, metadata)
    tables = list(source.yield_table("public"))
    assert len(tables) > 0
    assert all(isinstance(t, Table) for t in tables)

```

## Update CI for End-to-End Validation

Add your connector to the GitHub Actions matrix in [`.github/workflows/py-cli-e2e-tests.yml`](https://github.com/open-metadata/OpenMetadata/blob/main/.github/workflows/py-cli-e2e-tests.yml):

```yaml
strategy:
  matrix:
    database:
      - mysql
      - postgres
      - mypg  # your new connector

```

This ensures nightly runs validate your connector against real infrastructure.

## Verify UI Integration

Complete the development cycle by confirming UI functionality:

1. Start the UI locally: `yarn start` in `openmetadata-ui/...`
2. Navigate to **Settings → Services → Add Service**
3. Confirm your connector appears in the service type dropdown
4. Test configuration form rendering for all custom fields
5. Verify save operations persist correctly

## Summary

Adding a new connector to OpenMetadata follows a structured, schema-first approach:

- **Select the service type** and inherit from the appropriate base class (`DatabaseServiceSource`, `DashboardServiceSource`, etc.)
- **Scaffold with CLI** using `./run.sh scaffold-connector` to generate consistent boilerplate
- **Define JSON Schema** in `openmetadata-spec/` for UI-driven configuration
- **Implement three core Python modules**: [`connection.py`](https://github.com/open-metadata/OpenMetadata/blob/main/connection.py) (client wrapper), [`metadata.py`](https://github.com/open-metadata/OpenMetadata/blob/main/metadata.py) (entity generators), and [`service_spec.py`](https://github.com/open-metadata/OpenMetadata/blob/main/service_spec.py) (capability registration)
- **Register the connector** in both the `serviceType` enum and Python discovery imports
- **Add tests** using real-behavior patterns with `CliDBBase` or `CliCommonDB` test suites
- **Enable CI validation** by adding to [`.github/workflows/py-cli-e2e-tests.yml`](https://github.com/open-metadata/OpenMetadata/blob/main/.github/workflows/py-cli-e2e-tests.yml)
- **Verify UI integration** through local testing of the Add Service wizard

## Frequently Asked Questions

### What is the fastest way to start building a new OpenMetadata connector?

Use the scaffold CLI command: `./run.sh scaffold-connector --service-type database --name <your_connector>`. This generates the complete directory structure including [`connection.py`](https://github.com/open-metadata/OpenMetadata/blob/main/connection.py), [`metadata.py`](https://github.com/open-metadata/OpenMetadata/blob/main/metadata.py), [`service_spec.py`](https://github.com/open-metadata/OpenMetadata/blob/main/service_spec.py), and test boilerplate, ensuring you follow the canonical pattern from the start.

### Can I add a connector for a custom internal data source not listed in OpenMetadata's service types?

Yes. The framework supports any source that can be queried programmatically. Select the closest service type (e.g., `database` for SQL-based sources, `pipeline` for orchestration tools) and implement the required generator methods. For truly novel categories, you may need to extend the base class hierarchy in `ingestion/src/metadata/ingestion/source/`.

### How does OpenMetadata validate connector configurations before runtime?

Configuration validation uses JSON Schema definitions in `openmetadata-spec/`. When you run `make generate`, Pydantic models are created from these schemas. The ingestion framework validates incoming YAML configurations against these models at startup. Additionally, the `test_connection()` method in your [`connection.py`](https://github.com/open-metadata/OpenMetadata/blob/main/connection.py) provides live connectivity verification through the CLI `--validate` flag.

### What testing requirements must I meet before submitting a connector PR?

OpenMetadata requires real-behavior tests over mocking. For database connectors, subclass `CliDBBase.TestSuite` or `CliCommonDB.TestSuite` and implement `get_connector_name()`. Include source-level tests that instantiate your Source with a mock connection (e.g., SQLite in-memory) and assert that `yield_*` generators return expected entities. Finally, add your connector to the CI matrix in [`.github/workflows/py-cli-e2e-tests.yml`](https://github.com/open-metadata/OpenMetadata/blob/main/.github/workflows/py-cli-e2e-tests.yml) for nightly end-to-end validation.