How to Add a New Connector to OpenMetadata: A Complete Developer Guide
To add a new connector to OpenMetadata, you implement a Python Source class that yields metadata entities, define a JSON Schema for configuration, and register the connector with the ingestion framework—enabling automatic UI, CLI, and CI integration.
OpenMetadata's ingestion framework uses a schema-first, plug-in architecture that supports databases, dashboards, pipelines, messaging systems, and more. Each connector lives as a self-contained Python module with clear interfaces for connection handling, entity generation, and capability declaration. This guide walks through the complete workflow based on the actual source code structure in the open-metadata/OpenMetadata repository.
Choose Your Service Type and Base Class
OpenMetadata categorizes connectors by service type, each with a dedicated base class in ingestion/src/metadata/ingestion/source/:
| Service Type | Example Sources | Base Class |
|---|---|---|
| Database | MySQL, Postgres, Snowflake | DatabaseServiceSource (via CommonDbSourceService) |
| Dashboard | Tableau, Looker | DashboardServiceSource |
| Pipeline | Airflow, Dagster | PipelineServiceSource |
| Messaging | Kafka, Redpanda | MessagingServiceSource |
| ML Model | SageMaker, MLflow | MlModelServiceSource |
| Storage | S3, GCS | StorageServiceSource |
The MySQL connector demonstrates this pattern in ingestion/src/metadata/ingestion/source/database/mysql/metadata.py, where it implements DatabaseServiceSource to yield table and column metadata. For implementation standards, reference skills/standards/main.md in the repository.
Scaffold the Connector Structure
OpenMetadata provides a CLI helper that generates boilerplate files for new connectors. This accelerates development and ensures consistent directory structure.
Run the scaffold command in your development environment:
./run.sh scaffold-connector \
--service-type <service_type> \
--name <connector_name>
Parameters:
service_type— One of theserviceTypeenum values:database,dashboard,pipeline,messaging,mlmodel,storageconnector_name— Lower-case identifier, e.g.,mydb,customdash
The scaffolder creates this directory structure under ingestion/src/metadata/ingestion/source/<service_type>/<connector_name>/:
| File | Purpose |
|---|---|
connection.py |
Low-level client/SDK wrapper with connection testing |
metadata.py |
Source class that yields OpenMetadata entities |
service_spec.py |
ServiceSpec for framework registration |
models.py (optional) |
Pydantic models for connector-specific configuration |
The scaffolder logic is documented in skills/commands/scaffold-connector.md.
Define the JSON Schema for Configuration
All connectors require a JSON Schema that drives both UI form generation and configuration validation. These schemas live in openmetadata-spec/src/main/resources/json/schema/services/<service_type>/.
Create a new schema file named <connector_name>.json. For example, MySQL's schema is at openmetadata-spec/src/main/resources/json/schema/services/database/mysql.json.
Key implementation requirements:
- Use
$refto reference common authentication blocks (OAuth, username/password, SSL) rather than redefining them - Declare connector-specific properties such as
includeViews,maxResults, or custom timeout values - Follow the connection type patterns documented in
skills/connector-building/references/connection-type-guide.md
After defining the schema, run the generator to produce Pydantic models:
make generate
This outputs Python models to ingestion/src/metadata/ingestion/source/<service_type>/<connector_name>/models.py, ensuring type-safe configuration handling throughout the connector.
Implement the ServiceSpec
The service_spec.py file registers your connector with the ingestion engine and declares its capabilities. This class bridges configuration, connection, and source implementation.
Create service_spec.py with this structure:
from pydantic import Field
from metadata.ingestion.source.database.common_db_source import DatabaseServiceSource
from metadata.ingestion.source.database.mypg.connection import MyPgConnection
from metadata.ingestion.source.database.mypg.metadata import MyPgSource
from metadata.ingestion.api.common import WorkflowSource
class MyPgSpec(WorkflowSource):
config: MyPgConnection = Field(..., description="MyPG connector config")
serviceName: str = Field(..., description="Service name")
sourceType: str = "MyPg"
# Capability flags — enable only for implemented features
supportsMetadataExtraction: bool = True
supportsUsage: bool = True
supportsLineage: bool = False
Required fields:
config— Pydantic model for connection parametersserviceName— Display name for service instancessourceType— Unique identifier matching schema name- Capability booleans — Controls which pipeline stages execute
Export the spec class in __init__.py for framework discovery. The complete ServiceSpec template is in skills/standards/service_spec.md.
Build the Connection Wrapper
The connection.py module provides low-level client access and connection health verification. It accepts the generated Pydantic config and exposes standard methods for the ingestion framework.
For SQLAlchemy-based database connectors, implement this pattern:
from sqlalchemy import create_engine
from metadata.ingestion.source.database.common_db_source import CommonDbSourceService
class MyPgConnection(CommonDbSourceService):
def get_connection(self):
url = self.get_connection_url()
return create_engine(url, **self.connection_options)
def test_connection(self):
"""Validate connectivity before pipeline execution"""
engine = self.get_connection()
with engine.connect() as conn:
conn.execute("SELECT 1")
Key methods:
get_connection()— Returns live client object (engine, SDK client, etc.)test_connection()— Lightweight validation called by CLI--validateflag
For reference, see the MySQL implementation in ingestion/src/metadata/ingestion/source/database/mysql/connection.py.
Implement the Source Class for Entity Generation
The metadata.py module contains the Source class that transforms external system metadata into OpenMetadata entities. This class inherits from the service-type base and implements generator methods for each supported entity type.
For database connectors, override yield_table to emit Table entities:
from typing import Iterable
from metadata.ingestion.source.database.common_db_source import CommonDbSourceService
from metadata.generated.schema.entity.data.table import Table
class MyPgSource(CommonDbSourceService):
def yield_table(self, schema_name: str) -> Iterable[Table]:
engine = self.connection.get_connection()
with engine.connect() as conn:
rows = conn.execute(
f"""
SELECT tablename
FROM pg_catalog.pg_tables
WHERE schemaname = '{schema_name}'
"""
)
for row in rows:
yield Table(
name=row["tablename"],
database=self.context.get_database(),
schema=self.context.get_schema(),
)
Generator methods to implement by service type:
- Database:
yield_table,yield_view,yield_lineage(optional) - Dashboard:
yield_dashboard,yield_chart - Pipeline:
yield_pipeline,yield_task - Messaging:
yield_topic,yield_subscription
The MySQL source in ingestion/src/metadata/ingestion/source/database/mysql/metadata.py demonstrates complete implementation patterns for all supported operations.
Register the Connector with the Framework
Two registration steps enable discovery across the platform:
1. Enum Registration
Add the connector name to the serviceType enum in openmetadata-spec/src/main/resources/json/schema/services/service_type.json. This populates the Add Service wizard dropdown in the UI.
2. Python Discovery Import
Export the ServiceSpec in ingestion/src/metadata/ingestion/source/__init__.py:
from metadata.ingestion.source.database.mypg.service_spec import MyPgSpec
The complete registration procedure is documented in skills/standards/registration.md.
Add Unit and Integration Tests
OpenMetadata requires real-behavior tests over heavy mocking. Test patterns vary by connector type:
CLI Tests
Subclass CliDBBase.TestSuite for database connectors or CliCommonDB.TestSuite for CommonDbSourceService implementations:
# ingestion/tests/unit/cli_mypg/test_cli_mypg.py
from ingestion.tests.unit.test_cli_db import CliDBBase
class TestMyPg(CliDBBase.TestSuite):
@classmethod
def get_connector_name(cls) -> str:
return "mypg"
Copy from ingestion/tests/unit/cli_mysql/test_cli_mysql.py as a template.
Source Tests
Instantiate the Source with a mock connection (e.g., SQLite in-memory engine) and assert yield_* returns expected entities:
def test_yield_table():
source = MyPgSource(config, metadata)
tables = list(source.yield_table("public"))
assert len(tables) > 0
assert all(isinstance(t, Table) for t in tables)
Update CI for End-to-End Validation
Add your connector to the GitHub Actions matrix in .github/workflows/py-cli-e2e-tests.yml:
strategy:
matrix:
database:
- mysql
- postgres
- mypg # your new connector
This ensures nightly runs validate your connector against real infrastructure.
Verify UI Integration
Complete the development cycle by confirming UI functionality:
- Start the UI locally:
yarn startinopenmetadata-ui/... - Navigate to Settings → Services → Add Service
- Confirm your connector appears in the service type dropdown
- Test configuration form rendering for all custom fields
- Verify save operations persist correctly
Summary
Adding a new connector to OpenMetadata follows a structured, schema-first approach:
- Select the service type and inherit from the appropriate base class (
DatabaseServiceSource,DashboardServiceSource, etc.) - Scaffold with CLI using
./run.sh scaffold-connectorto generate consistent boilerplate - Define JSON Schema in
openmetadata-spec/for UI-driven configuration - Implement three core Python modules:
connection.py(client wrapper),metadata.py(entity generators), andservice_spec.py(capability registration) - Register the connector in both the
serviceTypeenum and Python discovery imports - Add tests using real-behavior patterns with
CliDBBaseorCliCommonDBtest suites - Enable CI validation by adding to
.github/workflows/py-cli-e2e-tests.yml - Verify UI integration through local testing of the Add Service wizard
Frequently Asked Questions
What is the fastest way to start building a new OpenMetadata connector?
Use the scaffold CLI command: ./run.sh scaffold-connector --service-type database --name <your_connector>. This generates the complete directory structure including connection.py, metadata.py, service_spec.py, and test boilerplate, ensuring you follow the canonical pattern from the start.
Can I add a connector for a custom internal data source not listed in OpenMetadata's service types?
Yes. The framework supports any source that can be queried programmatically. Select the closest service type (e.g., database for SQL-based sources, pipeline for orchestration tools) and implement the required generator methods. For truly novel categories, you may need to extend the base class hierarchy in ingestion/src/metadata/ingestion/source/.
How does OpenMetadata validate connector configurations before runtime?
Configuration validation uses JSON Schema definitions in openmetadata-spec/. When you run make generate, Pydantic models are created from these schemas. The ingestion framework validates incoming YAML configurations against these models at startup. Additionally, the test_connection() method in your connection.py provides live connectivity verification through the CLI --validate flag.
What testing requirements must I meet before submitting a connector PR?
OpenMetadata requires real-behavior tests over mocking. For database connectors, subclass CliDBBase.TestSuite or CliCommonDB.TestSuite and implement get_connector_name(). Include source-level tests that instantiate your Source with a mock connection (e.g., SQLite in-memory) and assert that yield_* generators return expected entities. Finally, add your connector to the CI matrix in .github/workflows/py-cli-e2e-tests.yml for nightly end-to-end validation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →