How to Implement Custom Fetchers for Integrating New Data Providers in OpenBB
You can integrate any external data source into OpenBB by subclassing the abstract Fetcher class located in openbb_platform/core/openbb_core/provider/abstract/fetcher.py, defining Pydantic models for query parameters and responses, and registering your implementation in the provider's __init__.py file to enable automatic discovery by the registry scanner.
The OpenBB Platform uses a provider-based architecture that allows developers to seamlessly add new data sources through a standardized fetcher pattern. Understanding how to implement custom fetchers for integrating new data providers in OpenBB enables you to extend the platform's capabilities with proprietary or niche financial data APIs. This guide walks through the core abstractions, registry mechanisms, and concrete implementation steps based on the actual source code in the OpenBB-finance/OpenBB repository.
Architecture Overview
OpenBB's data provider system relies on three core components that work together to standardize data ingestion across disparate sources.
The Fetcher Abstract Base Class
At the heart of the system lies the Fetcher abstract base class defined in openbb_platform/core/openbb_core/provider/abstract/fetcher.py. This class establishes a strict contract between the platform and data providers through generic type parameters <Query, Response>. Concrete implementations must override the asynchronous _fetch_data(self, params: Query) -> Response method, which contains the actual API call logic and data transformation. The public fetch_data method defined in the base class wraps this internal logic with error handling, caching checks, and validation, ensuring consistent behavior across all providers.
Provider Registry and Auto-Discovery
The registry_map.py file in openbb_platform/core/openbb_core/provider/ implements the auto-discovery mechanism that eliminates manual registration boilerplate. When OpenBB starts, the registry scans the openbb_platform.providers.<provider_name> namespace and automatically indexes any subclass of Fetcher it discovers. This registration links the fetcher to its provider identifier, making it available to the QueryExecutor and CLI without additional configuration.
Query Execution Pipeline
The QueryExecutor class in openbb_platform/core/openbb_core/provider/query_executor.py orchestrates the data retrieval workflow. It instantiates the appropriate fetcher class based on the user's provider selection, validates input parameters against the Query model using Pydantic, and invokes the fetcher's fetch_data method. This pipeline handles retries, timeouts, and caching transparently, allowing fetcher implementations to focus solely on data extraction.
Implementation Steps
Creating a functional custom fetcher requires four discrete steps that align with OpenBB's type-safe architecture.
Step 1: Define Query and Response Models
Every fetcher requires two Pydantic models: a Query model that defines input parameters and a Response model that structures the returned data. These models live in your provider's models/ directory and must inherit from BaseModel. The Query model drives CLI argument generation and API validation, while the Response model ensures type consistency across the platform.
Step 2: Subclass the Fetcher
Create a new class that inherits from Fetcher[YourQuery, YourResponse] and implements the required interface. The generic type parameters bind your specific models to the fetcher, enabling static type checking throughout the execution pipeline. You may optionally define a name property to aid in debugging, though the class name serves as the default identifier.
Step 3: Implement the Fetch Logic
Override the _fetch_data method with your API integration code. This method receives an instance of your Query model and must return an instance of your Response model. Use async/await syntax for non-blocking I/O operations, as the QueryExecutor expects an awaitable coroutine. Handle HTTP errors, rate limiting, and data normalization within this method, converting raw API responses into the structured format defined by your Response model.
Step 4: Register Your Fetcher
Import your fetcher class in your provider package's __init__.py file. The registry scanner imports all modules within openbb_platform.providers.your_provider, automatically detecting Fetcher subclasses and adding them to the available provider map. No explicit registration function call is required—simply ensuring the class loads into memory suffices.
Complete Working Example
The following implementation demonstrates a custom fetcher for a hypothetical "MyData" provider supplying daily equity prices. This example follows the exact patterns used in production fetchers like YFinanceEquityHistoricalFetcher in openbb_platform/providers/yfinance/openbb_yfinance/models/equity_historical.py.
# openbb_platform/providers/mydata/openbb_mydata/models/daily_price.py
from datetime import date
from typing import List
import httpx
import pandas as pd
from pydantic import BaseModel, Field
from openbb_core.provider.abstract.fetcher import Fetcher
class DailyPriceQuery(BaseModel):
"""Parameters for daily price data requests."""
symbol: str = Field(..., description="Ticker symbol, e.g. AAPL")
start_date: date = Field(..., description="Start of period (YYYY-MM-DD)")
end_date: date = Field(..., description="End of period (YYYY-MM-DD)")
class DailyPriceResponse(BaseModel):
"""Structured response for daily OHLCV data."""
date: List[date]
open: List[float]
high: List[float]
low: List[float]
close: List[float]
volume: List[int]
class MyDataDailyPriceFetcher(Fetcher[DailyPriceQuery, DailyPriceResponse]):
"""Custom fetcher integrating the MyData API."""
@property
def name(self) -> str:
return "mydata_daily_price"
async def _fetch_data(self, params: DailyPriceQuery) -> DailyPriceResponse:
"""Fetch and transform data from MyData REST API."""
url = (
f"https://api.mydata.com/v1/prices?"
f"symbol={params.symbol}"
f"&start={params.start_date.isoformat()}"
f"&end={params.end_date.isoformat()}"
)
async with httpx.AsyncClient() as client:
response = await client.get(url, timeout=30.0)
response.raise_for_status()
raw_data = response.json()
# Normalize JSON array into DataFrame for reliable column extraction
df = pd.DataFrame(raw_data)
return DailyPriceResponse(
date=pd.to_datetime(df["date"]).dt.date.tolist(),
open=df["open"].astype(float).tolist(),
high=df["high"].astype(float).tolist(),
low=df["low"].astype(float).tolist(),
close=df["close"].astype(float).tolist(),
volume=df["volume"].astype(int).tolist(),
)
Register the fetcher in your provider's initialization file:
# openbb_platform/providers/mydata/__init__.py
from openbb_mydata.models.daily_price import MyDataDailyPriceFetcher # noqa: F401
After registration, the fetcher becomes accessible via the OpenBB CLI:
openbb mydata daily-price --symbol AAPL --start-date 2024-01-01 --end-date 2024-01-31
Key Source Files Reference
Understanding the following files provides deep insight into the fetcher lifecycle and registration mechanics:
-
openbb_platform/core/openbb_core/provider/abstract/fetcher.py– Defines theFetcherabstract base class with thefetch_dataentry point and_fetch_dataabstract method requiring implementation. -
openbb_platform/core/openbb_core/provider/registry_map.py– Implements the auto-discovery scanner that maps provider namespaces to concrete fetcher classes at startup. -
openbb_platform/core/openbb_core/provider/query_executor.py– Orchestrates fetcher instantiation, parameter validation, and execution flow including caching and error handling layers. -
openbb_platform/providers/yfinance/openbb_yfinance/models/equity_historical.py– Production reference implementation showing a complete fetcher with complex query parameters and data transformation logic.
Summary
Implementing custom fetchers for integrating new data providers in OpenBB follows a standardized, type-safe pattern that minimizes boilerplate while maximizing flexibility. The architecture automatically handles validation, caching, and CLI integration once you provide the core data retrieval logic.
- Subclass
Fetcher[Query, Response]fromopenbb_platform/core/openbb_core/provider/abstract/fetcher.pyand implement the async_fetch_datamethod with your API-specific logic. - Define Pydantic models for query parameters and response data to enable automatic validation, CLI argument generation, and OpenAPI schema creation.
- Import your fetcher class in the provider's
__init__.pyto trigger auto-registration viaregistry_map.pywithout explicit function calls. - Reference existing implementations in the YFinance provider for complex data transformation patterns and best practices for error handling.
Frequently Asked Questions
What is the difference between fetch_data and _fetch_data in the Fetcher class?
The public fetch_data method is defined in the abstract base class and should not be overridden. It orchestrates caching, validation, and error handling before and after calling your implementation. You must implement the protected _fetch_data(self, params: Query) -> Response method, which contains only the API-specific logic for retrieving and formatting data. This separation of concerns ensures consistent behavior across all providers while allowing flexibility in data source integration.
How does OpenBB discover my custom fetcher without explicit registration?
OpenBB uses the registry_map.py scanner that imports all modules within openbb_platform.providers.<your_provider> at startup. Any class inheriting from Fetcher that loads into memory during this scan is automatically indexed and associated with its provider identifier. You only need to ensure your fetcher class is imported in your provider package's __init__.py file to trigger this discovery mechanism.
Can I use synchronous HTTP libraries like requests instead of httpx?
While the _fetch_data method is defined as async in the abstract base class, you can wrap synchronous calls using asyncio.get_event_loop().run_in_executor() or similar patterns. However, using native async libraries like httpx or aiohttp is strongly recommended to prevent blocking the event loop and maintain performance when multiple fetchers run concurrently. The QueryExecutor expects an awaitable coroutine and will properly schedule async fetchers.
What validation occurs on the Query and Response models?
OpenBB leverages Pydantic for strict type validation. The QueryExecutor validates user inputs against your Query model before calling _fetch_data, converting CLI arguments or API parameters to the correct Python types and raising validation errors for missing required fields or type mismatches. After your fetcher returns data, the Response model validates the structure, ensuring that all fields match their declared types and that the data conforms to the expected schema before passing it to the consumer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →