How to Extend ai-hedge-fund with a New Data Provider: A Complete Implementation Guide
To extend ai-hedge-fund with a new data provider, implement a typed provider function in src/tools/api.py that constructs requests, injects API keys, leverages the built-in _make_api_request helper for rate-limit handling, caches results via src/data/cache.py, and returns validated Pydantic models from src/data/models.py.
The virattt/ai-hedge-fund repository aggregates market data through a thin abstraction layer that standardizes external API interactions. This guide walks through the exact architecture used to integrate third-party data sources, ensuring your new provider benefits from automatic retries, in-memory caching, and type-safe model validation consumed by the agent system.
Understanding the Data Abstraction Architecture
The repository follows a consistent five-step pattern for all external data calls, as implemented in src/tools/api.py:
- Request Construction: Build the target URL with query parameters.
- Authentication: Inject API keys from environment variables (e.g.,
FINANCIAL_DATASETS_API_KEY). - Resilient Fetching: Call
_make_api_request, which handles rate-limit retries and HTTP errors. - Caching: Store raw JSON responses in
src/data/cache.pyto prevent duplicate API calls. - Model Validation: Wrap responses in Pydantic models defined in
src/data/models.pyfor type-safe consumption by agents.
Step-by-Step Implementation
1. Define the Pydantic Model in src/data/models.py
Create a schema that mirrors the provider's JSON response structure. Reference existing models at line 64 for the expected structure.
# src/data/models.py
from pydantic import BaseModel
class EconomicIndicator(BaseModel):
name: str
value: float
date: str # ISO-8601 format
class EconomicIndicatorResponse(BaseModel):
indicators: list[EconomicIndicator]
2. Configure Caching in src/data/cache.py (Optional)
If the provider returns data that benefits from dedicated cache logic, add getter and setter methods. For simple use cases, reuse generic methods like set_prices found at line 24.
# src/data/cache.py
class Cache:
# Existing initialization...
def get_economic_indicators(self, key: str) -> list[dict] | None:
return self._economic_indicators_cache.get(key)
def set_economic_indicators(self, key: str, data: list[dict]):
self._economic_indicators_cache = self._merge_data(
self._economic_indicators_cache.get(key), data, key_field="date"
)
3. Implement the Provider Function in src/tools/api.py
Add a new function following the established pattern. It should check the cache, build the request, call _make_api_request, parse the response, and update the cache.
# src/tools/api.py
import os
from src.data.models import EconomicIndicator, EconomicIndicatorResponse
from src.data.cache import _cache
def get_economic_indicators(
ticker: str,
start_date: str,
end_date: str,
api_key: str | None = None,
) -> list[EconomicIndicator]:
"""Fetch economic indicators with automatic caching and rate-limit handling."""
cache_key = f"econ_{ticker}_{start_date}_{end_date}"
# Check cache
if cached := _cache.get_economic_indicators(cache_key):
return [EconomicIndicator(**i) for i in cached]
# Build request
api_key = api_key or os.getenv("ECONOMIC_DATA_API_KEY")
headers = {"X-API-KEY": api_key} if api_key else {}
url = (
f"https://api.economicdata.io/indicators?"
f"ticker={ticker}&start={start_date}&end={end_date}"
)
# Execute with retry logic
response = _make_api_request(url, headers)
if response.status_code != 200:
return []
# Parse and validate
parsed = EconomicIndicatorResponse(**response.json())
indicators = parsed.indicators
# Update cache
_cache.set_economic_indicators(
cache_key, [i.model_dump() for i in indicators]
)
return indicators
4. Wire the Provider into Agents
Import your new function into any agent file (e.g., src/agents/technicals.py around line 56) and invoke it within the agent's execution logic.
# src/agents/economic_insight.py
from src.tools.api import get_economic_indicators
def run(state):
ticker = state["metadata"]["ticker"]
indicators = get_economic_indicators(
ticker,
state["metadata"]["start_date"],
state["metadata"]["end_date"]
)
state["data"]["economic_indicators"] = [i.model_dump() for i in indicators]
return state
5. Document the Environment Variable
Add the required API key to .env.example so users know to configure it:
# .env.example
FINANCIAL_DATASETS_API_KEY=your_key_here
ECONOMIC_DATA_API_KEY=your_new_key_here
Integrating with the CLI and Web UI
To expose the new data source via user interfaces:
- CLI: Modify
src/cli/input.py(around line 107) to add prompts or flags that toggle the new provider. - Web UI: Extend the FastAPI backend under
app/backend/servicesand create corresponding frontend components to display the data.
Testing Your Implementation
Validate your provider with two test layers:
- Unit Tests: Monkey-patch
_make_api_requestintests/to return mock JSON and assert that your function returns correctly typedEconomicIndicatorobjects. - Integration Tests: Run the existing back-testing harness to ensure your agent integration doesn't break the execution pipeline.
Summary
To successfully extend ai-hedge-fund with a new data provider:
- Define Pydantic models in
src/data/models.pyfor type-safe responses. - Implement caching logic in
src/data/cache.pyif dedicated storage is needed. - Create provider functions in
src/tools/api.pythat leverage_make_api_requestfor resilient fetching and automatic retries. - Import and call these functions from agent files in
src/agents/. - Document environment variables in
.env.exampleand add comprehensive tests.
Frequently Asked Questions
Do I need to modify the core caching system to add a new provider?
No, you can reuse existing generic cache methods like set_prices and get_prices if your data structure is compatible. Only implement dedicated getters and setters in src/data/cache.py if the provider requires specialized merge logic or cache invalidation rules.
How does ai-hedge-fund handle API rate limits when adding new providers?
The _make_api_request function in src/tools/api.py includes built-in retry logic with exponential backoff. When you call this helper instead of raw requests, your new provider automatically inherits rate-limit handling and error resilience.
Can I use the same Pydantic model for multiple similar providers?
Yes, but define distinct provider functions for each API source. Share the Pydantic model classes across functions if the JSON schemas match, allowing consistent type hints while maintaining separate request logic and cache keys for each data source.
Where should I add CLI options for toggling my new data source?
Add new prompts or argument flags in src/cli/input.py. Reference line 107 for examples of how the repository handles user input selection. Then pass these flags through to the agent initialization logic to conditionally fetch data from your new provider.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →