How to Extend ai-hedge-fund with a New Data Provider: A Complete Implementation Guide

To extend ai-hedge-fund with a new data provider, implement a typed provider function in src/tools/api.py that constructs requests, injects API keys, leverages the built-in _make_api_request helper for rate-limit handling, caches results via src/data/cache.py, and returns validated Pydantic models from src/data/models.py.

The virattt/ai-hedge-fund repository aggregates market data through a thin abstraction layer that standardizes external API interactions. This guide walks through the exact architecture used to integrate third-party data sources, ensuring your new provider benefits from automatic retries, in-memory caching, and type-safe model validation consumed by the agent system.

Understanding the Data Abstraction Architecture

The repository follows a consistent five-step pattern for all external data calls, as implemented in src/tools/api.py:

  1. Request Construction: Build the target URL with query parameters.
  2. Authentication: Inject API keys from environment variables (e.g., FINANCIAL_DATASETS_API_KEY).
  3. Resilient Fetching: Call _make_api_request, which handles rate-limit retries and HTTP errors.
  4. Caching: Store raw JSON responses in src/data/cache.py to prevent duplicate API calls.
  5. Model Validation: Wrap responses in Pydantic models defined in src/data/models.py for type-safe consumption by agents.

Step-by-Step Implementation

1. Define the Pydantic Model in src/data/models.py

Create a schema that mirrors the provider's JSON response structure. Reference existing models at line 64 for the expected structure.


# src/data/models.py

from pydantic import BaseModel

class EconomicIndicator(BaseModel):
    name: str
    value: float
    date: str  # ISO-8601 format

class EconomicIndicatorResponse(BaseModel):
    indicators: list[EconomicIndicator]

2. Configure Caching in src/data/cache.py (Optional)

If the provider returns data that benefits from dedicated cache logic, add getter and setter methods. For simple use cases, reuse generic methods like set_prices found at line 24.


# src/data/cache.py

class Cache:
    # Existing initialization...

    
    def get_economic_indicators(self, key: str) -> list[dict] | None:
        return self._economic_indicators_cache.get(key)

    def set_economic_indicators(self, key: str, data: list[dict]):
        self._economic_indicators_cache = self._merge_data(
            self._economic_indicators_cache.get(key), data, key_field="date"
        )

3. Implement the Provider Function in src/tools/api.py

Add a new function following the established pattern. It should check the cache, build the request, call _make_api_request, parse the response, and update the cache.


# src/tools/api.py

import os
from src.data.models import EconomicIndicator, EconomicIndicatorResponse
from src.data.cache import _cache

def get_economic_indicators(
    ticker: str,
    start_date: str,
    end_date: str,
    api_key: str | None = None,
) -> list[EconomicIndicator]:
    """Fetch economic indicators with automatic caching and rate-limit handling."""
    cache_key = f"econ_{ticker}_{start_date}_{end_date}"
    
    # Check cache

    if cached := _cache.get_economic_indicators(cache_key):
        return [EconomicIndicator(**i) for i in cached]
    
    # Build request

    api_key = api_key or os.getenv("ECONOMIC_DATA_API_KEY")
    headers = {"X-API-KEY": api_key} if api_key else {}
    url = (
        f"https://api.economicdata.io/indicators?"
        f"ticker={ticker}&start={start_date}&end={end_date}"
    )
    
    # Execute with retry logic

    response = _make_api_request(url, headers)
    if response.status_code != 200:
        return []
    
    # Parse and validate

    parsed = EconomicIndicatorResponse(**response.json())
    indicators = parsed.indicators
    
    # Update cache

    _cache.set_economic_indicators(
        cache_key, [i.model_dump() for i in indicators]
    )
    return indicators

4. Wire the Provider into Agents

Import your new function into any agent file (e.g., src/agents/technicals.py around line 56) and invoke it within the agent's execution logic.


# src/agents/economic_insight.py

from src.tools.api import get_economic_indicators

def run(state):
    ticker = state["metadata"]["ticker"]
    indicators = get_economic_indicators(
        ticker, 
        state["metadata"]["start_date"], 
        state["metadata"]["end_date"]
    )
    state["data"]["economic_indicators"] = [i.model_dump() for i in indicators]
    return state

5. Document the Environment Variable

Add the required API key to .env.example so users know to configure it:


# .env.example

FINANCIAL_DATASETS_API_KEY=your_key_here
ECONOMIC_DATA_API_KEY=your_new_key_here

Integrating with the CLI and Web UI

To expose the new data source via user interfaces:

  • CLI: Modify src/cli/input.py (around line 107) to add prompts or flags that toggle the new provider.
  • Web UI: Extend the FastAPI backend under app/backend/services and create corresponding frontend components to display the data.

Testing Your Implementation

Validate your provider with two test layers:

  1. Unit Tests: Monkey-patch _make_api_request in tests/ to return mock JSON and assert that your function returns correctly typed EconomicIndicator objects.
  2. Integration Tests: Run the existing back-testing harness to ensure your agent integration doesn't break the execution pipeline.

Summary

To successfully extend ai-hedge-fund with a new data provider:

  • Define Pydantic models in src/data/models.py for type-safe responses.
  • Implement caching logic in src/data/cache.py if dedicated storage is needed.
  • Create provider functions in src/tools/api.py that leverage _make_api_request for resilient fetching and automatic retries.
  • Import and call these functions from agent files in src/agents/.
  • Document environment variables in .env.example and add comprehensive tests.

Frequently Asked Questions

Do I need to modify the core caching system to add a new provider?

No, you can reuse existing generic cache methods like set_prices and get_prices if your data structure is compatible. Only implement dedicated getters and setters in src/data/cache.py if the provider requires specialized merge logic or cache invalidation rules.

How does ai-hedge-fund handle API rate limits when adding new providers?

The _make_api_request function in src/tools/api.py includes built-in retry logic with exponential backoff. When you call this helper instead of raw requests, your new provider automatically inherits rate-limit handling and error resilience.

Can I use the same Pydantic model for multiple similar providers?

Yes, but define distinct provider functions for each API source. Share the Pydantic model classes across functions if the JSON schemas match, allowing consistent type hints while maintaining separate request logic and cache keys for each data source.

Where should I add CLI options for toggling my new data source?

Add new prompts or argument flags in src/cli/input.py. Reference line 107 for examples of how the repository handles user input selection. Then pass these flags through to the agent initialization logic to conditionally fetch data from your new provider.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →