How to Implement a Custom Zipline Data Bundle for Intraday Trading Data

A custom Zipline data bundle for intraday trading data requires subclassing an exchange calendar to define market hours, implementing a generator that yields UTC-indexed minute-bar DataFrames paired with asset metadata, and registering the bundle via an extension.py file.

A Zipline data bundle is a packaged collection of asset metadata and OHLCV time-series that the Zipline back-testing engine can ingest and query. For intraday strategies, you must supply minute-resolution bars along with a custom calendar that defines your specific trading session hours. The stefan-jansen/machine-learning-for-trading repository provides a complete, production-ready implementation that converts Algoseek NASDAQ-100 minute-bar data into a Zipline-compatible format.

What Is a Custom Zipline Data Bundle?

A bundle encapsulates three core components required for back-testing:

  • Exchange calendar: Defines trading session open/close times and timezone.
  • Asset metadata: SID-to-symbol mapping, start/end dates, first-traded dates, auto-close dates, and exchange codes.
  • Minute-bar writer: Streams DataFrames indexed by UTC timestamps to Zipline's internal storage format.

For intraday data, you must also specify minutes_per_day during registration so Zipline correctly partitions the Bars matrix.

Step-by-Step Implementation

Define the Exchange Calendar

Create a subclass of XNYSExchangeCalendar (or another base calendar) and override open_times, close_times, and tz. In 08_ml4t_workflow/04_ml4t_workflow_with_zipline/01_custom_bundles/extension.py, the AlgoSeekCalendar defines extended hours for the Algoseek dataset:

from trading_calendars.exchange_calendar_xnys import XNYSExchangeCalendar
from datetime import time
from pytz import timezone

class AlgoSeekCalendar(XNYSExchangeCalendar):
    @property
    def name(self):
        return "AlgoSeek"
    
    @property
    def tz(self):
        return timezone("US/Eastern")
    
    open_times = ((None, time(4, 1)),)    # 04:01 AM ET

    close_times = ((None, time(19, 59)),) # 07:59 PM ET

Register the Calendar and Bundle

Register the calendar with trading_calendars and the bundle with zipline.data.bundles.register. The minutes_per_day parameter must match your session length (960 minutes for the 16-hour Algoseek session):

from trading_calendars import register_calendar
from zipline.data.bundles import register
from algoseek_1min_trades import algoseek_to_bundle

register_calendar('AlgoSeek', AlgoSeekCalendar())

register(
    'algoseek',
    algoseek_to_bundle(),
    calendar_name='AlgoSeek',
    minutes_per_day=960  # 16 hours × 60 minutes

)

Build Asset Metadata

The metadata_frame function in algoseek_1min_trades.py constructs an empty NumPy-structured DataFrame that will hold per-asset information:

import pandas as pd
import numpy as np

def metadata_frame():
    dtype = [
        ('symbol', 'object'),
        ('asset_name', 'object'),
        ('start_date', 'datetime64[ns]'),
        ('end_date', 'datetime64[ns]'),
        ('first_traded', 'datetime64[ns]'),
        ('auto_close_date', 'datetime64[ns]'),
        ('exchange', 'object')
    ]
    return pd.DataFrame(np.empty(len(load_equities()), dtype=dtype))

Create the Data Generator

The data_generator lazily yields tuples containing the minute-bar DataFrame and metadata fields. The DataFrame must be timezone-aware and converted to UTC:

def data_generator():
    for sid, symbol, asset_name in ticker_generator():
        # Read from HDF5 and normalize timezone

        df = (pd.read_hdf(custom_data_path / 'algoseek.h5', str(sid))
                .tz_localize('US/Eastern')
                .tz_convert('UTC'))
        
        # Compute metadata fields

        start_date = df.index[0]
        end_date = df.index[-1]
        first_traded = start_date.date()
        auto_close_date = end_date + pd.Timedelta(days=1)
        exchange = "AlgoSeek"
        
        yield (sid, df), symbol, asset_name, start_date, end_date, first_traded, auto_close_date, exchange

Implement the Ingest Function

The algoseek_to_bundle function returns the closure required by Zipline. It orchestrates writing minute bars, asset metadata, and optional adjustments:

def algoseek_to_bundle(interval='1m'):
    def ingest(environ, asset_db_writer, minute_bar_writer,
                 daily_bar_writer, adjustment_writer, calendar,
                 start_session, end_session, cache, show_progress,
                 output_dir):
        
        metadata = metadata_frame()
        
        # Write minute bars

        minute_bar_writer.write(
            (sid_df for (sid_df, *metadata.iloc[sid_df[0]]) in data_generator()),
            show_progress=True
        )
        
        # Drop empty rows and write asset metadata

        metadata.dropna(inplace=True)
        asset_db_writer.write(equities=metadata)
        
        # Optional: write splits/dividends

        adjustment_writer.write(splits=pd.read_hdf(custom_data_path / 'algoseek.h5', 'splits'))
    
    return ingest

Complete Working Example

Here is a minimal skeleton you can adapt for any intraday source (CSV, database, or API):


# extension.py

import sys
from pathlib import Path
sys.path.append(Path('~', '.zipline').expanduser().as_posix())

from datetime import time
from pytz import timezone
from trading_calendars import register_calendar
from trading_calendars.exchange_calendar_xnys import XNYSExchangeCalendar
from zipline.data.bundles import register
from my_bundle.ingest import my_intraday_to_bundle

class MyIntradayCalendar(XNYSExchangeCalendar):
    @property
    def name(self):
        return "MyIntraday"
    
    @property
    def tz(self):
        return timezone("US/Eastern")
    
    open_times = ((None, time(9, 30)),)   # 9:30 AM ET

    close_times = ((None, time(16, 0)),)  # 4:00 PM ET

register_calendar('MyIntraday', MyIntradayCalendar())

register(
    'my_intraday',
    my_intraday_to_bundle(),
    calendar_name='MyIntraday',
    minutes_per_day=390  # 6.5 hours × 60 minutes

)

# ingest.py

import pandas as pd
import numpy as np
from pathlib import Path
from os import getenv

DATA_ROOT = Path(getenv('ZIPLINE_ROOT', Path('~', '.zipline')).expanduser(), 'custom_data')
DATA_H5 = DATA_ROOT / 'my_intraday.h5'

def load_equities():
    return pd.read_hdf(DATA_H5, 'equities')

def ticker_generator():
    return (v for v in load_equities().values)

def data_generator():
    for sid, symbol, asset_name in ticker_generator():
        df = (pd.read_hdf(DATA_H5, str(sid))
                .tz_localize('US/Eastern')
                .tz_convert('UTC'))
        start, end = df.index[0], df.index[-1]
        yield (sid, df), symbol, asset_name, start, end, start.date(), end + pd.Timedelta(days=1), "NYSE"

def metadata_frame():
    dtype = [('symbol','object'),('asset_name','object'),
             ('start_date','datetime64[ns]'),('end_date','datetime64[ns]'),
             ('first_traded','datetime64[ns]'),('auto_close_date','datetime64[ns]'),
             ('exchange','object')]
    return pd.DataFrame(np.empty(len(load_equities()), dtype=dtype))

def my_intraday_to_bundle(interval='1m'):
    def ingest(environ, asset_db_writer, minute_bar_writer,
               daily_bar_writer, adjustment_writer, calendar,
               start_session, end_session, cache, show_progress,
               output_dir):
        meta = metadata_frame()
        minute_bar_writer.write(
            (sid_df for (sid_df, *meta.iloc[sid_df[0]]) in data_generator()),
            show_progress=True
        )
        meta.dropna(inplace=True)
        asset_db_writer.write(equities=meta)
        # adjustment_writer.write(splits=...) if needed

    return ingest

Key Files in the Repository

Path Purpose
08_ml4t_workflow/04_ml4t_workflow_with_zipline/01_custom_bundles/extension.py Registers the AlgoSeek calendar and bundle with Zipline.
08_ml4t_workflow/04_ml4t_workflow_with_zipline/01_custom_bundles/algoseek_1min_trades.py Implements ingest logic including metadata construction and minute-bar writing.
11_decision_trees_random_forests/00_custom_bundle/extension.py Minimal example for daily-resolution bundles (STOOQ Japan data).
11_decision_trees_random_forests/00_custom_bundle/stooq_jp_stocks.py Shows adaptation patterns for alternative data sources.
08_ml4t_workflow/04_ml4t_workflow_with_zipline/01_custom_bundles/algoseek_preprocessing.py Helper script that prepares the HDF5 source file (algoseek.h5).

Ingesting and Running the Bundle

  1. Set the ZIPLINE_ROOT environment variable (defaults to ~/.zipline):

    export ZIPLINE_ROOT=~/.zipline
  2. Ensure your extension.py is importable (place it in ~/.zipline/extension.py or run it explicitly):

    python -c "import extension"
  3. Ingest the data:

    zipline ingest -b algoseek
  4. Run a back-test using the custom bundle:

    zipline run -b algoseek -f my_strategy.py --start 2020-01-01 --end 2020-12-31

Summary

  • Subclass XNYSExchangeCalendar (or another base calendar) to define custom trading hours for your intraday data.
  • Register the calendar with trading_calendars before registering the bundle.
  • Supply minutes_per_day during bundle registration to match your session length (e.g., 390 for standard US equity hours, 960 for extended hours).
  • Yield UTC-indexed DataFrames from your data_generator, paired with SID and metadata tuples.
  • Use minute_bar_writer.write() to stream data, then asset_db_writer.write(equities=metadata) to store asset lifetimes.

Frequently Asked Questions

What is the difference between a daily and minute Zipline data bundle?

A daily bundle uses daily_bar_writer to ingest end-of-day OHLCV bars, while an intraday bundle uses minute_bar_writer to ingest minute-resolution DataFrames. Intraday bundles also require a custom calendar defining exact market open/close times and must specify minutes_per_day during registration so Zipline allocates the correct amount of storage per session.

Why must the DataFrame index be in UTC?

Zipline internally stores all timestamps in UTC to ensure consistency across different data sources and timezones. The data_generator must localize timestamps to the exchange timezone (e.g., US/Eastern) and then convert to UTC using .tz_convert('UTC') before yielding the DataFrame to minute_bar_writer.

How do I handle corporate actions like splits and dividends?

Pass a DataFrame of splits or dividends to adjustment_writer.write() inside the ingest function. According to the algoseek_1min_trades.py implementation, you can load these from your source HDF5 file: adjustment_writer.write(splits=pd.read_hdf(path, 'splits')). The DataFrame should contain columns for sid, effective_date, and ratio (for splits) or amount (for dividends).

Can I use this pattern for data sources other than HDF5?

Yes. The data_generator pattern is source-agnostic. You can replace pd.read_hdf with pd.read_csv, SQL queries via pandas.read_sql, or API calls. As long as you yield tuples of (sid, df) where df is a UTC-indexed minute-bar DataFrame, the bundle will function correctly with Zipline's back-testing engine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →