How to Use QLib for AI-Driven Quantitative Investment Research
QLib provides a modular data pipeline, feature engineering framework, and model-backtesting loop that enables researchers to move from raw market data to AI-driven trading strategies with minimal boilerplate code.
QLib is an open-source, AI-oriented quantitative investment platform maintained by Microsoft. According to the paperswithbacktest/awesome-systematic-trading repository—which lists QLib among its essential Machine Learning resources at line 278 of the README—this framework "aims to realize the potential, empower the research, and create the value of AI technologies in quantitative investment" while supporting many state-of-the-art research works.
Understanding QLib's Core Architecture
QLib's architecture revolves around three integrated components that streamline the research workflow:
- Data Server: Stores market data (prices, fundamentals, news) in a time-series key-value store and serves it through the unified
DataHandlerinterface using methods likeDataHandler.get_bar(). - Dataset / Feature Engine: Declares datasets and applies feature pipelines (rolling windows, EWMA, custom factors) to generate training tensors via
Dataset.prepare_data(). - Model / Backtest Engine: Wraps PyTorch-compatible models and provides a
predict()method that plugs into theBacktestengine for realistic portfolio simulation.
This design promotes research reproducibility through YAML-based configuration in qlib/config.yml, which specifies data sources, calendars, and benchmarks for consistent experimental setups across training and backtesting phases.
Setting Up the Environment
Begin by installing QLib and initializing the data provider:
# Install QLib
# pip install pyqlib
import qlib
from qlib.config import REG_CN
# Initialize with Chinese market data (use REG_US for US stocks)
qlib.init(provider_uri="~/.qlib/qlib_data", region=REG_CN)
The provider_uri parameter points to your cached market data directory, while region configures the appropriate trading calendar and benchmark settings defined in qlib/config.py.
Building the Data Pipeline
QLib abstracts data handling through the DataHandler class in qlib/data/data_handler.py. Define your dataset by extending DatasetH from qlib/data/dataset.py:
from qlib.data.dataset import DatasetH
class SimpleDataset(DatasetH):
def __init__(self):
super().__init__(
feature_col=["Ref($close, -1)", "Alpha101::RSI_6"],
label_col="Ref($close, 1)"
)
dataset = SimpleDataset()
train, test = dataset.prepare_data(split_ratio=0.7)
The feature_col parameter accepts QLib's expression syntax—Ref($close, -1) references the previous day's close, while Alpha101::RSI_6 computes a technical factor. The prepare_data() method handles the tensor generation and train/test splitting automatically.
Training AI Models
QLib supports any PyTorch-compatible architecture. Here is a minimal MLP implementation:
import torch
import torch.nn as nn
class MLP(nn.Module):
def __init__(self, input_dim):
super().__init__()
self.net = nn.Sequential(
nn.Linear(input_dim, 64),
nn.ReLU(),
nn.Linear(64, 1)
)
def forward(self, x):
return self.net(x)
model = MLP(input_dim=train["X"].shape[1])
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
# Training loop
for epoch in range(10):
optimizer.zero_grad()
pred = model(torch.from_numpy(train["X"]).float())
loss = criterion(pred.squeeze(), torch.from_numpy(train["y"]).float())
loss.backward()
optimizer.step()
While this example uses a simple neural network, you can substitute LightGBM, Transformers, or LSTM models, as QLib's base interface in qlib/model/base.py only requires fit() and predict() implementations.
Running Backtests
To evaluate your strategy, wrap the trained model using QLib's Model interface and execute a backtest via qlib/backtest/backtest.py:
from qlib.backtest import backtest
from qlib.model import Model
class TorchModel(Model):
def __init__(self, net):
self.net = net
def predict(self, x):
self.net.eval()
with torch.no_grad():
return self.net(torch.from_numpy(x).float()).numpy()
qlib_model = TorchModel(model)
result = backtest(
model=qlib_model,
dataset=dataset,
start_time="2008-01-01",
end_time="2020-12-31",
account=1000000,
benchmark="SH000300",
)
print(result.portfolio_stats())
The backtest engine handles order execution, slippage, transaction costs, and portfolio accounting. Performance metrics are available through qlib/analysis/analysis.py, which calculates IC (Information Coefficient), turnover, and Sharpe ratios.
Key Source Files
When extending QLib for custom research, focus on these core files:
qlib/config.py: Central configuration for regions, calendars, and data paths.qlib/data/data_handler.py: ImplementsDataHandlerfor fetching bars and fundamentals.qlib/data/dataset.py: AbstractDatasetclass for feature and label pipelines.qlib/model/base.py: BaseModelinterface requiringfit()andpredict()methods.qlib/backtest/backtest.py: Portfolio simulation engine with realistic execution logic.qlib/analysis/analysis.py: Performance attribution and risk metrics.
Summary
- QLib provides an end-to-end pipeline from data ingestion to backtesting through three core components: Data Server, Feature Engine, and Backtest Engine.
- Initialize the framework using
qlib.init()with appropriate region settings fromqlib/config.py. - Define datasets by extending
DatasetHand using QLib's expression syntax for feature engineering. - Wrap PyTorch or other ML models using the
Modelinterface fromqlib/model/base.pyto ensure compatibility with the backtester. - Execute realistic portfolio simulations using
backtest()fromqlib/backtest/backtest.py, which handles transaction costs and benchmark comparison.
Frequently Asked Questions
What types of AI models does QLib support?
QLib supports any model that implements the base interface in qlib/model/base.py, including PyTorch neural networks, LightGBM, XGBoost, and TensorFlow models. The framework is model-agnostic as long as you provide fit() and predict() methods that return predictions in the expected format.
How does QLib handle feature engineering?
QLib uses an expression-based syntax processed by the Dataset class in qlib/data/dataset.py. You can reference built-in fields like $close or $volume using operators like Ref() for time shifts, or invoke pre-built factor libraries like Alpha101. Custom features can be added by extending the feature pipeline or implementing new operators.
Can I use QLib for US equity markets?
Yes. Change the region parameter from REG_CN to REG_US when calling qlib.init(), and ensure your provider_uri points to US market data. The configuration in qlib/config.py automatically adjusts trading calendars and benchmark indices (such as S&P 500) based on the region setting.
Where does QLib store calculated metrics and backtest results?
Backtest results are returned as objects containing portfolio statistics accessible via methods like portfolio_stats(). For detailed analysis, use functions from qlib/analysis/analysis.py to calculate IC, turnover, and risk-adjusted returns. You can persist these metrics to disk using standard Python serialization or integrate them into experiment tracking tools.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →