Limitations of the aeon Time Series Library Compared to statsmodels for Forecasting

While aeon delivers high-performance machine learning forecasting through scikit-learn-compatible estimators, it lacks the rigorous statistical diagnostics, native multivariate VAR support, exact likelihood inference, and transparent coefficient reporting that statsmodels provides for classical econometric analysis.

The choice between aeon and statsmodels for time series forecasting depends fundamentally on whether your workflow prioritizes scalable machine learning or statistical rigor. According to the K-Dense-AI/scientific-agent-skills repository, aeon treats forecasting as a machine learning task using vectorized NumPy and numba pipelines, while statsmodels implements exact statistical estimators with comprehensive diagnostic capabilities. Understanding these limitations ensures you select the appropriate library for interpretable econometric analysis versus large-scale predictive modeling.

Architectural Design Philosophy

Machine Learning vs. Statistical Modeling

aeon adopts a scikit-learn-compatible design that treats time series problems as machine learning tasks. Forecasting is implemented through estimators following the fit/predict API pattern, including ARIMA, NaiveForecaster, and deep learning forecasters like DeepARNetwork and TCNForecaster. As documented in scientific-skills/aeon/SKILL.md, the library focuses on vectorized NumPy pipelines and optional deep learning back-ends (PyTorch, JAX) for high-dimensional feature extraction using transformations like ROCKET and Catch22.

statsmodels follows a statistical-modeling paradigm centered on inference, diagnostics, and classical econometric methods. The library provides dedicated statistical classes (ARIMA, SARIMAX, ETS, VAR) that expose model parameters, likelihoods, and rich diagnostic tools. According to scientific-skills/statsmodels/SKILL.md, these implementations use exact maximum likelihood estimation (MLE) and generalized least squares (GLS) written in Python with compiled Fortran routines via numpy.linalg and scipy.

Multivariate Modeling Limitations

Limited Native Multivariate Support

aeon's forecasting scope remains primarily univariate, with multivariate support limited to wrappers that treat each series independently or require manual concatenation of lagged features. The scientific-skills/aeon/references/forecasting.md file confirms that while deep learning forecasters exist, there is no equivalent to the vector autoregression (VAR) models essential for Granger causality testing and macro-economic forecasting.

statsmodels provides a complete multivariate ecosystem including VAR, VARMAX, and state-space models. These classes support cross-series relationships and simultaneous equation modeling, as demonstrated in scientific-skills/statsmodels/references/time_series.md with the VAR implementation that accepts DataFrame inputs and handles multiple endogenous variables natively.

Exogenous Regressor Handling

In aeon, exogenous support exists through RegressionForecaster but functions as a thin wrapper around scikit-learn regressors without built-in handling of lagged exogenous term alignment.

statsmodels offers native exog arguments across all classical models (ARIMA, SARIMAX, VARMAX) that automatically align exogenous series with endogenous lag structures. This integration is critical for transfer function modeling and intervention analysis.

Statistical Inference and Diagnostics

Absence of Diagnostic Toolboxes

aeon provides only generic machine learning metrics (MAE, MSE) and optional cross-validation utilities. The library lacks built-in Ljung-Box autocorrelation tests, ADF stationarity tests, and formal residual analysis plots. As noted in the source analysis, there are no likelihood-ratio tests or information-criterion based model selection utilities (AIC, BIC) within the core forecasting API.

statsmodels includes an extensive diagnostic toolbox accessible through scientific-skills/statsmodels/references/stats_diagnostics.md. Users can perform autocorrelation testing, heteroskedasticity analysis, formal seasonality decomposition, and obtain exact standard errors for every parameter estimate.

Confidence Intervals and Prediction Bands

aeon offers no built-in confidence intervals for forecasts beyond heuristic quantiles from deep learning models. The predict method returns point estimates without theoretically justified uncertainty bounds.

statsmodels derives prediction intervals from underlying state-space representations, providing theoretically justified confidence bounds through the get_forecast method. These intervals account for parameter estimation error and inherent stochasticity in the data generating process.

Parameter Interpretability

aeon model parameters typically represent feature-level weights (e.g., Ridge regression coefficients on ROCKET features) that lack direct time-series interpretability. There are no summary tables displaying lag coefficients, t-statistics, or p-values for individual terms.

statsmodels generates transparent coefficient tables with standard errors, z-statistics, and confidence intervals for every autoregressive and moving average term. This transparency enables rigorous hypothesis testing and academic reporting, as shown in the linear regression examples within scientific-skills/statsmodels/SKILL.md.

Performance Considerations for Small Data

aeon is optimized for large-scale, high-dimensional problems involving millions of series, where fast feature kernels justify the overhead. For small datasets with few observations, the fixed cost of feature extraction (e.g., computing Catch22 features) can dominate computation time, making it inefficient for short univariate series.

statsmodels remains highly efficient on modest datasets, where exact MLE calculation via Kalman filtering is fast for hundreds of observations and offers deterministic, reproducible results without feature extraction overhead.

Practical Code Comparison

aeon Implementation

The following example demonstrates aeon's machine learning API for ARIMA and naive forecasting:

from aeon.forecasting.arima import ARIMA
from aeon.forecasting.naive import NaiveForecaster
import numpy as np

# Simple synthetic series

y = np.arange(1, 21).astype(float)

# ARIMA model (machine-learning style)

arima = ARIMA(order=(1, 1, 1))
arima.fit(y)
forecast_arima = arima.predict(fh=[1, 2, 3])

# Naïve baseline for comparison

naive = NaiveForecaster(strategy="last")
naive.fit(y)
forecast_naive = naive.predict(fh=[1, 2, 3])

Source: scientific-skills/aeon/SKILL.md describes the forecasting API and available estimators.

statsmodels Implementation with Diagnostics

statsmodels provides identical forecasting capability with additional statistical output:

import numpy as np
import pandas as pd
from statsmodels.tsa.arima.model import ARIMA

# Same series wrapped in pandas Series (required for dates)

y = pd.Series(np.arange(1, 21).astype(float))

# Fit with full likelihood information

model = ARIMA(y, order=(1, 1, 1))
results = model.fit()

# Forecast with confidence intervals

forecast = results.get_forecast(steps=3)
pred = forecast.predicted_mean
ci = forecast.conf_int()

print(results.summary())

Source: scientific-skills/statsmodels/SKILL.md details the time-series API with statistical inference.

Multivariate VAR Example (statsmodels Only)

aeon cannot natively produce joint forecasts for multiple interacting series:

import pandas as pd
from statsmodels.tsa.vector_ar.var_model import VAR

# Simulated bivariate series

df = pd.DataFrame({
    "y1": np.random.randn(100).cumsum(),
    "y2": np.random.randn(100).cumsum(),
})

model = VAR(df)
results = model.fit(maxlags=2)

# Forecast 5 steps jointly

forecast = results.forecast(df.values[-results.k_ar:], steps=5)

Source: scientific-skills/statsmodels/references/time_series.md provides VAR class documentation.

Summary

  • aeon excels at scalable ML forecasting with scikit-learn compatibility but lacks rigorous statistical diagnostics.
  • statsmodels provides exact MLE estimation, hypothesis testing, and multivariate VAR models unavailable in aeon.
  • Parameter interpretability in aeon is limited to feature weights, while statsmodels offers transparent coefficient tables with standard errors.
  • For small datasets, aeon's feature extraction overhead provides no advantage over statsmodels' efficient analytical estimators.
  • Confidence intervals in statsmodels are theoretically derived from state-space representations, whereas aeon relies on heuristic approaches.

Frequently Asked Questions

Can aeon handle multivariate time series forecasting?

aeon offers limited multivariate support through wrappers that treat each series independently or require manual feature concatenation, but it lacks native vector autoregression (VAR) models. For true multivariate forecasting involving cross-series dependencies and Granger causality analysis, statsmodels provides dedicated VAR and VARMAX classes with proper joint estimation.

Does aeon provide statistical diagnostic tests like Ljung-Box or ADF?

No, aeon does not include built-in Ljung-Box autocorrelation tests, augmented Dickey-Fuller stationarity tests, or residual diagnostic plots. The library focuses on machine learning metrics such as MAE and MSE. statsmodels includes these diagnostics in scientific-skills/statsmodels/references/stats_diagnostics.md alongside likelihood-ratio tests and information criteria for model selection.

Which library is better for small datasets?

statsmodels is superior for small datasets because it uses exact maximum likelihood estimation without feature extraction overhead. aeon is optimized for large-scale problems where the fixed cost of transformations like ROCKET or Catch22 is justified by processing millions of series, but this overhead dominates for short univariate series.

Can aeon produce confidence intervals for forecasts?

aeon does not generate theoretically justified confidence intervals or prediction bands for most forecasters. Deep learning models in aeon may provide heuristic quantiles, but they lack the state-space derived intervals available in statsmodels through the get_forecast method, which accounts for both parameter uncertainty and innovation variance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →