Kronos-Mini vs Small vs Base vs Large: Complete Model Comparison and Selection Guide

Kronos-Mini, Kronos-Small, Kronos-Base, and Kronos-Large differ primarily in parameter count (4.1M to 499.2M), context window size (2048 vs 512), tokenizer selection, and hardware requirements, with only Mini, Small, and Base currently available on Hugging Face.

The shiyu-coder/Kronos repository provides a family of pre-trained Transformer models for time series forecasting. Understanding the distinctions between these four variants enables you to select the optimal model for your hardware constraints and accuracy requirements.

Architecture and Core Design

All variants share the identical architecture defined in model/kronos.py and model/module.py, utilizing a hierarchical binary-spherical quantizer combined with standard Transformer blocks. The differences stem purely from scaling factors—hidden dimensions, attention heads, and layer counts—rather than architectural changes.

Parameter Count and Model Specifications

Kronos-Mini (4.1M Parameters)

The lightweight Kronos-Mini contains 4.1 million parameters, making it suitable for edge devices and CPU-only inference. Unique among the family, it employs the Kronos-Tokenizer-2k and supports an extended 2048-token context window according to the model zoo table in README.md (lines 77-82). This variant is released as NeoQuasar/Kronos-mini on Hugging Face.

Kronos-Small (24.7M Parameters)

Kronos-Small scales to 24.7 million parameters, utilizing the Kronos-Tokenizer-base with a 512-token context length. Available at NeoQuasar/Kronos-small, this model fits comfortably on single GPUs with approximately 4GB VRAM, offering a balance between speed and predictive quality.

Kronos-Base (102.3M Parameters)

With 102.3 million parameters, Kronos-Base delivers the highest accuracy among released models. It shares the same tokenizer and 512-token context as Small, but requires approximately 8GB VRAM for efficient inference. The checkpoint is available at NeoQuasar/Kronos-base.

Kronos-Large (499.2M Parameters - Unreleased)

The Kronos-Large variant contains 499.2 million parameters but remains experimental. Marked with a ❌ in the model zoo documentation and webui/README.md (lines 79-82), this model has not been released to Hugging Face and would require multi-GPU or high-memory hardware for deployment.

Context Length and Tokenizer Differences

The primary functional distinction lies in sequence handling:

  • Kronos-Mini: 2048 time steps via Kronos-Tokenizer-2k, enabling longer look-back windows without truncation.
  • Kronos-Small, Base, Large: 512 time steps via Kronos-Tokenizer-base. The KronosPredictor class automatically truncates inputs exceeding this limit.

When instantiating the predictor for Mini, you must explicitly set max_context=2048 (the default is 512).

Hardware Requirements and Latency Trade-offs

Model Parameters Context Minimum VRAM Typical Use Case
Kronos-Mini 4.1M 2048 < 2GB Edge deployment, CPU inference, real-time dashboards
Kronos-Small 24.7M 512 ~4GB Single GPU, day-level forecasting
Kronos-Base 102.3M 512 ~8GB Production pipelines, high-accuracy research
Kronos-Large 499.2M 512 16GB+ Experimental scaling (unreleased)

Loading and Running Inference

The webui/app.py file maps model names to Hugging Face IDs, while model/kronos.py provides the loading interface. Below demonstrates switching between the three available variants:

from model import Kronos, KronosTokenizer, KronosPredictor
import pandas as pd

# Configuration template - uncomment desired variant

config = {
    "mini": {
        "model": "NeoQuasar/Kronos-mini",
        "tokenizer": "NeoQuasar/Kronos-Tokenizer-2k",
        "max_context": 2048
    },
    "small": {
        "model": "NeoQuasar/Kronos-small", 
        "tokenizer": "NeoQuasar/Kronos-Tokenizer-base",
        "max_context": 512
    },
    "base": {
        "model": "NeoQuasar/Kronos-base",
        "tokenizer": "NeoQuasar/Kronos-Tokenizer-base", 
        "max_context": 512
    }
}

# Select variant

variant = "mini"  # Change to "small" or "base"

cfg = config[variant]

# Initialize model and tokenizer

tokenizer = KronosTokenizer.from_pretrained(cfg["tokenizer"])
model = Kronos.from_pretrained(cfg["model"])
predictor = KronosPredictor(model, tokenizer, max_context=cfg["max_context"])

# Load example data from examples/prediction_example.py

df = pd.read_csv("examples/data/XSHG_5min_600977.csv")
df["timestamps"] = pd.to_datetime(df["timestamps"])

# Prepare input (ensure length ≤ max_context)

lookback = min(400, cfg["max_context"])
pred_len = 120

x_df = df.iloc[:lookback][["open", "high", "low", "close", "volume", "amount"]]
x_ts = df.iloc[:lookback]["timestamps"]
y_ts = df.iloc[lookback:lookback+pred_len]["timestamps"]

# Generate forecast

forecast = predictor.predict(
    df=x_df,
    x_timestamp=x_ts,
    y_timestamp=y_ts,
    pred_len=pred_len,
    T=1.0,
    top_p=0.9,
    sample_count=1
)

Summary

  • Kronos-Mini (4.1M) provides a 2048-token context using Kronos-Tokenizer-2k, optimized for CPU and edge deployment.
  • Kronos-Small (24.7M) and Kronos-Base (102.3M) utilize Kronos-Tokenizer-base with 512-token contexts, differing primarily in capacity (4GB vs 8GB VRAM requirements).
  • Kronos-Large (499.2M) remains unreleased but targets high-memory research applications.
  • All variants implement identical APIs in Kronos.from_pretrained() and KronosPredictor, requiring only adjustments to model IDs, tokenizer paths, and max_context parameters.

Frequently Asked Questions

What is the context length difference between Kronos-Mini and other variants?

Kronos-Mini supports 2048 tokens, while Kronos-Small, Base, and Large are limited to 512 tokens. This makes Mini particularly suitable for long time series sequences that would require truncation in larger models, as implemented in the KronosPredictor class.

Can I use Kronos-Base on a CPU?

While technically possible, Kronos-Base requires approximately 8GB of memory and performs significantly slower on CPUs compared to GPU acceleration. For CPU-only deployment, Kronos-Mini is the recommended choice due to its 4.1M parameter footprint and compatibility with the extended tokenizer.

Why does Kronos-Mini use a different tokenizer?

The Kronos-Tokenizer-2k matches Mini's compact architecture with a reduced vocabulary and embedding dimension. Larger variants share the Kronos-Tokenizer-base to ensure consistent token semantics across higher-capacity models, as documented in the repository's model zoo.

Is Kronos-Large available for download?

No, Kronos-Large is not yet released. According to the model zoo in README.md (lines 77-82) and webui/README.md (lines 79-82), the 499.2M parameter checkpoint is marked with a ❌, indicating it remains experimental and unpublished on the Hugging Face Hub.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →