# Kronos-Mini vs Small vs Base vs Large: Complete Model Comparison and Selection Guide

> Explore Kronos model differences: Mini, Small, Base, and Large. Compare parameter counts, context windows, and hardware needs to select the best fit for your project. Discover availability now.

- Repository: [ShiYu/Kronos](https://github.com/shiyu-coder/Kronos)
- Tags: comparison-guide
- Published: 2026-04-10

---

**Kronos-Mini, Kronos-Small, Kronos-Base, and Kronos-Large differ primarily in parameter count (4.1M to 499.2M), context window size (2048 vs 512), tokenizer selection, and hardware requirements, with only Mini, Small, and Base currently available on Hugging Face.**

The `shiyu-coder/Kronos` repository provides a family of pre-trained Transformer models for time series forecasting. Understanding the distinctions between these four variants enables you to select the optimal model for your hardware constraints and accuracy requirements.

## Architecture and Core Design

All variants share the identical architecture defined in [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py) and [`model/module.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/module.py), utilizing a **hierarchical binary-spherical quantizer** combined with standard Transformer blocks. The differences stem purely from scaling factors—hidden dimensions, attention heads, and layer counts—rather than architectural changes.

## Parameter Count and Model Specifications

### Kronos-Mini (4.1M Parameters)

The lightweight **Kronos-Mini** contains **4.1 million parameters**, making it suitable for edge devices and CPU-only inference. Unique among the family, it employs the `Kronos-Tokenizer-2k` and supports an extended **2048-token context window** according to the model zoo table in [`README.md`](https://github.com/shiyu-coder/Kronos/blob/main/README.md) (lines 77-82). This variant is released as `NeoQuasar/Kronos-mini` on Hugging Face.

### Kronos-Small (24.7M Parameters)

**Kronos-Small** scales to **24.7 million parameters**, utilizing the `Kronos-Tokenizer-base` with a **512-token context length**. Available at `NeoQuasar/Kronos-small`, this model fits comfortably on single GPUs with approximately 4GB VRAM, offering a balance between speed and predictive quality.

### Kronos-Base (102.3M Parameters)

With **102.3 million parameters**, **Kronos-Base** delivers the highest accuracy among released models. It shares the same tokenizer and 512-token context as Small, but requires approximately **8GB VRAM** for efficient inference. The checkpoint is available at `NeoQuasar/Kronos-base`.

### Kronos-Large (499.2M Parameters - Unreleased)

The **Kronos-Large** variant contains **499.2 million parameters** but remains experimental. Marked with a ❌ in the model zoo documentation and [`webui/README.md`](https://github.com/shiyu-coder/Kronos/blob/main/webui/README.md) (lines 79-82), this model has not been released to Hugging Face and would require multi-GPU or high-memory hardware for deployment.

## Context Length and Tokenizer Differences

The primary functional distinction lies in sequence handling:

- **Kronos-Mini**: **2048 time steps** via `Kronos-Tokenizer-2k`, enabling longer look-back windows without truncation.
- **Kronos-Small, Base, Large**: **512 time steps** via `Kronos-Tokenizer-base`. The `KronosPredictor` class automatically truncates inputs exceeding this limit.

When instantiating the predictor for Mini, you must explicitly set `max_context=2048` (the default is 512).

## Hardware Requirements and Latency Trade-offs

| Model | Parameters | Context | Minimum VRAM | Typical Use Case |
|-------|------------|---------|--------------|------------------|
| Kronos-Mini | 4.1M | 2048 | < 2GB | Edge deployment, CPU inference, real-time dashboards |
| Kronos-Small | 24.7M | 512 | ~4GB | Single GPU, day-level forecasting |
| Kronos-Base | 102.3M | 512 | ~8GB | Production pipelines, high-accuracy research |
| Kronos-Large | 499.2M | 512 | 16GB+ | Experimental scaling (unreleased) |

## Loading and Running Inference

The [`webui/app.py`](https://github.com/shiyu-coder/Kronos/blob/main/webui/app.py) file maps model names to Hugging Face IDs, while [`model/kronos.py`](https://github.com/shiyu-coder/Kronos/blob/main/model/kronos.py) provides the loading interface. Below demonstrates switching between the three available variants:

```python
from model import Kronos, KronosTokenizer, KronosPredictor
import pandas as pd

# Configuration template - uncomment desired variant

config = {
    "mini": {
        "model": "NeoQuasar/Kronos-mini",
        "tokenizer": "NeoQuasar/Kronos-Tokenizer-2k",
        "max_context": 2048
    },
    "small": {
        "model": "NeoQuasar/Kronos-small", 
        "tokenizer": "NeoQuasar/Kronos-Tokenizer-base",
        "max_context": 512
    },
    "base": {
        "model": "NeoQuasar/Kronos-base",
        "tokenizer": "NeoQuasar/Kronos-Tokenizer-base", 
        "max_context": 512
    }
}

# Select variant

variant = "mini"  # Change to "small" or "base"

cfg = config[variant]

# Initialize model and tokenizer

tokenizer = KronosTokenizer.from_pretrained(cfg["tokenizer"])
model = Kronos.from_pretrained(cfg["model"])
predictor = KronosPredictor(model, tokenizer, max_context=cfg["max_context"])

# Load example data from examples/prediction_example.py

df = pd.read_csv("examples/data/XSHG_5min_600977.csv")
df["timestamps"] = pd.to_datetime(df["timestamps"])

# Prepare input (ensure length ≤ max_context)

lookback = min(400, cfg["max_context"])
pred_len = 120

x_df = df.iloc[:lookback][["open", "high", "low", "close", "volume", "amount"]]
x_ts = df.iloc[:lookback]["timestamps"]
y_ts = df.iloc[lookback:lookback+pred_len]["timestamps"]

# Generate forecast

forecast = predictor.predict(
    df=x_df,
    x_timestamp=x_ts,
    y_timestamp=y_ts,
    pred_len=pred_len,
    T=1.0,
    top_p=0.9,
    sample_count=1
)

```

## Summary

- **Kronos-Mini** (4.1M) provides a **2048-token context** using `Kronos-Tokenizer-2k`, optimized for CPU and edge deployment.
- **Kronos-Small** (24.7M) and **Kronos-Base** (102.3M) utilize `Kronos-Tokenizer-base` with **512-token contexts**, differing primarily in capacity (4GB vs 8GB VRAM requirements).
- **Kronos-Large** (499.2M) remains unreleased but targets high-memory research applications.
- All variants implement identical APIs in `Kronos.from_pretrained()` and `KronosPredictor`, requiring only adjustments to model IDs, tokenizer paths, and `max_context` parameters.

## Frequently Asked Questions

### What is the context length difference between Kronos-Mini and other variants?

**Kronos-Mini supports 2048 tokens**, while Kronos-Small, Base, and Large are limited to **512 tokens**. This makes Mini particularly suitable for long time series sequences that would require truncation in larger models, as implemented in the `KronosPredictor` class.

### Can I use Kronos-Base on a CPU?

While technically possible, **Kronos-Base requires approximately 8GB of memory** and performs significantly slower on CPUs compared to GPU acceleration. For CPU-only deployment, **Kronos-Mini** is the recommended choice due to its 4.1M parameter footprint and compatibility with the extended tokenizer.

### Why does Kronos-Mini use a different tokenizer?

The **Kronos-Tokenizer-2k** matches Mini's compact architecture with a reduced vocabulary and embedding dimension. Larger variants share the **Kronos-Tokenizer-base** to ensure consistent token semantics across higher-capacity models, as documented in the repository's model zoo.

### Is Kronos-Large available for download?

**No, Kronos-Large is not yet released.** According to the model zoo in [`README.md`](https://github.com/shiyu-coder/Kronos/blob/main/README.md) (lines 77-82) and [`webui/README.md`](https://github.com/shiyu-coder/Kronos/blob/main/webui/README.md) (lines 79-82), the 499.2M parameter checkpoint is marked with a ❌, indicating it remains experimental and unpublished on the Hugging Face Hub.