How to Configure MiniMind to Use YaRN for Longer Context Windows

Enable YaRN in MiniMind by passing the --inference_rope_scaling flag or setting inference_rope_scaling=True in MiniMindConfig to extrapolate rotary positional embeddings and extend the usable context beyond the default 32,768 tokens.

MiniMind is a lightweight open-source language model that implements YaRN (Yet Another RoPE N-scaling) to handle longer input sequences without retraining or modifying model weights. Configuring MiniMind to use YaRN for longer contexts requires toggling a specific configuration flag that adjusts the rotary frequency computation in the model's attention mechanism.

Understanding YaRN in MiniMind

YaRN modifies how rotary positional embeddings (RoPE) calculate frequencies to support sequences longer than the training context. When enabled, MiniMind applies a scaling factor to the inverse frequency calculation, allowing the model to generalize to extended contexts.

According to the source code in model/model_minimind.py, setting inference_rope_scaling=True creates a rope_scaling dictionary with "type": "yarn" and a default factor of 16 (providing 4× extrapolation capability) in the MiniMindConfig class (lines 56-65).

Command-Line Configuration

The fastest way to activate YaRN is through MiniMind's CLI tools.

OpenAI-Compatible API Server

When launching the FastAPI inference server via scripts/serve_openai_api.py, append the --inference_rope_scaling argument:

python scripts/serve_openai_api.py \
    --load_from ../model \
    --weight full_sft \
    --hidden_size 512 \
    --max_seq_len 8192 \
    --inference_rope_scaling

The argument parser defines this flag at lines 70-73 of scripts/serve_openai_api.py, injecting the YaRN configuration into the model initializer.

Evaluation Scripts

For running benchmarks with extended contexts in eval_llm.py:

python eval_llm.py \
    --model_path ../model \
    --max_seq_len 16384 \
    --inference_rope_scaling

This flag is handled at lines 37-42 of eval_llm.py, enabling YaRN scaling during the evaluation phase without code modifications.

Programmatic Python Configuration

For custom inference pipelines, instantiate MiniMindConfig with YaRN enabled:

from model.model_minimind import MiniMindConfig, MiniMindModel

cfg = MiniMindConfig(
    hidden_size=512,
    num_hidden_layers=8,
    max_position_embeddings=32768,
    inference_rope_scaling=True  # Activate YaRN

)

model = MiniMindModel(cfg).to('cuda')

When inference_rope_scaling=True, the configuration object automatically populates the rope_scaling attribute with YaRN-specific parameters at lines 56-65 of model/model_minimind.py.

Customizing YaRN Parameters

To adjust the scaling behavior beyond the defaults, manually configure the rope_scaling dictionary after instantiation:

cfg = MiniMindConfig(inference_rope_scaling=True)
cfg.rope_scaling = {
    "type": "yarn",
    "factor": 32,        # 8× extrapolation

    "beta_fast": 64,
    "beta_slow": 2,
    "original_max_position_embeddings": 2048,
    "attention_factor": 1.0
}

These parameters control the frequency adjustment in precompute_freqs_cis at lines 117-124 of model/model_minimind.py, where the YaRN formula f'(i) = f(i)·((1‑γ) + γ/s) is applied to rescale rotary frequencies.

Technical Implementation Details

The YaRN implementation resides in the precompute_freqs_cis function within model/model_minimind.py. When rope_scaling is present, the code computes a linear ramp (γ) and modifies base frequencies according to the YaRN formula at lines 117-124. This rescaling allows MiniMind to maintain attention accuracy across sequences longer than the original training context.

Unlike truncation-based approaches, YaRN rescales rotary frequencies rather than limiting sequence length, preserving relative positional information for extended contexts while keeping the model weights frozen.

Summary

  • Activate YaRN using --inference_rope_scaling in CLI scripts or inference_rope_scaling=True in Python configs
  • Default scaling provides 4× extrapolation with a factor of 16 (configured in MiniMindConfig at lines 56-65)
  • Implementation modifies frequency computation in precompute_freqs_cis using the YaRN formula (lines 117-124 of model/model_minimind.py)
  • Compatibility works with serve_openai_api.py, eval_llm.py, and custom inference code without changing model weights

Frequently Asked Questions

What is the default YaRN scaling factor in MiniMind?

MiniMind defaults to a factor of 16 when YaRN is enabled, which typically supports 4× extrapolation beyond the base context length. This value is automatically set in the rope_scaling dictionary when inference_rope_scaling=True is specified in the configuration.

Can I use YaRN during training, or only for inference?

The configuration parameter is named inference_rope_scaling, indicating it is optimized for inference-time context extension. While the underlying RoPE scaling mechanism could theoretically apply to training, MiniMind's training scripts in the trainer/ directory do not expose this parameter, suggesting YaRN is primarily intended for extending context during inference.

How does YaRN differ from standard RoPE scaling methods?

YaRN (Yet Another RoPE N-scaling) calculates a linear interpolation factor (γ) and applies the formula f'(i) = f(i)·((1‑γ) + γ/s) to adjust frequencies, whereas standard linear scaling simply divides position indices by a constant factor. YaRN maintains better relative positional accuracy for longer contexts by specifically tuning the attention scale factors and frequency ramp parameters.

What is the maximum effective context length with YaRN enabled?

While MiniMind's default max_position_embeddings is 32,768 tokens, YaRN enables extrapolation significantly beyond this limit. With the default factor of 16, you can effectively process sequences up to approximately 131,072 tokens, though actual limits depend on available GPU memory and the specific scaling factor configured in the rope_scaling parameters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →