MiniMind Parameter Size Compared to GPT-3: From 26M to 145M Parameters Explained

MiniMind models range from 26 million to 145 million parameters, making them roughly 1/7000th the size of GPT-3's 175 billion parameters while maintaining functional chat capabilities.

The MiniMind repository by jingyaogong demonstrates how modern architectural optimizations can create ultra-lightweight large language models (LLMs) that run on consumer hardware. When analyzing the MiniMind parameter size compared to GPT-3, the scale difference is dramatic—from 0.015% to 0.08% of GPT-3's capacity—yet the architecture retains essential transformer capabilities through efficient design choices documented in the source code.

MiniMind Model Variants and Parameter Counts

The repository defines three primary model sizes in the Models List table located in README.md【README.md†L99-L105】. Each variant targets different computational budgets while sharing the same core architecture.

MiniMind2-small: The 26M Parameter Variant

MiniMind2-small contains 26 million parameters, making it the lightest production-ready variant. According to the repository documentation, this represents approximately ¼/7000 (or 1/6,730) of GPT-3's parameter count【README.md†L35-L36】. This variant uses a compact configuration with fewer hidden layers and smaller dimensionalities, defined in model/model_minimind.py where the base configuration specifies parameters like hidden_size and layer counts【model_minimind.py†L17-L22】.

MiniMind2 Base: The 104M Parameter Standard

The standard MiniMind2 (base) model scales to 104 million parameters, offering increased capacity for complex reasoning while remaining orders of magnitude smaller than GPT-3. This version balances inference speed with model capability, utilizing the same transformer decoder architecture but with expanded width and depth compared to the small variant.

MiniMind2-MoE: Scaling to 145M with Mixture-of-Experts

MiniMind2-MoE reaches 145 million parameters by implementing a Mixture-of-Experts (MoE) architecture. Unlike dense models where every parameter activates for every token, MoE routes tokens to specialized expert sub-networks. The implementation in model/model_minimind.py activates the MOEFeedForward class when use_moe=True【model_minimind.py†L71-L77】, allowing increased model capacity without linearly scaling computational requirements during inference.

The GPT-3 Comparison: 175 Billion Parameters

For context, GPT-3 contains 175 billion (175B) parameters according to the original OpenAI paper. This creates a size ratio where:

  • MiniMind2-small (26M) is approximately 0.015% of GPT-3
  • MiniMind2 base (104M) is approximately 0.06% of GPT-3
  • MiniMind2-MoE (145M) is approximately 0.08% of GPT-3

Architectural Innovations Enabling Miniaturization

MiniMind achieves this drastic parameter reduction through specific architectural optimizations that replace standard transformer components with more efficient alternatives.

RMSNorm Pre-normalization

Instead of the standard LayerNorm used in early GPT models, MiniMind implements RMSNorm (Root Mean Square Layer Normalization) for pre-normalization. This approach, noted in the README【README.md†L626-L628】, follows GPT-3's pre-norm strategy but eliminates the mean-centering computation, reducing parameter overhead while maintaining training stability.

Rotary Positional Embeddings (RoPE)

MiniMind replaces absolute positional embeddings with Rotary Positional Embeddings (RoPE). This technique encodes position information through rotation matrices applied to query and key vectors during attention computation, saving parameters compared to learned position embeddings while supporting better extrapolation to longer contexts than the original training sequence length.

SwiGLU Activation Function

The architecture utilizes SwiGLU activation functions instead of the ReLU or GELU activations common in GPT-3-style transformers. SwiGLU combines gating mechanisms with non-linear transformations in a parameter-efficient manner, offering improved gradient flow and representational capacity per parameter compared to standard feed-forward layers.

Mixture-of-Experts Architecture

The MoE variant demonstrates how MiniMind scales efficiently. By implementing sparse activation patterns where only a subset of experts process each token (controlled by routing logic in MOEFeedForward), the model increases total parameter count to 145M while keeping active parameters per token much lower. This architecture allows the model to specialize different experts for different linguistic patterns without requiring dense computation across all parameters.

Loading and Verifying MiniMind Parameter Counts

You can verify these parameter counts programmatically using the Hugging Face Transformers library and the MiniMind model implementations.

Loading a MiniMind Model

To load the 26M parameter variant and inspect its architecture:

from transformers import AutoModelForCausalLM, AutoTokenizer

# Load the smallest MiniMind variant (26M parameters)

model_name = "jingyaogong/MiniMind2-small"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    trust_remote_code=True,     # Required for custom MiniMind architecture

    torch_dtype="auto",
)

# Test generation

inputs = tokenizer("你好,请介绍一下自己。", return_tensors="pt")
output = model.generate(**ins, max_new_tokens=50)
print(tokenizer.decode(output[0], skip_special_tokens=True))

The model class is defined in model/model_minimind.py【model_minimind.py†L8-L23】 and registered with AutoModelForCausalLM through the config_class = MiniMindConfig assignment in the MiniMindForCausalLM class definition【model/model_minimind.py†L27-L33】.

Inspecting Parameters Programmatically

To calculate the approximate parameter count from configuration values:

from transformers import AutoConfig

cfg = AutoConfig.from_pretrained("jingyaogong/MiniMind2-small", trust_remote_code=True)

# Calculate approximate parameters (simplified estimation)

# For the small variant: hidden_size=512, num_hidden_layers=8

total_params = sum(p.numel() for p in model.parameters())
print(f"Model type: {cfg.model_type}")
print(f"Total parameters: {total_params:,}")

The exact architecture specifications (e.g., hidden_size=512, num_hidden_layers=8 for the small variant) are defined in the configuration classes within model/model_minimind.py【model_minimind.py†L17-L22】.

Switching to the MoE Variant

To load the larger 145M parameter MoE model:

model_name = "jingyaogong/MiniMind2-MoE"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_name, 
    trust_remote_code=True
)

# Verify the MoE layers are active

print(f"Using MoE: {model.config.use_moe}")

The MoE logic activates the MOEFeedForward layer when the configuration flag is set【model_minimind.py†L71-L77】, increasing the total parameter count through multiple expert networks while maintaining efficient inference.

Summary

  • MiniMind models range from 26M to 145M parameters, compared to GPT-3's 175B parameters.
  • The smallest variant is approximately 1/7000th the size of GPT-3, enabling deployment on resource-constrained devices.
  • Architectural optimizations including RMSNorm, RoPE, SwiGLU, and optional MoE architectures maximize capability per parameter.
  • Model configurations are defined in model/model_minimind.py with specific implementations for efficient attention and feed-forward layers.
  • All variants support the Hugging Face AutoModelForCausalLM interface via trust_remote_code=True.

Frequently Asked Questions

How many parameters does the smallest MiniMind model have?

MiniMind2-small contains 26 million parameters, as documented in the repository's model table【README.md†L99-L105】. This makes it small enough to run inference on CPUs and low-end GPUs while still providing coherent conversational capabilities.

What is the parameter size ratio between MiniMind and GPT-3?

The smallest MiniMind model is approximately 1/6,730 (or roughly 1/7000) the size of GPT-3, which has 175 billion parameters【README.md†L35-L36】. Even the largest MiniMind variant (MiniMind2-MoE at 145M) represents less than 0.1% of GPT-3's parameter count.

Why is MiniMind so much smaller than GPT-3?

MiniMind achieves its small footprint through architectural efficiency rather than just scaling down. It employs RMSNorm instead of LayerNorm, Rotary Positional Embeddings instead of absolute position encodings, and SwiGLU activations that provide better capacity per parameter. The MoE variant further optimizes by activating only subsets of parameters per token, allowing 145M total parameters without 145M active computations per inference step.

Can MiniMind run on consumer hardware?

Yes. With only 26 million to 145 million parameters, MiniMind models easily fit into consumer GPU memory (typically requiring less than 1GB VRAM for inference) and can even run on modern CPUs with reasonable latency. This contrasts with GPT-3, which requires specialized data center infrastructure with hundreds of gigabytes of memory distributed across multiple high-end GPUs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →