Recommended Hardware Configurations for Training Different Model Sizes on Hugging Face Jobs: The Complete Guide

Select GPU flavors based on model parameter count using the VRAM rule of thumb—t4-small for tiny models (<1B), a10g series for small-to-large models (1B–13B), and multi-GPU a10g-largex4 or a100-large for very-large models (>13B)—to optimize both cost efficiency and memory fit.

The huggingface/skills repository provides definitive guidance for matching compute resources to model sizes through its detailed hardware documentation. According to the source code in skills/hugging-face-jobs/references/hardware_guide.md and SKILL.md, selecting the proper hardware flavor ensures you avoid over-provisioning expensive GPUs while preventing out-of-memory errors during training.

Hardware Selection Matrix by Model Size

The hardware guide implements a tiered recommendation system based on parameter count and training methodology. The following matrix maps model sizes to specific Hugging Face Jobs flavors:

Model Size Minimum VRAM Recommended Flavor Budget Tier Use Case
Tiny (<1B) 2–4 GB t4-small (16 GB) <$5/hour Testing, prototyping
Small (1B–3B) 4–8 GB t4-medium (16 GB) or a10g-small (24 GB) $5–$20/hour Fine-tuning small LLMs
Medium (3B–7B) 8–16 GB a10g-small (24 GB) or a10g-large (24 GB) $20–$50/hour Production LoRA training
Large (7B–13B) 16–28 GB a10g-large (24 GB) or a100-large (40 GB) $50–$200/hour Full-parameter or LoRA fine-tuning
Very-Large (>13B) ≥28 GB a100-large (40 GB) or a10g-largex4 (96 GB) $200+/hour Multi-GPU distributed training

The VRAM Calculation Rule

As documented in references/hardware_guide.md (lines 131–140), memory requirements follow precise formulas based on training strategy:

  • Inference: Memory (GB) ≈ params_in_billions × 2–4
  • Full fine-tuning: Memory (GB) ≈ params_in_billions × 20
  • LoRA/QLoRA fine-tuning: Memory (GB) ≈ params_in_billions × 4

For example, a 7B model requires approximately 28 GB for LoRA fine-tuning (7 × 4), making the 24 GB a10g-large suitable with gradient checkpointing, while full fine-tuning requires 140 GB, necessitating multi-GPU configurations.

Flavor-Specific Recommendations

Tiny Models (<1B Parameters)

For models under 1 billion parameters, the t4-small flavor provides 16 GB VRAM—four times the minimum requirement. This configuration handles small-scale experiments and tokenizer training efficiently.

script = """

# /// script

# dependencies = ["transformers", "torch"]

# ///

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "facebook/opt-125m"  # 0.125B parameters

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id).to("cuda")

inputs = tokenizer("Hello, world!", return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0]))
"""

hf_jobs("uv", {
    "script": script,
    "flavor": "t4-small",
    "timeout": "30m",
    "secrets": {"HF_TOKEN": "$HF_TOKEN"}
})

Small to Medium Models (1B–7B Parameters)

The a10g-small and a10g-large flavors both provide 24 GB VRAM, supporting LoRA fine-tuning for models up to 7B parameters. According to the hardware guide (lines 114–120), these A10G instances represent the optimal price-performance balance for production training jobs.

script = """

# /// script

# dependencies = ["peft", "transformers", "datasets", "torch"]

# ///

from transformers import AutoModelForCausalLM, AutoTokenizer, Trainer, TrainingArguments
from peft import LoraConfig, get_peft_model
from datasets import load_dataset
import torch

model_id = "meta-llama/Meta-Llama-3-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

# LoRA reduces VRAM to ~4GB overhead

lora_config = LoraConfig(r=8, lora_alpha=32, target_modules=["q_proj", "v_proj"])
model = get_peft_model(model, lora_config)

dataset = load_dataset("json", data_files="train.json")["train"]

args = TrainingArguments(
    output_dir="/tmp/adapter",
    per_device_train_batch_size=2,
    num_train_epochs=1,
    learning_rate=2e-4,
    fp16=True
)

Trainer(model=model, args=args, train_dataset=dataset).train()
"""

hf_jobs("uv", {
    "script": script,
    "flavor": "a10g-large",
    "timeout": "8h",
    "secrets": {"HF_TOKEN": "$HF_TOKEN"}
})

Large and Very-Large Models (>7B Parameters)

For models exceeding 13 billion parameters, single-GPU configurations prove insufficient. The hardware guide recommends a100-large (40 GB) for single-GPU training or multi-GPU flavors like a10g-largex4 (96 GB total) for distributed training via Hugging Face Accelerate.

script = """

# /// script

# dependencies = ["accelerate", "transformers", "torch"]

# ///

from accelerate import Accelerator
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

accelerator = Accelerator()
model_id = "meta-llama/Meta-Llama-3-70B"

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"  # Distributes across 4 GPUs

)

tokenizer = AutoTokenizer.from_pretrained(model_id)

# Training loop with accelerator

inputs = tokenizer("Training data", return_tensors="pt")
inputs = {k: v.to(accelerator.device) for k, v in inputs.items()}
outputs = model(**inputs, labels=inputs["input_ids"])
accelerator.backward(outputs.loss)
"""

hf_jobs("uv", {
    "script": script,
    "flavor": "a10g-largex4",
    "timeout": "24h",
    "secrets": {"HF_TOKEN": "$HF_TOKEN"}
})

Compute Architecture Selection

The repository distinguishes between three compute categories in SKILL.md:

  • CPU flavors (cpu-basic, cpu-upgrade): Optimal for data preprocessing, statistical analysis, and tokenization workflows that don't require GPU acceleration
  • GPU flavors (t4-*, a10g-*, a100-*, l4-*): Required for all deep learning training; select based on the VRAM calculation matrix above
  • TPU flavors (v5e-1x1, v5e-2x2, v5e-2x4): Optimized for JAX/Flax pipelines requiring massive parallel computation

Key Reference Files in the Repository

The hardware recommendations derive from specific source files within huggingface/skills:

Summary

  • Calculate VRAM needs using the parameter-based formulas: ×4 for LoRA, ×20 for full fine-tuning, ×2–4 for inference
  • Start with the smallest sufficient GPU that meets your VRAM requirement, scaling up only when encountering OOM errors
  • Use t4-small (16 GB) for models under 1B parameters and initial prototyping
  • Select a10g-large (24 GB) for 3B–13B parameter models using LoRA techniques
  • Deploy multi-GPU flavors (a10g-largex4, a100-large) for models exceeding 13B parameters or full fine-tuning scenarios
  • Set explicit timeouts and configure Hub authentication to prevent data loss during long training runs

Frequently Asked Questions

What is the minimum VRAM required for fine-tuning a 7B parameter model?

Fine-tuning a 7B model requires approximately 28 GB VRAM for LoRA-based methods (7 × 4 GB) or 140 GB for full-parameter training (7 × 20 GB). The a10g-large flavor provides 24 GB VRAM, which accommodates 7B LoRA fine-tuning with gradient checkpointing enabled, while a100-large (40 GB) handles the workload without additional optimization techniques.

How do I choose between T4, A10G, and A100 GPUs for training?

Select based on model size and budget constraints. T4 GPUs (t4-small, t4-medium) offer 16 GB VRAM suitable for models under 3B parameters at the lowest cost tier. A10G GPUs (a10g-small, a10g-large) provide 24 GB VRAM and represent the optimal price-performance ratio for 3B–13B models. A100 GPUs (a100-large) deliver 40 GB VRAM necessary for models exceeding 13B parameters or when implementing full fine-tuning on 7B+ models.

Can I train large models on Hugging Face Jobs without multi-GPU setups?

Single GPU training is feasible for models up to approximately 13B parameters using memory-efficient techniques like LoRA, 4-bit quantization, and gradient checkpointing on a100-large (40 GB). However, models exceeding 13B parameters or requiring full fine-tuning necessitate multi-GPU flavors such as a10g-largex2 (48 GB) or a10g-largex4 (96 GB) to distribute layers across devices using Hugging Face Accelerate.

What timeout should I set for different model training jobs?

Timeout configuration depends on dataset size and training steps. Tiny models (<1B) typically require 30 minutes to 2 hours. Small-to-medium models (1B–7B) with LoRA need 4–8 hours for single-epoch fine-tuning. Large models (7B–13B) require 12–24 hours, while very-large models (>13B) on multi-GPU setups should specify 24+ hour timeouts to accommodate checkpoint saving and validation loops.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →