Key Parameters for Fine-Tuning LLMs with Hugging Face: A Complete Guide
Fine-tuning LLMs with Hugging Face requires configuring model loading identifiers, optimization hyperparameters in TrainingArguments, and data handling settings through the Trainer API.
Fine-tuning large language models (LLMs) with Hugging Face Transformers involves orchestrating multiple configuration layers, from checkpoint loading to gradient optimization. This guide examines the key parameters used in the Lordog/dive-into-llms repository, referencing actual implementations in documents/chapter1/dive-tuning.ipynb and documents/chapter4/sft_math.ipynb to provide actionable, reproducible configurations.
Model and Tokenizer Initialization
The model_name_or_path argument specifies which pre-trained checkpoint to load as the starting point for fine-tuning. This parameter accepts Hugging Face Hub identifiers like bert-base-uncased or local paths to saved checkpoints.
In documents/chapter1/dive-tuning.ipynb, the classification example uses --model_name_or_path bert-base-uncased to initialize the encoder【dive‑tuning.ipynb, L133‑L136】. The same identifier is reused when saving checkpoints, ensuring consistency between training iterations and model deployment.
Core Training Hyperparameters
The TrainingArguments class (or CLI equivalents) controls optimization, memory usage, and reproducibility. These parameters determine how the model learns from your dataset.
Optimization Parameters
learning_rate controls the step size for gradient updates. The repository demonstrates task-specific variations: 2e-5 for general classification【dive‑tuning.ipynb, L149‑L151】 and 1e-6 for math-specific fine-tuning【sft_math.ipynb, L806‑L808】.
lr_scheduler_type determines how the learning rate changes over time. The math fine-tuning example uses "cosine" annealing【sft_math.ipynb, L807‑L809】.
num_train_epochs specifies full passes over the dataset. Values range from 1 for quick classification experiments【dive‑tuning.ipynb, L150‑L151】 to 3 for thorough math instruction tuning【sft_math.ipynb, L803‑L806】.
Memory and Batch Management
per_device_train_batch_size sets the batch size per GPU/CPU. The classification script uses 32【dive‑tuning.ipynb, L147‑L149】 while memory-constrained math fine-tuning uses 4 with gradient_accumulation_steps=16 to simulate an effective batch size of 64【sft_math.ipynb, L802‑L804】.
max_seq_length truncates or pads input sequences to a fixed token count, directly impacting both computational cost and model performance. In documents/chapter1/dive-tuning.ipynb, the classification task sets max_seq_length=512 to balance context coverage with memory efficiency【dive‑tuning.ipynb, L146‑L148】.
gradient_checkpointing trades computation for memory by recomputing activations during backward passes. It defaults to False but can be enabled on low-memory GPUs【sft_math.ipynb, L820‑L822】.
Precision and Reproducibility
bf16 and fp16 enable mixed-precision training to reduce memory footprint and accelerate computation. The math fine-tuning script explicitly sets bf16=True for training the Qwen1.5-7B model【sft_math.ipynb, L817‑L819】. Use bf16 when your hardware supports it (e.g., NVIDIA A100, H100, or newer GPUs).
seed and data_seed guarantee reproducibility across runs. The repository consistently uses 42 for both parameters【sft_math.ipynb, L822‑L824】.
Data Handling and Trainer Orchestration
Training pipelines require explicit paths for datasets and checkpoints. The output_dir parameter specifies where checkpoints, logs, and final models are stored. The classification script uses experiments/【dive‑tuning.ipynb, L151‑L152】 while the math fine-tuning example uses ./checkpoints【sft_math.ipynb, L809‑L811】.
Dataset locations are specified via --train_file, --validation_file, and --test_file, pointing to CSV or JSON files as shown in the classification example【dive‑tuning.ipynb, L135‑L138】. The DataCollatorForSeq2Seq handles dynamic padding and batching when working with sequence-to-sequence models.
The Trainer class abstracts the training loop, evaluation, checkpointing, and logging into a unified API that orchestrates the model, data, and hyperparameters. In documents/chapter4/sft_math.ipynb, the Trainer is instantiated with the model, training_args, dataset, processing_class (tokenizer), and data_collator【sft_math.ipynb, L842‑L850】. For reinforcement learning from human feedback (RLHF), the repository demonstrates PPOTrainer in documents/chapter11/RLHF.ipynb.
Practical Implementation Examples
Text Classification via CLI
The following command from documents/chapter1/dive-tuning.ipynb demonstrates a complete fine-tuning workflow for classification tasks:
python run_classification.py \
--model_name_or_path bert-base-uncased \
--train_file data/train.csv \
--validation_file data/val.csv \
--test_file data/test.csv \
--shuffle_train_dataset \
--metric_name accuracy \
--text_column_name "text" \
--text_column_delimiter "\n" \
--label_column_name "target" \
--do_train \
--do_eval \
--do_predict \
--max_seq_length 512 \
--per_device_train_batch_size 32 \
--learning_rate 2e-5 \
--num_train_epochs 1 \
--output_dir experiments/
Source: dive-tuning.ipynb, lines 133‑151【dive‑tuning.ipynb, L133‑L151】.
Seq-to-Seq Fine-Tuning with Python API
For math instruction tuning, documents/chapter4/sft_math.ipynb shows the programmatic approach:
from transformers import (
AutoTokenizer, AutoModelForCausalLM,
Trainer, TrainingArguments, DataCollatorForSeq2Seq
)
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen1.5-7B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen1.5-7B")
training_args = TrainingArguments(
per_device_train_batch_size=4,
gradient_accumulation_steps=16,
num_train_epochs=3,
learning_rate=1e-6,
lr_scheduler_type="cosine",
output_dir="./checkpoints",
bf16=True,
gradient_checkpointing=False,
seed=42,
data_seed=42,
)
collator = DataCollatorForSeq2Seq(
tokenizer,
padding=True,
pad_to_multiple_of=8,
return_tensors="pt",
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
processing_class=tokenizer,
data_collator=collator,
)
trainer.train()
trainer.save_model("./checkpoints/final_model")
Source: sft_math.ipynb, lines 800‑822 and 842‑850【sft_math.ipynb, L800‑L822】, 【sft_math.ipynb, L842‑L850】.
Summary
model_name_or_pathspecifies the pre-trained checkpoint to initialize weights, used in bothAutoModelclasses and CLI scripts.learning_rateandlr_scheduler_typecontrol optimization dynamics, with typical values ranging from1e-6to2e-5depending on task complexity.per_device_train_batch_sizecombined withgradient_accumulation_stepsdetermines effective batch size and memory consumption.max_seq_lengthtruncates inputs to control computational cost, commonly set to512tokens for classification tasks.bf16/fp16andgradient_checkpointingenable mixed-precision training and memory optimization for large models.seedanddata_seedensure reproducibility across training runs.output_dirdefines checkpoint storage locations, whiletrain_file/**validation_file**specify dataset paths.
Frequently Asked Questions
What is the difference between per_device_train_batch_size and gradient_accumulation_steps?
per_device_train_batch_size sets the number of samples processed on each GPU or CPU at one time, while gradient_accumulation_steps simulates a larger batch size by accumulating gradients over multiple forward passes before updating weights. In documents/chapter4/sft_math.ipynb, the configuration uses per_device_train_batch_size=4 with gradient_accumulation_steps=16 to achieve an effective batch size of 64 without requiring proportional GPU memory【sft_math.ipynb, L802‑L804】.
When should I use bf16 instead of fp16 for fine-tuning?
bf16 (Bfloat16) provides the same dynamic range as FP32 while maintaining reduced precision, offering better numerical stability than fp16 for training large language models. According to documents/chapter4/sft_math.ipynb, bf16=True is explicitly set for math fine-tuning on the Qwen1.5-7B model to accelerate training while preventing gradient underflow【sft_math.ipynb, L817‑L819】. Use bf16 when your hardware supports it (e.g., NVIDIA A100, H100, or newer GPUs).
How does max_seq_length affect fine-tuning performance and memory?
max_seq_length truncates or pads input sequences to a fixed token count, directly impacting both computational cost and model performance. In documents/chapter1/dive-tuning.ipynb, the classification task sets max_seq_length=512 to balance context coverage with memory efficiency【dive‑tuning.ipynb, L146‑L148】. Longer sequences increase GPU memory consumption quadratically for attention mechanisms and reduce the number of samples processed per batch, while shorter sequences may truncate critical context.
What is the role of the Trainer class in Hugging Face fine-tuning?
The Trainer class abstracts the training loop, evaluation, checkpointing, and logging into a unified API that orchestrates the model, data, and hyperparameters. In documents/chapter4/sft_math.ipynb, the Trainer is instantiated with the model, training_args, dataset, processing_class (tokenizer), and data_collator【sft_math.ipynb, L842‑L850】. For reinforcement learning from human feedback (RLHF), the repository demonstrates PPOTrainer in documents/chapter11/RLHF.ipynb.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →