# Key Parameters for Fine-Tuning LLMs with Hugging Face: A Complete Guide

> Master fine-tuning LLMs with Hugging Face. Discover essential parameters for model loading, optimizer hyperparameters, and data handling using the Trainer API. Optimize your models efficiently.

- Repository: [Tongxin Yuan/dive-into-llms](https://github.com/Lordog/dive-into-llms)
- Tags: how-to-guide
- Published: 2026-04-16

---

**Fine-tuning LLMs with Hugging Face requires configuring model loading identifiers, optimization hyperparameters in `TrainingArguments`, and data handling settings through the `Trainer` API.**

Fine-tuning large language models (LLMs) with Hugging Face Transformers involves orchestrating multiple configuration layers, from checkpoint loading to gradient optimization. This guide examines the key parameters used in the `Lordog/dive-into-llms` repository, referencing actual implementations in `documents/chapter1/dive-tuning.ipynb` and `documents/chapter4/sft_math.ipynb` to provide actionable, reproducible configurations.

## Model and Tokenizer Initialization

The `model_name_or_path` argument specifies which pre-trained checkpoint to load as the starting point for fine-tuning. This parameter accepts Hugging Face Hub identifiers like `bert-base-uncased` or local paths to saved checkpoints.

In `documents/chapter1/dive-tuning.ipynb`, the classification example uses `--model_name_or_path bert-base-uncased` to initialize the encoder【[dive‑tuning.ipynb, L133‑L136](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter1/dive-tuning.ipynb#L133)】. The same identifier is reused when saving checkpoints, ensuring consistency between training iterations and model deployment.

## Core Training Hyperparameters

The `TrainingArguments` class (or CLI equivalents) controls optimization, memory usage, and reproducibility. These parameters determine how the model learns from your dataset.

### Optimization Parameters

**`learning_rate`** controls the step size for gradient updates. The repository demonstrates task-specific variations: `2e-5` for general classification【[dive‑tuning.ipynb, L149‑L151](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter1/dive-tuning.ipynb#L149)】 and `1e-6` for math-specific fine-tuning【[sft_math.ipynb, L806‑L808](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L806)】.

**`lr_scheduler_type`** determines how the learning rate changes over time. The math fine-tuning example uses `"cosine"` annealing【[sft_math.ipynb, L807‑L809](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L807)】.

**`num_train_epochs`** specifies full passes over the dataset. Values range from `1` for quick classification experiments【[dive‑tuning.ipynb, L150‑L151](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter1/dive-tuning.ipynb#L150)】 to `3` for thorough math instruction tuning【[sft_math.ipynb, L803‑L806](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L803)】.

### Memory and Batch Management

**`per_device_train_batch_size`** sets the batch size per GPU/CPU. The classification script uses `32`【[dive‑tuning.ipynb, L147‑L149](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter1/dive-tuning.ipynb#L147)】 while memory-constrained math fine-tuning uses `4` with `gradient_accumulation_steps=16` to simulate an effective batch size of 64【[sft_math.ipynb, L802‑L804](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L802)】.

**`max_seq_length`** truncates or pads input sequences to a fixed token count, directly impacting both computational cost and model performance. In `documents/chapter1/dive-tuning.ipynb`, the classification task sets `max_seq_length=512` to balance context coverage with memory efficiency【[dive‑tuning.ipynb, L146‑L148](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter1/dive-tuning.ipynb#L146)】.

**`gradient_checkpointing`** trades computation for memory by recomputing activations during backward passes. It defaults to `False` but can be enabled on low-memory GPUs【[sft_math.ipynb, L820‑L822](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L820)】.

### Precision and Reproducibility

**`bf16`** and **`fp16`** enable mixed-precision training to reduce memory footprint and accelerate computation. The math fine-tuning script explicitly sets `bf16=True` for training the Qwen1.5-7B model【[sft_math.ipynb, L817‑L819](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L817)】. Use `bf16` when your hardware supports it (e.g., NVIDIA A100, H100, or newer GPUs).

**`seed`** and **`data_seed`** guarantee reproducibility across runs. The repository consistently uses `42` for both parameters【[sft_math.ipynb, L822‑L824](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L822)】.

## Data Handling and Trainer Orchestration

Training pipelines require explicit paths for datasets and checkpoints. The **`output_dir`** parameter specifies where checkpoints, logs, and final models are stored. The classification script uses `experiments/`【[dive‑tuning.ipynb, L151‑L152](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter1/dive-tuning.ipynb#L151)】 while the math fine-tuning example uses `./checkpoints`【[sft_math.ipynb, L809‑L811](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L809)】.

Dataset locations are specified via **`--train_file`**, **`--validation_file`**, and **`--test_file`**, pointing to CSV or JSON files as shown in the classification example【[dive‑tuning.ipynb, L135‑L138](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter1/dive-tuning.ipynb#L135)】. The `DataCollatorForSeq2Seq` handles dynamic padding and batching when working with sequence-to-sequence models.

The `Trainer` class abstracts the training loop, evaluation, checkpointing, and logging into a unified API that orchestrates the model, data, and hyperparameters. In `documents/chapter4/sft_math.ipynb`, the `Trainer` is instantiated with the model, `training_args`, dataset, `processing_class` (tokenizer), and `data_collator`【[sft_math.ipynb, L842‑L850](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L842)】. For reinforcement learning from human feedback (RLHF), the repository demonstrates `PPOTrainer` in `documents/chapter11/RLHF.ipynb`.

## Practical Implementation Examples

### Text Classification via CLI

The following command from `documents/chapter1/dive-tuning.ipynb` demonstrates a complete fine-tuning workflow for classification tasks:

```bash
python run_classification.py \
    --model_name_or_path bert-base-uncased \
    --train_file data/train.csv \
    --validation_file data/val.csv \
    --test_file data/test.csv \
    --shuffle_train_dataset \
    --metric_name accuracy \
    --text_column_name "text" \
    --text_column_delimiter "\n" \
    --label_column_name "target" \
    --do_train \
    --do_eval \
    --do_predict \
    --max_seq_length 512 \
    --per_device_train_batch_size 32 \
    --learning_rate 2e-5 \
    --num_train_epochs 1 \
    --output_dir experiments/

```

*Source: `dive-tuning.ipynb`, lines 133‑151【[dive‑tuning.ipynb, L133‑L151](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter1/dive-tuning.ipynb#L133)】.*

### Seq-to-Seq Fine-Tuning with Python API

For math instruction tuning, `documents/chapter4/sft_math.ipynb` shows the programmatic approach:

```python
from transformers import (
    AutoTokenizer, AutoModelForCausalLM,
    Trainer, TrainingArguments, DataCollatorForSeq2Seq
)

tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen1.5-7B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen1.5-7B")

training_args = TrainingArguments(
    per_device_train_batch_size=4,
    gradient_accumulation_steps=16,
    num_train_epochs=3,
    learning_rate=1e-6,
    lr_scheduler_type="cosine",
    output_dir="./checkpoints",
    bf16=True,
    gradient_checkpointing=False,
    seed=42,
    data_seed=42,
)

collator = DataCollatorForSeq2Seq(
    tokenizer,
    padding=True,
    pad_to_multiple_of=8,
    return_tensors="pt",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=train_dataset,
    processing_class=tokenizer,
    data_collator=collator,
)

trainer.train()
trainer.save_model("./checkpoints/final_model")

```

*Source: `sft_math.ipynb`, lines 800‑822 and 842‑850【[sft_math.ipynb, L800‑L822](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L800)】, 【[sft_math.ipynb, L842‑L850](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L842)】.*

## Summary

- **`model_name_or_path`** specifies the pre-trained checkpoint to initialize weights, used in both `AutoModel` classes and CLI scripts.
- **`learning_rate`** and **`lr_scheduler_type`** control optimization dynamics, with typical values ranging from `1e-6` to `2e-5` depending on task complexity.
- **`per_device_train_batch_size`** combined with **`gradient_accumulation_steps`** determines effective batch size and memory consumption.
- **`max_seq_length`** truncates inputs to control computational cost, commonly set to `512` tokens for classification tasks.
- **`bf16`**/`fp16` and **`gradient_checkpointing`** enable mixed-precision training and memory optimization for large models.
- **`seed`** and **`data_seed`** ensure reproducibility across training runs.
- **`output_dir`** defines checkpoint storage locations, while **`train_file`**/`**validation_file**` specify dataset paths.

## Frequently Asked Questions

### What is the difference between per_device_train_batch_size and gradient_accumulation_steps?

**`per_device_train_batch_size`** sets the number of samples processed on each GPU or CPU at one time, while **`gradient_accumulation_steps`** simulates a larger batch size by accumulating gradients over multiple forward passes before updating weights. In `documents/chapter4/sft_math.ipynb`, the configuration uses `per_device_train_batch_size=4` with `gradient_accumulation_steps=16` to achieve an effective batch size of 64 without requiring proportional GPU memory【[sft_math.ipynb, L802‑L804](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L802)】.

### When should I use bf16 instead of fp16 for fine-tuning?

**`bf16`** (Bfloat16) provides the same dynamic range as FP32 while maintaining reduced precision, offering better numerical stability than **`fp16`** for training large language models. According to `documents/chapter4/sft_math.ipynb`, `bf16=True` is explicitly set for math fine-tuning on the Qwen1.5-7B model to accelerate training while preventing gradient underflow【[sft_math.ipynb, L817‑L819](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L817)】. Use `bf16` when your hardware supports it (e.g., NVIDIA A100, H100, or newer GPUs).

### How does max_seq_length affect fine-tuning performance and memory?

**`max_seq_length`** truncates or pads input sequences to a fixed token count, directly impacting both computational cost and model performance. In `documents/chapter1/dive-tuning.ipynb`, the classification task sets `max_seq_length=512` to balance context coverage with memory efficiency【[dive‑tuning.ipynb, L146‑L148](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter1/dive-tuning.ipynb#L146)】. Longer sequences increase GPU memory consumption quadratically for attention mechanisms and reduce the number of samples processed per batch, while shorter sequences may truncate critical context.

### What is the role of the Trainer class in Hugging Face fine-tuning?

The **`Trainer`** class abstracts the training loop, evaluation, checkpointing, and logging into a unified API that orchestrates the model, data, and hyperparameters. In `documents/chapter4/sft_math.ipynb`, the `Trainer` is instantiated with the model, `training_args`, dataset, `processing_class` (tokenizer), and `data_collator`【[sft_math.ipynb, L842‑L850](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter4/sft_math.ipynb#L842)】. For reinforcement learning from human feedback (RLHF), the repository demonstrates `PPOTrainer` in `documents/chapter11/RLHF.ipynb`.