# What Are the Dependencies for Fine-Tuning AI Agents in the ai-agent-book Repository?

> Discover the essential dependencies for fine-tuning AI agents in the ai-agent-book repository. Learn about PyTorch, Hugging Face Transformers, Accelerate, and more for efficient model training.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: tutorial
- Published: 2026-08-22

---

**Fine-tuning AI agents in the bojieli/ai-agent-book repository requires PyTorch, Hugging Face Transformers, Accelerate, Datasets, and TRL as core dependencies, with optional PEFT for LoRA adapters, BitsAndBytes or Unsloth for quantization, and Weights & Biases for experiment tracking.**

The repository implements a modular training pipeline for agent fine-tuning across multiple modalities. Whether you are performing prompt distillation, chain-of-thought training, or multilingual reasoning, the codebase relies on a standardized Hugging Face ecosystem with specific version constraints documented in per-experiment [`requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/requirements.txt) files.

## Core Machine Learning Stack

The foundation of agent fine-tuning rests on five essential libraries that handle tensor operations, model architecture, data processing, and training orchestration.

### PyTorch and Transformers

**`torch`** provides the underlying tensor computation engine, while **`transformers`** supplies the model and tokenizer classes. According to the source code in [`chapter8/prompt-distillation/train_sft_trl.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/prompt-distillation/train_sft_trl.py), models are instantiated using:

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    trust_remote_code=True,
    use_cache=False,
)

```

### Training Orchestration with Accelerate and TRL

**`accelerate`** handles distributed training across multiple GPUs, and **`trl`** (Transformer Reinforcement Learning) provides high-level trainers including `SFTTrainer`, `DPOTrainer`, and `GRPOTrainer`. The repository uses these to abstract away training loop boilerplate while maintaining full hyperparameter control.

### Data Pipeline with Datasets

**`datasets`** manages large-scale data loading. The training scripts load JSONL files using `Dataset.from_list`, as implemented in [`chapter8/prompt-distillation/train_sft_trl.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/prompt-distillation/train_sft_trl.py):

```python
from datasets import Dataset
import json

def load_jsonl(file_path: str) -> Dataset:
    data = []
    with open(file_path, "r", encoding="utf-8") as f:
        for line in f:
            if line.strip():
                data.append(json.loads(line))
    return Dataset.from_list(data)

```

## Parameter-Efficient Fine-Tuning (PEFT)

For memory-efficient training, the repository leverages **`peft`** to implement **LoRA** (Low-Rank Adaptation) adapters. This allows fine-tuning by updating only a small subset of model weights, drastically reducing GPU memory requirements.

In [`chapter8/prompt-distillation/train_sft_trl.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/prompt-distillation/train_sft_trl.py), LoRA configuration is applied as follows:

```python
from peft import LoraConfig, get_peft_model

peft_cfg = LoraConfig(
    r=32, 
    lora_alpha=16, 
    target_modules="all-linear", 
    lora_dropout=0.0,
    bias="none", 
    task_type="CAUSAL_LM"
)
model = get_peft_model(model, peft_cfg)

```

## Optional Optimization and Quantization

### BitsAndBytes for 4-Bit Training

**`bitsandbytes`** enables 4-bit quantized training, allowing very large models to fit on consumer GPUs. This is specifically used in the multilingual reasoning experiments located in `chapter8/MultilingualReasoning/`.

### Unsloth for Accelerated Training

**`unsloth`** provides optimized training kernels that significantly speed up LoRA fine-tuning. The speech SFT experiment at [`chapter8/speech-sft-experiment/run_sesame.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/speech-sft-experiment/run_sesame.py) includes `unsloth` as a primary dependency for faster convergence.

## Inference Acceleration with vLLM

When generating synthetic training data from teacher models, the repository uses **`vllm`** for fast batched inference. This is particularly relevant for prompt distillation workflows where a larger teacher model generates training examples for a smaller student model.

```python
from vllm import LLM, SamplingParams

llm = LLM(model="meta-llama/Meta-Llama-3-8B", tensor_parallel_size=2)
params = SamplingParams(temperature=0.7, max_tokens=512)
output = llm.generate([prompt], params)

```

## Experiment Tracking and Logging

**`wandb`** (Weights & Biases) is the primary tool for logging metrics, hyperparameters, and checkpoint artifacts. Some experiments also support **`trackio`** or TensorBoard as alternatives. The training scripts configure logging via the `report_to` parameter in `TrainingArguments`.

## Installation Requirements by Experiment

Different chapters in the repository specify exact version constraints in their respective [`requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/requirements.txt) files:

- **Prompt Distillation** ([`chapter8/prompt-distillation/requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/prompt-distillation/requirements.txt)): `torch>=2.0.0`, `transformers>=4.36.0`, `datasets>=2.14.0`, `accelerate>=0.28.0`, `trl>=0.9.0`, `peft>=0.11.0`, `wandb>=0.15.0`

- **CoT Distillation** ([`chapter8/cot-distillation/requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/cot-distillation/requirements.txt)): `torch>=2.3,<3`, `transformers==4.48.3`, `accelerate==1.2.1`, `peft==0.14.0`, `trl==0.22`

- **Multilingual Reasoning** ([`chapter8/MultilingualReasoning/requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/MultilingualReasoning/requirements.txt)): `torch>=2.0.0`, `transformers>=4.55.0`, `datasets>=2.14.0`, `accelerate>=0.20.0`, `trl>=0.20.0`, `peft>=0.17.0`, `bitsandbytes>=0.41.0`, plus optional `trackio`

- **Speech SFT** ([`chapter8/speech-sft-experiment/requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/speech-sft-experiment/requirements.txt)): `torch>=2.10,<2.11`, `transformers==4.57.6`, `datasets==3.6.0`, `trl==0.24.0`, `peft==0.19.0`, `unsloth==2026.8.2` with integrated `wandb` support

- **Premature-Completion DPO** ([`chapter8/premature-completion/requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/premature-completion/requirements.txt)): `torch>=2.4`, `transformers>=4.55,<6`, `trl>=0.22`, `peft>=0.17`, `datasets>=3.4`, `accelerate>=0.30`, `openai>=1.68`

## Complete Training Example

To fine-tune an agent using the standard pipeline, combine these components as shown in [`chapter8/prompt-distillation/train_sft_trl.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/prompt-distillation/train_sft_trl.py):

```python
from trl import SFTTrainer, TrainingArguments

trainer = SFTTrainer(
    model=model,
    tokenizer=tokenizer,
    train_dataset=train_ds,
    args=TrainingArguments(
        output_dir="./outputs",
        num_train_epochs=1,
        per_device_train_batch_size=4,
        gradient_accumulation_steps=4,
        learning_rate=2e-4,
        logging_steps=1,
        save_strategy="epoch",
        report_to="wandb",
        run_name="agent-finetune",
    ),
    peft_config=peft_cfg,
)

trainer.train()

```

## Summary

- **Core dependencies** for fine-tuning AI agents include `torch`, `transformers`, `accelerate`, `datasets`, and `trl` for the complete training pipeline.
- **PEFT** (`peft`) enables memory-efficient training through LoRA adapters, requiring only gigabytes instead of tens of gigabytes of VRAM.
- **Quantization** libraries like `bitsandbytes` and `unsloth` allow 4-bit training and accelerated kernels for large-scale experiments.
- **vLLM** provides high-throughput inference for teacher models generating synthetic training data.
- **Experiment tracking** relies on `wandb` or `trackio`, integrated directly into the `TrainingArguments` configuration.
- Each experiment in `chapter8/` maintains its own [`requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/requirements.txt) with tested version pins to ensure reproducibility across prompt distillation, CoT training, multilingual reasoning, and speech modalities.

## Frequently Asked Questions

### What is the minimum GPU memory required for fine-tuning?

With **LoRA** adapters via `peft` and 4-bit quantization using `bitsandbytes`, you can fine-tune 7B parameter models on GPUs with as little as 8GB VRAM. The repository's `SFTTrainer` configurations automatically handle gradient checkpointing and mixed precision to minimize memory footprint.

### Can I use the latest versions of Transformers and TRL?

The repository pins specific versions in each experiment's [`requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/requirements.txt) (e.g., `transformers==4.48.3` for CoT distillation). While newer versions may work, the tested configurations ensure compatibility between `trl` trainers and specific model architectures. Always check the relevant [`requirements.txt`](https://github.com/bojieli/ai-agent-book/blob/main/requirements.txt) file before upgrading.

### How do I switch from Weights & Biases to TensorBoard?

Modify the `report_to` parameter in the `TrainingArguments` initialization within your training script. Change `report_to="wandb"` to `report_to="tensorboard"` or `report_to=["wandb", "tensorboard"]` to use both simultaneously. The `SFTTrainer` automatically handles the appropriate callback initialization.

### What is the difference between SFTTrainer and DPOTrainer in the repository?

**`SFTTrainer`** performs supervised fine-tuning on instruction-following datasets, loading data via `Dataset.from_list` as shown in [`train_sft_trl.py`](https://github.com/bojieli/ai-agent-book/blob/main/train_sft_trl.py). **`DPOTrainer`** (Direct Preference Optimization) requires preference pairs (chosen/rejected responses) and is used in the premature-completion experiments to align agents with human preferences without explicit reward modeling.