What Are the Dependencies for Fine-Tuning AI Agents in the ai-agent-book Repository?

Fine-tuning AI agents in the bojieli/ai-agent-book repository requires PyTorch, Hugging Face Transformers, Accelerate, Datasets, and TRL as core dependencies, with optional PEFT for LoRA adapters, BitsAndBytes or Unsloth for quantization, and Weights & Biases for experiment tracking.

The repository implements a modular training pipeline for agent fine-tuning across multiple modalities. Whether you are performing prompt distillation, chain-of-thought training, or multilingual reasoning, the codebase relies on a standardized Hugging Face ecosystem with specific version constraints documented in per-experiment requirements.txt files.

Core Machine Learning Stack

The foundation of agent fine-tuning rests on five essential libraries that handle tensor operations, model architecture, data processing, and training orchestration.

PyTorch and Transformers

torch provides the underlying tensor computation engine, while transformers supplies the model and tokenizer classes. According to the source code in chapter8/prompt-distillation/train_sft_trl.py, models are instantiated using:

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.bfloat16,
    trust_remote_code=True,
    use_cache=False,
)

Training Orchestration with Accelerate and TRL

accelerate handles distributed training across multiple GPUs, and trl (Transformer Reinforcement Learning) provides high-level trainers including SFTTrainer, DPOTrainer, and GRPOTrainer. The repository uses these to abstract away training loop boilerplate while maintaining full hyperparameter control.

Data Pipeline with Datasets

datasets manages large-scale data loading. The training scripts load JSONL files using Dataset.from_list, as implemented in chapter8/prompt-distillation/train_sft_trl.py:

from datasets import Dataset
import json

def load_jsonl(file_path: str) -> Dataset:
    data = []
    with open(file_path, "r", encoding="utf-8") as f:
        for line in f:
            if line.strip():
                data.append(json.loads(line))
    return Dataset.from_list(data)

Parameter-Efficient Fine-Tuning (PEFT)

For memory-efficient training, the repository leverages peft to implement LoRA (Low-Rank Adaptation) adapters. This allows fine-tuning by updating only a small subset of model weights, drastically reducing GPU memory requirements.

In chapter8/prompt-distillation/train_sft_trl.py, LoRA configuration is applied as follows:

from peft import LoraConfig, get_peft_model

peft_cfg = LoraConfig(
    r=32, 
    lora_alpha=16, 
    target_modules="all-linear", 
    lora_dropout=0.0,
    bias="none", 
    task_type="CAUSAL_LM"
)
model = get_peft_model(model, peft_cfg)

Optional Optimization and Quantization

BitsAndBytes for 4-Bit Training

bitsandbytes enables 4-bit quantized training, allowing very large models to fit on consumer GPUs. This is specifically used in the multilingual reasoning experiments located in chapter8/MultilingualReasoning/.

Unsloth for Accelerated Training

unsloth provides optimized training kernels that significantly speed up LoRA fine-tuning. The speech SFT experiment at chapter8/speech-sft-experiment/run_sesame.py includes unsloth as a primary dependency for faster convergence.

Inference Acceleration with vLLM

When generating synthetic training data from teacher models, the repository uses vllm for fast batched inference. This is particularly relevant for prompt distillation workflows where a larger teacher model generates training examples for a smaller student model.

from vllm import LLM, SamplingParams

llm = LLM(model="meta-llama/Meta-Llama-3-8B", tensor_parallel_size=2)
params = SamplingParams(temperature=0.7, max_tokens=512)
output = llm.generate([prompt], params)

Experiment Tracking and Logging

wandb (Weights & Biases) is the primary tool for logging metrics, hyperparameters, and checkpoint artifacts. Some experiments also support trackio or TensorBoard as alternatives. The training scripts configure logging via the report_to parameter in TrainingArguments.

Installation Requirements by Experiment

Different chapters in the repository specify exact version constraints in their respective requirements.txt files:

Complete Training Example

To fine-tune an agent using the standard pipeline, combine these components as shown in chapter8/prompt-distillation/train_sft_trl.py:

from trl import SFTTrainer, TrainingArguments

trainer = SFTTrainer(
    model=model,
    tokenizer=tokenizer,
    train_dataset=train_ds,
    args=TrainingArguments(
        output_dir="./outputs",
        num_train_epochs=1,
        per_device_train_batch_size=4,
        gradient_accumulation_steps=4,
        learning_rate=2e-4,
        logging_steps=1,
        save_strategy="epoch",
        report_to="wandb",
        run_name="agent-finetune",
    ),
    peft_config=peft_cfg,
)

trainer.train()

Summary

  • Core dependencies for fine-tuning AI agents include torch, transformers, accelerate, datasets, and trl for the complete training pipeline.
  • PEFT (peft) enables memory-efficient training through LoRA adapters, requiring only gigabytes instead of tens of gigabytes of VRAM.
  • Quantization libraries like bitsandbytes and unsloth allow 4-bit training and accelerated kernels for large-scale experiments.
  • vLLM provides high-throughput inference for teacher models generating synthetic training data.
  • Experiment tracking relies on wandb or trackio, integrated directly into the TrainingArguments configuration.
  • Each experiment in chapter8/ maintains its own requirements.txt with tested version pins to ensure reproducibility across prompt distillation, CoT training, multilingual reasoning, and speech modalities.

Frequently Asked Questions

What is the minimum GPU memory required for fine-tuning?

With LoRA adapters via peft and 4-bit quantization using bitsandbytes, you can fine-tune 7B parameter models on GPUs with as little as 8GB VRAM. The repository's SFTTrainer configurations automatically handle gradient checkpointing and mixed precision to minimize memory footprint.

Can I use the latest versions of Transformers and TRL?

The repository pins specific versions in each experiment's requirements.txt (e.g., transformers==4.48.3 for CoT distillation). While newer versions may work, the tested configurations ensure compatibility between trl trainers and specific model architectures. Always check the relevant requirements.txt file before upgrading.

How do I switch from Weights & Biases to TensorBoard?

Modify the report_to parameter in the TrainingArguments initialization within your training script. Change report_to="wandb" to report_to="tensorboard" or report_to=["wandb", "tensorboard"] to use both simultaneously. The SFTTrainer automatically handles the appropriate callback initialization.

What is the difference between SFTTrainer and DPOTrainer in the repository?

SFTTrainer performs supervised fine-tuning on instruction-following datasets, loading data via Dataset.from_list as shown in train_sft_trl.py. DPOTrainer (Direct Preference Optimization) requires preference pairs (chosen/rejected responses) and is used in the premature-completion experiments to align agents with human preferences without explicit reward modeling.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →