# How to Use Python for Training and Evaluating LLMs: A Complete Workflow Guide

> Learn to train and evaluate LLMs in Python with Hugging Face Transformers and PEFT. Master data prep, fine-tuning, and standardized benchmarking for efficient LLM development.

- Repository: [Maxime Labonne/llm-course](https://github.com/mlabonne/llm-course)
- Tags: how-to-guide
- Published: 2026-03-01

---

**Training and evaluating large language models in Python requires Hugging Face Transformers for model handling, PEFT for efficient fine-tuning, and lm-evaluation-harness for standardized benchmarking, all orchestrated through a workflow of data preparation, tokenization, adapter-based training, and automated assessment.**

The **mlabonne/llm-course** repository provides a curated educational framework that teaches the complete LLM lifecycle using Python. According to the source code documentation in [`README.md`](https://github.com/mlabonne/llm-course/blob/main/README.md) (lines 103-110), the **Python for Machine Learning** chapter establishes why Python dominates LLM research, citing its readability and rich ecosystem of tensor operations, data wrangling, and visualization tools.

## Why Python Is the Lingua Franca for LLMs

Python’s dominance in machine learning stems from its unmatched ecosystem of specialized libraries. **NumPy** provides fast tensor operations, while **Pandas** handles complex data wrangling. **Matplotlib** and **Seaborn** visualize training metrics and dataset distributions, and **Scikit-learn** offers classical ML algorithms for baseline comparisons.

For modern LLM work, the **🤗 Transformers** library supplies state-of-the-art architectures and tokenizers. This stack enables rapid prototyping—from loading a `meta-llama/Meta-Llama-3.1-8B` checkpoint to deploying quantized models for inference.

## The Python LLM Workflow: Stage by Stage

The repository outlines a complete Python-centric pipeline for model development. Each stage leverages specific libraries optimized for GPU efficiency and distributed training.

### Data Preparation and Tokenization

Start with **Data Preparation** using `datasets`, `pandas`, and `json` to load raw text, clean outliers, and split into train/validation/test sets. For **Tokenization**, use `transformers.AutoTokenizer` to convert strings to integer IDs, handling special tokens and padding automatically.

Key concepts appear in [`README.md`](https://github.com/mlabonne/llm-course/blob/main/README.md) lines 669-671, which detail how text becomes numerical representations suitable for neural network inputs.

### Model Selection and Fine-Tuning

For **Model Selection**, instantiate architectures with `transformers.AutoModelForCausalLM`, loading pretrained checkpoints from the Hugging Face Hub. The **Fine-Tuning** stage offers two paths:

- **Standard Approach**: Use `transformers.Trainer` with custom `TrainingArguments` for full-parameter or gradient-checkpointed updates.
- **Efficient Approach**: Apply **Axolotl** or **TRL** (Transformer Reinforcement Learning) with **PEFT** (Parameter-Efficient Fine-Tuning) to run LoRA or QLoRA, reducing GPU memory requirements by training only adapter layers rather than full weights.

The repository covers pre-training versus fine-tuning distinctions in [`README.md`](https://github.com/mlabonne/llm-course/blob/main/README.md) lines 684-688, including distributed training tips and data pipeline construction.

### Evaluation and Deployment

For **Evaluation**, integrate **lm-evaluation-harness** to run automated benchmarks like MMLU, TruthfulQA, and HellaSwag. This framework computes accuracy, perplexity, and human-aligned scores without custom scripts. The evaluation methodologies are detailed in [`README.md`](https://github.com/mlabonne/llm-course/blob/main/README.md) lines 754-761, covering automated benchmarks, human evaluation protocols, and model-based judges.

Finally, **Deployment** uses **vLLM** or **text-generation-inference** for high-throughput serving, often paired with **FastAPI** for REST endpoints and quantization formats like GGUF or GPTQ.

## Fine-Tuning LLMs with LoRA and PEFT

The following Python implementation demonstrates adapter-based fine-tuning on instruction data, adapted from the practical examples linked in the repository’s notebook table (around line 33). This approach modifies only the attention projection layers (`q_proj`, `v_proj`) while freezing the base model.

```python

# Install required packages (run once)

# pip install transformers datasets peft tqdm

from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForCausalLM, Trainer, TrainingArguments
from peft import LoraConfig, get_peft_model

# 1️⃣ Load a tiny dataset – e.g. Alpaca style instructions

dataset = load_dataset("tatsu-lab/alpaca_eval", split="train[:1%]")

# 2️⃣ Tokenizer

model_name = "meta-llama/Meta-Llama-3.1-8B"
tokenizer = AutoTokenizer.from_pretrained(model_name, use_fast=True)

def tokenize_fn(example):
    return tokenizer(example["instruction"], truncation=True, max_length=512)

tokenized = dataset.map(tokenize_fn, batched=True)

# 3️⃣ Load the base model (in float16 for GPU memory efficiency)

base_model = AutoModelForCausalLM.from_pretrained(
    model_name, torch_dtype="auto", device_map="auto"
)

# 4️⃣ Attach LoRA adapters (low‑rank fine‑tuning)

lora_cfg = LoraConfig(
    r=8,               # rank

    lora_alpha=16,
    target_modules=["q_proj", "v_proj"],  # typical for Llama‑type models

    lora_dropout=0.1,
    bias="none",
)
model = get_peft_model(base_model, lora_cfg)

# 5️⃣ Trainer arguments

training_args = TrainingArguments(
    output_dir="./lora-llama",
    per_device_train_batch_size=2,
    gradient_accumulation_steps=8,
    learning_rate=2e-4,
    num_train_epochs=1,
    fp16=True,
    logging_steps=10,
    save_steps=200,
    eval_strategy="no",
)

# 6️⃣ Trainer

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized,
    tokenizer=tokenizer,
)

# 7️⃣ Kick off fine‑tuning

trainer.train()

```

## Evaluating Models with lm-evaluation-harness

After training, assess model performance using the **lm-evaluation-harness** framework. This tool standardizes benchmarking across the LLM community, supporting both Hugging Face models and local checkpoints.

```bash

# Install the harness

pip install lm-eval

# Run a quick benchmark on the newly saved checkpoint

lm_eval \
  --model hf-causal-experimental \
  --model_args pretrained=./lora-llama,trust_remote_code=True \
  --tasks mmlu,hellaswag \
  --batch_size 4

```

The harness calculates metrics comparable to published leaderboards, enabling objective assessment of your fine-tuning results.

## Advanced Concepts in the Repository

The [`README.md`](https://github.com/mlabonne/llm-course/blob/main/README.md) provides architectural context essential before writing training code:

- **Attention & Sampling** (lines 672-673): Explains self-attention mechanisms and generation strategies like temperature sampling and nucleus sampling.
- **Preference Alignment** (lines 735-742): Covers DPO (Direct Preference Optimization), GRPO, and PPO pipelines for aligning models with human feedback.
- **Pre-training Strategies** (lines 684-688): Details data pipelines, distributed training configurations, and memory optimization techniques for training from scratch versus fine-tuning.

## Summary

- Use **Hugging Face Transformers** as the core library for model architecture and tokenization, referencing `AutoModelForCausalLM` and `AutoTokenizer`.
- Apply **PEFT/LoRA** adapters via the `peft` library to reduce memory requirements during fine-tuning, targeting modules like `q_proj` and `v_proj`.
- Orchestrate training with `transformers.Trainer` or specialized frameworks like **Axolotl** and **TRL** for advanced techniques including RLHF.
- Benchmark performance using **lm-evaluation-harness** for standardized metrics like MMLU and HellaSwag.
- Reference the [`README.md`](https://github.com/mlabonne/llm-course/blob/main/README.md) in **mlabonne/llm-course** for conceptual foundations and linked Colab notebooks for hands-on practice.

## Frequently Asked Questions

### What Python libraries are essential for LLM training?

The core stack includes **Hugging Face Transformers** for model architectures, **datasets** for data handling, **PEFT** for adapter-based fine-tuning, and **TRL** or **Axolotl** for advanced training pipelines. According to the LLM Course repository's Python chapter (lines 103-110), NumPy and Pandas provide the foundational data manipulation layer required for preprocessing raw text into training tensors.

### How do I evaluate a fine-tuned LLM in Python?

Use the **lm-evaluation-harness** framework to run standardized benchmarks like MMLU and TruthfulQA. This command-line tool integrates with Hugging Face models via the `hf-causal-experimental` model type and computes accuracy, perplexity, and task-specific metrics without requiring custom evaluation scripts, as documented in the repository's evaluation section (lines 754-761).

### What is the difference between full fine-tuning and LoRA?

**Full fine-tuning** updates every model parameter, requiring substantial GPU memory and compute resources. **LoRA** (Low-Rank Adaptation), implemented via the `peft` library, injects trainable rank decomposition matrices into attention layers (typically `q_proj` and `v_proj`), reducing trainable parameters by up to 10,000x while maintaining comparable performance on downstream tasks.

### Where can I find runnable examples for LLM training?

The **mlabonne/llm-course** repository links to external Google Colab notebooks (referenced in the README around line 33) that contain production-ready Python scripts. These notebooks demonstrate specific implementations like "Fine-tune Llama 3.1 with Unsloth" and automated quantization workflows that complement the conceptual material in [`README.md`](https://github.com/mlabonne/llm-course/blob/main/README.md).