# Evaluating Instruction-Following Models with LLM-as-a-Judge: Implementation Guide

> Implement LLM-as-a-judge to evaluate instruction-following models. Learn how to score responses using Llama 3 and Ollama with the LLMs-from-scratch framework. Get started today!

- Repository: [Sebastian Raschka/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)
- Tags: how-to-guide
- Published: 2026-05-12

---

**The rasbt/LLMs-from-scratch repository provides a deterministic LLM-as-a-judge framework that scores instruction-following model responses on a 0-100 scale using an external LLM like Llama 3 via Ollama.**

Evaluating instruction-following models requires nuanced judgment beyond automated metrics. The rasbt/LLMs-from-scratch repository implements a complete LLM-as-a-judge pipeline that leverages a separate large language model to assess fine-tuned model outputs against ground-truth instructions with reproducible scoring.

## How the Evaluation Pipeline Works

The evaluation system treats an external LLM as a judge that assigns numerical scores to model responses. The pipeline consists of three core components implemented across specific modules in the repository.

### Dataset Preparation and Formatting

The pipeline begins by loading instruction-response pairs and formatting them into evaluation prompts. In [`pkg/llms_from_scratch/ch07.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch07.py), the `download_and_load_file` function handles JSON dataset downloading, while `format_input` structures each entry containing the instruction, optional input context, ground-truth output, and the candidate model response.

### Querying the Judge LLM

The judge LLM receives structured scoring prompts via HTTP POST requests. Implemented in [`ch07/01_main-chapter-code/ollama_evaluate.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/01_main-chapter-code/ollama_evaluate.py), the `query_model` function sends requests to an Ollama server endpoint (defaulting to Llama 3) with deterministic parameters: `seed=123`, `temperature=0`, and a fixed context length. This configuration ensures reproducible scores across evaluation runs.

### Score Aggregation and Analysis

The `generate_model_scores` function in the same file iterates over the test dataset, converts textual responses to integers, and handles edge cases where the judge returns non-numeric output. The system aggregates individual scores to compute average performance metrics, providing a quantitative view of instruction-following accuracy.

## Implementing the Evaluation in Python

To evaluate your fine-tuned model, first ensure your dataset contains the required fields: `instruction`, `input`, `output`, and `model_response`.

Load and prepare the dataset using the provided utilities:

```python
from llms_from_scratch.ch07 import download_and_load_file, format_input

file_path = "instruction-data.json"
url = "https://raw.githubusercontent.com/rasbt/LLMs-from-scratch/main/ch07/01_main-chapter-code/instruction-data.json"
data = download_and_load_file(file_path, url)

entry = data[0]
prompt = (
    f"Given the input `{format_input(entry)}` "
    f"and correct output `{entry['output']}`, "
    f"score the model response `{entry['model_response']}` "
    f"on a scale from 0 to 100, where 100 is the best score. "
    f"Respond with the integer number only."
)

```

Query the judge LLM to retrieve a numerical score:

```python
from ch07.01_main-chapter-code.ollama_evaluate import query_model

score = query_model(prompt)  # Returns string like "87"

print(int(score))            # Convert to integer: 87

```

For batch processing, use the `generate_model_scores` function which handles iteration and error handling:

```python
from tqdm import tqdm
from ch07.01_main-chapter-code.ollama_evaluate import query_model, format_input

def generate_model_scores(json_data, key="model_response", model="llama3"):
    scores = []
    for entry in tqdm(json_data, desc="Scoring"):
        if not entry[key]:
            scores.append(0)
            continue
        prompt = (
            f"Given the input `{format_input(entry)}` "
            f"and correct output `{entry['output']}`, "
            f"score the model response `{entry[key]}` "
            f"on a scale from 0 to 100, where 100 is the best score. "
            f"Respond with the integer number only."
        )
        raw = query_model(prompt, model=model)
        try:
            scores.append(int(raw))
        except ValueError:
            scores.append(0)
    return scores

```

## Running Batch Evaluations from the Command Line

The repository includes a command-line interface in [`ollama_evaluate.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ollama_evaluate.py) that automates the full evaluation workflow. The script first verifies the Ollama server is running via `check_if_running`, then loads the JSON file and computes average scores across all entries.

Execute the evaluation from your terminal:

```bash
python -m ch07.01_main-chapter-code.ollama_evaluate \
    --file_path path/to/instruction-data-with-response.json

```

The script outputs the total number of scored entries and the average score, enabling quick benchmarking of instruction-following performance.

## Summary

- The rasbt/LLMs-from-scratch repository implements a complete **LLM-as-a-judge** evaluation framework for instruction-following models.
- The pipeline uses **deterministic generation** (seed=123, temperature=0) to ensure reproducible 0-100 scores across runs.
- Core functions reside in [`pkg/llms_from_scratch/ch07.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch07.py) (data loading) and [`ch07/01_main-chapter-code/ollama_evaluate.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/01_main-chapter-code/ollama_evaluate.py) (scoring logic).
- The system supports both **programmatic Python access** and **command-line execution** for flexible integration into training workflows.

## Frequently Asked Questions

### What is LLM-as-a-judge evaluation?

LLM-as-a-judge is an evaluation methodology where a separate large language model assesses the quality of generated responses against reference outputs. Unlike traditional metrics such as BLEU or ROUGE, this approach captures semantic nuances and instruction-following accuracy by leveraging the judge model's understanding of natural language.

### How does the 0-100 scoring scale work?

The judge LLM receives a structured prompt containing the original instruction, ground-truth output, and candidate response, then returns an integer between 0 and 100 where 100 represents perfect adherence to instructions. The `query_model` function enforces this format by requesting integer-only responses, with the `generate_model_scores` function handling parsing and fallback to 0 for invalid outputs.

### Why use Ollama for the judge LLM?

Ollama provides local inference capabilities that eliminate API costs and latency associated with cloud-based models while maintaining evaluation quality. The implementation defaults to Llama 3 but supports any Ollama-compatible model, allowing researchers to run evaluations entirely on local hardware with consistent, reproducible results through fixed seed and temperature settings.

### How can I ensure my evaluation scores are reproducible?

The evaluation framework enforces determinism through specific generation parameters in [`ollama_evaluate.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ollama_evaluate.py): setting `seed=123` and `temperature=0` in the Ollama API request ensures identical outputs for identical inputs across runs. Additionally, validate that your instruction dataset includes all required fields (`instruction`, `input`, `output`, `model_response`) and that your Ollama server version remains consistent between evaluations.