Evaluating Instruction-Following Models with LLM-as-a-Judge: Implementation Guide

The rasbt/LLMs-from-scratch repository provides a deterministic LLM-as-a-judge framework that scores instruction-following model responses on a 0-100 scale using an external LLM like Llama 3 via Ollama.

Evaluating instruction-following models requires nuanced judgment beyond automated metrics. The rasbt/LLMs-from-scratch repository implements a complete LLM-as-a-judge pipeline that leverages a separate large language model to assess fine-tuned model outputs against ground-truth instructions with reproducible scoring.

How the Evaluation Pipeline Works

The evaluation system treats an external LLM as a judge that assigns numerical scores to model responses. The pipeline consists of three core components implemented across specific modules in the repository.

Dataset Preparation and Formatting

The pipeline begins by loading instruction-response pairs and formatting them into evaluation prompts. In pkg/llms_from_scratch/ch07.py, the download_and_load_file function handles JSON dataset downloading, while format_input structures each entry containing the instruction, optional input context, ground-truth output, and the candidate model response.

Querying the Judge LLM

The judge LLM receives structured scoring prompts via HTTP POST requests. Implemented in ch07/01_main-chapter-code/ollama_evaluate.py, the query_model function sends requests to an Ollama server endpoint (defaulting to Llama 3) with deterministic parameters: seed=123, temperature=0, and a fixed context length. This configuration ensures reproducible scores across evaluation runs.

Score Aggregation and Analysis

The generate_model_scores function in the same file iterates over the test dataset, converts textual responses to integers, and handles edge cases where the judge returns non-numeric output. The system aggregates individual scores to compute average performance metrics, providing a quantitative view of instruction-following accuracy.

Implementing the Evaluation in Python

To evaluate your fine-tuned model, first ensure your dataset contains the required fields: instruction, input, output, and model_response.

Load and prepare the dataset using the provided utilities:

from llms_from_scratch.ch07 import download_and_load_file, format_input

file_path = "instruction-data.json"
url = "https://raw.githubusercontent.com/rasbt/LLMs-from-scratch/main/ch07/01_main-chapter-code/instruction-data.json"
data = download_and_load_file(file_path, url)

entry = data[0]
prompt = (
    f"Given the input `{format_input(entry)}` "
    f"and correct output `{entry['output']}`, "
    f"score the model response `{entry['model_response']}` "
    f"on a scale from 0 to 100, where 100 is the best score. "
    f"Respond with the integer number only."
)

Query the judge LLM to retrieve a numerical score:

from ch07.01_main-chapter-code.ollama_evaluate import query_model

score = query_model(prompt)  # Returns string like "87"

print(int(score))            # Convert to integer: 87

For batch processing, use the generate_model_scores function which handles iteration and error handling:

from tqdm import tqdm
from ch07.01_main-chapter-code.ollama_evaluate import query_model, format_input

def generate_model_scores(json_data, key="model_response", model="llama3"):
    scores = []
    for entry in tqdm(json_data, desc="Scoring"):
        if not entry[key]:
            scores.append(0)
            continue
        prompt = (
            f"Given the input `{format_input(entry)}` "
            f"and correct output `{entry['output']}`, "
            f"score the model response `{entry[key]}` "
            f"on a scale from 0 to 100, where 100 is the best score. "
            f"Respond with the integer number only."
        )
        raw = query_model(prompt, model=model)
        try:
            scores.append(int(raw))
        except ValueError:
            scores.append(0)
    return scores

Running Batch Evaluations from the Command Line

The repository includes a command-line interface in ollama_evaluate.py that automates the full evaluation workflow. The script first verifies the Ollama server is running via check_if_running, then loads the JSON file and computes average scores across all entries.

Execute the evaluation from your terminal:

python -m ch07.01_main-chapter-code.ollama_evaluate \
    --file_path path/to/instruction-data-with-response.json

The script outputs the total number of scored entries and the average score, enabling quick benchmarking of instruction-following performance.

Summary

  • The rasbt/LLMs-from-scratch repository implements a complete LLM-as-a-judge evaluation framework for instruction-following models.
  • The pipeline uses deterministic generation (seed=123, temperature=0) to ensure reproducible 0-100 scores across runs.
  • Core functions reside in pkg/llms_from_scratch/ch07.py (data loading) and ch07/01_main-chapter-code/ollama_evaluate.py (scoring logic).
  • The system supports both programmatic Python access and command-line execution for flexible integration into training workflows.

Frequently Asked Questions

What is LLM-as-a-judge evaluation?

LLM-as-a-judge is an evaluation methodology where a separate large language model assesses the quality of generated responses against reference outputs. Unlike traditional metrics such as BLEU or ROUGE, this approach captures semantic nuances and instruction-following accuracy by leveraging the judge model's understanding of natural language.

How does the 0-100 scoring scale work?

The judge LLM receives a structured prompt containing the original instruction, ground-truth output, and candidate response, then returns an integer between 0 and 100 where 100 represents perfect adherence to instructions. The query_model function enforces this format by requesting integer-only responses, with the generate_model_scores function handling parsing and fallback to 0 for invalid outputs.

Why use Ollama for the judge LLM?

Ollama provides local inference capabilities that eliminate API costs and latency associated with cloud-based models while maintaining evaluation quality. The implementation defaults to Llama 3 but supports any Ollama-compatible model, allowing researchers to run evaluations entirely on local hardware with consistent, reproducible results through fixed seed and temperature settings.

How can I ensure my evaluation scores are reproducible?

The evaluation framework enforces determinism through specific generation parameters in ollama_evaluate.py: setting seed=123 and temperature=0 in the Ollama API request ensures identical outputs for identical inputs across runs. Additionally, validate that your instruction dataset includes all required fields (instruction, input, output, model_response) and that your Ollama server version remains consistent between evaluations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →