# Building Evaluation Frameworks for LLM Agents: A Complete Guide to the AI Engineering from Scratch Curriculum

> Learn to build LLM agent evaluation frameworks with this AI Engineering from Scratch guide. Explore unit tests, MCP servers, and automated metrics like RAGAS and DeepEval for robust AI development.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: tutorial
- Published: 2026-07-26

---

**The AI Engineering from Scratch repository provides a comprehensive, code-first curriculum for building LLM agent evaluation frameworks using unit tests, MCP servers, and automated metrics like RAGAS, DeepEval, and G-Eval.**

The rohitg00/ai-engineering-from-scratch repository offers a battle-tested approach to building evaluation frameworks for LLM agents through its 20-phase, 503-lesson curriculum. This open-source project teaches AI engineering from first principles, covering everything from mathematical foundations to production agent deployment. By following the repository's structured methodology, developers can construct reproducible evaluation pipelines that validate agent behavior using concrete test suites and industry-standard metrics.

## The Build-It/Use-It Pedagogy

The curriculum organizes every lesson around a **six-step workflow** that ensures deep understanding before production use:

- **MOTTO** – A one-line core concept defining the lesson's objective
- **PROBLEM** – A concrete pain point that motivates the solution
- **CONCEPT** – Intuitive diagrams and theoretical explanation
- **BUILD IT** – Minimal implementation written from raw math without external ML libraries
- **USE IT** – The same algorithm implemented using production-grade frameworks like PyTorch or JAX
- **SHIP IT** – Exportable artifacts including prompts, skills, and agent configurations

Each lesson ships with a `code/tests/` directory containing at least five pytest cases. This structure ensures that every component of your evaluation framework can be validated independently before integration.

## Architectural Pillars for Agent Evaluation

The repository structures its evaluation capabilities across three critical phases:

| Pillar | Location | Evaluation Function |
|--------|----------|---------------------|
| **NLP Foundations to Advanced** | Phase 5, Lesson 27 | Implements **RAGAS**, **DeepEval**, and **G-Eval** metrics for quantitative assessment |
| **Tools and Protocols** | Phase 13 | Deploys the **Model-Context-Protocol (MCP)** for standardized tool calling and agent-server communication |
| **Agent Engineering** | Phase 14 | Houses the reference **ReAct-style** agent loop used for testing decision sequences |

These phases work sequentially: Phase 5 provides the metrics, Phase 13 supplies the communication infrastructure, and Phase 14 delivers the agent runtime.

## Implementing the Evaluation Pipeline

### Step 1: Validating Components with Unit Tests

Before evaluating end-to-end agent behavior, verify individual implementations using the built-in test suites. Each lesson directory contains pytest-compatible tests:

```bash
cd phases/14-agent-engineering/01-the-agent-loop
python3 -m unittest discover -v

```

This command discovers and executes all test cases in the `code/tests/` subdirectory, ensuring your agent logic passes regression checks.

### Step 2: Deploying the MCP Server Infrastructure

The evaluation framework relies on the **Model-Context-Protocol (MCP)** to standardize tool interactions. Start the MCP server from Phase 13 to expose tools that your agent will call during evaluation:

```bash
python3 phases/13-tools-and-protocols/07-building-an-mcp-server/server.py \
    --host 0.0.0.0 --port 8000

```

This server registers available tools and handles structured tool calls from the agent.

### Step 3: Executing the Agent Loop

The reference implementation for the agent runtime lives in [`phases/14-agent-engineering/01-the-agent-loop/code/agent_loop.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/01-the-agent-loop/code/agent_loop.py). The `run()` function implements a **ReAct-style** decision loop that manages the conversation between the LLM and available tools:

```python
def run(query, tools):
    """ReAct‑style loop for any tool list."""
    history = [user(query)]
    for step in range(MAX_STEPS):
        msg = llm(history)                       # LLM generates response / tool calls

        if msg.tool_calls:                       # If the model wants to call a tool …

            for call in msg.tool_calls:
                result = tools[call.name](**call.args)   # … invoke the Python function

                history.append(tool_result(call.id, result))
            continue
        return msg.content                        # Final answer when no more tool calls

    raise StepLimitExceeded

```

Connect to the running MCP server using the client from `phases/13-tools-and-protocols/08-building-an-mcp-client/`:

```python
from mcp_client import MCPClient

client = MCPClient("http://localhost:8000")
response = client.run("What is the capital of France?")
print(response)   # -> "Paris"

```

### Step 4: Computing Evaluation Metrics

Phase 5 Lesson 27 (`phases/05-nlp-foundations-to-advanced/27-llm-evaluation-frameworks/`) provides Python implementations for industry-standard metrics. After generating agent outputs, compute scores against ground truth references:

```python
from evaluation import ragas_score, deepeval_score, geval_score

generated = client.run("Summarize the attached document.")
reference = "... ground truth summary ..."
print("RAGAS:", ragas_score(generated, reference))
print("DeepEval:", deepeval_score(generated, reference))
print("G‑Eval:", geval_score(generated, reference))

```

These functions evaluate retrieval accuracy, factual consistency, and generative quality respectively.

## Key Source Files for Agent Evaluation

The following files constitute the core evaluation infrastructure in the repository:

- [`phases/14-agent-engineering/01-the-agent-loop/code/agent_loop.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/01-the-agent-loop/code/agent_loop.py) – Reference **ReAct-style** agent loop with the `run()` function
- [`phases/13-tools-and-protocols/07-building-an-mcp-server/server.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/13-tools-and-protocols/07-building-an-mcp-server/server.py) – **MCP server** implementation for tool exposure
- `phases/13-tools-and-protocols/08-building-an-mcp-client/` – **MCP client** for querying agents over HTTP
- `phases/05-nlp-foundations-to-advanced/27-llm-evaluation-frameworks/` – Evaluation metrics including **RAGAS**, **DeepEval**, and **G-Eval**
- [`scripts/install_skills.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py) – Utility to install generated **SKILL** files into Claude, Cursor, and other AI agents
- [`AGENTS.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/AGENTS.md) – Operating manual specifying commit conventions and evaluation contracts for AI contributors

## Summary

- The curriculum employs a **six-step Build-It/Use-It workflow** that teaches evaluation concepts from scratch before introducing production libraries
- **Phase 14** delivers the reference `run()` function in [`agent_loop.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/agent_loop.py) for implementing **ReAct-style** agent evaluation
- **Phase 13** implements the **Model-Context-Protocol (MCP)** enabling standardized tool invocation and monitoring
- **Phase 5** provides automated scoring through **RAGAS**, **DeepEval**, and **G-Eval** metrics
- Every lesson includes `code/tests/` directories with pytest suites for continuous validation
- The [`scripts/install_skills.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py) utility enables direct integration with existing AI assistants like Claude and Cursor

## Frequently Asked Questions

### What evaluation metrics does the framework support?

The repository implements **RAGAS**, **DeepEval**, and **G-Eval** in `phases/05-nlp-foundations-to-advanced/27-llm-evaluation-frameworks/`. **RAGAS** evaluates retrieval-augmented generation pipelines, **DeepEval** checks for factual consistency and answer relevance, while **G-Eval** uses LLM-as-a-judge patterns for flexible quality assessment.

### How does the Model-Context-Protocol enable agent evaluation?

The **Model-Context-Protocol (MCP)** standardizes how agents discover and invoke tools. By running the MCP server from [`phases/13-tools-and-protocols/07-building-an-mcp-server/server.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/13-tools-and-protocols/07-building-an-mcp-server/server.py), evaluators can inject controlled tool responses, monitor the exact sequence of `tool_calls` in the agent loop, and verify that the `run()` function handles edge cases like `StepLimitExceeded` correctly.

### Can I use this framework with existing AI agents like Claude or Cursor?

Yes. The repository includes [`scripts/install_skills.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py) which consumes the **SKILL** definitions exported to `outputs/skill-*.md` files. These skills are compatible with Claude, Cursor, Codex, OpenClaw, and Hermes, allowing you to evaluate these agents using the repository's test suites and metrics.

### Where are the unit tests located for validating agent behavior?

Each lesson contains a `code/tests/` subdirectory with pytest-compatible test cases. For example, the agent loop tests reside in `phases/14-agent-engineering/01-the-agent-loop/code/tests/` and can be executed using `python3 -m unittest discover -v` to validate the implementation before deployment.