Building Evaluation Frameworks for LLM Agents: A Complete Guide to the AI Engineering from Scratch Curriculum

The AI Engineering from Scratch repository provides a comprehensive, code-first curriculum for building LLM agent evaluation frameworks using unit tests, MCP servers, and automated metrics like RAGAS, DeepEval, and G-Eval.

The rohitg00/ai-engineering-from-scratch repository offers a battle-tested approach to building evaluation frameworks for LLM agents through its 20-phase, 503-lesson curriculum. This open-source project teaches AI engineering from first principles, covering everything from mathematical foundations to production agent deployment. By following the repository's structured methodology, developers can construct reproducible evaluation pipelines that validate agent behavior using concrete test suites and industry-standard metrics.

The Build-It/Use-It Pedagogy

The curriculum organizes every lesson around a six-step workflow that ensures deep understanding before production use:

  • MOTTO – A one-line core concept defining the lesson's objective
  • PROBLEM – A concrete pain point that motivates the solution
  • CONCEPT – Intuitive diagrams and theoretical explanation
  • BUILD IT – Minimal implementation written from raw math without external ML libraries
  • USE IT – The same algorithm implemented using production-grade frameworks like PyTorch or JAX
  • SHIP IT – Exportable artifacts including prompts, skills, and agent configurations

Each lesson ships with a code/tests/ directory containing at least five pytest cases. This structure ensures that every component of your evaluation framework can be validated independently before integration.

Architectural Pillars for Agent Evaluation

The repository structures its evaluation capabilities across three critical phases:

Pillar Location Evaluation Function
NLP Foundations to Advanced Phase 5, Lesson 27 Implements RAGAS, DeepEval, and G-Eval metrics for quantitative assessment
Tools and Protocols Phase 13 Deploys the Model-Context-Protocol (MCP) for standardized tool calling and agent-server communication
Agent Engineering Phase 14 Houses the reference ReAct-style agent loop used for testing decision sequences

These phases work sequentially: Phase 5 provides the metrics, Phase 13 supplies the communication infrastructure, and Phase 14 delivers the agent runtime.

Implementing the Evaluation Pipeline

Step 1: Validating Components with Unit Tests

Before evaluating end-to-end agent behavior, verify individual implementations using the built-in test suites. Each lesson directory contains pytest-compatible tests:

cd phases/14-agent-engineering/01-the-agent-loop
python3 -m unittest discover -v

This command discovers and executes all test cases in the code/tests/ subdirectory, ensuring your agent logic passes regression checks.

Step 2: Deploying the MCP Server Infrastructure

The evaluation framework relies on the Model-Context-Protocol (MCP) to standardize tool interactions. Start the MCP server from Phase 13 to expose tools that your agent will call during evaluation:

python3 phases/13-tools-and-protocols/07-building-an-mcp-server/server.py \
    --host 0.0.0.0 --port 8000

This server registers available tools and handles structured tool calls from the agent.

Step 3: Executing the Agent Loop

The reference implementation for the agent runtime lives in phases/14-agent-engineering/01-the-agent-loop/code/agent_loop.py. The run() function implements a ReAct-style decision loop that manages the conversation between the LLM and available tools:

def run(query, tools):
    """ReAct‑style loop for any tool list."""
    history = [user(query)]
    for step in range(MAX_STEPS):
        msg = llm(history)                       # LLM generates response / tool calls

        if msg.tool_calls:                       # If the model wants to call a tool …

            for call in msg.tool_calls:
                result = tools[call.name](**call.args)   # … invoke the Python function

                history.append(tool_result(call.id, result))
            continue
        return msg.content                        # Final answer when no more tool calls

    raise StepLimitExceeded

Connect to the running MCP server using the client from phases/13-tools-and-protocols/08-building-an-mcp-client/:

from mcp_client import MCPClient

client = MCPClient("http://localhost:8000")
response = client.run("What is the capital of France?")
print(response)   # -> "Paris"

Step 4: Computing Evaluation Metrics

Phase 5 Lesson 27 (phases/05-nlp-foundations-to-advanced/27-llm-evaluation-frameworks/) provides Python implementations for industry-standard metrics. After generating agent outputs, compute scores against ground truth references:

from evaluation import ragas_score, deepeval_score, geval_score

generated = client.run("Summarize the attached document.")
reference = "... ground truth summary ..."
print("RAGAS:", ragas_score(generated, reference))
print("DeepEval:", deepeval_score(generated, reference))
print("G‑Eval:", geval_score(generated, reference))

These functions evaluate retrieval accuracy, factual consistency, and generative quality respectively.

Key Source Files for Agent Evaluation

The following files constitute the core evaluation infrastructure in the repository:

Summary

  • The curriculum employs a six-step Build-It/Use-It workflow that teaches evaluation concepts from scratch before introducing production libraries
  • Phase 14 delivers the reference run() function in agent_loop.py for implementing ReAct-style agent evaluation
  • Phase 13 implements the Model-Context-Protocol (MCP) enabling standardized tool invocation and monitoring
  • Phase 5 provides automated scoring through RAGAS, DeepEval, and G-Eval metrics
  • Every lesson includes code/tests/ directories with pytest suites for continuous validation
  • The scripts/install_skills.py utility enables direct integration with existing AI assistants like Claude and Cursor

Frequently Asked Questions

What evaluation metrics does the framework support?

The repository implements RAGAS, DeepEval, and G-Eval in phases/05-nlp-foundations-to-advanced/27-llm-evaluation-frameworks/. RAGAS evaluates retrieval-augmented generation pipelines, DeepEval checks for factual consistency and answer relevance, while G-Eval uses LLM-as-a-judge patterns for flexible quality assessment.

How does the Model-Context-Protocol enable agent evaluation?

The Model-Context-Protocol (MCP) standardizes how agents discover and invoke tools. By running the MCP server from phases/13-tools-and-protocols/07-building-an-mcp-server/server.py, evaluators can inject controlled tool responses, monitor the exact sequence of tool_calls in the agent loop, and verify that the run() function handles edge cases like StepLimitExceeded correctly.

Can I use this framework with existing AI agents like Claude or Cursor?

Yes. The repository includes scripts/install_skills.py which consumes the SKILL definitions exported to outputs/skill-*.md files. These skills are compatible with Claude, Cursor, Codex, OpenClaw, and Hermes, allowing you to evaluate these agents using the repository's test suites and metrics.

Where are the unit tests located for validating agent behavior?

Each lesson contains a code/tests/ subdirectory with pytest-compatible test cases. For example, the agent loop tests reside in phases/14-agent-engineering/01-the-agent-loop/code/tests/ and can be executed using python3 -m unittest discover -v to validate the implementation before deployment.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →