Best Practices for Prompt Engineering: A Complete Guide from the AI Engineering from Scratch Curriculum

Prompt engineering is the cornerstone of reliable LLM behavior and is implemented as a first-class engineering skill throughout Phase 11 of the AI Engineering from Scratch curriculum.

The AI Engineering from Scratch repository by rohitg00 treats prompt engineering as a rigorous discipline rather than an ad-hoc craft. According to the source code, effective prompt engineering combines role definition, context selection, explicit constraints, and iterative refinement to control model behavior without modifying weights. This guide distills the curriculum's implementation patterns, common pitfalls, and evaluation frameworks into actionable best practices.

Core Principles of Prompt Engineering

The curriculum defines four foundational pillars that every engineered prompt should address. These principles appear in phases/11-llm-engineering/01-prompt-engineering/docs/en.md (line 13) and form the basis of the pattern catalogue.

Role Definition

Role definition establishes a stable persona for the model, preventing role-play jailbreaks and providing consistent behavioral context. Rather than assuming the model knows its job, explicitly state the persona: "You are a climate-expert assistant" or "You are a medical assistant."

This technique is covered in phases/11-llm-engineering/01-prompt-engineering/docs/en.md and serves as the first checkpoint in the prompt optimization pipeline.

Context Selection

Context selection determines what information enters the prompt window, directly controlling token budget and preventing context overflow. The curriculum emphasizes separating retrieval, compression, and tool selection into a dedicated Context Engineering pipeline.

Detailed implementation guidance appears in phases/11-llm-engineering/05-context-engineering/docs/en.md (line 571), which teaches engineers to treat context management as distinct from prompt structure.

Constraints and Output Formatting

Explicit constraints reduce ambiguity by limiting format, style, or safety boundaries. Pairing constraints with output format specifications—such as JSON schemas, bullet points, or tables—prevents decoding errors.

The structured outputs lesson in phases/11-llm-engineering/03-structured-outputs/docs/en.md (line 30) clarifies that format specification is a prompt engineering responsibility distinct from the model's generation logic.

Iterative Refinement

Iterative refinement uses few-shot examples and chain-of-thought reasoning to boost performance without fine-tuning. The curriculum positions this as the baseline optimization technique before advancing to parameter updates.

This approach is documented in phases/11-llm-engineering/08-fine-tuning-lora/docs/en.md (line 200), establishing that you should exhaust prompt engineering strategies before modifying model weights.

Common Pitfalls in Prompt Engineering

The repository identifies three critical failures that undermine prompt reliability.

Vagueness causes models to guess wildly. A prompt like "Write a story" lacks boundaries. The fix involves adding concrete constraints such as length limits, tone specifications, and JSON schemas.

Over-reliance on token-level tricks ignores higher-level structure. Engineers should use the Core Prompt Engineering Patterns (role, context, constraints, output format) as a checklist rather than chasing low-level optimizations.

Ignoring context engineering treats the entire prompt as a monolith. The curriculum recommends separating retrieval, compression, and tool selection into distinct pipeline stages, as implemented in Phase 11-05.

Architectural Implementation

The curriculum implements prompt engineering as reusable software components rather than static strings.

Skill Definition and Pattern Catalogue

The skill definition resides in phases/11-llm-engineering/01-prompt-engineering/outputs/skill-prompt-patterns.md, containing the complete pattern catalogue. This file serves as the single source of truth for prompt structures across the curriculum.

Prompt Optimizer

The prompt optimizer in phases/11-llm-engineering/01-prompt-engineering/outputs/prompt-prompt-optimizer.md exposes a callable interface that rewrites draft prompts using the pattern catalogue. This programmatic approach ensures consistency across different LLM interactions.

Language Implementations

Python and TypeScript implementations demonstrate production usage:

These artifacts integrate into downstream lessons, including the Safety Gate capstone, demonstrating how engineered prompts propagate through full LLM pipelines.

Evaluation and Regression Testing

Prompt engineering changes require rigorous validation. The curriculum employs promptfoo (Phase 10-10-evaluation) and built-in refusal evaluation metrics.

Promptfoo Integration

The evaluation harness validates that engineered prompts do not regress across model versions or use cases. Configuration files specify prompts, datasets, and provider endpoints for systematic testing.

Refusal Evaluation

The refusal evaluation lesson in phases/19-capstone-projects/84-refusal-evaluation/docs/en.md measures under-refusal versus over-refusal, quantifying how prompt changes affect safety boundaries. The jailbreak taxonomy in phases/19-capstone-projects/82-jailbreak-taxonomy/docs/en.md provides adversarial test cases for sanity-checking engineering decisions.

Practical Code Examples

Using the Prompt Optimizer (Python)

from prompt_engineering import optimize_prompt

draft = "Write a helpful answer about climate change."
optimized = optimize_prompt(draft)
print(optimized)

# → "You are a climate‑expert assistant. \

#    Provide a concise, fact‑checked summary (max 150 words) \

#    about the current state of climate change, using bullet points."

Source: phases/11-llm-engineering/01-prompt-engineering/code/prompt_engineering.py

Defining a Few-Shot Template (TypeScript)

const system = "You are a helpful tutor for Python programming.";
const examples = [
  { prompt: "Explain list comprehensions.", completion: "A list comprehension …" },
  { prompt: "What is async/await?", completion: "Async/await …" }
];
const userPrompt = "How do I read a CSV file?";

const fullPrompt = `${system}\n\n${examples
  .map(e => `Q: ${e.prompt}\nA: ${e.completion}`).join('\n\n')}\n\nQ: ${userPrompt}\nA:`;

Source: phases/11-llm-engineering/01-prompt-engineering/code/main.ts

Running Regression Tests (Promptfoo)


# tests/prompt_qa.yaml

prompts:
  - "You are a medical assistant. Provide a brief, factual answer to: {query}"
datasets:
  - name: medical_qa
    path: ./datasets/medical_qa.json
providers:
  - name: gpt-4
    model: openai/gpt-4

Execute with promptfoo eval tests/prompt_qa.yaml to validate prompt stability. Reference: Phase 10-10-evaluation docs (phases/10-llms-from-scratch/10-evaluation/docs/en.md).

Summary

  • Prompt engineering is a first-class engineering skill in the AI Engineering from Scratch curriculum, not an afterthought.
  • Four core principles govern effective prompts: role definition, context selection, explicit constraints, and output format specification.
  • Programmatic optimization through the optimize_prompt function and skill definitions ensures consistency across Python and TypeScript implementations.
  • Evaluation frameworks using promptfoo and refusal metrics catch regressions before deployment.
  • Context engineering remains distinct from prompt structuring, requiring separate pipeline stages as covered in Phase 11-05.

Frequently Asked Questions

What are the four core components of a well-engineered prompt?

According to the AI Engineering from Scratch repository, every prompt should explicitly define the role, manage context selection, impose constraints, and specify the output format. These components work together to reduce ambiguity and guide the model toward deterministic, reproducible behavior without requiring fine-tuning.

How does the curriculum recommend testing prompt changes?

The repository recommends promptfoo for regression testing and the refusal evaluation framework for safety metrics. By running promptfoo eval against standardized datasets and measuring under-refusal versus over-refusal rates, engineers can quantify how prompt modifications affect both performance and safety boundaries.

When should prompt engineering be exhausted before fine-tuning?

The curriculum in phases/11-llm-engineering/08-fine-tuning-lora/docs/en.md (line 200) establishes iterative refinement as the mandatory baseline before modifying model weights. You should implement few-shot examples and chain-of-thought reasoning through prompt engineering first, only proceeding to fine-tuning when architectural constraints require behavior changes beyond what prompt patterns can achieve.

What is the difference between prompt engineering and context engineering?

Prompt engineering focuses on the structure, role definition, and constraints within the immediate prompt window, while context engineering manages what information enters that window through retrieval, compression, and tool selection. The repository treats these as separate pipeline stages, with context engineering covered in Phase 11-05 and prompt engineering in Phase 11-01.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →