Academic Papers Linked in the Prompt Engineering Guide: A Complete Bibliography
The Prompt Engineering Guide references over 50 peer-reviewed arXiv papers spanning reliability research, Chain-of-Thought techniques, ChatGPT evaluations, and adversarial safety methods.
The dair-ai/Prompt-Engineering-Guide repository serves as a comprehensive open-source resource for prompt engineering practitioners. This guide embeds direct citations to foundational academic papers within its markdown documentation, organizing them by technical domain across six primary guide sections. Every reference links directly to arXiv abstracts or PDFs, providing immediate access to the underlying research.
Reliability and Safety Research
The reliability section in guides/prompts-reliability.md anchors the guide's safety-focused content with seven critical papers on model calibration and harm reduction:
- Constitutional AI: Harmlessness from AI Feedback (arXiv:2212.08073) – Introduces self-critique and revision methods for reducing harmful outputs without human feedback labels.
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? (arXiv:2202.12837) – Analyzes the mechanics behind few-shot prompting effectiveness, challenging assumptions about demonstration ground truth.
- Prompting GPT-3 To Be Reliable (arXiv:2210.09150) – Examines consistency and truthfulness in large language model responses under varying prompt conditions.
- On the Advance of Making Language Models Better Reasoners (arXiv:2206.02336) – Surveys techniques for improving logical reasoning capabilities through structured prompting.
- Unsolved Problems in ML Safety (arXiv:2109.13916) – Outlines persistent challenges in deploying safe machine learning systems at scale.
- Red Teaming Language Models to Reduce Harms (arXiv:2209.07858) – Documents scaling behaviors and systematic methods for adversarial testing of language models.
- Calibrate Before Use: Improving Few-Shot Performance of Language Models (arXiv:2102.09690v2) – Addresses probability calibration issues inherent in few-shot learning scenarios.
Advanced Prompting Techniques
Located in guides/prompts-advanced-usage.md, this section contains ten foundational papers that define modern prompt engineering methodologies:
- Demonstrating few-shot prompting (Brown et al., 2020, arXiv:2005.14165) – The seminal paper establishing few-shot in-context learning patterns in large language models.
- Chain-of-Thought (COT) prompting (Wei et al., 2022, arXiv:2201.11903) – Introduces step-by-step reasoning prompts to improve arithmetic, commonsense, and symbolic reasoning performance.
- Zero-Shot COT (Kojima et al., 2022, arXiv:2205.11916) – Extends chain-of-thought techniques to zero-shot settings using simple prompt augmentation.
- Self-Consistency (Wang et al., 2022, arXiv:2203.11171) – Utilizes majority voting across multiple sampled reasoning paths to enhance accuracy over greedy decoding.
- Automatic Prompt Engineer (APE) (Zhou et al., 2022, arXiv:2211.01910) – Automates prompt generation and selection using large language models as inference engines.
- AutoPrompt (Shin et al., 2020, arXiv:2010.15980) – Gradient-based discrete optimization for automatically constructing prompt tokens.
- Prefix Tuning (Li & Liang, 2021, arXiv:2101.00190) – Prepends trainable vectors to input representations for efficient task-specific adaptation.
- Prompt Tuning (Lester et al., 2021, arXiv:2104.08691) – Refines soft prompts through backpropagation while freezing the pre-trained model weights.
ChatGPT-Specific Studies
The guides/prompts-chatgpt.md file aggregates approximately thirty recent (2023) empirical studies evaluating ChatGPT across diverse domains:
Key evaluations include "Is ChatGPT a Good NLG Evaluator? A Preliminary Study" (arXiv:2303.04048), "Can ChatGPT Assess Human Personalities? A General Evaluation Framework" (arXiv:2303.01248), and "Exploring the Feasibility of ChatGPT for Event Extraction" (arXiv:2303.03836).
Domain-specific applications feature "ChatGPT is on the horizon: Could a large language model be all we need for Intelligent Transportation?" (arXiv:2303.05382), "Making a Computational Attorney" (arXiv:2303.05383), and "Does Synthetic Data Generation of LLMs Help Clinical Text Mining?" (arXiv:2303.04360).
Technical analysis papers such as "On the Robustness of ChatGPT: An Adversarial and Out-of-Distribution Perspective" (arXiv:2302.12095) and "A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT" (arXiv:2302.11382) provide implementation guidance for practitioners building robust applications.
Adversarial and Application-Specific Research
The adversarial section in guides/prompts-adversarial.md links three papers addressing safety and verification:
- Knowledge generation with LLMs (Liu et al., 2022, arXiv:2110.08387) – Available via direct PDF link, this paper explores generating structured knowledge from language models.
- Self-verification with LLMs (Weng et al., 2022, arXiv:2212.09561v1) – Investigates internal consistency checking mechanisms for verification-free reasoning.
- Machine Generated Text: A Comprehensive Survey of Threat Models and Detection Methods (arXiv:2210.07321) – Systematizes risks from synthetic text generation and corresponding detection strategies.
For practical implementations, guides/prompts-applications.md references Program-Aided Language Models (PAL) (arXiv:2211.10435), which combines natural language understanding with programmatic execution to solve quantitative reasoning problems.
Miscellaneous and Multimodal Research
The guides/prompts-miscellaneous.md file collects six papers on emergent techniques:
- Active-Prompt (arXiv:2302.12246) – Available as PDF, this paper introduces uncertainty-aware example selection for chain-of-thought prompting.
- Language Is Not All You Need: Aligning Perception with Language Models (arXiv:2302.14045) – Explores multimodal integration research connecting vision and language.
- GraphPrompt (Liu et al., 2023, arXiv:2302.08043) – Proposes graph neural network prompting frameworks for structured data.
- Additional 2023 papers by Li et al. (arXiv:2302.11520), Yao et al. (arXiv:2210.03629), and Zhang et al. (arXiv:2302.00923) covering specialized prompting variants for specific architectural adaptations.
Programmatic Extraction of Paper References
You can dynamically harvest all academic citations from the repository using Python to generate bibliographies or verify link integrity. The following script scans markdown files for arXiv and PDF links:
import pathlib
import re
import json
# Root of the cloned repository
repo_root = pathlib.Path("/cache/repos/github.com/dair-ai/Prompt-Engineering-Guide/main")
# Regex for markdown links that point to arXiv or a PDF on arXiv
LINK_RE = re.compile(r"\[([^\]]+)\]\((https?://arxiv\.org/(abs|pdf)/[^)]+)\)")
papers = []
for md_file in repo_root.rglob("*.md"):
for ln, line in enumerate(md_file.read_text().splitlines(), start=1):
for match in LINK_RE.finditer(line):
title, url, _ = match.groups()
papers.append(
{
"file": str(md_file.relative_to(repo_root)),
"line": ln,
"title": title.strip(),
"url": url,
}
)
# Pretty-print as JSON for further processing or reporting
print(json.dumps(papers, indent=2))
Executing this script produces structured JSON output:
[
{
"file": "guides/prompts-reliability.md",
"line": 161,
"title": "Constitutional AI: Harmlessness from AI Feedback",
"url": "https://arxiv.org/abs/2212.08073"
}
]
This extraction method enables automated citation management, allowing researchers to import the guide's references into reference managers like Zotero or generate static bibliography pages for derived documentation.
Summary
- The Prompt Engineering Guide references over 50 academic papers across six documentation files in the
guides/directory. - Papers span foundational techniques (Chain-of-Thought, few-shot learning) through contemporary ChatGPT evaluations (2023).
- Source files organize research by domain:
guides/prompts-reliability.mdfor safety,guides/prompts-advanced-usage.mdfor methodologies, andguides/prompts-chatgpt.mdfor model-specific studies. - All citations link directly to arXiv abstracts or PDFs for immediate access to full text.
- Python scripts can programmatically extract the entire bibliography using regex patterns on markdown link syntax.
Frequently Asked Questions
How many academic papers are linked in the Prompt Engineering Guide?
The guide contains approximately 50 distinct academic papers distributed across six markdown files. The guides/prompts-chatgpt.md file alone references roughly 30 empirical studies focused on ChatGPT evaluations and applications, while guides/prompts-advanced-usage.md contains 10 foundational methodological papers.
Where can I find the Chain-of-Thought prompting paper in the repository?
The seminal Chain-of-Thought paper by Wei et al. (2022) appears in guides/prompts-advanced-usage.md with the arXiv identifier 2201.11903. This file also houses related techniques including Zero-Shot COT (Kojima et al., arXiv:2205.11916) and Self-Consistency (Wang et al., arXiv:2203.11171).
Are all papers in the guide hosted on arXiv?
Nearly all references link to arXiv.org, though some entries provide direct PDF links (such as Active-Prompt in guides/prompts-miscellaneous.md and Self-Consistency in guides/prompts-advanced-usage.md). The extraction script specifically targets arxiv.org domains to capture the complete bibliography.
Can I programmatically generate a citation list from the guide?
Yes. The repository structure supports automated extraction using Python's pathlib and regular expressions to parse markdown link syntax. The provided script outputs JSON containing file paths, line numbers, paper titles, and URLs, enabling integration with reference management tools or static site generators.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →