# What Checks Are Performed During the INTEGRITY Stage of the ARS Pipeline: The Complete 5-Phase Audit

> Discover the crucial checks performed during the INTEGRITY stage of the ARS pipeline. This detailed audit ensures no fabricated sources, data errors, or plagiarism before peer review.

- Repository: [Edward Cheng-I Wu/academic-research-skills](https://github.com/Imbad0202/academic-research-skills)
- Tags: deep-dive
- Published: 2026-05-13

---

**The INTEGRITY stage of the ARS pipeline executes a mandatory two-point audit (Stage 2.5 and Stage 4.5) across five verification phases—references, citations, data, originality, and claims—to eliminate fabricated sources, data errors, and plagiarism before peer review or publication.**

The INTEGRITY stage serves as the gatekeeper in the Academic Research Skills (ARS) pipeline developed in the `Imbad0202/academic-research-skills` repository. Enforced by the `integrity_verification_agent`, these checks are organized into five distinct phases (A through E) that audit every aspect of academic integrity from bibliography accuracy to claim verification. Understanding these checks is essential for researchers and developers integrating the pipeline, as they represent a zero-tolerance barrier that blocks progression unless strict verification standards are met.

## Phase A: Reference Verification

According to [`academic-pipeline/agents/integrity_verification_agent.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/academic-pipeline/agents/integrity_verification_agent.md) (lines 84-154), Phase A validates that every bibliographic entry exists and matches authoritative sources. The phase runs four sequential sub-checks:

- **A0. Semantic Scholar Batch Lookup**: An optional fast API pass that pre-screens references against the Semantic Scholar database for rapid validation.
- **A1. Existence WebSearch**: For each reference, the agent performs a web search using the author, title, and year to confirm the source exists, returning a status of `VERIFIED`, `NOT_FOUND`, or `MISMATCH`.
- **A2. Bibliographic Accuracy**: For verified references, the system compares every metadata field—author list, publication year, title, venue, DOI, and page numbers—against the discovered source, flagging discrepancies as **SERIOUS**, **MEDIUM**, or **MINOR**.
- **A3. Ghost-Citation Check**: The agent detects orphan references (entries never cited in the manuscript) and dangling citations (in-text citations missing from the bibliography).

## Phase B: Citation-Context Verification

Phase B ensures that citations accurately reflect their source material and follow consistent formatting rules. As defined in [`integrity_verification_agent.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/integrity_verification_agent.md) (lines 54-67), this phase includes:

- **B1. Citation Accuracy**: The system performs spot-checks to verify that cited passages match the source content. During **Stage 2.5 (Initial Verification)**, the agent samples at least 30% of citations; during **Stage 4.5 (Final Verification)**, this increases to 100%. The check flags cherry-picked data, misrepresented findings, or incorrect statistics.
- **B2. Citation Format Consistency**: Validation of APA 7th edition compliance, proper use of "et al.," and consistent page-number formatting across the manuscript.

## Phase C: Data Verification

According to [`integrity_verification_agent.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/integrity_verification_agent.md) (lines 78-100), Phase C audits quantitative claims through two mechanisms:

- **C1. Statistical Data Cross-Reference**: Each numerical claim is traced to its cited source, and the reported values are compared against the original publication. Discrepancies are flagged by severity level.
- **C2. Internal Consistency**: The system verifies that the same data point remains identical across all manuscript locations—including tables, figures, and body text—preventing contradictory statistics within the document.

## Phase D: Originality Verification

Phase D detects plagiarism and improper text reuse through paragraph-level analysis. As specified in [`integrity_verification_agent.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/integrity_verification_agent.md) (lines 121-165):

- **D1. Paragraph-Level Originality**: The agent extracts 1-2 characteristic sentences per paragraph and searches them against the open web. During Initial Verification, the system samples at least 30% of paragraphs; during Final Verification, this increases to 50% (with 100% coverage on newly-added paragraphs). Matches are graded from **ORIGINAL** to **VERBATIM**.
- **D2. Self-Plagiarism Detection**: Using the author-provided name, the agent searches for overlap with the author's prior publications, flagging unacknowledged text recycling.

## Phase E: Claim Verification

The final verification phase enumerates and validates every factual assertion. According to [`integrity_verification_agent.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/integrity_verification_agent.md) (lines 120-148):

- **E1. Claim Extraction**: All quantitative and factual claims are systematically enumerated from the manuscript text.
- **E2. Source Tracing**: The system locates the specific supporting passage within each cited source.
- **E3. Cross-Referencing**: Claim text is compared against source text, with verdicts assigned as **VERIFIED**, **MINOR_DISTORTION**, **MAJOR_DISTORTION**, **UNVERIFIABLE**, or **UNVERIFIABLE_ACCESS** (for paywalled sources).

## Operating Modes: Initial vs Final Verification

The `integrity_verification_agent` operates in two distinct modes that determine sampling depth and strictness, as documented in the pipeline state machine:

**Mode 1: Initial Verification (Stage 2.5)**
- Runs before peer review
- Full Phase A execution
- Spot-check 30% of citations (Phase B)
- Full Phase C execution
- Sample 30% of paragraphs for originality (Phase D)
- Sample 30% of claims (Phase E)
- Must achieve **PASS** to enter Stage 3 (Review)

**Mode 2: Final Verification (Stage 4.5)**
- Runs after revision and before finalization
- Fresh full Phase A execution
- 100% citation verification (Phase B)
- Full Phase C execution
- Minimum 50% paragraph sampling plus 100% on new content (Phase D)
- 100% claim verification (Phase E)
- Must achieve **PASS** with zero SERIOUS/MEDIUM/MAJOR_DISTORTION/UNVERIFIABLE issues to enter Stage 5 (Finalisation)

## Verdict Synthesis and Pipeline Gates

The final integrity verdict aggregates severity across all phases using criteria defined in [`integrity_verification_agent.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/integrity_verification_agent.md) (lines 36-42):

- **PASS**: Zero SERIOUS or MEDIUM issues, zero MAJOR_DISTORTION, and zero UNVERIFIABLE claims.
- **PASS WITH NOTES**: Only MINOR issues or UNVERIFIABLE_ACCESS restrictions (requires explicit user acknowledgment).
- **FAIL**: Any SERIOUS/MEDIUM issue, MAJOR_DISTORTION, or UNVERIFIABLE claim detected.

The pipeline orchestrator ([`pipeline_orchestrator_agent.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/pipeline_orchestrator_agent.md), lines 119-140) enforces these checks as **mandatory gates** that cannot be skipped. Stage 2.5 requires an Integrity Report (Schema 5) attached to the hand-off before entering Review, while Stage 4.5 requires a Final Integrity Report before the manuscript can be finalized. The state machine ([`pipeline_state_machine.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/pipeline_state_machine.md), lines 202-234) specifies an Integrity Check FAIL Loop that returns the manuscript for correction when gates fail.

## Implementation Example

Developers can invoke the integrity verification programmatically using the agent interface. The pipeline automatically wires these calls, but the following examples demonstrate expected inputs and outputs:

```python

# Example: Running the integrity verification agent on a draft

from agents.integrity_verification_agent import IntegrityVerificationAgent

draft_path = "paper_draft.md"
agent = IntegrityVerificationAgent(mode="initial")   # mode="final" for Stage 4.5

report = agent.run(draft_path)

print(report.verdict)            # PASS / PASS WITH NOTES / FAIL

print(report.summary)            # Tabulated counts per phase

print(report.correction_list)   # Items to fix if verdict != PASS

```

The agent emits an Integrity Report following **Schema 5** (defined in [`shared/handoff_schemas.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/shared/handoff_schemas.md), line 247):

```yaml

# Example: Expected Integrity Report (Schema 5) – excerpt

integrity_pass_date: "2026-03-08T15:45:00Z"
verification_status: "VERIFIED"
reference_existence:
  total: 62
  passed: 62
  issues: 0
bibliographic_accuracy:
  serious: 0
  medium: 0
  minor: 1
citation_context:
  spot_check_rate: 30%
  serious: 0
  medium: 0
  minor: 0
claim_verification:
  verified: 58
  minor_distortion: 2
  major_distortion: 0
  unverifiable: 0

```

## Summary

- The INTEGRITY stage mandates two checkpoints (Stage 2.5 and Stage 4.5) enforced by the `integrity_verification_agent`.
- Five verification phases (A-E) cover references, citation contexts, data consistency, originality, and factual claims.
- Initial Verification uses sampling rates of 30% for citations, originality, and claims, while Final Verification requires 100% coverage for citations and claims, and 50%+ for originality.
- Verdicts derive from severity aggregation: PASS requires zero serious issues; FAIL triggers a mandatory correction loop.
- The pipeline gates in [`pipeline_orchestrator_agent.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/pipeline_orchestrator_agent.md) block progression until integrity reports meet Schema 5 standards.

## Frequently Asked Questions

### What is the difference between Stage 2.5 and Stage 4.5 integrity checks?

Stage 2.5 (Initial Verification) runs before peer review with reduced sampling (30% spot-checks) to catch major issues early. Stage 4.5 (Final Verification) runs after revisions with comprehensive 100% verification of citations and claims, plus 50% paragraph sampling for originality, ensuring zero tolerance for errors before publication.

### What constitutes a FAIL verdict in the INTEGRITY stage?

A FAIL verdict occurs when the audit detects any **SERIOUS** or **MEDIUM** severity issue in reference or data verification, any **MAJOR_DISTORTION** in claim verification, or any **UNVERIFIABLE** claim without source access. Only **MINOR** issues or **UNVERIFIABLE_ACCESS** restrictions allow a PASS WITH NOTES verdict.

### How does the pipeline handle ghost citations?

During Phase A3 (Ghost-Citation Check), the `integrity_verification_agent` cross-references the bibliography against in-text citations to identify orphan references (uncited bibliography entries) and dangling citations (in-text cites missing from references). Both conditions trigger severity flags that can prevent pipeline progression.

### Can the INTEGRITY stage be skipped or bypassed?

No. According to [`pipeline_orchestrator_agent.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/pipeline_orchestrator_agent.md) and the pipeline state machine in [`pipeline_state_machine.md`](https://github.com/Imbad0202/academic-research-skills/blob/main/pipeline_state_machine.md), the integrity gates are **mandatory checkpoints** that cannot be disabled. The orchestrator blocks hand-off to subsequent stages until the agent emits a valid Integrity Report (Schema 5) meeting the minimum PASS criteria.