# How Graphify Performs PR Impact Analysis and Triage: A Technical Deep Dive

> Discover how Graphify analyzes PR impact and prioritizes reviews. Learn how it maps code changes to a knowledge graph, calculates blast radius, and uses LLM triage for efficient review ranking.

- Repository: [Safi/graphify](https://github.com/safishamsi/graphify)
- Tags: deep-dive
- Published: 2026-06-15

---

**Graphify treats pull requests as enriched objects that map code changes against a knowledge graph to calculate blast radius, then leverages LLM-powered triage to automatically rank PRs by review priority.**

The `safishamsi/graphify` repository provides an intelligent PR impact analysis and triage system that transforms raw GitHub pull requests into actionable intelligence. By combining graph topology analysis with AI-driven prioritization, Graphify reveals exactly which communities and nodes each PR touches, enabling maintainers to focus review efforts where they matter most.

## The PRInfo Data Model as a First-Class Object

Graphify's architecture centers on the `PRInfo` class, which aggregates three distinct data streams:

- **GitHub metadata**: CI status, review decisions, and draft flags fetched via the `gh` CLI.
- **Work-tree mapping**: A dictionary linking branch names to local work-tree paths for rapid filesystem access.
- **Graph-impact data**: Community intersections, node counts, and blast-radius calculations derived from [`graphify-out/graph.json`](https://github.com/safishamsi/graphify/blob/main/graphify-out/graph.json).

## The 8-Step Impact Analysis Workflow

The PR analysis pipeline executes through a coordinated sequence of functions defined in [`graphify/prs.py`](https://github.com/safishamsi/graphify/blob/main/graphify/prs.py).

### Fetching PRs and Detecting Base Branches

The process begins with `fetch_prs()` (lines 89-118), which executes `gh pr list` to retrieve open pull requests and instantiate `PRInfo` objects. Concurrently, `_detect_default_branch()` (lines 48-69) auto-detects the repository's default branch using `gh repo view` or `git symbolic-ref`, ensuring accurate base comparisons.

### Work-Tree Mapping and Graph Loading

To map local development contexts, `fetch_worktrees()` (lines 92-106) parses `git worktree list --porcelain` output into a `{branch: path}` dictionary. The system then loads topological data via `_load_graph_json()` (lines 18-27), which reads [`graphify-out/graph.json`](https://github.com/safishamsi/graphify/blob/main/graphify-out/graph.json) while applying a size-cap security check via `check_graph_file_size_cap` from [`graphify/security.py`](https://github.com/safishamsi/graphify/blob/main/graphify/security.py).

### Computing Graph Impact and Blast Radius

The core analysis occurs in `attach_graph_impact()` (lines 34-92). This function:

1. Fetches changed files concurrently using `fetch_pr_files()`.
2. Matches paths against the graph index using `_path_match`, which tolerates both absolute and relative path formats.
3. Populates `communities_touched`, `nodes_affected`, and generates a human-readable `blast_radius` string (e.g., "12 nodes / 3 communities").

### Rendering the Impact Dashboard

Finally, `render_dashboard()` (lines 96-114) outputs a color-coded terminal table displaying each PR's blast radius and impact summary.

## O(nodes + files) Graph Lookup Algorithm

Graphify optimizes impact calculation through one-time indexing of `file_comms` and `file_count`, ensuring that changed file lookups operate at **O(nodes + files)** complexity rather than quadratic time. The `_path_match` function normalizes path representations, treating [`src/foo.py`](https://github.com/safishamsi/graphify/blob/main/src/foo.py) and [`foo.py`](https://github.com/safishamsi/graphify/blob/main/foo.py) as identical entities during community matching.

## Opus-Powered Triage and Backend Resolution

For automated prioritization, `triage_with_opus()` (lines 42-56) constructs textual prompts describing each actionable PR's impact characteristics, then streams ranked priorities through an LLM backend.

The backend selection logic in `_resolve_triage_backend()` (lines 58-73) follows this priority order:

1. Explicit `GRAPHIFY_TRIAGE_BACKEND` environment variable.
2. First available backend with valid API credentials (Claude, Kimi, OpenAI, Gemini, Ollama).
3. Fallback to local `claude-cli` if installed.

Model selection uses `GRAPHIFY_TRIAGE_MODEL` or defaults from `_TRIAGE_MODEL_DEFAULTS` defined in [`graphify/llm.py`](https://github.com/safishamsi/graphify/blob/main/graphify/llm.py).

## Command-Line Usage Examples

Execute impact analysis using the following commands:

```bash

# Display dashboard of all open PRs with impact calculations

graphify prs

# Generate AI-ranked triage list for review prioritization

graphify prs --triage

# Inspect specific PR details including changed files and impact

graphify prs 1234

```

Programmatic access to raw impact data:

```python
from pathlib import Path
from graphify.prs import fetch_prs, attach_graph_impact

prs = fetch_prs()
graph_path = Path("graphify-out/graph.json")
labels = attach_graph_impact(prs, graph_path)

for pr in prs:
    if pr.communities_touched:
        print(f"PR #{pr.number} touches communities {pr.communities_touched} "
              f"affecting {pr.nodes_affected} nodes → {pr.blast_radius}")

```

## Summary

- Graphify's **PR impact analysis and triage** system treats pull requests as enriched `PRInfo` objects combining GitHub metadata, work-tree mappings, and graph topology data.
- Impact calculation runs in **O(nodes + files)** time via indexed lookups in `attach_graph_impact()` within [`graphify/prs.py`](https://github.com/safishamsi/graphify/blob/main/graphify/prs.py).
- The **blast radius** metric quantifies PR scope as human-readable strings like "12 nodes / 3 communities".
- **Opus-powered triage** automatically ranks PRs by review priority using configurable LLM backends resolved via `_resolve_triage_backend()`.
- Integration endpoints in [`graphify/serve.py`](https://github.com/safishamsi/graphify/blob/main/graphify/serve.py) expose the same functionality over HTTP for CI/CD pipelines.

## Frequently Asked Questions

### How does Graphify determine which graph communities a PR affects?

Graphify matches changed files against the knowledge graph index using `_path_match` in [`graphify/prs.py`](https://github.com/safishamsi/graphify/blob/main/graphify/prs.py), which identifies intersecting communities while normalizing absolute and relative path formats. The results populate `PRInfo.communities_touched` and `nodes_affected`.

### What LLM backends does Graphify support for triage ranking?

The system supports Claude (Opus), Kimi, OpenAI, Gemini, and Ollama, with selection controlled by `GRAPHIFY_TRIAGE_BACKEND` and model selection via `GRAPHIFY_TRIAGE_MODEL`. If no API keys are configured, it falls back to local `claude-cli` according to the logic in `_resolve_triage_backend()`.

### Why is the impact lookup O(nodes + files) instead of O(n²)?

Graphify pre-indexes the graph into `file_comms` and `file_count` structures during `_load_graph_json()`, allowing `attach_graph_impact()` to perform hashtable lookups rather than nested iterations when matching changed files against graph nodes.

### Can I use Graphify's PR analysis in a CI/CD pipeline?

Yes. While the CLI provides the primary interface, [`graphify/serve.py`](https://github.com/safishamsi/graphify/blob/main/graphify/serve.py) exposes the same PR impact analysis and triage functionality via HTTP endpoints, enabling integration with automated build systems and custom dashboards.