# How to Customize or Extend Site Audit Functionality in OpenSEO: 6 Developer Methods Explained

> Developers can customize OpenSEO site audit functionality using 6 methods: modify workflow phases, extend AuditConfig schema, inject crawl logic, add issue detectors, and more. Learn how to extend OpenSEO today.

- Repository: [Every App/open-seo](https://github.com/every-app/open-seo)
- Tags: how-to-guide
- Published: 2026-08-15

---

**Developers can customize OpenSEO's site audit functionality by modifying the modular workflow phases in `src/server/workflows/`, extending the `AuditConfig` schema, injecting custom logic into crawl policies, or adding new multipage issue detectors—all without altering the core orchestration infrastructure.**

OpenSEO is an open-source SEO platform that runs site audits through a durable, step-based workflow architecture. Because each audit phase is implemented as a pure function that receives a typed `WorkflowStep` and configuration object, the system provides clear extension points for custom business logic, data sources, and reporting requirements.

## Understanding the Site Audit Workflow Architecture

The audit pipeline in OpenSEO is explicitly divided into isolated phases defined in [`src/server/workflows/siteAuditWorkflowPhases.ts`](https://github.com/every-app/open-seo/blob/main/src/server/workflows/siteAuditWorkflowPhases.ts). Each phase handles a specific concern, making the system predictable to extend.

| Phase | Responsibility | Primary Implementation |
|-------|---------------|----------------------|
| **Discovery** | Fetches sitemap/robots.txt and seeds the crawl frontier | `runDiscoveryPhase` in [`siteAuditWorkflowPhases.ts`](https://github.com/every-app/open-seo/blob/main/siteAuditWorkflowPhases.ts) |
| **Crawl** | Fetches pages in chunks, persists data, and records internal links | `runCrawlPhase` / `runCrawlChunk` in [`siteAuditWorkflowCrawl.ts`](https://github.com/every-app/open-seo/blob/main/siteAuditWorkflowCrawl.ts) |
| **Lighthouse** | Runs configurable Lighthouse checks on sampled pages | `runLighthousePhase` in [`siteAuditWorkflowPhases.ts`](https://github.com/every-app/open-seo/blob/main/siteAuditWorkflowPhases.ts) |
| **Multipage Checks & Finalization** | Runs aggregate issue detectors and closes the audit | `runMultipageChecks` and `finalizeAudit` in [`siteAuditWorkflowPhases.ts`](https://github.com/every-app/open-seo/blob/main/siteAuditWorkflowPhases.ts) |

Because these phases communicate via well-defined parameter contracts rather than shared mutable state, you can override or augment any specific behavior while preserving the surrounding orchestration.

## Method 1: Configure Audit Behavior via AuditConfig

The simplest way to customize audit behavior is through the **`AuditConfig`** object passed when creating an audit via the `POST /api/audit` endpoint. The configuration flows through `runAuditPhases` to every downstream phase.

Key configuration fields include:

- **`maxPages`** – Upper bound on pages the crawler visits (default: `1000`)
- **`lighthouseStrategy`** – Sampling strategy: `"none"`, `"sample"`, or `"full"`
- **`crawlDepth`** – Optional limit for internal link depth

To adjust a specific audit, modify the JSON payload:

```json
{
  "projectId": "abc123",
  "startUrl": "https://example.com",
  "config": {
    "maxPages": 5000,
    "lighthouseStrategy": "none"
  }
}

```

The `runDiscoveryPhase`, `runCrawlPhase`, and `runLighthousePhase` functions receive this config object and adjust their behavior accordingly.

## Method 2: Modify Crawl Policies with Custom Queuing Logic

The crawl phase decides whether to queue a URL via **`shouldQueueCrawlLink`** in [`src/server/workflows/siteAuditWorkflowCrawl.ts`](https://github.com/every-app/open-seo/blob/main/src/server/workflows/siteAuditWorkflowCrawl.ts) (lines 60-68). This function performs origin checks, robots.txt validation, and URL crawlability tests.

To implement a custom filtering rule—such as excluding specific path patterns—modify the logic after the base checks:

```typescript
// src/server/workflows/siteAuditWorkflowCrawl.ts
function shouldQueueCrawlLink(
  link: string,
  origin: string,
  robots: RobotsResult,
): boolean {
  const baseChecks = isSameOrigin(link, origin) 
    && isCrawlableUrl(link) 
    && robots.isAllowed(link);
  
  // Custom exclusion: skip admin paths
  const isAdminPath = /\/admin\//.test(link);
  
  return baseChecks && !isAdminPath;
}

```

Because `runCrawlPhase` calls `shouldQueueCrawlLink` for every discovered link, your custom logic immediately affects crawl scope without touching the chunking or persistence layers.

## Method 3: Extend Lighthouse Sampling Strategies

Lighthouse execution is controlled by **`runLighthousePhase`**, which delegates page selection to **`selectLighthouseSample`** in [`src/server/lib/audit/lighthouse.ts`](https://github.com/every-app/open-seo/blob/main/src/server/lib/audit/lighthouse.ts). By default, this supports `"sample"` and `"full"` strategies.

To implement a custom sampling algorithm:

1. Update the `AuditConfig` type definition in [`src/types/schemas/audit.ts`](https://github.com/every-app/open-seo/blob/main/src/types/schemas/audit.ts) to include your new strategy literal (e.g., `"full-slow"`)
2. Add a branch in `selectLighthouseSample` to handle the new case

```typescript
// src/server/lib/audit/lighthouse.ts
export function selectLighthouseSample(
  pages: { url: string; statusCode: number }[],
  startUrl: string,
  strategy: AuditConfig["lighthouseStrategy"]
) {
  if (strategy === "full-slow") {
    // Return all crawled pages without sampling
    return pages.map(p => p.url);
  }
  // Existing strategies...
  if (strategy === "sample") {
    // Default sampling logic
  }
}

```

`runLighthousePhase` automatically forwards the strategy value, so no orchestration changes are required.

## Method 4: Add Custom Multipage Issue Detectors

After crawling completes, **`runMultipageChecks`** aggregates issues by calling detectors exported from [`src/server/lib/audit/issues/multipage/index.ts`](https://github.com/every-app/open-seo/blob/main/src/server/lib/audit/issues/multipage/index.ts). Each detector must implement the `DetectedIssue` contract.

To add a custom check that flags pages with slow load times:

```typescript
// src/server/lib/audit/issues/multipage/slow-pages.ts
import type { DetectedIssue } from "@/server/lib/audit/issues/page-reporters";

export async function detectSlowPages(auditId: string): Promise<DetectedIssue[]> {
  const pages = await AuditRepository.getPagesForAudit(auditId);
  
  return pages
    .filter(p => p.loadTimeMs && p.loadTimeMs > 3000)
    .map(p => ({
      issueType: "slow-page" as const,
      pageId: p.id,
      pageUrl: p.url,
      dedupeKey: p.url,
      details: { loadTimeMs: p.loadTimeMs },
    }));
}

// Export from the multipage index
// src/server/lib/audit/issues/multipage/index.ts
export { detectSlowPages } from "./slow-pages";

```

Once exported, `runMultipageChecks` (called within `finalizeAudit`) automatically includes your detector in the next audit execution.

## Method 5: Hook Into the Finalization Step

The **`finalizeAudit`** function in [`src/server/workflows/siteAuditWorkflowPhases.ts`](https://github.com/every-app/open-seo/blob/main/src/server/workflows/siteAuditWorkflowPhases.ts) handles the last database writes, emits analytics events, and clears the scratchpad (lines 71-86). This is the ideal location to add custom side effects like notifications or webhook calls.

To send a Slack notification when an audit completes:

```typescript
// Inside finalizeAudit in siteAuditWorkflowPhases.ts
await captureServerEvent({
  distinctId: billingCustomer.userId,
  event: "site_audit:complete",
  organizationId: billingCustomer.organizationId,
  properties: { auditId, status: "completed" },
});

// Custom side-effect
await fetch("https://hooks.slack.com/services/...", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ text: `✅ Audit ${auditId} finished successfully.` }),
});

await AuditProgressKV.clear(auditId);

```

Because `finalizeAudit` runs inside the same Cloudflare Worker durable execution, additional async operations are automatically checkpointed and resumed if the workflow pauses.

## Method 6: Create Entirely New Workflow Phases

If the built-in phases don't cover your requirements—for example, adding a pre-crawl HTML validation step—you can create a new phase that follows the established function signature.

New phases must accept a **`WorkflowStep`** and a typed parameter object, returning `Promise<void>`:

```typescript
// src/server/workflows/siteAuditWorkflowPhases.ts (new function)
export async function runHtmlValidationPhase(
  step: WorkflowStep,
  params: {
    auditId: string;
    workflowInstanceId: string;
    startUrl: string;
    config: AuditConfig;
  }
): Promise<void> {
  // Custom validation logic here
  const validationErrors = await validateHtml(startUrl);
  await AuditRepository.saveValidationErrors(params.auditId, validationErrors);
}

```

Then insert the phase into the main orchestration function `runAuditPhases` between existing calls, and add a corresponding step configuration in [`src/server/workflows/auditStepConfigs.ts`](https://github.com/every-app/open-seo/blob/main/src/server/workflows/auditStepConfigs.ts):

```typescript
await runDiscoveryPhase(step, { auditId, workflowInstanceId, startUrl, config });
await runHtmlValidationPhase(step, { auditId, workflowInstanceId, startUrl, config }); // New
await runCrawlPhase(step, { auditId, workflowInstanceId, config });

```

## Summary

- **Configuration-driven customization**: Pass custom values via `AuditConfig` to control crawl limits and Lighthouse sampling without code changes.
- **Crawl policy extension**: Modify `shouldQueueCrawlLink` in [`siteAuditWorkflowCrawl.ts`](https://github.com/every-app/open-seo/blob/main/siteAuditWorkflowCrawl.ts) to inject custom URL filtering logic.
- **Lighthouse sampling**: Add new strategies to `selectLighthouseSample` in [`lighthouse.ts`](https://github.com/every-app/open-seo/blob/main/lighthouse.ts) for custom page selection algorithms.
- **Multipage issue detection**: Export new detectors from the [`multipage/index.ts`](https://github.com/every-app/open-seo/blob/main/multipage/index.ts) aggregate to run custom cross-page validations.
- **Finalization hooks**: Extend `finalizeAudit` to add notifications, webhooks, or custom reporting triggers.
- **New workflow phases**: Implement functions matching the `(step, params) => Promise<void>` signature to add entirely new audit stages.

## Frequently Asked Questions

### How do I limit the crawl to specific subdirectories?

Modify the `shouldQueueCrawlLink` function in [`src/server/workflows/siteAuditWorkflowCrawl.ts`](https://github.com/every-app/open-seo/blob/main/src/server/workflows/siteAuditWorkflowCrawl.ts) to test URL patterns against a whitelist. For example, add a check like `new URL(link).pathname.startsWith('/allowed-path')` before returning `true`. This runs during the crawl phase for every discovered link without affecting the discovery or Lighthouse phases.

### Can I disable Lighthouse checks entirely for certain audits?

Yes. Set `lighthouseStrategy` to `"none"` in the `AuditConfig` when calling the audit creation endpoint. The `runLighthousePhase` function in [`siteAuditWorkflowPhases.ts`](https://github.com/every-app/open-seo/blob/main/siteAuditWorkflowPhases.ts) checks this value and skips execution when the strategy is `"none"`, allowing faster audits that focus solely on crawl data.

### Where should I store custom audit data that persists between workflow steps?

Use the `AuditScratchpad` durable object located in [`src/server/features/audit/AuditScratchpad.ts`](https://github.com/every-app/open-seo/blob/main/src/server/features/audit/AuditScratchpad.ts) to store temporary state across workflow steps. For data that needs to survive the audit completion, write to the `AuditRepository` in [`src/server/features/audit/repositories/AuditRepository.ts`](https://github.com/every-app/open-seo/blob/main/src/server/features/audit/repositories/AuditRepository.ts), which provides the database abstraction for persisting pages, links, and issue records.