# Wigolo Byte-Offset Source Span Tracking: How It Enables Verifiable Citations

> Discover Wigolo's byte-offset source span tracking for verifiable citations. Learn how exact character ranges and SHA-1 hashes ensure independent verification of your sources.

- Repository: [Towhid Khan/wigolo](https://github.com/KnockOutEZ/wigolo)
- Tags: how-to-guide
- Published: 2026-07-19

---

**Wigolo tracks exact character ranges in source documents using byte-offset source spans, then generates deterministic SHA-1 hashes to create citation IDs that can be independently verified.**

Wigolo is an open-source search and citation engine that maps generated answers back to their original source text with character-level precision. By implementing **byte-offset source span tracking** for every evidence snippet, the system creates an immutable audit trail from output claims to specific character positions in the source markdown. This design ensures that any citation can be cryptographically verified and traced to its origin.

## What Is Byte-Offset Source Span Tracking?

At the core of Wigolo's verification system is the `SourceSpan` interface, defined in [[`src/types.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/types.ts)](https://github.com/KnockOutEZ/wigolo/blob/main/src/types.ts#L35-L38):

```typescript
export interface SourceSpan {
  start: number;   // inclusive offset in the source markdown
  end:   number;   // exclusive offset
}

```

This structure records the exact byte positions where a snippet begins and ends. The **start** offset is inclusive, marking the first character of the passage, while the **end** offset is exclusive, pointing to the character immediately following the snippet. By storing these precise coordinates, Wigolo creates a mathematical mapping between generated content and source documents.

## How Wigolo Captures Source Spans

The system populates `SourceSpan` objects during the highlight extraction pipeline, ensuring every candidate passage retains its positional metadata.

### Highlight Extraction Pipeline

When the search pipeline identifies relevant passages, each candidate records its character offsets as `charStart` and `charEnd`. These values are copied into the public `Highlight` shape in [[`src/search/highlights.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/highlights.ts)](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/highlights.ts#L172-L174):

```typescript
source_span: { start: cand.charStart, end: cand.charEnd },

```

This assignment preserves the exact byte positions for downstream processing, allowing the system to reference the original text location regardless of subsequent transformations.

### Fallback Highlight Generation

If the reranker is unavailable, Wigolo invokes the **`fallbackHighlights`** function to create spans that cover the chosen text. Located at [[`src/search/highlights.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/highlights.ts)](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/highlights.ts#L86-L92), this logic ensures that even degraded operation modes maintain source span integrity.

## Generating Verifiable Citations

Once spans are captured, Wigolo converts them into deterministic citation identifiers that serve as cryptographic proofs of origin.

### From Spans to Evidence Items

The conversion from highlights to evidence objects occurs in [[`src/search/evidence.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/evidence.ts)](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/evidence.ts#L90-L99). The function copies the span unchanged into the `EvidenceItem` structure:

```typescript
source_span: input.sourceSpan,

```

This preservation ensures that the original byte offsets travel intact through the entire processing pipeline.

### Creating Stable Citation IDs

Wigolo derives deterministic identifiers using the **`stableCitationId`** helper in [[`src/search/evidence.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/evidence.ts)](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/evidence.ts#L70-L72):

```typescript
export function stableCitationId(url: string, start: number): string {
  return createHash('sha1')
    .update(`${url}#${start}`)
    .digest('hex')
    .slice(0, 12);
}

```

The function computes a SHA-1 hash of the string `${url}#${start}`, then truncates the result to a 12-character hexadecimal string. Because the hash input combines the immutable URL with the span's start offset, the same passage always yields the same ID across requests. This **deterministic citation ID** becomes the verifiable fingerprint for the evidence.

### Assembling the Final Citation List

After evidence collection, the **`buildCitationsFromEvidence`** function ([[`src/search/evidence.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/evidence.ts)](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/evidence.ts#L95-L107)) selects the highest-scoring passage for each source URL. It copies the `citation_id` into the corresponding `Citation` object, creating the final link between evidence spans and output citations.

## Verifying Citations in Practice

The resulting API response contains two critical arrays: `evidence` objects with `source_span` and `citation_id`, and `citations` entries that reference these IDs. Clients can locate the original text by applying the byte offsets and verify authenticity by recomputing the SHA-1 hash on `url#start`.

### Building Evidence from Markdown

To extract evidence with verifiable spans:

```typescript
import { buildEvidenceFromMarkdown } from './search/evidence.js';

const evidence = await buildEvidenceFromMarkdown(
  'what is quantum computing',
  'Quantum Computing – Wikipedia',
  'https://en.wikipedia.org/wiki/Quantum_computing',
  fetchedMarkdown,
  { maxItems: 3, maxTokensOut: 500 }
);

console.log(evidence[0].source_span);   // { start: 124, end: 250 }
console.log(evidence[0].citation_id);   // "a3f9c2d1b4e5"

```

### Generating Citation IDs Manually

You can reproduce the citation ID calculation:

```typescript
import { stableCitationId } from './search/evidence.js';

const url = 'https://example.com/article';
const startOffset = 342;
const id = stableCitationId(url, startOffset);
console.log(id);                       // "7b9f1c4e2a6d"

```

### Verifying API Responses

Implement client-side verification:

```typescript
function verifyCitation(citation) {
  const expected = stableCitationId(citation.url, citation.source_span.start);
  return expected === citation.citation_id;
}

resp.evidence.forEach(ev => {
  console.assert(
    verifyCitation({ url: ev.url, source_span: ev.source_span, citation_id: ev.citation_id })
  );
});

```

## Summary

- **SourceSpan Interface**: Defined in [`src/types.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/types.ts), stores inclusive start and exclusive end byte offsets for every snippet.
- **Highlight Tracking**: The extraction pipeline in [`src/search/highlights.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/highlights.ts) captures exact character ranges during search processing.
- **Deterministic IDs**: The `stableCitationId` function generates 12-character SHA-1 hashes from the URL and start offset.
- **Evidence Pipeline**: [`src/search/evidence.ts`](https://github.com/KnockOutEZ/wigolo/blob/main/src/search/evidence.ts) propagates spans through to `EvidenceItem` objects and final `Citation` arrays.
- **Verifiable Output**: Clients can locate original text using byte offsets and verify authenticity by recomputing citation hashes.

## Frequently Asked Questions

### How does Wigolo handle text that moves within a document?

Wigolo's citation IDs are based on byte offsets relative to the fetched markdown at search time. If the source document changes, the same logical text will likely occupy different byte positions, generating a different citation ID. This design treats citations as snapshots tied to specific document versions, ensuring that verifications always reference the exact text that was retrieved.

### Can citation IDs be predicted without accessing Wigolo's API?

Yes. Since `stableCitationId` uses a standard SHA-1 hash of `${url}#${start}`, any client can precompute the expected citation ID given the URL and start offset. This property enables third-party verification systems to confirm citations without requiring access to Wigolo's internal state or database.

### What happens to citations when passages exceed token budgets?

When the `buildCitationsFromEvidence` function trims evidence to fit token constraints, citations that survive the budget retain their `citation_id`, while trimmed passages result in source-level citations without specific IDs. This ensures that the system never invents span data for truncated content, maintaining the integrity of the byte-offset tracking system.

### Why does Wigolo use byte offsets instead of line numbers?

Byte offsets provide character-level precision that remains stable regardless of formatting changes, line wrapping, or display variations. Unlike line numbers, which shift when documents are reformatted, byte offsets in the raw markdown create an immutable coordinate system that cryptographic hashes can reliably reference.