# How the Prompt Optimizer Comparison Service Enables Prompt A/B Testing and Evaluation

> Discover how the Prompt Optimizer comparison service enables prompt A/B testing. It generates structured diffs and summary statistics to power your evaluation workflow effectively.

- Repository: [且炼时光/prompt-optimizer](https://github.com/linshenkx/prompt-optimizer)
- Tags: how-to-guide
- Published: 2026-02-23

---

**The comparison service in Prompt Optimizer wraps the jsdiff library to generate structured diffs between original and optimized prompts, producing fragment-level changes and summary statistics that power the evaluation engine's A/B testing workflow.**

The linshenkx/prompt-optimizer repository provides a sophisticated system for refining LLM prompts through iterative optimization. At the heart of its evaluation pipeline sits the **comparison service**, a dedicated module that quantifies textual differences between prompt versions to enable data-driven A/B testing decisions.

## Core Architecture and Entry Points

The comparison functionality resides in [`packages/core/src/services/compare/service.ts`](https://github.com/linshenkx/prompt-optimizer/blob/main/packages/core/src/services/compare/service.ts) and exposes a factory function `createCompareService()` that returns an object implementing the `ICompareService` interface. This service acts as a thin, type-safe wrapper around the **jsdiff** library, abstracting the complexity of text diffing while providing domain-specific processing for prompt comparison.

The primary public method is `compareTexts(original, optimized, options?)`, which accepts two prompt strings and an optional configuration object. This method validates inputs, merges user-provided options with sensible defaults (word-level granularity, case-sensitive matching, whitespace preservation), and orchestrates the diffing workflow before returning a structured `CompareResult`.

## The Comparison Workflow

The service processes prompt pairs through a six-stage pipeline that transforms raw text differences into actionable metrics.

### Input Validation and Configuration

Before processing begins, the service validates inputs through the `validateInput` function. If either argument is not a string, the service throws a `CompareValidationError`, ensuring type safety at runtime. Valid inputs trigger an options merge sequence that combines user preferences with defaults: word-level diffing, case-sensitive comparison, and no whitespace ignore flags.

### Text Preprocessing

Depending on the `ignoreWhitespace` and `caseSensitive` options passed to `compareTexts()`, the service normalizes both strings. When enabled, whitespace is collapsed and text is converted to lowercase in [`packages/core/src/services/compare/service.ts`](https://github.com/linshenkx/prompt-optimizer/blob/main/packages/core/src/services/compare/service.ts) lines 84-96, ensuring that stylistic differences do not obscure semantic changes during the A/B test.

### Diff Calculation with jsdiff

The core diffing logic leverages the **jsdiff** library's algorithms. Based on the `granularity` option, the service dispatches to either `diffChars` for character-level precision or `diffWords` for word-level analysis. This allows developers to tune the sensitivity of comparisons—character mode catches typos and small edits, while word mode focuses on meaningful vocabulary changes between prompt versions.

### Result Processing and Fragment Generation

Raw `Change[]` objects from jsdiff are converted into the library-specific `TextFragment` shape defined in [`packages/core/src/services/compare/types.ts`](https://github.com/linshenkx/prompt-optimizer/blob/main/packages/core/src/services/compare/types.ts). Each fragment contains the text content, change type (added, removed, or unchanged), and positional index. Critically, unchanged fragments retain exact slices from the original string to preserve formatting and spacing.

The service then employs a fragment merging pass that collapses consecutive fragments of the same `ChangeType`. This compaction reduces visual noise in the final output while maintaining the semantic integrity of the diff.

### Summary Statistics Generation

In the final stage, the service walks the merged fragment list to populate the `summary` object. It tallies total **additions**, **deletions**, and **unchanged** pieces, exposing these counts in the `CompareResult` output. These metrics allow evaluation templates to quantify precisely how much an optimized prompt diverges from its original version.

## Data Models and Type Safety

The comparison output adheres strictly to the `CompareResult` interface defined in [`packages/core/src/services/compare/types.ts`](https://github.com/linshenkx/prompt-optimizer/blob/main/packages/core/src/services/compare/types.ts). This contract specifies an ordered `fragments` array containing `TextFragment` objects and a numeric `summary` object tracking change statistics.

The type definitions include the `ChangeType` enum (typically 'added', 'removed', 'unchanged') and optional configuration interfaces allowing fine-grained control over diffing behavior. This type safety ensures that downstream consumers—whether UI components or evaluation engines—can depend on a consistent data structure.

## Integration with Evaluation Templates

The comparison service feeds directly into the optimizer's evaluation system. Templates such as [`packages/core/src/services/template/default-templates/evaluation/pro/user/evaluation-compare_en.ts`](https://github.com/linshenkx/prompt-optimizer/blob/main/packages/core/src/services/template/default-templates/evaluation/pro/user/evaluation-compare_en.ts) consume the `CompareResult` to render A/B testing interfaces.

When running a prompt evaluation, the service's output is embedded in a JSON payload containing the fragment list and summary statistics. This structured data enables automated scoring logic to assess whether additions and deletions represent genuine improvements or regressions, providing empirical backing for optimization decisions.

## Practical Implementation Example

The following example demonstrates how to instantiate the service and compare two prompt variants:

```typescript
import { createCompareService } from '@prompt-optimizer/core/src/services/compare/service';

// Initialize the comparison service
const comparer = createCompareService();

// Define prompt variants for A/B testing
const originalPrompt = `Generate a short summary of the input.`;
const optimizedPrompt = `Provide a concise 2‑sentence summary of the given text.`;

// Configure comparison options
const options = { 
  granularity: 'word', 
  ignoreWhitespace: true,
  caseSensitive: false 
};

// Execute comparison
const result = comparer.compareTexts(originalPrompt, optimizedPrompt, options);

console.log('Diff fragments:', result.fragments);
console.log('Summary:', result.summary);
/*
  Diff fragments → [
    { text: 'Generate a ', type: 'unchanged', index: 0 },
    { text: 'short ', type: 'removed', index: 1 },
    { text: 'concise 2‑sentence ', type: 'added', index: 2 },
    ...
  ]
  Summary → { additions: 2, deletions: 1, unchanged: 3 }
*/

```

For integration with evaluation pipelines, pass the result to the evaluation engine alongside test outputs:

```typescript
const evaluationPayload = {
  comparison: result,
  originalTestResult: originalRes,
  optimizedTestResult: optimizedRes
};

await evaluatePromptComparison(evaluationPayload);

```

## Summary

- The **CompareService** in [`packages/core/src/services/compare/service.ts`](https://github.com/linshenkx/prompt-optimizer/blob/main/packages/core/src/services/compare/service.ts) provides a type-safe wrapper around the jsdiff library for prompt comparison.
- The `compareTexts()` method validates inputs, preprocesses text, and executes either character-level or word-level diffs based on the `granularity` option.
- Results are standardized into the **CompareResult** interface containing **TextFragment** arrays and numeric summaries of additions, deletions, and unchanged content.
- The service integrates with evaluation templates to support automated scoring of prompt A/B tests, enabling data-driven optimization decisions.
- Configuration options include whitespace normalization, case sensitivity toggles, and diff granularity controls.

## Frequently Asked Questions

### What library powers the diff calculation in the comparison service?

The service delegates to **jsdiff**, a popular JavaScript diffing library. Depending on configuration, it invokes either `diffChars()` for character-level granularity or `diffWords()` for word-level analysis, as implemented in [`packages/core/src/services/compare/service.ts`](https://github.com/linshenkx/prompt-optimizer/blob/main/packages/core/src/services/compare/service.ts) lines 99-106.

### How does the service handle whitespace differences?

Whitespace handling is controlled via the `ignoreWhitespace` option. When set to `true`, the service normalizes whitespace through collapsing and trimming operations during the preprocessing stage, preventing formatting variations from registering as semantic changes in the diff output.

### What granularity levels are supported for prompt comparison?

The service supports two granularity modes specified via the `granularity` option: **'char'** for character-level diffs that catch typos and minor edits, and **'word'** for word-level diffs that focus on vocabulary and phrasing changes. These map directly to jsdiff's `diffChars` and `diffWords` methods.

### How is the comparison result used in the A/B testing workflow?

Evaluation templates consume the **CompareResult** object—specifically the `fragments` array and `summary` statistics—to construct prompts for the evaluation engine. This allows automated scoring systems to quantify textual differences and assess whether the optimized prompt represents an improvement over the original version according to predefined metrics.