How the Prompt Optimizer Comparison Service Enables Prompt A/B Testing and Evaluation
The comparison service in Prompt Optimizer wraps the jsdiff library to generate structured diffs between original and optimized prompts, producing fragment-level changes and summary statistics that power the evaluation engine's A/B testing workflow.
The linshenkx/prompt-optimizer repository provides a sophisticated system for refining LLM prompts through iterative optimization. At the heart of its evaluation pipeline sits the comparison service, a dedicated module that quantifies textual differences between prompt versions to enable data-driven A/B testing decisions.
Core Architecture and Entry Points
The comparison functionality resides in packages/core/src/services/compare/service.ts and exposes a factory function createCompareService() that returns an object implementing the ICompareService interface. This service acts as a thin, type-safe wrapper around the jsdiff library, abstracting the complexity of text diffing while providing domain-specific processing for prompt comparison.
The primary public method is compareTexts(original, optimized, options?), which accepts two prompt strings and an optional configuration object. This method validates inputs, merges user-provided options with sensible defaults (word-level granularity, case-sensitive matching, whitespace preservation), and orchestrates the diffing workflow before returning a structured CompareResult.
The Comparison Workflow
The service processes prompt pairs through a six-stage pipeline that transforms raw text differences into actionable metrics.
Input Validation and Configuration
Before processing begins, the service validates inputs through the validateInput function. If either argument is not a string, the service throws a CompareValidationError, ensuring type safety at runtime. Valid inputs trigger an options merge sequence that combines user preferences with defaults: word-level diffing, case-sensitive comparison, and no whitespace ignore flags.
Text Preprocessing
Depending on the ignoreWhitespace and caseSensitive options passed to compareTexts(), the service normalizes both strings. When enabled, whitespace is collapsed and text is converted to lowercase in packages/core/src/services/compare/service.ts lines 84-96, ensuring that stylistic differences do not obscure semantic changes during the A/B test.
Diff Calculation with jsdiff
The core diffing logic leverages the jsdiff library's algorithms. Based on the granularity option, the service dispatches to either diffChars for character-level precision or diffWords for word-level analysis. This allows developers to tune the sensitivity of comparisons—character mode catches typos and small edits, while word mode focuses on meaningful vocabulary changes between prompt versions.
Result Processing and Fragment Generation
Raw Change[] objects from jsdiff are converted into the library-specific TextFragment shape defined in packages/core/src/services/compare/types.ts. Each fragment contains the text content, change type (added, removed, or unchanged), and positional index. Critically, unchanged fragments retain exact slices from the original string to preserve formatting and spacing.
The service then employs a fragment merging pass that collapses consecutive fragments of the same ChangeType. This compaction reduces visual noise in the final output while maintaining the semantic integrity of the diff.
Summary Statistics Generation
In the final stage, the service walks the merged fragment list to populate the summary object. It tallies total additions, deletions, and unchanged pieces, exposing these counts in the CompareResult output. These metrics allow evaluation templates to quantify precisely how much an optimized prompt diverges from its original version.
Data Models and Type Safety
The comparison output adheres strictly to the CompareResult interface defined in packages/core/src/services/compare/types.ts. This contract specifies an ordered fragments array containing TextFragment objects and a numeric summary object tracking change statistics.
The type definitions include the ChangeType enum (typically 'added', 'removed', 'unchanged') and optional configuration interfaces allowing fine-grained control over diffing behavior. This type safety ensures that downstream consumers—whether UI components or evaluation engines—can depend on a consistent data structure.
Integration with Evaluation Templates
The comparison service feeds directly into the optimizer's evaluation system. Templates such as packages/core/src/services/template/default-templates/evaluation/pro/user/evaluation-compare_en.ts consume the CompareResult to render A/B testing interfaces.
When running a prompt evaluation, the service's output is embedded in a JSON payload containing the fragment list and summary statistics. This structured data enables automated scoring logic to assess whether additions and deletions represent genuine improvements or regressions, providing empirical backing for optimization decisions.
Practical Implementation Example
The following example demonstrates how to instantiate the service and compare two prompt variants:
import { createCompareService } from '@prompt-optimizer/core/src/services/compare/service';
// Initialize the comparison service
const comparer = createCompareService();
// Define prompt variants for A/B testing
const originalPrompt = `Generate a short summary of the input.`;
const optimizedPrompt = `Provide a concise 2‑sentence summary of the given text.`;
// Configure comparison options
const options = {
granularity: 'word',
ignoreWhitespace: true,
caseSensitive: false
};
// Execute comparison
const result = comparer.compareTexts(originalPrompt, optimizedPrompt, options);
console.log('Diff fragments:', result.fragments);
console.log('Summary:', result.summary);
/*
Diff fragments → [
{ text: 'Generate a ', type: 'unchanged', index: 0 },
{ text: 'short ', type: 'removed', index: 1 },
{ text: 'concise 2‑sentence ', type: 'added', index: 2 },
...
]
Summary → { additions: 2, deletions: 1, unchanged: 3 }
*/
For integration with evaluation pipelines, pass the result to the evaluation engine alongside test outputs:
const evaluationPayload = {
comparison: result,
originalTestResult: originalRes,
optimizedTestResult: optimizedRes
};
await evaluatePromptComparison(evaluationPayload);
Summary
- The CompareService in
packages/core/src/services/compare/service.tsprovides a type-safe wrapper around the jsdiff library for prompt comparison. - The
compareTexts()method validates inputs, preprocesses text, and executes either character-level or word-level diffs based on thegranularityoption. - Results are standardized into the CompareResult interface containing TextFragment arrays and numeric summaries of additions, deletions, and unchanged content.
- The service integrates with evaluation templates to support automated scoring of prompt A/B tests, enabling data-driven optimization decisions.
- Configuration options include whitespace normalization, case sensitivity toggles, and diff granularity controls.
Frequently Asked Questions
What library powers the diff calculation in the comparison service?
The service delegates to jsdiff, a popular JavaScript diffing library. Depending on configuration, it invokes either diffChars() for character-level granularity or diffWords() for word-level analysis, as implemented in packages/core/src/services/compare/service.ts lines 99-106.
How does the service handle whitespace differences?
Whitespace handling is controlled via the ignoreWhitespace option. When set to true, the service normalizes whitespace through collapsing and trimming operations during the preprocessing stage, preventing formatting variations from registering as semantic changes in the diff output.
What granularity levels are supported for prompt comparison?
The service supports two granularity modes specified via the granularity option: 'char' for character-level diffs that catch typos and minor edits, and 'word' for word-level diffs that focus on vocabulary and phrasing changes. These map directly to jsdiff's diffChars and diffWords methods.
How is the comparison result used in the A/B testing workflow?
Evaluation templates consume the CompareResult object—specifically the fragments array and summary statistics—to construct prompts for the evaluation engine. This allows automated scoring systems to quantify textual differences and assess whether the optimized prompt represents an improvement over the original version according to predefined metrics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →