# ULTRAPLINIAN Scoring Algorithm: How G0DM0D3 Ranks Parallel LLM Responses

> Discover the ULTRAPLINIAN scoring algorithm used by G0DM0D3 to rank parallel LLM responses. Analyze model outputs with a deterministic five-axis function for interpretable scores.

- Repository: [pliny/G0DM0D3](https://github.com/elder-plinius/G0DM0D3)
- Tags: deep-dive
- Published: 2026-07-19

---

**The ULTRAPLINIAN scoring algorithm evaluates model outputs using a deterministic, five-axis function that returns an interpretable integer score from 0 to 100 based on length, structure, anti-refusal patterns, directness, and relevance.**

The ULTRAPLINIAN scoring algorithm powers the selection logic in elder-plinius/G0DM0D3, an open-source framework that runs parallel LLM races to identify the highest-quality response. Unlike opaque neural evaluators, this algorithm implements a transparent, policy-driven scoring system in [`api/lib/ultraplinian.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/lib/ultraplinian.ts) that weighs specific quality dimensions to rank outputs deterministically.

## How the ULTRAPLINIAN Scoring Algorithm Works

At the core of the system lies the `scoreResponse(content, userQuery)` function implemented in [`api/lib/ultraplinian.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/lib/ultraplinian.ts) (lines 71-101). This function analyzes a candidate response against five orthogonal criteria, each weighted to reflect specific quality priorities. The cumulative result is bounded to a 0-100 integer scale using `Math.round(Math.min(score, 100))`, ensuring consistent comparability across different model outputs.

## The Five Axes of ULTRAPLINIAN Scoring

The algorithm allocates 100 points across five distinct axes, with **length** and **anti-refusal** receiving the highest weights (25 points each) to prioritize substantive answers that avoid hedging.

### Length: Measuring Substance (25 Points)

Longer responses tend to contain more information. The algorithm awards up to 25 points based on character count using the formula `Math.min(content.length / 40, 25)`, creating a linear scale that caps at 1,000 characters.

### Structure: Rewarding Formatting Richness (20 Points)

Well-formatted responses receive up to 20 points based on structural elements. The calculation `headers * 3 + listItems * 1.5 + codeBlocks * 5` rewards Markdown headers, bullet lists, and fenced code blocks, with the total capped at 20 points.

### Anti-Refusal: Penalizing Hedge Phrases (25 Points)

To combat evasive model behavior, the algorithm scans for patterns defined in `REFUSAL_PATTERNS`. Each detected match subtracts 8 points from the 25-point allocation, with a floor of 0, ensuring that refusals significantly damage the final score.

### Directness: Valuing Immediate Answers (15 Points)

Responses that jump straight to solutions outperform those with lengthy preambles. If `PREAMBLE_PATTERNS` detects introductory fluff, the response receives only 8 points; otherwise, it earns the full 15 points.

### Relevance: Matching Query Intent (15 Points)

The final axis measures topical alignment by calculating the fraction of significant query words (≥4 characters) that appear in the response. This ratio is multiplied by 15 to determine the relevance score.

## Implementation Details and Code Examples

The deterministic scoring logic resides in [`api/lib/ultraplinian.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/lib/ultraplinian.ts), where the `scoreResponse` function combines these axes into a final ranking metric. Here is how the algorithm evaluates different response types:

```typescript
import { scoreResponse } from './api/lib/ultraplinian'

// Example 1 – a concise but on-topic reply
const short = `Sure, here's the answer.`
console.log(scoreResponse(short, 'Explain how ULTRAPLINIAN scores responses'))
// → low score (few length points, short structure, possible preamble)

// Example 2 – a rich, well-structured answer
const rich = `

# How ULTRAPLINIAN Scores

- **Length**: 500+ words
- **Structure**: Headers, lists, and code blocks

\`\`\`ts
function foo() { return 42 }
\`\`\`
`
console.log(scoreResponse(rich, 'Explain how ULTRAPLINIAN scores responses'))
// → high score (maxed length, structure, no refusal, high relevance)

```

## Integration with the G0DM0D3 Pipeline

The scoring system integrates deeply into the application's architecture. The live-race state management in [`src/store/index.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/store/index.ts) tracks the current leader's score across parallel model streams, while [`src/components/ChatMessage.tsx`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/components/ChatMessage.tsx) renders the final score adjacent to each response using `{activeResponse.score}pts`. For calibration and validation, the repository includes [`research/eval_scoring_calibration.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/research/eval_scoring_calibration.ts), which provides scripts to tune the axis weights and verify scoring behavior against reference datasets.

## Summary

- The ULTRAPLINIAN scoring algorithm uses a deterministic, five-axis function (`scoreResponse`) in [`api/lib/ultraplinian.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/lib/ultraplinian.ts) to rank LLM outputs.
- Scores range from 0-100, with **length** and **anti-refusal** weighted highest (25 points each) to prioritize substantive, non-hedged answers.
- **Structure** (20 points), **directness** (15 points), and **relevance** (15 points) provide granular quality signals based on formatting, preamble absence, and query-word overlap.
- The algorithm explicitly avoids neural scoring in favor of interpretable, policy-driven metrics that can be audited and calibrated via [`research/eval_scoring_calibration.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/research/eval_scoring_calibration.ts).

## Frequently Asked Questions

### What is the maximum possible score in ULTRAPLINIAN?

The maximum score is **100 points**, achieved when a response maximizes all five axes: substantial length (25/25), rich structure (20/20), no refusal patterns detected (25/25), no preamble detected (15/15), and perfect query-word overlap (15/15). The final result is explicitly bounded using `Math.round(Math.min(score, 100))`.

### How does ULTRAPLINIAN detect and penalize refusal patterns?

The algorithm maintains a `REFUSAL_PATTERNS` array of regular expressions that match common hedge phrases like "I cannot" or "I'm not able." Each match subtracts 8 points from the 25-point anti-refusal allocation, with a hard floor at 0 points, ensuring multiple refusal indicators do not produce negative scores.

### Where is the ULTRAPLINIAN scoring function implemented?

The primary implementation resides in [`api/lib/ultraplinian.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/lib/ultraplinian.ts) at lines 71-101, which exports the `scoreResponse(content, userQuery)` function. Supporting infrastructure appears in [`src/store/index.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/store/index.ts) for state management and [`src/components/ChatMessage.tsx`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/components/ChatMessage.tsx) for UI rendering.

### Why does ULTRAPLINIAN use a deterministic algorithm instead of neural scoring?

The designers prioritized **interpretability** and **policy alignment** over opaque neural evaluation. By using explicit axis weights (25-20-25-15-15) and transparent regex-based detection for preambles and refusals, the system allows operators to audit exactly why a specific response won the parallel race, enabling fine-grained calibration via [`research/eval_scoring_calibration.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/research/eval_scoring_calibration.ts) without retraining opaque models.