# How llmfit-core Assembles a PR-Ready Payload from Locally Cached Benchmark Results

> Discover how llmfit-core assembles PR-ready payloads from cached benchmark results. Learn how it sanitizes paths, adds metadata, and creates pull requests.

- Repository: [Alex Jones/llmfit](https://github.com/AlexsJones/llmfit)
- Tags: internals
- Published: 2026-09-12

---

**The [`share.rs`](https://github.com/AlexsJones/llmfit/blob/main/share.rs) module in `llmfit-core` reads pending benchmark JSON files from the local store, sanitizes absolute GGUF paths, enriches results with hardware metadata, and serializes them into a camelCase JSON payload that is uploaded to a new GitHub branch and submitted as a pull request.**

The `llmfit-core` crate in the [AlexsJones/llmfit](https://github.com/AlexsJones/llmfit) repository manages the entire lifecycle of community benchmark submissions, from local caching to GitHub integration. When users run benchmarks using the `bench` command, results are stored in a pending queue before being transformed into a standardized submission format. This article explains how [`llmfit-core/src/share.rs`](https://github.com/AlexsJones/llmfit/blob/main/llmfit-core/src/share.rs) orchestrates the assembly of these cached results into a PR-ready payload that conforms to the community schema.

## Loading and Sanitizing Pending Benchmarks

The process begins by reading benchmark data from the local filesystem. The `read_store` function (located at approximately line 70 in [`share.rs`](https://github.com/AlexsJones/llmfit/blob/main/share.rs)) walks the directory returned by `store_root()/pending` (typically `~/.local/share/llmfit/benchmarks/pending`) and deserializes each `*.json` file into a `StoredBenchmark` struct.

Files are sorted by creation time to ensure deterministic ordering. As benchmarks are parsed, the `sanitize_stored_payload` function (line 57) strips any absolute GGUF paths from the `model` field to prevent leaking user-specific directory structures. This sanitization step ensures that community submissions contain only relative model identifiers rather than local file system paths.

### Removing Sensitive Path Information

The `strip_gguf_path` helper specifically targets absolute paths embedded in benchmark results. By removing these before payload construction, the module guarantees that submissions contain portable model references suitable for public sharing. Sanitized results are collected into a `Vec<StoredBenchmark>` ordered by file-creation time.

## Building the Submission Payload

With raw benchmarks loaded, the `build_submission` function (line 41) constructs the final payload structure. This function gathers system information via `SystemSpecs::detect()`, classifies the hardware into one of three tiers—`UNIFIED`, `DISCRETE_GPU`, or `CPU_ONLY`—and computes a `declared_mem_tier` for the hardware specification.

Each `BenchResult` is mapped to a `ResultPayload` with numeric values rounded to two decimal places. The resulting `Submission` struct contains:
- `schema_version`: A constant set to `1` (`SCHEMA_VERSION`)
- `submitted_at_unix`: Current epoch seconds from `now_unix()`
- `tool`: Static metadata identifying the tool as `llmfit`
- `hardware`: An `HwPayload` populated from `SystemSpecs` containing `hw_class`, `vram_gb`, `cpu`, and other system details
- `results`: A vector of per-model `ResultPayload` entries

The struct derives `Serialize` with `#[serde(rename_all = "camelCase")]`, ensuring the final JSON follows community naming conventions.

### Hardware Classification and Metadata

The hardware classification logic determines whether the system uses unified memory (Apple Silicon), discrete GPU architecture, or CPU-only inference. This classification affects how the community interprets performance metrics. The `SystemSpecs` detection provides accurate VRAM calculations and CPU identification that feed into the `HwPayload` serialization.

## Preparing Files for GitHub Upload

Before interacting with the GitHub API, the `prepare_files` function (line 56) processes each stored benchmark into a upload-ready tuple. For every entry, it clones the JSON payload, overwrites `submittedAtUnix` with the current timestamp, and re-serializes the content.

The function extracts the original file name or falls back to `<timestamp>-<index>.json`, then computes a hardware slug via `payload_slug` (line 38). This slug, derived from the hardware name and sanitized for filesystem safety, pairs with a `short_hash` (FNV-1a hash of the JSON plus timestamp) to produce the final output: `Vec<(slug, file_name, json)>`.

## GitHub Branch Management and File Upload

The submission logic checks for existing open PRs using `find_open_bench_pr` (line 113). If a branch matching `bench/*` already exists with an open PR, the system reuses that branch rather than creating duplicates. Otherwise, `upstream_head_sha` (line 77) queries the upstream repository state, and `create_branch` (line 90) establishes a new branch named `bench/<slug>-<hash>`.

Files are uploaded via `put_files` (line 78), which delegates to `put_file` (line 106) for each tuple. The `put_file` function sends a PUT request to the GitHub contents endpoint with content base64-encoded using `base64::engine::general_purpose::STANDARD`. If GitHub returns a `422` status with a "sha" error, indicating the file already exists, the upload is counted as skipped rather than failing the entire operation.

## Opening the Pull Request

Once files are staged, `open_pr` (line 158) creates the pull request with the title format `bench: community results for <slug>`. The PR body contains a markdown table summarizing every model with columns for model name, provider, average TPS, and average TTFT. A footnote indicates the PR was created programmatically without the `gh` CLI.

The entire workflow is orchestrated by `share_all_pending`, which handles interactive dry-run prompts and confirmation logic before delegating to `submit_stored` for the actual GitHub operations.

## Practical Code Examples

To store benchmarks locally before sharing:

```rust
// Run a benchmark and cache results locally
let bench_results = llmfit_core::bench::run("llama3.1:8b", /* config */)?;
let specs = llmfit_core::hardware::SystemSpecs::detect()?;
let path = llmfit_core::share::store_local(&bench_results, &specs)?;
println!("Stored submission at {}", path.display());

```

To submit all pending benchmarks in a CI environment:

```rust
// Share all pending benchmarks non-interactively
let opts = llmfit_core::share::ShareOptions { 
    dry_run: false, 
    assume_yes: true 
};
let token = llmfit_core::share::resolve_token_noninteractive()
    .ok_or("GitHub token not found")?;
let outcome = llmfit_core::share::share_all_pending(&opts, Some(token))?;
println!("PR opened: {:?}", outcome?.pr_url);

```

## Summary

- **Three-stage pipeline**: The module loads raw JSON from `store_root()/pending`, sanitizes absolute paths, builds a `Submission` struct with hardware metadata, and uploads via the GitHub contents API.
- **Path sanitization**: `sanitize_stored_payload` and `strip_gguf_path` remove user-specific directory information before community submission.
- **Hardware classification**: Automatic detection categorizes systems as `UNIFIED`, `DISCRETE_GPU`, or `CPU_ONLY` with accurate memory tier reporting.
- **Idempotent uploads**: The system handles existing files gracefully (HTTP 422) and reuses existing PR branches when available.
- **Schema compliance**: JSON output uses camelCase field names and schema version 1 for compatibility with community benchmark aggregators.

## Frequently Asked Questions

### What JSON format does the PR-ready payload use?

The payload follows a strict schema with camelCase field names as specified by the `#[serde(rename_all = "camelCase")]` attribute on the `Submission` struct. The root object includes `schemaVersion` (always `1`), `submittedAtUnix`, `tool` metadata, a `hardware` object describing the system specs, and a `results` array containing per-model benchmarks with rounded numeric values.

### How does llmfit-core prevent leaking local file system paths?

According to the source code in [`share.rs`](https://github.com/AlexsJones/llmfit/blob/main/share.rs), the `sanitize_stored_payload` function processes every loaded benchmark and calls `strip_gguf_path` to remove absolute GGUF paths from the `model` field. This ensures that only portable model identifiers reach the public repository, protecting user privacy by eliminating local directory structures from community submissions.

### Can the tool update an existing pull request with new benchmarks?

Yes. The `find_open_bench_pr` function checks for existing open PRs whose head branch matches the `bench/*` pattern. If found, `submit_stored` reuses that existing branch rather than creating a new one, allowing users to accumulate multiple benchmark results into a single community submission. New files are added to the existing branch via the contents API.

### What happens if a benchmark file already exists on the remote branch?

The `put_file` implementation handles HTTP 422 responses from the GitHub API that indicate a file conflict (sha mismatch). When this occurs, the file is counted as "skipped" in the upload statistics, and the process continues with the remaining files. This idempotent behavior prevents duplicate uploads from failing the entire submission workflow.