How llmfit-core Assembles a PR-Ready Payload from Locally Cached Benchmark Results

The share.rs module in llmfit-core reads pending benchmark JSON files from the local store, sanitizes absolute GGUF paths, enriches results with hardware metadata, and serializes them into a camelCase JSON payload that is uploaded to a new GitHub branch and submitted as a pull request.

The llmfit-core crate in the AlexsJones/llmfit repository manages the entire lifecycle of community benchmark submissions, from local caching to GitHub integration. When users run benchmarks using the bench command, results are stored in a pending queue before being transformed into a standardized submission format. This article explains how llmfit-core/src/share.rs orchestrates the assembly of these cached results into a PR-ready payload that conforms to the community schema.

Loading and Sanitizing Pending Benchmarks

The process begins by reading benchmark data from the local filesystem. The read_store function (located at approximately line 70 in share.rs) walks the directory returned by store_root()/pending (typically ~/.local/share/llmfit/benchmarks/pending) and deserializes each *.json file into a StoredBenchmark struct.

Files are sorted by creation time to ensure deterministic ordering. As benchmarks are parsed, the sanitize_stored_payload function (line 57) strips any absolute GGUF paths from the model field to prevent leaking user-specific directory structures. This sanitization step ensures that community submissions contain only relative model identifiers rather than local file system paths.

Removing Sensitive Path Information

The strip_gguf_path helper specifically targets absolute paths embedded in benchmark results. By removing these before payload construction, the module guarantees that submissions contain portable model references suitable for public sharing. Sanitized results are collected into a Vec<StoredBenchmark> ordered by file-creation time.

Building the Submission Payload

With raw benchmarks loaded, the build_submission function (line 41) constructs the final payload structure. This function gathers system information via SystemSpecs::detect(), classifies the hardware into one of three tiers—UNIFIED, DISCRETE_GPU, or CPU_ONLY—and computes a declared_mem_tier for the hardware specification.

Each BenchResult is mapped to a ResultPayload with numeric values rounded to two decimal places. The resulting Submission struct contains:

  • schema_version: A constant set to 1 (SCHEMA_VERSION)
  • submitted_at_unix: Current epoch seconds from now_unix()
  • tool: Static metadata identifying the tool as llmfit
  • hardware: An HwPayload populated from SystemSpecs containing hw_class, vram_gb, cpu, and other system details
  • results: A vector of per-model ResultPayload entries

The struct derives Serialize with #[serde(rename_all = "camelCase")], ensuring the final JSON follows community naming conventions.

Hardware Classification and Metadata

The hardware classification logic determines whether the system uses unified memory (Apple Silicon), discrete GPU architecture, or CPU-only inference. This classification affects how the community interprets performance metrics. The SystemSpecs detection provides accurate VRAM calculations and CPU identification that feed into the HwPayload serialization.

Preparing Files for GitHub Upload

Before interacting with the GitHub API, the prepare_files function (line 56) processes each stored benchmark into a upload-ready tuple. For every entry, it clones the JSON payload, overwrites submittedAtUnix with the current timestamp, and re-serializes the content.

The function extracts the original file name or falls back to <timestamp>-<index>.json, then computes a hardware slug via payload_slug (line 38). This slug, derived from the hardware name and sanitized for filesystem safety, pairs with a short_hash (FNV-1a hash of the JSON plus timestamp) to produce the final output: Vec<(slug, file_name, json)>.

GitHub Branch Management and File Upload

The submission logic checks for existing open PRs using find_open_bench_pr (line 113). If a branch matching bench/* already exists with an open PR, the system reuses that branch rather than creating duplicates. Otherwise, upstream_head_sha (line 77) queries the upstream repository state, and create_branch (line 90) establishes a new branch named bench/<slug>-<hash>.

Files are uploaded via put_files (line 78), which delegates to put_file (line 106) for each tuple. The put_file function sends a PUT request to the GitHub contents endpoint with content base64-encoded using base64::engine::general_purpose::STANDARD. If GitHub returns a 422 status with a "sha" error, indicating the file already exists, the upload is counted as skipped rather than failing the entire operation.

Opening the Pull Request

Once files are staged, open_pr (line 158) creates the pull request with the title format bench: community results for <slug>. The PR body contains a markdown table summarizing every model with columns for model name, provider, average TPS, and average TTFT. A footnote indicates the PR was created programmatically without the gh CLI.

The entire workflow is orchestrated by share_all_pending, which handles interactive dry-run prompts and confirmation logic before delegating to submit_stored for the actual GitHub operations.

Practical Code Examples

To store benchmarks locally before sharing:

// Run a benchmark and cache results locally
let bench_results = llmfit_core::bench::run("llama3.1:8b", /* config */)?;
let specs = llmfit_core::hardware::SystemSpecs::detect()?;
let path = llmfit_core::share::store_local(&bench_results, &specs)?;
println!("Stored submission at {}", path.display());

To submit all pending benchmarks in a CI environment:

// Share all pending benchmarks non-interactively
let opts = llmfit_core::share::ShareOptions { 
    dry_run: false, 
    assume_yes: true 
};
let token = llmfit_core::share::resolve_token_noninteractive()
    .ok_or("GitHub token not found")?;
let outcome = llmfit_core::share::share_all_pending(&opts, Some(token))?;
println!("PR opened: {:?}", outcome?.pr_url);

Summary

  • Three-stage pipeline: The module loads raw JSON from store_root()/pending, sanitizes absolute paths, builds a Submission struct with hardware metadata, and uploads via the GitHub contents API.
  • Path sanitization: sanitize_stored_payload and strip_gguf_path remove user-specific directory information before community submission.
  • Hardware classification: Automatic detection categorizes systems as UNIFIED, DISCRETE_GPU, or CPU_ONLY with accurate memory tier reporting.
  • Idempotent uploads: The system handles existing files gracefully (HTTP 422) and reuses existing PR branches when available.
  • Schema compliance: JSON output uses camelCase field names and schema version 1 for compatibility with community benchmark aggregators.

Frequently Asked Questions

What JSON format does the PR-ready payload use?

The payload follows a strict schema with camelCase field names as specified by the #[serde(rename_all = "camelCase")] attribute on the Submission struct. The root object includes schemaVersion (always 1), submittedAtUnix, tool metadata, a hardware object describing the system specs, and a results array containing per-model benchmarks with rounded numeric values.

How does llmfit-core prevent leaking local file system paths?

According to the source code in share.rs, the sanitize_stored_payload function processes every loaded benchmark and calls strip_gguf_path to remove absolute GGUF paths from the model field. This ensures that only portable model identifiers reach the public repository, protecting user privacy by eliminating local directory structures from community submissions.

Can the tool update an existing pull request with new benchmarks?

Yes. The find_open_bench_pr function checks for existing open PRs whose head branch matches the bench/* pattern. If found, submit_stored reuses that existing branch rather than creating a new one, allowing users to accumulate multiple benchmark results into a single community submission. New files are added to the existing branch via the contents API.

What happens if a benchmark file already exists on the remote branch?

The put_file implementation handles HTTP 422 responses from the GitHub API that indicate a file conflict (sha mismatch). When this occurs, the file is counted as "skipped" in the upload statistics, and the process continues with the remaining files. This idempotent behavior prevents duplicate uploads from failing the entire submission workflow.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →