What Is the BackendDescriptor in orx and How Is It Serialized?

The BackendDescriptor is a serde-powered data structure defined in src/jobs/mod.rs that captures backend-specific job metadata and provides parse() and to_json() methods for JSON serialization to SQLite.

The orx CLI (OpenResearch) uses the BackendDescriptor as the single source of truth for recording how a run was launched on a particular compute backend. It is created when a job is submitted and later deserialized by local launch modules to reconnect, monitor, or clean up remote executions.

Core Purpose and Structure in src/jobs/mod.rs

What BackendDescriptor Stores

According to the alphaXiv/OpenResearch source code, the BackendDescriptor struct (defined starting at line 46 in src/jobs/mod.rs) serves as the central data structure for persisting backend-specific metadata. Its primary role is to store the information required to recreate the execution context for any supported backend—whether Modal, Slurm, Kubernetes, SSH, or others—after the initial launch process has exited.

The struct uses serde derives for Serialize and Deserialize, ensuring that any instance can be converted to a JSON string for storage and reliably reconstructed later.

Required vs Optional Fields

The descriptor uses a fixed schema where only the kind field is mandatory. All other fields are Option<T> and are omitted from the JSON output when None:

  • kind (required): Identifies the backend type using string identifiers such as "modal_job", "slurm_job", "ssh_job", "ray_job", "k8s_job", "hf_job", "openresearch_job", "local_job", or "tinker_job".
  • Optional fields: namespace, job_id, flavor, image, url, context, manifest, resources, ssh_host, ssh_port, ssh_user, timeout_secs, source_digest, source_path, and source_size.

The implementation applies #[serde(skip_serializing_if = "Option::is_none")] to optional fields, keeping the serialized JSON compact and backend-agnostic.

JSON Serialization Implementation

The parse() and to_json() Methods

The BackendDescriptor implementation provides two convenience methods for JSON round-trips, located at lines 98–104 in src/jobs/mod.rs:

  • BackendDescriptor::parse(json: &str) -> Result<Self>: Uses serde_json::from_str to parse a JSON string into a validated BackendDescriptor. This method is used throughout src/local/* modules when reading from the SQLite store.
  • BackendDescriptor::to_json(&self) -> String: Uses serde_json::to_string to serialize the struct into a compact JSON string suitable for storage.

These methods guarantee that any descriptor stored in the backend_json column of the SQLite database can be recovered into an identical Rust struct with all backend-specific data intact.

Serde Configuration

The struct relies on standard serde attributes to ensure clean JSON output. The skip_serializing_if attribute prevents null fields from cluttering the JSON, while the fixed field list ensures that deserialization fails explicitly if required keys are missing. This strict schema provides a round-trip guarantee: a descriptor that successfully parses will always serialize back to an equivalent JSON representation, preventing data loss between storage and retrieval.

Usage in the OpenResearch Codebase

Persistence in SQLite

In src/compute.rs, the BackendDescriptor is serialized via to_json() and stored in the backend_json column of the local SQLite database. When a user queries the status of a run or attempts to reconnect to a remote job, the CLI retrieves this JSON string and reconstructs the descriptor using BackendDescriptor::parse(). This architecture decouples job submission from job management, allowing the CLI to operate statelessly while retaining all backend-specific context.

Backend-Specific Accessors

The descriptor implements typed accessor helpers that validate the kind field and return relevant tuples for each backend:

  • hf_ref() → (namespace, job_id)
  • k8s_ref() → (namespace, job_id)
  • modal_ref() → (namespace, job_id)
  • slurm_ref() → (namespace, job_id)
  • ray_ref() → (namespace, job_id)
  • ssh_ref() → (ssh_host, ssh_port, ssh_user)
  • local_ref() and openresearch_ref() for local execution contexts

These helpers are used in src/local/modal.rs, src/local/slurm.rs, and other backend modules to extract connection parameters without manual JSON parsing.

Code Example: Serializing a Modal Job Descriptor

The following Rust example demonstrates creating a BackendDescriptor for a Modal job, serializing it to JSON (as stored in SQLite), and parsing it back:

use orx::jobs::BackendDescriptor;

// Create a descriptor for a Modal job
let descriptor = BackendDescriptor {
    kind: "modal_job".to_string(),
    namespace: Some("my-app".to_string()),
    job_id: Some("sandbox-1234".to_string()),
    flavor: None,
    image: None,
    url: None,
    context: None,
    manifest: None,
    resources: None,
    ssh_host: None,
    ssh_port: None,
    ssh_user: None,
    timeout_secs: Some(3600),
    source_digest: Some("sha256:abcd...".to_string()),
    source_path: Some("/path/to/source".to_string()),
    source_size: Some(42_000_000),
};

// Serialize to JSON (what gets stored in the run row)
let json = descriptor.to_json();
println!("Stored JSON: {}", json);

// Later, read the JSON back into a descriptor
let parsed = BackendDescriptor::parse(&json).expect("Failed to parse descriptor");
assert_eq!(parsed.kind, "modal_job");
assert_eq!(parsed.job_id.unwrap(), "sandbox-1234");

Summary

  • The BackendDescriptor in src/jobs/mod.rs is the canonical representation of backend-specific job metadata in the orx CLI.
  • It uses serde with skip_serializing_if to produce compact JSON containing only populated fields.
  • The parse() and to_json() methods at lines 98–104 provide reliable serialization round-trips for SQLite storage.
  • Only the kind field is required; all other fields are optional and backend-specific.
  • Accessor methods like modal_ref() and slurm_ref() allow type-safe extraction of connection parameters in src/local/* modules.

Frequently Asked Questions

What is the difference between BackendDescriptor and the job manifest in orx?

The BackendDescriptor contains runtime metadata needed to reconnect to a job (such as job_id, namespace, and url), while the job manifest defines the execution payload (commands, environment variables, and resource requirements). The descriptor is stored as JSON in SQLite via src/compute.rs, whereas the manifest is typically handled separately during the submission phase.

How does orx handle missing optional fields during deserialization?

Because optional fields in BackendDescriptor use Option<T> with #[serde(skip_serializing_if = "Option::is_none")], missing keys in the JSON are automatically deserialized as None. This allows backward compatibility when new fields are added to the struct—older stored descriptors will simply lack the new fields without causing parse errors.

Where is the BackendDescriptor stored after a job is launched?

The JSON representation generated by to_json() is stored in the backend_json column of the SQLite database managed by src/compute.rs. This local storage persists the descriptor independently of the remote backend, enabling the CLI to reconnect to jobs even after process restarts.

Can I manually construct a BackendDescriptor for testing purposes?

Yes. As shown in the code example above, you can instantiate BackendDescriptor directly in Rust code, providing only the kind field and any relevant optional fields. You can then use to_json() to generate the JSON format expected by the CLI's storage layer, or parse() to read a stored JSON string back into a struct for testing recovery logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →