What Is the BackendDescriptor in orx and How Is It Serialized?
The BackendDescriptor is a serde-powered data structure defined in src/jobs/mod.rs that captures backend-specific job metadata and provides parse() and to_json() methods for JSON serialization to SQLite.
The orx CLI (OpenResearch) uses the BackendDescriptor as the single source of truth for recording how a run was launched on a particular compute backend. It is created when a job is submitted and later deserialized by local launch modules to reconnect, monitor, or clean up remote executions.
Core Purpose and Structure in src/jobs/mod.rs
What BackendDescriptor Stores
According to the alphaXiv/OpenResearch source code, the BackendDescriptor struct (defined starting at line 46 in src/jobs/mod.rs) serves as the central data structure for persisting backend-specific metadata. Its primary role is to store the information required to recreate the execution context for any supported backend—whether Modal, Slurm, Kubernetes, SSH, or others—after the initial launch process has exited.
The struct uses serde derives for Serialize and Deserialize, ensuring that any instance can be converted to a JSON string for storage and reliably reconstructed later.
Required vs Optional Fields
The descriptor uses a fixed schema where only the kind field is mandatory. All other fields are Option<T> and are omitted from the JSON output when None:
kind(required): Identifies the backend type using string identifiers such as"modal_job","slurm_job","ssh_job","ray_job","k8s_job","hf_job","openresearch_job","local_job", or"tinker_job".- Optional fields:
namespace,job_id,flavor,image,url,context,manifest,resources,ssh_host,ssh_port,ssh_user,timeout_secs,source_digest,source_path, andsource_size.
The implementation applies #[serde(skip_serializing_if = "Option::is_none")] to optional fields, keeping the serialized JSON compact and backend-agnostic.
JSON Serialization Implementation
The parse() and to_json() Methods
The BackendDescriptor implementation provides two convenience methods for JSON round-trips, located at lines 98–104 in src/jobs/mod.rs:
BackendDescriptor::parse(json: &str) -> Result<Self>: Usesserde_json::from_strto parse a JSON string into a validatedBackendDescriptor. This method is used throughoutsrc/local/*modules when reading from the SQLite store.BackendDescriptor::to_json(&self) -> String: Usesserde_json::to_stringto serialize the struct into a compact JSON string suitable for storage.
These methods guarantee that any descriptor stored in the backend_json column of the SQLite database can be recovered into an identical Rust struct with all backend-specific data intact.
Serde Configuration
The struct relies on standard serde attributes to ensure clean JSON output. The skip_serializing_if attribute prevents null fields from cluttering the JSON, while the fixed field list ensures that deserialization fails explicitly if required keys are missing. This strict schema provides a round-trip guarantee: a descriptor that successfully parses will always serialize back to an equivalent JSON representation, preventing data loss between storage and retrieval.
Usage in the OpenResearch Codebase
Persistence in SQLite
In src/compute.rs, the BackendDescriptor is serialized via to_json() and stored in the backend_json column of the local SQLite database. When a user queries the status of a run or attempts to reconnect to a remote job, the CLI retrieves this JSON string and reconstructs the descriptor using BackendDescriptor::parse(). This architecture decouples job submission from job management, allowing the CLI to operate statelessly while retaining all backend-specific context.
Backend-Specific Accessors
The descriptor implements typed accessor helpers that validate the kind field and return relevant tuples for each backend:
hf_ref()→ (namespace, job_id)k8s_ref()→ (namespace, job_id)modal_ref()→ (namespace, job_id)slurm_ref()→ (namespace, job_id)ray_ref()→ (namespace, job_id)ssh_ref()→ (ssh_host, ssh_port, ssh_user)local_ref()andopenresearch_ref()for local execution contexts
These helpers are used in src/local/modal.rs, src/local/slurm.rs, and other backend modules to extract connection parameters without manual JSON parsing.
Code Example: Serializing a Modal Job Descriptor
The following Rust example demonstrates creating a BackendDescriptor for a Modal job, serializing it to JSON (as stored in SQLite), and parsing it back:
use orx::jobs::BackendDescriptor;
// Create a descriptor for a Modal job
let descriptor = BackendDescriptor {
kind: "modal_job".to_string(),
namespace: Some("my-app".to_string()),
job_id: Some("sandbox-1234".to_string()),
flavor: None,
image: None,
url: None,
context: None,
manifest: None,
resources: None,
ssh_host: None,
ssh_port: None,
ssh_user: None,
timeout_secs: Some(3600),
source_digest: Some("sha256:abcd...".to_string()),
source_path: Some("/path/to/source".to_string()),
source_size: Some(42_000_000),
};
// Serialize to JSON (what gets stored in the run row)
let json = descriptor.to_json();
println!("Stored JSON: {}", json);
// Later, read the JSON back into a descriptor
let parsed = BackendDescriptor::parse(&json).expect("Failed to parse descriptor");
assert_eq!(parsed.kind, "modal_job");
assert_eq!(parsed.job_id.unwrap(), "sandbox-1234");
Summary
- The
BackendDescriptorinsrc/jobs/mod.rsis the canonical representation of backend-specific job metadata in theorxCLI. - It uses serde with
skip_serializing_ifto produce compact JSON containing only populated fields. - The
parse()andto_json()methods at lines 98–104 provide reliable serialization round-trips for SQLite storage. - Only the
kindfield is required; all other fields are optional and backend-specific. - Accessor methods like
modal_ref()andslurm_ref()allow type-safe extraction of connection parameters insrc/local/*modules.
Frequently Asked Questions
What is the difference between BackendDescriptor and the job manifest in orx?
The BackendDescriptor contains runtime metadata needed to reconnect to a job (such as job_id, namespace, and url), while the job manifest defines the execution payload (commands, environment variables, and resource requirements). The descriptor is stored as JSON in SQLite via src/compute.rs, whereas the manifest is typically handled separately during the submission phase.
How does orx handle missing optional fields during deserialization?
Because optional fields in BackendDescriptor use Option<T> with #[serde(skip_serializing_if = "Option::is_none")], missing keys in the JSON are automatically deserialized as None. This allows backward compatibility when new fields are added to the struct—older stored descriptors will simply lack the new fields without causing parse errors.
Where is the BackendDescriptor stored after a job is launched?
The JSON representation generated by to_json() is stored in the backend_json column of the SQLite database managed by src/compute.rs. This local storage persists the descriptor independently of the remote backend, enabling the CLI to reconnect to jobs even after process restarts.
Can I manually construct a BackendDescriptor for testing purposes?
Yes. As shown in the code example above, you can instantiate BackendDescriptor directly in Rust code, providing only the kind field and any relevant optional fields. You can then use to_json() to generate the JSON format expected by the CLI's storage layer, or parse() to read a stored JSON string back into a struct for testing recovery logic.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →