# What Is the BackendDescriptor in orx and How Is It Serialized?

> Understand the BackendDescriptor in orx, a data structure for backend job metadata. Learn how it serializes to JSON using serde for SQLite integration in this technical guide.

- Repository: [alphaXiv/OpenResearch](https://github.com/alphaXiv/OpenResearch)
- Tags: internals
- Published: 2026-09-13

---

**The `BackendDescriptor` is a serde-powered data structure defined in [`src/jobs/mod.rs`](https://github.com/alphaXiv/OpenResearch/blob/main/src/jobs/mod.rs) that captures backend-specific job metadata and provides `parse()` and `to_json()` methods for JSON serialization to SQLite.**

The `orx` CLI (OpenResearch) uses the `BackendDescriptor` as the single source of truth for recording how a run was launched on a particular compute backend. It is created when a job is submitted and later deserialized by local launch modules to reconnect, monitor, or clean up remote executions.

## Core Purpose and Structure in src/jobs/mod.rs

### What BackendDescriptor Stores

According to the alphaXiv/OpenResearch source code, the `BackendDescriptor` struct (defined starting at line 46 in [`src/jobs/mod.rs`](https://github.com/alphaXiv/OpenResearch/blob/main/src/jobs/mod.rs)) serves as the central data structure for persisting backend-specific metadata. Its primary role is to store the information required to recreate the execution context for any supported backend—whether Modal, Slurm, Kubernetes, SSH, or others—after the initial launch process has exited.

The struct uses **serde** derives for `Serialize` and `Deserialize`, ensuring that any instance can be converted to a JSON string for storage and reliably reconstructed later.

### Required vs Optional Fields

The descriptor uses a fixed schema where only the `kind` field is mandatory. All other fields are `Option<T>` and are omitted from the JSON output when `None`:

- **`kind`** (required): Identifies the backend type using string identifiers such as `"modal_job"`, `"slurm_job"`, `"ssh_job"`, `"ray_job"`, `"k8s_job"`, `"hf_job"`, `"openresearch_job"`, `"local_job"`, or `"tinker_job"`.
- **Optional fields**: `namespace`, `job_id`, `flavor`, `image`, `url`, `context`, `manifest`, `resources`, `ssh_host`, `ssh_port`, `ssh_user`, `timeout_secs`, `source_digest`, `source_path`, and `source_size`.

The implementation applies `#[serde(skip_serializing_if = "Option::is_none")]` to optional fields, keeping the serialized JSON compact and backend-agnostic.

## JSON Serialization Implementation

### The parse() and to_json() Methods

The `BackendDescriptor` implementation provides two convenience methods for JSON round-trips, located at lines 98–104 in [`src/jobs/mod.rs`](https://github.com/alphaXiv/OpenResearch/blob/main/src/jobs/mod.rs):

- **`BackendDescriptor::parse(json: &str) -> Result<Self>`**: Uses `serde_json::from_str` to parse a JSON string into a validated `BackendDescriptor`. This method is used throughout `src/local/*` modules when reading from the SQLite store.
- **`BackendDescriptor::to_json(&self) -> String`**: Uses `serde_json::to_string` to serialize the struct into a compact JSON string suitable for storage.

These methods guarantee that any descriptor stored in the `backend_json` column of the SQLite database can be recovered into an identical Rust struct with all backend-specific data intact.

### Serde Configuration

The struct relies on standard serde attributes to ensure clean JSON output. The `skip_serializing_if` attribute prevents null fields from cluttering the JSON, while the fixed field list ensures that deserialization fails explicitly if required keys are missing. This strict schema provides a **round-trip guarantee**: a descriptor that successfully parses will always serialize back to an equivalent JSON representation, preventing data loss between storage and retrieval.

## Usage in the OpenResearch Codebase

### Persistence in SQLite

In [`src/compute.rs`](https://github.com/alphaXiv/OpenResearch/blob/main/src/compute.rs), the `BackendDescriptor` is serialized via `to_json()` and stored in the `backend_json` column of the local SQLite database. When a user queries the status of a run or attempts to reconnect to a remote job, the CLI retrieves this JSON string and reconstructs the descriptor using `BackendDescriptor::parse()`. This architecture decouples job submission from job management, allowing the CLI to operate statelessly while retaining all backend-specific context.

### Backend-Specific Accessors

The descriptor implements typed accessor helpers that validate the `kind` field and return relevant tuples for each backend:

- `hf_ref()` → (namespace, job_id)
- `k8s_ref()` → (namespace, job_id)
- `modal_ref()` → (namespace, job_id)
- `slurm_ref()` → (namespace, job_id)
- `ray_ref()` → (namespace, job_id)
- `ssh_ref()` → (ssh_host, ssh_port, ssh_user)
- `local_ref()` and `openresearch_ref()` for local execution contexts

These helpers are used in [`src/local/modal.rs`](https://github.com/alphaXiv/OpenResearch/blob/main/src/local/modal.rs), [`src/local/slurm.rs`](https://github.com/alphaXiv/OpenResearch/blob/main/src/local/slurm.rs), and other backend modules to extract connection parameters without manual JSON parsing.

## Code Example: Serializing a Modal Job Descriptor

The following Rust example demonstrates creating a `BackendDescriptor` for a Modal job, serializing it to JSON (as stored in SQLite), and parsing it back:

```rust
use orx::jobs::BackendDescriptor;

// Create a descriptor for a Modal job
let descriptor = BackendDescriptor {
    kind: "modal_job".to_string(),
    namespace: Some("my-app".to_string()),
    job_id: Some("sandbox-1234".to_string()),
    flavor: None,
    image: None,
    url: None,
    context: None,
    manifest: None,
    resources: None,
    ssh_host: None,
    ssh_port: None,
    ssh_user: None,
    timeout_secs: Some(3600),
    source_digest: Some("sha256:abcd...".to_string()),
    source_path: Some("/path/to/source".to_string()),
    source_size: Some(42_000_000),
};

// Serialize to JSON (what gets stored in the run row)
let json = descriptor.to_json();
println!("Stored JSON: {}", json);

// Later, read the JSON back into a descriptor
let parsed = BackendDescriptor::parse(&json).expect("Failed to parse descriptor");
assert_eq!(parsed.kind, "modal_job");
assert_eq!(parsed.job_id.unwrap(), "sandbox-1234");

```

## Summary

- The `BackendDescriptor` in [`src/jobs/mod.rs`](https://github.com/alphaXiv/OpenResearch/blob/main/src/jobs/mod.rs) is the canonical representation of backend-specific job metadata in the `orx` CLI.
- It uses serde with `skip_serializing_if` to produce compact JSON containing only populated fields.
- The `parse()` and `to_json()` methods at lines 98–104 provide reliable serialization round-trips for SQLite storage.
- Only the `kind` field is required; all other fields are optional and backend-specific.
- Accessor methods like `modal_ref()` and `slurm_ref()` allow type-safe extraction of connection parameters in `src/local/*` modules.

## Frequently Asked Questions

### What is the difference between BackendDescriptor and the job manifest in orx?

The `BackendDescriptor` contains runtime metadata needed to reconnect to a job (such as `job_id`, `namespace`, and `url`), while the job manifest defines the execution payload (commands, environment variables, and resource requirements). The descriptor is stored as JSON in SQLite via [`src/compute.rs`](https://github.com/alphaXiv/OpenResearch/blob/main/src/compute.rs), whereas the manifest is typically handled separately during the submission phase.

### How does orx handle missing optional fields during deserialization?

Because optional fields in `BackendDescriptor` use `Option<T>` with `#[serde(skip_serializing_if = "Option::is_none")]`, missing keys in the JSON are automatically deserialized as `None`. This allows backward compatibility when new fields are added to the struct—older stored descriptors will simply lack the new fields without causing parse errors.

### Where is the BackendDescriptor stored after a job is launched?

The JSON representation generated by `to_json()` is stored in the `backend_json` column of the SQLite database managed by [`src/compute.rs`](https://github.com/alphaXiv/OpenResearch/blob/main/src/compute.rs). This local storage persists the descriptor independently of the remote backend, enabling the CLI to reconnect to jobs even after process restarts.

### Can I manually construct a BackendDescriptor for testing purposes?

Yes. As shown in the code example above, you can instantiate `BackendDescriptor` directly in Rust code, providing only the `kind` field and any relevant optional fields. You can then use `to_json()` to generate the JSON format expected by the CLI's storage layer, or `parse()` to read a stored JSON string back into a struct for testing recovery logic.