# How to Handle Large Outputs with Runtime Artifacts in OpenAI Plugins

> Learn how to handle large outputs in OpenAI plugins by writing to disk as runtime artifacts. Keep LLM responses lightweight and read data on demand.

- Repository: [OpenAI/plugins](https://github.com/openai/plugins)
- Tags: how-to-guide
- Published: 2026-09-13

---

**Skills should write large outputs to disk as runtime artifacts and return only compact metadata, allowing downstream phases to read the data on demand while keeping the LLM response stream lightweight.**

The openai/plugins repository demonstrates this pattern in the NVIDIA Omniverse USD Performance-Tuning skill, where multi-gigabyte files and heavy binary blobs are persisted locally rather than buffered in memory or passed through the LLM context window.

## The Runtime Artifact Pattern

When a skill generates very large outputs—such as multi-gigabyte USD stages, extensive log files, or heavyweight binary blobs—the plugin framework treats these as **runtime artifacts**. Instead of serializing the entire payload into the skill's return value, the skill writes the data to a local file and passes a lightweight descriptor to subsequent phases.

According to the [`SKILL.md`](https://github.com/openai/plugins/blob/main/SKILL.md) in the NVIDIA Omniverse USD Performance-Tuning skill, the design mandates that you "Invoke downstream skill bodies only when their phase is reached, and **keep raw runtime artifacts on disk while reading compact summaries**" [/plugins/nvidia/skills/omniverse-usd-performance-tuning/SKILL.md]. This prevents the LLM from processing massive strings and avoids hitting token limits.

## Implementation Guidelines

### Write Artifacts to Local Disk

The skill must persist large data to a designated output directory on the local filesystem. The [`invocation.md`](https://github.com/openai/plugins/blob/main/invocation.md) reference specifies that skills should "**Write optimized stages and runtime artifacts under the local output**" [/plugins/nvidia/skills/omniverse-usd-performance-tuning/references/so-run-operations/references/invocation.md]. This ensures that heavy data remains outside the LLM context window and can be streamed by downstream processes.

### Return Compact Summaries Only

Instead of returning file contents, return a JSON descriptor containing metadata such as file path, size, checksum, and a brief description. The [`runtime-artifact-token-budget.md`](https://github.com/openai/plugins/blob/main/runtime-artifact-token-budget.md) guide explicitly warns: "**Keep large runtime artifacts on disk**. Do not read, paste, or summarize full raw data" [/plugins/nvidia/skills/omniverse-usd-performance-tuning/references/runtime-artifact-token-budget.md]. This keeps the skill response within token budgets and reduces latency.

### Defer Reading to Downstream Phases

Downstream skill bodies should read artifacts only when their execution phase is reached. This phased execution model allows the pipeline to process large datasets asynchronously without blocking the LLM or consuming excessive memory. The artifact remains on disk until explicitly requested by a subsequent phase.

## Code Example: Phased Execution with Artifacts

The following Python pattern demonstrates writing a 2 GB artifact in the first phase and consuming it lazily in the second:

```python
import json
import os
import hashlib

def run_phase_one(input_args):
    """Generate a large artifact and write it to disk."""
    output_path = "/tmp/runtime_artifacts/large_output.bin"
    os.makedirs(os.path.dirname(output_path), exist_ok=True)

    # Simulate heavy computation writing a large file

    with open(output_path, "wb") as f:
        f.write(b"\0" * 2_000_000_000)  # 2 GB placeholder

    # Return lightweight descriptor

    return {
        "artifact_path": output_path,
        "size_bytes": os.path.getsize(output_path),
        "sha256": hashlib.sha256(open(output_path, "rb").read()).hexdigest(),
        "summary": "Binary payload generated by phase one."
    }

def run_phase_two(descriptor):
    """Consume the artifact only when needed."""
    path = descriptor["artifact_path"]
    
    # Stream file in 8 MiB chunks to avoid memory pressure

    with open(path, "rb") as f:
        while chunk := f.read(8 * 1024 * 1024):
            process_chunk(chunk)  # Replace with actual processing logic

```

This approach keeps the inter-phase communication lightweight (just the JSON descriptor) while allowing efficient streaming of large binary data when required.

## Key Source Files

The following files in the openai/plugins repository define the runtime artifact handling specifications:

- **[`/plugins/nvidia/skills/omniverse-usd-performance-tuning/SKILL.md`](https://github.com/openai/plugins/blob/main//plugins/nvidia/skills/omniverse-usd-performance-tuning/SKILL.md)** – Outlines the phased execution model and mandates keeping raw artifacts on disk while passing compact summaries between phases.

- **[`/plugins/nvidia/skills/omniverse-usd-performance-tuning/references/runtime-artifact-token-budget.md`](https://github.com/openai/plugins/blob/main//plugins/nvidia/skills/omniverse-usd-performance-tuning/references/runtime-artifact-token-budget.md)** – Defines the token-budget policy prohibiting the inclusion of full raw data in skill responses.

- **[`/plugins/nvidia/skills/omniverse-usd-performance-tuning/references/so-run-operations/references/invocation.md`](https://github.com/openai/plugins/blob/main//plugins/nvidia/skills/omniverse-usd-performance-tuning/references/so-run-operations/references/invocation.md)** – Specifies the local output directory structure for writing optimized stages and runtime artifacts.

## Summary

- **Persist large outputs to disk** rather than returning them directly in skill responses to avoid token limit violations and memory pressure.
- **Return compact metadata descriptors** containing file paths, sizes, and checksums, allowing downstream skills to locate and validate artifacts.
- **Defer artifact reading** until the specific execution phase that requires the data, enabling asynchronous processing and efficient resource utilization.
- **Follow the NVIDIA Omniverse USD Performance-Tuning skill pattern** documented in the openai/plugins repository for production implementations.

## Frequently Asked Questions

### What qualifies as a "large output" requiring runtime artifact handling?

Any output that risks exceeding LLM token limits or causing memory pressure qualifies as large. According to the runtime artifact token budget guidelines, this includes multi-gigabyte files, extensive logs, and heavyweight binary blobs that would consume excessive context window space if returned as string literals.

### Why can't large outputs be returned directly in the skill response?

Returning large outputs directly violates token budget constraints and can cause truncation or errors. The runtime artifact token budget documentation explicitly prohibits reading, pasting, or summarizing full raw data, as this would overwhelm the LLM context window and increase latency for downstream processing.

### Where should runtime artifacts be stored on the filesystem?

Artifacts should be written to the skill's designated local output directory, typically under `/tmp/runtime_artifacts/` or a path specified by the invocation reference. The [`invocation.md`](https://github.com/openai/plugins/blob/main/invocation.md) file specifies that optimized stages and runtime artifacts must be placed under the local output directory to ensure discoverability by downstream phases.

### How do downstream skills access data from runtime artifacts?

Downstream skills receive a compact JSON descriptor containing the `artifact_path`, `size_bytes`, and `sha256` checksum. They open the file path specified in the descriptor and stream the contents in manageable chunks (e.g., 8 MiB) only when their execution phase is reached, keeping memory usage constant regardless of file size.