How to Handle Large Outputs with Runtime Artifacts in OpenAI Plugins
Skills should write large outputs to disk as runtime artifacts and return only compact metadata, allowing downstream phases to read the data on demand while keeping the LLM response stream lightweight.
The openai/plugins repository demonstrates this pattern in the NVIDIA Omniverse USD Performance-Tuning skill, where multi-gigabyte files and heavy binary blobs are persisted locally rather than buffered in memory or passed through the LLM context window.
The Runtime Artifact Pattern
When a skill generates very large outputs—such as multi-gigabyte USD stages, extensive log files, or heavyweight binary blobs—the plugin framework treats these as runtime artifacts. Instead of serializing the entire payload into the skill's return value, the skill writes the data to a local file and passes a lightweight descriptor to subsequent phases.
According to the SKILL.md in the NVIDIA Omniverse USD Performance-Tuning skill, the design mandates that you "Invoke downstream skill bodies only when their phase is reached, and keep raw runtime artifacts on disk while reading compact summaries" [/plugins/nvidia/skills/omniverse-usd-performance-tuning/SKILL.md]. This prevents the LLM from processing massive strings and avoids hitting token limits.
Implementation Guidelines
Write Artifacts to Local Disk
The skill must persist large data to a designated output directory on the local filesystem. The invocation.md reference specifies that skills should "Write optimized stages and runtime artifacts under the local output" [/plugins/nvidia/skills/omniverse-usd-performance-tuning/references/so-run-operations/references/invocation.md]. This ensures that heavy data remains outside the LLM context window and can be streamed by downstream processes.
Return Compact Summaries Only
Instead of returning file contents, return a JSON descriptor containing metadata such as file path, size, checksum, and a brief description. The runtime-artifact-token-budget.md guide explicitly warns: "Keep large runtime artifacts on disk. Do not read, paste, or summarize full raw data" [/plugins/nvidia/skills/omniverse-usd-performance-tuning/references/runtime-artifact-token-budget.md]. This keeps the skill response within token budgets and reduces latency.
Defer Reading to Downstream Phases
Downstream skill bodies should read artifacts only when their execution phase is reached. This phased execution model allows the pipeline to process large datasets asynchronously without blocking the LLM or consuming excessive memory. The artifact remains on disk until explicitly requested by a subsequent phase.
Code Example: Phased Execution with Artifacts
The following Python pattern demonstrates writing a 2 GB artifact in the first phase and consuming it lazily in the second:
import json
import os
import hashlib
def run_phase_one(input_args):
"""Generate a large artifact and write it to disk."""
output_path = "/tmp/runtime_artifacts/large_output.bin"
os.makedirs(os.path.dirname(output_path), exist_ok=True)
# Simulate heavy computation writing a large file
with open(output_path, "wb") as f:
f.write(b"\0" * 2_000_000_000) # 2 GB placeholder
# Return lightweight descriptor
return {
"artifact_path": output_path,
"size_bytes": os.path.getsize(output_path),
"sha256": hashlib.sha256(open(output_path, "rb").read()).hexdigest(),
"summary": "Binary payload generated by phase one."
}
def run_phase_two(descriptor):
"""Consume the artifact only when needed."""
path = descriptor["artifact_path"]
# Stream file in 8 MiB chunks to avoid memory pressure
with open(path, "rb") as f:
while chunk := f.read(8 * 1024 * 1024):
process_chunk(chunk) # Replace with actual processing logic
This approach keeps the inter-phase communication lightweight (just the JSON descriptor) while allowing efficient streaming of large binary data when required.
Key Source Files
The following files in the openai/plugins repository define the runtime artifact handling specifications:
-
/plugins/nvidia/skills/omniverse-usd-performance-tuning/SKILL.md– Outlines the phased execution model and mandates keeping raw artifacts on disk while passing compact summaries between phases. -
/plugins/nvidia/skills/omniverse-usd-performance-tuning/references/runtime-artifact-token-budget.md– Defines the token-budget policy prohibiting the inclusion of full raw data in skill responses. -
/plugins/nvidia/skills/omniverse-usd-performance-tuning/references/so-run-operations/references/invocation.md– Specifies the local output directory structure for writing optimized stages and runtime artifacts.
Summary
- Persist large outputs to disk rather than returning them directly in skill responses to avoid token limit violations and memory pressure.
- Return compact metadata descriptors containing file paths, sizes, and checksums, allowing downstream skills to locate and validate artifacts.
- Defer artifact reading until the specific execution phase that requires the data, enabling asynchronous processing and efficient resource utilization.
- Follow the NVIDIA Omniverse USD Performance-Tuning skill pattern documented in the openai/plugins repository for production implementations.
Frequently Asked Questions
What qualifies as a "large output" requiring runtime artifact handling?
Any output that risks exceeding LLM token limits or causing memory pressure qualifies as large. According to the runtime artifact token budget guidelines, this includes multi-gigabyte files, extensive logs, and heavyweight binary blobs that would consume excessive context window space if returned as string literals.
Why can't large outputs be returned directly in the skill response?
Returning large outputs directly violates token budget constraints and can cause truncation or errors. The runtime artifact token budget documentation explicitly prohibits reading, pasting, or summarizing full raw data, as this would overwhelm the LLM context window and increase latency for downstream processing.
Where should runtime artifacts be stored on the filesystem?
Artifacts should be written to the skill's designated local output directory, typically under /tmp/runtime_artifacts/ or a path specified by the invocation reference. The invocation.md file specifies that optimized stages and runtime artifacts must be placed under the local output directory to ensure discoverability by downstream phases.
How do downstream skills access data from runtime artifacts?
Downstream skills receive a compact JSON descriptor containing the artifact_path, size_bytes, and sha256 checksum. They open the file path specified in the descriptor and stream the contents in manageable chunks (e.g., 8 MiB) only when their execution phase is reached, keeping memory usage constant regardless of file size.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →