What Is the Purpose of the `cans` Directory in Soup CLI: A Complete Guide to Soup Cans

The cans directory in Soup CLI implements Soup Cans—a shareable .can artifact format that captures, transports, and replays complete training iterations as portable, reproducible packages.

The cans directory is the core engine behind Soup Cans, the artifact system that transforms individual training runs into self-contained, versionable units. According to the MakazhanAlpamys/Soup source code, this mechanism enables users to pack an entire fine-tuning iteration—including configuration, data references, and metadata—into a single tar.gz archive that can be verified, shared, and executed elsewhere.

What Is a .can File?

A .can file is a compressed archive with a strict internal structure. When you create or inspect a can, you're working with four standardized entries:

Entry Purpose
manifest.yaml Format version, name, author, and metadata—governed by the Manifest model in src/soup_cli/cans/schema.py
config.yaml The complete SoupConfig used for the training run
data_ref.yaml Data provenance: hash plus URL or Hugging Face dataset ID
recipe.md Optional human-readable documentation of the run

This structure ensures that any .can can be unpacked and executed with full reproducibility, regardless of where it originated.

Key Files in the cans Directory

Understanding the cans directory requires examining its implementation files. Each module handles a specific lifecycle phase of a can:

schema.py: Data Models

The Manifest and DataRef Pydantic models in src/soup_cli/cans/schema.py enforce schema validation for all can contents. These models guarantee that every .can follows a consistent, forward-compatible format.

pack.py: Creating Cans

The pack_entry() function in src/soup_cli/cans/pack.py assembles the archive. It serializes the manifest, config, and data reference into the standardized structure. For derived works, fork_can() creates lineage-aware copies.

unpack.py: Inspection and Extraction

src/soup_cli/cans/unpack.py provides inspect_can() for reading can metadata without full extraction, plus utilities for unpacking to disk.

run.py: Execution

The run_can() function in src/soup_cli/cans/run.py orchestrates replaying a can—fetching data, loading config, and executing the training iteration.

publish.py and verify.py: Distribution and Trust

publish_can() in src/soup_cli/cans/publish.py handles remote repository uploads, while src/soup_cli/cans/verify.py performs attestation verification to cryptographically validate can integrity.

How to Create a .can from the CLI

The simplest way to generate cans is through the loop watch command with automatic packing enabled:


# Run training loop and create a .can for each successful iteration

soup loop watch --pack-cans

This flag triggers can creation immediately after each completed training step, ensuring no successful run goes unarchived.

How to Work with Cans Programmatically

For custom pipelines, the cans module exports a complete Python API.

Creating a Can Manually

from soup_cli.cans.pack import pack_entry
from soup_cli.cans.schema import Manifest, DataRef

manifest = Manifest(name="my-model", author="alice")
data_ref = DataRef(url="hf://my/dataset", sha256="abc123")

pack_entry(
    manifest=manifest,
    data_ref=data_ref,
    config=my_soup_config,
    out_path="my_model.can",
)

Inspecting a Can Without Extraction

from soup_cli.cans.unpack import inspect_can

info = inspect_can("my_model.can")
print(info.manifest)
print(info.config)  # Full SoupConfig from the original run

Publishing to a Remote Repository

from soup_cli.cans.publish import publish_can

publish_can(
    can_path="my_model.can",
    repo_id="myorg/my-models",
    token="YOUR_HF_TOKEN",  # Falls back to environment variable if omitted

)

Why the cans Directory Matters for Reproducibility

The cans directory transforms Soup from a training tool into a reproducible research platform. By capturing not just weights but the complete execution context—data provenance, configuration, and documentation—Soup Cans solve three critical problems in machine learning engineering:

  • Provenance tracking: The data_ref.yaml SHA-256 hash guarantees the exact dataset version used
  • Configuration auditing: config.yaml eliminates "works on my machine" discrepancies
  • Collaborative iteration: Published cans become forkable, versioned training recipes

Summary

  • The cans directory implements the Soup Can artifact system for portable, reproducible training runs
  • A .can file is a tar.gz archive containing manifest.yaml, config.yaml, data_ref.yaml, and optional recipe.md
  • Key modules in src/soup_cli/cans/ handle packing (pack.py), inspection (unpack.py), execution (run.py), publishing (publish.py), and verification (verify.py)
  • CLI users generate cans automatically via --pack-cans or programmatically through pack_entry()
  • The schema-defined structure in schema.py ensures forward compatibility and validation

Frequently Asked Questions

What file format does a Soup Can use?

A Soup Can uses the .can extension and is internally a tar.gz archive. The compression preserves space while the tar structure maintains readable, extractable contents without special tooling.

How does Soup verify data integrity in a can?

The DataRef model in schema.py requires a SHA-256 hash of the training data. When verify.py processes a can, it recomputes this hash and validates against the stored value, detecting any data corruption or substitution.

Can I execute a can without installing the full Soup CLI?

No—the run_can() function in run.py requires the Soup runtime environment to interpret the SoupConfig and orchestrate training. However, you can inspect any can's metadata using inspect_can() with minimal dependencies.

What's the difference between pack_entry() and fork_can()?

pack_entry() in pack.py creates a fresh can from constituent components (manifest, config, data reference). fork_can() copies an existing can while updating its lineage metadata—useful for creating derivative works while preserving provenance history.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →