What Is the Purpose of the `cans` Directory in Soup CLI: A Complete Guide to Soup Cans
The cans directory in Soup CLI implements Soup Cans—a shareable .can artifact format that captures, transports, and replays complete training iterations as portable, reproducible packages.
The cans directory is the core engine behind Soup Cans, the artifact system that transforms individual training runs into self-contained, versionable units. According to the MakazhanAlpamys/Soup source code, this mechanism enables users to pack an entire fine-tuning iteration—including configuration, data references, and metadata—into a single tar.gz archive that can be verified, shared, and executed elsewhere.
What Is a .can File?
A .can file is a compressed archive with a strict internal structure. When you create or inspect a can, you're working with four standardized entries:
| Entry | Purpose |
|---|---|
manifest.yaml |
Format version, name, author, and metadata—governed by the Manifest model in src/soup_cli/cans/schema.py |
config.yaml |
The complete SoupConfig used for the training run |
data_ref.yaml |
Data provenance: hash plus URL or Hugging Face dataset ID |
recipe.md |
Optional human-readable documentation of the run |
This structure ensures that any .can can be unpacked and executed with full reproducibility, regardless of where it originated.
Key Files in the cans Directory
Understanding the cans directory requires examining its implementation files. Each module handles a specific lifecycle phase of a can:
schema.py: Data Models
The Manifest and DataRef Pydantic models in src/soup_cli/cans/schema.py enforce schema validation for all can contents. These models guarantee that every .can follows a consistent, forward-compatible format.
pack.py: Creating Cans
The pack_entry() function in src/soup_cli/cans/pack.py assembles the archive. It serializes the manifest, config, and data reference into the standardized structure. For derived works, fork_can() creates lineage-aware copies.
unpack.py: Inspection and Extraction
src/soup_cli/cans/unpack.py provides inspect_can() for reading can metadata without full extraction, plus utilities for unpacking to disk.
run.py: Execution
The run_can() function in src/soup_cli/cans/run.py orchestrates replaying a can—fetching data, loading config, and executing the training iteration.
publish.py and verify.py: Distribution and Trust
publish_can() in src/soup_cli/cans/publish.py handles remote repository uploads, while src/soup_cli/cans/verify.py performs attestation verification to cryptographically validate can integrity.
How to Create a .can from the CLI
The simplest way to generate cans is through the loop watch command with automatic packing enabled:
# Run training loop and create a .can for each successful iteration
soup loop watch --pack-cans
This flag triggers can creation immediately after each completed training step, ensuring no successful run goes unarchived.
How to Work with Cans Programmatically
For custom pipelines, the cans module exports a complete Python API.
Creating a Can Manually
from soup_cli.cans.pack import pack_entry
from soup_cli.cans.schema import Manifest, DataRef
manifest = Manifest(name="my-model", author="alice")
data_ref = DataRef(url="hf://my/dataset", sha256="abc123")
pack_entry(
manifest=manifest,
data_ref=data_ref,
config=my_soup_config,
out_path="my_model.can",
)
Inspecting a Can Without Extraction
from soup_cli.cans.unpack import inspect_can
info = inspect_can("my_model.can")
print(info.manifest)
print(info.config) # Full SoupConfig from the original run
Publishing to a Remote Repository
from soup_cli.cans.publish import publish_can
publish_can(
can_path="my_model.can",
repo_id="myorg/my-models",
token="YOUR_HF_TOKEN", # Falls back to environment variable if omitted
)
Why the cans Directory Matters for Reproducibility
The cans directory transforms Soup from a training tool into a reproducible research platform. By capturing not just weights but the complete execution context—data provenance, configuration, and documentation—Soup Cans solve three critical problems in machine learning engineering:
- Provenance tracking: The
data_ref.yamlSHA-256 hash guarantees the exact dataset version used - Configuration auditing:
config.yamleliminates "works on my machine" discrepancies - Collaborative iteration: Published cans become forkable, versioned training recipes
Summary
- The
cansdirectory implements the Soup Can artifact system for portable, reproducible training runs - A
.canfile is atar.gzarchive containingmanifest.yaml,config.yaml,data_ref.yaml, and optionalrecipe.md - Key modules in
src/soup_cli/cans/handle packing (pack.py), inspection (unpack.py), execution (run.py), publishing (publish.py), and verification (verify.py) - CLI users generate cans automatically via
--pack-cansor programmatically throughpack_entry() - The schema-defined structure in
schema.pyensures forward compatibility and validation
Frequently Asked Questions
What file format does a Soup Can use?
A Soup Can uses the .can extension and is internally a tar.gz archive. The compression preserves space while the tar structure maintains readable, extractable contents without special tooling.
How does Soup verify data integrity in a can?
The DataRef model in schema.py requires a SHA-256 hash of the training data. When verify.py processes a can, it recomputes this hash and validates against the stored value, detecting any data corruption or substitution.
Can I execute a can without installing the full Soup CLI?
No—the run_can() function in run.py requires the Soup runtime environment to interpret the SoupConfig and orchestrate training. However, you can inspect any can's metadata using inspect_can() with minimal dependencies.
What's the difference between pack_entry() and fork_can()?
pack_entry() in pack.py creates a fresh can from constituent components (manifest, config, data reference). fork_can() copies an existing can while updating its lineage metadata—useful for creating derivative works while preserving provenance history.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →