# How Layer Streaming Enables Fine‑Tuning an 8B Model on a 4 GB GPU: Soup's Memory‑Efficient Training Architecture

> Discover how layer streaming slashes GPU memory needs from 30 GB to 4 GB, enabling 8B model fine-tuning on low-spec hardware. Learn about Soup's efficient training.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: deep-dive
- Published: 2026-08-16

---

**Layer streaming reduces GPU memory usage from ~30 GB to under 4 GB by loading only the active layer's weights during each forward‑backward pass, keeping remaining layers in system RAM or compressed on disk.**

The **Soup** framework implements layer streaming as a transparent, drop‑in replacement for standard model training. Instead of holding an entire 8‑billion‑parameter model in VRAM, the runtime orchestrates a continuous flow of weight chunks between storage tiers—enabling fine‑tuning on consumer‑grade GPUs that would otherwise be impossible.

## What Layer Streaming Actually Does

Traditional fine‑tuning loads all model parameters into GPU memory once. For an 8B model in FP16, this requires approximately 16 GB for weights alone, plus activation memory that pushes requirements past 30 GB. Layer streaming inverts this assumption.

According to the Soup source code, the technique rests on a single principle: **GPU memory need only contain the parameters currently being computed**. The framework implements this through a streaming‑aware model wrapper that intercepts every layer call, a tiered storage system, and quantized weight serialization.

## Core Components of the Streaming Pipeline

### Streaming‑Aware Model Wrapper

The wrapper patches each layer's forward pass in [`src/soup_cli/trainer/stream_setup.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/stream_setup.py). When a layer executes:

1. The wrapper requests the weight chunk from the streaming runtime
2. The runtime loads from `RamSource` or `DiskSource` onto GPU
3. Computation proceeds normally
4. GPU memory for that chunk is immediately released

This pattern repeats for every layer in both forward and backward passes. The wrapper ensures training code requires no modifications—`stream_layers=True` toggles the behavior transparently.

### Tier Selection and Stream‑Fit Analysis

The runtime decides storage placement through functions defined in [`src/soup_cli/utils/layer_stream.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/layer_stream.py):

- `estimate_stream_peak_vram()` – calculates per‑layer memory footprint including activations
- `decide_stream_fit()` – determines if a candidate streaming plan respects GPU limits
- `choose_tier()` – assigns each layer to GPU, RAM, or disk based on access patterns and bandwidth

The **stream‑fit** analysis considers layer depth (earlier layers compute more frequently during back‑propagation), tensor sizes, and measured disk/RAM throughput to minimize stalling.

### NF4 Quantized Storage

Weights stored on disk use the **NF4** (Normal Float 4) format, reducing storage to roughly ¼ of FP16 size with near‑lossless reconstruction. The quantization preserves weight distributions better than naive 4‑bit schemes, critical for maintaining fine‑tuning quality when layers repeatedly stream from cold storage.

### Activation Off‑Loading

Intermediate activations that would persist for backward passes are optionally moved to RAM or disk. The `build_stream_plan()` function in [`layer_stream.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/layer_stream.py) precomputes which tensors require preservation and schedules their evacuation before they compete with incoming weight chunks.

### Throughput Forecasting

Before training begins, `forecast_stream_throughput()` simulates the I/O cost of the proposed streaming plan against the system's measured bandwidth. If projected throughput falls below a threshold, the planner adjusts tier assignments—typically promoting frequently accessed layers to RAM cache.

## Runtime Execution Flow

When `StreamingSetupMixin` initializes (in [`src/soup_cli/trainer/stream_setup.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/stream_setup.py)), it:

1. Registers the streaming runtime with the trainer
2. Invokes `build_stream_plan()` to generate a layer‑to‑tier mapping
3. Patches the model's forward method with the streaming wrapper

During training, [`src/soup_cli/utils/layer_stream_runtime.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/layer_stream_runtime.py) manages the actual data movement:

```python

# Conceptual flow — actual implementation handles async prefetching

class LayerStreamRuntime:
    def load_chunk(self, layer_id: int) -> Tensor:
        source = self._plan[layer_id]  # RamSource or DiskSource

        chunk = source.read(layer_id)   # Decompress if from disk

        return chunk.cuda()             # Move to GPU

        
    def release_chunk(self, layer_id: int):
        self._gpu_allocator.free(layer_id)
        self._stats.track_release(layer_id)

```

The runtime maintains statistics accessible via `trainer._stream_runtime.stats()`, reporting actual VRAM consumption, cache hit rates, and achieved throughput.

## Enabling Layer Streaming in Practice

### Command‑Line Interface

```bash

# Fine-tune Llama 3 8B on a 4 GB GPU

soup finetune \
  --model meta-llama/Meta-Llama-3-8B-Instruct \
  --dataset instruction_data.jsonl \
  --learning-rate 5e-5 \
  --epochs 3 \
  --stream-layers

```

### Python Configuration

```python
from soup_cli.config.schema import Config
from soup_cli.trainer.pretrain import Trainer

cfg = Config(
    model_name="meta-llama/Meta-Llama-3-8B-Instruct",
    dataset="instruction_data.jsonl",
    epochs=3,
    learning_rate=5e-5,
    training=dict(stream_layers=True),  # Enable layer streaming

)

trainer = Trainer(cfg)
trainer.run()

```

### Verifying Memory Savings

```python

# After training completes

runtime = trainer._stream_runtime
print(runtime.stats())

# Example output:

# {

#   "total_vram_used": "3.8 GB",

#   "layers_streamed": 28,

#   "disk_reads": 28,

#   "ram_cache_hits": 12,

#   "throughput_gb_per_s": 1.4

# }

```

The `total_vram_used` field confirms the streaming plan's effectiveness—well below the 4 GB threshold for compatible GPUs.

## Key Implementation Files

| Path | Responsibility |
|------|---------------|
| [`src/soup_cli/utils/layer_stream.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/layer_stream.py) | VRAM estimation, tier selection (`decide_stream_fit`, `choose_tier`), plan construction (`build_stream_plan`) |
| [`src/soup_cli/utils/layer_stream_runtime.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/layer_stream_runtime.py) | Chunk loading/unloading, `RamSource`/`DiskSource` abstractions, statistics collection |
| [`src/soup_cli/trainer/stream_setup.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/trainer/stream_setup.py) | `StreamingSetupMixin` integrating runtime with trainer pipeline |
| [`tests/test_v07203.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_v07203.py) | Unit tests for VRAM estimation and throughput forecasting |
| [`tests/test_v07204.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_v07204.py) | Integration tests validating streaming with DPO/ORPO/SimPO/KTO preference learning |

## Performance Characteristics

Layer streaming trades compute efficiency for memory capacity. The overhead consists of:

- **I/O latency**: Disk→RAM→GPU transfer per layer access
- **Decompression**: NF4→FP16 conversion on load
- **Synchronization**: Ensuring weight arrival before computation

The Soup runtime mitigates this through:

- **Prefetching**: Loading layer *n+1* while computing layer *n*
- **RAM caching**: Hot‑start layers (embeddings, early transformers) resident in system memory
- **Async copies**: Overlapping GPU computation with host‑side I/O

For typical fine‑tuning workloads, throughput remains within 30‑50% of non‑streaming baselines—acceptable given the alternative of training impossibility.

## Summary

- **Layer streaming** in Soup enables 8B model fine‑tuning on 4 GB GPUs by holding only active layer weights in VRAM
- The `StreamingSetupMixin` patches models transparently—no training code changes required
- **Stream‑fit analysis** (`decide_stream_fit`, `choose_tier`) optimizes layer placement across GPU/RAM/disk tiers
- **NF4 quantization** reduces disk storage by 4× with minimal quality degradation
- Runtime statistics via `trainer._stream_runtime.stats()` verify memory constraints are met

## Frequently Asked Questions

### How much slower is layer streaming compared to full GPU training?

Throughput typically degrades 30‑50% depending on storage bandwidth. NVMe SSDs with 3+ GB/s sequential read achieve better results than SATA drives. The `forecast_stream_throughput()` function in [`layer_stream.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/layer_stream.py) predicts this before training begins.

### Does layer streaming work with LoRA or QLoRA?

Yes. The streaming wrapper operates below adapter layers—LoRA parameters remain in GPU memory while base model weights stream. This combination enables larger effective batch sizes than either technique alone.

### Can I stream to CPU RAM instead of disk?

`RamSource` in [`layer_stream_runtime.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/layer_stream_runtime.py) provides exactly this. The tier selection logic automatically promotes frequently accessed layers to RAM when capacity permits, falling back to `DiskSource` for overflow.

### What models are compatible with Soup's layer streaming?

Any PyTorch transformer implementing standard `nn.Module` semantics. The test suite in [`tests/test_v07204.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_v07204.py) validates Llama, Mistral, and Qwen architectures. Custom architectures require no changes if they follow sequential layer composition.