# Audio Slicing Pipeline in GPT-SoVITS: How threshold, min_length, and min_interval Control Output

> Understand the GPT-SoVITS audio slicing pipeline. Learn how threshold, min_length, and min_interval parameters control audio output for better voice conversion.

- Repository: [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS)
- Tags: internals
- Published: 2026-03-07

---

**The GPT-SoVITS audio slicing pipeline uses RMS-based silence detection to split audio only when silent intervals exceed `min_interval` and resulting segments meet `min_length`, while `threshold` defines what counts as "silent."**

The **audio slicing pipeline** in the [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS) repository is a critical preprocessing tool for voice cloning workflows. Located in [`tools/slicer2.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/slicer2.py), this pipeline converts long recordings into training-ready chunks by detecting silence and cutting at optimal points. Understanding how the `threshold`, `min_length`, and `min_interval` parameters interact allows you to fine-tune segmentation for speech, music, or noisy environments.

## How the Audio Slicing Pipeline Works

The pipeline processes audio through a series of deterministic steps implemented in the `Slicer` class. Each stage transforms the raw waveform into discrete segments based on energy levels.

### Step 1: Audio Loading and RMS Calculation

The pipeline begins by loading audio via `librosa.load` in `main()` (lines 55-66 of [`tools/slicer2.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/slicer2.py)). The `get_rms` function (lines 4-35) then computes **Root Mean Square (RMS)** energy using a sliding window defined by `win_size` and `hop_size`. This produces a frame-wise energy profile where each value represents the loudness of a short audio segment.

### Step 2: Silence Detection and Frame Marking

Inside `Slicer.slice` (starting at line 78), the pipeline converts the `threshold` parameter from decibels to linear amplitude using `self.threshold = 10 ** (threshold/20)` (line 53). Any frame with RMS below this value is marked as **silent**. This binary mask determines candidate locations for audio cuts.

### Step 3: Cut Decision Logic

The core slicing algorithm (lines 90-96) applies two constraints before creating a cut:

1.  **Minimum Silence Duration (`min_interval`)**: The silent stretch must exceed `min_interval` (converted to frames at line 57). Brief pauses shorter than this are ignored, preventing fragmentation during natural speech breathing.
2.  **Minimum Segment Length (`min_length`)**: The audio accumulated since the last cut must exceed `min_length` (converted to frames at line 56). If the segment is too short, the slicer skips the cut and merges the content with the next segment.

### Step 4: Intelligent Silence Trimming

When a valid cut is found, the pipeline selects the **lowest-RMS point** within the silent region as the exact split position (lines 95-122). It preserves at most `max_sil_kept` milliseconds of surrounding silence to avoid harsh cuts, then returns chunks as tuples of `[audio_segment, start_sample, end_sample]`.

## The Role of threshold, min_length, and min_interval

These three parameters form a control system that balances granularity against coherence. Adjusting them requires understanding their unit conversions and interaction effects.

### threshold: Defining the Silence Floor

The `threshold` parameter (CLI: `--db_thresh`, default `-40` dB) sets the **RMS energy level** below which audio is considered silent.

-   **Higher values** (e.g., `-30` dB): Treat louder audio as silence, creating more aggressive cuts. Useful for noisy recordings where background hum exists.
-   **Lower values** (e.g., `-50` dB): Require true silence before cutting, preserving quiet speech or breath sounds. Better for studio-quality voice recordings.

In [`tools/slicer2.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/slicer2.py) line 53, the conversion occurs: `self.threshold = 10 ** (threshold / 20)`.

### min_length: Enforcing Minimum Clip Duration

The `min_length` parameter (CLI: `--min_length`, default `5000` ms) guarantees that **no output segment is shorter than this duration**.

-   If the audio between two silence regions is shorter than `min_length`, the slicer **merges** the segments by skipping the intermediate cut.
-   This prevents the creation of tiny, unusable clips (e.g., single words or click sounds) that would degrade TTS training data quality.

The conversion to frames happens at line 56: `self.min_length = round(sr * min_length / 1000 / self.hop_size)`.

### min_interval: Filtering Brief Pauses

The `min_interval` parameter (CLI: `--min_interval`, default `300` ms) specifies the **minimum duration of silence required to trigger a cut**.

-   Natural speech contains brief pauses between phrases (often 100-200 ms). Setting `min_interval` to `300` ms prevents the slicer from cutting at every micro-pause.
-   For music or dialogue with distinct gaps, increasing `min_interval` to `1000` ms ensures cuts only occur between tracks or scenes.

The frame conversion is at line 57: `self.min_interval = round(min_interval / self.hop_size)`.

## Practical Usage Examples

### Command-Line Interface

The [`tools/slicer2.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/slicer2.py) script operates as a standalone CLI tool. Here is an optimized configuration for voice cloning datasets:

```bash
python -m tools.slicer2 \
    input.wav \
    --out ./sliced_audio \
    --db_thresh -35 \
    --min_length 3000 \
    --min_interval 400 \
    --hop_size 10 \
    --max_sil_kept 200

```

This configuration:
-   Uses a **higher threshold** (`-35` dB) to handle room noise
-   Sets **shorter minimum length** (`3000` ms) for more granular training samples
-   Requires **longer silence** (`400` ms) to avoid cutting mid-sentence

### Python API Integration

For programmatic control within a data processing pipeline:

```python
import librosa
from tools.slicer2 import Slicer

# Load stereo or mono audio

waveform, sr = librosa.load("interview.wav", sr=None, mono=False)

# Initialize slicer with millisecond parameters

slicer = Slicer(
    sr=sr,
    threshold=-40,      # dB threshold

    min_length=5000,  # 5 seconds minimum

    min_interval=300,   # 300ms silence required

    hop_size=10,        # 10ms frame hop

    max_sil_kept=500    # keep 0.5s silence padding

)

# Execute slicing

chunks = slicer.slice(waveform)

# Process results: each chunk is [audio, start_sample, end_sample]

for idx, (segment, start, end) in enumerate(chunks):
    duration = (end - start) / sr
    print(f"Chunk {idx}: {duration:.2f}s (samples {start}-{end})")
    # librosa.output.write_wav(f"chunk_{idx}.wav", segment, sr)

```

This approach gives you direct access to sample indices for timestamp metadata while leveraging the same RMS-based detection logic used by the web interface.

## Key Source Files to Explore

Understanding the implementation requires examining these specific files in the [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS) repository:

| File | Purpose | Key Components |
|------|---------|----------------|
| [`tools/slicer2.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/slicer2.py) | Core slicing implementation | `Slicer` class, `get_rms` function, `slice` method (lines 4-152) |
| [`tools/slice_audio.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/tools/slice_audio.py) | Web UI wrapper | High-level interface that instantiates `Slicer` with Gradio parameters |
| [`webui.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/webui.py) | Main application entry | Integration point showing default parameter values for the GUI |

The `Slicer.__init__` method (lines 38-58) contains the unit conversion logic that transforms millisecond inputs into frame counts based on the hop size, while the `slice` method (lines 60-152) implements the state machine that tracks silence duration and enforces minimum length constraints.

## Summary

The GPT-SoVITS audio slicing pipeline uses **RMS energy detection** and a three-parameter control system to segment long recordings into training-ready chunks:

-   **`threshold`** sets the silence detection floor in decibels; higher values create more aggressive cuts by treating louder audio as silence.
-   **`min_length`** enforces a minimum duration for every output segment, merging short clips to prevent fragmentation.
-   **`min_interval`** filters out brief pauses by requiring a minimum duration of continuous silence before allowing a cut.

Together, these parameters prevent pathological splits while preserving natural boundaries, making the pipeline suitable for voice cloning datasets, podcast editing, and music segmentation.

## Frequently Asked Questions

### What is the difference between min_length and min_interval in the GPT-SoVITS slicer?

**`min_length`** controls the duration of the **output audio segments**, ensuring each clip is at least that many milliseconds long by merging short segments with their neighbors. **`min_interval`** controls the duration of the **silence itself**, requiring a gap of at least that length before the slicer will insert a cut. You need both because `min_interval` determines *where* cuts can happen, while `min_length` determines *whether* those cuts are actually executed.

### How do I choose the right threshold value for noisy audio?

For noisy recordings or audio with background hum, increase the **`threshold`** to a less negative value such as `-35` dB or `-30` dB (default is `-40` dB). This higher threshold treats louder background noise as "silence," preventing the slicer from creating cuts within continuous noise. However, be careful not to set it too high, or you risk cutting into quiet speech or breath sounds that fall below the elevated threshold.

### Why does the slicer sometimes merge segments even when there is silence between them?

This occurs when the audio between two silence regions is shorter than the **`min_length`** parameter. The slicer's state machine (implemented in `Slicer.slice` around lines 90-96) checks if the accumulated audio since the last cut meets the minimum length requirement before finalizing a new cut. If the segment is too short, the slicer ignores the intervening silence and continues accumulating audio until the minimum length is satisfied, effectively merging what would have been separate clips into a single longer segment.

### Can I use the slicer for music or only for speech?

The slicer works for any audio type, but you must adjust the parameters accordingly. For **music**, which often has gradual decays and shorter natural gaps, you might reduce **`min_interval`** to `100-200` ms to catch brief pauses between notes, and lower **`min_length`** to `2000-3000` ms to isolate musical phrases. For **speech**, the defaults (`300` ms interval, `5000` ms length) work well because human conversation contains longer pauses between sentences and requires context-rich segments for TTS training.