How to Analyze Disk Usage with Mole: Fast Concurrent Scanning on macOS

Mole’s analyze command performs fast, concurrent disk‑usage scans on macOS, caching results for 24 hours and offering both an interactive TUI and JSON output for automation.

Mole is an open‑source macOS CLI tool that ships with a powerful disk‑usage analyzer. The mo analyze command recursively walks directory trees, aggregates file sizes using concurrent workers, and presents results through an interactive Bubble Tea interface or machine‑readable JSON. According to the tw93/Mole source code, the scanner uses bounded goroutines, dual‑channel result streaming, and Spotlight integration to deliver sub‑second overviews of multi‑terabyte volumes.

Getting Started with Mole Analyze

The basic syntax targets either a specific directory or the entire system overview.


# Analyze home directory with interactive TUI

mo analyze ~/

# System‑wide overview (Home, /Applications, /Library)

mo analyze

# Scan external volumes only

mo analyze /Volumes

# Force fresh scan ignoring cache

mo analyze -r ~/Projects

# JSON output for scripting

mo analyze --json ~/Documents > usage.json

The -r (refresh) flag bypasses the disk cache stored in cmd/analyze/cache.go, forcing scanPathConcurrent to rebuild the size model from scratch.

Architecture of the Disk Scanner

At the heart of Mole’s analyzer is a concurrent pipeline defined in cmd/analyze/scanner.go. The entry point scanPathConcurrent (lines 63‑95) orchestrates workers, channels, and heaps to stream results in real time.

The Scanner Engine (scanPathConcurrent)

When invoked, the scanner:

  1. Reads immediate children of the target path.
  2. Spawns workers bounded by a semaphore to prevent overwhelming the OS.
  3. Streams directory entries through entryChan and large files through largeFileChan.

Each worker processes either a subdirectory via calculateDirSizeConcurrent (lines 95‑120) or measures a single file’s actual block usage via getActualFileSize to correctly account for sparse files.

Concurrent Size Calculation

The scanner maintains three atomic counters—filesScanned, dirsScanned, and bytesScanned—that update lock‑free during the walk. Every 100 ms the UI layer reads these counters to render a progress spinner (tickMsg).

Results flow into two min‑heaps:

  • entryHeap – stores top‑N directories by size (bound by maxEntries in cmd/analyze/constants.go).
  • largeFileHeap – stores the biggest regular files (bound by maxLargeFiles).

Large File Discovery via Spotlight

While the primary walk executes, a background goroutine runs findLargeFilesWithSpotlight (lines 101‑124) using macOS mdfind. This supplements the heap with any large files that the traditional du‑based scan might have missed due to permission errors or hidden paths.

Caching Mechanism for Faster Re‑scans

Mole persists scan results to disk to avoid redundant I/O. The cache implementation in cmd/analyze/cache.go provides loadCacheFromDisk and storeOverviewSize functions that serialize the model state.

Cache entries expire after 24 hours of directory modification time, controlled by the cacheReuseWindow constant defined in cmd/analyze/constants.go (lines 20‑22). When a cached entry is valid, Mole skips the scanPathConcurrent walk entirely and loads the previous scanResult directly into the Bubble Tea model struct defined in cmd/analyze/main.go (lines 24‑33).

Output Modes: Interactive TUI vs. JSON

The CLI entry point in cmd/analyze/main.go branches into two execution paths based on flags:

Interactive Exploration

By default, runTUIMode launches a Bubble Tea interface. The model struct holds:

  • Current path and navigation stack
  • []dirEntry slice for directory entries
  • []fileEntry slice for large files
  • Progress counters and viewport state

Users navigate with arrow keys, enter directories with l or →, and return with h or ←. The view renders bars proportional to size using definitions from cmd/analyze/view.go.

Machine‑Readable JSON Output

When invoked with --json, the program executes runJSONMode (lines 58‑66), marshaling the scanResult struct directly to stdout. This bypasses the TUI entirely, making Mole ideal for CI pipelines or shell scripts that parse disk usage thresholds.

Programmatic Usage in Go

You can import Mole’s scanner into other Go applications. The ScanPath helper in cmd/analyze/scanner.go exposes the internal engine as a public API.

package main

import (
    "context"
    "fmt"
    
    "github.com/tw93/mole/cmd/analyze"
)

func main() {
    // Scan concurrently with automatic worker pooling
    result, err := analyze.ScanPath(context.Background(), "/Users/me/Downloads")
    if err != nil {
        panic(err)
    }
    
    fmt.Printf("Total: %d bytes\n", result.TotalSize)
    for _, e := range result.Entries {
        fmt.Printf("%s – %d B\n", e.Path, e.Size)
    }
}

ScanPath is a thin wrapper around scanPathConcurrent that manages context cancellation and channel cleanup automatically.

Summary

  • Concurrent scanning: scanPathConcurrent uses semaphore‑bounded workers and dual‑channel streaming via entryChan and largeFileChan to saturate disk I/O without overwhelming the OS.
  • Smart caching: Results persist for 24 hours via storeOverviewSize in cmd/analyze/cache.go, with TTL enforced by cacheReuseWindow.
  • Dual output: The model state drives either an interactive Bubble Tea TUI (runTUIMode) or JSON serialization (runJSONMode) based on the --json flag.
  • Spotlight integration: findLargeFilesWithSpotlight runs concurrently with the main walk to catch large files that standard recursion might miss.
  • Programmatic access: Go developers can call analyze.ScanPath to embed Mole’s disk scanner into larger automation tools.

Frequently Asked Questions

How does Mole handle sparse files and actual disk usage?

Mole uses getActualFileSize inside cmd/analyze/scanner.go to measure actual block usage rather than logical file size. This ensures that sparse files or files with extended attributes report the true bytes consumed on APFS or HFS+ volumes.

What is the performance impact of the concurrent scanner?

The scanner limits concurrent workers using a semaphore defined in cmd/analyze/constants.go, preventing file‑descriptor exhaustion. Atomic counters (filesScanned, dirsScanned, bytesScanned) update without locks, and results stream through buffered channels into min‑heaps, keeping memory usage bounded by maxEntries and maxLargeFiles regardless of directory depth.

Can I disable the cache when running Mole analyze?

Yes. Pass the -r flag (refresh) or the --no-cache equivalent to bypass loadCacheFromDisk. This forces scanPathConcurrent to perform a fresh walk and overwrite the stale cache entry via storeOverviewSize.

Is Mole’s disk analyzer available only as a CLI tool?

No. While the primary interface is the mo analyze command, the underlying ScanPath function in cmd/analyze/scanner.go is exported as a Go API. You can import github.com/tw93/mole/cmd/analyze and call it programmatically with a standard context.Context for cancellation and timeout control.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →