How to Analyze Disk Usage with Mole: Fast Concurrent Scanning on macOS
Mole’s analyze command performs fast, concurrent disk‑usage scans on macOS, caching results for 24 hours and offering both an interactive TUI and JSON output for automation.
Mole is an open‑source macOS CLI tool that ships with a powerful disk‑usage analyzer. The mo analyze command recursively walks directory trees, aggregates file sizes using concurrent workers, and presents results through an interactive Bubble Tea interface or machine‑readable JSON. According to the tw93/Mole source code, the scanner uses bounded goroutines, dual‑channel result streaming, and Spotlight integration to deliver sub‑second overviews of multi‑terabyte volumes.
Getting Started with Mole Analyze
The basic syntax targets either a specific directory or the entire system overview.
# Analyze home directory with interactive TUI
mo analyze ~/
# System‑wide overview (Home, /Applications, /Library)
mo analyze
# Scan external volumes only
mo analyze /Volumes
# Force fresh scan ignoring cache
mo analyze -r ~/Projects
# JSON output for scripting
mo analyze --json ~/Documents > usage.json
The -r (refresh) flag bypasses the disk cache stored in cmd/analyze/cache.go, forcing scanPathConcurrent to rebuild the size model from scratch.
Architecture of the Disk Scanner
At the heart of Mole’s analyzer is a concurrent pipeline defined in cmd/analyze/scanner.go. The entry point scanPathConcurrent (lines 63‑95) orchestrates workers, channels, and heaps to stream results in real time.
The Scanner Engine (scanPathConcurrent)
When invoked, the scanner:
- Reads immediate children of the target path.
- Spawns workers bounded by a semaphore to prevent overwhelming the OS.
- Streams directory entries through
entryChanand large files throughlargeFileChan.
Each worker processes either a subdirectory via calculateDirSizeConcurrent (lines 95‑120) or measures a single file’s actual block usage via getActualFileSize to correctly account for sparse files.
Concurrent Size Calculation
The scanner maintains three atomic counters—filesScanned, dirsScanned, and bytesScanned—that update lock‑free during the walk. Every 100 ms the UI layer reads these counters to render a progress spinner (tickMsg).
Results flow into two min‑heaps:
entryHeap– stores top‑N directories by size (bound bymaxEntriesincmd/analyze/constants.go).largeFileHeap– stores the biggest regular files (bound bymaxLargeFiles).
Large File Discovery via Spotlight
While the primary walk executes, a background goroutine runs findLargeFilesWithSpotlight (lines 101‑124) using macOS mdfind. This supplements the heap with any large files that the traditional du‑based scan might have missed due to permission errors or hidden paths.
Caching Mechanism for Faster Re‑scans
Mole persists scan results to disk to avoid redundant I/O. The cache implementation in cmd/analyze/cache.go provides loadCacheFromDisk and storeOverviewSize functions that serialize the model state.
Cache entries expire after 24 hours of directory modification time, controlled by the cacheReuseWindow constant defined in cmd/analyze/constants.go (lines 20‑22). When a cached entry is valid, Mole skips the scanPathConcurrent walk entirely and loads the previous scanResult directly into the Bubble Tea model struct defined in cmd/analyze/main.go (lines 24‑33).
Output Modes: Interactive TUI vs. JSON
The CLI entry point in cmd/analyze/main.go branches into two execution paths based on flags:
Interactive Exploration
By default, runTUIMode launches a Bubble Tea interface. The model struct holds:
- Current path and navigation stack
[]dirEntryslice for directory entries[]fileEntryslice for large files- Progress counters and viewport state
Users navigate with arrow keys, enter directories with l or →, and return with h or ←. The view renders bars proportional to size using definitions from cmd/analyze/view.go.
Machine‑Readable JSON Output
When invoked with --json, the program executes runJSONMode (lines 58‑66), marshaling the scanResult struct directly to stdout. This bypasses the TUI entirely, making Mole ideal for CI pipelines or shell scripts that parse disk usage thresholds.
Programmatic Usage in Go
You can import Mole’s scanner into other Go applications. The ScanPath helper in cmd/analyze/scanner.go exposes the internal engine as a public API.
package main
import (
"context"
"fmt"
"github.com/tw93/mole/cmd/analyze"
)
func main() {
// Scan concurrently with automatic worker pooling
result, err := analyze.ScanPath(context.Background(), "/Users/me/Downloads")
if err != nil {
panic(err)
}
fmt.Printf("Total: %d bytes\n", result.TotalSize)
for _, e := range result.Entries {
fmt.Printf("%s – %d B\n", e.Path, e.Size)
}
}
ScanPath is a thin wrapper around scanPathConcurrent that manages context cancellation and channel cleanup automatically.
Summary
- Concurrent scanning:
scanPathConcurrentuses semaphore‑bounded workers and dual‑channel streaming viaentryChanandlargeFileChanto saturate disk I/O without overwhelming the OS. - Smart caching: Results persist for 24 hours via
storeOverviewSizeincmd/analyze/cache.go, with TTL enforced bycacheReuseWindow. - Dual output: The
modelstate drives either an interactive Bubble Tea TUI (runTUIMode) or JSON serialization (runJSONMode) based on the--jsonflag. - Spotlight integration:
findLargeFilesWithSpotlightruns concurrently with the main walk to catch large files that standard recursion might miss. - Programmatic access: Go developers can call
analyze.ScanPathto embed Mole’s disk scanner into larger automation tools.
Frequently Asked Questions
How does Mole handle sparse files and actual disk usage?
Mole uses getActualFileSize inside cmd/analyze/scanner.go to measure actual block usage rather than logical file size. This ensures that sparse files or files with extended attributes report the true bytes consumed on APFS or HFS+ volumes.
What is the performance impact of the concurrent scanner?
The scanner limits concurrent workers using a semaphore defined in cmd/analyze/constants.go, preventing file‑descriptor exhaustion. Atomic counters (filesScanned, dirsScanned, bytesScanned) update without locks, and results stream through buffered channels into min‑heaps, keeping memory usage bounded by maxEntries and maxLargeFiles regardless of directory depth.
Can I disable the cache when running Mole analyze?
Yes. Pass the -r flag (refresh) or the --no-cache equivalent to bypass loadCacheFromDisk. This forces scanPathConcurrent to perform a fresh walk and overwrite the stale cache entry via storeOverviewSize.
Is Mole’s disk analyzer available only as a CLI tool?
No. While the primary interface is the mo analyze command, the underlying ScanPath function in cmd/analyze/scanner.go is exported as a Go API. You can import github.com/tw93/mole/cmd/analyze and call it programmatically with a standard context.Context for cancellation and timeout control.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →