Performance Considerations for croc: Optimizing Peer-to-Peer File Transfers

croc achieves fast, secure file transfers through chunked streaming, optional compression, and configurable hashing, with performance tunable via CLI flags and buffer settings.

Understanding the performance considerations for croc helps you maximize throughput whether you are transferring terabytes across a LAN or sending files over high-latency internet connections. The schollz/croc repository implements several architectural optimizations that balance CPU usage, memory consumption, and network efficiency. By adjusting specific parameters and understanding the underlying mechanics in the source code, you can significantly improve transfer speeds for your specific environment.

Chunked Streaming Architecture

croc splits files into fixed-size chunks to enable reliable streaming and recovery. In src/models/constants.go, the TCP_BUFFER_SIZE is set to 64 KB by default, defining the chunk size used throughout the transfer pipeline.

Smaller chunks reduce latency on unreliable networks by allowing granular retransmission, though they increase packet overhead. The chunk handling logic resides in src/croc/croc.go (lines 183-204), where the sender reads files in these discrete blocks before transmission.

Compression and CPU Trade-offs

By default, croc uses compress.Compress defined in src/message/message.go (line 56) to reduce data size before transmission. While compression benefits high-latency or low-bandwidth links, it adds significant CPU overhead.

When CPU saturation becomes the bottleneck rather than network bandwidth, disable compression entirely:


# Disable compression for CPU-bound systems or fast LANs

croc --no-compress send large.iso

The compression implementation in src/compress/compress.go (lines 28-30) streams data to keep memory usage low, with performance benchmarks available in src/compress/compress_test.go.

Hash Algorithm Selection

The default imohash algorithm prioritizes speed over cryptographic strength, reducing CPU load during integrity verification. You can specify alternative hashing algorithms using the --hash flag parsed in src/cli/cli.go (line 381).

For CPU-constrained devices, selecting a faster hash (or staying with the default imohash) minimizes processing overhead:


# Explicitly select SHA-256 (slower but different collision profile)

croc --hash sha256 send dataset.csv

TCP Buffering and Throughput

The TCP layer in src/tcp/tcp.go (line 466) reads data in bursts sized according to models.TCP_BUFFER_SIZE. This burst size directly impacts throughput on high-speed networks, though the code caps bursts to the buffer size to prevent excessive memory consumption.

Larger buffers improve throughput on stable, high-bandwidth connections but increase memory usage per connection.

Resuming Interrupted Transfers

croc implements efficient resume capability through missing-chunk detection. The utils.MissingChunks function in src/utils/utils.go (lines 352-360) identifies which chunks the receiver still needs, while ChunkRangesToChunks optimizes the resume process.

This mechanism, utilized in the transfer loop at src/croc/croc.go (lines 2633-2640), ensures that interrupted transfers resume without retransmitting successfully received data, saving significant bandwidth and time on unstable connections.

Relay vs. Direct Connection Performance

When NAT traversal prevents direct peer-to-peer establishment, croc falls back to the relay server defined in src/croc/croc.go (lines 42-55). Relaying adds an extra network hop, increasing latency and potentially reducing throughput, though it enables connectivity in restricted network environments.

Hosting a geographically close relay server minimizes this performance penalty. Use the built-in relay binary or Docker deployment to position infrastructure near both peers.

Memory Efficiency in Large File Transfers

Unlike tools that load entire files into RAM, croc maintains constant memory usage regardless of file size. The decompress flow in src/compress/compress.go (lines 28-30) and the streaming architecture ensure only the current chunk (or small buffer) resides in memory at any moment.

This design allows transferring multi-terabyte files on systems with limited RAM without swapping or memory pressure.

Optimizing croc for Different Network Conditions

Tailor croc performance to your specific constraints using these validated configurations:

High-Latency, Low-Bandwidth Links

  • Keep compression enabled (default)
  • Use default imohash for faster integrity checks
  • Allow default chunk size for reliable delivery

CPU-Bound Systems (Old Hardware or Embedded Devices)

  • Use --no-compress to eliminate compression overhead
  • Consider --hash sha256 only if checksum speed becomes problematic (though imohash is usually fastest)

Local LAN, High-Speed Transfers

  • Disable compression to achieve raw network throughput
  • Default buffer sizes typically saturate gigabit links

Unstable or Intermittent Connections

  • Rely on built-in missing-chunk recovery (no flags required)
  • Use --quiet to reduce terminal I/O overhead during recovery

Batch File Operations

  • Run multiple croc send instances in parallel for many files
  • Monitor CPU and network saturation to identify bottlenecks

Custom Relay Deployment

  • Host croc relay near both endpoints to minimize RTT
  • Use Docker or pre-built binaries for quick deployment

Programmatic Configuration

When integrating croc into Go applications, configure performance options directly:

// Example: Disabling compression programmatically
opts := croc.Options{
    NoCompress: true, // Bypasses compress.Compress
}
c, _ := croc.New(opts)
c.Send("file.txt")

Summary

  • Chunked streaming (64 KB default in src/models/constants.go) balances latency and overhead for most networks.
  • Compression helps on slow links but costs CPU; disable with --no-compress when processor-bound.
  • Hash selection affects CPU usage; default imohash in src/cli/cli.go prioritizes speed.
  • Resume capability via MissingChunks in src/utils/utils.go prevents redundant data retransmission.
  • Relay fallback adds latency but enables NAT traversal; proximity hosting mitigates impact.
  • Memory efficiency remains constant through streaming, handling files of any size in minimal RAM.

Frequently Asked Questions

How does croc handle interrupted transfers without starting over?

croc tracks which chunks have been successfully received using the MissingChunks function in src/utils/utils.go. When a connection drops and reconnects using the same code phrase, the receiver communicates which chunks are still needed, and the sender resumes from that point. This process happens automatically without requiring special flags.

Will disabling compression always improve transfer speed?

Disabling compression with --no-compress improves speed only when CPU is the bottleneck, such as on older hardware or during LAN transfers where network bandwidth exceeds compression throughput. On high-latency or low-bandwidth internet connections, compression typically improves effective throughput by reducing total bytes transmitted, even with the CPU overhead.

What is the optimal TCP buffer size for croc?

The default 64 KB TCP_BUFFER_SIZE defined in src/models/constants.go works well for most scenarios. While larger buffers can improve throughput on very high-speed, stable networks, they increase memory usage per connection. The current implementation in src/tcp/tcp.go caps burst sizes to this buffer limit, making it generally unnecessary to modify unless you are compiling from source for specific high-performance environments.

When should I use a custom relay server instead of the default public relay?

Use a custom relay when both peers are geographically close to each other but far from the default public infrastructure, or when transferring sensitive data that should not traverse third-party servers. Deploy croc relay on a VPS near both endpoints to minimize added latency while maintaining the ability to penetrate NATs when direct P2P connection fails.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →