How Files Are Transferred in Chunks Using Croc: Architecture and Workflow

croc transfers files by splitting them into fixed-size chunks (default 64 KB), detecting missing segments on the receiver, and transmitting only the specific chunks requested, enabling efficient resumable transfers over unreliable networks.

When transferring large files over intermittent connections, breaking data into manageable pieces prevents complete retransmission on failure. In the schollz/croc repository, files are transferred in chunks using a sophisticated resume protocol that minimizes redundant data transmission. This article examines the exact mechanism behind croc's chunked file transfer, referencing the actual source code implementation to explain how partial files are detected, requested, and assembled.

Step 1: Discovering Missing Chunks with utils.MissingChunks

The process begins on the receiver side. When a transfer starts, croc creates a temporary file sized to match the expected total file size. The function MissingChunks in src/utils/utils.go (lines 352-410) scans this file to determine which data is already present.

The implementation compares each block of the existing file against a zero-filled buffer. Blocks containing non-zero data are considered present; blocks that are entirely zero (or outside the current file size) are marked as missing. If the file does not exist or its size differs from the expected total, the function reports that all chunks are missing, initiating a fresh transfer.

Step 2: Converting Ranges to Absolute Offsets with ChunkRangesToChunks

Once MissingChunks identifies the gaps, it returns a compact range encoding in the format [chunkSize, start, count, start, count …]. The utility function ChunkRangesToChunks in src/utils/utils.go (lines 426-435) expands this compressed representation into a flat slice of absolute chunk offsets.

This transformation is crucial because the sender needs a simple list of indices to determine which blocks to read from disk and transmit. The flat slice eliminates range arithmetic during the actual file streaming, reducing CPU overhead on the sender side.

Step 3: Sender-Side Chunk Selection and Mapping

On the sender side, the croc.Croc struct manages the transfer state in src/croc/croc.go (lines 2275-2290). When the sender receives the resume request containing the missing chunk ranges, it performs two critical operations:

  1. Expands ranges: Calls ChunkRangesToChunks to populate c.CurrentFileChunks with the flat list of required indices.
  2. Builds lookup map: Creates c.chunkMap (lines 2278-2280) as a map[uint64]struct{} for O(1) existence checks.

During the streaming loop, the sender reads the source file sequentially but checks each chunk index against c.chunkMap. If the index exists, the chunk is read, optionally compressed and encrypted, and transmitted. After successful transmission, the index is deleted from the map. When the map is empty, the transfer is complete, allowing the sender to stop early even if the entire file hasn't been streamed.

Step 4: Receiver-Side Data Assembly with comm.ChunkedReader

The receiver handles incoming data through comm.ChunkedReader in src/comm/comm.go (lines 27-30). This component reads from the underlying TCP connection into a buffer sized to models.TCP_BUFFER_SIZE, which may contain partial chunks or multiple chunks depending on network conditions.

Because the receiver already knows exactly which chunks it requested (from Step 1), it can write each received byte range to the correct offset in the temporary file. The implementation ignores any data belonging to chunks already marked as present, providing resilience against duplicate or out-of-order packets.

The Complete Transfer Workflow

Putting these components together, here is how files are transferred in chunks using croc in practice:

  1. Handshake: The sender announces the file size and a unique fingerprint (hash) to establish the transfer parameters.
  2. Resume request: The receiver opens or creates a file of the announced size and executes MissingChunks. The resulting range list is transmitted back to the sender.
  3. Chunk selection: The sender expands the ranges with ChunkRangesToChunks and populates c.chunkMap for fast lookup.
  4. Streaming: The sender reads the source file sequentially, transmitting only chunks whose indexes exist in c.chunkMap. Each chunk is written to the exact byte offset on the receiver side.
  5. Completion: When c.chunkMap becomes empty, the sender knows the receiver possesses every required chunk and closes the stream.

Practical Configuration and Usage

The default chunk size is 64 KB, configurable via the --chunksize flag. Smaller chunks improve resume granularity but increase protocol overhead; larger chunks reduce overhead but may retransmit more data on interruption.


# Send a file using default 64 KB chunks

croc send large-movie.mkv

# Receive with automatic resume capability

croc receive

# If the file exists partially, croc automatically requests only missing chunks
// Minimal implementation showing chunk discovery on the receiver
func requestResume(fname string, totalSize int64, chunkSize int) []int64 {
    // Returns flat slice of chunk offsets the sender must transmit
    ranges := utils.MissingChunks(fname, totalSize, chunkSize)
    return utils.ChunkRangesToChunks(ranges)
}
// Sender preparation: building the fast-lookup map
func (c *Croc) prepareChunkMap() {
    c.CurrentFileChunks = utils.ChunkRangesToChunks(c.CurrentFileChunkRanges)
    c.chunkMap = make(map[uint64]struct{})
    for _, ch := range c.CurrentFileChunks {
        c.chunkMap[uint64(ch)] = struct{}{}
    }
}
// Inside the sending loop: selective transmission
if _, ok := c.chunkMap[uint64(chunkIdx)]; ok {
    // Read, compress, encrypt, and send chunk
    sendChunk(chunkBuffer)
    delete(c.chunkMap, uint64(chunkIdx)) // Mark as transmitted
}

Summary

  • Chunk discovery occurs via MissingChunks in src/utils/utils.go, which scans existing files to identify zero-filled (missing) blocks.
  • Range expansion converts compact [start, count] pairs into absolute offsets using ChunkRangesToChunks (lines 426-435).
  • Sender optimization uses a chunkMap (lines 2278-2280 in src/croc/croc.go) for O(1) lookup, enabling the sender to skip unwanted chunks efficiently.
  • Receiver assembly writes incoming data directly to file offsets via comm.ChunkedReader in src/comm/comm.go, supporting out-of-order delivery.
  • Resumability is inherent: any interruption simply restarts the handshake, and only missing chunks are retransmitted.

Frequently Asked Questions

What is the default chunk size in croc?

The default chunk size is 64 KB (65,536 bytes). You can configure this using the --chunksize flag when sending files. Larger chunks reduce metadata overhead but decrease granularity for resume operations, while smaller chunks improve resume efficiency at the cost of increased protocol overhead.

How does croc resume an interrupted transfer?

When a transfer resumes, the receiver executes MissingChunks to scan the existing partial file. This function compares each block against a zero-filled buffer to identify which chunks are already present. Only the missing chunks are requested from the sender, who then uses chunkMap to transmit exactly those specific blocks, appending them to the existing file rather than starting over.

What prevents the sender from transmitting chunks the receiver already has?

The sender maintains a chunkMap (implemented as a Go map with empty struct values for memory efficiency) that serves as a whitelist. Before reading each chunk from disk, the sender checks if the chunk index exists in this map. If the index is not present—meaning the receiver did not request it—the chunk is skipped entirely, ensuring no redundant data traverses the network.

How does the receiver handle out-of-order chunk delivery?

The receiver uses comm.ChunkedReader to read data from the TCP stream into a buffer, then writes each byte range to the specific file offset corresponding to that chunk's index. Because chunks are written directly to their final positions rather than sequentially, the order of arrival does not matter. The receiver ignores any data for chunks not listed in the original request, providing resilience against duplicate transmissions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →